How I Run Claude Code
The first maintenance pass found a sync job failing 84% of its runs
Built with
Not public
The setup sits inside my notes vault and names client folders, and Life OS holds my own health and financial data, so neither can be published. The meeting-capture code is public, though it's behind the version I run.
I use Claude Code for nearly everything, from coursework to client work. Used as one long chat it kept drifting, and I retyped the same setup every Monday. So I wrote down what to do instead, in a folder every session reads, and built the rest on top of it.
Where a session opens
Where I open a session matters more than anything else here. Opening at the root of my notes loads every project's context at once, and the work gets worse. Opening in the project folder loads one context file and that folder's skills. Most of the decision tree comes down to picking the narrower option.
Skills, agents and the verifier
There are 29 skills I can call from anywhere, plus more scoped to single folders. One of my two read-only subagents is a verifier. It grades finished work against a fixed rubric and never did any of the work itself, and it has failed my own write-ups twice, correctly, both times for claims about files I hadn't re-checked.
Two workflows written in JavaScript drive the subagents. `iterate-until-good` sends a draft to three critics, any of which can block it, and loops until the lowest score clears a bar. Then a judge that didn't write or critique it decides. `deep-work` does the research first and hands a brief to that loop.
Scheduled checks
Two checks run through launchd with no model in them. Sunday's covers system health and deadlines, and Friday's covers client admin. I added them after the first full maintenance pass found my personal-data sync had logged 1,113 failed runs against 209 clean, and nothing had flagged it except a log file I never read. Most of that was one job macOS wouldn't run from my iCloud folder, which is fixed now.
A writing voice with its own judge
When Claude Code writes as me, it pulls two to five samples of my own writing, matched on register before topic. There are 34 samples so far, across five registers. A separate judge agent that can read but not edit scores each draft on six things, and one fail fails it. When it's torn, it picks the shorter, rougher option.
I checked the judge by ranking pairs of drafts blind. It matched me on 6 of 8 blind pairs (75%), under its own 80% bar, so I still read everything. A change to the voice rules only stays if it doesn't make a regression set of 28 cases worse. That set's one full run hit a usage limit partway through, so the newer rules rest on my own feedback.
Meeting capture
I built this so calls could be written up without a notetaker service in the audio path. It records the far end and my microphone on separate channels, and transcription (whisper.cpp) and speaker labels (pyannote) both run on the laptop.
Most of the code is defensive. If the system output isn't pinned to the right device, the far-end channel records digital silence and nothing complains at the time. Whisper then fills the silence with stock phrases like "Thank you". So the recorder plays a test tone and picks the device itself before it starts. During the call, a watchdog re-pins the output every ten seconds and sends a notification if the far end goes quiet while I'm still talking.
The two channels are also how I debug it. On one call the backup mode, which plays the call out loud, let the microphone pick up everyone in the room, and the script labelled all of them as me. That gave 1,056 turns attributed to one speaker on a 53-minute call. The far-end channel was about 25 times quieter than the microphone, which is how I spotted it. It's fixed now, and the fix is in the public repository.
Life OS
Life OS pulls my own data, from Hevy and Apple Health to Toggl and my bank statements, into one SQLite database on a schedule. Then it writes it back into my notes as markdown. Every source table has a unique key, so running a sync twice never adds a row twice.
Transactions go through 73 substring rules first, and a model only sees what they miss. Anything it isn't sure about waits for me. I paused the scheduler on 18 September to rebuild it around a new routine, and some sources had been failing for weeks before that.
What this does not show
- No measurement of the whole. I haven't compared working with it against working without it, so 'the work is better' is my judgement.
- The judge in the workflows is always a Claude model, because the workflow can only call Claude models, so it shares the drafter's blind spots.
- iterate-until-good was checked end to end once, in June. I have no record of deep-work finishing a run.
- The scheduled checks count failures in the logs, so they can't see a job that has stopped altogether, and the Sunday one has produced a report on one of the three Sundays since I installed it.
- The voice judge is below its own 80% bar, and it was calibrated on four pairs a round against the fifteen to twenty its rubric asks for.
- Meeting capture is macOS only, and its 29 tests cover the two pure-logic modules, not the audio, transcription or diarisation code.
- Life OS is paused, built for one user on one machine, and has never been pushed anywhere.