Field notes
Notes from the build.
Working notes on the AI systems we design, build, and run — what we tried, what held up, and what we changed afterwards.
- August 30, 2026
Daily Wire — Aug 30
Eleven merges landed across two repositories in a day — and nearly half of them were about deciding who gets credited when a person and a model both touch the same piece of writing.
- August 29, 2026
What Fine-Tuning Actually Cost: A Full Audit
Prepaid cash is $214.90 after tax per Claude Max and per ChatGPT Pro. Fine-tuning used 25.4% of one weekly Max bar and 64.2% of one weekly Pro bar ($2,754 API-equivalent on the $8,000 / $14,000 square). Counted training tokens still price at about $75 list. The gap is untracked session work, not a second cash bill.
- August 29, 2026
The Borg Memory System: One Brain for Every Agent
How a single-operator estate wired Claude Code, Codex, and Grok into one shared, self-consolidating memory — a recall layer, a judgment layer, and one rule that keeps an aggressive automation posture safe: memory is never the authority. Then it taught tiny local models to run the expensive part.
- August 29, 2026
Conductors: Turning ChatGPT Seats into a Steerable Worker Fleet
A worker you cannot talk to until it exits is not a fleet. A ~200-line HTTP control plane wraps each account seat so lanes can be started, watched, steered mid-flight, interrupted, and resumed — plus the seven gotchas that cost real time, so they cost you none.
- August 28, 2026
The exam it passed, the gate it failed
Our small model scored 400 out of 400 on the held-out exam, then finished 1 of 25 real notes in the pipeline it was meant to run. What the exam could not see, why the promotion gate caught it, and the recorder that is now building the training set for the fix.
- August 28, 2026
Thursday they published the week, then emptied the drawer
The daily beat, upgraded: the 11-day merge scoreboard charted, Friday morning's three line incidents hour by hour, and three rules any agent operator can steal.
- August 28, 2026
The 47-hour day: training a small model to replace a big one
Three always-on loops kept one 27-billion-parameter model busy for 47.02 hours in a single day. What we measured, the small model we are training to take the job, and the four things that have to be true before it ships.
- August 27, 2026
Five days, five stops, 147 merges: our week with a model that had no name
Five days of running a fleet of coding agents against a preview model, told as a route with five stops. What broke, where our own instruments lied to us, and why one merged pull request was the only number worth counting.
- August 27, 2026
Your Cursor bill has four tanks. They do not mix.
Live Cursor Ultra bill (Aug 18–Sep 18 cycle) shows four separate usage tanks that do not mix: Cursor Models (monthly), Other Models (monthly), Grok Bot (weekly), and Grok CLI (weekly). Display rows are not quota. Estimates and hypotheses are labeled.
- August 25, 2026
Local AI for a service business: which open-weight models actually fit
DRAM shortage made the luxury box a luxury. Here is the Mac Studio unified-memory ladder, the open-weight models that fit, and why Utlyze builds custom AI systems on a box that can hold the work.