Usage audit · 30 days to Sep 14 · 8,900 transcripts, verified twice
We don't run out of capacity because sessions are too big. We run out because subagents wait, then pay full price to reread everything they already knew.
Roughly half of everything we spent in the last 30 days is three things: subagents rewriting their cache after a short wait, spin/poll loops, and a handful of runaway subagents. That share holds (39–54%) no matter how you weigh the tokens.
All three got a fix today. Two more fixes are ready and waiting on your go-ahead.
A subscription doesn't show a dollar bill, so we had to reverse-engineer what the 5-hour/weekly meter actually charges for. We compared 26,686 meter readings against the tokens each session was using at that exact minute.
Answer: the meter is not counting raw tokens. A reused chunk of context ("cache read") is nearly free. Writing a new chunk of context (a "cache write") costs full price. That's the whole story.
| Token type | Weight on the limit | Note |
|---|---|---|
| Cache writes (new context) | Full weight | Counts like typing it fresh |
| Fresh input | Full weight | |
| Output | Full weight | |
| Cache reads (context we already had) | 0.0063 per token | 95% CI 0.0031–0.0132 — cheaper even than the API's own 0.1 discount |
Practical rate: about 0.5% of a 5-hour session per million cache-discounted tokens. One sentence version: a big context that stays cached is nearly free; re-writing it after the cache expires is what costs.
The input/output multipliers (82x and 13x) came back with very wide error bars — real, but not precise enough to use as a price list. Only the cache-read discount is a sharp, repeatable number.
There is one weekly pool per account, not two. Anthropic's support article "Claude Fable models on your plan" says Fable may use up to 50% of the weekly limit and that "your use of other models draws from the same usage limits." The "weekly Fable" meter is that 50% sub-cap. A half-empty Fable meter next to a full "weekly all" meter is not spare capacity.
Our own meter history agrees (per million cache-discounted tokens, non-negative least squares over the same windows as above; details in tools/token-audit/meter_by_model.md):
| Model | 5-hour meter, relative to Opus | Weekly meter, relative to Opus |
|---|---|---|
| Fable | 2.4x (95% CI 1.65–3.26) | 3.1x (95% CI 2.02–4.47) |
| Opus | 1x | 1x |
| Sonnet | 0.76x | 0.79x |
Consequence: moving subagent work onto Fable would empty the shared pool two to four times faster. Fable stays where it pays, judgment and orchestration.
Sonnet and Opus are, for practical purposes, 100% subagent work. Fable is, for practical purposes, 100% Fede's main sessions. The rule that keeps orchestration on Fable and hands worker jobs to subagents is being followed. The problem sits inside how those Sonnet/Opus workers behave, not in who's driving.
| Model | Share of weighted spend | Runs mostly as |
|---|---|---|
| Sonnet 5 | 48.2% | Subagent worker (99.5% of its spend) |
| Opus 5 | 36.7% | Subagent worker (99.6% of its spend) |
| Fable 5 + 5.1 | 14.1% | Fede's main sessions (95%+ of its spend) |
| Opus 4.8 | 0.8% | Mostly main sessions |
| Haiku 4.5 | 0.2% | Barely used at all |
4 in 10 subagent launches (42.1%) don't pass a model at all and fall back to whatever the default is — that's still a gap in the "route by task" rule.
| Rank | What | Size | Evidence |
|---|---|---|---|
| #1 | Idle subagents re-writing their whole context. Subagents keep their cache for 5 minutes; main sessions get an hour. Any pause past 5 minutes — a long wait, watching CI — forces a full, pricier re-write. | 2.87B weighted tokens (16.2% of everything) | A subagent idle 5–15 minutes rewrites its cache 86% of the time, vs. 7.8% for a main session in the same gap. |
| #2 | Spin/poll turns — the same command run 100+ times in a row inside one subagent. | 2.64B weighted tokens (14.9%) | 64,623 such turns across 30 days. The Sep 12 guard has refused 358 of them since; almost all the damage predates it. |
| #3 | Ten runaway Opus subagents, all spawned from Fede's main sessions, that stayed alive near the 1M-token ceiling. | 1.52B weighted tokens (8.6%) | Six of the ten hit over 850k context; five hit over 966k. The worst ran 4,057 turns over 20 hours. |
Fede pushed back on the ~40% headline, so we re-ran it under the meter's real cache-read discount instead of the API price list. The share went up, not down — the API price list was actually the conservative estimate.
| Weighting | #1 idle rewrites | #2 spin/poll | #3 top-10 subagents | Union (no double-count) |
|---|---|---|---|---|
| Cache-writes only | 53.7% | 1.6% | 3.1% | 54.4% |
| Meter-realistic (best estimate) | 42.4% | 5.0% | 4.5% | 47.6% |
| Pessimistic bound | 34.3% | 4.6% | 3.8% | 39.1% |
The defensible number: these three things are 39–54% of 30-day spend, best estimate ~48%. The ranking that matters more than the total: under a realistic weighting, idle cache rewrites alone dwarf the other two (42.4% vs. 5.0% and 4.5%). Spinning looked expensive under a raw-token view because it generates huge cache-read volume — which the meter barely charges for. The cost was never the polling. It was the cache rewrites the polling caused.
Short answer: it stopped the spinning completely, but total spend did not go down, because the subagents switched from spinning to waiting through one long pause instead — and every pause over 5 minutes still rewrites the whole cache.
| Date | Spend that day | Spin turns | Spin cost | Cache-rewrite cost | Active subagents |
|---|---|---|---|---|---|
| Sep 1 | 179M | 843 | 31M | 11M | 145 |
| Sep 2 | 252M | 1,530 | 81M | 19M | 104 |
| Sep 3 | 1,228M | 6,290 | 346M | 124M | 342 |
| Sep 4 | 423M | 1,714 | 78M | 36M | 220 |
| Sep 5 | 649M | 6,834 | 285M | 43M | 94 |
| Sep 6 | 794M | 2,803 | 141M | 51M | 418 |
| Sep 7 | 655M | 4,124 | 152M | 47M | 259 |
| Sep 8 | 623M | 3,778 | 85M | 55M | 288 |
| Sep 9 | 1,364M | 3,902 | 172M | 160M | 354 |
| Sep 10 | 1,010M | 4,107 | 112M | 125M | 295 |
| Sep 11 | 1,286M | 10,377 | 511M | 107M | 304 |
| Sep 12 guard shipped mid-day | 1,400M | 6,865 | 239M | 197M | 486 |
| Sep 13 | 744M | 0 | 0 | 174M | 250 |
| Sep 14 (partial) | 300M | 0 | 0 | 37M | 176 |
| Date | Cache writes | Subagent idle re-writes | Main-session re-writes | Spin turns | Fable cache writes | Guard denials (wait / turn-cap / read / bash-output / idle-resume) |
|---|---|---|---|---|---|---|
| Sep 6 | 117.1M | 31.5M | 4.3M | 2,316 | 10.8M | — |
| Sep 7 | 93.6M | 34.5M | 3.0M | 3,813 | 8.1M | — |
| Sep 8 | 109.1M | 36.3M | 6.2M | 2,005 | 16.4M | — |
| Sep 9 | 253.4M | 110.1M | 11.5M | 2,775 | 37.5M | — |
| Sep 10 | 193.3M | 75.4M | 5.6M | 1,983 | 30.8M | — |
| Sep 11 | 168.6M | 78.7M | 4.3M | 8,018 | 15.5M | — |
| Sep 12 | 286.6M | 144.2M | 3.4M | 5,811 | 19.5M | — |
| Sep 13 | 237.4M | 152.3M | 8.4M | 0 | 24.1M | — |
| Sep 14 | 99.4M | 40.8M | 1.8M | 0 | 15.4M | — |
| Sep 15 | 98.3M | 30.6M | 2.4M | 0 | 11.5M | — |
| Sep 16 | 89.9M | 19.2M | 1.1M | 53 | 14.7M | — |
| Sep 17 | 122.6M | 28.7M | 2.3M | 0 | 17.6M | — |
| Sep 18 | 210.5M | 40.8M | 0.8M | 0 | 24.4M | 0 / 0 / 52 / 548 / 5 |
| Sep 19 | 6.4M | 0.9M | 0.3M | 0 | 0.5M | — |
Targets: subagent idle re-writes under 45M/day (one third of the Sep 12–14 average of 136M), spin turns 0.
Offender #2 turned into offender #1. The next fix has to remove the wait itself from big-context subagents, not just stop them from spinning while they wait — that's exactly what today's wait-guard hook (section 8) does.
Most subagents are short. A median subagent runs 32 turns. But the distribution has a long tail, and that tail is where the money is.
| p50 | p75 | p90 | p95 | p99 | Longest |
|---|---|---|---|---|---|
| 32 turns | 73 turns | 214 turns | 363 turns | 950 turns | 4,057 turns |
What a hard turn cap would have removed, if it only cut off the tail past that point (gross removal — doesn't count the cost of re-spawning a fresh agent to finish the work):
| Cap at | Subagents that ran past it | Spend past the cap | % of 30-day spend |
|---|---|---|---|
| 100 turns | 933 of 4,647 (20%) | 2.24B | 51.5% |
| 200 turns | 505 | 1.60B | 36.8% |
| 300 turns | 296 | 1.17B | 26.8% |
| 500 turns | 147 | 0.64B | 14.7% |
One subagent in five produced half the month's bill, in the part of its life after turn 100 — the part where it had stopped building and started waiting. This is the direct evidence behind today's 300-turn cap (section 8).
| Claim | Verdict | Why |
|---|---|---|
| Fixed per-session overhead (CLAUDE.md, memory, skills) | Small | ~26k tokens per session (main sessions 29.5k, subagents 26.2k); after the first turn, near-zero grows in |
| The 1.77MB portfolio-architecture doc, read in full every session | Refuted | Sessions read only the slices they need, never the whole file |
| Headless bots (Smith, Trinity, cron) | Small | 2.9% of all spend |
| Guard hooks (poll-guard, load-guard, etc.) | Zero cost | They emit no output on a normal pass-through |
Two real audits happened before this one, a month apart, both triggered by Fede hitting a wall rather than by anyone checking proactively.
| Audit | What it found | What happened to it |
|---|---|---|
| Aug 15 | Every subagent defaulted to Opus | Fixed in config — global default is now Sonnet |
| Aug 15 | Route work by task: Haiku for pulls/extraction, Sonnet as the workhorse, Opus for hard cases; drop effort for routine agents | Rule only — written in CLAUDE.md, no hook enforces it; 42% of launches still skip the model param |
| Sep 12–13 | Subagents burning turns spinning on CI/status checks | Fixed with a hook — poll-guard shipped same night, and it works (see section 4) |
| Sep 12–13 | "Auto-compact is now automatic at 250k" | Never actually set — no such setting exists anywhere on disk; today's audit fixed the real setting (section 8) |
| Sep 4 | CLAUDE.md + CONSTITUTION.md cost ~57k tokens/session before pruning | Partially done — pruning is still pending review |
The pattern across both prior audits: the finding that got a hook (poll-guard) actually stuck. Everything left as "a rule in CLAUDE.md" — model routing, effort levels, orchestrators not doing worker jobs — kept happening anyway, because nothing on the machine enforces it.
All five below are live and verified on disk — not proposals.
CLAUDE_CODE_AUTO_COMPACT_WINDOW). It was effectively at ~967k, right up against the 1M ceiling. Fede tried 400k first, then brought it down to 300k; 200k was tried and rejected as too aggressive for orchestrator sessions.gh --watch, run watch — inside a subagent whose context is already 100k tokens or bigger. Tested 8/8, and verified against a real 732k-context transcript.tools/token-audit/, so the next audit doesn't start from scratch.The infrastructure to do this the right way already exists — it just isn't being used consistently.
These are recommendations, not done yet. Both are cheap and low-risk; neither touches anything customer-facing.
| head / jq discipline in prompts, is the highest-leverage lever left on the table.| Claim | Verdict | What the data actually says |
|---|---|---|
| An account can go from a fresh reset to capped in under a day | False | Fastest observed in 30 days: 27.6 hours. Most accounts take 55–142 hours. |
| The 1.77MB portfolio-architecture doc is read in full every session | False | Sessions only ever read slices of it |
| Fixed per-session overhead is a major driver of spend | False | ~26k tokens/session — small next to the 2.87B-token rewrite category |
Everything below is the original Aug 15 audit, unchanged, kept for the record. Where an item from it is now live, it's marked.
A subscription doesn't bill dollars, so to compare consumption across models we weighted every token by its published API price — the best available proxy at the time for how hard each token pushes on the 5-hour and weekly limits.
Live now By the Sep 14 audit, Sonnet and Opus have real usage (48.2% and 36.7% of weighted spend) — the "essentially unused" cheap-model problem from Aug 15 no longer exists at the model-choice level. The remaining leak is behavioral (idling, spinning, runaway turns), not "everything defaults to Opus."
The subagent model was pinned to Opus globally, on the assumption that agents are always coding. Most fan-outs are actually research, extraction, search-pulling, summarizing, verifying — work a smaller model does at a fraction of the pressure.
| Task the agent is doing | Model | Why |
|---|---|---|
| Pull / collate raw search results, scrape, extract fields | Haiku | Routine, verifiable, high volume |
| Summarize, grep / explore a codebase, first-pass drafting | Haiku | Cheap, fast, good enough |
| Research a profile, verify facts, normal implementation | Sonnet | The workhorse; far lighter than Opus |
| Adjudicate ambiguity, hard debugging, architecture, final synthesis | Opus | Worth the pressure — but the minority of work |
Live now Global default is Sonnet, not Opus. Still rule-only Per-task routing to Haiku/Opus depends on each launch remembering to pass a model — 42% still don't.
Across the Aug 15 window we found roughly 1,500 limit hits clustered into 25 "walls" — moments where an account capped and everything on it died at once, 85–99% of a swarm's workers within five minutes of each other.
| Assumption | Verdict | How we knew |
|---|---|---|
| Subagent swarms all ran Opus | Confirmed | 100% of 6,685 subagent runs, 30-day scan |
| The rotation tool refreshes the account a live session holds, killing it | Wrong | Code check: it never refreshes the active account |
| Several long-running sessions on one account revoke each other | Strongly inferred | Refresh tokens are single-use families; stacking 4 sessions on one account killed it |
| Keychain is the source of truth, the file is secondary | Confirmed | The switch writes Keychain first — the two can desync, Keychain wins |
| cron can't do the account switch | Confirmed | Keychain write exit 195, crontab "operation not permitted" — twice |
| A session launched from inside another hangs before its first request | Confirmed | Reproduced 3x |
| Hot-swapping the account under a running session works | Promising, unproven | New account's meter rose while old held, but subagents confound it |
| apiKeyHelper accepts a subscription token (clean rotation path) | Untested | Couldn't test from a nested session |
| A per-request pool proxy | Avoid | Needs cloaking; failure mode is an account ban |
budget tool reads live pool headroom and sizes fleets against it.| Piece | What it does | State as of Sep 14 |
|---|---|---|
| budget | Reads live pool headroom, sizes a fan-out, downshifts the model when short | Live in production use |
| ccd | The supervisor loop — proactive rotate, park when capped, sleep-until-reset, spread across accounts | Live under launchd, 38 sessions resumed |
| checkpoint contract | The brief block every long worker carries so progress lives on disk | Now standing in the fleet skill |
| exp1 / exp2 | Terminal experiments settling live-swap and apiKeyHelper | Still not run |
message.id, cross-checked against a from-scratch rewrite of the original parser · 26,686 usage-meter readings, Aug 15–Sep 14 · Aug 15 appendix: 9,159 transcript files, rotation-tool cache + logs, live swap test on the workstation.