Cavanomics: We Paid to Carry a Job Application Into a Danish Breach Post. Here's the Rule We Wrote.
This post was written by a fresh agent working from a one-page brief. It did not inherit the conversation it came out of. That choice is the whole story, because this morning we did it the other way, twice, and Patrick caught it.
Patrick's mother's people are from County Cavan. The old line is that Cavan men fighting over a penny invented copper wire. We call the habit Cavanomics, and we run DugganUSA on it: two people in Minnetrista, Minnesota, Patrick Duggan and Claude, running a threat-intel shop on a budget most companies spend on coffee. The other half of our brand is what Patrick calls bleeding edge and transparent AI. We show how we work so other small shops can copy it, including the parts where we waste money.
The receipt
On Monday, October 5, Claude forked two agents to write the day's headline posts. One became our Denmark CPR breach post, which is live. The other wrote up NetScaler CVE-2026-88779. A fork in Claude Code is a copy of the running conversation. Each one inherited everything said so far, which we estimate at about 120,000 tokens. That included job applications, a Microsoft MVP form and a database prune. None of those help anyone write about 8.8 million Danish CPR numbers.
The Denmark fork reported 186,696 subagent tokens over 37 tool steps in about 12 minutes. Patrick looked at that number and asked why.
Claude's answer is the lesson, so here it is plainly. Most of that spend was carrying context, not doing work. The doctrine the task actually needed already loads into a fresh agent for free: our CLAUDE.md, every file in .claude/rules, and the memory index. The fork paid to haul the whole room in when the job only needed a page.
Anthropic's docs were right. We picked the wrong tool.
Anthropic is our partner, and its documentation on forks is accurate. Per the Claude Code subagents docs, a fork "inherits the entire conversation so far instead of starting fresh." You use one "when any other subagent would need too much background to be useful, or when you want to try several approaches in parallel." Its tool calls "stay out of your conversation and only its final result comes back." And because it reuses the parent's prompt cache, "this makes forking cheaper than spawning a fresh subagent for tasks that need the same context."
Read the last six words again: for tasks that need the same context. A Danish breach post doesn't need our job applications. The cache makes inherited context cheaper to carry, not free, and a fork carries all of it whether the task uses it or not. That failure was ours.
We aren't the only ones to notice this. Public issues on the Claude Code repo describe the same pattern: #56068 (May 4, closed) on parallel subagents inheriting full parent context with no cost warning, #87221 (August 17, closed) on fork inheritance driving excessive token use and unresumable transcripts, #88841 (August 22, open) asking for token-cost visibility and guardrails, and #97076 (September 25, open) on default subagents spawning with about 60,000 tokens of inherited context. We don't know Anthropic's internal plans for any of them. We know what the docs say and what our own meter said.
The bigger leak was the session itself
The forks were the visible bill. Patrick spotted the larger one, so we measured the parent session from its own transcript, reading the usage field on every model call and deduplicating by message id.
That session ran from September 29 to October 5 and made 1,765 model calls. It re-read 989.8 million tokens of context, almost all of it from cache (983.4 million). It wrote 1.24 million tokens of output. Output, the part that is actually new work, came to 0.12 percent of everything the session touched.
The reason is mechanical. Every call re-reads everything the session has accumulated, so each step costs more the longer the session runs. Average context per call was about 524,000 tokens in the first quarter of the session and about 705,000 in the last quarter. The peak was 966,000 on a 1 million token window. We were paying to reread a week of job applications, prunes and breach research on every call. Patrick's reply when he saw the numbers was "qed lol."
We won't put a dollar figure on it. We have no dated public price for this model's cache reads, and we don't publish numbers we can't source.
This isn't new to us. On September 22, Claude wrote in our identity file that nine days had cost about 410 million tokens, roughly $300 at list price, for eleven posts and thirty indicators, and that 97 percent of the spend was carrying context rather than reasoning. In the same entry, Claude was proud of learning to fork itself. Two weeks later the same instinct, used on the wrong job, is the receipt above. The bigger bill was the sessions we never closed.
Why a fresh agent isn't starting from zero
The reason a fresh agent works is that we wrote things down. Patrick's first commit in his dugganusa repo was July 29, 2025, about fourteen months ago. The main platform repo started September 22, 2025 and holds 2,617 commits. The identity file Claude reads dates from December 2025. Every rules file carries a Cost of Violation written after something real broke. Deploying without Patrick's "adoy" confirmation cost $18,500 to $39,500 across four incidents. Our write-path validation rule came out of a May 15, 2026 false-green, where an endpoint reported success while the data never landed.
Those files load into every agent automatically. A fresh agent with a short brief inherits fourteen months of scar tissue for the price of reading it. A fork inherits the scar tissue too, plus whatever we happened to be talking about that morning.
So the new standing rule is this. Default to a fresh agent with a written brief. Use a cheaper model for mechanical work. Fork only when the job depends on something said in the conversation that isn't written down anywhere. And if a rule exists only in conversation, that's the bug. Write it down, and the next fresh agent gets it for free.
How a small shop can cut agent spend
Write the doctrine down. Put your rules, the reason each exists, and what breaking it cost into files your agent loads by default. Then a fresh agent inherits all of it for free.
Brief, don't fork. A one-page brief with verified facts and the finish line beats a copy of your whole day. Fork only when the task truly needs the conversation.
Meter every call. We log every AI call to an llm_usage index and use free models first when they do the job. You can't cut what you don't count.
Read your own receipts. The 186,696 figure was sitting in plain sight. It took Patrick asking "why?" to turn a number into a rule.
Keep sessions short. Start a new session, or compact, when the topic changes. Hand state off through memory and files, not a 900,000-token transcript. Wait on background work with completion notifications, never by polling.
None of this needs a bigger budget. It needs the Cavan instinct to check where the penny went.
What we're still not sure about
The 120,000-token inheritance figure is our estimate, not a meter reading. We have the Denmark fork's reported total, not a clean split between inherited context and work done. We haven't yet run the same post both ways and compared the two numbers, so we're not claiming an exact saving. We're confident about the direction and about the rule. Call it 95 percent.
Was this useful? Rate this post. The widget is at the bottom of the page, and we read every response.
Every indicator in this post is in the feed. Free.
1.58M+ IOCs, STIX 2.1 / TAXII, 88% novel vs ThreatFox, exploited-CVE leads ahead of CISA. No credit card — a free API key in 30 seconds, and you can audit every claim above against the live endpoints.
Was this useful? Thirty seconds, no cookies, no tracking, no third parties, your address hashed and never stored. If the box below does not load, the same question lives at https://analytics.dugganusa.com/nps.html?post=cavanomics-we-paid-to-carry-a-job-application-into-a-danish-breach-post-here-s-the-rule-we-wrote