We Metered Our Own AI Context. Forked Agents Cost 3x Per Result, and 78% of Our Memory Index Was Never Used
Everyone running AI coding agents is paying for context. Every call re-reads the instructions, the memory, the skills and the conversation so far. Almost nobody measures which of that context actually gets used. So we measured ours: 218 Claude Code sessions over about 30 days, in a shop of two people, using our own meter. Two numbers came back that changed how we work the same afternoon.
Number one: the fork tax was 16% of everything we spent
When an agent hands work to a helper, it can start that helper fresh with a short brief, or it can fork itself, so the helper starts with a full copy of the conversation. Forking feels free. It is not.
Across 349 helper agents, the 46 forks cost $1.58 for every thousand tokens of output. The 303 fresh ones cost $0.51. The forks were 13% of the helpers and 63% of helper spend. If their output had been produced at the fresh rate, it would have cost $120 instead of $376. That $256 difference is 16% of everything we metered. A fork's first call already carried a median 435 thousand tokens it had inherited, and it paid for those again on every call after.
The honest caveat: forks tended to get bigger jobs (a median of 20 calls against 5), so this is not a perfectly matched comparison. But the gap per call is too large for task size to explain away. Our rule was already "fresh agent plus a brief, never a fork." Now the rule has a dollar figure behind it.
Carrying context inside one conversation is a different thing, and it is fine. In the session where we found all this, 98.7% of the context we read was cache reads, the cheapest tier, and it was our own memory of the work we had just done. Remembering is cheap. Copying the memory into a second agent is what costs money.
Number two: most of our memory index was never used
We keep a memory of decisions, rules and lessons, with an index that loads into every session. We counted which memories the work actually reached for, meaning read them or named them. In 30 days that was 79 of 424. Of the 314 in the always-on index, 245 were never touched, and their index lines rode along on every single call.
So we tiered the index by use. Hard rules and anything that was actually used stay always-on. The rest moved to a second index that is one search away. Nothing was deleted, and the 189 index lines that existed before still exist. The always-on index went from 25.3 KB to 17.4 KB. A side effect we did not plan: the index had been over the size where the tail gets cut off, so two whole sections (our business and people notes) had silently stopped loading. They load again.
The caveat here is the kind we put in every post: "used" counts only explicit reads and mentions. Following a rule without naming it does not show up. So "never used" is an upper bound on waste, not proof of it. We are now running the before-and-after comparison: corrections, retractions and missed memories over the next two weeks.
What we shipped: context reuse, on finops.dugganusa.com
The meter at finops.dugganusa.com now has a Context reuse table. Point it at your Claude Code logs and it lists which skills, memories, rules and project instruction files your agents actually reached for, counted by sessions. It runs in your browser and uploads nothing. It keeps only the name of each item, never the file path and never any text, and we have tests that fail if a username or a path ever leaks into the result.
Why counted by sessions: reuse across work is the value. A skill that shows up in twenty sessions is carrying your team. One nobody reaches for is cruft, and if you do the job right, cruft sorts itself out. You stop maintaining what nothing uses.
There is a second use for the same number, and it is the one we find most interesting. Reusable context is written by people. How often the rest of a team's sessions reach for the skills and memories one engineer wrote is a direct measure of that engineer's reach across the team, the kind of impact that never shows up in a commit count. Our command-line version can attribute each item to whoever first committed it. We report that per contribution, never as a ranking of people. The point is to find the context that carries a team and the people who wrote it, not a leaderboard to game.
The third cost we started counting
One more number from the same afternoon, because it belongs in the same ledger. Coding agents now run behind safety gates that decide whether an action is allowed. We like gates: our own rule is that nothing deploys until a human types a confirmation word. But a gate that does not know your protocol stops work you already authorized. We counted: 35 such stops in 7 sessions over the last month, 23 of them in the first eight days of October. We taught the gate our confirmation word, and within the hour the deploy step it had blocked went through. That is a ledger now too, because the only way to know whether a fix worked is to count before and after.
What this means if you run agents
Measure which context gets used before you add more. Start helpers fresh with a brief instead of forking them. Tier what loads by default, and keep the rest searchable. And count the reuse: the skills that show up everywhere are the cheapest, most valuable things your team owns. They are worth far more than a better prompt, because a prompt gets used once and a skill keeps paying off.
Try it on your own logs at finops.dugganusa.com. The dollars are list-price equivalents, not an invoice; a flat subscription pays something else, and we say so every time.
Was this useful? Rate this post. The widget is at the bottom of the page, and we read every response.
How do AI models see YOUR brand?
AIPM has audited 250+ domains. 15 seconds. Free while still in beta.
Was this useful? Thirty seconds, no cookies, no tracking, no third parties, your address hashed and never stored. If the box below does not load, the same question lives at https://analytics.dugganusa.com/nps.html?post=we-metered-our-own-ai-context-fork-tax-context-reuse



Comments