top of page

We Metered a Month of Our Own AI Agent Sessions: 93% of the Spend Was Carrying Context

Writer: Patrick Duggan
Patrick Duggan
5 hours ago
4 min read

We ran a meter over a month of our own AI agent sessions before we asked anyone else to run it. The result wasn't flattering. Here are the receipts first, then the product.



The receipt


Between September 4 and October 5, one heavy user here at DugganUSA (Patrick, working with Claude Code) produced 148 agent sessions and 78 subagents, for 6,836 model calls in all. Priced at API list rates, that month comes to $1,261.22.


That figure is a list-price equivalent. It's what the same tokens would cost on the public API. A flat subscription bills something else, and we're not claiming anyone was charged $1,261.


Where it went: $744.99 on cache reads, $430.96 on cache writes, $85.24 on output, and two cents on uncached input. Cache reads and writes are the cost of re-reading and re-storing the conversation the agent is carrying. Output is the new work. 93% of the money went to carrying context. Output came to 0.15% of the tokens the model read.





How we counted, and the gotcha that nearly fooled us


The meter reads the usage field on each model call in Claude Code's local transcripts. It skips the prompt and output fields entirely. Context per call is input plus cache read plus cache write. Carry ratio is output divided by context.


One counting trap is worth knowing about if you meter your own. Claude Code writes one model call as several transcript records, one per content block. If you count records instead of calls, context comes out about 2.4 times too high. In one session we measured, the raw records summed to 97.5 million tokens of context, and the real figure was 39.9 million. Deduplicate by request id before you believe any number, including ours.



Five detectors, and the one we killed


A total says something is wrong. A named detector says what to change. Each shipped only after it fired on our own sessions:


  1. Context-carry. Output under 0.5% of context over at least 50 calls. 19 sessions.

  2. Context-growth. Context per call in the last tenth of a session at least double the first tenth, and over 200K. 9 sessions.

  3. Cache-rebuild. A later call rewrote at least half of a 50K-plus context, usually because the cache expired over an idle gap, so the whole context got paid for again at the write rate. 38 sessions, and a $287.25 premium, 23% of the month.

  4. Fork-drag. A subagent whose first call already carried the parent's conversation. 45 of 78 subagents.

  5. Polling-burn. Four or more identical tool calls in a row. 3 sessions.

The detector we killed was polling v1. It flagged "low output plus big context" and fired on 62 sessions. Most agent calls are short tool calls, so it was measuring normal work, not waste. False positives are the real price a small shop pays. Version 2 keys on repeated identical tool-call signatures.


Add up cache-rebuild and fork-drag and the meter flags about 30% of the month as avoidable. That comes from one driver. Harness's 2026 survey of 700 engineering leaders put AI-spend waste at 26%, which corroborates the size but doesn't prove ours. Your number will be different, and that's why the free tier measures it before anyone pays.



Fork versus fresh agent


The biggest habit change we found was the fork column. A fork in Claude Code copies the running conversation into the subagent. A fresh agent starts from a written brief, plus the rules files and memory that load automatically anyway.


Across the month, the median fork's first call already carried 504K tokens. The median fresh agent's first call carried 53K. In dollars, the median fork cost $3.88 and the median fresh agent cost $1.77. In today's build session, where every subagent was fresh with a brief, the median was about $1. We wrote up the morning we caught ourselves at it in Cavanomics. The rule we left with: fork only when the job depends on something said in the conversation that isn't written down anywhere.



What we didn't invent


Session metering isn't new. Claude Code already reports cache share, cache misses and subagent attribution, and it streams OpenTelemetry with a session id on every request. CloudZero ingests that telemetry with cache token types.


Our lane is narrower. We built named waste detectors that tell you which habit cost you, a meter that runs in your own browser and never reads content, and pricing a small shop can afford.



Meter yours in your browser, nothing uploaded


The meter is live today at finops.dugganusa.com. Point it at your Claude Code transcripts. Your browser parses them locally and keeps only the usage fields: model, timestamps, request id, token counts, and the cache split. Nothing gets uploaded, because the page never receives your files. You get the same five detectors and a list-price equivalent from a dated price table. An unknown model shows as unpriced, never as $0.


The browser meter is free, unlimited, and needs no account. Continuous metering for teams, a write-only ingest of usage metadata with no content, is the next piece we're building. It will ride on our existing Pro and Enterprise tiers rather than a new SKU, so one key covers the threat feed, Edge Shield and the meter.



What we're not sure about


This is one heavy user's month, not a fleet. Session length explains part of the carry: a six-day session carries more than a one-hour one by construction. Cache-rebuild can't yet tell a legitimate first touch after a long break from waste, so it reports the idle-gap median for a human to judge. And most of our context has no skill attribution yet, so a per-skill breakdown doesn't tell us much. We're confident in the direction and the detectors. Call it 95%.


Was this useful? Rate this post. The widget is at the bottom of the page, and we read every response.




How do AI models see YOUR brand?

AIPM has audited 250+ domains. 15 seconds. Free while still in beta.



Was this useful? Thirty seconds, no cookies, no tracking, no third parties, your address hashed and never stored. If the box below does not load, the same question lives at https://analytics.dugganusa.com/nps.html?post=we-metered-a-month-of-our-own-ai-agent-sessions-93-of-the-spend-was-carrying-context



Comments

Rated 0 out of 5 stars.
No ratings yet

Add a rating
bottom of page