I've been running AI coding agents against real codebases for a while now, and one thing became obvious fast: the API bills don't scale linearly with the work done. They explode. A task that feels small — "fix this one function" — can quietly chew through an order of magnitude more tokens than you'd ever expect from a chat completion.
Then a recent paper, "How Do AI Agents Spend Your Money? Analysing and Predicting Token Consumption in Agentic Coding Tasks", put numbers on the thing I'd been watching in my own billing dashboard. The headline figures are worth sitting with.
So I dug into the paper to figure out exactly where the tokens go — and how Synapse MCP, which I've been building for exactly this reason, attacks the problem at the source.
Where the tokens actually go
The naive mental model is that an agent writes a lot of code, so output tokens are the cost driver. The paper shows the opposite. When an agent works a real issue, it doesn't just emit a solution — it explores the repo, reads multiple files, runs tests, and iterates. Every file it pulled in during turn 1 comes back as input context in turn 2, turn 3, and every turn after that. The cost is dominated by what gets fed in, not what comes out.
Here's the cost dynamic the paper lays out:
- Input tokens dominate. Roughly 80% of agentic spend is input, not output. As the trajectory grows, the full conversation history — every file read, every tool response — is replayed into the model on every turn. Each turn compounds on the last.
- Redundant exploration burns money for nothing. The paper found that higher token usage doesn't buy you higher accuracy. Performance often degrades at the top of the cost range. The expensive failures are the ones where the agent re-reads the same files over and over, spinning in place without making progress.
- Prompt caching helps, but doesn't save you. Even with caching making re-reads cheaper per token, the sheer volume of accumulated context means the cheap cache reads still outweigh the expensive output generation in aggregate. The volume problem swamps the per-unit discount.
Agents spend your money by pulling raw file content into their context window while stumbling through a codebase without a map. The problem isn't the price per token — it's the quantity of tokens needed to navigate blind.
What Synapse MCP does about it
The paper makes one thing clear: you can't fix this by just caching harder. You have to stop the context from bloating in the first place. That's the design principle behind Synapse MCP — a Model Context Protocol server that exposes a persistent code-knowledge graph. Three mechanisms target the exact failure modes the paper identifies.
Precision retrieval: kill the grep tax
The paper pins a lot of cost on inefficient search. Agents waste millions of tokens running crude grep and find commands, dumping irrelevant text blobs into context, and trying to orient themselves in an unfamiliar codebase by brute force.
Synapse sidesteps this entirely. Because it maintains a persistent code-knowledge graph, agents don't grope around the file system. Tools like ask_synapse and synapse_search_codebase retrieve precise context in a single turn — a semantic search for a concept, an exact symbol lookup for a function definition, or a graph query for all dependencies. The right code lands in context on the first try, so tokens go to problem-solving instead of navigation.
AST-based outliner: scan structure, skip the bodies
During exploration, agents often need to scan multiple files to map API boundaries, module definitions, and function signatures. The naive approach — dumping entire raw files into context — is catastrophically expensive. I've watched agents do this and burn a full context window before they've even started reasoning.
Synapse lets agents request an outline format instead of full file content. It uses language-specific parsing (Sourceror for Elixir, indentation tracking for Python, bracket matching for JS/TS) to strip out function bodies entirely, keeping only the structural skeleton.
An outlined Elixir function collapses to this:
def compare_versions(v1, v2) do
...
end
Keeping only function heads, docstrings, and type specs means an agent can scan 20+ files in a single turn for a fraction of the cost. It gets the architectural picture without polluting the context window with implementation details — a 60–80% token reduction per file.
JSON SmartCrusher: zero-loss payload minification
Every tool response an agent receives becomes input context on the next turn. Since input tokens dominate agentic cost, even small per-response savings compound hard across a long session.
Synapse ships with a SmartCrusher enabled by default on every tool response. It minifies the JSON payload before it reaches the LLM:
- Key shortening: verbose keys get abbreviated (
file_path→fp,chunk_id→cid,language→lang). - Pruning: empty lists, empty maps, and null fields are stripped automatically — no tokens wasted on absent data.
- Path relativisation: absolute workspace paths become short relative paths.
The result is 30–60% fewer response tokens with zero information loss. The minification happens inside the tool dispatcher, so the agent gets all the data it needs while occupying less space in the context window.
| Response Type | Without SmartCrusher | With SmartCrusher | Saving |
|---|---|---|---|
| Symbol lookup | ~1,200 tokens | ~480 tokens | 60% |
| File inspection (outline) | ~3,400 tokens | ~680 tokens | 80% |
| Graph traversal (callers) | ~890 tokens | ~445 tokens | 50% |
| Codebase overview | ~2,100 tokens | ~840 tokens | 60% |
Why this compounds, not just adds up
The paper's most important finding, and the one that matches what I see in practice: token costs in agentic tasks aren't linear, they compound. A file read in turn 3 becomes part of the context that gets re-processed in turns 4, 5, 6, and every turn after. An agent that reads ten unnecessary files early in a session pays for those files dozens of times over by the time it finishes.
This is exactly why Synapse's approach works. It doesn't just trim tokens on individual queries — it keeps the context window from bloating in the first place. A leaner context at turn 3 means a leaner context at turns 10, 20, and 50.
| Approach | Discovery Method | Tokens (single turn) | Compounding Impact |
|---|---|---|---|
| Shell grep loop | 5–10 grep+read cycles | ~18,000–25,000 | High — all context replayed each turn |
| Synapse MCP (full) | Single graph query | ~2,000–5,000 | Low — compressed context stays lean |
| Synapse MCP (outline) | Single graph query + AST outline | ~400–900 | Minimal — structural skeleton only |
Context is capital
The paper frames token consumption as capital expenditure, and that framing is the right one. Just as you wouldn't read every file in a repo before fixing a bug, an agent shouldn't have to — but without the right tooling, that's exactly what it does.
At frontier-model pricing of $3–$15 per million tokens, and with agentic tasks eating 1,000× more tokens than simple reasoning, the gap between a well-tooled agent and a poorly-tooled one isn't marginal. It's the difference between a productive workflow and a runaway API bill. Synapse MCP Pro at $19/month pays for itself in the first week of active development through raw API savings alone — before you even count the less visible costs the paper calls out: agent failures from context exhaustion, redundant re-exploration, and the opportunity cost of slower, less accurate sessions.
The bottom line
Context is king, but context is expensive. The paper shows that throwing more raw tokens at a problem leads to diminishing returns and bloated bills from inefficient, redundant exploration. Synapse MCP combines a code-knowledge graph for precision retrieval with an AST outliner and the SmartCrusher to keep the context window lean — so an agent's token budget goes to reasoning and problem-solving, not re-reading the same verbose JSON keys, unneeded function bodies, and irrelevant grep results turn after turn.
If you're building or deploying coding agents, smart context compression isn't a nice-to-have. It's a requirement for economic viability at scale.
Stop Paying Your Agent to Grep
Full AST indexing, 50+ languages, SmartCrusher compression, worktree overlay, and unlimited repos — free, forever. Add safe writes, crash resolution, and dead code decisions for $19/mo.
Download Now — for FREE →