Tomás Alcalde Leer en español

Context economy

What makes an agent expensive is not what it does, it is what it remembers

313 real working sessions with a coding agent, 2.36 billion tokens. 94 % of that went into re-reading old conversation. Here are the numbers, and what I did about them.

August 2026 · Tomás Alcalde · 8 min read

The bill started to hurt before I understood why. The work was the usual: queries against a production database, reports, some automation, email. Nothing that sounded expensive. But the consumption kept climbing and it did not match the feeling of asking for very little.

So I measured it. Coding agents store every session as a .jsonl file on disk, with the token accounting of each call to the model. They are sitting right there, for free, and almost nobody opens them. I wrote two scripts that read those files and answer two different questions: how much did each session spend, and which specific call inflated the context. What came out changed how I work, and it forced me to correct a rule I had written down wrong myself.

313
sessions measured, the 120 days before 27 August 2026
2.36 B
tokens consumed in total
94 %
of that total is re-reading context already accumulated
0.8 %
is what the model actually wrote

First: what you pay for is not the work

The natural intuition is that an agent costs in proportion to what it produces. It is exactly the other way around. Out of 2.36 billion tokens, what the model wrote (the code, the answers, the analysis) came to 19 million. Less than one percent.

Stacked bar: 94.4 % re-reading accumulated context, 4.8 % new context, 0.8 % output.
The 2.36 billion tokens split by type. The red band is cache_read: context that was already there and gets read again on every turn.

The reason is structural, and it holds for any conversational agent, not just this one. Models have no memory between calls: on every turn the whole conversation is sent again. Turn 40 means reading the previous 39 turns one more time. Prompt caching makes that re-reading a lot cheaper, but it does not remove it, and above all it does not change the fact that the amount grows on its own while you work.

The mechanism: the same work, more expensive each time

This is the chart that organised the diagnosis. I took the 45 sessions longer than 120 turns, split each one into ten equal stretches, and averaged how much context the model read in each stretch.

Rising curve: context per turn goes from 69k in the first tenth to 200k in the last.
Context read on each turn, by progress through the session. Average of 45 long sessions. The last stretch costs 2.9 times the first one, doing the same kind of work.

A session starts out reading around 69 thousand tokens per turn and ends reading 200 thousand. Nobody asked for anything bigger at the end than at the beginning. It is the same question, with more history behind it.

Averaged over the whole curve: each turn of a long session costs close to twice what that same turn would cost with the session freshly started. That is the tax, and it is paid without ever showing up anywhere as a separate line.

What I believed, and got wrong

After the first measurement I wrote down a rule that sounded right: cost grows with the square of session length. The reasoning was clean. If context grows linearly with turns and every turn re-reads all of it, the sum is quadratic. A 300-turn session should cost five times what four 75-turn sessions cost.

When I fitted the curve against the 313 real sessions, the exponent did not come out at 2. It came out at 1.20.

Log-log scatter of tokens per session against turns, with a fit of exponent 1.20 and R² 0.97.
Each dot is a session. On a log scale a power law shows up as a straight line and the slope is the exponent. Fit over 313 sessions, R² 0.97.

The gap between 2 and 1.2 is not a blackboard detail, it changes the recommendation. At exponent 2, splitting one long session into ten saves 90 %. At 1.2 it saves about a third. Still a lot, and a third of the bill is a third of the bill, but it is a different conversation.

Why did the theory overshoot? Two things it ignored. There is a large fixed floor: the first call of every session already starts at 59 thousand tokens on median, before you type anything, because the preamble with the project instructions and the tool definitions always travels. That floor is paid by every session, short or long, and it dilutes the part that grows. And long sessions do not grow forever either: they hit the context window and the history gets compacted.

I am leaving this in because it is the part that normally does not get published: the rule I was applying came from correct reasoning on an incomplete assumption. The number corrected it. A mental model that nobody checks against their own data turns into technical folklore fairly quickly.

Where the volume comes from

Knowing that you pay for re-reading does not tell you what to stop reading. That is what the second measurement is for, and it is the one that actually drives decisions.

Every tool result (the contents of a file, the output of a command, an email search) enters the context and stays there. Its real cost is not its tokens: it is its tokens multiplied by the turns that came after it. A 20-thousand-token dump on turn 3 of an 80-turn session costs more than a 100-thousand-token one on the last turn. I call that carry, and sorting by that column instead of by size completely rearranges the list of culprits.

Horizontal bars: reading whole files 54 %, shell commands 28 %, connectors 7 % and 5 %, the rest 3 %, filtered search 2 %.
380 million tokens of carry over the window measured, split by type of call. Reading whole files takes more than half.

54 % is reading whole files. And inside that, the biggest surprise: a good part of it was not code files but the agent's own memory files. The notes you accumulate so the assistant knows how the business works. Every time one of those notes becomes relevant it gets loaded in full, and the ones I had let grow unchecked were over 9 thousand tokens each. The system that existed to make the agent more efficient was the main source of variable cost.

The pattern: filter before it enters, not after

All of the above converges on a single operating rule, and it is boringly simple: whatever enters the context is paid for many times, so it has to be filtered before it gets in. Not afterwards, not by summarising it on the next turn. Once the file is inside, the damage is done for the rest of the session.

In practice that meant writing a thin filtered-reading layer. A reader that searches by regular expression with surrounding context, with a hard character cap on what reaches the screen and the full content dumped to a temporary file in case it is needed. It is very little code and has no intellectual charm whatsoever. The effect does:

Listing the last six emailstokens entering the context
Standard mail connector2,730 (average of 190 calls)
Local filtered reader258

An order of magnitude, for the same data on screen. And because the saving multiplies across every turn that comes after, in a 100-turn session that single decision is around 250 thousand tokens you never pay.

The other three that stuck as habits, in order of how much they returned:

One I evaluated and dropped: compacting the history pre-emptively at half the window. It does save, but it compresses exactly the fine detail (precise figures, file state) and in work with production data that detail is the one thing I cannot afford to lose to a mediocre summary. It also treats the symptom. If a session is getting heavy, ending it and opening another beats summarising it blind.

How to measure your own

None of this needs special instrumentation or a vendor that hands you telemetry. The transcripts are on your disk and they carry the usage block of every call, with input, output and tokens read from cache. A two-hundred-line script walks them and builds the session ranking. Another one attributes each tool result to its call and multiplies it by the turns that followed.

What matters is not the script but the order of the questions, and there are three:

  1. What share of the spend is re-reading rather than producing? (If you get less than 80 %, check the measurement.)
  2. How much does context per turn grow within a single session, start to finish? That is the tax, and it is what decides whether cutting sessions is worth it.
  3. Which specific calls carry the most? That is your list of things to learn to read filtered.

Mine are in this repository, together with the script that draws these four charts. It runs with a fixed cut-off date, so the window is reproducible rather than drifting on every run: python graficos_articulo.py --dias 120 --hasta 2026-08-27. None of the three is long, around two hundred lines each, and the questions above matter more than the code: what is worth having is the ranking of your own transcripts, not mine.

Why this matters outside my machine

These measurements come from one person working at a company in southern Chile. They are not a study. But the mechanism does not depend on scale, and the companies that started putting agents to real work this year are going to run into the same thing when the first bill arrives that does not match the intuition.

What I take away is that context discipline looks far more like 1990s memory management than like prompt engineering. It is not about writing better instructions. It is about deciding, deliberately and sometimes painfully, what does not get in.

And about measuring it, because whoever does not measure their context ends up paying to read it.

Frequently asked questions

The questions that came back after publishing this, each with the figure that answers it. All of them come from the same measurement described above.

Why does a coding agent get more expensive as the session goes on?

Because the model has no memory between calls: every turn resends the whole conversation. Measured across 45 long sessions, the context read per turn goes from 69 thousand tokens in the first tenth of the session to 200 thousand in the last one, doing the same kind of work. That is 2.9 times more expensive purely for carrying more history behind it.

How do you cut token usage in Claude Code or any AI agent?

Four measures, ordered by what they actually returned: start a new session when the topic changes; send large outputs to a file instead of into the conversation; use grep and read the exact range rather than opening whole files; and filter what enters the context before bringing it in, not by summarising it afterwards. Replacing a mail connector with a filtered local reader took the same on-screen data from 2,730 tokens down to 258.

Is it better to end the session and open a new one, or keep going?

Better to end it. Cost per session grows with turns raised to 1.20 (fit over 313 sessions, R² 0.97), so splitting one long session into ten saves around a third. Beware the rule of thumb that growth is quadratic: an exponent of 2 would promise 90 %, and against real data the exponent is not 2.

Does prompt caching not solve this?

It makes it cheaper, it does not remove it. In the measurement, 94 % of the 2.36 billion tokens is cache_read: context that was already there and gets read again every turn. Caching lowers the unit price, but it does not change the fact that the volume grows on its own while you work.

Do the agent's memory files help, or do they cost?

Both, and it has to be measured. Reading whole files accounts for 54 % of the carry, and much of it was not source code but the agent's own memory notes: each one loads in full whenever it becomes relevant, and the ones left to grow unchecked were over 9 thousand tokens each. Split them into sections and read only the section you need.

How do you measure where an agent's tokens go, without vendor telemetry?

The transcripts are on your own disk as .jsonl files, carrying the usage block of every call with input, output and cache-read tokens. A script of about two hundred lines builds the session ranking; another attributes each tool result to its call and multiplies it by the turns that came after, which is the carry. All three are published in context-economy.