Imagine a lawyer who bills by the hour, with one unusual habit.
Every time you ask her a new question, she re-reads the entire case file from page one before answering. Not the last page. Not her notes. The whole file — every document, every previous question, every answer she already gave you.
Question one is cheap. The file is thin.
By question ten, the file has swollen with nine rounds of discussion, and she’s reading all of it, again, to answer one more small thing. You asked ten questions. You paid for something much closer to fifty.
That’s not a strange lawyer. That’s how every AI agent you run is billed.

You’re not paying per answer. You’re paying per re-read.
Here’s the single fact that explains almost every surprising agent invoice: language model APIs bill for the entire conversation history on every call.
A chatbot mostly gets away with this. You ask, it answers, the exchange is short, and the history grows slowly. One question, one bill.
An agent doesn’t work that way. An agent reasons, calls a tool, reads the result, reasons again, calls another tool — and every one of those steps is a fresh API call that re-sends everything that came before it. The system prompt. The tool definitions. The original question. Every previous thought. Every tool result, in full.
So when people say “agents are expensive,” they’re describing a symptom. The cause is more specific and more useful to know: you’re paying to re-read the case file, over and over, and the file gets fatter every step.
This is why your estimate was wrong. You multiplied a per-token price by what you thought one task costs. But a task isn’t one call — it’s a dozen calls, each one bigger than the last.
The shape nobody expects: it’s not a straight line
Most people intuitively model agent cost as linear. Ten steps should cost roughly ten times one step. Reasonable guess. Wrong shape.
Because each step carries everything before it, the cost of a naive agent loop grows closer to quadratically with the number of steps. Step 10 isn’t the same size as step 1 — it’s carrying nine steps of accumulated baggage. Double the steps and you don’t double the bill; you roughly quadruple it.
The published multipliers bear this out. Gartner estimates agentic workloads consume 5–30x the tokens of a single chat exchange. Independent measurements put a typical single agent at around 4x a chat interaction, and multi-agent systems at roughly 15x. The spread is wide because it depends entirely on how many steps your task takes and how much junk each step drags along.
The calculator below makes the shape visible. Put in your own numbers — your provider’s rates, your typical context size, your step count — and watch what happens as you push the steps up.
Your numbers, your provider. Change the step count and watch the shape change.
Total input tokens
—
Cost per task
—
If it were linear
—
You're paying
—
Cost accumulating, step by step
Play with the step count first. That’s the lever that surprises people.
Where the tokens actually hide
Knowing the cost is quadratic tells you that it grows. Knowing where the tokens hide tells you what to cut. Four places, and only one of them is obvious.
Tool definitions, paid for on every single call. This is the one that catches everybody. Your agent has five tools available. On a given task it only uses three. You might assume you pay for three — but the model needs to know about all five to choose between them, so all five tool schemas are sent on every call. Five schemas, every step, whether or not they’re used. Give an agent twenty tools “just in case” and you’ve bought a standing charge on every step of every task.
Retries, which re-send the whole file. A tool call fails or returns something unusable, so the agent tries again. That retry doesn’t just cost one extra tool call — it re-sends the entire accumulated context along with it. A three-retry pattern on a single database read roughly triples the token cost of that step. Flaky tools aren’t a reliability problem with a cost side effect; they’re a cost problem wearing a reliability costume.
Raw tool output, piling up forever. An API returns a 50 KB JSON blob. If you feed that straight back into the agent’s context, you haven’t paid for it once — you’ve paid for it on every subsequent step of the task, because it’s now part of the history being re-read. A tool returning 50 KB instead of 5 KB inflates cost tenfold every time that context is sent again.
The reasoning trail itself. The agent’s own intermediate thoughts accumulate alongside everything else. Useful for coherence. Not free.
Stanford’s Digital Economy Lab put a number on the aggregate: re-sent context accounts for roughly 62% of total agent inference bills. Most of what you’re paying for isn’t new thinking. It’s re-reading old thinking.
The counterintuitive part: it’s input, not output
Here’s where most optimization instincts go wrong.
When we think about making a model cheaper, we reach for output: shorter answers, tighter responses, less verbosity. That’s the lever we’re used to.
For agents, it’s largely the wrong lever. An agent doesn’t generate dramatically more text than a chatbot does — a tool call and a short reasoning step are usually a few hundred tokens. What it does is re-read a context window ten or a hundred times longer, on every single step.
The cost lives in the input. Which means the question isn’t “how do I make my agent say less?” It’s “how do I make my agent carry less?”
That single reframe changes what you optimize, and it’s why teams that tune output length see their bills barely move.
What actually brings the bill down
Six things, roughly in order of effort-to-payoff.
1. Summarize tool output before it enters context. The highest-leverage fix, and the most skipped. Never feed a raw API response back to the model. Put a deterministic parser or a cheap summarizer in between that extracts only the fields that matter. You’re not just saving tokens once — you’re saving them on every subsequent re-read of that context.
2. Cap the iterations. A hard ceiling on tool calls per task (three to five is a sane start) turns an unbounded cost into a bounded one. This is the same fix that prevents the runaway-loop failure mode — one limit, two problems solved.
3. Turn on prompt caching. Most providers now let you cache a stable conversation prefix so it isn’t re-billed at full price on every call. With Anthropic, cached tokens cost about 10% of the input price, though cache writes cost about 25% more than normal input. For an agent re-sending the same system prompt and tool definitions a dozen times a task, that trade is usually overwhelmingly in your favour.
4. Trim the tool list. Every tool you attach is a schema paid for on every call. Give each agent only the tools it genuinely needs for its job. This improves accuracy too — over-broad tool access is a known source of both higher cost and worse tool selection.
5. Right-size the model per step. Not every step deserves your most capable model. Routine classification, routing, and summarizing can run on a small fast model; save the expensive reasoning for the steps that actually need to reason.
6. Isolate context with sub-agents. When a task has genuinely separable parts, delegating to child agents keeps each context small. A parent orchestrating three children pays for three small contexts instead of one enormous one, because the parent never has to carry the children’s intermediate reasoning and tool chatter.
The honest limits
Every one of those six has a case where it backfires, and you should know them before you go optimizing.
Caching isn’t free money. It helps when a large prefix is genuinely stable across calls. If your context churns every step, you’ll pay the write premium repeatedly and save little. And there are minimum-size thresholds below which caching simply doesn’t engage.
Sub-agents can cost more. This is the trap. Splitting a task into child agents adds coordination overhead, and multi-agent systems measure around 15x a chat baseline versus roughly 4x for a single agent. Context isolation only pays off when the sub-tasks are genuinely separable. Distribution should be earned, not assumed.
The cheapest model can be the most expensive. A weak model that fails, times out, or produces output that triggers retries can cost more in total than a stronger model that completes the task once. Per-token price is not cost. Cost is price times how many times you had to do it.
Summarizing has a quality floor. Compress tool output too aggressively and you starve the agent of what it needed, which produces… retries. Which re-send the full context. You can optimize yourself straight into a bigger bill.
Budget is a design decision
The uncomfortable thing about agent costs is that they aren’t really a billing problem. By the time the invoice arrives, every decision that caused it was made months earlier — how many tools you attached, whether you summarized tool output, whether you capped the loop, whether you reached for the big model by default.
A chatbot’s cost is roughly a function of usage. An agent’s cost is a function of architecture.
Which means the useful question isn’t “why is this bill so high?” It’s the one worth asking before you build:
How much is this task allowed to cost — and what in my design enforces that?
Set the budget first. Let the architecture answer to it.