Appearance
Cost and Performance Engineering
After an agent ships, the first thing to blow up is usually not quality — it's the bill. A decent chatbot costs a few thousand tokens per conversational turn; a coding agent that actually works can burn hundreds of thousands to over a million tokens on a single task — a gap not of tens of percent but of one to two orders of magnitude. Gartner's March 2026 analysis put the magnitude at: agentic tasks consume 5 to 30 times the tokens of an ordinary chatbot per task; measured data for multi-step tool-calling agents generally lands in the 10-100× range.
This page covers two things: where the money actually goes (token economics), and how to compress it without visibly sacrificing quality (routing, caching, context slimming, latency and rate-limit engineering). It is not a generic "cost-cutting checklist" — every section tries to give numbers and code you can copy directly.
Price timeliness
All unit prices in this page reflect public pricing verified as of August 2026 (Anthropic's official pricing page was checked item by item). LLM prices change every six months; verify against the vendor's site before citing.
1. Token Economics: Where an Agent Task's Money Goes
Why agents burn money "quadratically"
A chatbot's cost model is linear: each turn's input ≈ system prompt + history + the user message, and history grows slowly. An agent's cost model is approximately quadratic — that's the root cause of most blown budgets.
Every iteration of the agent loop (see Agent Loop) re-sends the complete trajectory so far to the model:
Turn 1 input: system prompt + tool schemas + the task
Turn 2 input: all of turn 1 + model output 1 + tool result 1
Turn 3 input: all of turn 2 + model output 2 + tool result 2
...
Turn N input: the sum of everything from all previous turnsIf each turn adds Δ tokens (model thinking + tool returns), the total input tokens for an N-turn task are roughly N × initial context + Δ × N(N+1)/2. The trajectory-accumulation term is O(N²). Augment Code's analysis illustrated this with a concrete example: a 20-step loop consumes more than 10× the tokens of the naive linear estimate of "per-step estimate × 20."
A worked estimate
Take a typical coding agent task (fixing one bug):
- System prompt + tool schemas: 3,000 tokens (tool schemas aren't cheap themselves — Anthropic's official docs disclose that tool calling alone adds 286-804 tokens of built-in system prompt, plus a few hundred tokens per tool definition)
- Per-turn model output (thinking + tool calls): ~500 tokens
- Per-turn tool returns (file reads, test output): ~2,000 tokens
- 20 turns total
Total input ≈ 20 × 3,000 + (500 + 2,000) × 20×21/2 = 60,000 + 525,000 = 585,000 tokens; total output ≈ 10,000 tokens.
At Claude Sonnet 4.5 pricing ($3 / MTok input, $15 / MTok output):
- Input: 585k × $3/M = $1.76
- Output: 10k × $15/M = $0.15
- Total per task ≈ $1.9
For comparison: an ordinary chatbot Q&A (2,000 input + 500 output) costs about $0.014. One bug-fix task is 130× more expensive — and it fixed just "one" bug. On Opus 4.5 ($5/$25), the input portion is $2.9, totaling about $3.2.
Output tokens are the unit-price killer
The output/input price ratio is typically 5× (Sonnet 4.5: $3 vs $15). Long thinking chains, rewriting whole files verbatim instead of emitting diffs, and verbose final summaries all burn money at the 5× rate. The ROI of optimizing output tokens is often higher than optimizing input.
The cost structure, decomposed
Breaking the ledger above apart, an agent task's token consumption has roughly four components:
| Component | Share (typical) | Character | Main optimization |
|---|---|---|---|
| System prompt + tool schemas | A fixed 3-10k per turn | Sent identically every time | prompt caching (nearly free) |
| Trajectory accumulation (history) | 60-85% of the total | Grows as N² | Context slimming, sub-agent isolation |
| Tool return content | The main source of increments | Highly redundant (logs, full files) | Trimming and masking tool results |
| Model output | Often over 30% of total cost | 5× unit price | Cap output length; diff-style edits |
The conclusion first: the main battleground of agent cost optimization is "trajectory accumulation" and "tool returns," not the system prompt. Plenty of people heroically trim a few hundred tokens out of the system prompt while letting the full output of one npm test (8,000 tokens) enter the context untouched, billed repeatedly every turn.
2. Model Routing and Tiering: Don't Let Opus Do Haiku's Job
Tiering is the default architecture of agent systems
One model to rule them all is the most common waste. The mature approach is tiering by task difficulty:
┌─────────────────┐
planning / hard │ flagship model │ Opus 4.5 ($5/$25)
reasoning → │ (Planner) │ GPT-5 class
└─────────────────┘
┌─────────────────┐
execution / │ mid-tier model │ Sonnet 4.5 ($3/$15)
tool calls → │ (Executor) │ the price-performance sweet spot
└─────────────────┘
┌─────────────────┐
classification / │ small model │ Haiku 4.5 ($1/$5)
routing → │ (Classifier) │ GPT-4.1 nano (~$0.10 class)
summarization └─────────────────┘The price gaps are real: roughly 1.7× between flagship and mid-tier on input, 3× between mid-tier and small, and 50× between flagship and nano-class. In a multi-agent system (see Multi-Agent Architecture), the planner is called a handful of times per task while executors and classifiers are called dozens to hundreds of times — pushing the high-frequency calls onto small models transforms the bill's structure immediately.
Learned routing: RouteLLM and FrugalGPT
Beyond static tiering ("this class of task always uses model X"), there's learned dynamic routing — a trained classifier predicts "does this query need a strong model, or is a weak one enough":
- RouteLLM (LMSYS / UC Berkeley, arXiv 2406.18665, ICLR 2025): trains four routers on Chatbot Arena preference data (similarity weighting, matrix factorization, BERT, causal LLM). The paper reports over 85% cost reduction on MT-Bench while retaining about 95% of GPT-4-level quality; with data augmentation, the matrix-factorization router sends only 14% of calls to the strong model.
- FrugalGPT (Stanford, 2023): LLM cascades — start with the cheapest model, escalate tier by tier when quality checks fail; reports up to 98% cost reduction on benchmarks like HEADLINES.
Stay skeptical of 85%-98% headline numbers
Those headline numbers are dataset-specific. The honest conclusion from multiple independent replications in 2026 (on mixed real production traffic): routers deliver roughly 20-35% cost savings in practice, with quality loss held under 2%, and some commercial routers fail to beat the baseline on mixed traffic. Treat RouteLLM's paper numbers as a ceiling, not your expected return. Routing's real value is inside agent systems — structural tiering that separates planning from execution delivers far more solid gains than "bolt a classifier onto chatbot traffic."
The quality trade-offs of routing strategies
- Tier first, route second: structural tiering (planning/execution/classification on different models) is an architecture decision — no training needed, predictable gains. Dynamic routing is an optimization that needs eval data to back it.
- The routing decision itself must be cheap: using a $0.10/M classifier to decide between a $3/M and $5/M model is a good trade; using a $1/M model to route may not be.
- Misrouting costs are asymmetric: sending a hard task to a small model costs a failed task and an entire wasted trajectory re-run — far more than the savings. Routers should be biased toward escalation (when unsure, use the strong model).
- Always back routing with evals: before a routing strategy ships, run an offline evaluation on your own task distribution (methodology in Evaluation Systems); confirm the quality loss is acceptable before talking savings.
A minimal usable Python routing example (current APIs as of 2026):
python
from anthropic import Anthropic
client = Anthropic()
# Tiered model table: pick models by role, not by mood
MODELS = {
"planner": "claude-opus-4-5", # planning: flagship
"executor": "claude-sonnet-4-5", # execution: the sweet spot
"classifier": "claude-haiku-4-5", # classification/summarization: small model
}
def call(role: str, messages: list, **kwargs):
return client.messages.create(
model=MODELS[role],
max_tokens=kwargs.pop("max_tokens", 4096),
messages=messages,
**kwargs,
)
# Example: use the small model to judge task difficulty; only hard tasks get escalated to the flagship
def route_task(task: str) -> str:
resp = call("classifier", [{
"role": "user",
"content": f"Judge this task's difficulty; answer only 'simple' or 'hard': {task}"
}], max_tokens=10)
return "planner" if resp.content[0].text.strip() == "hard" else "executor"3. Caching: the Agent System's Free Lunch (or at Least the First One)
Prompt caching: the mechanism and the real pricing
Prompt caching is the highest-ROI optimization in the agent scenario, full stop. The mechanism: the attention states of an already-processed prompt prefix are cached, and when a later request's prefix hits the cache, it's billed at a discount — and the agent loop's "system prompt + trajectory history," re-sent every turn, is naturally a prefix structure.
The main vendors' mechanisms, verified August 2026:
| Provider | How to enable | TTL | Cache write | Cache read |
|---|---|---|---|---|
| Anthropic | explicit cache_control breakpoints (or top-level automatic caching) | 5 min / 1 hour | 1.25x (5 min) / 2x (1 hour) of base input price | 0.1x (90% off) |
| OpenAI | automatic, prefix ≥ 1024 tokens | ~5-10 min | no write fee | 0.5x (half price) |
| Google Gemini | explicit context caching API | 1 min - 24 hours, configurable | billed by storage duration | ~0.25x |
| DeepSeek and others | usually automatic | varies | — | usually 0.1-0.2x |
Anthropic's math is easy: 5-minute cache writes cost 1.25× and reads 0.1× — a single hit pays for itself; the 1-hour cache writes at 2× pays back on the second hit. The agent loop appends on top of the previous turn each round, so the prefix almost certainly hits — caching isn't an "optimization" for agents, it's "must be switched on."
Back to Section 1's estimate: of 585k input tokens, assume that after the first turn's write, 80% of subsequent input hits at 0.1×:
- Input cost without cache: $1.76
- With cache: writes ≈ 585k×20%×$3×1.25/M ≈ $0.44; reads ≈ 585k×80%×$0.3/M ≈ $0.14; total ≈ $0.58, a 67% saving
In code, you only need one breakpoint at the end of the stable prefix:
python
resp = client.messages.create(
model="claude-sonnet-4-5",
max_tokens=4096,
system=[
{
"type": "text",
"text": LONG_SYSTEM_PROMPT, # the system prompt goes first — it's the most stable
"cache_control": {"type": "ephemeral"}, # cache breakpoint
}
],
tools=TOOLS, # tool schemas are also in the prefix, cached along with the system prompt
messages=history, # the trajectory appends turn by turn; the prefix structure hits naturally
)The key engineering constraints:
- Prefix order determines hit rate: the system prompt, tool definitions, and stable documents must come first and stay byte-for-byte identical; timestamps, random IDs, and anything that changes every turn must never appear in the prefix.
- TTL and heartbeats: Anthropic's 5-minute cache expires mid-task when an agent pauses (say, waiting for user confirmation); the 1-hour cache (2× write) suits this human-in-the-loop pattern, paying back on the second hit.
- Monitor the cache metrics: the
cache_read_input_tokens/cache_creation_input_tokensfields in the response tell you the hit rate directly; a dropping hit rate usually means someone put dynamic content into the system prompt.
KV cache and semantic caching
- KV cache: at the inference-engine layer (vLLM, SGLang, and other self-hosted setups), the attention KV of a prefix gets reused. You can't touch it through an API, but when self-hosting open models, prefix caching / RadixAttention can cut the first-token latency and compute of multi-agent systems sharing a system prompt by more than half — the single most important throughput optimization in self-hosted deployments.
- Semantic caching: query-level caching — embed historical queries, and if a new query's similarity to a past one exceeds a threshold, return the cached answer. Works well for FAQ-style agents; nearly useless for coding/research agents (queries almost never repeat), so don't misuse it. Note the "wrong answer to a similar question" precision risk — the threshold needs to be tuned with evals.
4. Context Slimming Is Cost Slimming
Every technique of context engineering (see Context Engineering) is simultaneously a cost optimization — tokens are money. Ordered by ROI:
- Trim tool results: the single largest source of savings. A
grepreturning 200 lines keeps only ±5 lines around the hits; test output keeps only the failure summary; file reads use line-number windows instead of whole files. One principle: what a tool returns to the model should be "the conclusion," not "the raw data." - Trajectory compaction: as you approach the context window limit, use a small model to compress the early trajectory into a structured summary (keeping file paths, error history, decisions made, todos) and discard the original text. Claude Code's auto-compact is exactly this mechanism (see Claude Code Anatomy).
- Observation masking: keep the tool-call records but replace old tool returns with placeholders (
[output omitted]); the model knows "what was looked up here" without paying tokens for the old output. - Sub-agent isolation: delegate the dirty work — searching, exploring — to sub-agents; the lead agent receives only final conclusions. The sub-agent's dozens of trajectory turns never enter the main context — the most underrated cost benefit of multi-agent architecture.
- Diff-style file editing: have the model emit edit blocks instead of rewriting whole files — saves output tokens (at the 5× rate) and saves next-turn input tokens from re-reading the file.
Slimming has a quality cost — measure it
Trimming and compression both lose information, and losing the wrong bits means task failure. The right posture: treat the context strategy as a hyperparameter and put it into your eval-set regression tests (see Evals in Practice). "40% fewer tokens, success rate down 8 points" is usually a bad trade; "40% fewer tokens, success rate unchanged" is the one you want.
5. Latency Engineering: Users Are Waiting for More Than Money
Beyond cost, the second engineering bottleneck is latency — 20 serial LLM calls at 5-15 seconds each means the user waits several minutes. Techniques, ordered by implementation difficulty:
TTFT and streaming
- TTFT (time to first token) is the decisive metric for perceived latency. Ways to cut it: shorten the input prefix (context slimming again), enable prompt caching (cache hits skip prefill — Anthropic officially claims up to 85% TTFT reduction for long prompts), use a smaller model.
- Streaming costs almost no engineering effort and delivers the biggest perceived win: push the model's thinking and intermediate conclusions to the user in real time, and a 30-second task feels like 5 seconds. Streaming is a one-parameter affair in every mainstream SDK — there's no excuse not to enable it.
Parallel tool calls
The agent loop is serial by default: think → call one tool → wait → think again. But many tool calls have no dependency on each other (reading three files, searching two keywords) and can fire in parallel within one turn. Mainstream models support returning multiple tool_use blocks in one response, which the client executes in parallel and feeds back together:
python
import concurrent.futures
# when the model returns multiple tool_use blocks in one turn, run them in parallel
# instead of a serial for loop
def execute_tool_calls(tool_uses: list) -> list:
with concurrent.futures.ThreadPoolExecutor(max_workers=8) as pool:
futures = {pool.submit(run_tool, tu.name, tu.input): tu for tu in tool_uses}
results = []
for future in concurrent.futures.as_completed(futures):
tu = futures[future]
results.append({
"type": "tool_result",
"tool_use_id": tu.id,
"content": future.result(),
})
return resultsFor I/O-bound tools (HTTP, databases, files), a thread pool is enough; the parallelism gains during I/O waits are typically 3-5×.
Speculative execution
Two levels:
- Inference layer: speculative decoding (a small model drafts, the large model verifies) is an acceleration technique for self-hosted inference engines — API users can't use it, but it's a fair question when picking an inference provider.
- Agent layer: the more practical idea is speculative prefetching — will the model probably read a certain file next? Read it while waiting for the model's response. And speculative branching: for slow operations, try the two most likely branches in parallel up front and discard the loser once the model picks. Spending a few tokens to buy latency is often a good trade in interactive scenarios.
Latency and cost are often the same problem
Fewer turns, smaller context, parallel tools — these cut cost and latency simultaneously. The only genuine conflict is routing to a smaller model (cheaper but possibly slower and more failure-prone) versus long thinking chains (more expensive, slower, more accurate). Engineering-wise, be explicit about which objective function your product is optimizing.
6. Concurrency and Rate Limiting: Designing Throughput Under Rate Limits
Production APIs are not infinite pipes. Anthropic limits RPM (requests per minute) and TPM (tokens per minute) by usage tier (Build/Scale, etc.); OpenAI is similar. And agent systems are especially hungry for rate limit — one task is 20 turns, so 100 concurrent tasks mean 2,000 requests and hundreds of millions of tokens in a short window. Design essentials:
- Client-side token bucket + queue: don't fire raw requests and let the API reject you. Implement a token bucket at the gateway layer (meter by TPM, not RPM — agent requests vary too widely in size; rate limiting by request count goes out of control), and queue the overflow. The queue must carry task priority: interactive user requests first; batch jobs yield.
- Exponential backoff + jitter: 429s and 529s (overloaded) are the norm; you must retry. Standard practice: exponential backoff (1s, 2s, 4s... capped at 60s) with random jitter; after 5-8 retries, give up and mark the task resumable. Agent retries are safer than chatbot retries — as long as tool calls are idempotent, you just continue from the breakpoint.
- Checkpoint recovery: being rate-limited mid-task doesn't mean failure. Persist the trajectory (checkpoint) and resume from the last turn instead of re-running the whole task — a re-run means the first 500k tokens were burned for nothing. This is also a point reliability design keeps hammering.
- Batch work goes through the Batch API: non-real-time tasks (evaluations, data labeling, overnight batches) should use Anthropic / OpenAI's Batch API — a flat 50% discount, and it doesn't consume real-time rate-limit quota. The most overlooked half-price channel.
- Multi-provider fallback: when the primary provider rate-limits, degrade to a backup provider or a smaller model. The routing layer (Section 2) is the natural home for this logic.
7. Cost Observability and Budget Alerts
Cost you can't measure, you can't optimize. The minimum bar: attribute every LLM call's tokens and cost to the trace level (one task = one trace, each turn = one span), then aggregate by task type, user, and model. This shares instrumentation with observability; only the aggregation dimension is money.
Implementation options (the state of things in 2026):
- Langfuse (open source, self-hostable): automatically records token usage per call and converts it to cost via per-model price lists, with aggregation by trace/user/session; the Monitors feature (still beta as of mid-2026) supports threshold alerts on cost, latency, and scores, routable to Slack/webhooks.
- LangSmith: tracing + cost tracking inside the LangChain ecosystem; managed and worry-free.
- LiteLLM Proxy: as a unified gateway it records the cost of every call automatically and supports per-API-key budget caps (max budget) — the most direct "hard gate."
- DIY: multiply the
usagefields in responses (input/output/cache_read/cache_creation token counts) by your price list and store — a few dozen lines of code.
The minimal budget-alert practice:
- Three gates: an hourly cost-spike alert (say, 3× the same hour's 7-day average), a daily soft budget cap (alert), and a monthly hard cap (the gateway rejects outright or degrades to a small model).
- Attribute by task type: discovering that "80% of cost comes from 5% of task types" is the norm — usually a runaway retry loop or one tool returning giant outputs. Without trace-level attribution you'll never find it.
- Cap the cost of a single task: build
max_cost_per_taskinto the agent loop; when cumulative cost exceeds the limit, force compaction + a small-model finish, or simply terminate and return partial results. Stories of runaway loops burning hundreds of dollars show up in the community every month.
8. Case Study: Cost Before and After Optimization on a Real Task
Scenario: a code-review agent doing a full review of one PR. The numbers below are estimates, with these assumptions (re-measure against your own distribution before citing):
- Task size: average 25 agent-loop turns; 600 tokens of model output per turn; tool returns average 3,000 tokens (diffs, file fragments, lint output)
- Model: before optimization, Claude Opus 4.5 ($5/$25) throughout; no caching; tool results untrimmed
- After optimization: planning on Opus 4.5, execution on Sonnet 4.5 ($3/$15), summarization/classification on Haiku 4.5 ($1/$5); prompt caching on; tool output trimmed 60%; overnight bulk reviews processed via Batch
Before optimization:
- Input ≈ 25×5,000 (system+schemas) + 3,600×25×26/2 ≈ 125k + 1,170k ≈ 1.3M tokens, all at Opus input price: $6.5
- Output ≈ 15k × $25/M = $0.38
- ≈ $6.9 per task; at 200 PRs a day = $1,380/day
After optimization (each row maps to a section above):
| Measure | Effect | Saving |
|---|---|---|
| Trim tool results 60% (Section 4) | Per-turn increment 3,600 → 1,440; total input down to ~550k | -58% input volume |
| Prompt caching (Section 3) | ~75% of remaining input hits at 0.1× | input cost down another ~65% |
| Model tiering (Section 2) | Only 3 planning turns on Opus, 22 execution turns on Sonnet, 2 summaries on Haiku | blended input unit price down ~35% |
| 50% of overnight tasks via Batch (Section 6) | That portion gets another 50% off | total down another 25% |
Back-of-envelope: after optimization, ≈ $0.6-0.8 per task — total cost down to about 10% of the original, with latency halved too (parallel tools + cache hits cutting TTFT). The quality cost: on the internal eval, review-issue recall dropped from 91% to 88%, which the team judged acceptable — this trade-off was measured, not guessed.
This case uses no black magic — it's the combination of the seven sections above. That's the current state of agent cost engineering: the techniques are all known; the gap is entirely in execution — did you trim, are you monitoring cache hit rates, does routing have an eval backstop.
References
- Anthropic official pricing page (docs.anthropic.com) — the first-hand source for every Claude unit price here, cache multipliers (write 1.25x/2x, read 0.1x), the Batch 50% discount, and tool token overhead; verified 2026-08.
- RouteLLM: Learning to Route LLMs with Preference Data (arXiv 2406.18665) — the LMSYS/Berkeley routing paper, accepted at ICLR 2025; the original source of the 85% cost-reduction claim.
- The honest guide to LLM routing — an independent re-check of the various routing solutions' headline numbers, finding real-world gains of roughly 20-25% on mixed traffic.
- Prompt caching complete guide 2026 — a comparison of Anthropic/OpenAI/Gemini cache mechanisms and TTLs.
- The Bill Arrives: How to Manage Agentic AI Costs at Scale (Cockroach Labs) — cites the Gartner 2026-03 analysis: agentic tasks consume 5-30× the tokens of chatbots.
- Why AI Agents Can Cost 50x More Than Expected (Larridin) — the analysis of quadratic token growth from trajectory accumulation, including Augment Code's 20-step loop example.
- Langfuse Token & Cost Tracking docs — the implementation reference for trace-level cost tracking and alerting.