Skip to content

Model Selection & Cost Engineering

At a glance Which model to assign to which step is a design decision that directly shapes both capability and the bill — the capability/price/latency three-way trade-off across flagship and small models, three routing modes (rule-based, cascade, and learned routers), fallback and retry discipline, cost knobs such as prompt caching and batch processing, and the eval discipline that every model swap must run regressions.

Model Selection & Cost Engineering ​

Earlier chapters keep coming back to one formula: Agent = Model + Harness. But the "Model" slot has never been singular — as of mid-2026, every major vendor's lineup splits into at least three tiers (flagship, mid-range, lightweight) with prices more than an order of magnitude apart. Which step goes to which tier is a decision the harness engineer makes; the model will not make it for you. The decision lands on three axes at once: capability (what does it cost when this step goes wrong), cost (how many times does this step run), and latency (does this step stand between the user and their goal).

This is exactly the question model routing answers. Together with context engineering, it is one of the only two parts of a harness that spend money on every turn: context decides how big each call is; routing decides how expensive it is.

The economics of model tiers ​

Start with the bill. As of mid-2026, the pricing structure of the major APIs looks roughly like this (the specific numbers churn fast with every release; treat the official pricing pages as the source of truth — Anthropic, OpenAI, Google):

TierInput price range (per million tokens)Output price rangeTypical uses
Flagship (Opus / GPT flagship tier)~$5–15~$25–75Complex planning, hard debugging, final review
Mid-range (Sonnet tier)~$2–3~$10–15The workhorse for everyday coding and writing
Lightweight (Haiku / mini / Flash tier)~$0.1–1~$0.8–5Classification, summarization, format conversion, monitoring

Three observations worth keeping:

The price spread is real money. Flagship and lightweight tiers are commonly 10–30x apart on unit price. Agent workloads are characterized by extremely long inputs and extremely short outputs (the context engineering chapter cites Manus's data: an average input-to-output ratio of about 100:1), so the input price is the bulk of an agent bill — routing decisions should be made on the input column.

Capability gaps are not uniform. On tasks with a fixed answer shape — classification, extraction, rewriting — small models trail the flagship by very little; the gap concentrates in long-chain reasoning, ambiguity resolution, and recovery from rare errors, the places without a map. All of routing's margin hides in that unevenness.

Latency is tiered too, and it runs opposite to price. Small models deliver the first token sooner and generate faster. For interactive steps that stand between the user and their goal, "fast and good enough" often beats "strong but slow" — latency itself belongs in routing as an input dimension, as we'll expand on later.

Different models for different steps of the same task ​

A single agent task is not a homogeneous stream of calls; it's a mix of steps at different difficulty levels. A sensible default tiering:

text
┌────────────────────── model tiers across one agent task ─────────────────────┐
│                                                                              │
│  Planning / re-planning ────────────► Flagship        (few calls; can't fail)  │
│  Key decision points (ambiguity, final review) ──► Flagship                  │
│  Routine execution steps ────────► Mid-range  (most frequent; bulk of bill)  │
│  Tool-result classification / summary ──► Lightweight  (fixed answer shape)  │
│  Sub-agent narrow tasks ──► Lightweight/Mid-range (see /components/subagents)│
│  Context compaction ────► Specialized / Lightweight (high-frequency, mechanical) │
│                                                                              │
│  Principle: pay flagship prices per decision; cheap models per throughput    │
└──────────────────────────────────────────────────────────────────────────────┘

This tiering has plenty of industry precedent. Planning and task decomposition covered a hidden dividend of Plan-and-Execute: the expensive model as planner, the cheap one as executor. Cognition goes further — as shown in the Devin case study, they trained a dedicated compression model for compaction, because a long-running agent can trigger compaction several times an hour: a textbook "high-frequency, mechanical, fixed answer shape" step that flagship rates shouldn't be paying for.

A counterintuitive lesson

The biggest resistance to tiering isn't technical, it's psychological: engineers worry the small model will be "a little bit worse". The right question isn't "is the small model good enough" but "what does it cost when this step fails, and can a later step catch it?" Errors a downstream step or validator can catch: let a cheap model take the shot. Decision points where one wrong move loses the whole run: those are worth flagship prices.

Three routing modes ​

Mechanisms for deciding "who gets this step", in three tiers of sophistication.

Mode 1: Rule-based routing ​

The simplest and most common: hard-code it in the harness — "flagship for planning, mid-range for execution, lightweight for classification" — or dispatch on static task features (issues tagged bug go to model A, issues tagged refactor go to model B).

  • Pros: zero added latency; behavior is fully predictable and auditable; when something breaks you can attribute it to an exact rule.
  • Cons: rules are gut-feel priors that never adjust to the real distribution; task difficulty is a continuum, and a rule's boundary will inevitably slice part of it wrong.

Mode 2: Cascade — start small, escalate when needed ​

A cascade swaps "rules" for "confidence": every request goes to the cheapest model first, and a scorer judges whether the output is trustworthy. Trustworthy, use it; not, escalate to the next tier.

text
┌───────────────────── LLM cascade ───────────────────┐
│                                                     │
│  Request ──► Lightweight ──► Scorer ──trust──► return it   │
│                   │            │                       │
│                   │            └─distrust──► Mid-range ──► │
│                   │                    score / escalate     │
│                   ▼                                        │
│       A few hard cases eventually reach the flagship model │
│                                                     │
│  Expected cost = Σ (probability × tier unit price)         │
└────────────────────────────────────────────────────┘

The theoretical foundation of this line is Stanford's FrugalGPT (arXiv:2305.05176, published in TMLR 2024): on narrow classification tasks such as HEADLINES, a cascade matched GPT-4 accuracy while cutting cost by up to 98%, or beat GPT-4 by about 4 points at the same cost. Cite those numbers with the mandatory footnote: they come from narrow classification tasks on 2023-era model generations, where the scorer is cheap and easy to train; the "98%" does not extrapolate to open-ended tasks like agentic coding. The real gate for a cascade is the scorer — judging "is this output trustworthy" is itself a hard problem, and on open-ended generation a reliable confidence estimator can cost more than the money it saves.

Mode 3: Learned router ​

One step further: no hand-written rules and no step-by-step escalation — instead, train a dedicated routing model that predicts, in one shot as the request arrives, whether the weak model would produce a satisfactory answer for this query, and dispatch directly. LMSYS's RouteLLM (arXiv:2406.18665, open-source implementation) is the representative work: trained on human preference data, it reported over 85% cost reduction on MT-Bench with essentially no quality loss, and the router shows some transfer ability when the strong/weak pair is swapped at test time.

  • Pros: a single prediction, none of the cascade's serial latency; routing boundaries learned from data, not gut feel.
  • Cons: needs training data (preference pairs, quality labels over historical trajectories); every new model generation means retraining; and the router is itself a component that has to be evaluated and maintained.
DimensionRule-basedCascadeLearned router
Decision timingStatic (code/config)Per request, small firstPer request, single dispatch
Added latencyZeroHigh (serial, multi-stage)Low (one forward pass)
Required assetsThe engineer's priorA reliable scorerTraining data + a retraining pipeline
Failure modeBoundary cuts the wrong casesScorer misses or false-positivesBreaks under distribution drift
Best fitHarnesses with fixed step typesServices with high query-difficulty varianceHigh-traffic products with trajectory data

Industry reference: Devin Fusion's dynamic switching ​

As of mid-2026, the most candid piece of public engineering literature on model routing for agentic workloads is Cognition's Devin Fusion (details in the Devin case study). It folds the three modes above into one production-grade scheme, and two of its design choices are worth stealing:

  1. Don't dispatch at task start; switch dynamically mid-session. A lightweight classifier keeps judging whether the current phase should escalate or downgrade — because an agent task's difficulty unfolds along the timeline, the task label at step zero can't predict how hard step 40 will be.
  2. Schedule the switch at compaction time. Switching models invalidates the entire KV cache prefix (the mechanism is covered in the KV cache section of context engineering; not repeated here), while compaction rebuilds context anyway — align the model switch with compaction and the cache cost of switching drops to zero. A superb demonstration that "routing decisions must understand the inference infrastructure."

The result: Fusion held frontier-level performance at roughly 35% lower cost (later figures put it at up to 60%). Cognition is also blunt about the lesson: general-purpose routing schemes that shine on benchmarks produce "code you wouldn't merge" — agentic routing must be evaluated on real task trajectories; routing gains on static Q&A benchmarks say nothing about gains in agent scenarios.

Fallback and retry design ​

The other side of routing is failure handling. Model APIs are not stable infrastructure: rate limits, overload, timeouts, content refusals, over-long context — each demands a different response. Writing fallback as a bare try: call() except: retry() is one of the most common places an agent bill bleeds out.

Triage by error type:

ErrorNatureCorrect response
Rate limit (429) / overloadTransientExponential backoff with jittered retries; or downgrade to a mid-range model / backup provider
TimeoutTransient, but you may already be billedRetry a limited number of times; note that the timeout for long-output calls should be measured against the first streamed token, not total duration
Context over the limitDeterministic; retrying is uselessDon't retry — compact first or drop to a long-context tier, then resend
Content refusalDeterministicRetrying is futile; rewrite the input or switch models, and log it as an observable event
Malformed outputModel-sideA verbatim retry occasionally works; better to feed the parse failure and error message back to the model and let it self-correct

Three disciplines:

  • Fallback chains need a terminus. Flagship → mid-range → backup provider → a clear error, three or four hops at most. Unlimited downgrading quietly turns "a hard task failed" into "a hard task was answered wrongly by a lightweight model" — far more dangerous, because it doesn't log an error; it just produces a wrong result.
  • Budget your retries. Set a retry count and a total-cost ceiling per task; when the ceiling trips, follow the escalation path in permissions and human-in-the-loop and ask a human. A retry loop with no cost ceiling can burn straight through the bill in a rate-limit storm.
  • Every switch is an invalidation point. Any fallback that swaps model or provider changes system behavior (tool-call style and formatting habits differ across models); the switch event must go into the trajectory log — it's one of the most important anomaly signals in observability.

A hidden billing black hole

Stack a cascade on top of fallback and you get "double billing": the lightweight model answers wrong → escalate to mid-range → timeout → retry → escalate again to flagship. One user request, billed by four models. Design the escalation to remember prior attempts (pass the lightweight model's output along as context for the next tier — it saves tokens and hands the next model clues), and set a task-level cost breaker.

The full landscape of cost knobs ​

Routing is the most visible cost lever, but not the only one. The complete knob list, ordered by leverage:

KnobMagnitudeKey points
Model routingAn order of magnitudeThis chapter's topic; the bill is mostly input tokens, so keep your eye on the input column
Prompt cachingCached portion of input price drops to ~1/10Mechanics and discipline (stable prefix, append-only) in context engineering; the routing-side lesson is don't switch models casually — a model switch wipes the cache entirely
Batch APIs50% off across the boardOpenAI's Batch API and Anthropic Message Batches both give 50% off with 24-hour turnaround; workloads that don't need a real-time response — eval suites, overnight batch runs, offline data labeling — should always go through batching
Context slimmingLinearEvery token saved saves input money; compaction, tool-result cleanup, and just-in-time retrieval all live here
max_tokens disciplinePrevents runawayCap each step's output so the model can't lapse into a long monologue; compress it to the minimum for "return JSON only" steps
Structured outputIndirectStrict schema constraints cut retries, saving fallback money

How the knobs interact

These knobs are not independent. Cache discipline constrains routing (model-switch timing should align with compaction points); batching stacks with routing (batch jobs should lean even harder on cheap models, since with no real-time pressure they can absorb a cascade's serial latency); max_tokens discipline shapes fallback design (truncated output gets caught by format validation, triggering retries). Treat cost engineering as a system of interlocking parts, not a checklist to tick off item by item.

Latency and experience trade-offs ​

Beyond cost, the second axis of routing is latency. An agent's latency experience is decided by three things:

Streaming is the cheapest experience optimization there is. It changes no cost structure at all; it just turns "wait 30 seconds for the full answer" into "see the first token within 1 second". The cost to the harness is parsing logic that gets more complex — tool-call arguments have to be parsed as they stream in, and format errors only surface at the end of the stream.

Speculative execution. At two levels. On the inference side, speculative decoding (Leviathan et al. 2023) has a small model draft and a large model verify in parallel, multiplying the large model's effective generation speed — a provider-side optimization, but it explains why "large model + small draft model" serving architectures can be faster without a meaningful price premium. The application-side counterpart is OpenAI's Predicted Outputs: when most of the output is predictable (say, a code rewrite that touches a few lines), pass the predicted content in as a reference and the model only needs to "confirm and patch"; official figures report latency reductions of several-fold in these scenarios. At bottom, this is the harness cashing in its own prior — I know roughly what the output will look like — for a latency win.

Fast models up front. For latency-sensitive interactive steps (intent recognition, drafting the next action), a small model gives instant feedback first, and the slow, heavy steps move to the background — which is naturally isomorphic to the cost tiering, so the routing table should annotate each tier's expected TTFT and throughput.

The core of the trade-off: latency and cost can often be traded on the same knob, but turning it in opposite directions cancels out. Cascades save money but carry the highest serial latency; all-flagship is the fastest and the most expensive; Fusion-style "sidekick parallelism" tries to have it both ways — the main agent and a cheap agent run side by side, trading compute redundancy for the hope of "fast and cheap". No free lunch, only trade-offs with clear price tags.

Evals first: every model swap must run regressions ​

Last, and the easiest to skip: any change to model routing — swapping a model, editing routing rules, tuning cascade thresholds — is a harness change and must run regression evals.

The reason is hard-nosed. An agent's output is the joint product of model and harness (research cited in Model vs Harness shows that swapping the harness alone moves the score by nearly 10 points on the same model), and a model swap's effect on the harness is asymmetric: the new model may be smarter, but its tool-call formatting habits differ, its robustness to post-compaction summaries differs, and its failure modes inside the fallback chain differ. The prompt discipline you tuned for the old model may be a patch for that model, not a universal best practice.

The minimum engineering bar:

  • Run a frozen regression set on every routing change (task set, graders, and cost accounting all locked down), and compare two metrics, pass rate and per-task cost — quality alone misses cost regressions, cost alone misses quality regressions;
  • Pin model versions in config and disable latest aliases — silent vendor upgrades are the canonical source of "your harness didn't change but behavior did";
  • Log the routing decisions themselves: which step went to which tier, why, and what it cost — without that data, an anomalous bill leaves you with nowhere to start.

For methodology, see Evals in Practice; how to build eval sets and write graders is fully developed in that chapter.

Default advice for a self-built harness

Start with rule-based routing (flagship for planning, mid-range for execution, lightweight for chores) and get per-step cost logging in place first; only once you've banked a few hundred real trajectories do you have the data to judge whether cascades or a learned router are worth it. Training a router without trajectory data is overfitting your own prior. For the hands-on path, see Build Your Own Harness and Design Principles.

Further reading ​

References ​