Skip to content

Case Study: Cursor

At a glance An anatomy of the harness design behind an IDE-native agent — how Cursor supplies context for large codebases with Merkle tree indexing and embedding retrieval, feeds the model editor state only an IDE can provide, solves the engineering problem of getting edits right with a specially trained apply model, and how Tab completion and agent mode divide the work inside a single harness.

Case Study: Cursor ​

If Claude Code represents the terminal-native route to agent harness design — generic tools, explicit file operations, leaving as much of the intelligence as possible to the model — then Cursor represents another equally successful route: burying the harness deep inside the editor itself. It wasn't satisfied with "a model plus generic tools"; instead it built its own indexing pipeline, specialized models, and editing protocol around the IDE as the host environment.

Cursor is worth dissecting in its own right because its team has published an unusual amount of harness engineering detail — from Merkle tree synchronization for codebase indexing, to why they don't use diffs to edit code, to the division of labor between the Tab model and the agent model. These official blog posts and docs make one thing unmistakably clear: when the host environment shifts from terminal to IDE, the harness design space moves with it — but none of the problems go away.

The Design Space: IDE-Native vs. Terminal-Native ​

Start with the big-picture comparison. At the bottom, both products run the same species of agent loop (think → tool calling → observe → think again; see Agent Loop); the differences all live at the harness layer:

DimensionTerminal-native (Claude Code)IDE-native (Cursor)
Host environmentA shell; the model explores everything itselfA VS Code fork, with editor state close at hand
Context sourcesExplicit tool calls (Read/Grep/Glob)Implicit editor state + index retrieval + explicit tool calls
Interaction rhythmDelegate a task, then watch from the sidelinesHuman in the loop, with editing, completion, and agent work coexisting at three granularities
Applying editsThe model generates full new file content directly; exact string matchingA two-stage pipeline: planning model + specialized apply model
Large codebasesAutonomous agent exploration (grep/find/read)A pre-built embedding index, with retrieval feeding the model
Rollbackgit + user review of diffsCheckpoint snapshots, one-click restore

Nobody in this table is the "more advanced" side — these are two different bets. The terminal camp bets on the model's ability to explore, keeping the harness thin and generic; the IDE camp bets on the value of environmental signals, letting the harness grow thicker and more specialized. Cursor's official docs are disarmingly frank about what an agent is made of — "instructions (system prompts and rules), tools, and models" — and note that instructions and tools are "specifically optimized for each frontier model" (Agent docs). That's the classic harness vendor's confession: the model is replaceable; the harness is the product.

Why Cursor forked VS Code instead of building a plugin

The decision is itself a harness architecture decision. IDE plugins run inside the host's extension sandbox, limited by the plugin API in what editor state they can read and what UI rendering they can control; Tab's ghost text, cross-file jump portals, block-by-block diff approval, and the checkpoint timeline all require deep surgery on the editor itself. Forking means carrying the maintenance burden of an entire editor; the payoff is that the harness gains complete control over both the "context collection surface" and the "human interaction surface." It's the mirror image of Claude Code's choice of the terminal: one thickens the host, the other thins it to nearly zero.

A perspective worth remembering

The essential advantage of an IDE-native harness isn't "having a graphical interface" — it's that the editor is a process that continuously emits high-value structured signals: which file the user is looking at right now, where the cursor is parked, what was just changed, which files are flagged with errors. In a terminal, those signals either don't exist or have to be reconstructed by the model at the cost of tool calls. Cursor's entire harness design revolves around feeding those signals to the model at low cost.

Context Supply: Codebase Indexing Is an Infrastructure Problem ​

Faced with a hundred-thousand-file monorepo, a terminal-native agent gets by on "look while you work" — grepping keywords, reading directories, crawling along imports. That works, but every step burns turns. Cursor's choice was to turn the codebase into a searchable index ahead of time, an industrial-grade specimen of the "offline preprocessing" route in context engineering.

According to Cursor's official blog post Secure Codebase Indexing, the pipeline looks like this:

text
┌─────────────────── Client (local) ────────────────────┐
│                                                       │
│  Scan workspace ──> chunk by syntax                   │
│  (syntactic chunks)                                   │
│       │                                               │
│       ▼                                               │
│  SHA-256 per file ──> build Merkle tree               │
│  (directory hash = hash of child hashes)              │
│       │                                               │
│       ▼ upload hashes only, never code                │
└───────┼───────────────────────────────────────────────┘
        ▼
┌──────────────────── Server ───────────────────────────┐
│                                                       │
│  Compare client and server Merkle trees:              │
│    · matching root hash → nothing to do               │
│    · walk only divergent branches → pinpoint          │
│      changed files                                    │
│       │                                               │
│       ▼                                               │
│  Chunk changed files ──> embed ──> write to vector    │
│  store (unchanged chunks hit the content cache and    │
│  aren't recomputed)                                   │
│       │                                               │
│       ▼                                               │
│  At query time: embedding search finds relevant code  │
│  chunks and returns them to the client for decryption │
│  (obfuscated filenames, encrypted chunks, no          │
│  plaintext source left on the server)                 │
└───────────────────────────────────────────────────────┘

Several engineering details deserve unpacking:

The Merkle tree solves synchronization, not retrieval. For a workspace of fifty thousand files, filenames plus SHA-256 hashes alone come to roughly 3.2 MB — without a tree, every sync would have to haul that full manifest around. With a tree, a matching root hash skips the whole thing; only subtrees with divergent hashes need to be descended into, and change detection drops from O(everything) to O(number of changes × tree depth). This is twenty-year-old git technology, carried over wholesale to incremental updates of an embedding index.

Chunking and caching determine maintenance cost. When a file changes, it's re-chunked along its syntactic structure, and embeddings are cached per chunk — most edits touch only a few chunks, and untouched chunks hit the cache directly. The index is therefore "fast to update, cheap to maintain," which is exactly what lets it run as always-on infrastructure rather than a one-off batch job.

Team-level index reuse. Within the same organization, different clones of the same codebase are on average 92% similar (Cursor's own figures). When a new member joins, the client derives a simhash (a similarity hash) from the Merkle tree and uploads it; the server looks in the team's simhash vector store for an existing index above the threshold and copies it over as the starting index — compressing first-time indexing on the largest repos from "hours" down to "seconds."

Is it worth building? The official numbers say yes. Cursor's internal evaluations report that semantic retrieval improves answer accuracy by 12.5% on average, and that the resulting code changes are more likely to survive into the codebase. They name semantic retrieval "one of the biggest drivers of agent performance."

Indexing isn't free

The cost of this pipeline is on the privacy side: code has to be chunked, uploaded, and embedded server-side. Cursor's mitigations are filename obfuscation, chunk encryption, and client-side decryption (see its codebase indexing docs), plus support for .cursorignore. Architecturally, though, this remains a "code leaves your machine" design — precisely the trade-off that terminal-native agents sidestep by default through local tool calls and on-demand reads. Which route to take depends on your tolerance for code leaving the premises.

The index handles semantic retrieval; exact lookup runs on its own track. Cursor also ships an in-house Instant Grep search engine (officially claimed to beat ripgrep on large codebases), and the agent automatically switches to exact matching when it references a specific symbol. Semantic retrieval handles "discovery," exact matching handles "pinpointing," and the two channels complement each other. The agent can also dispatch an Explore subagent — one that runs a large volume of searches in parallel on a faster model inside a separate context window and brings back only the conclusions, keeping raw file contents from polluting the main context (this is the textbook use of subagents as a context-isolation mechanism).

On top of automatic retrieval sits one more layer: manual context supplied by the user. The @ symbol system lets users specify context explicitly — @Files pins down specific files, @Code references code symbols, @Docs injects third-party library documentation, @Web triggers a web search, @Git pulls commit history. This isn't redundancy bolted onto the retrieval system; it's a correction mechanism for it. Embedding retrieval is recall-oriented ("possibly relevant"), while users often know exactly where the answer is. Making the human the last link in the context supply chain is the pragmatic IDE-native answer to the reality that retrieval is always noisy — the terminal-native version of the same move is the user pasting paths and file contents straight into the conversation.

IDE-Exclusive Context: Editor State Is the Prompt ​

Indexing solves supply for "the whole codebase," but the IDE-native camp holds a second card: the editor process is itself a context-producing machine. The signals Cursor's agent and Tab can read include:

  • Open files and recent browsing history — the shape of the user's attention, a strong prior on "which code is relevant";
  • Cursor position and selection — the focal coordinates of the task, and where the product's name comes from;
  • Edit history — what the user just did, which directly hints at what comes next;
  • Diagnostics (linter errors) — type errors and undefined symbols already computed by the language server, structured feedback at zero cost;
  • Terminal history and command output — immediate results from builds and tests.

Notice the nature of this list: none of it is "queried" by the model through tool calls — it's all collected in passing by the harness. The language server is already running, the edit history is already being recorded — the harness merely serializes those in-process states into the prompt. The contrast with terminal-native agents is sharp: for Claude Code to learn whether type errors exist, the model has to run tsc and read the output; in Cursor, the language server computed the diagnostics long ago, and the cost is one serialization. This is the strongest argument for the "thick harness" route: the environment is already producing high-value signals for free, and not harvesting them is waste.

Diagnostics have a second layer of value: they are an immediate post-edit verification signal. The moment the agent finishes editing a file, the language server re-checks it nearly in step — newly introduced type errors can be fed straight back to the model, forming an "edit → feedback" micro-loop with latency in the hundreds of milliseconds, no dedicated build required. In harness terms, this turns verification from an "explicit tool call" into a "receipt the environment issues on its own," and the quality of the "observe" step in the agent loop rises accordingly.

Tab and Agent: Two Latency Budgets, One Goal ​

Cursor's most famous feature isn't the agent — it's Tab, the gray ghost text that predicts your next edit. From a harness perspective, the relationship between Tab and Agent is worth spelling out: they aren't two products; they are the same goal (officially, in-flow Next Action Prediction — "predicting the next action within the flow") deployed under two different latency budgets.

The facts (from Cursor's official blog post A New Tab Model):

  • Since March 2024, Tab has been driven by an in-house sparse language model, trained on billions of tokens for the specialized task of "predicting edits";
  • It produces over a billion characters of edits per day, with request volume up roughly 100x from the first-generation model — by Cursor's own account, "almost no LLM in the world generates more code than the Tab model";
  • The Fusion model, shipped with the 0.45.0 client in early 2025, improved on the original with over 25% higher line-by-line prediction accuracy on hard edits, more than 10x longer stretches of consecutive changes per suggestion, p50 server-side latency down from 475ms to 260ms, and context length up from 5,500 to 13,000 tokens;
  • Tab doesn't just complete text: it predicts edits (rewrites and deletions, not just insertions) and jumps — after you accept a suggestion, it predicts where you'll edit next and sends the cursor there, across files.

In another official blog post, More problems, Cursor described the end state of this mechanism: the tab-tab-tab sequence. Many code edits are a chain of low-entropy actions — change a function signature, and what must follow is fixing call sites one by one, updating type definitions, adjusting tests — each step tightly constrained by the last. When single-step prediction is accurate enough, you can keep pressing Tab and let the model "autoplay" the entire chain (the official demo completes a whole run of changes with 11 consecutive Tab presses). At that point the boundary between Tab and Agent genuinely blurs: the former is an agent that still needs a human keypress at every step; the latter is a Tab that does away with the keypress too. The gap between them isn't intelligence; it's trust calibration — how large a change the user is willing to accept without step-by-step confirmation.

This yields a clean two-tier structure:

TabAgent
Latency budgetHundreds of milliseconds (fail if it breaks flow)Seconds to minutes
Task granularityThe next edit actionA complete task
ModelIn-house small model (sparse, specialized)Frontier large model
ContextCursor neighborhood + edit history + diagnostics (13k tokens)Whole-repo retrieval + tool calls (hundreds of thousands of tokens)
Human in the loopEvery Tab is a micro-approvalCheckpoints + diff review

Model tiering is a standard harness weapon

The Tab/Agent split is the extreme version of the "model tiering" principle in Planning and Task Decomposition: assign subtasks with different latency, cost, and capability requirements to models of different sizes. The lesson also runs in reverse — don't expect one model to be simultaneously good at "predicting the next cursor position within 260ms" and "refactoring a module autonomously." Those are two different inference workloads, and the harness's job is to split them apart and give each its own engine.

Making the Model Edit Code Correctly: Apply Is an Engineering Problem in Its Own Right ​

The harness component end users never see, but where Cursor has invested the most, is the pipeline that lands code edits on disk. The official blog post Editing Files at 1000 Tokens per Second gives the full account of this problem, and it's one of the best case studies anywhere for "why the harness matters."

The problem statement: having a frontier model directly output large stretches of modified code is slow and unreliable — the model "cuts corners" (eliding untouched code with // ... rest of code), makes unrelated changes, and regularly falls into multi-round repair loops. Cursor's solution was to split code editing into two stages: planning and applying. The frontier large model, in the conversation, only works out what to change (producing a "sketch of the edit"); turning that into a complete new file is handed to a specially trained fast apply model. Model tiering again.

Even more instructive are their experimental findings on edit formats — why have the model rewrite the whole file rather than output a diff:

  1. More output tokens = more thinking. An autoregressive model performs one more forward pass for every token it emits; a diff format forces the model to "think" in fewer tokens, which actually hurts correctness.
  2. Diffs are out of distribution. In pretraining and post-training corpora, complete code files vastly outnumber diffs — asking a model to output a diff is asking it to do something it isn't good at.
  3. Line numbers are the model's Achilles' heel. Standard diffs require line numbers, and models are notoriously bad at counting; if the tokenizer treats "123" as one token, the model has to bet on the right line number on its very first output token.

Cursor also tested Aider-style search/replace blocks (which eliminate the line-number problem, as long as the search text matches uniquely in the file) and concluded that full-file rewrites perform better overall on files under 400 lines. But note the flip side of the official finding that "with the exception of Claude Opus, most models can't produce accurate diffs": the choice of edit format depends on which model you're using. This isn't a format debate — it's a mapping of how model capabilities are distributed.

On the speed side, they fine-tuned a 70B model and paired it with speculative edits, a variant of speculative decoding adapted for code editing — the original file's contents act as the "draft," and the model only verifies and modifies the parts that differ — reaching roughly 1,000 tokens/s (about 3,500 characters/s), around 13x vanilla Llama-3-70b inference. A 70B model running at small-model speed comes at a price: this inference optimization can only be deployed on self-hosted, in-house models (closed-source API models can't do speculative edits), which explains why Cursor trains its own models: not to chase benchmarks, but because a capability gap in one harness stage had no off-the-shelf model to fill it.

The lesson this pipeline holds for tool system design at large is universal: on the surface, the "edit_file" tool is a model emitting edit instructions; underneath, it's an entire body of engineering — format choice, tolerance for parse failures, retries, specialized models. The various camps bet differently — Aider on search/replace blocks, Claude Code on "read the old text + exact replacement + full rewrite as the fallback," Cursor on a two-stage pipeline — but all of them are answering the same question: between the model's output format and the environment's actual state sits a need for a layer of robust adaptation.

The connection to permission design

The two-stage pipeline also reshapes permissions and human-machine collaboration: changes arrive as visual diffs, accepted or rejected block by block, and checkpoint snapshots created automatically during agent sessions can restore the entire workspace in one click — approval partially shifts from "whether to ask before executing" to "whether to keep it afterward." That's a different risk balance from the terminal-native agent's front-loaded permission prompts.

What Cursor Teaches Us About Harness Design ​

Taken apart, Cursor contributes four transferable judgments to harness engineering:

  1. The host environment determines the context strategy. If the IDE hands you free structured signals (diagnostics, cursor, edit history), build pipelines to harvest them; if the terminal only hands you a byte stream, sharpen the exploration tools. The first step of context engineering is taking inventory of the signals that already exist in the environment.
  2. Large codebases need offline infrastructure, not just in-the-loop exploration. Embedding indexes, incremental synchronization (Merkle trees), and content-addressed caching are the three-piece toolkit that turns "retrieval" from a runtime burden on the agent into standing infrastructure; the privacy cost has to be priced in as well.
  3. Tier models by latency budget. Hundreds-of-milliseconds prediction, second-scale edit application, and minute-scale task planning are three distinct inference workloads, each worth its own (possibly in-house) model.
  4. "Getting edits right" deserves a dedicated component. Edit format is not a detail; it's a correctness variable. When no off-the-shelf model handles some harness stage well, training a small model to fill the gap is a route the leading vendors have already validated.

Further Reading ​

References ​