Modern AI Engineering

Lesson 12.8 · 27 min

Cursor: Inside an AI-Native Code Editor

When we type in Cursor and a grey suggestion appears before we have finished thinking, or the agent edits five files in one go, what is actually happening behind the screen?

In short: Cursor is a code editor (built on VS Code) with AI woven into every part of it. It indexes our codebase into embeddings so it can find relevant code by meaning, uses a small fast model for Tab predictions, larger models for Chat and the Agent, and a specialised apply step that merges suggested edits into files quickly. Different jobs use different models because each job has a different trade-off between speed, cost and intelligence, and privacy settings control what is stored.

What is Cursor? Code editor + AI

Cursor is an AI code editor made by the company Anysphere. It is a fork of Visual Studio Code: it started from VS Code's open-source code, so it looks and feels familiar, and most VS Code extensions, themes and keyboard shortcuts work. On top of that editor, Cursor adds AI features at several levels: next-edit prediction as we type (Tab), a conversation panel (Chat/Ask), and an autonomous Agent that can search, edit many files and run commands.

A simple way to remember it: Cursor = code editor + AI everywhere. The editor provides files, cursor position, open tabs, recent edits, terminal and diagnostics; the AI uses all of that as context. In this lesson we describe publicly documented behaviour; Cursor ships fast, so feature names and details change often.

Think of a co-pilot at three distances Sometimes the co-pilot finishes your sentence (Tab), sometimes you turn and ask them a question (Chat), and sometimes you hand them the controls for a while with a destination (Agent). Same co-pilot knowledge of the route, three levels of autonomy.

The big idea behind Cursor

Two ideas run through Cursor's design. First, context is everything: a model's answer about our code is only as good as the code it sees, so the editor works hard to gather the right files, symbols and recent changes automatically. Second, match the model to the moment: a suggestion while typing must appear in a fraction of a second, while a multi-file refactor can take minutes. One model cannot be both the fastest and the smartest, so Cursor uses several.

How does Cursor understand our code?

A model has no idea what is in our repository unless the editor puts it into the prompt. Cursor gathers context from several sources:

  • What we are looking at: the current file, cursor position, selection and open tabs.
  • What we just did: recent edits, which hint at what we will do next (vital for Tab).
  • What we point at: @ mentions in Chat or Agent to include specific files, folders, docs or web results.
  • What the editor knows: errors and warnings from language servers and linters, terminal output.
  • What search finds: results of semantic search over the codebase index, plus exact text search (grep).
  • What we told it: project rules (for example files in .cursor/rules or an AGENTS.md) that are added to prompts.

Semantic search needs an index, which is the next topic.

How does Cursor index our codebase?

When codebase indexing is on, Cursor builds a semantic index, so it can answer “where do we check the password?” even if no file contains that exact phrase. Cursor's documentation describes the pipeline roughly as follows:

Building and updating the index

  1. Chunk: Files are split into smaller pieces (chunks), aiming for meaningful units such as functions or classes. Files matched by .gitignore or .cursorignore are skipped.
  2. Embed: Each chunk is turned into an embedding, a vector of numbers where similar meaning gives nearby vectors.
  3. Store: Embeddings are stored in a remote vector database together with metadata such as (obfuscated) file paths and line ranges, so results can be mapped back to local files.
  4. Sync with a Merkle tree: To stay current without re-uploading everything, Cursor hashes files into a Merkle tree (each folder's hash is built from its children's hashes) and periodically compares it with the server's copy. Only branches whose hashes differ are walked, and only changed files are re-embedded.
  5. Search: At question time, the query is embedded and the nearest chunks are returned; the editor then reads the actual code from our local files to put into the prompt.

mini_index.py: chunk, embed, Merkle-sync, search

import hashlib, re
import numpy as np
files = {"auth/login.py": "def login(user, password): check password hash and create session token",
"auth/logout.py": "def logout(session): delete session token from store",
"billing/invoice.py": "def make_invoice(order): sum line items add tax return pdf"}
def h(text): return hashlib.sha256(text.encode()).hexdigest()[:8]
def embed(text, dim=256):   # toy embedding: hashed bag of words, unit length
v = np.zeros(dim)
for w in re.findall(r"[a-z]+", text.lower()):
v[int(hashlib.md5(w.encode()).hexdigest(), 16) % dim] += 1
return v / np.linalg.norm(v)
def merkle(files):         # folder hash = hash of its children's hashes
leaves = {p: h(src) for p, src in files.items()}
dirs = {}
for p, x in sorted(leaves.items()):
dirs.setdefault(p.split("/")[0], []).append(x)
folders = {d: h("".join(xs)) for d, xs in dirs.items()}
return h("".join(folders.values())), folders, leaves
root1, folders1, leaves1 = merkle(files)
index = {p: embed(src) for p, src in files.items()}          # initial full index
files["billing/invoice.py"] += " with discount"                # we edit one file
root2, folders2, leaves2 = merkle(files)
changed = [d for d in folders2 if folders2[d] != folders1[d]]
stale = [p for p in leaves2 if leaves2[p] != leaves1[p]]
print("root changed:", root1 != root2, "| folders to walk:", changed, "| re-embed:", stale)
for p in stale: index[p] = embed(files[p])
q = embed("where do we check the password at login")
for p, s in sorted(((p, float(q @ v)) for p, v in index.items()), key=lambda t: -t[1]):
print(f"{s:.2f}  {p}")

Output:

root changed: True | folders to walk: ['billing'] | re-embed: ['billing/invoice.py']
0.49  auth/login.py
0.22  auth/logout.py
0.00  billing/invoice.py

Pause and think: A project has 20,000 files and we edit 3. Why does the Merkle tree make syncing cheap?

Comparing the root hash tells us immediately that something changed, and then we only descend into folders whose hashes differ. Unchanged folders (almost all of them) are skipped as a whole, so we find and re-embed just the 3 changed files instead of re-processing 20,000.

How does Tab autocomplete work?

Classic autocomplete completes the word at the cursor. Cursor's Tab predicts our next edit: it can suggest several lines, modify existing code around the cursor (not only insert), and suggest a jump to the next place that likely needs the same change. For example, after we rename a parameter in a function signature, Tab can propose updating its uses below, and pressing Tab moves through them.

To do this, Tab uses Cursor's own specialised model, trained for this task, with context such as the code around the cursor and our recent edits. Speed dominates the design: a suggestion that appears after we have already typed the line is useless, so the model is small enough to answer within a fraction of a second on every keystroke pause. Cursor has said it improves Tab using signals from which suggestions people accept or reject.

Why a frontier model is not used for Tab A large chat model might predict slightly better, but it would be too slow and too expensive to call on nearly every keystroke for millions of users. For Tab, a fast good guess beats a slow perfect one.

How do Chat and Agent mode work?

Chat (Ask) is a conversation in a side panel. Cursor builds a prompt from our question, the current file or selection, anything we @-mention, project rules, and (when useful) semantic search results. A larger model answers, often with code blocks. In Ask mode it only reads and explains.

Agent mode turns the chat into an agent loop like the ones in earlier lessons. The model is given tools: semantic search, grep, read file, edit file, run terminal commands, search the web, and tools from connected MCP servers. It decides which to call, sees the results, and continues until the task is done. Terminal commands can require our approval (or be allowed by a list we configure), and Cursor keeps checkpoints so we can roll back the agent's changes.

How does Cursor apply the changes?

A chat model often writes a change as a sketch: the new function body, with comments like // ... existing code ... for unchanged parts. Somebody has to turn that sketch into the exact new file. Asking the big model to rewrite the whole file is slow and risks accidental changes. Cursor instead uses a separate apply step with a specialised model that takes the original file and the suggested edit and produces the full updated file, which is then shown as a diff we can accept or reject.

Speed comes from a trick Cursor has written about called speculative edits, a cousin of speculative decoding. Most of the new file is identical to the old file, so the old file is used as the “draft”: the model verifies long runs of unchanged tokens in one parallel step and only generates token by token where the code actually changes.

Always read the diff The apply step can misplace or drop code, especially with ambiguous sketches. Accepting every diff without reading it is the most common way AI edits introduce bugs.

Why different models, and how privacy works

Different models. Each feature has a different budget. Tab needs very low latency and runs constantly, so it uses a small custom model. Apply needs fast, faithful rewriting, so it uses a specialised model. Chat and Agent need deep reasoning and tool use, so they use frontier models from providers such as Anthropic, OpenAI and Google, or Cursor's own agent model (Cursor introduced one called Composer in late 2025). Indexing needs an embedding model. Users can usually pick the chat or agent model, or let an automatic setting choose.

Privacy. Requests go from the editor through Cursor's servers to the model providers. Cursor offers a Privacy Mode; with it enabled, Cursor states that code is not stored by Cursor or its model providers for training (it relies on zero-data-retention agreements with providers). For indexing, embeddings and obfuscated metadata are stored, while the plain code is read from our local machine when needed. .cursorignore keeps files out of indexing and AI features. Organisations should check the current security documentation, because these policies are what really matter for compliance.

Pause and think: A teammate worries that turning on codebase indexing uploads the whole repo as plain text to be stored. How would you answer, based on Cursor's documented design?

Chunks are sent to compute embeddings, and what is stored remotely is the embeddings plus obfuscated paths and line ranges; the actual code shown to the model is read locally at request time. Files can be excluded with .cursorignore, and Privacy Mode controls retention. For strict compliance, verify against Cursor's current security docs.

Worked example, step by step

The mini index above had three files, so its Merkle tree saved almost nothing. Let us count on a larger project. Take 20,000 files spread evenly over 200 folders, 100 files each, and a two-level tree like the one in our code: one root hash, 200 folder hashes and 20,000 file hashes. We edit 3 files that sit in 2 folders. The layout is illustrative.

Finding the 3 changed files

  1. Compare the root: 1 comparison. The roots differ, so something changed.
  2. Compare the folder hashes: 200 comparisons. 198 folders match and are skipped whole. 2 folders differ.
  3. Open the 2 changed folders: 2 × 100 = 200 file-hash comparisons. They reveal the 3 changed files.
  4. Re-embed: Only those 3 files are chunked and embedded again.
  5. Add it up: 1 + 200 + 200 = 401 comparisons, against 20,000 if we compared every file hash. That is about 50 times fewer.

Two details are worth noticing. The count depends mostly on how many folders the edits touch, not on the size of the project: a project ten times larger, with the same edits in 2 folders, needs 1 + 2,000 + 200 = 2,201 comparisons, not ten times 401. And comparing hashes is the cheap part. The costly work is embedding, and that is done for 3 files either way, as long as we know which 3. The tree's job is to find them quickly.

When nothing has changed at all, the answer costs a single comparison of the root.

Practice: try it yourself

We will build a plain-code stand-in for the apply step. It takes the original file and a sketch made of changed lines plus # ... existing code ... markers, and it produces the full new file. Then it prints the diff a person would review. Cursor uses a trained model for this job; our version uses a simple copying rule, which is enough to see why the step exists and how it can go wrong.

practice_apply_sketch.py

import difflib
MARK = "# ... existing code ..."
original = ["def total(prices):",
"    subtotal = sum(prices)",
"    tax = subtotal * 0.2",
"    return subtotal + tax",
"",
"def label(name):",
"    return name.upper()"]
# The sketch a chat model might write: changed lines plus markers
sketch = ["def total(prices, discount=0):",
"    subtotal = sum(prices) - discount",
"    " + MARK,
"",
"def label(name):",
"    " + MARK]
def apply(original, sketch):
# Plain-code stand-in for the apply step: build the full new file
out, i = [], 0                        # i points into the original file
for j, line in enumerate(sketch):
if line.strip() == MARK:
# Copy original lines until the sketch's next line shows up
stop = sketch[j + 1] if j + 1 < len(sketch) else None
while i < len(original) and original[i] != stop:
out.append(original[i]); i += 1
else:
out.append(line); i += 1      # a line the sketch rewrote
return out
new = apply(original, sketch)
copied = sum(a == b for a, b in zip(original, new))
print(f"{copied} of {len(new)} lines copied unchanged from the old file")
for d in difflib.unified_diff(original, new, "before", "after", lineterm="", n=0):
print(d)

Output:

5 of 7 lines copied unchanged from the old file
--- before
+++ after
@@ -1,2 +1,2 @@
-def total(prices):
-    subtotal = sum(prices)
+def total(prices, discount=0):
+    subtotal = sum(prices) - discount

Now change it:

  • Delete the empty string "" from sketch. Predict the new file before running. Look at which line the first marker now waits for.
  • Insert a new line, " prices = list(prices)", as the second line of sketch. Predict which original line goes missing from the result, then find it in the diff.
  • Replace the last marker line with " return name.lower()". Predict the hunks of the diff and the “copied unchanged” count.

Pause and think: The program reports that 5 of 7 lines were copied unchanged. How does that number relate to the speculative edits idea from this lesson?

Speculative edits rest on the same observation: most of the new file is identical to the old one. Here 5 of 7 lines need no new writing at all, and only 2 must be produced fresh. That is why the old file makes a good draft: long unchanged stretches can be checked and accepted quickly, and slow token-by-token generation is needed only where the code really changes.

Pause and think: Our copying rule stops when it meets the sketch's next line in the original file. Suppose that next line were return x, and the original file had return x in three functions. What could go wrong, and who catches it?

The rule would stop at the first return x it meets, which may not be the one the sketch meant, so code could be dropped or land in the wrong function. The sketch is ambiguous, and no merging rule can fully repair that. The safety net is the diff: a reviewer who reads it sees lines removed that nobody asked to remove.

The complete flow of Cursor

From opening a project to merged change

  1. Open the project: Cursor indexes it (chunk → embed → store) and keeps the index in sync using Merkle-tree hashing.
  2. Type code: Tab's fast model predicts next edits from local context and recent changes.
  3. Ask a question: Chat gathers context (current file, @-mentions, rules, semantic search) and a large model answers.
  4. Give a goal to the Agent: The agent loops: search the index and grep, read files, edit, run commands (with approval), check results.
  5. Apply and review: Edits are merged by the apply step and shown as diffs; we accept, reject or roll back to a checkpoint.

Compared with Claude Code from the previous lesson, the agent loop is the same idea. The differences are in the harness: Cursor lives inside an editor and offers a pre-built semantic index plus Tab, while Claude Code started in the terminal and leans on agentic search with grep and file reads. Both depend on our tests and reviews for a real definition of done.

Key takeaways

  • Cursor is VS Code plus AI at three levels: Tab, Chat and Agent.
  • A semantic index (chunk, embed, store) lets it find code by meaning; Merkle-tree hashing keeps it in sync cheaply.
  • Tab predicts the next edit with a small fast model because latency is everything while typing.
  • Agent mode is a tool loop: search, read, edit, run commands, check, with our approval and checkpoints.
  • A separate apply step merges sketches into files quickly using speculative edits; always review the diff.
  • Different features use different models because their speed, cost and intelligence needs differ.

Key terms

  • Fork: A new project built from a copy of another project's source code, here VS Code.
  • Codebase index: Embeddings of code chunks stored for semantic search over a project.
  • Merkle tree: A tree of hashes where each parent's hash is computed from its children's, so changes can be located quickly.
  • Tab (next-edit prediction): Cursor's feature that predicts and suggests the next code edit as we type.
  • Apply model: A specialised model that merges a suggested code change into the full file.
  • Speculative edits: Speeding up file rewriting by using the original file as a draft that the model verifies in parallel.

← 12.7 Claude Code: AI-Powered Software Engineering at the CLI · 13.1 LLM Inference Optimization: The Full Landscape →