Lesson 12.7 · 27 min
Claude Code: AI-Powered Software Engineering at the CLI
How does an AI go from “here is a code snippet, good luck pasting it” to actually opening our files, running our tests and fixing the bug itself?
In short: Claude Code is Anthropic's agentic coding tool. It runs a Claude model in an agent loop with real tools (read and edit files, search with patterns, run shell commands) inside our project, so it can gather context, make changes and verify them by running tests. Project instructions live in CLAUDE.md, permission rules keep us in control of risky actions, and plan mode, subagents and hooks let us shape how it works. It is a clear real-world example of harness and loop engineering.
What is Claude Code?
Claude Code is an agentic coding tool made by Anthropic. “Agentic” means it does not just suggest text: it takes actions toward a goal. It started as a command-line program that runs in our terminal (released as a research preview in February 2025 and generally available a few months later) and is now also available through IDE extensions, a desktop app and the web. In every form, the idea is the same: we describe a task in plain language, and Claude reads our code, edits files, runs commands, and reports back.
Under the hood it is exactly what the previous lessons described: a strong model (Claude) plus a carefully engineered harness: tools, a loop, context management, project memory, permissions and extension points. In this lesson we describe only behaviour that Anthropic documents publicly; details such as exact tool names can change between versions.
Think of a new teammate at your desk A chatbot is like a friend on the phone: you read your code aloud, they suggest a fix, you type it in, you tell them the error. Claude Code is like a teammate sitting at your keyboard: they open the files themselves, run the tests themselves, and ask you before doing anything risky.
Running example: a small shop project where total([5, 10]) returns 10 instead of 15, and a test is failing.
The problem with a normal AI chatbot
With a chat assistant in a browser, we are the harness. We copy code in, the model guesses at the parts it cannot see, we paste its answer back, run the tests, copy the error back, and repeat. Three problems follow:
- Missing context: the model sees only what we pasted. It does not know how
totalis called elsewhere, what the tests expect, or our project's conventions. - No verification: it cannot run the code, so it cannot know whether its fix works. Plausible but wrong answers look identical to right ones.
- We do the busywork: every loop iteration (copy, paste, run, report) is manual, slow and error-prone.
The agent loop
Anthropic describes Claude Code's way of working as a loop of three phases: gather context, take action, verify results, repeated until the task is done, with us able to interrupt and steer at any point.
One task through the loop
- Receive the task: Our request, plus the system instructions, CLAUDE.md contents and the list of available tools, form the starting context.
- Gather context: The model calls search and read tools to find the relevant files, as a developer would.
- Take action: It proposes edits or commands. The harness checks permissions, asks us if needed, then runs them.
- Verify: It runs tests, builds or linters and reads the output.
- Repeat or finish: If verification fails, the error goes back into context and the loop continues. When it passes, Claude summarises what changed.
The tools of Claude Code
The model can only do what its tools allow. The built-in tool set covers what a developer does at a terminal. The documented set includes tools along these lines (names may evolve):
| Tool | What it does | Needs permission by default? |
|---|---|---|
| Read | Read a file's contents (also images and PDFs) | No |
| Glob | Find files by name pattern, e.g. src/*/.py | No |
| Grep | Search file contents with regular expressions | No |
| Edit / Write | Change part of a file by exact string replacement, or write a whole file | Yes |
| Bash | Run a shell command: tests, builds, git, package managers | Yes |
| WebFetch / WebSearch | Read a web page or search the web | Yes |
| Agent (subagents) | Hand a sub-task to a separate agent with its own context | No |
| To-do list | Keep a visible task checklist for multi-step work | No |
Two design choices are worth noticing. First, the Edit tool replaces an exact snippet of text with new text, and fails if the snippet is not found or not unique. That forces the model to have actually read the file and makes each change small and reviewable. Second, Bash is a universal tool: anything we can do in a shell (run pytest, npm test, git diff), the agent can do, which is powerful and is exactly why it is permission-gated.
Beyond built-ins, Claude Code can connect to external tools through the Model Context Protocol (MCP), such as issue trackers, databases or browsers.
Example: fixing a bug
Here is a runnable simulation of the loop for our shop bug. The “model decisions” are scripted so the run is repeatable, but the tools work on a real in-memory project, the edit uses exact-match replacement, and edits go through a permission check.
mini_claude_code.py
import re
repo = {"cart.py": "def total(prices):\n return sum(prices[1:])\n",
"test_cart.py": "assert total([5, 10]) == 15\n",
"README.md": "Shop demo"}
def grep(pattern): return [f for f, src in repo.items() if re.search(pattern, src)]
def read(path): return repo[path]
def edit(path, old, new):
assert repo[path].count(old) == 1, "old text must be unique" # exact-match edit
repo[path] = repo[path].replace(old, new)
def run_tests():
env = {}; exec(repo["cart.py"], env)
try: exec(repo["test_cart.py"], env); return "1 passed"
except AssertionError: return "1 failed"
TOOLS = {"grep": grep, "read": read, "edit": edit, "run_tests": run_tests}
ALLOW = {"grep", "read", "run_tests"} # read-only tools run freely
# What a model might decide, turn by turn, after seeing each result
plan = [("run_tests",), ("grep", r"def total"), ("read", "cart.py"),
("edit", "cart.py", "prices[1:]", "prices"), ("run_tests",)]
for name, *args in plan:
if name not in ALLOW:
print(f" [permission] allow {name}{tuple(args)}? -> yes (user approved)")
result = TOOLS[name](*args)
shown = repr(result)[:40] if result is not None else "ok"
print(f"{name:9} -> {shown}")
print("final cart.py:", repr(repo["cart.py"]))Output:
run_tests -> '1 failed'
grep -> ['cart.py']
read -> 'def total(prices):\n return sum(pric
[permission] allow edit('cart.py', 'prices[1:]', 'prices')? -> yes (user approved)
edit -> ok
run_tests -> '1 passed'
final cart.py: 'def total(prices):\n return sum(prices)\n'Notice the order: the agent reproduced the failure first (so it knows the test is a valid signal), then found and read the code, made the smallest change, and verified with the same test. That is good engineering practice, and it is what a well-instructed coding agent tries to do.
Pause and think: What would the Edit tool do if the model tried to replace prices (instead of prices[1:]) in cart.py?
It would refuse. prices appears more than once in the file (in the parameter list and in the sum), so the exact-match rule fails with “old text must be unique”. The model must include enough surrounding text to identify one location.
Searching a big project and verifying the work
Searching. A real repository may have thousands of files, far too many to put in the context window. Anthropic has explained that Claude Code relies on agentic search: the model uses Glob, Grep and Read step by step, the way a developer explores unfamiliar code (find files named like cart, grep for def total, read the hits, follow imports), rather than requiring a pre-built embedding index of the codebase. The benefits are that search always reflects the current files and there is no index to build or sync; the cost is extra tool calls on very large codebases.
For broad investigations, Claude Code can hand searches to subagents (see below), which explore in their own context and return only a summary, keeping the main conversation focused.
Verifying. The model's confidence is not evidence. Claude Code verifies by running the project's own checks through Bash: unit tests, type checkers, linters, builds. Failures come back as text the model can read and act on, which is the try–check–retry loop from the definition-of-done lesson. Its effectiveness therefore depends heavily on us: a project with good tests and a clear “how to run tests” instruction gives the agent a reliable definition of done.
Make verification easy Tell Claude how to verify in CLAUDE.md (“run pytest -q”, “run npm run typecheck after edits”) and, where possible, ask for a failing test first. Agents are much better at hitting a target they can run.
CLAUDE.md: the project memory
Each session starts with no memory of previous ones. To give Claude lasting knowledge about a project, we write a Markdown file called CLAUDE.md. Claude Code automatically loads it into context at the start of a session. Typical contents: how to build and test, code style rules, important directories, things to avoid (“never edit generated files in gen/”).
- Project level:
CLAUDE.mdin the repository root, committed so the whole team shares it. - User level:
~/.claude/CLAUDE.mdfor personal preferences that apply to all projects. - Subdirectories: CLAUDE.md files in sub-folders are pulled in when Claude works with files there.
- Helpers: the
/initcommand drafts a CLAUDE.md by exploring the project; files can import others with@pathsyntax.
Keep it short and true CLAUDE.md is sent with every session, so it costs context each time, and stale instructions mislead the agent. Write concise, specific, current rules; prune what no longer applies.
Permissions: how we stay in control
An agent that can run shell commands could also delete files or push to production. Claude Code's answer is a permission system. By default, read-only actions run without asking, while file edits and shell commands ask for approval the first time (we can approve once or allow for the session).
- Permission rules in settings files can allow or deny specific tools and patterns, for example allow
Bash(npm run test:*)but deny reading.envfiles. Deny rules take precedence. - Permission modes change the default: a normal mode that asks, a mode that auto-accepts file edits, plan mode that is read-only, and a mode that skips prompts entirely, which is intended only for isolated environments such as containers.
- Interrupting: we can stop Claude at any time and redirect it, and file changes can be reviewed (and rewound via checkpoints) before we commit.
Pause and think: A team wants Claude to run tests freely but never run git push. What permission setup fits?
An allow rule for the test command pattern (for example Bash(npm run test:)) and a deny rule for Bash(git push:). Tests then run without prompts, and pushes are blocked even if the model tries.
Plan mode, subagents and hooks
Plan mode is a read-only mode: Claude may explore and read but not edit or run changing commands. It produces a plan that we review and approve before any change is made. It is ideal for large or risky tasks where we want to agree on the approach first.
Subagents are specialised helpers, each running in its own context window with its own instructions and allowed tools. Custom ones are defined as Markdown files (for example in .claude/agents/). The main agent delegates a sub-task (“find every place we parse dates”), the subagent does many searches and reads, and only its summary returns. This keeps the main context clean and lets work be split.
Hooks are our own shell commands that Claude Code runs automatically at defined moments in its lifecycle, such as before a tool runs (PreToolUse), after a tool runs (PostToolUse), when we submit a prompt, or when Claude finishes. Unlike instructions in CLAUDE.md, which the model may or may not follow, hooks are deterministic: a hook can always run the formatter after every edit, or block an edit to a protected file.
Worked example, step by step
This lesson said that subagents keep the main context clean. Let us put illustrative numbers on that. We ask: “find every place we parse dates”. Answering takes 12 search and read calls, and each call returns about 1,500 tokens of file content.
Where do the tokens go?
- Search in the main conversation: 12 × 1,500 = 18,000 tokens of file content are added to the main context. Most of it is code we looked at once and will not need again.
- Carry it on every turn: An agent loop sends its context again on each later turn. If the task runs for 20 more turns, those 18,000 tokens are carried 20 more times.
- Delegate instead: A subagent makes the same 12 calls in its own context window. That window fills with the same 18,000 tokens, but it is set aside when the subagent finishes.
- Only the summary returns: The main agent receives a short report, say 300 tokens: the files and line numbers where dates are parsed.
- Compare: 18,000 tokens against 300 in the main context: 60 times less. The main conversation stays focused on the fix.
There is a price on the other side. The subagent starts with none of the main conversation, so the task we hand over must stand on its own: what to look for, where, and what to report back. A summary can also leave out a detail we need later, and then the main agent must read that file itself. Delegation suits broad searches with a short result. For one file we are about to edit, reading it directly is simpler.
Practice: try it yourself
We will build a small permission checker in the spirit of this lesson: read-only tools run freely, deny rules beat allow rules, anything the rules do not cover asks the user, and a hook that we wrote can block a call. The rule format is a simplified sketch, not the tool's exact syntax.
practice_permission_rules.py
from fnmatch import fnmatchcase
# A simplified sketch of permission rules: deny beats allow, the rest asks.
ALLOW = ["Read(*)", "Grep(*)", "Bash(npm run test*)"]
DENY = ["Bash(git push*)", "Read(.env)"]
def pre_tool_hook(tool, arg):
# Our own deterministic check, run before every tool call
if tool == "Edit" and arg.startswith("migrations/"):
return "blocked by hook: migrations are protected"
return None
def decide(tool, arg):
call = f"{tool}({arg})"
blocked = pre_tool_hook(tool, arg)
if blocked:
return blocked
if any(fnmatchcase(call, rule) for rule in DENY): # deny is checked first
return "denied by rule"
if any(fnmatchcase(call, rule) for rule in ALLOW):
return "allowed, runs without asking"
return "ask the user first"
# Tool calls a model might propose while fixing the cart bug
calls = [("Read", "cart.py"), ("Read", ".env"), ("Bash", "npm run test:unit"),
("Edit", "cart.py"), ("Edit", "migrations/001.sql"),
("Bash", "git push origin main"), ("Bash", "rm -rf build")]
for tool, arg in calls:
print(f"{tool + '(' + arg + ')':27} -> {decide(tool, arg)}")Output:
Read(cart.py) -> allowed, runs without asking Read(.env) -> denied by rule Bash(npm run test:unit) -> allowed, runs without asking Edit(cart.py) -> ask the user first Edit(migrations/001.sql) -> blocked by hook: migrations are protected Bash(git push origin main) -> denied by rule Bash(rm -rf build) -> ask the user first
Now change it:
- Add
"Edit(*)"toALLOW. Predict the result for both Edit calls. Which one still does not run, and why? - Add
"Bash(*)"toALLOW. Predict the results forgit push origin mainand forrm -rf build. Is this a rule we would want? - Remove
"Read(.env)"fromDENY. Predict what happens toRead(.env). Why does a broad allow rule make a specific deny rule necessary?
Pause and think: Read(.env) matches the allow rule Read(*) and also the deny rule Read(.env). Why is “deny wins” the safer way to settle the tie?
Allow rules are usually broad, so that routine work runs without prompts. Deny rules are narrow and protect specific things. If allow won, every broad allow rule would silently cancel the protections behind it, and we would have to remember each secret file whenever we wrote one. With deny winning, a mistake errs on the side of blocking, which costs us a prompt and not a leaked secret.
Pause and think: Bash(rm -rf build) matched no rule and got “ask the user first”. Why is asking a better default for unmatched calls than allowing them or denying them?
We cannot list every command in advance. If unmatched calls were allowed, anything we forgot to deny would run, including destructive commands. If they were denied, the agent would be stopped by every new, harmless command. Asking puts a person in front of exactly the cases the rules did not foresee, and each answer shows us which rule to add next.
Putting it all together
When we type “fix the failing cart test” into Claude Code, the harness assembles a context from the system instructions, our CLAUDE.md and the tool definitions. The model reproduces the failure with Bash, searches with Grep and Glob, reads the relevant files, proposes an exact-match Edit (which may trigger a permission prompt and hooks), re-runs the tests, and loops until they pass or it needs our input. Long sessions are kept within the context window by compaction (summarising earlier parts), and side investigations can go to subagents.
What it is good and less good at Strong: tasks with a runnable definition of done (failing tests, type errors, build failures), codebase exploration, multi-file refactors with tests, writing tests. Weaker: vague goals (“make it better”), tasks needing context that is not in the repo or docs, and projects with no way to verify. As always, review the diff before merging.
Every idea from this module appears here: a harness (tools, permissions, hooks, CLAUDE.md), a loop with verification (gather, act, verify), and a definition of done supplied by our tests. The next lesson looks at Cursor, which wraps similar ideas inside a code editor.
Key takeaways
- Claude Code is a Claude model in an agent loop with real tools in our project: gather context, act, verify.
- Agentic search (Glob, Grep, Read) finds relevant code on demand instead of pasting it in.
- Verification through tests, builds and linters is what makes its changes trustworthy.
- CLAUDE.md gives persistent project instructions; keep it short and current.
- Permission rules and modes keep risky actions under our control.
- Plan mode, subagents and deterministic hooks let us shape and constrain how it works.
Key terms
- Agentic coding tool: A tool where an AI model acts in a codebase (reads, edits, runs commands) in a loop toward a goal.
- Agentic search: Finding relevant code by having the model call search and read tools step by step.
- CLAUDE.md: A Markdown file of project instructions that Claude Code loads into context at session start.
- Permission mode: A setting that controls which actions Claude Code may take without asking, such as plan mode or auto-accepting edits.
- Subagent: A helper agent with its own context window and tools that handles a delegated sub-task and returns a summary.
- Hook: A user-defined shell command the harness runs automatically at a lifecycle event, able to add checks or block actions.
← 12.6 LangGraph: Graph-Based Agent Orchestration Explained · 12.8 Cursor: Inside an AI-Native Code Editor →