Modern AI Engineering

Lesson 11.9 · 23 min

Agent Skills: Reusable Capabilities in Agentic Systems

How can one agent know how to fill PDF forms, follow your brand guide and run your release checklist, without stuffing all of that into every prompt?

In short: An Agent Skill is a folder with a SKILL.md file (a short name and description, then instructions) plus any scripts or reference files the task needs. The agent sees only each skill's name and description at first, and opens the full instructions and files only when a task needs them. This "progressive disclosure" lets one agent carry many skills while keeping its context window small.

The problem before Agent Skills

Let us follow one running example through this lesson. We run a small company and use an AI agent for office work. We want it to fill in PDF forms the way our finance team likes, make slides in our brand style, and follow our 12-step release checklist.

Before skills, we had three weak options. We could paste every instruction into the system prompt (the fixed text the model reads before every conversation). We could paste the right instructions into each chat by hand. Or we could fine-tune a model, which is slow, costly and hard to update.

  • A giant system prompt wastes the context window (the limited number of tokens the model can read at once) on instructions that are irrelevant to most requests, costs money on every call, and can distract the model.
  • Copy-pasting by hand is error-prone. Every person keeps a slightly different version of the "right" prompt.
  • Prompts cannot run code. A careful procedure like "extract every form field, then validate the dates" is more reliable as a small script than as prose the model must imitate.

Think of it like a new employee's bookshelf A new hire does not memorise every company manual on day one. They read the spines on the shelf: "Expense policy", "Brand guide", "Release checklist". When a task needs one, they pull that binder down and read it. Agent Skills give an agent the same shelf: spines always visible, binders opened only when needed.

What are Agent Skills, and what is inside one?

An Agent Skill is a folder that packages instructions, scripts and resources for one kind of task, so an agent can discover it and load it on demand. Anthropic introduced Skills for Claude in October 2025 and published the format as an open specification in December 2025 (at agentskills.io), so other agent tools can read the same folders. Support details still differ between products, so always check the docs of the agent you use.

The key point: a skill is not a new model and not a new API. It is plain files on disk. The agent already knows how to read files and run commands; a skill just tells it which files to read, when, and what to do with them.

So "skill" here means packaged know-how: a procedure, a style, domain knowledge, or a tested script. The model stays the same; what it knows how to do grows as we add folders.

What is inside a Skill? The only required file is SKILL.md. It begins with YAML frontmatter (a small block of key: value metadata between two --- lines) followed by normal Markdown instructions. Everything else is optional.

pdf-forms/SKILL.md

---
name: pdf-forms
description: Fill in and read PDF forms. Use when the user mentions a PDF form, form fields, or filling a PDF.
---
# Filling PDF forms
1. Run `python scripts/list_fields.py <file>` to list every field.
2. Map the user's data to the field names. Ask if anything is missing.
3. Run `python scripts/fill.py <file> <data.json>` and open the result.
4. Dates must be DD/MM/YYYY. See reference/finance-rules.md for other rules.
A typical skill folder
PathWhat it holdsRequired?
SKILL.mdFrontmatter (name, description) + main instructionsYes
scripts/Code the agent can run, e.g. fill.pyNo
reference/ or other .md filesLonger docs read only when neededNo
assets/Templates, fonts, images, sample dataNo

In the open spec, name must be short (up to 64 characters, lowercase letters, numbers and hyphens) and normally matches the folder name. description can be up to 1,024 characters. A few optional fields exist, such as license and tool permissions, but which ones an agent honours varies by product.

The description is the trigger

There is no keyword router or classifier hidden in the system. At startup the agent puts each skill's name and description into its context. When we ask for something, the model itself reads those descriptions and decides whether a skill is relevant. So the description is the only thing that decides whether the skill ever gets used.

A good description says two things: what the skill does and when to use it, using the words users are likely to type.

Pause and think: Our skill has perfect instructions, but the agent never uses it when people say "fill this form". What is the first thing to fix?

The description. The model only sees name and description until it decides to load the skill, so if the description does not mention forms or filling, the agent has no reason to open it. The body does not matter until then.

Progressive disclosure, the main idea

Progressive disclosure means showing information in layers: a little first, more only when needed. Skills use three levels.

The three levels of a skill

  1. Level 1: metadata (always loaded): At startup the agent reads only name + description of every installed skill. That is roughly a hundred tokens per skill, so even dozens of skills cost little.
  2. Level 2: the SKILL.md body (on trigger): When a request matches, the agent reads the full SKILL.md instructions into context. Guidance from Anthropic is to keep this body fairly short (a few thousand tokens, under about 500 lines) and move detail elsewhere.
  3. Level 3: extra files (only if needed): The body points to other files, like reference/finance-rules.md or scripts/fill.py. The agent opens a reference file only if this task needs it, and it can run a script and read only its output, without loading the code into context.
  4. Work and finish: The agent follows the steps, runs the scripts, checks the result and answers. Skills that never matched stayed at Level 1 the whole time.

A Skill can carry real code

Language models are good at judgement and bad at exact, repetitive work like counting fields or validating a date format a hundred times. A script does that exactly the same way every run. So a skill can ship scripts, and the instructions tell the agent when to run them.

This also saves context: when the agent runs scripts/fill.py, the code itself does not need to be read; only the printed result comes back. To use script skills, the agent needs a code execution environment (a sandbox or a shell). Without one, only the text parts of a skill work.

Below we simulate the idea in plain Python: parse frontmatter, build the Level 1 list, let a simple word-overlap rule stand in for the model's choice, and compare token costs.

progressive_disclosure.py

# Simulate progressive disclosure for Agent Skills (stdlib only).
SKILLS = {
"pdf-forms": """---
name: pdf-forms
description: Fill in and read PDF forms. Use when the user mentions a PDF form, fields, or filling a PDF.
---
# Filling PDF forms
1. Run scripts/list_fields.py on the file.
2. Map the user's data to field names.
3. Run scripts/fill.py and check the result.
""" + "Detailed notes about edge cases. " * 300,
"brand-slides": """---
name: brand-slides
description: Make slide decks in our company style. Use for presentations, decks, or slides.
---
# Brand slides
Use the colours in assets/palette.json and the layouts in reference/layouts.md.
""" + "Style rules and examples. " * 200,
}
def tokens(text):            # rough rule of thumb: ~0.75 words per token
return round(len(text.split()) / 0.75)
def frontmatter(md):         # read the YAML block between the two '---' lines
block = md.split("---")[1]
return dict(line.split(": ", 1) for line in block.strip().splitlines())
meta = {k: frontmatter(v) for k, v in SKILLS.items()}
startup = "\n".join(f"- {m['name']}: {m['description']}" for m in meta.values())
print("Level 1 (always loaded):", tokens(startup), "tokens")
print("If we loaded every full SKILL.md:", sum(tokens(v) for v in SKILLS.values()), "tokens")
request = "please fill this pdf form with my address"
words = set(request.split())
score = {}
for name, m in meta.items():   # stand-in for the model reading descriptions
desc = set(m["description"].lower().replace(",", "").replace(".", "").split())
score[name] = sorted(words & desc)
print(f"{name:12s} overlap={score[name]}")
best = max(score, key=lambda n: len(score[n]))
print("Level 2 loads:", best, "->", tokens(SKILLS[best]), "tokens")

Output:

Level 1 (always loaded): 48 tokens
If we loaded every full SKILL.md: 3173 tokens
pdf-forms    overlap=['fill', 'form', 'pdf']
brand-slides overlap=[]
Level 2 loads: pdf-forms -> 2065 tokens

The output shows the trade: 48 tokens of metadata let the agent know about both skills, and only the needed body (about 2,000 tokens) is paid for when a matching task arrives.

Where Skills live and how to create one

Because a skill is just a folder, "installing" one means putting the folder where the agent looks. The exact places depend on the product. For Claude, as documented at the time of writing:

WhereLocationWho sees it
Claude Code, personal~/.claude/skills/<skill-name>/SKILL.mdYou, in every project
Claude Code, project.claude/skills/<skill-name>/SKILL.md in the repoEveryone who clones the repo
Claude Code pluginsBundled inside an installed pluginAnyone with the plugin
Claude appsUploaded as a zip in settings (plus built-in skills)Your account / organisation
Claude APIUploaded via the API and attached to requests that use code executionYour app

Creating our own skill

  1. Pick one repeatable task: For example "write release notes from merged pull requests". One skill = one job. Narrow skills trigger more reliably.
  2. Make the folder and SKILL.md: Create release-notes/SKILL.md with name: release-notes and a description that says what it does and when ("Use when the user asks for release notes or a changelog").
  3. Write short, ordered instructions: Numbered steps, the format of the result, and examples of good output. Move long material into separate files and link them by relative path.
  4. Add scripts for exact work: Anything deterministic (collect PR titles, sort by label) goes into scripts/. Tell the agent when to run them and what the output means.
  5. Test with real prompts: Try requests that should trigger it and ones that should not. Fix the description first if it misfires, then the body.

Agent Skills vs MCP

The Model Context Protocol (MCP) is an open protocol that connects an agent to external tools and data through a server: a database, GitHub, a calendar. People often ask whether skills replace MCP. They do different jobs: MCP gives the agent new abilities to reach things, skills give it know-how about doing things.

A real example and why Skills matter

Built-in document skills Anthropic ships skills for creating and editing Word, Excel, PowerPoint and PDF files. When we ask Claude for "a spreadsheet of these numbers", it triggers the spreadsheet skill, reads its instructions and runs its helper scripts in a sandbox to produce a real file. The same mechanism is open to us: a team can write a "quarterly-report" skill containing its template and a script that pulls the numbers.

  • Composable: many skills can be installed; the agent may use two together (e.g. a data skill and a slides skill).
  • Portable: the same folder can work in several agents that follow the open format.
  • Reviewable: it is plain text and code in git, so changes get code review like anything else.
  • Cheap: thanks to progressive disclosure, an unused skill costs about one line of context.
  • Reliable: scripts make exact steps exact, and instructions encode what experts already know.

Pause and think: We install 40 skills. Roughly how much context do they cost on a simple "hello" message, and why?

Only their metadata: on the order of 40 × ~100 tokens, a few thousand tokens at most. No skill matches "hello", so no SKILL.md body or extra file is loaded.

Things we must be careful about

A skill is code you are trusting Installing a skill means letting its instructions steer the agent and letting its scripts run in your environment. A malicious skill could tell the agent to send data somewhere or run a harmful script. Only install skills from sources you trust, and read the SKILL.md and scripts first, just as you would review a dependency.

  • Vague or overlapping descriptions: two skills that both claim "documents" confuse the model. Make triggers distinct.
  • Bloated SKILL.md: putting everything in the body defeats progressive disclosure. Keep the body lean and link to reference files.
  • No execution environment: script-based skills silently degrade where code cannot run.
  • Stale skills: when the procedure changes, the skill must change too. Version them in git.
  • Not a security boundary: instructions in a skill are guidance to the model, not enforced permissions. Real limits belong in the sandbox and tool permissions.

When not to use a skill: for a one-off task, just write the instruction in the chat. For reaching a live external system, you need a tool or MCP server (possibly plus a skill). For changing the model's core behaviour across everything, a system prompt or fine-tuning may fit better.

Common mistakes and how to spot them

Most skill problems are trigger problems, and they come in two kinds. A skill under-triggers when a request needed it and the agent did not load it. A skill over-triggers when it is loaded for a request that had nothing to do with it. Both are invisible unless we test for them.

The test is simple: write a short list of prompts, note which skill each one should load (or none), run them, and compare. Here is such a test for our office agent with three skills: pdf-forms, brand-slides and release-notes. The results are illustrative.

A small trigger test (illustrative results)
PromptShould loadDid loadVerdict
“Fill this PDF form with my address”pdf-formspdf-formsCorrect
“Complete the attached tax document”pdf-formsnoneUnder-trigger: the user said “document”, the description only says “form”
“Make a deck for Monday's meeting”brand-slidesbrand-slidesCorrect
“Summarise this PDF report”nonepdf-formsOver-trigger: the description says “read PDF”, which is too broad
“Write a changelog for version 2.3”release-notesrelease-notesCorrect
“Hello, how are you?”nonenoneCorrect

Diagnosing a failed row

  1. Did the right skill load at all?: If not, the body is not the problem; it was never read. Look only at the description.
  2. Under-trigger: add the user's words: The description lacked the words real users type. Add them: “Use for PDF forms, applications, or documents with fields to fill in.”
  3. Over-trigger: narrow the claim: The description promised more than the skill does. Say what it is not for: “Not for summarising or reading ordinary PDFs.”
  4. Two skills fight for one prompt: If both load, or the wrong one wins, their descriptions overlap. Rewrite them so each names a different task.
  5. Right skill, wrong result: Only now look at the body: unclear steps, a missing example, or a script that fails without an execution environment.
  6. Re-run the whole list: A wider description can fix one row and break another. Always run every prompt again after a change.

In this test, 4 of 6 prompts behaved correctly. Both failures point at one description, so one careful edit to pdf-forms may fix both. Keep the prompt list next to the skill in git; it is the cheapest test suite we will ever write.

Practice: try it yourself

The earlier code compared Level 1 and Level 2. Now we follow one request through all three levels, including the part that is easy to miss: a script is run, and only its output enters the context. We keep a small skill folder in memory and print the context size after every load. We count words, as a rough stand-in for tokens.

practice_skill_levels.py

import re
# A skill folder held in memory. We track every word that enters the context.
FILES = {
"date-check/SKILL.md": "---\nname: date-check\n"
"description: Check dates in a form. Use when the user asks to validate dates.\n---\n"
"1. Run scripts/check.py on the dates.\n"
"2. Only if a date fails, read reference/rules.md and explain the rule.",
"date-check/reference/rules.md": "Dates must be DD/MM/YYYY. " + "Another finance rule. " * 40,
"date-check/scripts/check.py": "# imagine careful date parsing here\n" * 60,
}
def run_script(path, dates):             # stands in for a sandbox running the file
bad = [d for d in dates if not re.fullmatch(r"\d\d/\d\d/\d{4}", d)]
return f"{len(dates) - len(bad)} ok, bad: {bad}"
context = []                             # everything the model has read so far
def load(label, text):
context.append(text)
total = sum(len(c.split()) for c in context)
print(f"{label:30s} +{len(text.split()):3d} words  (context now {total})")
_, front, body = FILES["date-check/SKILL.md"].split("---\n")
meta = dict(line.split(": ", 1) for line in front.strip().splitlines())
assert re.fullmatch(r"[a-z0-9-]{1,64}", meta["name"])      # the naming rule
# The steps below are the choices a model would make; here they are scripted.
dates = ["03/04/2026", "2026-04-05"]
load("Level 1: name + description", f"{meta['name']}: {meta['description']}")
load("user request", "Please validate these dates: " + " ".join(dates))
load("Level 2: SKILL.md body", body)                       # the description matched
result = run_script("date-check/scripts/check.py", dates)  # body step 1
load("Level 3: script output", result)
if "bad: []" not in result:                                # body step 2: only if needed
load("Level 3: reference/rules.md", FILES["date-check/reference/rules.md"])
print("words in scripts/check.py, never loaded:", len(FILES["date-check/scripts/check.py"].split()))

Output:

Level 1: name + description    + 14 words  (context now 14)
user request                   +  6 words  (context now 20)
Level 2: SKILL.md body         + 18 words  (context now 38)
Level 3: script output         +  4 words  (context now 42)
Level 3: reference/rules.md    +124 words  (context now 166)
words in scripts/check.py, never loaded: 360

Now change it:

  • Make both dates valid: change the second date to "05/04/2026". Predict which line disappears from the output and the final context size.
  • Change name: date-check to name: Date Check inside the SKILL.md text. Predict what happens and on which line. Why is it helpful that this fails at load time and not during a task?
  • Pretend there is no execution environment: replace the run_script call with result = FILES["date-check/scripts/check.py"], so the model has to read the code instead. Predict the size of the “script output” load. What else is lost besides context space?

Pause and think: The final context holds 166 words, and 124 of them came from one file. If this skill were used a hundred times a day, what one change to the skill would save the most context, and what is the trade-off?

Shorten what gets loaded when a date fails: split rules.md so the date rule sits in its own tiny file, or have the script print the rule that was broken. Then a failed date costs a few words instead of 124. The trade-off is more files or a smarter script to maintain. This is progressive disclosure applied inside Level 3: load the smallest piece that answers the need.

Pause and think: The script is 360 words and the reference file is 124 words, yet only the reference file ever enters the context. Why is that the right way round?

A script is meant to be executed: the agent needs its result, not its text, and the result is 4 words. A reference file is meant to be read: its value is the information itself, so it has to enter the context to be used. That is the reason exact, repeatable work belongs in scripts and explanations belong in reference files.

Summary

An Agent Skill is a folder: SKILL.md with a name and a trigger-style description, the main instructions, and optional scripts and reference files. The agent keeps only the descriptions in mind, loads a skill's body when a task matches, and opens deeper files or runs scripts only when needed. That progressive disclosure is why one agent can carry many skills cheaply. Skills add know-how; MCP adds connections; the two work well together. Treat skills like code: review them, version them, and install only from trusted sources.

Key takeaways

  • An Agent Skill is a folder: SKILL.md (name, description, instructions) plus optional scripts and reference files.
  • The description is the trigger: the model decides to load a skill only from its name and description.
  • Progressive disclosure: metadata always, body on trigger, extra files and scripts only when needed.
  • Scripts make exact steps reliable and keep code out of the context window.
  • Skills add know-how; MCP adds connectivity; use them together.
  • Treat skills as trusted code: review, version and install only from trusted sources.

Key terms

  • Agent Skill: A folder of instructions, scripts and resources that an agent discovers and loads on demand for one kind of task.
  • SKILL.md: The required file of a skill: YAML frontmatter with name and description, then Markdown instructions.
  • Frontmatter: A block of key: value metadata between two --- lines at the top of a Markdown file.
  • Progressive disclosure: Loading information in layers: a short summary first, details only when the task needs them.
  • Context window: The maximum amount of text, in tokens, a model can read in one call.
  • MCP (Model Context Protocol): An open protocol for connecting agents to external tools and data through servers.

← 11.8 Model Context Protocol: A Standard Interface for Agent Tools · 11.10 Open Knowledge Format: Structured Agent-to-Agent Communication →