Skip to main content

Why Your AI Agent Needs SKILL.md: Progressive Disclosure Explained

Calculating read time…

SKILL.md is the file format behind Agent Skills — a folder containing a single Markdown file with YAML frontmatter (a name and a description) plus a body of plain-English instructions, which an AI agent like Claude loads on demand to gain a specific capability it didn't have a moment ago. 🗂️

This matters because the old way of teaching an agent your team's conventions was to cram everything into one always-loaded prompt — and that approach collapses the moment you have more than a handful of workflows to teach it. A always-loaded instructions file that covers forty different situations burns context on the thirty-nine that don't apply to the current task, crowding out the model's actual reasoning room and degrading answer quality as the file grows. SKILL.md solves this with progressive disclosure: an agent can have dozens of skills installed and pay almost nothing for the ones it isn't using right now. Anthropic published the format as an open standard, and it's since been adopted well beyond Claude itself — which means a skill you write once is a skill you can likely reuse across tools, not a one-off dead end tied to a single vendor. 📈

Diagram of the three levels of progressive disclosure in Agent Skills: metadata, instructions, and resources

The three-level loading model that makes SKILL.md context-cheap at scale.

🔀 Quick Comparison

Approach When Loaded Context Cost When Idle Reusable Across Chats? Scales Past ~10 Workflows?
SKILL.md (Agent Skills) On demand, when description matches ~100 tokens (metadata only) Yes, automatically Yes — near-zero marginal cost
Always-loaded instructions file Every single request Full file, every time Yes, but not selectively No — degrades with size
One-off conversation prompt Only in that conversation N/A — not persistent No — must be repeated No — manual repetition
Custom tool / function Declared in every request's tool list Schema tokens per tool Yes Good for actions, not know-how

1. What Is SKILL.md?

A Skill is a folder. Inside that folder sits one required file, SKILL.md, and everything else the skill needs — scripts, templates, reference documents — sits alongside it. The file itself has exactly two parts: a YAML frontmatter block delimited by --- lines, and a Markdown body underneath it. The frontmatter is metadata about the skill; the body is the actual instructions Claude follows once the skill is relevant.

Real, production-grade example — Anthropic's own reference structure

---
name: pdf-processing
description: Extract text and tables from PDF files, fill forms, merge
  documents. Use when working with PDF files or when the user mentions
  PDFs, forms, or document extraction.
---

# PDF Processing

## Quick start

Use pdfplumber to extract text from PDFs:

```python
import pdfplumber
with pdfplumber.open("document.pdf") as pdf:
    text = pdf.pages[0].extract_text()
```

For advanced form filling, see FORMS.md.

Read this the way a new Claude Code user would, one piece at a time:

  1. The frontmatter block between the two --- lines contains only name and description here — the two required fields. This is the only part of the file loaded automatically, for every skill, at the start of a session.
  2. The description does double duty: it tells a human what the skill does and it's the exact text Claude matches against a request to decide whether to trigger the skill at all. Get this wrong and the skill either never fires or fires on the wrong requests.
  3. The body — everything after the closing --- — is ordinary Markdown. It only enters Claude's context once the skill has actually been triggered by a matching request.
  4. The reference to FORMS.md at the bottom is a pointer to a second file in the same folder — Claude reads it only if the current task actually needs form-filling logic.

What breaks without this structure: a team that instead pastes all of this into one giant always-on system prompt pays the token cost of the PDF instructions on every single request, even a request about spreadsheets — and as more workflows get added the same way, the model has less and less room left to actually think.

✅ Worked example: tying this back to the frontmatter above — if a user asks "extract the text from this PDF and summarize it," Claude's system prompt already contains the pdf-processing name and description at near-zero cost. The description's wording ("Use when working with PDF files") is what makes the match, and only then does Claude run cat SKILL.md to load the rest.

🎯 Use this when you have a repeatable workflow — a house style, a checklist, a domain-specific procedure — that you don't want to re-explain in every conversation.

2. Why Skills Exist: The Context Economics Problem

Every token an agent's context window holds is a token it isn't spending on reasoning about your actual request. An always-loaded instructions file — the kind many teams write early on, covering everything from commit-message format to database migration rules — has to be read on every single turn whether it's relevant or not. That's fine at one page. It stops being fine once a team wants the agent to know forty different procedures, because now most of what's loaded on any given turn is irrelevant noise competing for the model's attention.

Agent Skills flip that model. Claude's system prompt only ever holds the lightweight name and description pair for every installed skill — independent measurement of Anthropic's own official skill set puts that median discovery cost at well under 100 tokens per skill. The full instructional body, which can run into the thousands of tokens, is read from the filesystem only after a request actually matches that skill's description. A skill that's never triggered in a given session costs almost nothing, no matter how detailed its instructions are.

💡 Contrast this with a custom tool definition: a function tool's JSON schema is also declared on every request, similar to a skill's metadata — but a tool describes an action Claude can take, while a skill describes know-how: a procedure, a set of conventions, a domain-specific way of working. The two are complementary, not competing — a skill's instructions frequently tell Claude which tool or script to reach for.

🎯 Use this framing when deciding whether something belongs in a skill, an always-on system prompt, or a one-off message: if it's used occasionally and is verbose, it's a skill candidate.

3. Anatomy of a SKILL.md File

The frontmatter accepts a small, defined set of fields. name and description are mandatory; everything else is optional metadata that most skills never need. Anthropic's field rules are specific enough to validate against directly:

  1. name — up to 64 characters, lowercase letters, numbers, and hyphens only. It cannot contain XML tags, and it cannot contain the reserved words "anthropic" or "claude."
  2. description — must be non-empty, up to 1024 characters, and cannot contain XML tags. It has to describe both what the skill does and when Claude should reach for it, in the same sentence or two.
  3. license, metadata, compatibility (open-standard extensions) — optional fields that record authorship, versioning, or hints about required tooling. Most skills skip these entirely.

Real example — full skeleton from Anthropic's own skill structure guidance

---
name: your-skill-name
description: Brief description of what this Skill does and when to use it
---

# Your Skill Name

## Instructions
[Clear, step-by-step guidance for Claude to follow]

## Examples
[Concrete examples of using this Skill]

The body section headings — Instructions and Examples — aren't enforced by the format itself; Markdown gives you total freedom below the frontmatter. But this pattern earns its place: an "Instructions" section gives Claude a clear procedure to follow rather than background trivia to interpret, and an "Examples" section anchors abstract guidance in something concrete the model can pattern-match against.

✅ Worked example: the pdf-processing skill from Section 1 fits this skeleton exactly — its "Quick start" heading plays the role of "Instructions," and the fenced Python snippet plays the role of "Examples," giving Claude both the what and the how in one glance.

🎯 Use this skeleton as your default starting point for any new skill — you can always add sections, but starting minimal keeps the file's Level 2 token cost low.

4. The Three Levels of Progressive Disclosure

Progressive disclosure is the architectural idea that makes SKILL.md work at scale, and it operates in three distinct stages, each with a different loading trigger and a different token budget:

  1. Level 1 — Metadata (always loaded). Just name and description, read into the system prompt at session start. Anthropic's own guidance puts this at roughly 100 tokens per skill — install fifty skills and you're still only a few thousand tokens deep before any of them fire.
  2. Level 2 — Instructions (loaded when triggered). The full SKILL.md body, recommended to stay under 5,000 tokens. This only enters context once a request matches the description closely enough that Claude decides to act on the skill.
  3. Level 3 — Resources and code (loaded as needed). Bundled files — additional Markdown references, templates, executable scripts — that Claude reads or runs only if the Level 2 instructions actually point to them for the task at hand. A skill can bundle a hundred-page API reference and pay zero tokens for it on a task that never touches that reference.
Diagram showing a skill folder anatomy: SKILL.md, FORMS.md, REFERENCE.md, and a scripts folder, and which files load when

Only SKILL.md is guaranteed to load — everything else is opt-in, per task.

The mechanism behind Level 3 is worth understanding precisely, because it's the part beginners most often get wrong: Claude accesses bundled files by running ordinary bash commands — cat for a reference file, direct execution for a script. When a script runs, its source code never enters the context window at all; only whatever it prints to stdout does. That's why a validation script that returns "Validation passed" is dramatically cheaper than asking Claude to write and reason through equivalent validation logic inline, every time.

💡 Key warning: a Level 2 body that balloons past a few thousand tokens defeats the purpose of the format — you've effectively rebuilt an always-loaded file, just gated behind a trigger. If your instructions are getting long, that's the signal to split detail out into a Level 3 reference file and keep SKILL.md itself as a lean table of contents with pointers.

🎯 Use this three-level mental model every time you're deciding where a piece of content belongs inside a skill — SKILL.md for the procedure, separate files for depth.

5. Writing an Effective Description Field

The description is the single highest-leverage sentence in the entire file, because it's the only thing Claude sees before deciding whether the rest of the skill is worth reading at all. A weak description either under-triggers (the skill never fires when it should) or over-triggers (it fires on unrelated requests and wastes a context-loading pass).

Real example — the actual description shipped with Anthropic's pdf-processing sample

description: Extract text and tables from PDF files, fill forms, merge
  documents. Use when working with PDF files or when the user mentions
  PDFs, forms, or document extraction.

Notice the two-part shape: the first clause ("Extract text and tables… merge documents") is a capability summary — what the skill does. The second clause ("Use when…") is a trigger condition — when Claude should reach for it, phrased around the words a user is actually likely to type. Both halves matter; a description with only the first half reads like a product blurb, and Claude has nothing concrete to match a request against.

  1. Lead with concrete verbs describing the capability, not vague category words.
  2. List the specific nouns a user's request is likely to contain — file types, tool names, task words — since that's literally the matching surface.
  3. Keep it under the 1,024-character ceiling, but don't pad it just because you have room; a bloated description doesn't trigger more reliably, it triggers less precisely.

✅ Worked example: "Use when working with PDF files or when the user mentions PDFs, forms, or document extraction" is doing real work — it covers three different phrasings a user might actually type, not just the one the author had in mind while writing it.

🎯 Use this when a skill you wrote "just isn't triggering" — before touching the body, rewrite the description around the exact words a real user would type.

6. Bundling Scripts, References, and Assets

A skill folder can hold three different kinds of Level 3 content, and each one exists for a different reason:

  1. Instructions — additional Markdown files, like FORMS.md or REFERENCE.md, holding guidance too detailed or too situational to belong in the main SKILL.md body.
  2. Code — executable scripts Claude runs through bash for deterministic operations. A script guarantees the same output every time given the same input, which prompted instructions alone can't guarantee, since the model could phrase or execute an equivalent procedure slightly differently on different runs.
  3. Resources — factual reference material: schemas, templates, sample data, or API documentation Claude reads only when a task specifically calls for it.

This is also where the format's biggest practical advantage shows up: because none of this content costs tokens until it's actually accessed, there's effectively no ceiling on how much a skill can bundle. A skill can ship a genuinely comprehensive reference — the full API surface of a library, say — without that size ever penalizing the sessions that don't need it.

💡 Contrast this with inlining logic in the prompt: asking Claude to "generate and reason through" validation code every time is both slower and less reliable than pointing it to a small, already-tested validate_form.py it simply runs. The script's code stays out of context entirely — only its output, like a pass/fail message, consumes tokens.

🎯 Use scripts for anything that must be exactly correct every time — validation, format conversion, calculations — and reserve prose instructions for judgment calls a script can't make.

7. Where Skills Live: Claude Code, claude.ai, and the API

SKILL.md is the same file format everywhere, but installation and sharing differ meaningfully by surface — a distinction that trips up teams trying to standardize on one workflow across a whole organization:

  1. Claude Code — fully filesystem-based. Skills live at ~/.claude/skills/ for personal use or .claude/skills/ for project-scoped skills checked into a repo. No upload step; Claude Code discovers them directly, and they can also be distributed through Claude Code Plugins.
  2. claude.ai — custom skills are uploaded as zip files through Settings, and they belong to the individual user, not the organization; admins currently can't centrally manage or push custom skills org-wide on this surface.
  3. Claude API — skills upload through dedicated Skills API endpoints and are shared workspace-wide, meaning every member of that API workspace can access an uploaded custom skill, unlike the per-user model on claude.ai.

✅ Worked example: a pdf-processing skill you build and test locally in Claude Code, under .claude/skills/pdf-processing/, does not automatically appear in claude.ai or in your API workspace — each surface needs its own separate install of that same folder.

🎯 Use this when planning a skill's rollout — decide up front which surfaces need it, because "build once" still means "install everywhere it's used," not "install once, everywhere."

8. Security Considerations

A skill isn't inert documentation — it's instructions and code that direct an agent's behavior, and Claude generally follows those instructions with the same trust it gives any other input, which means a poorly sourced skill can steer Claude toward actions the user never intended. Anthropic's own guidance is unambiguous: use skills only from sources you created yourself or obtained directly from Anthropic.

💡 Key warning: skills that fetch data from external URLs carry particular risk — fetched content can contain instructions the skill's author never wrote or approved, and even a skill that started trustworthy can be compromised later if an external dependency it relies on changes. A skill you audited once isn't necessarily still safe six months later if it reaches out to the internet.

Before trusting an unfamiliar skill, review every file it bundles — not just SKILL.md, but every script, image, and reference file — and watch specifically for behavior that doesn't match the skill's stated purpose: unexpected network calls, unusual file-access patterns, or logic that reads more broadly than the task needs. Treat installing a skill the same way you'd treat installing any other third-party software with access to your data.

🎯 Use this checklist before installing any skill you didn't author yourself — especially before giving it access to a system with sensitive data or production credentials.

9. Enterprise Rollout at Scale

Standardizing skills across a team is a governance exercise, not just a technical one. A pattern that holds up in practice:

  1. Pick an ownership model per skill. A skill encoding a regulated procedure (financial reporting rules, a legal review checklist) needs a named owner and a review process; a skill that's just a personal shortcut doesn't.
  2. Standardize the skeleton, not the content. Give every author the same starting template — frontmatter fields, an Instructions section, an Examples section — so skills are consistent to review even as their subject matter varies wildly.
  3. Version-control project-scoped skills. Skills under .claude/skills/ live in the repo like any other config, which means code review, git history, and rollback all apply for free.
  4. Match the surface to the sharing need. If a skill needs to reach an entire API workspace automatically, the Claude API's workspace-wide sharing model fits; if it's genuinely personal, claude.ai's per-user model is the right — not the wrong — choice.
  5. Budget Level 2 tokens deliberately. Track SKILL.md body length the way you'd track a function's cyclomatic complexity — a body creeping past the ~5,000-token guidance is a signal to split content into Level 3 reference files, not a cue to write more tersely at the cost of clarity.

✅ Worked example: a platform team ships a shared pdf-processing skill through the Claude API's Skills API so every workspace member gets it automatically, while individual analysts keep small personal skills — a favorite report format, say — installed only in their own claude.ai account, exactly matching each skill's real audience.

🎯 Use this checklist when skills move from one engineer's convenience to the team's shared way of working.

🚫 Common Mistakes

1. Writing a description that only says what, never when. "Handles PDF documents" tells Claude the capability but gives it nothing concrete to match a user's actual wording against — the "Use when…" clause is what makes triggering reliable, and skipping it is the single most common cause of a skill that silently never fires.

2. Stuffing everything into SKILL.md itself. Once the body creeps past a few thousand tokens, every trigger of that skill pays the same heavy cost the format was designed to avoid. Detailed reference material belongs in a separate Level 3 file the instructions point to, not in the main body.

3. Assuming a skill installed on one surface is available everywhere. A skill built in Claude Code doesn't automatically appear on claude.ai or in the API, and a skill uploaded to claude.ai is scoped to that one user, not the organization — each surface needs its own explicit install.

4. Using reserved or invalid characters in the name field. A name over 64 characters, containing uppercase letters or underscores, or containing the words "anthropic" or "claude" fails validation outright — small formatting mistakes here block the whole skill from loading.

5. Installing a skill from an unverified source without auditing it. Because a skill's instructions can direct tool use and code execution, an unreviewed skill from an unknown source carries the same risk as unreviewed third-party software — not a lesser one just because it's "only a Markdown file."

6. Putting logic that must be exactly correct into prose instead of a script. Asking Claude to "carefully calculate" something repeatable, instead of bundling a small deterministic script for it, trades reliability for a token cost you didn't even save — a script's output is usually cheaper than the reasoning trace of reproducing the same logic inline.

❓ FAQ

Do I need to know YAML to write a SKILL.md file?

Only a tiny amount. The frontmatter is just two lines at minimum — name: and description: — wrapped between two --- markers. Everything below that is ordinary Markdown, which most beginners already know.

What's the difference between a Skill and a custom tool?

A tool gives Claude a specific action it can take, declared with a formal schema on every request. A Skill gives Claude know-how — a procedure, conventions, or context — loaded only when relevant. Skills frequently instruct Claude to use particular tools as part of their procedure, so the two work together rather than replacing one another.

How big can a SKILL.md file's bundled resources be?

There's effectively no practical limit on bundled Level 3 content, because those files cost zero tokens until Claude actually reads or runs them. The 5,000-token guidance applies specifically to the SKILL.md body itself, which is the one part guaranteed to load whenever the skill triggers.

Do custom Skills sync automatically between Claude Code, claude.ai, and the API?

No. Custom Skills do not sync across surfaces — a skill uploaded to claude.ai stays on claude.ai, one uploaded through the API isn't visible on claude.ai, and Claude Code's filesystem-based skills are separate from both. Each surface needs its own explicit install of the same folder.

Is it safe to install a Skill someone else shared with me?

Only after auditing it. A Skill can direct Claude to run code and invoke tools, so an unreviewed Skill from an unfamiliar source carries the same risk as any other unreviewed software — check every bundled file for behavior that doesn't match its stated purpose before trusting it, especially if it fetches anything from external URLs.

🔗 References & Further Reading

All product names, trademarks, and registered trademarks (including Claude and Anthropic) are the property of their respective owners and are referenced here for identification and educational purposes only. Every explanation above is original synthesis written to fact-check against the sources listed — no text is reproduced verbatim from any source.

📝 Summary

  • SKILL.md is a folder with a required Markdown file: YAML frontmatter plus a Markdown body of instructions.
  • Skills exist to solve context economics — avoiding the cost of always-loaded instruction files.
  • The frontmatter has two required fields, name and description, each with strict validation rules.
  • Progressive disclosure loads content in three levels: metadata, instructions, and resources.
  • The description field is the highest-leverage sentence in the file — it must state what and when.
  • Bundled scripts, references, and assets cost nothing until a task actually needs them.
  • Skills don't sync automatically across Claude Code, claude.ai, and the API — each surface needs its own install.
  • Skills carry real security weight — audit any skill you didn't author yourself before trusting it.
  • Rolling skills out at scale means owning taxonomy, templates, version control, and token budgets.
  • The most common mistakes trace back to one root cause: treating SKILL.md like a prompt instead of like a loaded-on-demand program.

That's the full anatomy of SKILL.md — from the two required lines of frontmatter to the enterprise rollout checklist that keeps dozens of skills manageable at once. The fastest way to really understand the format is to write one: pick a task you explain to Claude the same way every week, and turn it into a five-line frontmatter block and a short instructions section. 🧩

Comments