What Are Agent Skills? SKILL.md Explained (2026)
Vorec Team · 2026-09-28 · About 8 min read
An AI coding agent is capable in general and uninformed about your team in particular. It does not know your release checklist, your brand rules, or what went wrong the last time someone recorded a product demo. You can paste that context into every conversation, or you can package it once so the agent picks it up when the task calls for it.
That package is an Agent Skill. This post explains what the format is, what the specification actually requires, how agents load skills, and the security caveats the vendors themselves publish — using a real skill we maintain as the worked example.
Checked on 28 September 2026. Every factual claim below links to the primary source it came from. Agent tooling changes quickly, so check the linked pages for anything newer.
What are Agent Skills?
Anthropic's engineering post introducing them (16 October 2025) describes Agent Skills as organized folders of instructions, scripts and resources that agents can discover and load dynamically to perform better at specific tasks. The same post compares a skill to an onboarding guide for a new hire.
In practice, a skill is a directory with a `SKILL.md` file at its root. The open specification shows the layout:
Anthropic's post carries an update note that an open standard for Agent Skills was published on 18 December 2025. The standard's site, agentskills.io, says the format was originally developed by Anthropic and released as an open standard, and it lists a number of agent products as supporting clients. Two we checked in their own documentation: Anthropic's docs cover Skills in Claude Code, and OpenAI's Codex skills page says its skills are a directory with a `SKILL.md` file and build on the open agent skills standard. For any other client, check that client's docs; we have not tested each one.
What goes in SKILL.md
`SKILL.md` is YAML frontmatter followed by Markdown instructions. The specification defines two required fields and several optional ones:
| Field | Required | Constraint (per the spec) |
|---|---|---|
| `name` | Yes | 1–64 characters; lowercase letters, numbers and hyphens; no leading, trailing or consecutive hyphens; must match the parent directory name |
| `description` | Yes | 1–1024 characters; should say what the skill does and when to use it |
| `license` | No | License name or bundled license file |
| `compatibility` | No | Up to 500 characters of environment requirements |
| `metadata` | No | String key–value map for client-specific properties |
| `allowed-tools` | No | Pre-approved tools; marked experimental, support varies by agent |
A minimal valid skill is just this:
Individual products can add rules of their own. Anthropic's Agent Skills overview, for example, says a `name` cannot contain XML tags or the reserved words "anthropic" and "claude". If you target a specific product, read its docs as well as the spec.
How do Agent Skills work? Progressive disclosure
The design idea is that an agent should not read every skill in full at the start of every conversation. The spec calls this progressive disclosure and describes three levels. Exactly when and how each level loads is up to the client; the table is the spec's model, which Anthropic's overview documents for its own products:
| Level | What loads | When | Size guidance (per the spec) |
|---|---|---|---|
| 1. Metadata | `name` + `description` | At startup, for every installed skill | About 100 tokens |
| 2. Instructions | The full `SKILL.md` body | When the agent decides the skill applies | Under 5,000 tokens recommended |
| 3. Resources | Files in `scripts/`, `references/`, `assets/` | Only when needed | — |
Anthropic's overview gives the same breakdown and adds a detail about scripts: when Claude runs a bundled script through bash, the script's code does not enter the context window — only its output does.
Two consequences follow, and both are practical rather than theoretical:
- The description does the routing. Anthropic's overview says the `description` is what Claude matches a request against when deciding whether to trigger a skill. A vague description ("Helps with PDFs") can make discovery less reliable — the agent has less to match against. The spec's own example of a good description names the tasks and the words a user might say.
- Length is a cost you choose. The spec recommends keeping the main `SKILL.md` under 500 lines and moving detailed reference material into separate files, kept one level deep.
A worked example: a skill that records tutorials
We maintain a skill called `record-tutorial` in our public plugin repository. Its frontmatter follows the format above; the exact file revision we refer to here is pinned so the numbers below stay checkable.
The description does what the spec asks: it says what the skill does, then lists the phrases — "tutorial", "demo video", "screencast", "walkthrough" — that should trigger it.
The body tells an agent such as Claude Code or Codex how to plan a recording, drive `vorec run` to capture locally through the Vorec Recorder app, and stop so the user reviews the take before `vorec analyze` uploads it and generates narration. That review boundary is written into the workflow on purpose: an agent should not upload or spend credits on a recording nobody has looked at.
It also illustrates the trade-off in the previous section honestly. At that pinned revision (the current version when we checked on 28 September 2026), our `SKILL.md` is 4,203 lines long — far past the spec's under-500-lines recommendation. The skill tells the agent to read all of it before recording, and its own text attributes past failed recordings to agents skipping sections they judged irrelevant. That is a deliberate choice with a context cost. The spec's alternative — splitting reference material into files the agent loads on demand — would reduce that cost. Treat our skill as an example of the format, not of the recommended length.
For what the recording workflow produces, see the Claude Code plugin guide and using any coding agent to record your app demo.
Where Agent Skills run, and how the environment differs
Anthropic's overview is specific about differences between its own surfaces, and they matter if you expect one skill to behave the same everywhere:
| Surface (Anthropic) | Where custom skills live | Runtime notes in the docs |
|---|---|---|
| Claude Code | `~/.claude/skills/` (personal) or `.claude/skills/` (project); also via plugins | Same network access as any other program on the user's computer |
| Claude API | Uploaded through the Skills API; shared workspace-wide | No network access; no runtime package installation |
| claude.ai | Uploaded as a zip in settings; individual to each user | Network access varies with user and admin settings |
The same page states that custom skills do not sync across surfaces — a skill uploaded to claude.ai is not available through the API, and Claude Code's filesystem skills are separate from both.
Other clients that support the open format document their own locations and behaviour. Read the client's docs — the spec defines the file format, not where each product stores or runs it.
Security: treat a skill like software
A skill is instructions plus, optionally, code the agent may run. Anthropic's overview is direct about what that means: use skills only from trusted sources, and a malicious skill "can direct Claude to invoke tools or execute code in ways that don't match the Skill's stated purpose." It recommends auditing every bundled file, and flags skills that fetch data from external URLs as particularly risky, because fetched content can contain malicious instructions.
A short checklist based on that guidance:
- Read the whole `SKILL.md`, not just the description.
- Open every script. Look for network calls, file access outside the task, and anything the description doesn't mention.
- Be wary of skills that pull instructions from URLs at run time — the content can change after you reviewed it.
- Pin a version, or vendor the skill into your repository, so an upstream change doesn't reach you unreviewed.
Agent Skills vs MCP vs a long prompt
People often search for how these compare. This is our reading of how they divide the work, not a statement from either specification:
| Agent Skill | MCP server | Pasted prompt | |
|---|---|---|---|
| What it adds | Procedural knowledge + optional scripts | Access to an external tool or data source | Instructions for one conversation |
| How the agent finds it | Available-skill metadata; body loaded on activation (per the spec's model) | Tool discovery through the client's MCP connection (client-dependent) | Pasted in, in full |
| Portable across clients | Per the open spec, where clients support it | Per the MCP protocol, where clients support it | Copy and paste |
| Good for | "How we do X here" | "Let the agent reach system Y" | One-off tasks |
They are not exclusive. A skill can also point an agent at other tooling. Ours, for example, describes the recording workflow and tells the agent to call the Vorec CLI — a command-line tool, not an MCP server.
FAQ
What are Claude Skills?
It's the name people often search for. Anthropic's documentation calls them Agent Skills, and the format is the same `SKILL.md` folder described above.
Do I need to know how to code to write a skill?
No. The required parts are a name, a description and Markdown instructions. Scripts are optional.
How many skills can an agent have installed?
The spec doesn't set a number. What it describes is a cost: in its model, every available skill's name and description load at startup, so each adds a little to every conversation.
Are Agent Skills safe?
That depends on more than the files. It depends on the instructions and code inside the skill, on what the agent is allowed to access and run where the skill executes, and on any external content the skill fetches, which can change after you reviewed it. Anthropic's guidance is to use skills only from trusted sources and to audit anything else thoroughly.
Related reading: What Is the Model Context Protocol?.
Want an agent to record your tutorials? Record with Vorec — it has its own macOS recorder, and Claude Code or Codex can drive it through the `record-tutorial` skill, capturing locally for you to review before anything is uploaded — or upload a recording you already have. Vorec drafts narration matched to the workflow it captured and generates the voiceover, so nothing is spoken into a microphone. The same capture can also produce a written step-by-step guide. Start free — 7-day trial, 100 credits, no credit card required. Trial includes up to 3 projects; exports carry a watermark.