Beyond AGENTS.md: The Enterprise Markdown Stack Powering Claude, Codex & Google Stitch
The .md file ecosystem is the informal but increasingly standardized set of Markdown files — AGENTS.md, DESIGN.md, SKILL.md, SOUL.md, PLANS.md, and WORKFLOW.md — that engineering teams use to give AI coding agents durable, version-controlled context about behavior, visual identity, procedures, persona, roadmaps, and orchestration. Each file answers a different question an agent silently asks before it acts: what are the rules here, what should this look like, how do I do this specific task, who am I while I do it, what are we building toward, and who else is involved. 🗂️
Why this matters: by mid-2026, AGENTS.md alone is read natively by more than two dozen agent runtimes and sits inside over 60,000 public repositories under Linux Foundation governance — but teams that stop at AGENTS.md are leaving real failure modes on the table. An agent with perfect behavioral rules but no design tokens still ships inconsistent UI. An agent with a flawless SKILL.md but no PLANS.md loses the thread on anything that spans more than one session. The stakes are concrete: rework hours, brand-inconsistent interfaces, and agents that make confident, wrong assumptions because nobody told them what "done" looks like. ⚙️
📑 In This Post
- A 60-second recap: what AGENTS.md still doesn't cover
- DESIGN.md — visual identity agents can't hallucinate their way around
- SKILL.md — procedural knowledge that loads only when needed
- SOUL.md — persona, voice, and why it's also a security surface
- PLANS.md — the execution roadmap for work that outlives one session
- WORKFLOW.md — orchestrating more than one agent at a time
- Bringing it together: how the files load in one real repo
- Enterprise rollout at scale
- Common mistakes (and the reasoning behind each one)
- FAQ
- References & further reading
- Summary
🔀 Quick Comparison
| File | Answers | Loads when | Notable real-world driver |
|---|---|---|---|
| AGENTS.md | What are the rules here? | Session/repo start | Agentic AI Foundation / Linux Foundation |
| DESIGN.md | What should this look like? | Any UI-generation task | Google Labs / Stitch |
| SKILL.md | How do I do this task? | Description match, on demand | Anthropic Agent Skills |
| SOUL.md | Who am I while doing it? | Every session, first in prompt order | OpenClaw / Hermes Agent |
| PLANS.md | What are we building toward? | Multi-hour / multi-session work | OpenAI Codex Execution Plans |
| WORKFLOW.md | Who else is involved, and in what order? | Multi-agent handoffs | Emerging orchestration frameworks (no single steward yet) |
1. A 60-Second Recap: What AGENTS.md Still Doesn't Cover
AGENTS.md has effectively won the "one behavioral file per repo" argument. It is now governed by the Agentic AI Foundation under the Linux Foundation and read natively by dozens of runtimes — Cursor, Codex CLI, Aider, Devin, Copilot, Gemini CLI, Windsurf, Amazon Q, and more. The one confirmed holdout is Claude Code, which reads its own CLAUDE.md but is commonly configured to import AGENTS.md as its single source of truth via an @AGENTS.md reference or a symlink.
my-project/
├── AGENTS.md # cross-tool shared rules
├── CLAUDE.md # thin file: "@AGENTS.md" + Claude-only config
├── frontend/
│ └── AGENTS.md # nearest-file-wins override
└── backend/
└── AGENTS.md
What AGENTS.md was never designed to hold is the point of this whole ecosystem. It's a README for agents, covering build commands, test commands, and coding conventions — not brand color hex values, not a step-by-step PDF-merging procedure, not a 300-step migration plan, and not which of three agents should hand off to which other agent next. Trying to cram all of that into one file is exactly the anti-pattern researchers keep flagging: bloated instruction files compete for the same attention budget and measurably increase inference cost without improving task success. Splitting concerns across purpose-built files is what keeps each one short, current, and actually followed.
💡 Key warning: Research comparing developer-written AGENTS.md files against LLM-auto-generated ones found the auto-generated versions reduced task success in the majority of tested scenarios and added multiple extra steps per task — largely because they duplicated information the agent already had from the codebase itself. The same discipline applies to every file below: write what the agent cannot infer, not what it can.
🎯 Use this when: you're deciding whether a new instruction belongs in AGENTS.md at all — if it's about visual output, a repeatable multi-step task, tone of voice, a long-horizon plan, or coordination between agents, it belongs in one of the five files below instead.
2. DESIGN.md — Visual Identity Agents Can't Hallucinate Their Way Around
Without design guidance, coding agents default to statistically common training-data patterns: generic blues, default system fonts, and uniform spacing with no intentional rhythm. DESIGN.md exists to close that gap by giving agents exact, enforceable values instead of vague prose like "make it look modern."
Real example — Google Labs' DESIGN.md specification. In April 2026, Google Labs published DESIGN.md as the format specification behind Stitch, Google's AI-driven UI generation product. A DESIGN.md file has two layers in a single document: YAML front matter holding machine-readable design tokens, and a Markdown body underneath explaining the design rationale in plain language.
---
name: Heritage
colors:
primary: "#1A1C1E"
secondary: "#6C7278"
tertiary: "#B8422E"
neutral: "#F7F5F2"
typography:
h1:
fontFamily: Public Sans
fontSize: 3rem
body-md:
fontFamily: Public Sans
fontSize: 1rem
rounded:
sm: 4px
md: 8px
spacing:
sm: 8px
md: 16px
---
## Overview
Architectural minimalism meets journalistic gravitas...
## Colors
The palette is rooted in high-contrast neutrals with
a single accent color reserved for calls to action.
Line-by-line, why enterprises include this: the YAML block is deliberately the part an agent "can't argue with" — primary: "#1A1C1E" is a fact, not a suggestion, the way a natural-language color description would be. The Markdown prose underneath isn't decorative; it's what stops the agent from technically matching the tokens while violating the intent behind them — for instance using the accent color on a decorative element instead of reserving it for calls to action. Google ships an official CLI, npx @google/design.md lint, plus export commands that turn the same file into a Tailwind theme or a CSS variable block, so the design file and the shipped stylesheet never drift apart.
What breaks in production without it: every new screen an agent generates picks its own "reasonable-looking" palette and spacing scale, so a five-screen feature ships looking like five different products. Design review becomes a permanent bottleneck instead of an occasional check, because there's no shared source of truth to review against.
✅ Worked example: a fintech team documents its dashboard's dark-mode tokens once in DESIGN.md, referenced automatically from CLAUDE.md and from .cursor/rules. Every new component — whether generated by Cursor on a Tuesday or Claude Code on a Friday — pulls the same border-radius scale and the same neutral-vs-accent logic, without either tool needing a repeated prompt.
3. SKILL.md — Procedural Knowledge That Loads Only When Needed
AGENTS.md tells an agent about a project. SKILL.md tells an agent how to do something — and, critically, that capability travels with the agent across any project, the way a colleague's expertise doesn't disappear when they change teams.
Real example — Anthropic's Agent Skills. A Skill is a folder containing a SKILL.md file with exactly two required YAML frontmatter fields, name and description, followed by a Markdown body of instructions:
--- name: pdf-processing description: Extract text and tables from PDF files, fill forms, merge documents. Use when working with PDF files or when the user mentions PDFs, forms, or extraction. --- # PDF Processing ## Guidelines 1. Check whether the target is scanned or text-based 2. Use the extraction library for text-based PDFs 3. Fall back to OCR only for scanned pages ...
Why this line-by-line design matters in production: at startup, the agent only loads the name and description field from every installed Skill — roughly a hundred tokens each — into the system prompt, building a lightweight catalog. The full body, and any bundled scripts or reference files, load only when the description matches the current request. This three-tier progressive disclosure is what lets a team install hundreds of Skills without paying a context-window tax for the ones not in use on a given task. The description field carries the real weight: it is the only thing the agent matches against to decide whether to trigger the Skill at all, so a vague description means a Skill that never fires, and an overly broad one means a Skill that fires when it shouldn't.
What breaks in production without it: teams re-explain the same multi-step procedure — say, a company's specific PDF form-filling process — in every single prompt, in slightly different wording each time, which produces slightly different (and sometimes wrong) results each time. SKILL.md turns a one-off instruction into a certified, reusable procedure that behaves the same way whether it's Claude Code, Cursor, or Codex CLI invoking it, since the format's two required fields are honored across every major runtime.
💡 Contrasting example / key warning: unlike AGENTS.md, several SKILL.md frontmatter fields — allowed-tools, context: fork, and hooks — are enforced by Claude Code and largely ignored elsewhere. Rule of thumb: if a skill only uses name, description, and plain Markdown, it is portable across runtimes; anything beyond that is a Claude Code-specific enhancement layered on top of a shared format, not a universal guarantee.
🎯 Use this when: the same multi-step procedure — PDF processing, a deployment checklist, a data-cleaning routine — is being explained fresh in every conversation instead of being packaged once and reused everywhere.
4. SOUL.md — Persona, Voice, and Why It's Also a Security Surface
SOUL.md is the file that answers a question none of the others do: not what the agent should do, but who it should be while doing it — tone, formality, how it handles disagreement, and its base level of caution.
Real example — OpenClaw and Hermes Agent's persona architecture. Both frameworks load SOUL.md first in prompt order, ahead of tool guidance, memory snapshots, and project context, so identity anchors everything that follows rather than competing with it for attention late in the prompt. In OpenClaw specifically, SOUL.md defines behavioral philosophy while separate files — IDENTITY.md, USER.md — hold factual metadata and user-specific context, keeping "who I am" cleanly split from "what I know about you."
# SOUL.md - Who You Are ## Voice You are direct. You prefer short answers over long ones. You explain your reasoning briefly before doing anything irreversible. You write in a measured, unhurried register. ## Authority Bounds [What you cannot commit to without escalation] ## Behavioral Examples <example>...</example> ## Guardrails [Confidentiality, spotlighting, non-modulation]
Why this matters in production, and what breaks without it: persona consistency degrades measurably over long conversations — a phenomenon researchers call persona drift — and SOUL.md's fixed position early in the system prompt is a deliberate re-anchoring technique to fight that drift. Skip it, and a support agent that opens formally can end a 40-turn conversation sounding like a completely different assistant, with no single file a reviewer can check to see what "in character" was even supposed to mean.
💡 Key warning: academic testing across 162 personas found that persona framing largely does not improve factual accuracy — SOUL.md's value case is brand consistency and traceable tone, not raw capability. More seriously, adversarial "persona modulation" attacks have been shown to cut refusal rates substantially by pressuring the model to over-commit to an assigned character. Treat SOUL.md as a tone-and-brand layer that sits on top of the model's safety behavior, never as a mechanism that overrides it — real security still requires architectural controls like tool restrictions and read-only mounts, not persona instructions.
🎯 Use this when: the same agent is deployed across multiple surfaces — a Slack bot, a support widget, an internal CLI — and needs one consistent voice a brand or support lead can review in a single file, separate from the project rules in AGENTS.md.
5. PLANS.md — The Execution Roadmap for Work That Outlives One Session
AGENTS.md and SKILL.md both assume the agent already knows roughly what "done" looks like. PLANS.md exists for the opposite case: multi-hour or multi-session work where the destination has to be specified up front, because there's no memory of prior sessions to fall back on.
Real example — OpenAI Codex Execution Plans (ExecPlans). Codex's PLANS.md convention treats the reader — the agent — as a complete beginner to the repository who has only the current working tree and a single plan file, with no external context and no memory of anything that came before. The plan itself is written to be followed milestone by milestone, without the agent needing to stop and ask what to do next at every step.
# Migration: Legacy Auth -> OAuth2 (ExecPlan) ## Milestone 1: Parallel auth paths - [ ] Add OAuth2 provider config - [ ] Route 5% of traffic through new path - [ ] Verify session parity in logs ## Milestone 2: Cutover - [ ] Increase traffic to 100% - [ ] Remove legacy auth middleware - [ ] Delete deprecated endpoints ## Milestone 3: Cleanup - [ ] Archive legacy auth tests - [ ] Update AGENTS.md auth section
Line-by-line, why enterprises include this: each milestone is self-contained and checkable, so an agent picking up the plan mid-way — whether that's the same agent three days later or a different agent entirely — can verify what's actually been done from the state of the repository rather than trusting a stale conversation history. The instruction to "not prompt the user for next steps" and simply proceed to the next milestone is what turns a plan into something an agent executes rather than a document it merely reads.
What breaks in production without it: long-running work stalls out every time context resets, because nothing in the repo records intent — only the diff of what happened, not why, and not what comes next. Teams end up re-deriving the plan from git history, which is slow and error-prone for anything with more than a handful of steps.
✅ Worked example: a platform team migrating a payments service to a new auth provider keeps PLANS.md in the repo root, checked into version control alongside the code it describes. Three separate agent sessions, spread across a week, each start by reading PLANS.md, checking off completed milestones against the actual codebase, and continuing from the first unchecked item — no human has to re-explain the migration goal each time.
🎯 Use this when: a task is too large for one session — a migration, a large refactor, a multi-week feature — and you need the "why" and "what's next" to survive a context reset the same way the code itself survives a git commit.
6. WORKFLOW.md — Orchestrating More Than One Agent at a Time
WORKFLOW.md is the newest and least standardized file in this ecosystem — there is no single steward comparable to the Linux Foundation's role for AGENTS.md, and conventions vary by orchestration framework. What's consistent across implementations is the underlying job: declaring which specialized agents exist, what each one is responsible for, and the order or conditions under which control passes between them.
Representative pattern, seen across current multi-agent orchestration setups:
.claude/workflows/release-prep/ ├── definition.md # objectives + workflow_id ├── workflow.md # phases + agent assignments └── start.md # executable entry-point prompt ## Phase 1 — Requirements (orchestrator agent) Goal: clarify scope, write docs/requirements.md ## Phase 2 — Implementation (backend agent) Input: docs/requirements.md Goal: write docs/plan.md, then implement ## Phase 3 — Verification (validator agent) Input: implementation diff Goal: run tests, report pass/fail to orchestrator
Why enterprises include this, and what breaks without it: the moment a task needs more than one specialized agent — a planner, an executor, a validator — someone has to define the handoff contract: what one agent produces, what the next one consumes, and what happens on failure. Without a written WORKFLOW.md, that contract lives only in the orchestrator's prompt for a single run and has to be reconstructed by hand for the next one, which is exactly the kind of undocumented, unreviewable coordination that makes multi-agent systems fragile in practice.
💡 Key warning: because this file type is still emerging, don't assume portability across tools the way you reasonably can with AGENTS.md or SKILL.md. Treat WORKFLOW.md as an internal, team-owned convention — document your own schema clearly at the top of the file — rather than expecting an external runtime to parse it out of the box.
🎯 Use this when: a single task routinely gets handed between more than one agent role (planning, coding, review) and the handoff logic needs to be reviewable and repeatable, not re-improvised in every orchestrator prompt.
7. Bringing It Together: How the Files Load in One Real Repo
None of these files compete with each other — they load on different triggers, so a well-structured repository can carry all six without any single one becoming bloated:
- Session start: SOUL.md (if used) loads first to anchor identity, followed by AGENTS.md for project rules.
- Any UI-generation request: DESIGN.md tokens load, referenced from AGENTS.md or a tool-specific rules file so there's exactly one source of truth.
- A recognized task pattern: the matching SKILL.md loads on demand, pulling in its bundled scripts only if actually invoked.
- Multi-session work: PLANS.md is read and updated at the start and end of each session, independent of the length of any single conversation.
- Multi-agent tasks: WORKFLOW.md is read once by the orchestrator to determine which specialized agent — and by extension, which of the other files — applies to the current phase.
The practical takeaway: keep AGENTS.md thin and referential rather than encyclopedic. If a rule is really about colors, procedures, tone, milestones, or coordination, it belongs in one of the other five files, linked from AGENTS.md rather than duplicated inside it.
8. Enterprise Rollout at Scale
Individually adopting one of these files is easy. Making six of them behave consistently across dozens of repositories and hundreds of engineers requires actual governance:
- Ownership: assign a named owner per file type — typically a staff engineer for AGENTS.md, a design systems lead for DESIGN.md, and a platform or DevEx team for SKILL.md and WORKFLOW.md. Ambiguous ownership is why these files rot within a quarter.
- Templates: ship a starter template for each file type in your internal scaffolding tool so new repos begin compliant instead of retrofitted later.
- CI enforcement: lint DESIGN.md token schemas and SKILL.md required frontmatter fields in CI, the same way you'd lint code style; a malformed frontmatter field silently breaks discovery rather than throwing a visible error.
- Size budgets: keep AGENTS.md under roughly 150 lines per current best-practice research, and treat any file approaching a hard technical limit (Codex enforces a 32 KiB cap on AGENTS.md, for instance) as a signal to split, not compress.
- Metrics: track agent task success rate and rework hours before and after rollout per repository; a file nobody is measuring is a file nobody will maintain.
9. Common Mistakes
- Auto-generating AGENTS.md wholesale from an LLM prompt. Research shows this can reduce task success compared to developer-written files, because auto-generated content tends to duplicate information the agent could already infer from the codebase, adding cost without adding signal. Use auto-generation only for narrow, governed data sections pulled from an API, not the whole file.
- Treating SOUL.md as a jailbreak-resistant safety layer. It isn't one, and adversarial persona-modulation research shows exactly why: persona instructions can be pressured into overriding intended behavior. Keep real safety controls architectural — tool permissions, not prompt text.
- Letting DESIGN.md drift from the shipped stylesheet. Without an export or lint step wired into CI, the file becomes documentation nobody trusts within a few sprints. Export tokens programmatically rather than hand-copying hex values into both places.
- Writing a vague SKILL.md description. Because the description field is the only thing the agent matches against to decide whether to trigger a Skill, a generic description means the Skill either never fires or fires on the wrong tasks. Be as specific about "when to use this" as about "what this does."
- Skipping PLANS.md for anything that "should just take a day." Multi-day work has a way of becoming multi-week work, and by the time a plan is obviously needed, reconstructing intent from commit history is far more expensive than writing three milestones up front would have been.
❓ FAQ
Do I need all six files, or can I start with just one or two?
Start with AGENTS.md — it's the foundation almost every runtime reads. Add DESIGN.md the moment agents start generating UI, and SKILL.md the moment you catch yourself re-explaining the same procedure across conversations. SOUL.md, PLANS.md, and WORKFLOW.md are situational: add them when persona consistency, multi-session work, or multi-agent handoffs actually become a problem, not preemptively.
Does Claude Code read AGENTS.md directly?
Claude Code natively reads CLAUDE.md, not AGENTS.md. The common bridge pattern is a thin CLAUDE.md that imports AGENTS.md as its single source of truth, keeping only Claude-specific configuration — permissions, MCP servers — in CLAUDE.md itself.
Is DESIGN.md tied specifically to Google Stitch, or is it a general standard?
The specification originated with Google Labs as the format behind Stitch, but it's published as an open format with a public CLI validator and export tooling, and community catalogs of DESIGN.md files modeled on well-known brands have grown independently of Stitch itself.
Can SKILL.md scripts run arbitrary code?
A Skill can bundle executable scripts alongside SKILL.md, and the frontmatter's allowed-tools field restricts what tools the skill may use when it activates — but that field is enforced in some runtimes (like Claude Code) and silently ignored in others, so treat it as a permission boundary only where it's actually honored, and review bundled scripts before trusting them in a new environment.
Is WORKFLOW.md an official standard like AGENTS.md?
No — as of mid-2026 it's a recurring pattern across different orchestration frameworks rather than a single governed specification. Treat any WORKFLOW.md schema as project- or framework-specific until (and unless) a comparable neutral steward emerges for it.
🔗 References & Further Reading
- AGENTS.md open specification — agents.md
- Google Labs, DESIGN.md format specification — github.com/google-labs-code/design.md
- Anthropic, Agent Skills overview — platform.claude.com/docs
- Anthropic, public Agent Skills repository — github.com/anthropics/skills
- OpenAI Cookbook, Codex Execution Plans (PLANS.md) — developers.openai.com/cookbook
- soul.md portable persona specification — github.com/rokoss21/soul.md
All product names, trademarks, and specifications referenced above (including AGENTS.md, DESIGN.md, Google Stitch, and Claude) belong to their respective owners. This article synthesizes and explains publicly available information in original wording; it does not reproduce source text verbatim.
📝 Summary
- AGENTS.md handles behavioral rules but was never meant to hold visual, procedural, persona, planning, or orchestration detail.
- DESIGN.md gives agents enforceable design tokens plus the rationale behind them, closing the gap that produces inconsistent UI.
- SKILL.md packages reusable, portable procedures that load only when their description matches the task at hand.
- SOUL.md anchors persona and tone across long conversations, but is a brand tool, not a safety mechanism.
- PLANS.md carries execution intent across sessions the way version control carries code.
- WORKFLOW.md documents multi-agent handoffs, though it remains the least standardized file in the set.
- At scale, none of this works without named ownership, CI enforcement, and size discipline per file.
That's the ecosystem beyond AGENTS.md — six files, six different jobs, one shared goal of giving agents context they don't have to guess at. Happy building!
Comments
Post a Comment