Skip to content

Writing Good AGENTS.md / CLAUDE.md

At a glance AGENTS.md has become the cross-tool open standard adopted by 60,000+ open-source projects, and CLAUDE.md is Claude Code's counterpart. This guide gives a structure template you can apply directly, side-by-side good and bad examples, layered scope rules, and maintenance discipline, so the "onboarding document" in your repo genuinely changes the agent's behavior in every session.

Writing Good AGENTS.md / CLAUDE.md ​

If you've used Claude Code, Codex, or Cursor, you've probably felt this frustration: the agent starts every session like a first-day hire—it doesn't know the test command, doesn't know which directories are off-limits, doesn't know the team uses pnpm instead of npm. You correct the same issues over and over in conversation, when they should have lived in a document.

That document is AGENTS.md (or CLAUDE.md). It is the onboarding document your codebase writes for coding agents: everything you'd tell a new colleague that "doesn't fit in the README but must be known before doing work" belongs here. Writing it well is the single highest-leverage step toward turning a coding agent from "a clever intern" into "a reliable team member."

1. What It Is ​

AGENTS.md: the cross-tool open standard ​

AGENTS.md was released in August 2025 as an open format for giving coding agents project-level instructions. It isn't one vendor's private design—according to the official site agents.md, it came out of collaboration across the AI coding tool ecosystem, with participants including OpenAI Codex, Amp, Google's Jules, Cursor, Factory, and others. Adoption was immediate: when OpenAI announced in December 2025 that it was co-founding the Agentic AI Foundation under the Linux Foundation, it disclosed that AGENTS.md had been adopted by over 60,000 open-source projects and agent frameworks; the format is now maintained under that foundation.

The problem it solves is practical: before it, every tool had its own private convention—Claude Code read CLAUDE.md, Cursor read .cursor/rules, Gemini read GEMINI.md. Getting multiple tools to "behave" in the same repo meant maintaining several nearly identical files. AGENTS.md converges these into one shared file: write once, works across agents. For tools without native support, there's usually a fallback configuration (e.g. Gemini CLI can set the context filename to AGENTS.md in .gemini/settings.json, and Aider accepts read: AGENTS.md in .aider.conf.yml).

How it divides work with the README

The README is for humans: project intro, quick start, contribution guide. AGENTS.md is for agents: build steps, test commands, code conventions, safety red lines. The official framing is blunt—"AGENTS.md is the README for agents." Keeping the two separate is deliberate: the README stays lean and human-contributor-facing, while verbose operational detail goes to the agent-specific place.

CLAUDE.md: Claude Code's counterpart ​

CLAUDE.md is the project memory file read by Anthropic's Claude Code. Same mechanism: automatically loaded into context at the start of every session. Notably, as of now Claude Code's official documentation says it reads CLAUDE.md rather than AGENTS.md (the community issue requesting native AGENTS.md support has been open for over seven months without an official response). The common workaround in practice: keep AGENTS.md as the single source of truth and write a single line @AGENTS.md in CLAUDE.md to import it—CLAUDE.md natively supports the @path import syntax.

On how to write CLAUDE.md, Anthropic's engineering blog gives the "golden rule": keep it concise and human-readable, and stresses that the file is worth iterating on like a prompt (tune it). The principle applies fully to AGENTS.md, and everything below covers both.

2. What It Affects: Why This File's Leverage Is So Large ​

Understanding this file's value starts with the agent's context mechanism. If you've read the Context Engineering chapter, you'll remember the core fact: what an agent can "see" in a session determines what it can "get right."

AGENTS.md / CLAUDE.md is special because of where it loads:

Session start
  │
  ├─ system prompt (written by the tool vendor, out of your control)
  ├─ AGENTS.md / CLAUDE.md  ← your file sits here, entering context every session
  │     ├─ root file (always loaded)
  │     ├─ global file (e.g. ~/.claude/CLAUDE.md, applies across projects)
  │     └─ other files pulled in via imports
  ├─ the user's conversation input
  └─ context produced by the agent's tool calls (file contents, command output…)

That means three things:

  1. It shapes the behavioral baseline of every session. Not a one-off instruction for one task, but a persistently effective "team norm." Write in "run pnpm test before committing," and the agent runs it every time it wraps up—the official AGENTS.md FAQ states explicitly that test commands listed in the file will be proactively executed by the agent, with failures fixed before finishing the task.
  2. It spends your context window budget. The flip side of the coin: every token in the file is paid for on every session, squeezing the space needed for real work. A 5,000-word essay-style AGENTS.md loses to a 300-word one that cuts to the bone. Length is itself a quality metric.
  3. Its priority sits below explicit user instructions and above the agent's "default guesses." The official FAQ is clear: on conflict, the AGENTS.md nearest the file being edited wins, and explicit user instructions in the conversation outrank everything. So it's responsible for "defaults," not "iron laws."

In one sentence: this file is your cheapest, most persistent intervention point on agent behavior. No model changes, no code, no harness configuration—just editing one markdown file systematically eliminates a whole class of repeated errors.

3. Structure Template: An AGENTS.md You Can Apply Directly ​

Community practice and official examples converge on this consensus structure: project overview, tech stack, directory conventions, build & test commands, code conventions, off-limits zones, workflow preferences. Below is a complete template, tailored to a pnpm + TypeScript monorepo, ready to adapt:

markdown
# AGENTS.md

## Project Overview
- This is a VitePress-driven documentation site about AI Agents.
- Site source lives in docs/, build output in docs/.vitepress/dist/ (never hand-edit).

## Tech Stack
- Node 20+, pnpm 9 (do not use npm/yarn; the lockfile is pnpm-lock.yaml)
- VitePress 1.x, Vue 3, TypeScript strict mode

## Directory Conventions
- docs/<section>/<page>.md: content pages; section names are fixed; new sections require changing config.mts first
- docs/public/: static assets; images go here, referenced as /xxx.png
- Any new page must also be registered in the sidebar in docs/.vitepress/config.mts

## Common Commands
- Install dependencies: `pnpm install`
- Local dev: `pnpm docs:dev`
- Build: `pnpm docs:build` (must pass before committing)
- Unit tests: `pnpm test`; run a single case: `pnpm vitest run -t "<name>"`

## Code Conventions
- TypeScript strict; no new `any`
- 2-space indent, single quotes, no semicolons
- Markdown content is written in English; keep technical terms in standard form

## Off-Limits (do not do)
- Do not modify theme files under docs/.vitepress/theme/ unless the task explicitly requires it
- Do not commit package-lock.json or yarn.lock
- Do not fabricate URLs, version numbers, or benchmark numbers in content; flag uncertainty instead

## Workflow Preferences
- Run tests before committing code; do not commit with a red test suite
- Commit messages in English, format: `<scope>: <what>` (e.g. `docs: add rag page`)
- PR title format: `[<section>] <title>`
- For changes spanning more than 3 files, post the plan before starting

The judgment calls behind the template matter more than the template itself:

  • Only write what the agent doesn't know or will guess wrong. "This project uses TypeScript" is a fact the agent gets from glancing at package.json—don't write it; "must use pnpm, not npm" is a wrong-guess waiting to happen—worth writing.
  • Commands must be complete and executable. Write pnpm test, not "run the tests." The agent executes literally; vague descriptions only force it to guess.
  • Off-limits zones must be specific to paths and actions. "Be careful" is filler; "do not modify already-applied migration files under migrations/" actually constrains.
  • Workflow preferences should state "definition of done." Things like "build passes before committing" and "post the plan before large changes" define when the agent can consider itself "finished"—the key to less rework.

Generate a first draft, then polish by hand

Mainstream coding agents (Claude Code's /init, Codex, etc.) can scan the repo and generate a first-draft AGENTS.md/CLAUDE.md. That's a good starting point, but generated drafts are usually verbose and off-target—it can see the code, but not your preferences and pain points. The right move: let it generate, then delete and rewrite line by line, keeping only "mistakes the agent has actually made" and "defaults you know it will guess wrong." This file is worth an hour of careful polishing; the payoff is every session thereafter.

4. Good vs Bad: Three Typical Ways to Write It Badly ​

After reading dozens of real AGENTS.md / CLAUDE.md files in the wild, bad writing falls into three buckets. Here are the anti-examples and the rewrites.

Anti-example 1: vague platitudes, all correct nonsense ​

markdown
## Code Quality
- Write high-quality code
- Maintain good code style
- Pay attention to performance and security
- Write necessary comments and documentation

The problem: none of these instructions carries executable information. "High quality" doesn't constrain the agent—it already believes every line it writes is high quality. The only function of such content is burning context budget and possibly diluting the instructions that matter—the more irrelevant content in context, the lower the probability that key instructions get followed.

The rewrite: replace adjectives with verifiable criteria.

markdown
## Code Conventions
- Functions over 40 lines must be split
- All public API inputs validated with zod schemas; validation failures return 400 + error details
- No new dependencies unless the task explicitly requires it; use existing utilities in `src/utils/` first

Anti-example 2: stale or wrong commands ​

markdown
## Testing
- Run `npm run test:unit` to execute unit tests
- Build the project with `make build`

The problem: the project migrated from npm to pnpm and from Makefile to turbo six months ago, and nobody updated the file. The consequence is worse than "not written": the agent faithfully executes npm run test:unit, gets "command not found," and starts improvising—guessing other commands, even inventing its own test script. A stale instruction isn't neutral information; it's active misinformation.

The rewrite: only write commands you've personally verified, and give key commands a "how to verify" trail.

markdown
## Testing (verified 2026-06)
- Full test suite: `pnpm turbo run test` (~2 minutes)
- Single package: `pnpm test --filter @app/server`
- If any of the above errors out, check the scripts field in package.json first; do not invent commands

That last line—"what to do when it errors"—is an underrated trick: giving the agent an explicit failure path keeps it from improvising when things go wrong.

Anti-example 3: prose storytelling, conventions written as a blog ​

markdown
## About This Project's Architecture

We originally chose a microservice architecture in 2023 because the team expected traffic to grow quickly.
Practice later showed a monolith suited our scale better, so 2024 began the long consolidation journey.
We're currently in a transitional state: most business logic has merged into apps/main,
but the orders module remains independently deployed for historical reasons. Speaking of the orders module, it involves…

The problem: the agent needs "what to do now," not "how we got here." Historical narrative has no operational value; after reading it the agent still doesn't know: where should code changes go? Can the orders module be touched? And this style tends to be long and winding—the biggest waster of context budget.

The rewrite: keep only conclusions from history that affect current decisions, one sentence max.

markdown
## Current Architecture
- Code is consolidating from microservices back into a monolith. Write all new business logic in apps/main/; no new standalone services
- apps/orders/ is a legacy independently deployed module: bug fixes only, no new features, migration planned for 2026 Q4
- All database access goes through the repository layer in apps/main/db/; no direct SQL in business code

The common root disease across all three anti-examples: the author treats the file as "a document for humans" instead of "operating instructions for the agent." The test is simple—ask yourself, item by item: if I deleted this line, would the agent's behavior get worse? If the answer is no, delete it.

5. Layering and Scope: Making the Right Instructions Take Effect in the Right Place ​

Real projects rarely run on a single file. Both systems support layering, but the rules differ, and mixing them up is a classic trap.

AGENTS.md's nesting rules ​

The rule is clean: the agent automatically reads the AGENTS.md nearest the file being edited; nearest wins. This means a monorepo can place its own AGENTS.md in every subpackage:

repo/
├── AGENTS.md                  ← global conventions: tech stack, commit rules, off-limits zones
├── apps/
│   ├── web/
│   │   └── AGENTS.md          ← frontend-specific: component conventions, styling approach
│   └── server/
│       └── AGENTS.md          ← backend-specific: DB migration process, API conventions
└── packages/
    └── ui/
        └── AGENTS.md          ← component library-specific: release process

This isn't theoretical design—the official site mentions that OpenAI's own main repository had 88 AGENTS.md files at the time of writing. The layering principle: root files hold what's "true for the whole repo," child files hold what's "only true in this directory." Child files shouldn't repeat root-file content (duplication means maintaining both, which inevitably drifts); write only the increment.

CLAUDE.md's hierarchy and imports ​

Claude Code's system is slightly different, distinguishing several sources:

LevelLocationScope
Global user memory~/.claude/CLAUDE.mdAll your projects: personal preferences, general habits
Project memory<repo>/CLAUDE.mdThe current project, committed to git, shared with the team
Project local memory<repo>/CLAUDE.local.mdThe current project, not committed, personal experimental preferences
Subdirectory filesCLAUDE.md in a subdirectoryLoaded on demand when the agent touches files in that directory

CLAUDE.md also supports the @path import syntax, letting you split rules across multiple files and pull them in, e.g.:

markdown
# CLAUDE.md
@AGENTS.md
@docs/conventions/api-design.md

Imports are a good way to manage large files: the main file stays a lean "skeleton," and detail documents come in on demand. But beware deep nesting—every extra layer of indirection adds risk of "you think the agent read it, but it didn't." In Claude Code you can use the /memory command to check which instruction files actually loaded in the current session; after writing rules, always look—"did it load?" outranks "is it well written?"

Should both files coexist? ​

If the team uses only Claude Code, one CLAUDE.md is enough. If the toolchain is mixed (increasingly common in 2026), the recommended approach: AGENTS.md as the single source of truth, with CLAUDE.md containing just one line @AGENTS.md plus a little Claude Code-specific configuration. Never maintain two content-parallel, independently evolving files—drift is inevitable, and the agent won't tell you which file it read the stale instruction from.

The migration trap for old files

When migrating from old conventions like AGENT.md or .cursorrules, the officially recommended approach is renaming and leaving a symlink (mv AGENT.md AGENTS.md && ln -s AGENTS.md AGENT.md) for backward compatibility. But periodically check whether any tool still reads the old path; symlinks are a transition, not a long-term solution—the long-term solution is all tools converging on AGENTS.md.

6. Maintenance Discipline: This File Is Alive ​

AGENTS.md's biggest failure mode isn't poor writing—it's write once, forget forever. Code evolves, the file doesn't, and three months later it's the "active misinformation" from anti-example 2. Some enforceable disciplines:

  1. Make "update AGENTS.md" a required item on certain PRs. Migrating package managers, changing test commands, restructuring directories, adding off-limits zones—add a checkbox to these PRs' templates: "Does this PR change any facts stated in AGENTS.md?" Spread the maintenance cost across every change instead of saving it up until the file is fully stale.
  2. Every agent "mistake" is an update signal. This is the most important mindset of all. When the agent uses npm again, edits the wrong directory again, forgets to run lint again—don't just correct it in conversation; ask yourself: will this error happen again? If yes, write the correction into AGENTS.md. This file's content should grow mainly from "errors the agent has actually made," not from you armchairing what it might do wrong.
  3. Prune regularly. Read it through each quarter and delete: what models can now infer on their own (as models get stronger, much early-explicit common knowledge becomes redundant), what no longer holds for the project, what never had any effect. A shorter file is usually a better file.
  4. Review it like code. If it's in git and affects every session, it goes through the PR process. If anyone on the team can casually add "I prefer XXX style" without consensus, the file becomes a noise landfill.

Use Evals to Verify the Effect ​

"Did writing AGENTS.md actually help?" shouldn't rest on feelings. The rigorous approach is to treat it as a prompt change and validate with evaluation—the same thinking as the Agent Evaluation chapter:

  • Collect a set of tasks where the agent genuinely failed in this repo ("used the wrong package manager," "finished without running tests," "touched files it shouldn't have") as a regression set.
  • Run several passes with and without AGENTS.md (or before/after changes), comparing task success rates and violation counts.
  • Re-run whenever the file changes substantially. If a change brings no measurable behavioral improvement, roll it back.

This sounds heavy, but a regression set of a dozen or so cases takes half a day to set up, and what you get is evidence—not vibes—for every future edit to this file. For eval implementation details, see Evals in Practice.

7. Generalizing: Beyond Coding Agents ​

The "write an onboarding document for the agent" pattern has long since spilled out of coding. Grasp its essence—externalizing persistent, cross-session behavioral agreements into a readable, versionable file—and you'll see it applies to almost any long-running agent:

  • Personality and voice files. The personal agent framework OpenClaw uses SOUL.md to define the agent's voice: tone, opinions, humor boundaries, default directness. The official advice matches this page exactly—write only what changes behavior, shorter over longer, "don't write a life story or a wall of feelings with no behavioral effect." Alongside SOUL.md there are role-split files like USER.md (who the agent serves).
  • Team convention files. Research and operations agents benefit the same way: a TEAM.md spelling out "output formats, citation rules, what counts as a trustworthy source, which operations need human confirmation" works just like AGENTS.md does in coding.
  • Your own system-prompt asset. If you build agents directly on the API, the AGENTS.md idea maps to the "project context" layer of your system prompt—same discipline: concise, specific, verifiable, evolving with the code. See Prompt Engineering and Context Engineering.

You can even push the pattern to its limit: the harness design discussed in the Claude Code Case Study chapter is, at bottom, systems engineering for "the agent's environment and documentation," and AGENTS.md is just its lightest component. If you're preparing a project portfolio, a carefully maintained AGENTS.md is itself a strong exhibit—it demonstrates your meta-skill of steering agents, which in interviews is far more persuasive than "I've used XX framework"; see Portfolio Projects.

A closing standard: a good AGENTS.md should read like the best engineer on your team spending 15 minutes on an onboarding briefing for a new hire—no filler, all "traps you'll hit if nobody tells you." If your file reads like the company wiki's "About Us" page, rewrite it.

References ​