Markdown instruction files give coding agents persistent project guidance. They can describe a required test command or a convention that would otherwise need explaining in each task. Research documents their adoption, but does not establish that adding an AGENTS.md file generally improves correctness or reduces costs. The contents, agent and task all need consideration.
This chapter of The Markdown Layer separates three questions: whether a file exists, whether an agent loads it, and whether its instructions help. Those are different measurements.
What is an agent instruction file?
An agent instruction file is a text artifact intended to guide a coding agent's work in a project. The AGENTS.md format project describes a shared location for guidance such as build steps, tests and coding conventions. Markdown provides headings, lists and code blocks that a maintainer can read and edit alongside the code.
A repository might need to explain that a generated directory must be updated through its generator, or that a fast test suite excludes integration tests. Those statements supply working context. They also create maintenance obligations: when a command or directory changes, the instruction may need to change with it.
| Artifact | Question it should answer | Example content |
|---|---|---|
| Human README | What is this project and how do I start? | Purpose, installation, a working example |
| Agent instruction file | What must an agent know while changing this project? | Required checks, scoped conventions, generated-file boundaries |
| Task description | What should change in this request? | Expected behavior and acceptance criteria |
| Historical decision record | Why was a choice made? | Alternatives considered and the dated decision |
These are editorial roles, not a universal directory structure. A small project can keep related material together. The practical test is whether a reader can identify which statement governs the current task.
How common are these files?
Stride Research reported instruction files in 2,463 of 7,370 eligible repositories on September 27, 2026, or 33.4%. Its population was public, non-fork, non-archived GitHub repositories with at least 5,000 stars and a push in the preceding 90 days. Root AGENTS.md, including filename case variants, appeared in 25.3%; root or .claude/CLAUDE.md in 19.2%. A repository could contain both. Study and methodology.
MD2FILE recomputed those counts from the published CSV. That checks the reported arithmetic, not the completeness of the GitHub census or whether every detected file was used. Stride is a commercial publisher, and its population excludes smaller and private repositories. The repository adoption chapter explains why percentages from different samples should not be combined into a global adoption rate.
Do AGENTS.md files improve agent performance?
The answer depends on which outcome is measured. A fast run can produce an incorrect patch; a correct run can spend more time checking constraints. The two studies below therefore need to remain separate.
| Study | Evaluation | Reported result | Limit |
|---|---|---|---|
| Lulla and colleagues, March 2026 revision | 124 small pull-request tasks from 10 repositories; paired Codex runs with and without the existing root file | Median runtime fell 28.64%; median output tokens fell 16.58% | Comprehensive correctness evaluation was outside scope |
| Gloaguen and colleagues, September 29 revision | SWE-bench Lite and CTXbench; multiple model/harness setups, with absent, generated or developer-committed context | No statistically significant success improvement against absent context; generated files increased average inference cost | Python benchmark tasks and the tested configurations |
The efficiency study used gpt-5.2-codex and manually sanity-checked outputs for 50 tasks. Its token reduction concerns output tokens; median total tokens did not fall. It provides a reason to investigate efficiency in a similar workflow, not a promise of lower bills or equivalent code quality.
The revised effectiveness study used 300 SWE-bench Lite tasks and 138 CTXbench tasks. Generated context increased average costs by 20% and 23% respectively. Developer-committed files performed better than generated files, but did not significantly outperform having no file. A nonsignificant difference does not prove that every instruction is useless.
These results do not supply a universal word limit or an ideal template. They suggest that adopting the filename and demonstrating an improvement are separate steps. A team still needs to decide what result would justify keeping a particular instruction.
A file must be loaded before it can influence a task
Discovery rules depend on the agent and version. For example, Anthropic's current documentation says Claude Code can read AGENTS.md directly from version 2.1.277, with default behavior affected by CLAUDE.md files in the project hierarchy. Imports and settings can change that behavior. Verify the actual loaded context before diagnosing an ignored instruction. Claude Code documentation.
Conceptual diagram by MD2FILE. It separates guidance, loading and execution controls; it is not a precedence specification for every agent.
Keep permissions explicit outside the prose. Anthropic describes instruction files as context rather than enforced configuration. A sentence asking an agent to avoid an action should not be the only control preventing that action. Review the tool's permission settings separately from the wording of its instructions. Memory and instruction behavior.
Write an instruction that can be checked
Consider a fictional package whose client code is generated from an API schema. “Keep the client consistent” leaves the expected work unclear. A more useful draft names the source and the verification step:
## Generated API client
Edit api/schema.yaml when changing an endpoint.
Regenerate packages/client with the project's client generator.
Include the schema and generated-client diff in the same review.
Do not describe the change as verified if generation fails.
This example deliberately avoids inventing a command for your repository. Replace the generator reference with a command that actually exists, and run it before publishing the instruction. If regeneration changes unrelated files, investigate that difference instead of teaching the agent to accept it automatically.
The instruction has a defined subject, scope and observable result. A reviewer can check whether the schema changed, whether generation ran and whether the resulting diff belongs to the task. It is still possible for an agent to misunderstand or disregard it, so those checks remain necessary.
Avoid adding a permanent rule for every unusual incident. A temporary migration exception can belong in the migration task. A recurring constraint on all client changes belongs in maintained project guidance. This choice reduces the chance that an old exception will be mistaken for a current requirement.
Evaluate one change to the instructions
Use a small set of representative tasks before expanding the file. Include an ordinary change and a case where the proposed instruction should matter, such as modifying the schema in the example above. Keep the repository revision, agent version, model and permissions the same across comparison runs.
Record each outcome separately:
- Did the patch satisfy the task and pass the relevant tests?
- Did it respect the particular requirement being evaluated?
- What were the elapsed time, input tokens, output tokens and charged cost, where available?
- How much correction or review did a person need to perform?
Repeat tasks where practical and retain failed runs. Report the number of trials and variation, rather than selecting the best result. A single successful demonstration can show that a workflow is possible; it cannot establish a reliable average improvement.
Keep a useful instruction with its evidence and owner. Revise an ambiguous instruction, and remove one whose purpose has expired. The companion chapter on maintaining agent documentation develops that review process, including how to separate current requirements from archived decisions.
Evidence and publisher disclosure
MD2FILE publishes this research series and provides Markdown editing and conversion tools. The cited studies evaluate repository practices and coding agents, not MD2FILE products. We did not rerun their agent experiments. Source versions and sampling limits are recorded in the series methodology.
Previous: Repository adoption evidence · Overview: The Markdown Layer · Next: Maintaining agent documentation
