Markdown for AI agents: what instruction files can and cannot do

Sep 30, 2026

6 min read

MD2FILE Team
Markdown for AI agents: what instruction files can and cannot do

Markdown instruction files give coding agents persistent project guidance. They can describe a required test command or a convention that would otherwise need explaining in each task. Research documents their adoption, but does not establish that adding an AGENTS.md file generally improves correctness or reduces costs. The contents, agent and task all need consideration.

This chapter of The Markdown Layer separates three questions: whether a file exists, whether an agent loads it, and whether its instructions help. Those are different measurements.

What is an agent instruction file?

An agent instruction file is a text artifact intended to guide a coding agent's work in a project. The AGENTS.md format project describes a shared location for guidance such as build steps, tests and coding conventions. Markdown provides headings, lists and code blocks that a maintainer can read and edit alongside the code.

A repository might need to explain that a generated directory must be updated through its generator, or that a fast test suite excludes integration tests. Those statements supply working context. They also create maintenance obligations: when a command or directory changes, the instruction may need to change with it.

ArtifactQuestion it should answerExample content
Human READMEWhat is this project and how do I start?Purpose, installation, a working example
Agent instruction fileWhat must an agent know while changing this project?Required checks, scoped conventions, generated-file boundaries
Task descriptionWhat should change in this request?Expected behavior and acceptance criteria
Historical decision recordWhy was a choice made?Alternatives considered and the dated decision

These are editorial roles, not a universal directory structure. A small project can keep related material together. The practical test is whether a reader can identify which statement governs the current task.

How common are these files?

Stride Research reported instruction files in 2,463 of 7,370 eligible repositories on September 27, 2026, or 33.4%. Its population was public, non-fork, non-archived GitHub repositories with at least 5,000 stars and a push in the preceding 90 days. Root AGENTS.md, including filename case variants, appeared in 25.3%; root or .claude/CLAUDE.md in 19.2%. A repository could contain both. Study and methodology.

MD2FILE recomputed those counts from the published CSV. That checks the reported arithmetic, not the completeness of the GitHub census or whether every detected file was used. Stride is a commercial publisher, and its population excludes smaller and private repositories. The repository adoption chapter explains why percentages from different samples should not be combined into a global adoption rate.

Do AGENTS.md files improve agent performance?

The answer depends on which outcome is measured. A fast run can produce an incorrect patch; a correct run can spend more time checking constraints. The two studies below therefore need to remain separate.

StudyEvaluationReported resultLimit
Lulla and colleagues, March 2026 revision124 small pull-request tasks from 10 repositories; paired Codex runs with and without the existing root fileMedian runtime fell 28.64%; median output tokens fell 16.58%Comprehensive correctness evaluation was outside scope
Gloaguen and colleagues, September 29 revisionSWE-bench Lite and CTXbench; multiple model/harness setups, with absent, generated or developer-committed contextNo statistically significant success improvement against absent context; generated files increased average inference costPython benchmark tasks and the tested configurations

The efficiency study used gpt-5.2-codex and manually sanity-checked outputs for 50 tasks. Its token reduction concerns output tokens; median total tokens did not fall. It provides a reason to investigate efficiency in a similar workflow, not a promise of lower bills or equivalent code quality.

The revised effectiveness study used 300 SWE-bench Lite tasks and 138 CTXbench tasks. Generated context increased average costs by 20% and 23% respectively. Developer-committed files performed better than generated files, but did not significantly outperform having no file. A nonsignificant difference does not prove that every instruction is useless.

These results do not supply a universal word limit or an ideal template. They suggest that adopting the filename and demonstrating an improvement are separate steps. A team still needs to decide what result would justify keeping a particular instruction.

A file must be loaded before it can influence a task

Discovery rules depend on the agent and version. For example, Anthropic's current documentation says Claude Code can read AGENTS.md directly from version 2.1.277, with default behavior affected by CLAUDE.md files in the project hierarchy. Imports and settings can change that behavior. Verify the actual loaded context before diagnosing an ignored instruction. Claude Code documentation.

Conceptual flow from discovered repository instructions, the user task and higher-level rules to task context, actions, tests and review. Tool permissions separately constrain execution; review leads to revised instructions when a guidance problem is found.
Conceptual flow from discovered repository instructions, the user task and higher-level rules to task context, actions, tests and review. Tool permissions separately constrain execution; review leads to revised instructions when a guidance problem is found.

Open diagram at full size

Conceptual diagram by MD2FILE. It separates guidance, loading and execution controls; it is not a precedence specification for every agent.

Keep permissions explicit outside the prose. Anthropic describes instruction files as context rather than enforced configuration. A sentence asking an agent to avoid an action should not be the only control preventing that action. Review the tool's permission settings separately from the wording of its instructions. Memory and instruction behavior.

Write an instruction that can be checked

Consider a fictional package whose client code is generated from an API schema. “Keep the client consistent” leaves the expected work unclear. A more useful draft names the source and the verification step:

## Generated API client

Edit api/schema.yaml when changing an endpoint.
Regenerate packages/client with the project's client generator.
Include the schema and generated-client diff in the same review.
Do not describe the change as verified if generation fails.

This example deliberately avoids inventing a command for your repository. Replace the generator reference with a command that actually exists, and run it before publishing the instruction. If regeneration changes unrelated files, investigate that difference instead of teaching the agent to accept it automatically.

The instruction has a defined subject, scope and observable result. A reviewer can check whether the schema changed, whether generation ran and whether the resulting diff belongs to the task. It is still possible for an agent to misunderstand or disregard it, so those checks remain necessary.

Avoid adding a permanent rule for every unusual incident. A temporary migration exception can belong in the migration task. A recurring constraint on all client changes belongs in maintained project guidance. This choice reduces the chance that an old exception will be mistaken for a current requirement.

Evaluate one change to the instructions

Use a small set of representative tasks before expanding the file. Include an ordinary change and a case where the proposed instruction should matter, such as modifying the schema in the example above. Keep the repository revision, agent version, model and permissions the same across comparison runs.

Record each outcome separately:

  • Did the patch satisfy the task and pass the relevant tests?
  • Did it respect the particular requirement being evaluated?
  • What were the elapsed time, input tokens, output tokens and charged cost, where available?
  • How much correction or review did a person need to perform?

Repeat tasks where practical and retain failed runs. Report the number of trials and variation, rather than selecting the best result. A single successful demonstration can show that a workflow is possible; it cannot establish a reliable average improvement.

Keep a useful instruction with its evidence and owner. Revise an ambiguous instruction, and remove one whose purpose has expired. The companion chapter on maintaining agent documentation develops that review process, including how to separate current requirements from archived decisions.

Evidence and publisher disclosure

MD2FILE publishes this research series and provides Markdown editing and conversion tools. The cited studies evaluate repository practices and coding agents, not MD2FILE products. We did not rerun their agent experiments. Source versions and sampling limits are recorded in the series methodology.

Previous: Repository adoption evidence · Overview: The Markdown Layer · Next: Maintaining agent documentation

Found this post interesting? Please help us and share it!