Markdown Repository Study

Sep 30, 2026

6 min read

MD2FILE Team
Markdown in GitHub Repositories: A 2026 Snapshot

All 21 complete repository trees in MD2FILE's September 30, 2026 snapshot contain Markdown and a root Markdown README. Fifteen contain at least one recognized agent-context filename. We selected 24 repositories; three file trees exceeded our collection limit and remain unknown. These counts describe a small panel selected by GitHub stars, not Markdown adoption across GitHub.

The repository dataset records the exact selection, commit identifiers and missing coverage. This study is part of The Markdown Layer, which examines how structured text moves between documentation, software and AI workflows. MD2FILE publishes the research and sells document tools; no repository in the panel was selected for using an MD2FILE product.

What can a repository tell us about Markdown use?

A repository tree can establish that a file exists at a particular commit. It can show the file's name, location and stored size. Those observations support questions about documentation structure: whether a project has a README, whether instructions appear in several directories, and whether multiple instruction-file conventions coexist.

The same tree cannot establish how often someone reads a document, whether an agent loads it, or whether it helps complete a task. A large collection of Markdown files may be a course, translated documentation, generated material or test data. A filename count needs a definition before it becomes a useful statistic.

Our study records regular tracked files with .md or .markdown extensions, ignoring extension case. MDX is separate because it permits embedded components and belongs to a different rendering contract. We also record several exact context-file names, without assuming that every matching file is active configuration.

How we selected the 24 repositories

We ran four public GitHub repository searches, using Python, TypeScript, Rust and Go as language filters. Each query required at least 1,000 stars, a push on or after January 1, 2026, and a public repository that was neither a fork nor archived. We retained the first six results sorted by descending star count, including repositories whose names or purposes did not fit a conventional application project.

This selection favors highly visible projects. It also includes curated lists and coding-agent software, both of which can affect the results. GitHub's language metadata describes repositories; it does not identify the languages spoken by their contributors or every technology they use. The exact queries and reproduction procedure make those choices inspectable.

Collection ran on September 30 between 13:52 and 13:55 UTC. Each default branch was resolved once to an immutable commit and tree identifier. We retrieved 21 complete trees. The responses for openclaw/openclaw, rust-lang/rust and microsoft/TypeScript exceeded our 12 MiB client limit. We kept all three in the selected panel with unknown counts and made no replacement selection.

README and contribution files remain visible beside agent files

The 21 complete trees contain Markdown files serving several recognizable naming conventions. A repository can contribute to several rows, and most rows allow matches anywhere in its tree.

Counts of document filenames in the 21 complete repository trees; the three incomplete trees are excluded from these counts
Counts of document filenames in the 21 complete repository trees; the three incomplete trees are excluded from these counts

Open chart at full size

Observed filename or directoryComplete repositories containing it
Root README.md or README.markdown21 of 21
CONTRIBUTING.md or .markdown, anywhere19 of 21
SECURITY.md or .markdown, anywhere11 of 21
CODE_OF_CONDUCT.md or .markdown, anywhere10 of 21
CHANGELOG.md or .markdown, anywhere6 of 21
Exact root docs/ directory11 of 21
A recognized context-name candidate15 of 21

Source: processed panel, September 30, 2026. Human-document basenames are case-insensitive. Context-name matching follows the stricter rules below. The counts omit three unknown trees and must not be presented as 24 successful inspections.

A project without the exact docs/ directory can still have extensive documentation elsewhere. Likewise, release notes can live in GitHub releases, a differently named file or another website. Absence from one filename rule is narrower than absence of the underlying practice.

For practical README authoring, the existing GitHub-flavored Markdown guide includes a worked source file and its PDF. That tutorial addresses export behavior; this repository study did not render or test third-party documents.

Context-file conventions overlap

The main context-name rule recognizes case-sensitive AGENTS.md, CLAUDE.md and GEMINI.md basenames anywhere in the tree, plus the exact path .github/copilot-instructions.md. Among the 21 complete trees, 14 contain an AGENTS.md candidate, seven contain CLAUDE.md, two contain GEMINI.md, and one contains the Copilot path.

Seven repositories contain more than one of these categories. Nine contain a recognized candidate below the repository root. These observations make coexistence and directory scope useful questions for documentation maintainers. They do not show whether the instructions agree or which one a particular agent reads.

Mutually exclusive context filename combinations in 21 complete trees; three selected trees remain unknown
Mutually exclusive context filename combinations in 21 complete trees; three selected trees remain unknown

Open chart at full size

The paths expose reasons for caution. Kubernetes contains named candidates inside vendor/; n8n includes candidate files inside a template directory. Other projects have instructions near tests. A name match in one of these locations may describe a dependency, an example or a directory-specific workflow. The classified path CSV retains those locations instead of silently treating every match as a repository-wide policy.

SKILL.md, GitHub instruction files and Cursor .mdc files are separate fields in the dataset. They are not added to the headline context count. The chapter on AI-agent instructions explains why loading rules matter, while maintaining agent documentation addresses ownership and conflicting instructions.

Why a total file count can mislead

The 21 complete trees contain 35,904 Markdown files. Two repositories account for 27,715 of them: freeCodeCamp/freeCodeCamp has 16,810, and nilbuild/developer-roadmap has 10,905. Their size has far more influence on the total than a small project with a handful of Markdown documents.

Those files total 152,448,261 stored Git-blob bytes across the complete panel. Bytes measure source size, including syntax and whatever content the files hold. They are not words read, documents exported or evidence of Markdown's share of all software work.

For that reason, we emphasize repository-level presence and provide the per-repository counts. A researcher interested in documentation volume would need to separate generated content, translations, educational material and duplicated examples. This snapshot makes no such content classification.

How this relates to larger research

Hora, Montandon and Costa's repository-content study, version 2 reports .md files in 9,860 of 10,000 sampled repositories and README.md in 9,532. Their sampling frame contains 116,013 non-fork repositories active in 2026 with at least 100 stars and 100 commits. Those eligibility rules and their random selection differ substantially from our small popularity-based panel.

The paper's historical slices use repositories from the 2026 sample that existed in earlier years. Changes across those slices should be read with that surviving-sample limitation. Our own snapshot supplies no historical comparison at all. Combining the two datasets into a single adoption rate would obscure their different populations and file definitions.

The larger study supports discussion of documentation patterns within its stated sample. Our panel supplies a smaller, inspectable set of current paths and immutable references. Neither dataset measures all private repositories, local notes, AI conversations or the behavior of all developers.

What another snapshot should preserve

A useful follow-up can revisit the same repository identities at new pinned commits on a later date and record additions, removals and renamed files. It should retain the three current unknowns as missing baseline observations. Substituting new popular repositories would answer a different question about a changing leaderboard.

Readers can download the repository CSV and reproduction script. The methodology explains the byte ceiling, matching rules, API request counts and limits on reuse. A correction should identify the repository and commit so that another reader can inspect the same evidence.

Part of The Markdown Layer: State of Markdown 2026. Research overview · Next: AI-agent instructions · Methodology and limitations.

Found this post interesting? Please help us and share it!