The Markdown Layer combines published research, platform documentation and a new, deliberately narrow GitHub repository snapshot. These sources answer different questions. We keep their populations and dates separate, identify missing observations, and distinguish measured file presence from interpretations about how people or agents use Markdown.
This methodology accompanies the research hub and the repository study. Version 1.0 is dated September 30, 2026. MD2FILE publishes the report and offers document-conversion products; the research is not an independent product certification, user survey or market-size estimate.
Evidence classes and source review
Published studies retain their authors' populations, dates and measurement definitions. An eligible-repository sample does not automatically describe all GitHub projects. A website listing, a developer announcement and an experiment with an AI agent also measure different things. We use vendor documentation to describe documented behavior and research papers to describe their reported observations or experiments.
The source inventory identifies references used across the series. For a quantitative claim, the relevant record is the original paper, dataset or platform publication wherever available. Access to a paper does not imply that we independently reran its analysis. Where a source dataset could not be retrieved, the distinction remains explicit.
For example, Hora, Montandon and Costa's 2026 repository-content paper uses a random sample within an eligibility-filtered population. Its historical observations are subsets of repositories surviving into its 2026 sample. Our new panel uses current star ordering and one pinned commit per repository. The two designs cannot be pooled into a trend line or a shared adoption percentage.
Conceptual diagrams in the report explain workflows. They are not measured flows or estimates of how many people move through each stage. Product references are examples whose claims require separate verification; they do not contribute observations to the repository study.
Original panel selection
We selected six repositories from each of four GitHub language-filtered searches. The exact query template was:
language:{language} stars:>=1000 pushed:>=2026-01-01 fork:false archived:false is:public
{language} took the values Python, TypeScript, Rust and Go. The REST search parameters were sort=stars, order=desc, per_page=6 and page=1. We retained GitHub's returned order, including any tie order, and required incomplete_results=false. All 24 selected repository identifiers were distinct. See GitHub's repository-search definitions.
These are named, highly starred repositories satisfying the search criteria on September 30, 2026. The sample includes curated lists, educational material and agent-related projects. Its membership is therefore unsuitable for estimating prevalence among ordinary projects, all developers or private repositories. Language filters are not demographic categories.
The search and file-tree collection ran from 13:52:29 to 13:54:49 UTC. For each selected repository, we resolved the current default branch to its latest commit once, retained that commit's tree SHA and requested the recursive tree by that SHA. Stars, language and other repository metadata belong to the search observation; tree contents belong to the pinned commit. They are not a historical time series.
API budget and incomplete observations
The original collection made 53 read-only GET requests: one rate-limit preflight, four repository searches, 24 commit lookups and 24 tree requests. It used GitHub REST API version 2022-11-28, optional existing authentication, a 40-second request timeout and no automatic retries. Tokens were held in memory and were not written to the dataset or logs.
The collector limited each response to 12 MiB, reading one extra byte to detect an oversized body. Three tree responses exceeded that limit: openclaw/openclaw, rust-lang/rust and microsoft/TypeScript. No complete tree was parsed for these repositories. They remain in the selected panel with unknown file counts; the other 21 trees were complete.
The first oversized response stopped the initial collector before its received prefix was saved. We retained the stopped-run record and recorded that request without inventing an HTTP status, full response size or body hash. The bounded continuation reused saved results and collected only untouched repositories. It did not repeat the oversized request or select a replacement.
GitHub also documents its own recursive-tree limits and truncated flag. A client byte ceiling and GitHub's truncation flag are different checks. Here the three unknown trees were stopped by the client ceiling. We do not infer whether their complete server responses would have carried a truncation flag.
File definitions
Counts use regular tracked Git blobs with modes 100644 or 100755. Symlink targets and submodule contents are outside the count. Git LFS pointer size, where present, describes the tracked pointer rather than the linked object's full size.
| Field | Definition |
|---|---|
| Markdown files | Paths ending .md or .markdown, case-insensitive |
| Markdown bytes | Sum of those blobs' reported sizes |
| MDX files | .mdx paths, counted separately |
| Root Markdown README | Root README.md or README.markdown, case-insensitive |
| Human-document names | README, CONTRIBUTING, CHANGELOG, CODE_OF_CONDUCT or SECURITY with either Markdown extension, anywhere, case-insensitive |
| Root docs directory | Exact tree path docs |
| Named context candidate | Exact AGENTS.md, CLAUDE.md or GEMINI.md basename anywhere, or .github/copilot-instructions.md |
| Nested context candidate | A recognized candidate with at least one slash in its path |
The dataset also records SKILL.md, .github/instructions/*.instructions.md and .cursor/rules/*.mdc separately. These names are not added to the headline context-candidate count. The file-path CSV flags several directory names that may indicate tests, examples or vendored material. That flag is a reading aid, not a validated classification of whether a file is active.
A repository can contain several categories. Counting AGENTS.md and CLAUDE.md repositories separately and adding the results would double-count repositories containing both. The dataset includes exact category combinations so that charts can show overlap honestly.
Denominators and missing data
The selected panel has 24 repositories, of which 21 have complete trees. Aggregate file counts and category-presence counts use those 21. JSON null and blank CSV fields mean unknown; a recorded zero in a complete tree means the exact matching rule found no such file.
The 35,904 Markdown files and 152,448,261 source bytes are totals across complete trees. Larger documentation repositories dominate those totals. The study does not estimate prose quality, readership, active use, agent effectiveness, conversion demand or historical growth. No third-party Markdown was rendered as part of this panel.
Downloads and reproduction
The public artifacts contain repository metadata, identifiers, paths and derived counts. They omit third-party file contents, commit-author emails, credentials and private operating notes.
- Repository dataset and definitions, JSON
- One row per selected repository, CSV
- Classified filenames with immutable source links, CSV
- Python reproduction script
The script uses Python 3.10 or newer and its standard library. Its default mode checks the published metadata without making a network request:
python3 reproduce-panel.py --manifest repository-panel.json
To recompute the original complete observations from immutable GitHub trees, use an empty output directory and explicitly enable the network mode:
python3 reproduce-panel.py --manifest repository-panel.json \
--fetch-pinned --output reproduced-panel
This mode makes at most 22 GET requests: one rate check and 21 pinned-tree requests. It preserves the three original unknowns and does not retry failures. Optional credentials can come from GH_TOKEN, GITHUB_TOKEN or an existing GitHub CLI login. A separate offline mode accepts privately saved raw tree responses.
We ran that offline tree-recomputation mode against the original saved responses. All 21 complete trees matched the published derived fields. This verifies the file-count calculation, not the correctness of repository content. Repeating today's search later may select different projects; use the pinned identifiers to reproduce this observation.
Citation, reuse and corrections
Cite the publisher, report title, September 30, 2026 date, version 1.0 and the relevant study or methodology URL. BibTeX and plain-text citation downloads accompany the full PDF and report source.
Repository license metadata is retained as provenance. The download does not relicense underlying repositories, and this release makes no separate license grant for the dataset or scripts. Check the relevant terms before redistributing content or code.
A correction should identify the dataset version, repository, commit and disputed field. Future revisions should record changed definitions or coverage rather than silently replacing an unknown with zero. Any 2027 outlook in the series is editorial interpretation, not a measurement collected in 2027.
Part of The Markdown Layer: State of Markdown 2026. Previous: AI conversations as documents · Research overview · Repository study.
