Skip to main content
Back to timeline
ORBi UMONSSource publication:

GHAgentFiles: A Dataset of Coding Agent File Histories in GitHub Repositories

Synopsis

This work builds and releases the GHAgentFiles dataset together with the open-source extraction tool cofee, identifying 27,717 repositories with coding agent files out of 165,281 candidates and extracting 142,294 context, skill and subagent file histories (401,873 file revisions across 185,795 commits, spanning 22 November 2022 to 1 July 2026), and it presents a preliminary observational analysis showing that coding agent files emerge around early 2025, peak in additions and modifications in early 2026, and that the generic AGENTS.md has become the most popular naming practice.

Source-provided article image: GHAgentFiles: A dataset of coding agent file histories in GitHub repositories
Figure 1

The heatmap of Figure 1 quantifies how file revisions in the

· Page 3

Interpretation

It provides the largest dataset to date of coding agent context file histories, containing 142,294 file histories and 401,873 revisions from 27,717 GitHub repositories, covering six agents: Claude, Codex, GitHub Copilot, Gemini, Windsurf and Cursor. Compared with prior studies (e.g., 453 repositories, 2.3K+ files, 4,768 files, or 4,738 repositories), this dataset expands the number of repositories, file histories and the time span; the authors describe it as the "largest dataset of context files covering the largest period of time". The scale is supported by an explicit extraction pipeline and statistics (Table III lists repositories, histories, commits and revisions per snapshot); the dataset was validated on a random sample of 666 files independently reviewed by three authors, with 92.4% interrater agreement and no "No" responses.

It introduces and open-sources cofee (COntext File Extraction Engine), a CLI tool that automatically identifies and extracts context, skill and subagent files plus their git metadata using glob-like path rules in a .toml configuration. Prior studies largely relied on their own limited datasets; this tool makes the dataset reproducible, extensible and reusable, which the authors say helps by "alleviating the error-prone and time-consuming nature of producing such datasets". The tool's behavior is described concretely through its command form (cofee -c config.toml repo), a configuration example (Listing 1) and traversal rules (first-parent rule, defaulting to git HEAD), and it is released on GitHub and PyPI.

A preliminary observational analysis finds that coding agent files emerge roughly from early 2025, with a sharp peak in additions and modifications in early 2026, more modifications than additions, and infrequent removals or renames; context files grow slowly until June 2025 then accelerate, with another steeper phase in December 2025, skill files take off around November 2025, and subagents are least used. These evolutionary patterns are observed on a dataset far larger than prior work; the authors state the analysis "only scratches the surface" and call for replicating and extending prior studies on it. The findings rest on weekly and monthly count curves in Figures 2, 3 and 4 with exponential regression fits, and are explicitly framed by the authors as a "preliminary observational analysis".

In the latest snapshot, the generic AGENTS.md has become the most popular naming practice (45.7% of all file revisions), Claude is the leading non-generic group (33.7%), Copilot and Cursor are less prevalent, and the most frequent second-level headings in context files are project overview, architecture, testing and commands. This supplies large-scale counting evidence that AGENTS.md is becoming a cross-tool open standard and that context file content centers on project structure, architecture, testing and commands. Based on 111,727 file histories in the latest snapshot and the heading frequency counts in Table IV; the heading analysis covers only second-level headings (present in 86.6% of context files) and is descriptive.

Perspective

The dataset is intended for empirical research on the evolution of coding agent files in collaborative GitHub repositories, suited to settings that need cross-repository and cross-time comparison of adoption and maintenance patterns for context, skill and subagent files; its repository scope is limited to projects retrieved via SEART that were created at least a year ago, are not forks, have at least 300 commits and at least one commit in the last year, and history traversal follows only the main branch under the first-parent rule, so results apply to this class of relatively active, mature public repositories.

Readers should still watch: the file identification configuration is based on manual analysis of six agents' documentation, which may be outdated, incomplete or changed over time, so completeness cannot be guaranteed; the dataset follows only the main branch and first-parent paths, excluding commits on other branches; symbolic links (1,086) are not modified when their target changes and should be treated separately in analyses; the observational analysis is descriptive and does not address causality; and this reading was of incomplete scope, so full figure and table details may affect further verification of specific numbers.

Sources