Skip to main content
Back to timeline
arXivSource publication:

Full development history of a wholly AI-authored codebase released: 14.3% of AI code-generation events in a 21,000-line Python tool contained real errors, and roughly 1 in 4-5 interactive responses contained factual errors

Synopsis

The work releases a new dataset consisting of the full development history of a 21,000-line Python tool built entirely by Claude AI with no human-authored code or tests, together with two code-provenance tracing tools and three taxonomies for instruction intent, commit provenance, and response reliability; applying these to the dataset, it finds that user coding agent CLI instructions differ in kind from IDE-chat instructions with a greater focus on comprehension, planning and consultation, that code development is mainly proactive, that 14.3% of AI code-generation events contain a real error later caught by the AI-authored test suite, and that roughly 1 in 4-5 of the AI's interactive responses contains one or more factual errors.

Source-provided article image: Between the Commits: Process, Error, and Claim Reliability in a Wholly AI-Authored Codebase
Fig. 3 ·

Fig. 3 : Behavioral-intent subcategory distribution: Tang et al.’s published data (left, IDE-chat messages) vs. this study’s Extractor corpus (right). Bars are coloured by top-level category, categories below 1% in both panels are omitted.

arXiv

Interpretation

The work constructs and releases a new dataset recording the full development history of a 21,000-line Python tool built entirely by Claude AI, with no human-authored code or tests. Unlike prior work on AI-assisted programming, this dataset covers the complete development process of a codebase written entirely by AI with no human code involvement, rather than human-AI mixed or fragment-level generation settings. The evidence comes from recording and organizing the full development history, with a dataset scale of 21,000 lines of Python code and an explicit statement that there is no human-authored code or tests.

The work presents two code-provenance tracing tools and three taxonomies for characterizing instruction intent, commit provenance, and response reliability. These tools and taxonomies provide a reusable annotation and analysis framework for systematically studying the process of AI-authored code, rather than a one-off observation. The evidence comes from the authors' presentation of the tools and taxonomies and their application to analyze the dataset.

The analysis finds that user coding agent CLI instructions differ in kind from IDE-chat instructions, with the former focusing more on comprehension, planning and consultation, and that code development is mainly proactive. This finding distinguishes the instruction-intent differences between the two interaction channels and indicates that development is mainly proactive, offering a new empirical description of how AI coding agents are actually used. The evidence comes from applying the instruction-intent taxonomy to the instruction records in the dataset.

14.3% of AI code-generation events contain a real error later caught by the AI-authored test suite, and roughly 1 in 4-5 of the AI's interactive responses contains one or more factual errors. The work quantifies with specific proportions the error rate in a wholly AI-authored codebase and the factual error rate in responses, providing direct empirical data for assessing the reliability of autonomous AI coding. The evidence comes from statistics over the AI code-generation events and interactive responses in the dataset, with errors caught by the AI-authored test suite and response reliability judged by the corresponding taxonomy.

Perspective

The work is aimed at researchers concerned with the process and reliability of autonomous AI coding and at developers using AI coding agents, and it applies to the specific setting of analyzing a codebase written entirely by AI with no human code or tests. Its provenance tools and three taxonomies can be used for process analysis of other similar codebases, and the 14.3% code-generation error rate and the roughly 1 in 4-5 response factual error rate provide comparable baseline references for assessing the reliability of AI coding agents.

The work is based on the development history of a single 21,000-line Python tool, and whether its 14.3% code-generation error rate and roughly 1 in 4-5 response factual error rate apply to other programming languages, other AI models, or other project scales remains to be tested with more data. In addition, errors are caught by the AI-authored test suite, so the coverage of that test suite affects error detection; the annotation consistency of the three taxonomies for instruction intent, commit provenance, and response reliability also warrants further attention.

Sources