Skip to main content
Back to timeline
arXivSource publication:

15-day field study in an industrial C++ repository: after AI mass remediation, CI and review became the main bottlenecks, and directory-based batching with a per-change file cap restored throughput

Synopsis

In a 15-day exploratory single-case field study in a closed-source industrial C++ repository, an experienced developer used a command-line AI coding buddy to remediate widespread issues, triangulating Gerrit metadata with a developer diary and team chat through descriptive statistics and qualitative coding; AI-assisted remediation rapidly generated hundreds of commits touching thousands of lines, saturating build-on-commit CI and reviewer attention, naive per-file commits overloaded CI, and switching to directory-based batching with a cap on files per change restored throughput, yet still required explicit review solicitation, negotiation of acceptable commit granularity, and iterative follow-up to resolve build and static-analysis failures, showing that when mechanical editing becomes che

AI-generated editorial illustration: Orchestrating AI-Assisted Code Remediation: Socio-Technical Bottlenecks in a Large Industrial Repository

Interpretation

The study characterizes how AI-assisted mass code remediation affects real workflows in an industrial repository: AI-assisted remediation rapidly generated hundreds of commits touching thousands of lines, saturating build-on-commit CI and reviewer attention. Prior work on LLM coding assistants often focuses on generation capability or controlled experiments; this work observes a closed-source industrial C++ repository's real collaboration process and measures CI and review as affected parties. A 15-day exploratory single-case field study with one experienced developer using a command-line AI coding buddy, triangulating Gerrit metadata with a developer diary and team chat, analyzed through descriptive statistics and qualitative coding.

The study identifies commit granularity as a central moderating variable: naive per-file commits overloaded build-on-commit CI, while switching to directory-based batching and capping the number of files per change restored throughput. It elevates commit granularity from an engineering habit to a key mechanism for whether AI-assisted remediation is sustainable, and offers actionable batching and file-cap practices. Based on comparative observation of two commit strategies within the same case, supported by Gerrit metadata and the developer diary, as an empirical finding within a single case.

Even after throughput recovered, AI-assisted remediation still required explicit review solicitation, negotiation of acceptable commit granularity, and iterative follow-up to resolve build and static-analysis failures. It shows the bottleneck is not only technical capacity but also socio-technical steps such as review coordination and negotiation, reframing 'remediation done' as an ongoing orchestration process. Qualitative coding of the developer diary and team chat, alongside Gerrit metadata, as single-case qualitative evidence.

The study proposes treating semantic change sets, such as 'fix all instances of warning X', as first-class units of work that can be sliced differently for developers, reviewers, and CI. It redefines the unit of work from files or commits to semantic change sets, offering an organizational and process-level design direction for coordinating AI-assisted remediation in very large repositories. A concluding claim based on case observations, supported by case experience and qualitative analysis, without cross-repository validation.

Perspective

The study targets engineering teams and process owners introducing AI-assisted mechanical remediation in large, long-lived industrial repositories, especially organizations using build-on-commit CI and Gerrit-style review. Its conclusions apply to mechanical, batchable remediation scenarios, such as uniformly handling a class of warnings; in that setting, it suggests deliberately controlling commit, review, and CI batch granularity and treating semantic change sets as units of work that can be sliced separately for developers, reviewers, and CI.

As an exploratory single-case study, its findings come from one closed-source industrial C++ repository, one experienced developer, and a 15-day window, so behavior under other repository sizes, languages, CI configurations, and team structures remains to be observed. In addition, this is abstract-level material and does not include numeric details such as specific commit counts, lines of code, CI queueing, or failure rates, nor quantitative comparisons of different batching strategies, so the strength of these mechanisms across settings remains an open question.

Sources