Skip to main content
Back to timeline
arXivSource publication:

GPT models change more lines than human patches when fixing Codeforces submissions, and solve more problems from scratch than by patching

Synopsis

This work collects roughly 3000 submissions from a couple of Codeforces users, pairs each buggy submission with its corresponding human fix, uses the similarity between the buggy solution and the human fix as a baseline, evaluates the quality of LLM-generated bug fixes on three OpenAI GPT models (gpt-5-nano, gpt-5-mini, gpt-5.1), and checks whether generated solutions solve the problem using the Codeforces-R1 dataset; the findings suggest LLMs tend to modify more lines than human fixes and sometimes generate entirely new solutions, and that LLMs solve more problems correctly when allowed to generate from scratch rather than patch buggy submissions, even when those submissions are close to the human patch.

Source-provided article image: Large Language Models for Programming: Actually Fixing or Reimplementing Incorrect Code?
Fig. 1 ·

Fig. 1: For a given user u u and problem p p , we order the submissions in descending order. There are two anchors that are related to 3, respectively 1 buggy submission.

arXiv

Interpretation

The paper places problem solving and bug fixing in a single framework, examining how far LLM fixes deviate from human patches and whether there is a bias toward generating entirely new solutions. Prior work evaluated LLM performance in problem solving or bug fixing independently and did not explore the relationship between these two capabilities; this work links them and uses the similarity between the buggy solution and the human fix as a baseline for measuring deviation. Built on roughly 3000 submissions from a couple of Codeforces users, with each buggy submission matched to its corresponding human fix, forming comparable paired data.

Across three OpenAI GPT models (gpt-5-nano, gpt-5-mini, gpt-5.1), LLM-generated fixes tend to modify more lines than human fixes, and in some cases generate entirely new solutions. This observation turns the question of whether LLMs actually fix bugs from qualitative discussion into a measurable deviation comparison, indicating a bias toward rewriting rather than minimal edits. Quality of LLM-generated fixes is evaluated against a baseline of similarity between the buggy solution and the human fix.

When allowed to generate solutions from scratch, LLMs solve more problems correctly than when patching buggy submissions, even when those submissions are close to the human patch. The result reveals a capability gap between bug fixing and from-scratch solving, showing that being close to correct does not mean the model will follow the original approach with incremental edits. Correctness of generated solutions is checked using the Codeforces-R1 dataset, whose tests were generated with the DeepSeek-R1 model.

The paper states that these findings have direct implications for the design of AI-assisted programming tools, particularly in supporting user debugging and promoting incremental problem-solving rather than solution replacement. It translates observations of model behavior into a design orientation for tools, emphasizing debugging support and incremental strategies over replacing the whole solution. Conclusions are supported jointly by the deviation measures on the paired data and the correctness checks on Codeforces-R1.

Perspective

The work targets competitive programming, uses roughly 3000 submissions from a couple of Codeforces users, evaluates three OpenAI GPT models (gpt-5-nano, gpt-5-mini, gpt-5.1), and relies on the Codeforces-R1 dataset for correctness checks. It enables researchers and tool designers to compare model behavior using deviation from human patches and the success gap between from-scratch generation and patching, and to design AI-assisted programming tools that support debugging and encourage incremental solving.

Readers should note that the data come from only two Codeforces users, limiting representativeness; correctness checks depend on the Codeforces-R1 dataset and its DeepSeek-R1-generated tests; and evaluation covers only three OpenAI GPT models. In addition, the loaded text is at the abstract level and does not include specific figures or per-item statistics, so the exact magnitude of line changes, differences among models, and the success-rate gap between from-scratch generation and patching cannot be quantified here; these are open questions to confirm when reading the full paper.

Sources