Skip to main content
Back to timeline
arXivSource publication:

CodePori uses a multi-agent system to automate code generation and surveys practitioners on its usability in real-world software development

Synopsis

The study proceeded in two phases: it built a multi-agent system named CodePori to automate code generation, then used a survey in which participants evaluated the agents' performance on autonomous software development tasks; the results indicate that LLM-based multi-agent systems show potential for autonomous code generation, but successful integration requires addressing specific challenges such as short-term memory limitations, hallucinations, and code smells, and incorporating a practitioner-centric perspective.

Source-provided article image: CodePori: Large-Scale System for Autonomous Software Development Using Multi-Agent Technology
Figure 1 ·

Figure 1: Multi-agent system and survey workflow.

arXiv

Interpretation

The work proposes and builds CodePori, a multi-agent system for automating code generation, as the first phase of a two-phase approach. Compared with prior studies that evaluated agents only on benchmark datasets, this work combines system construction with evaluation oriented toward real-world software development tasks. The abstract explicitly states a two-phase approach whose first phase is the development of the CodePori multi-agent system to automate code generation; architectural details are not given in the abstract.

The study evaluates agent performance through a survey in which participants assessed the agents on autonomous software development tasks. Whereas prior work often measured agents by binary pass-or-fail results on benchmark datasets, this study uses a survey to examine practical applicability. The abstract states that the study conducted a survey with participants evaluating the agents' performance in autonomous software development tasks; participant counts and statistical results are not reported in the abstract.

The results indicate that LLM-based multi-agent systems show potential for autonomous code generation, but successful integration requires addressing specific challenges such as short-term memory limitations, hallucinations, and code smells. The conclusion shifts evaluation emphasis from benchmark scores toward the conditions needed for real-world applicability and stresses a practitioner-centric perspective. The abstract lists short-term memory limitations, hallucinations, and code smells as examples of challenges and notes the need for a practitioner-centric perspective; no quantitative effect sizes are provided.

The study argues for moving beyond standard benchmarks to evaluate real-world applicability and identifies new opportunities for broader adoption in industry and academia. The work treats the evaluation approach itself as an object of discussion, arguing that benchmark results are insufficient to reflect practical usability. The abstract explicitly states the need to move beyond standard benchmarks for evaluating real-world applicability and mentions opportunities for broader adoption; no concrete adoption pathways or data are given.

Perspective

The study targets automated code generation as a real-world software development task and is intended for practitioners and researchers who want to understand how LLM-based multi-agent systems actually perform in autonomous software development. Its two-phase design, first building the CodePori multi-agent system and then evaluating agent performance through a survey, means the conclusions come primarily from participants' assessments of agent performance on autonomous software development tasks rather than from pass-or-fail results on benchmark datasets. The application setting the conclusions point to is broader adoption of such systems in industry and academia, conditional on addressing specific challenges such as short-term memory limitations, hallucinations, and code smells and on incorporating a practitioner-centric perspective.

The abstract does not state the number of survey participants, the sampling approach, the evaluation instrument, or the statistical analysis, nor does it detail CodePori's agent composition and collaboration mechanism, so the robustness of the survey conclusions cannot be judged. It mentions short-term memory limitations, hallucinations, and code smells as examples without indicating how often or how severely these challenges appeared in the evaluation. It argues for moving beyond standard benchmarks to assess real-world applicability but does not present a concrete alternative evaluation design. In addition, the available text consists of the abstract and submission history, with the body's figures and experimental details absent, so reproduction or comparison would still require the full text.

Sources