Skip to main content
Back to timeline
arXivSource publication:

An empirical analysis of 2,042 Python repositories finds 11.2% show OS-dependent test failures, yielding a 7-category taxonomy and an LLM repair evaluation

Synopsis

This study presents a large-scale empirical analysis of cross-OS portability issues across 2,042 open-source Python repositories, combining cross-OS test reexecution (500 projects, 11.2% showing OS-dependent test failures) with manual analysis of 240 GitHub issues (confirming 102 genuine portability problems across 95 additional projects), and develops a taxonomy of 7 primary categories, 24 sub-categories, 15 diagnostic signatures, and 4 systematic repair patterns, while evaluating existing static analysis tools as providing minimal support and large language models achieving 40-79% accuracy in identifying issues and 50-77% success in generating fixes under structured guidance, with practical applicability demonstrated through 33 contributed pull requests (17 merged, zero rejected).

Source-provided article image: An Empirical Analysis of Cross-OS Portability Issues in Python Projects
arXiv · Page 3

Interpretation

The study provides the first large-scale empirical baseline for cross-OS portability issues in Python: cross-OS test reexecution covered 500 projects, of which 11.2% exhibited OS-dependent test failures; manual analysis of 240 GitHub issues confirmed 102 genuine portability problems spanning 95 additional projects. Prior work lacked systematic quantification of Python cross-OS portability failures; this work uses two complementary approaches (test reexecution and manual issue analysis) to cover both runtime failures and developer-reported problems. Based on screening 2,042 open-source repositories, cross-platform test reexecution of 500 projects, and manual confirmation of 240 issues, the sample scale and dual-method design support the baseline conclusions.

The study builds a taxonomy of cross-OS portability issues comprising 7 primary failure categories (file/directory operations, process management, and library dependencies being most prevalent), 24 sub-categories, 15 diagnostic signatures, and 4 systematic repair patterns. It organizes scattered failure phenomena into a reusable classification and diagnostic structure, enabling developers and tool designers to locate and fix issues by category. The taxonomy derives from systematic synthesis of test results across 500 projects and 240 issues, and is paired with diagnostic signatures and repair patterns, making it actionable.

The study evaluates existing tools and large language models on portability issues: existing static analysis tools provide minimal support for portability detection, while large language models achieve 40-79% accuracy in identifying issues and 50-77% success in generating fixes when provided with structured guidance. It is the first to compare static analysis tools and LLMs on cross-OS portability tasks within the same empirical framework, and highlights the role of structured guidance in LLM performance. The evaluation is based on the aforementioned project and issue sets, reporting ranges for identification accuracy and fix success rates, though the abstract does not disclose detailed evaluation protocols.

The study demonstrates practical applicability through 33 contributed pull requests, of which 17 were merged and zero rejected, indicating developer acceptance of the findings. It advances empirical findings to real-world project fixes, providing closed-loop evidence from analysis to landed contributions. Merged/rejected pull request counts serve as a direct indicator of developer acceptance, with 33 submissions, 17 merged, and zero rejected.

Perspective

The study targets open-source Python projects and applies to developers, static analysis tool designers, and researchers studying portability in cross-OS deployment scenarios. Its taxonomy, diagnostic signatures, and repair patterns can serve as a foundation for building portability detection and automated repair tools, while the LLM evaluation offers reference for using large language models with structured guidance. The conclusions apply to the range of the Python ecosystem represented by the analyzed repositories and issue sets.

The currently loaded text is an incomplete abstract and page fragment, lacking the full paper, figures, and detailed evaluation protocols, so the specific environment configuration for cross-OS test reexecution, the prompt design and judgment criteria for the LLM evaluation, and the distribution details of the 33 pull requests cannot be confirmed. Readers who wish to reproduce or extend this work will need to consult the methods and experiments sections of the original paper. Additionally, whether the taxonomy and the LLM performance ranges remain stable across different Python versions, dependency ecosystems, or project types remains an open question worth watching.

Sources