Skip to main content
Back to timeline
Studies in health technology and informaticsSource publication:

Benchmarking AI Vibe Coding for Clinical Statistical Analysis: A Structured Evaluation in Pulmonary Hypertension Research

Synopsis

This study had five current large language models (GPT 5.3, Claude Sonnet 4.6, Gemini 2.5 Flash, Perplexity, and Grok) reproduce three statistical tasks from a published clinical workflow under identical datasets and standardised prompts—descriptive table generation, Kaplan–Meier survival analysis, and Cox proportional hazards modelling—and compared their outputs with analyses by two experts trained in mathematical statistics on grouping correctness, numerical accuracy, missing-value handling, and quality of generated R code, finding that all five models produced correct descriptive statistics once dataset variables were specified explicitly, that two models failed the initial descriptive benchmark because of variable-name ambiguity but recovered after prompt clarification, that all models

AI-generated editorial illustration: Benchmarking AI Vibe Coding for Clinical Statistical Analysis: A Structured Evaluation in Pulmonary Hypertension Research.

Interpretation

The study builds a benchmark for LLM statistical reproduction anchored to a published clinical workflow, covering three task types: descriptive table generation, Kaplan–Meier survival analysis, and Cox proportional hazards modelling. Relative to generic code-generation evaluations, it anchors assessment to real clinical analysis results validated by two experts in mathematical statistics and standardises the datasets and prompts. A structured comparison across five named models, three statistical tasks, and expert reference analyses, with evaluation dimensions including grouping correctness, numerical accuracy, missing-value handling, and R code quality.

Once dataset variables were specified explicitly, all five models produced correct descriptive statistics; two models initially failed the descriptive benchmark because of variable-name ambiguity but recovered after prompt clarification. This indicates that the clarity of variable definitions in the prompt directly affects whether a model passes the foundational descriptive step. Based on direct observation of five models under identical datasets and standardised prompts, with the failure and subsequent recovery after prompt clarification documented.

All models reproduced the correct Kaplan–Meier p-value, though figure completeness differed; in the Cox benchmark four models reproduced all required hazard-ratio terms while one omitted the interaction results. It extends evaluation from numerical correctness to figure completeness and to differences among models in handling interaction terms. Based on item-by-item comparison with expert-validated analyses, covering p-values, hazard-ratio terms, and interaction results.

The overall conclusion is that LLMs may support clinical research workflows, but their outputs still require careful validation before use. It provides an initial reproducibility-centred evidence base for introducing LLM assistance in clinical statistical settings. The conclusion is supported jointly by the between-model differences across the three tasks and the expert reference comparisons.

Perspective

The evaluation concerns reproducing three statistical tasks from a published clinical workflow under given datasets and standardised prompts, and applies to clinical research teams wishing to use LLM assistance for descriptive statistics, survival analysis, and Cox modelling; its conclusions pertain to analysis settings where variables are specified explicitly and expert-validated results serve as the reference.

The loaded text is at the abstract level and lacks figures and full methodological detail, so the specific performance of each model on missing-value handling and R code quality, the concrete content of the figure-completeness differences, and how much the omitted interaction results affect the conclusions remain open questions a reader would need to confirm against the original; moreover, changes in model versions and prompt design may affect reproduction performance.

Sources