Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

Studies in health technology and informatics

Benchmarking AI Vibe Coding for Clinical Statistical Analysis: A Structured Evaluation in Pulmonary Hypertension Research

This study had five current large language models (GPT 5.3, Claude Sonnet 4.6, Gemini 2.5 Flash, Perplexity, and Grok) reproduce three statistical tasks from a published clinical workflow under identical datasets and standardised prompts—descriptive table generation, Kaplan–Meier survival analysis, and Cox proportional hazards modelling—and compared their outputs with analyses by two experts trained in mathematical statistics on grouping correctness, numerical accuracy, missing-value handling, and quality of generated R code, finding that all five models produced correct descriptive statistics once dataset variables were specified explicitly, that two models failed the initial descriptive benchmark because of variable-name ambiguity but recovered after prompt clarification, that all models