Public articles linked to the same research event.
Studies in health technology and informatics This study had five current large language models (GPT 5.3, Claude Sonnet 4.6, Gemini 2.5 Flash, Perplexity, and Grok) reproduce three statistical tasks from a published clinical workflow under identical datasets and standardised prompts—descriptive table generation, Kaplan–Meier survival analysis, and Cox proportional hazards modelling—and compared their outputs with analyses by two experts trained in mathematical statistics on grouping correctness, numerical accuracy, missing-value handling, and quality of generated R code, finding that all five models produced correct descriptive statistics once dataset variables were specified explicitly, that two models failed the initial descriptive benchmark because of variable-name ambiguity but recovered after prompt clarification, that all models
This study had five current large language models (GPT 5.3, Claude Sonnet 4.6, Gemini 2.5 Flash, Perplexity, and Grok) reproduce three statistical tasks from a published clinical workflow under identical datasets and standardised prompts—descriptive table generation, Kaplan–Meier survival analysis, and Cox proportional hazards modelling—and compared their outputs with analyses by two experts trained in mathematical statistics on grouping correctness, numerical accuracy, missing-value handling, and quality of generated R code, finding that all five models produced correct descriptive statistics once dataset variables were specified explicitly, that two models failed the initial descriptive benchmark because of variable-name ambiguity but recovered after prompt clarification, that all models
This study had five current large language models (GPT 5.3, Claude Sonnet 4.6, Gemini 2.5 Flash, Perplexity, and Grok) reproduce three statistical tasks from a published clinical workflow under identical datasets and standardised prompts—descriptive table generation, Kaplan–Meier survival analysis, and Cox proportional hazards modelling—and compared their outputs with analyses by two experts trained in mathematical statistics on grouping correctness, numerical accuracy, missing-value handling, and quality of generated R code, finding that all five models produced correct descriptive statistics once dataset variables were specified explicitly, that two models failed the initial descriptive benchmark because of variable-name ambiguity but recovered after prompt clarification, that all models
This study had five current large language models (GPT 5.3, Claude Sonnet 4.6, Gemini 2.5 Flash, Perplexity, and Grok) reproduce three statistical tasks from a published clinical workflow under identical datasets and standardised prompts—descriptive table generation, Kaplan–Meier survival analysis, and Cox proportional hazards modelling—and compared their outputs with analyses by two experts trained in mathematical statistics on grouping correctness, numerical accuracy, missing-value handling, and quality of generated R code, finding that all five models produced correct descriptive statistics once dataset variables were specified explicitly, that two models failed the initial descriptive benchmark because of variable-name ambiguity but recovered after prompt clarification, that all models