Artificial Intelligence in Evaluating Undergraduate Dental Students: Benchmarking AI Platforms Against Student Performance
Synopsis
In this cross-sectional comparative study, 49 second-year dental students completed a 45-item multiple-choice final examination in prosthodontic technology under standardized conditions, the same items were submitted to three AI systems (ChatGPT, Gemini, and DeepSeek), and analysis using ANOVA, Cronbach Alpha, and Tukey HSD showed that DeepSeek achieved perfect scores on high-difficulty questions while ChatGPT showed adaptive improvement across attempts, both significantly outperforming the student cohort, whereas Gemini demonstrated lower and consistent accuracy, with significant differences between models (P=0.036, eta-squared=0.891), leading the authors to argue that unsupervised online summative examinations require critical reevaluation of assessment integrity.
Interpretation
Within a concrete subject setting, a standardized final examination in prosthodontic technology, AI models can equal or exceed the performance of undergraduate dental students. Prior discussion of LLM potential in educational assessment has often remained general; this study anchors the comparison to one course, the same 45 multiple-choice items, and the same cohort of 49 second-year students, providing a direct between-group contrast. Cross-sectional comparative design with 49 students and 45 items, using ANOVA, Cronbach Alpha, and Tukey HSD for intergroup differences, reporting P=0.036 and eta-squared=0.891.
Different AI systems performed differently on the same examination: DeepSeek achieved perfect scores on high-difficulty questions, ChatGPT showed adaptive improvement across attempts, and Gemini showed lower but consistent accuracy. Rather than treating AI as a single entity, the study compares three platforms side by side, revealing a capability stratification among models that the statistical tests confirm as significant. Three models answered the same items, with ANOVA and Tukey HSD used for between-group comparison; the difference between models was statistically significant (P=0.036) with a large effect size (eta-squared=0.891).
The fact that AI can match or surpass student performance points directly to integrity risks in unsupervised digital evaluation. The study translates a performance comparison into policy implications for assessment format, proposing that summative examinations should not be conducted online without strict proctoring or secure, verified environments. This conclusion rests on the observed performance advantage of AI over the student cohort in this study and is presented as an inference-based recommendation.
Perspective
This study addresses the specific setting of a standardized final examination in prosthodontic technology for second-year undergraduate dental students, and it is relevant to educators, examination administrators, and curriculum designers concerned with digital assessment integrity. It indicates that in comparable subject examinations AI may match or exceed student performance, so online summative evaluation requires strict proctoring or secure, verified environments; the scope of this conclusion should be limited to item types, difficulty distributions, and candidate populations similar to those studied here.
Readers should note that the loaded text is summary-level and does not include item-by-item performance, difficulty-stratified scores, the specific number of ChatGPT attempts and the magnitude of its improvement, or the run configurations of each model, so the precise sources of between-model differences remain unclear. In addition, whether the AI performance advantage varies with item type, language, or course content, and whether feasible assessment alternatives exist beyond strict proctoring, are open questions worth watching.
