Skip to main content
Back to timeline
Medical teacherSource publication:

An AI-driven immediate feedback system for observation-based clinical placements: a design-based research study

Synopsis

Using a design-based research framework, this study implemented an AI-driven immediate feedback system with 89 undergraduate judo therapy students during a four-day observation-based clinical placement, where students submitted daily reflective notes and received rubric-based AI scores and feedback within one minute; all 356 system requests were processed successfully, daily note submission rates exceeded 98% and self-assessment completion rates exceeded 95%, and student questionnaires indicated favorable perceptions; as an exploratory external check, three blinded clinical educators independently rated 120 notes from 30 randomly selected students, showing a moderate rank association (Spearman's rho = .495) but limited absolute agreement (ICC = .

AI-generated editorial illustration: Development and feasibility of an AI-driven immediate feedback system for observation-based clinical placements: A design-based research study.

Interpretation

The study built and deployed an AI-driven immediate feedback system that provides rubric-based scores and feedback on students' daily reflective notes within one minute, operating stably in a real clinical placement setting. Timely, individualized formative feedback in early observation-based placements is often constrained by supervisory capacity; this work embeds AI immediate feedback into the daily reflective-note workflow and realizes it in a real placement through design-based research. Direct evidence comes from system performance: all 356 system requests were processed successfully; across 89 students over four days, daily note submission rates exceeded 98% and self-assessment completion rates exceeded 95%, and student questionnaires indicated favorable perceptions of usability and feedback.

AI-generated rubric total scores showed a moderate rank association but limited absolute agreement with blinded human ratings from clinical educators, indicating that AI scores should be treated as system-internal indicators. Beyond reporting usability, the study introduced independent ratings of 120 notes by three blinded educators as an exploratory external check, placing AI scores side by side with external human judgment. Based on 120 notes from 30 randomly selected students independently rated by three blinded clinical educators using the same rubric, yielding Spearman's rho = .495 and ICC = .399; the authors accordingly state that AI-generated scores should not be treated as independent evidence of educational improvement.

AI total scores increased from Day 1 to Day 2 and then remained at similar levels, but this early increase was not reproduced in human total ratings, while human ratings of observational specificity increased from Day 1 to Day 4. The study presents the temporal trends of AI scores and human ratings separately, revealing differences in short-term patterns rather than treating AI scores alone as representing learning progress. The trend descriptions come from comparing AI scoring records with blinded human rating records across the four-day placement; the authors frame these as exploratory short-term patterns and call for controlled studies using external outcome measures.

Perspective

This study aimed to examine the feasibility and short-term patterns of the system in observation-based clinical placements, and it applies to formative feedback settings centered on daily reflective notes and rubric scoring under clinical educator oversight; its results reflect the implementation experience of 89 undergraduate judo therapy students over a four-day placement and can inform other programs or placement arrangements, but AI-generated scores should be understood as system-internal indicators rather than independent evidence of educational improvement.

Readers may still watch: how far the limited absolute agreement between AI and human scores (ICC = .399) is influenced by the rubric, the raters, or note content; how to interpret the divergence between the early rise in AI total scores and the rise in human observational specificity; how students' favorable questionnaire responses relate to changes in reflection quality; and, absent controlled studies with external outcome measures, the system's effect on learning outcomes remains an open question.

Sources