Twenty-seven middle school teachers configured pedagogical intent in a teacher-facing chatbot authoring tool, and log-based evaluation showed 88.9% alignment for responsiveness and 81.5% for persona versus 70.4% for rules and 59.3% for purpose
Synopsis
In professional development workshops, 27 middle school teachers used a teacher-facing chatbot authoring tool, and by analyzing focus-group interviews alongside configuration and interaction logs the study examined how pedagogical intentions are translated into chatbot configurations and reflected in chatbot behavior, finding that teachers envisioned chatbots as instructional scaffolds offering differentiated support, extending access to assistance, and preserving student thinking within teacher-defined boundaries; configuration analysis showed Purpose primarily captured instructional goals and content focus while Rules more often specified pedagogical behavior, guardrails, and learner-specific adaptations; and log-based evaluation showed stronger alignment for responsiveness (88.
Figure 1. The chatbot teacher-facing authoring interface. Teachers can configure a chatbot’s Purpose , Character and Personality , Communication Tone , and Rules and Guidelines , as well as adjustable behavioral traits such as confidence, transparency, formality, and assertiveness. The interface allows teachers to define both the chatbot’s instructional role and the behavioral constraints that guide its responses. Screenshot of the chatbot authoring interface showing configurable sections for chatbot purpose, character and personality, communication tone, rules and guidelines, and adjustable behavioral traits.
arXivInterpretation
The study characterizes how teachers write pedagogical intent into chatbot configurations: the Purpose field primarily captured instructional goals and content focus, whereas the Rules field more often specified pedagogical behavior, guardrails, and learner-specific adaptations. Prior work on educational chatbots often sets bot behavior from a designer or researcher perspective, whereas this work hands configuration to frontline teachers and systematically compares what different configuration fields actually carry, revealing that intent is not evenly distributed across fields. Based on configuration logs and focus-group interviews with 27 middle school teachers in professional development workshops, it is direct observation of real teacher configuration behavior in a middle school teacher sample.
Log-based evaluation showed that alignment between chatbot behavior and teacher intent varies by dimension: responsiveness reached 88.9% and persona 81.5%, while rules reached 70.4% and purpose 59.3%. The study goes beyond how teachers configure by using logs to compare configuration with actual chatbot behavior, providing per-dimension alignment proportions that make pedagogical fidelity a measurable object. The alignment proportions come from interaction-log evaluation across four dimensions—responsiveness, persona, rules, and purpose—and the numerical differences are explicitly stated in the text.
The study concludes that configurable controls alone do not ensure pedagogical fidelity, highlighting the need for authoring tools that help teachers express, test, and refine intended chatbot behavior. This conclusion relocates the problem from model capability to authoring-tool design, suggesting that the bottleneck for educational AI in practice may lie in how teachers express and verify intent rather than in generation quality alone. The conclusion is supported jointly by configuration analysis and log-based evaluation, and is an induction from tool use in workshop settings rather than a large-scale controlled experiment.
Perspective
The study addresses middle school teachers using a teacher-facing chatbot authoring tool, and its conclusions apply to configuration and use processes within professional development workshop settings. It enables follow-up work to design authoring tools around helping teachers express, test, and refine intended chatbot behavior, and makes pedagogical fidelity something measurable per dimension from logs, with reference value for educational AI tool designers, teacher educators, and researchers.
The alignment proportions come from log-based evaluation in workshop settings, and the criteria and coding details behind them are not expanded in the abstract, so readers interested in how alignment is defined would still need the full text. In addition, the sample of 27 middle school teachers and the professional development workshop setting mean that generalization to other grade levels, subjects, or non-workshop contexts still requires further verification.
