Generative UI Learning Interactives for Teachers: Letting Educators Generate Guided Simulations Aligned to Curriculum Goals
Synopsis
This research experiment lets teachers enter a curriculum topic and have generative UI, under instructional-design guardrails, dynamically produce interactive simulations with leveled challenges, hints, feedback, and worked solutions, and it releases a library of over 30 teacher-reviewed STEM learning interactives; UK STEM teachers rated 40 interactives overall as good or excellent, with physics and chemistry most amenable to simulation creation, and 12 US teachers who each requested three custom interactives gave an average quality rating of 8 out of 10.
Interpretation
It proposes a generative-UI pipeline for education: the educator leads by suggesting a topic, the system generates precise and coherent learning objectives that are modifiable and must be teacher-approved, and those objectives then drive generation of the learning interactive. Compared with pre-coded off-the-shelf simulations, the interface itself is dynamically constructed by the model, while approval of learning objectives stays with the teacher. Presented as a process description and a list of generation requirements at the system-design level, without a controlled comparison.
It combines game-based levels with AI-generated scaffolding: each interactive contains progressively harder challenges plus an introduction to prime prior knowledge, a toolbox of formulas and theories, multiple levels of hints, tailored feedback, and worked solutions. It operationalizes active learning and the interactive behaviors emphasized by the ICAP framework into concrete interface elements rather than leaving them as principles. The text states these elements 'were all tested and iterated upon with teachers and students' but gives no test scale or quantitative results.
It embeds self-correcting loops and pedagogical guardrails in generation, covering pedagogy, mechanics, and visual aspects, including agentic auto-evaluation that opens a Chrome instance and interacts with the simulation as a user, plus adversarial actions such as taking knobs to extreme values. Quality checking moves from after-the-fact human review into the generation loop itself, which repeats until the outcome meets all required criteria. Described as an iterative process that increases generation time; no loop counts or pass rates are reported.
It releases a library of over 30 English STEM learning interactives and reports teacher evaluations: UK STEM teachers rated 40 interactives overall as good or excellent, with physics and chemistry most amenable to simulation creation, and 12 US teachers each requested three custom interactives and gave an average quality rating of 8 out of 10. It provides usable samples and initial teacher feedback from generation through teacher vetting, not just a concept demonstration. The sample is 40 interactives and 12 teachers, and the evaluation is mainly teacher subjective ratings; no student learning gains are reported.
Perspective
The result is aimed at middle and high school STEM teachers, especially those who want simulations customized to their own curriculum goals and who face the difficulty of differentiating with off-the-shelf simulations; the current library is in English, and physics and chemistry are reported as the subjects most amenable to simulation creation. Its value lies in shifting teachers from 'you get what you get' to requesting and vetting interactives against learning objectives, with expansion planned through the Google for Education Pilot Program and UX research and field studies to evaluate learning gains and student engagement.
The text does not yet report classroom learning gains or student engagement data, which are listed as future UX research and field studies; full details of the UK teacher evaluation are in the tech report, with only the overall conclusion given here; the number of self-correcting iterations, failure rates, and generation time are not quantified; teacher ratings rest on a sample of 40 interactives and 12 teachers and are mainly subjective; whether generation quality remains stable across subjects, grade levels, and languages remains an open question.
