OPGAgent uses a multi-tool agent with consensus to read panoramic dental X-rays, reaching 42.3% exact-match F1 on its OPG-Bench while holding false positives to 4.89 per case
Synopsis
The work proposes OPGAgent, a multi-tool dental agent planned by GPT-5.2 under the ReAct paradigm that orchestrates hierarchical evidence gathering, a specialized toolbox, and a consensus subagent, together with OPG-Bench, a structured-report protocol built on (Location, Field, Value) triples derived from real clinical reports; on OPG-Bench, comprising 1,009 anonymized OPGs and 5,219 VQA pairs, OPGAgent reaches 42.3% exact-match F1, a 49.7% aggregate score, 43.1% precision and 4.89 false positives per case, and leads MMOral-OPG at 62.53% accuracy.
Fig. 1. Overview of OPGAgent. The Agent orchestrates three modules: Hierarchical Evidence Gathering, Specialized Toolbox, and Consensus Subagent.
· Page 4Interpretation
OPGAgent is a multi-tool agent for OPG interpretation composed of Hierarchical Evidence Gathering, a Specialized Toolbox, and a Consensus Subagent, refining analysis through global, quadrant, and tooth-level phases and storing findings in Memory. The authors describe it as the first agentic system specifically designed for OPG interpretation, embedding dental domain mechanisms such as FDI notation and dynamic ROI cropping, whereas existing medical agents (MedAgents, MDAgents, MedAgent-Pro) are not designed for dental nuances. The paper details the three-phase pipeline, four toolbox categories (spatial, detection, utility, expert zoos), and the consensus rule (a finding is confirmed when at least 3 sources agree or at least 2 sources report the same finding), and validates each module through cumulative ablation.
The Consensus Subagent resolves conflicts through multi-source voting and anatomical constraints: when a majority confirms a finding but sources disagree on tooth number or severity, it consults the detection tool's FDI coordinate map to assign the finding to the spatially correct tooth. Compared with single-pass generative VLM output, this mechanism uses deterministic coordinates as hard constraints to correct VLM attribute errors and lets rare conditions missed by rule-based detectors still pass through VLM votes. Ablation shows adding Expert Zoos raises precision to 38.62% and cuts false positives from 6.17 to 2.37 but drops recall to 16.64%; adding Spatial Tools restores recall to 35.98% and F1 to 36.55%; adding Detection Tools reaches the peak 42.30% F1.
OPG-Bench formalizes dental reports as (Location, Field, Value) triples, with locations following FDI (ISO 3950) notation and values built on standards such as ICDAS, AAP/EFP 2017, and PAI that map numerical scales into semantic severity levels, recording only anomalous findings. The authors argue existing VQA benchmarks measure only the questions asked, leaving finding types absent from the question set invisible and hallucinations in unprompted regions unquantified; the triple protocol audits both pathological findings and hallucinations and provides exact-match plus step-wise detection, localization, and classification metrics. The dataset contains 1,009 anonymized OPGs with paired clinical and structured reports and 5,219 unguided VQA pairs, sourced from multiple clinics, restricted to patients aged 16 and above, with ground truths from real clinical reports validated by manual spot-checks.
On OPG-Bench and the public MMOral-OPG, OPGAgent outperforms dental VLMs and medical agent frameworks across both structured-report and VQA evaluation. The authors report that OPGAgent keeps precision relatively high while substantially reducing false positives, whereas Gemini-3-Flash has higher recall (45.1%) but precision drops to 27.6% with 10.58 false positives per case, and OralGPT-Omni has the fewest false positives (4.02) but low coverage (F1 6.2%). Comparisons span general VLMs, dental VLMs, and medical agents; to isolate the architectural contribution, MedAgent-Pro is configured with the same LLM and tools as OPGAgent, and all models receive identical prompts (temperature 0.3) and pass through a shared Parser Agent.
Perspective
The result targets multi-task screening and structured report generation on panoramic radiographs, for patients aged 16 and above with predominantly permanent dentition; its value lies in organizing scattered specialized models and VLM experts into an auditable pipeline and grounding findings as FDI location, clinical field, and graded value. For readers building dental or other imaging agents, the reusable parts are the three layers of hierarchical evidence gathering, tool wrapping, and multi-source consensus; for clinical and evaluation readers, the reusable parts are the triple-based reporting protocol and the step-wise detection, localization, and classification metrics. The authors state that code will be released upon acceptance, so what is currently reproducible is the method description and evaluation design.
The authors note that real-world reports typically prioritize chief complaints over incidental findings, so reported false positive rates may be slightly inflated when the agent detects undocumented yet valid pathologies. All models are structured by the same Parser Agent and manually filtered, so the influence of the parser and filtering steps on conclusions remains worth watching. In addition, dental-specific VLMs score much lower on the authors' OPG-Bench than general VLMs while performing well on their own MMOral-OPG, leading the authors to suggest that models trained on generated QA pairs may not fully capture the distribution of real clinical reports, an explanation that still needs more data to verify. The sensitivity of the consensus thresholds (at least 3 sources agreeing or at least 2 reporting the same finding) and distribution differences across contributing clinics are further directions a careful reader may watch.
