IOSVLM diagnoses multiple dental diseases directly from native 3D intraoral scan point clouds, reaching 77.23% macro accuracy, 9.58 points above Gemini 3 Pro
Synopsis
The authors present IOSVLM, an end-to-end 3D vision-language model that represents intraoral scans (IOS) as point clouds and follows a 3D encoder-projector-LLM design for unified diagnosis and generative VQA, together with IOSVQA, a dataset of 19,002 cases and 249,055 VQA pairs over 23 oral diseases and heterogeneous scan types, and with a geometry-to-chromatic proxy plus two-stage curriculum training it reaches 77.23% macro accuracy and 50.39% macro F1 on IOSVQA, outperforming baselines including GPT-5 and Gemini 3 Pro.
Interpretation
Construction of IOSVQA, a large-scale multi-source IOS diagnostic VQA dataset supporting multiple scan-type inputs, covering 23 oral diseases with 19,002 cases and 249,055 QA pairs. Prior dental vision-language work largely operated on 2D images or multi-view renderings of IOS, leaving a gap in large paired resources built on native 3D geometry that explicitly capture co-existing diseases and heterogeneous scan forms. Data were aggregated from MaloccIOS, DiseaseIOS, and the public Bits2Bites; a randomly selected subset of 557 MaloccIOS cases was manually corrected by 28 orthodontists yielding 7,628 high-quality samples, while the remaining 200,742 samples contain partial label noise; DiseaseIOS and Bits2Bites labels were annotated by 5 and 1 orthodontic experts respectively.
IOSVLM, described as the first end-to-end VLM that takes native 3D IOS geometry as input for unified multi-disease diagnosis and generative VQA. Unlike pipelines that render 3D scans into multi-view images for 2D VLMs, this model does not rely on view selection and models 3D surface geometry directly. On the IOSVQA test set (5,884 samples) it attains the highest macro accuracy 77.23%, macro F1 50.39%, and recall 52.96%; it exceeds open-source 2D MLLMs by at least +16.02% accuracy/+11.25% F1, open-source 3D MLLMs by at least +34.20% accuracy/+16.21% F1, GPT-5 by +14.97% accuracy/+5.66% F1, and Gemini 3 Pro by +9.58% accuracy/+1.46% F1.
A geometry-to-chromatic proxy (GCP) that uses surface normals to build flip-robust pseudo-color channels, bridging the distribution gap between color-free IOS and color-dependent point-cloud pretraining. The work reframes RGB as a local separability cue rather than semantic color and substitutes that component with a purely geometry-derived quantity, allowing reuse of color-pretraining priors. In ablation, replacing GCP with a constant white color gives 67.02% accuracy/43.10% F1, while enabling GCP raises this to 72.28%/48.06%, i.e. +5.26% accuracy and +4.96% F1.
A two-stage curriculum: Stage-1 trains the 3D encoder and projectors with the LLM frozen, and Stage-2 freezes the encoder and fine-tunes projectors and LLM with LoRA, with GPT-4o-generated chain-of-thought rationales added to part of the high-quality samples. The schedule first builds geometry-language alignment under mixed-quality supervision and then uses higher-quality annotations to improve diagnostic and generative reliability. The default setting yields the best overall accuracy and F1 (77.23%/50.39%); rationale supervision gives macro F1 comparable to label-only tuning while reaching 100% macro parsing rate, with qualitatively fewer degenerate behaviors such as repeated labels.
Perspective
The work targets intraoral scans in routine dental settings, covering single-arch and occluded-arch inputs, with the goal of unified multi-disease diagnosis and generative natural-language question answering for clinical documentation and dentist-patient communication. The dataset integrates MaloccIOS, DiseaseIOS, and Bits2Bites and applies global registration to standardize orientation and relative maxillomandibular pose, so its conclusions apply to IOS data under this label system and preprocessing pipeline. For researchers and clinical informatics teams seeking to bring 3D surface geometry into diagnostic workflows, IOSVQA and IOSVLM offer a reusable training and evaluation starting point, and GCP plus the two-stage curriculum can serve as reference designs for other color-free 3D medical surface tasks.
Open questions remain: 200,742 MaloccIOS samples in IOSVQA contain partial label noise, which the curriculum training is designed to mitigate, but the specific contribution of that noise to the final metrics is not separately quantified; rationale supervision mainly improves output parsability and reduces degenerate behavior while adding limited macro F1, so its clinical usability still needs testing in real workflows; evaluation is concentrated on the single IOSVQA dataset with macro-averaged metrics, and external validation across institutions or scanners is not presented; GCP is instantiated with surface normals, and while the text notes that other geometry descriptors such as curvature may serve as alternative proxies, corresponding results are not given.
