Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

arXiv

IOSVLM diagnoses multiple dental diseases directly from native 3D intraoral scan point clouds, reaching 77.23% macro accuracy, 9.58 points above Gemini 3 Pro

The authors present IOSVLM, an end-to-end 3D vision-language model that represents intraoral scans (IOS) as point clouds and follows a 3D encoder-projector-LLM design for unified diagnosis and generative VQA, together with IOSVQA, a dataset of 19,002 cases and 249,055 VQA pairs over 23 oral diseases and heterogeneous scan types, and with a geometry-to-chromatic proxy plus two-stage curriculum training it reaches 77.23% macro accuracy and 50.39% macro F1 on IOSVQA, outperforming baselines including GPT-5 and Gemini 3 Pro.