Public articles linked to the same research event.
arXiv The authors present IOSVLM, an end-to-end 3D vision-language model that represents intraoral scans (IOS) as point clouds and follows a 3D encoder-projector-LLM design for unified diagnosis and generative VQA, together with IOSVQA, a dataset of 19,002 cases and 249,055 VQA pairs over 23 oral diseases and heterogeneous scan types, and with a geometry-to-chromatic proxy plus two-stage curriculum training it reaches 77.23% macro accuracy and 50.39% macro F1 on IOSVQA, outperforming baselines including GPT-5 and Gemini 3 Pro.
The authors present IOSVLM, an end-to-end 3D vision-language model that represents intraoral scans (IOS) as point clouds and follows a 3D encoder-projector-LLM design for unified diagnosis and generative VQA, together with IOSVQA, a dataset of 19,002 cases and 249,055 VQA pairs over 23 oral diseases and heterogeneous scan types, and with a geometry-to-chromatic proxy plus two-stage curriculum training it reaches 77.23% macro accuracy and 50.39% macro F1 on IOSVQA, outperforming baselines including GPT-5 and Gemini 3 Pro.
The authors present IOSVLM, an end-to-end 3D vision-language model that represents intraoral scans (IOS) as point clouds and follows a 3D encoder-projector-LLM design for unified diagnosis and generative VQA, together with IOSVQA, a dataset of 19,002 cases and 249,055 VQA pairs over 23 oral diseases and heterogeneous scan types, and with a geometry-to-chromatic proxy plus two-stage curriculum training it reaches 77.23% macro accuracy and 50.39% macro F1 on IOSVQA, outperforming baselines including GPT-5 and Gemini 3 Pro.
The authors present IOSVLM, an end-to-end 3D vision-language model that represents intraoral scans (IOS) as point clouds and follows a 3D encoder-projector-LLM design for unified diagnosis and generative VQA, together with IOSVQA, a dataset of 19,002 cases and 249,055 VQA pairs over 23 oral diseases and heterogeneous scan types, and with a geometry-to-chromatic proxy plus two-stage curriculum training it reaches 77.23% macro accuracy and 50.39% macro F1 on IOSVQA, outperforming baselines including GPT-5 and Gemini 3 Pro.