TSegAgent achieves zero-shot tooth segmentation and FDI identification by pairing SAM3 with geometric reasoning, reaching mIoU 93.37 and TIR=1 of 87.17% on Teeth3DS while generalizing to a private dataset
Synopsis
The work proposes TSegAgent, which reformulates tooth instance segmentation and FDI labeling of intra-oral scanned 3D models as a zero-shot geometric reasoning problem: multi-view renderings with curvature heatmaps and the SAM3 text prompt "tooth" produce candidate masks, which are merged into face-level instance labels via IoU and containment relations, after which a geometry-aware vision-language agent performs non-tooth region identification, central incisor localization, full-arch classification, and error correction through multi-round conversation; it reports mIoU 93.37, TLA 96.40, TSA 96.76, TIR 97.46, and TIR=1 87.17 on Teeth3DS with 1200 3D tooth models, and mIoU 82.10, TLA 96.68, TSA 95.99, TIR 85.41, and TIR=1 51.
Interpretation
It proposes a zero-shot framework for tooth instance segmentation and identification that outputs FDI labels directly from intra-oral scanned 3D models without task-specific training or dense annotation. Most existing methods treat tooth segmentation and identification as supervised direct prediction problems relying on task-specific 3D networks trained on densely annotated data such as Teeth3DS; this work instead casts them as zero-shot geometric reasoning, replacing learned dental-specific features with general-purpose foundation models plus dental anatomy priors. The paper compares against MeshSegNet, PTv3, TSGCNet, DBGANet, CrossTooth, and 3DTeethSAM on Teeth3DS (1200 3D tooth models with per-vertex FDI labels) and on a private dataset (340 test models, scanned plaster models with more severe defects and artifacts), reporting best or near-best results for TSegAgent on both.
It designs a geometry-aware and interpretable vision-language agent that integrates dental arch priors, multi-view visual evidence, and volumetric information for tooth classification. Whereas SAM-based methods such as IOSSAM and 3DTeethSAM still require training a teeth detector or refiner, this work reorders instance IDs along a fitted dental arch curve and writes each tooth's oriented bounding box 3D size into the text prompt, guiding the VLM to determine central incisors from dental arch symmetry and size correspondence rather than left-right pixel balance in the image. Ablation shows the segmentation-only baseline at mIoU 75.21 and TLA 96.45; adding VLM and curvature raises mIoU to 85.49 with TIR 89.64 and TIR=1 66.72; adding arch reordering gives TIR 90.43 and TIR=1 74.83; adding OBB volumetric cues gives TIR 86.36 and TIR=1 71.17; the full configuration reaches mIoU 93.37, TIR 97.46, and TIR=1 87.17.
It introduces explicit geometric reasoning and verification mechanisms that detect and correct identification errors through multi-round conversation, reducing ambiguity in tooth labeling. The work decomposes identification into interpretable sub-tasks including non-tooth region identification, central incisor identification, full-arch classification, and error detection and correction, triggering self-correction from bilateral symmetry, sequential consistency along the dental arch, and relative volumetric similarity between corresponding teeth, without external supervision or post-hoc optimization. The paper states that segmentation gains mainly come from non-tooth filtering and over-segmentation removal while error correction chiefly improves FDI accuracy; on the private dataset, supervised methods such as 3DTeethSAM and DBGANet show notable degradation while TSegAgent maintains robust overall accuracy.
It demonstrates stronger generalization on a private dataset with substantial domain shift. Most compared methods are trained from scratch to convergence on the Teeth3DS training set and degrade sharply under cross-domain testing; because TSegAgent does not rely on task-specific training, it maintains higher metrics on scanned plaster models with more severe defects, interference, and artifacts. On the private dataset TSegAgent reports mIoU 82.10, TLA 96.68, TSA 95.99, TIR 85.41, and TIR=1 51.96, versus 3DTeethSAM at mIoU 81.42, TLA 70.98, TIR 82.42, and TIR=1 37.06, and DBGANet at mIoU 41.59, TIR 61.83, and TIR=1 22.06.
Perspective
The result targets intra-oral scanned 3D models in digital dentistry, suited to settings requiring tooth instance segmentation and FDI labeling such as orthodontic diagnosis, treatment planning, and prosthodontic design; because the method runs zero-shot without task-specific training or dense annotation, it is more attractive to teams with limited annotation budgets or diverse scan sources. The paper states TSegAgent is compatible with various VLM services, uses ChatGPT 5.2 in experiments, and has released code, making it feasible to reproduce and extend on one's own scan data. The authors propose extending the reasoning framework to other anatomical structures and incorporating additional domain-specific priors to further enhance robustness and clinical applicability.
The paper does not report inference latency, memory footprint, or per-case cost details, nor a quantitative analysis of the number of conversation rounds; the private dataset comes from a collaborating clinic with limited description of acquisition and annotation, so cross-institution reproduction requires independently assessing domain differences. Some ablation settings lack TSA or TIR values, leaving room to further characterize interactions between the full configuration and individual components. In addition, experiments use a single VLM service, so behavior with other models remains an open question; extending to other anatomical structures is listed as future work without validation results in that direction.
