Reconstructing 3D Oral Models from Just Ten 2D Intraoral Images Reaches 77.49% Nearest-Neighbor Accuracy on Teeth3DS
Synopsis
The study proposes a software-only method that reconstructs a 3D oral model from only ten 2D intraoral images captured from different angles, requiring no dedicated hardware; the model is trained on the public Teeth3DS dataset of 950 upper jaw samples and uses MobileNetV2 as the image encoder with Multi-head Attention for multi-view feature fusion, achieving 77.49% accuracy under nearest-neighbor matching with a distance threshold of 0.035, while predicted vertices tend to concentrate in high-density regions of the ground truth, producing uneven point distribution in the reconstructed model.
Figure 2: Ten fixed viewpoints rendered from a preprocessed upper jaw mesh.
arXivInterpretation
A software-based approach that reconstructs a 3D oral model using only ten 2D intraoral images taken from different angles, with no dedicated hardware devices required. Compared with impression taking, which uses trays and alginate or silicone and suffers from patient discomfort, material deformation errors, and storage and transport difficulties, and with intraoral scanners that rely on structured light or laser technology at substantially high equipment cost, this approach shifts 3D modeling from hardware dependence to a purely software pipeline. The abstract states the input is ten 2D intraoral images and that the method requires 'no dedicated hardware devices', listing reduced cost, no physical scanning equipment, minimized patient discomfort, and automated reconstruction as goals; implementation details are not expanded in the provided text.
The model uses MobileNetV2 as the image encoder and Multi-head Attention for multi-view feature fusion, trained on the public Teeth3DS dataset. The work applies attention-based multi-view fusion to the task of reconstructing 3D oral structure from intraoral images, pairing a public dataset with a lightweight encoder. The abstract states training on the publicly available Teeth3DS dataset comprising 950 upper jaw samples and specifies that the model 'employs MobileNetV2 as the image encoder combined with Multi-head Attention for multi-view feature fusion'.
Under nearest-neighbor matching with a distance threshold of 0.035, reconstruction accuracy reaches 77.49%. This number gives a quantified result for a software-only, ten-image-input approach under a stated metric, offering a reference point for later comparisons. The abstract directly reports that the model 'achieves an accuracy of 77.49%, measured by nearest-neighbor matching with a distance threshold of 0.035', without confidence intervals or cross-dataset validation.
Predicted vertices tend to concentrate in high-density regions of the ground truth, resulting in uneven point distribution across the reconstructed model. The authors report a spatial distribution bias in the output point set rather than only an aggregate accuracy figure. The abstract states that 'predicted vertices tend to concentrate in high-density regions of the ground truth, resulting in uneven point distribution across the reconstructed model', a qualitative description without quantified distribution metrics.
Perspective
The result is aimed at dental oral 3D modeling, particularly software-based workflows that seek to avoid impression materials and costly intraoral scanning equipment; its setting is ten 2D intraoral images from different angles as the only input, with training and evaluation on 950 upper jaw samples from the public Teeth3DS dataset. For readers, this positions the work as a reference point for an image-only, low-cost alternative to hardware scanning and as a starting point for improving multi-view fusion and point distribution uniformity.
The provided text is an abstract plus bibliographic information and does not include network architecture details, training hyperparameters, ablation studies, comparison tables with other methods, or a quantified analysis of the uneven point distribution, so the stability of the 77.49% accuracy across thresholds or datasets cannot be judged. The abstract also does not describe the acquisition conditions for the ten images, whether real patient data were involved, or validation within a clinical workflow; these remain open questions a reader may follow up on.
