Public articles linked to the same research event.
arXiv The authors introduce TomoTransformer, a transformer architecture that treats each local filtered projection as an individual token and predicts missing views via self-attention in a back-projection space, yielding a single foundation model that can process any number of input projections at arbitrary angular locations and detector dimensions and query any number of target angles without retraining; trained on a large-scale dataset spanning diverse medical CT anatomies and natural images, it significantly outperforms concurrent multi-purpose models like ViewTrans and matches or exceeds strong protocol-specific baselines on several sparse-view benchmarks, while showing robust zero-shot generalization on real experimental nanoscale brain data collected from an X-ray synchrotron.
The authors introduce TomoTransformer, a transformer architecture that treats each local filtered projection as an individual token and predicts missing views via self-attention in a back-projection space, yielding a single foundation model that can process any number of input projections at arbitrary angular locations and detector dimensions and query any number of target angles without retraining; trained on a large-scale dataset spanning diverse medical CT anatomies and natural images, it significantly outperforms concurrent multi-purpose models like ViewTrans and matches or exceeds strong protocol-specific baselines on several sparse-view benchmarks, while showing robust zero-shot generalization on real experimental nanoscale brain data collected from an X-ray synchrotron.
The authors introduce TomoTransformer, a transformer architecture that treats each local filtered projection as an individual token and predicts missing views via self-attention in a back-projection space, yielding a single foundation model that can process any number of input projections at arbitrary angular locations and detector dimensions and query any number of target angles without retraining; trained on a large-scale dataset spanning diverse medical CT anatomies and natural images, it significantly outperforms concurrent multi-purpose models like ViewTrans and matches or exceeds strong protocol-specific baselines on several sparse-view benchmarks, while showing robust zero-shot generalization on real experimental nanoscale brain data collected from an X-ray synchrotron.
The authors introduce TomoTransformer, a transformer architecture that treats each local filtered projection as an individual token and predicts missing views via self-attention in a back-projection space, yielding a single foundation model that can process any number of input projections at arbitrary angular locations and detector dimensions and query any number of target angles without retraining; trained on a large-scale dataset spanning diverse medical CT anatomies and natural images, it significantly outperforms concurrent multi-purpose models like ViewTrans and matches or exceeds strong protocol-specific baselines on several sparse-view benchmarks, while showing robust zero-shot generalization on real experimental nanoscale brain data collected from an X-ray synchrotron.