Public articles linked to the same research event.
arXiv ELSA3D introduces elastic semantic anchoring, representing geometry with a scale-aware octree tokenizer and adding sparse cross-modal Anchor Tokens that select semantic cues, route them to the most relevant 3D scale, and write the fused signal back into the unified representation, with a lightweight per-block router making computation and reasoning elastic; it achieves state-of-the-art performance across image-to-3D generation, text-to-3D generation, and 3D captioning, outperforming the strongest unified baseline while roughly halving FLOPs and inference latency relative to the non-elastic version of the same model.
ELSA3D introduces elastic semantic anchoring, representing geometry with a scale-aware octree tokenizer and adding sparse cross-modal Anchor Tokens that select semantic cues, route them to the most relevant 3D scale, and write the fused signal back into the unified representation, with a lightweight per-block router making computation and reasoning elastic; it achieves state-of-the-art performance across image-to-3D generation, text-to-3D generation, and 3D captioning, outperforming the strongest unified baseline while roughly halving FLOPs and inference latency relative to the non-elastic version of the same model.
ELSA3D introduces elastic semantic anchoring, representing geometry with a scale-aware octree tokenizer and adding sparse cross-modal Anchor Tokens that select semantic cues, route them to the most relevant 3D scale, and write the fused signal back into the unified representation, with a lightweight per-block router making computation and reasoning elastic; it achieves state-of-the-art performance across image-to-3D generation, text-to-3D generation, and 3D captioning, outperforming the strongest unified baseline while roughly halving FLOPs and inference latency relative to the non-elastic version of the same model.
ELSA3D introduces elastic semantic anchoring, representing geometry with a scale-aware octree tokenizer and adding sparse cross-modal Anchor Tokens that select semantic cues, route them to the most relevant 3D scale, and write the fused signal back into the unified representation, with a lightweight per-block router making computation and reasoning elastic; it achieves state-of-the-art performance across image-to-3D generation, text-to-3D generation, and 3D captioning, outperforming the strongest unified baseline while roughly halving FLOPs and inference latency relative to the non-elastic version of the same model.