Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

arXiv

OctLLM treats octree occupancy bytes as an explicit 3D language, improving 3D generation and understanding in one unified multimodal model while preserving language ability

OctLLM converts meshes into coordinate- and depth-anchored Sparse Octree (S-Octree) occupancy bytes as an explicit 3D language and places 3D capacity in full-rank branches separate from the pretrained weights, training about 2.80B parameters; it achieves the best 3D generation and understanding among unified multimodal LLMs, lowering image-to-3D Inception FID from ShapeLLM-Omni's 38.21 to 31.58 and raising render-grounded captioning from 18.14 to 46.88, while exactly reproducing the backbone on MMLU and HellaSwag and staying within one point on GSM8K.