Public articles linked to the same research event.
arXiv OctLLM converts meshes into coordinate- and depth-anchored Sparse Octree (S-Octree) occupancy bytes as an explicit 3D language and places 3D capacity in full-rank branches separate from the pretrained weights, training about 2.80B parameters; it achieves the best 3D generation and understanding among unified multimodal LLMs, lowering image-to-3D Inception FID from ShapeLLM-Omni's 38.21 to 31.58 and raising render-grounded captioning from 18.14 to 46.88, while exactly reproducing the backbone on MMLU and HellaSwag and staying within one point on GSM8K.
OctLLM converts meshes into coordinate- and depth-anchored Sparse Octree (S-Octree) occupancy bytes as an explicit 3D language and places 3D capacity in full-rank branches separate from the pretrained weights, training about 2.80B parameters; it achieves the best 3D generation and understanding among unified multimodal LLMs, lowering image-to-3D Inception FID from ShapeLLM-Omni's 38.21 to 31.58 and raising render-grounded captioning from 18.14 to 46.88, while exactly reproducing the backbone on MMLU and HellaSwag and staying within one point on GSM8K.
OctLLM converts meshes into coordinate- and depth-anchored Sparse Octree (S-Octree) occupancy bytes as an explicit 3D language and places 3D capacity in full-rank branches separate from the pretrained weights, training about 2.80B parameters; it achieves the best 3D generation and understanding among unified multimodal LLMs, lowering image-to-3D Inception FID from ShapeLLM-Omni's 38.21 to 31.58 and raising render-grounded captioning from 18.14 to 46.88, while exactly reproducing the backbone on MMLU and HellaSwag and staying within one point on GSM8K.
OctLLM converts meshes into coordinate- and depth-anchored Sparse Octree (S-Octree) occupancy bytes as an explicit 3D language and places 3D capacity in full-rank branches separate from the pretrained weights, training about 2.80B parameters; it achieves the best 3D generation and understanding among unified multimodal LLMs, lowering image-to-3D Inception FID from ShapeLLM-Omni's 38.21 to 31.58 and raising render-grounded captioning from 18.14 to 46.88, while exactly reproducing the backbone on MMLU and HellaSwag and staying within one point on GSM8K.