Skip to main content
Back to timeline
arXivSource publication:

OctLLM treats octree occupancy bytes as an explicit 3D language, improving 3D generation and understanding in one unified multimodal model while preserving language ability

Related research and updates

Synopsis

OctLLM converts meshes into coordinate- and depth-anchored Sparse Octree (S-Octree) occupancy bytes as an explicit 3D language and places 3D capacity in full-rank branches separate from the pretrained weights, training about 2.80B parameters; it achieves the best 3D generation and understanding among unified multimodal LLMs, lowering image-to-3D Inception FID from ShapeLLM-Omni's 38.21 to 31.58 and raising render-grounded captioning from 18.14 to 46.88, while exactly reproducing the backbone on MMLU and HellaSwag and staying within one point on GSM8K.

AI-generated editorial illustration: Octrees as an Explicit 3D Language

Interpretation

Octree occupancy bytes serve as an explicit 3D language that preserves spatial locality and hierarchy. Earlier 3D LLMs mostly compress shapes into VQVAE codebook indices or serialize meshes as OBJ coordinate text, so spatial organization is no longer explicit; OctLLM serializes the octree in Z-order, packs every eight occupancy bits into one byte-level mesh token, and randomly empties penultimate-level nodes together with their finest-level descendants to obtain the S-Octree. In an ablation that replaces the S-Octree with ShapeLLM-Omni's 1,024 VQVAE tokens while keeping all other settings fixed, FID drops from 33.10 to 26.31 (20.5% lower) at the same KID of 0.80, and Sentence-BERT and SimCSE rise from 67.15/68.87 to 69.47/70.28; this ablation runs on the ShapeNet airplane subset, and the authors state its results are not directly comparable with the full Toys4K evaluation.

A token-routed dual-stream architecture puts 3D capacity in separate trainable branches, avoiding any rewrite of the language pathway. Full fine-tuning is costly and exposes pretrained weights to 3D gradients, while LoRA still alters the language pathway once merged; OctLLM routes mesh bytes and position-query <MASK> tokens through full-rank 3D branches in a subset of decoder blocks, keeps text and image tokens on the frozen vision-language pathway, shares causal self-attention between the streams, and gives the mesh vocabulary its own input embeddings and output head. 3D branches are routed through 12 of 28 decoder blocks (0-indexed layers 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 22, 24), adding about 2.80B trainable parameters, roughly a quarter of the total and about a third of what full fine-tuning updates; in Table 2(b) OctLLM matches Qwen2.5-VL-7B exactly on MMLU and HellaSwag (64.65, 67.44) and scores 83.01 versus 83.54 on GSM8K, whereas ShapeLLM-Omni and LLaMA-Mesh lose between a third and a half of the backbone's accuracy on MMLU and HellaSwag.

OctLLM achieves the best 3D generation and 3D understanding within the unified multimodal LLM family. Prior unified models were clearly weaker on one side of generation or understanding; OctLLM uses one backbone to both read and write S-Octrees, with position-aware mask modeling for generation and direct S-Octree input for understanding. Image-to-3D Inception FID is 31.58 and KID 0.56, better than ShapeLLM-Omni's 38.21 and 1.02; text-to-3D Inception FID is 45.61 and KID 0.98, better than ShapeLLM-Omni's 52.36 and 1.38 and OctGPT's 46.40 and 1.13; on PointLLM-200 captioning GPT-img is 46.88 versus ShapeLLM-Omni's 18.14, GPT-ref 26.61 versus 22.27, and S-BERT 38.80 versus 36.08; the authors note that the dedicated generator TRELLIS still scores best on most metrics, with OctLLM second on most of them.

Sparsification lowers training and decoding cost while preserving generation quality. Full octree sequences grow rapidly with depth; OctLLM exploits redundancy at fine levels by randomly emptying penultimate-level nodes and omitting their descendants, then restores the missing fine occupancy with a 3D U-Net. For the illustrated example the sequence shrinks from 3,751 to 2,297 mesh tokens; on the same four 80GB A100 GPUs peak memory falls from 58GB to 44GB and epoch time from 2.17 to 1.37 hours; under the same ShapeNet airplane protocol a complete-octree model reaches FID 27.54 and KID 0.83 versus 26.31 and 0.80 for the S-Octree; in the completion ablation the 3D U-Net with binary cross-entropy gives the best FID 36.86 and KID 1.60.

Perspective

The work targets object-level 3D generation and 3D understanding inside a single backbone: geometry enters as S-Octree occupancy bytes on a six-level octree with a 64³ finest grid, generation is driven by position-aware <MASK> queries, understanding reads the S-Octree directly, and the final mesh is decoded by a frozen TRELLIS flow model conditioned on the completed sparse occupancy coordinates. Training uses the union of the Objaverse-XL Sketchfab subset, HSSD, ABO, and ShapeNet, giving 195K assets and 585K instructions, with the completion network trained separately on 175K sparse-complete occupancy pairs. For researchers and engineering teams who want to add a new modality without retraining the language backbone, or who need one model to output both shapes and text descriptions, this representation-plus-routing scheme offers a directly reusable interface; the authors have released source code and model weights.

The authors list three points in Appendix E: the S-Octree is variable-length and averages 1,793 tokens across the training set, more than ShapeLLM-Omni's fixed 1,024, increasing training time and GPU memory; geometry passes through two learned predictions, S-Octree generation and 3D U-Net completion, so upstream occupancy errors can propagate or be amplified into holes and other artifacts; and the model does not support appearance understanding or 3D editing, with future work planned on native color and material modeling to remove the dependence on an external 3D generator. In addition, the ablations run on the ShapeNet airplane subset and the authors explicitly state those results are not directly comparable with the full Toys4K evaluation; in understanding, the S-Octree encodes geometry alone, so OctLLM describes shape and part structure and says nothing about appearance, and PointLLM still leads on GPT-ref and the embedding metrics because it reads colored point clouds. Readers can watch how variable-length sequences are compressed, how error is controlled across the two-stage prediction, and whether language preservation still holds once appearance and editing are added.

Sources