Braco reframes visual-token compression from picking tokens to re-parameterizing them, reaching 95.2% normalized accuracy at 23–64x compression while cutting prefill FLOPs by 84.2%–86.7%
Synopsis
The work recasts extreme visual-token compression as a token-parameterization problem, separating basis transformation and structured truncation (which fix the retained subspace and compressibility) from coordinate organization (which affects optimization and cross-modal alignment, i.e. learnability), and designs Braco, a lightweight four-step coder combining transform-basis truncation, input-independent basis-coordinate embeddings, budget-dependent orthogonal re-parameterization, and learned spatial residual tokens from lightweight pooling; experiments show Braco forms the favorable empirical accuracy–efficiency frontier under 23x–64x compression and remains competitive at 144x, reaching 95.2% accuracy while reducing prefill FLOPs by 84.2%–86.
Interpretation
The paper introduces a token-parameterization view of extreme visual-token compression, splitting it into basis transformation plus structured truncation, which determines the retained subspace (compressibility), and coordinate organization, which determines how coordinates inside that same subspace are arranged (learnability), formalized as unified functionals. Prior work largely treated compression as token selection (pruning/merging) or learned resampling; this work makes 'what to retain' and 'how to represent what is retained' two separately diagnosable design axes. The paper gives explicit functionals for compressibility (energy retention and task-direction readability) and learnability (statistical conditioning and geometric compatibility), with derivations and threshold conditions in the appendices; diagnostics are run on frozen token grids and CelebA linear probes.
The resulting Braco has four steps: a transform-domain backbone from DCT low-frequency block truncation, input-independent basis-coordinate embeddings, budget-dependent orthogonal coordinate organization (coefficient coordinates at very small budgets, inverse-DCT coarse grid once the backbone is larger), and spatial residual tokens from lightweight sparse pooling. Unlike pure low-pass transform coding such as Fourier-VLM's inverse-DCT reconstruction, Braco restores the token identity lost after the transform and uses spatial residuals to recover localized evidence dropped by low-pass truncation. Ablations show the hybrid c2s5 reaches 93.2 normalized accuracy at 9 tokens versus 89.8 for backbone-only c3s0 and 92.0 for residual-only c0s9; at 16 tokens, replacing the DCT backbone with spatial/Haar costs 2.3–3.0 accuracy points, replacing the selected coordinate organization costs 2.5–3.1 points, and replacing Polar Fourier embeddings costs 0.8–0.9 points.
In end-to-end evaluation under matched extreme token budgets, Braco forms a favorable accuracy–efficiency frontier: relative to the 576-token Vanilla model it reduces full-pipeline prefill FLOPs from 8.67T to 1.09–1.37T while retaining 91.2–95.2 Vanilla-normalized accuracy across 4–25 tokens. It attains the highest normalized accuracy among evaluated methods at 25, 16, and 9 tokens and stays within 0.2 of QueCC at 4 tokens; against QueCC it matches within 0.1 accuracy at 16/9 tokens while reducing latency by about 36%. The main table reports per-benchmark scores and normalized accuracy on GQA, MMBench (EN/CN), MME, POPE, ScienceQA, TextVQA, and MMVet, with all baselines sharing the retraining/evaluation harness and matched budgets; at the compressor level Braco matches QueCC-level accuracy with 16.6x lower module latency and 78.8x fewer compressor FLOPs.
Braco remains effective at larger inputs and with a smaller LLM: in the 2880-input-token Vicuna-7B setting it keeps 98.1 normalized accuracy while cutting full-pipeline FLOPs from 40.57T to 3.45T and latency from 261.08ms to 44.36ms; on Qwen2.5-3B it retains 93.1/92.6 accuracy at 729/1024 input tokens and 90.8 accuracy at 3645 tokens. The paper notes that the 576-token setting mainly tests whether accuracy survives a very small retained budget, whereas as the visual interface grows the same small budget removes a much larger uncompressed prefill burden and yields larger end-to-end savings. Table 8 gives per-benchmark scores, FLOPs, and latency for Vicuna-7B and Qwen2.5-3B at 576/2880/729/1024/3645 input tokens; the training-time table shows Braco is no slower than QueCC.
Perspective
The result targets single-image multimodal inference in vision-encoder–LLM pipelines, suited to mobile inference, low-latency interactive agents, and long-context multimodal settings that must operate within dozens of visual tokens; for engineering practice that wants to swap the visual interface without retraining the downstream LLM, the paper offers a lightweight compressor whose cost can be measured independently and a budget-dependent coordinate rule. It also provides a label-free calibration procedure: estimate the crossover on the visual-interface training distribution, or re-estimate with a small unlabeled sample when the target domain changes, adapting the backbone–residual split, coordinate organization, and structured basis to new encoders or resolutions.
The paper states that its evaluation scope is mainly LLaVA-family pipelines and standard single-image benchmarks, that the coordinate choice is calibrated for the tested encoders and resolutions, and that new backbones, video inputs, dense localization tasks, domain-specific distributions, irregular region features, or dynamic visual tokens may require a small calibration sweep; three-seed accuracy comparisons are reported for the key Braco–QueCC and Braco–TokenPacker settings, while the remaining accuracy comparisons use one trained checkpoint per setting. In addition, some tables and figures in the loaded text (such as batch-scaling latency and parts of the cost breakdown) appear with empty values, so those specific numbers cannot be restated in this summary; readers interested in batch-scaling and stage-cost details should consult the original appendices.
