Liquid AI pairs LFM2.5-VL-3B with a 280M-parameter DSpark drafter, reaching up to 3.13x faster decoding on M5 Max and 2.66x on H100
Synopsis
Liquid AI released the LFM2.5-VL-DSpark vision drafter, which reuses its text DSpark architecture: image patches and text tokens are projected into a shared representation, and a 4-layer attention-only drafter predicts blocks of candidate tokens, adding roughly 280M parameters (8.9% on top of the 3B target) and delivering 2.30x to 3.13x faster decoding with 1.56x to 2.62x end-to-end gains on M5 Max, 1.57x to 2.14x and 1.30x to 1.77x on M3 Ultra, and 1.64x to 2.27x end-to-end on H100, shipped with day-one llama.cpp, MLX-VLM, and SGLang integrations.
Interpretation
It extends a text speculative-decoding drafter to vision-language models: the drafter reads the target model's hidden states at a fixed set of tapped layers, and image patches and text tokens are projected into a shared representation beforehand, so the drafter operates on hidden-state vectors of identical dimensionality regardless of input modality and the inference algorithm is unchanged from the text models. Relative to the earlier text-only LFM2.5-DSpark drafters, this provides a working implementation of the same DSpark recipe for vision-language input, showing that the modality difference can be absorbed at the projection stage without changing the speculative decoding algorithm. The text describes the architecture and inference flow and states that the drafter is trained on a mixture of vision-language SFT data; no controlled comparison against the text drafter is reported.
The drafter's size and structure were set by ablation: after ablations across 3, 4, and 5 layers, the chosen design is a simplified attention-only drafter with 4 layers and a block size of 9; training ran 10 epochs on the final mixture, with acceptance improving as training tokens increased before diminishing returns, and a block size of 8 or 9 is recommended at inference depending on hardware. It reports the concrete design trade-offs and training budget rather than only final metrics, so a reimplementer can see the basis for the layer count, block size, and training volume. Grounded in the 3/4/5-layer ablation and per-epoch acceptance measurements; absolute ablation numbers are not tabulated in the text.
The parameter overhead is small: the drafter has about 280M parameters, comprising a 193.0M decoder stack, 21.0M hidden-state projection, 65.5M Markov head, and 6.4k norms plus confidence head, totaling 279.5M, which increases the deployed model's parameter count by just 8.9%. It ties the speedup to an explicit parameter budget, indicating that vision speculative decoding need not come with a large increase in deployment footprint. The text provides a per-component parameter table, which is checkable model-scale data.
Measurements cover on-device and GPU inference across six vision-based tasks (general VQA, text VQA, image captioning, chart VQA, complex reasoning, and multi-turn conversation, following the MMSpec benchmark): MLX on M5 Max gives 2.30x to 3.13x faster decoding and 1.56x to 2.62x end-to-end; llama.cpp on M3 Ultra gives 1.57x to 2.14x and 1.30x to 1.77x; H100 gives 1.64x to 2.27x end-to-end. It reports speedups separately by hardware and task and gives both decode and end-to-end figures, helping readers estimate the gain on their own setup. Based on measurements across two on-device configurations and one GPU configuration over six task types; per-task values and baseline absolute latencies are not given.
Perspective
The result targets deployments using LFM2.5-VL-3B that want lower decode latency, applies to on-device Apple silicon and single-H100 settings, and comes with executable configurations for three inference paths: llama.cpp, MLX-VLM, and SGLang. The text notes that speculative decoding is exact: the target verifies every proposed token, so greedy output equals the target alone, which makes the approach usable where output consistency matters. It does not address reducing vision encoding or prefill time, since the text states that speculative decoding speeds up only decode.
The text does not give per-task speedup values, baseline absolute latencies, or acceptance-rate curves, so the source of variation across tasks is hard to judge; the H100 decode range is written as 20.4x to 2.66x, whose lower and upper bounds run in inconsistent directions and should be checked against the original or later material; and the drafter's vision-language training details, data composition, and absolute ablation results are not expanded, which are directions a follow-up reader might watch.
