EvSpark: Lossless Speculative Decoding for Hybrid DNA Foundation Models
Synopsis
EvSpark is a speculative decoding system for hybrid convolutional-recurrent-attention DNA foundation models such as Evo2: it verifies draft blocks in parallel and restores all three classes of inference state by selecting retained intermediate states without replay, reaching 2.96x on 43 real-sequence prompts and 3.27x including five synthetic controls on an Evo2 7B 48-prompt benchmark with three training seeds, retaining 1.84x-2.43x at 262k context, and yielding 2.18x-2.46x on real sequences and 2.51x-2.78x on the full suite for 20B and 40B targets.
Figure 1: EvSpark architecture and speculative decoding. (a) Frozen Evo2 supplies intermediate features and offline teacher supervision; the 7B configuration extracts layer 27 (HCL). (b) A projected window of up to 64 target states forms the K/V prefix of each of two causal drafter layers (d = 1024). Anchor and mask embeddings produce all draft logits in one trunk pass; a previous-token Markov bias then supports serial sampling. Input and output embeddings are tied and frozen; projections and the two-layer trunk are trained. Dashed paths denote offline supervision using teacher- forced draft probabilities before the sampling transform T; serial sampling is shown for decoding. The confidence head does not control decoding. (c) The same Evo2 verifies the old anchor and six drafts in one block. Three drafts are accepted; a sampled correction becomes the next unconsumed anchor. (d) State returns to block index 3: shorten the readable KV prefix, select the last two outer-FIR inputs u2, u3, and retain IIR state s3. The feature buffer commits the same consumed prefix for the next round; no target computation is replayed.
bioRxiv · Page 4Interpretation
It introduces replay-free verification for a hybrid target: one initial-state block forward plus a common state-slicing protocol restores FIR, IIR, and KV state consistently. Prior state recovery for speculative decoding focused mainly on truncating a Transformer KV cache, or on state recovery for state-space or hybrid models separately; this work brings StripedHyena2's finite convolution windows, modal IIR states, and attention caches under a single slicing contract. The paper proves Lemma 2 (block-forward contract), Lemma 3 (state restoration), and Theorem 1 (sequence-level exactness), and reports that an identity-drafter control improves from 0.54x to 1.05-1.16x native throughput once replay is removed; slicing itself takes about 0.62 ms and replaces a second target forward costing about 22 ms.
It designs a low-latency parallel, feature-conditioned distilled drafter that proposes a whole block in one trunk forward pass. The drafter adapts a DSpark-style architecture to DNA: it conditions on the anchor and a W=64 window of target hidden states, uses a two-layer causal attention trunk of width d=1024 with a full-matrix Markov bias and a confidence head, and the paper evaluates feature injection sites and training budget on a hybrid target. Draft cost is 2.5-4.3 ms over the tested lengths and speedup rises from 1.77x to 3.07x; in comparison, First-8 falls to 0.85x at gamma=12, First-16 is below native speed throughout, and the independent 1B autoregressive drafter reaches tau-bar=8.33 at gamma=12 (versus 4.37 for EvSpark) but takes 395 ms to draft a block and achieves 0.43x.
It shows the implementation and training recipe transfer across scale and operating conditions, and characterizes operating limits. The same width-1024 drafter recipe is retrained for Evo2 20B and 40B (capturing layer 20 of 24 and layer 45 of 50), with measurements spanning long context, sustained generation, batching, and multiple candidates. 20B/40B give 2.18x-2.46x on real sequences and 2.51x-2.78x on the full suite; draft latency stays at 3.6-4.0 ms across targets, 9.3% (40B) versus 18.4% (20B) of a native token step; at 262k context E. coli and B. subtilis give 1.84-2.43x; with C=1/2/4 candidates tau-bar rises from 4.37 to 5.40 while all-48 speedup falls from 3.07x to 2.64x.
In a complete regulatory-DNA design workflow case study, the decoding speedup translates into shorter end-to-end design time. Prior work mostly reports single-sequence decode speedup; this study measures complete design time under a fixed predictor-guided search policy, including prefill, candidate generation, cache handling, scoring, and final checking. With native batch 8 and EvSpark batch 4 calibrated on the same devices, mean design time fell from 13.63 to 8.73 minutes, with paired speedups having a median of 1.57x (range 1.42x-1.70x); all four outputs from each backend met the prespecified predictive qualification criterion (search-ensemble and final-checker AUROC at or above 0.90), with mean checker AUROCs of 0.975 for native and 0.977 for EvSpark.
Perspective
The results target single-sequence decode latency, and the token-generation benchmarks exclude prefill; the regulatory-design case study additionally covers prefill, batched independent candidates, and scoring under a fixed search policy. Acceleration is measured from 7B to 40B targets, at 51k to 262k context, and over 32k generated tokens, but the 262k tests require the 48 GB 4090 variant and larger targets are evaluated on one H20 with SDPA. For readers aiming to reduce latency in long-sequence DNA generation or design workflows, the work offers a reusable state-slicing contract and drafter recipe; performance under arbitrary mixtures of concurrent requests is explicitly left for future evaluation.
The exactness guarantee assumes exact arithmetic; the bf16 implementation has near-tie greedy cases and measurable sampling deviations, most visibly in repeats, and a full-model fp32 control and a sequence-level bound on these numerical differences remain open. The prespecified sequence-statistic equivalence criterion was not met, and the post hoc intervals include zero without establishing equivalence. The 48-prompt suite covers a limited set of regions with overlapping loci and unequal representation of biological diversity, and decode trajectory variation can shift a suite mean by up to 0.44x. The 20B/10M cell received no ncRNA windows despite their planned allocation, and hg38 is revisited 1.50x at 40B/80M, so the realized training mixture differs from the plan in these cells. The regulatory-design evaluation covers two patterns and two design seeds, its quality and diversity summaries describe observed outputs, predictive qualification is assessed computationally, and regulatory function has not been experimentally validated.
