Public articles linked to the same research event.
arXiv BitNest introduces a bit-nested speculative decoding framework that embeds a low-precision draft directly into the higher-precision target representation and recovers the target through residual refinement, so draft and target share a single physical weight representation, with the progressive-precision design extended to the KV cache for long-context inference; across multiple 7B–8B edge-friendly LLMs and diverse workloads it reports an average speculative acceptance rate of 95.2% while closely preserving higher-precision model quality, and 1.48–1.61x end-to-end speedup over FP16 autoregressive decoding, with consistently competitive or higher decoding speedup on the LLaMA models supported by all representative self-speculative baselines.
BitNest introduces a bit-nested speculative decoding framework that embeds a low-precision draft directly into the higher-precision target representation and recovers the target through residual refinement, so draft and target share a single physical weight representation, with the progressive-precision design extended to the KV cache for long-context inference; across multiple 7B–8B edge-friendly LLMs and diverse workloads it reports an average speculative acceptance rate of 95.2% while closely preserving higher-precision model quality, and 1.48–1.61x end-to-end speedup over FP16 autoregressive decoding, with consistently competitive or higher decoding speedup on the LLaMA models supported by all representative self-speculative baselines.
BitNest introduces a bit-nested speculative decoding framework that embeds a low-precision draft directly into the higher-precision target representation and recovers the target through residual refinement, so draft and target share a single physical weight representation, with the progressive-precision design extended to the KV cache for long-context inference; across multiple 7B–8B edge-friendly LLMs and diverse workloads it reports an average speculative acceptance rate of 95.2% while closely preserving higher-precision model quality, and 1.48–1.61x end-to-end speedup over FP16 autoregressive decoding, with consistently competitive or higher decoding speedup on the LLaMA models supported by all representative self-speculative baselines.
BitNest introduces a bit-nested speculative decoding framework that embeds a low-precision draft directly into the higher-precision target representation and recovers the target through residual refinement, so draft and target share a single physical weight representation, with the progressive-precision design extended to the KV cache for long-context inference; across multiple 7B–8B edge-friendly LLMs and diverse workloads it reports an average speculative acceptance rate of 95.2% while closely preserving higher-precision model quality, and 1.48–1.61x end-to-end speedup over FP16 autoregressive decoding, with consistently competitive or higher decoding speedup on the LLaMA models supported by all representative self-speculative baselines.