Allspark lets a weak teacher pass reasoning to frozen strong models via alternating chains of thought, lifting ARC-AGI-2 accuracy from 78.1% to 81.6%
Synopsis
The work introduces Allspark, a training and inference framework in which a weak teacher and a frozen copy of the same model alternate reasoning segments during training while only the teacher is updated from weak-model rollouts, and at inference a stronger frozen student replaces the frozen partner; across Qwen3-1.7B/4B math and reasoning tasks and Inkling-Small teacher evaluations on ARC-AGI-2 with Inkling, Kimi-K2.6, and Nemotron-3-Ultra students, accuracy gains are observed along with an explicit accuracy-token tradeoff.
Interpretation
Allspark places the weak-to-strong transfer interface in shared text: during training a weak teacher and a frozen copy of the same model take turns extending one reasoning trace, the frozen copy produces the final answer, and the task reward updates only teacher-generated reasoning tokens while student tokens and inserted boundary markers are masked out. Unlike on-policy distillation methods that sample from and update the strong student (Direct-OPD, W2S-OPD, OPRD), no strong-student rollouts are used in training; unlike frozen-student guidance methods such as weak-to-strong search, CoWeST, and Co-LLM, the output is a reusable trained weak teacher rather than a one-off search or preference signal. The paper gives a formal definition of alternating reasoning, an expected continuation-value expression for teacher chunks, and full training protocols for Qwen and Inkling (LoRA, masking rules, asynchronous bounded lag, length ceilings).
In controlled Qwen3 experiments, steering a frozen Qwen3-4B student with the Allspark teacher is roughly neutral in math and positive in reasoning: 96.1 versus 96.9 alone in math and 82.8 versus 81.2 alone in reasoning, while an untrained teacher (93.0/79.7) and an ordinary-RL teacher (89.8/79.6) both fall below the student alone in both domains. This ablation separates whether the teacher was trained in the alternating setup from whether it was merely RL-trained, indicating the gain is not from arbitrary weak-model prompting but from a teacher trained alongside a frozen partner. 128 development questions per domain with one sampled answer each, and paired conditions share questions, student decoding settings, and reasoning/answer limits; the authors themselves call this a sign of life, so it is small-scale controlled evidence.
On 96 ARC-AGI-2 development tasks, the Inkling-Small teacher steering a frozen Inkling student reaches 81.6% mean pass@1 versus 78.1% for the student alone, and reaches 79.2% at 10.8k retained tokens, exceeding the student's best observed 78.1% at 12.3k tokens. This moves weak-to-strong transfer onto a benchmark known for abstract reasoning and reports accuracy together with retained tokens, making the accuracy-token tradeoff a comparable object rather than a single accuracy point. Each configuration uses 96 tasks with three attempts per task and shared task-attempt seeds, reporting mean pass@1 over 288 attempts; the authors note token totals are counted in each writer's native tokenizer and are not a common-tokenizer budget.
The same Inkling-Small teacher transfers directly to other model families: Kimi-K2.6 improves by 11.5 percentage points while using 7.7% fewer retained tokens, and Nemotron-3-Ultra improves by 17.7 and 5.2 percentage points in medium and full thinking modes respectively, with increased token use. Cross-family transfer relies on text handoffs, where each reasoning chunk is decoded with its writer's tokenizer, stripped of model-specific control markers, and re-encoded with the receiving model's tokenizer into its native reasoning channel, so no shared vocabulary or architecture is required. Three operating points each use 288 attempts with paired task-attempt seeds and report task-level bootstrap intervals; the authors state that benefits vary with the target student's inference settings and that token totals are not a common-tokenizer budget.
Perspective
The result targets research and engineering settings that want to reuse RL gains from a weak model without generating strong-model rollouts: training needs only weak-model compute, and once trained the teacher can be fixed and applied to multiple frozen strong students, provided the interface exposes an interruptible reasoning stream. The reported scope covers Qwen3-1.7B/4B math and reasoning tasks and Inkling-Small teacher evaluations on ARC-AGI-2 with Inkling, Kimi-K2.6, and Nemotron-3-Ultra students; cross-family transfer re-encodes reasoning text with the receiving tokenizer, so a shared vocabulary is not required. The authors position this route as lowering the barrier to RL research and hope it extends to alignment and AI safety research.
The authors state that experiments used limited compute and that evaluation covers a restricted set of models, tasks, and training runs, so broader studies are needed to establish robustness and generality; inference requires serving both teacher and student, and the extra model calls and handoffs may increase cost and latency, so fewer retained tokens do not necessarily translate into lower deployment cost; the implementation relies on interrupting and resuming a shared reasoning stream and remains untested for non-thinking models or interfaces that do not expose a suitable reasoning stream. The Qwen portion is a single-seed controlled study that the authors call a sign of life; in cross-family evaluation the target students use their native thinking renderers, so equal numerical effort or equal reasoning compute across families is not assumed, and token totals are counted in each writer's native tokenizer rather than a common-tokenizer budget. The case studies are explicitly described as illustrating exchanged content and do not establish the causal effect of an individual teacher span.
