Training-free SFT-as-context lets a parent model use the SFT model's answer as context, preserving fine-tuned and general capabilities across 19 model pairs and 11 benchmarks
Synopsis
The work introduces SFT-as-context, a training-free inference scheme in which the SFT model first answers the query and the parent model then generates the final answer using that response as context, acquiring fine-tuned capabilities through in-context learning while keeping its own general capabilities; across 19 parent-SFT model pairs and 11 benchmarks it stays close to the SFT models on fine-tuned capabilities (gaps of only 2.2 and 2.1 percentage points on AIME 2024 and LiveCodeBench, and 2.0 macro MAE on NutriBench-English) and within 2.2 percentage points of the parent models on general capabilities on average, and it can answer queries that need both capability types even when neither model succeeds alone.
Figure 1: SFT-as-context recovers both fine-tuned and general capabilities. The task is to predict the nutritional values (carbohydrates, fat, protein, and energy) from a meal description, in the same language as the given query. Only the carbohydrate value is shown here as an example. Subfigure (a) shows that an SFT model trained on English data excels in fine-tuned capability. However, subfigures (b) and (c) show that it fails at acknowledging non-food queries and responding in the same language as the query. When a query requires both fine-tuned nutrition estimation capability and general multilingual capability, only SFT-as-context satisfies both requirements. Subfigure (d) shows that the parent model attends more to useful English SFT responses than to hallucinated SFT responses to non-food queries, suggesting that selective attention helps the parent model use the SFT response through in-context learning. We discuss details of the attention analysis in Section 4.2 .
arXivInterpretation
SFT-as-context is a training-free, text-interface-only inference method: a first pass has the SFT model generate a response, and a second pass gives the parent model the original query, that response, and fixed SFT-as-context instructions to produce the final answer. Prior forgetting-mitigation approaches change the fine-tuning process (regularization, replay, constraints on parameter updates, choice of training method) and cannot be applied without further training to ready-to-use checkpoints on platforms such as Hugging Face; this method leaves training untouched and needs neither model weights nor the fine-tuning process, so it also applies to closed-source models. The method and prompts are specified in the main text and Appendix B.2; the authors report evaluation over 19 parent-SFT pairs and 11 benchmarks and describe fallback conditions (e.g., when the SFT response exhausts its budget or shows pathological repetition, the parent answers alone), with a MathIF fallback rate of 3.73% (376 of 10,080 responses).
SFT substantially damages general capabilities, lowering general-capability scores by 33.2 percentage points on average (for example, Qwen3-8B's nutrition model drops from 99.6% to 10.7% in acknowledging non-food queries), while SFT-as-context restores general capabilities close to the parent level while retaining fine-tuning gains. The authors quantify this with a recovery rate: SFT-as-context preserves 92.0% of fine-tuning gains on average and recovers 94.2% of the general-capability gap, with recovery rates of 95.0% and 94.1% on AIME 2024 and LiveCodeBench. Results are presented as three-way comparisons of parent, SFT, and SFT-as-context on the same benchmarks, spanning mathematics, coding, and nutrition tasks plus general-capability benchmarks (IFEval, MGSM, SQuAD2.0, FQuAD2.0, CoQA, ACPBench).
SFT-as-context can satisfy fine-tuned and general requirements within a single response, which neither the parent nor the SFT model achieves alone; on MathIF, LiveCodeBenchIF, and NutriBench-Non-English it even outperforms an oracle router that picks the better of the parent and SFT responses. Earlier inference-time combinations of pretrained and fine-tuned models (such as EFT and Proxy-Tuning) rely on token-level probabilities or logits and do not specifically target forgetting, while routing answers each query with only one model and so cannot meet both requirement types. Joint success rates show SFT-as-context averaging 30.7 versus 27.2 for the oracle router on MathIF and 20.7 versus 18.7 on LiveCodeBenchIF; on NutriBench-Non-English it reaches recovery rates of 84.6% for nutrition estimation and 95.2% for language consistency.
A Bayesian framework for in-context learning yields error guarantees: under stated assumptions, SFT-as-context has lower error than the parent on in-domain queries and stays close to the SFT model, while on out-of-domain queries its error difference from the parent is upper bounded; attention analysis shows the parent attends more to useful SFT responses than to irrelevant ones (37.44% versus 15.11%, a 22.33 percentage point average difference). This supplies a theoretical characterization and mechanism-level empirical evidence for the claim that the parent selectively acquires fine-tuned capabilities from the SFT response while preserving general capabilities, rather than only reporting performance. Theorems and proofs appear in Appendix K and rely on assumptions including domain posterior bounds, output separability, in-domain realizability, and out-of-domain posterior retention; the attention analysis uses Gemma-4-E4B-it and its SFT version with 100 randomly sampled nutrition queries and 100 non-food queries, averaged over response tokens, layers, and attention heads.
Perspective
The method targets deployment settings where a parent model and an SFT checkpoint already exist: any standard text input-output interface suffices, so it applies to closed-source models and allows a stronger non-parent model to produce the final answer. The authors report gains across three fine-tuning domains (mathematical reasoning, coding, nutrition estimation) and general capabilities including instruction following, multilingual tasks, question answering, and planning; on queries needing both capability types it merges them into one response. On overhead, the second pass uses at most 13.0% as many tokens as the first for math and coding, while nutrition second-pass responses are longer but still range from a few hundred to slightly over one thousand tokens. The authors note the same framework could naturally extend beyond SFT to other post-training methods such as reinforcement learning, which they leave as a future direction.
Several open questions remain for a careful reader. The theoretical results rest on explicit assumptions (domain posterior bounds, output separability, in-domain realizability, out-of-domain posterior retention), and the text does not verify how broadly these hold for real models. The attention analysis covers only one model pair, Gemma-4-E4B-it, with 100 queries per category, so generalization to other scales and families is untested. The safety case shows that when fine-tuned and general capabilities directly conflict, SFT-as-context follows the fine-tuned capability (average recovery of 16.44 on HarmBench), and the authors leave optimizing the instructions to prioritize specific general capabilities to future work. In addition, nutrition second-pass responses are markedly longer than the first pass (198.1% ratio on NutriBench-English), so latency-sensitive settings need per-task evaluation; the authors also report that RL post-training induces negligible forgetting, so they do not apply the method there.
