Skip to main content
Back to timeline
arXivSource publication:

Rewriting public documents into about 10k synthetic samples lifts a 35B student from 13.7% to 24.6% on CL-bench, near a trillion-parameter frontier model

Synopsis

The work proposes a human-annotation-free synthetic supervision pipeline: it deeply rewrites 3,515 public documents via entity renaming and numeric perturbation, has LLMs generate questions and rubrics that require reasoning over the document, and admits only samples that genuinely depend on the document via a with-document versus without-document rubric gap check, yielding about 9,625 training samples; SFT raises Qwen3.6-35B-A3B on CL-bench from 13.7% to 22.8%, a subsequent rubric-reward RL stage reaches 24.6%, comparable to Qwen3.8-2.4T at 23.9%, with broad transfer to long-context understanding, instruction following, and reasoning, while code generation and knowledge stay mostly flat.

AI-generated editorial illustration: Learning to Learn from Context: Synthetic Training from Perturbed Public Documents

Interpretation

The pipeline turns public documents already consumed in pretraining into context-learning supervision through deep rewriting plus a gap check: rewriting includes proper-noun renaming, numeric perturbation, section renumbering, and list-order shuffling, and the gap check compares rubric pass rates with and without the document, admitting only samples that are answerable from the document, not trivially answerable from memory, and genuinely require the document. Prior synthetic long-context supervision either grounds in real documents, which carries memorization risk, or synthesizes the documents themselves, which yields logically simpler and more homogeneous contexts; this work keeps real documents but perturbs them so the teacher cannot answer from parametric memory and must extract and reason over the document content. The paper reports a memorization audit in which a model answered questions about a well-known technical standard perfectly with no document in context but scored nearly zero once the standard was rewritten; the pipeline filters out 32.5% of generated samples before retry, and an ablation shows training on unfiltered data reaches 17.75% on CL-bench versus 19.17% with the gap check.

Without any human annotators, the pipeline generates 9,625 training samples from 3,515 public documents, covering 16 of CL-bench's 18 subcategories, with sources including IETF RFCs, SEC filings, court opinions, arXiv papers, public handbooks, wikis, and encyclopedias. Compared with human annotation, where an expert must design a document of tens of thousands of words, verify the answer against it, and specify what a correct answer must contain, this pipeline automates supervision construction and documents its sources and subcategory distribution. The paper provides a per-subcategory document table (e.g., Humanities 376, Science 370, Technical Standards 250, Healthcare 244, totaling 3,515) and states admission criteria of at least 4,000 characters and global URL deduplication.

Training gains are substantial and partly transfer beyond the target benchmark: SFT raises Qwen3.6-35B-A3B on CL-bench from 13.7% to 22.8% and on CL-bench Life from 8.4% to 12.6%, and the subsequent rubric-reward RL stage reaches 24.6% on CL-bench and 13.8% on CL-bench Life; long-context understanding (AALCR 62.6 to 69.8), instruction following (IFBench 57.7 to 71.7), and reasoning (ARC-AGI-1 47.4 to 63.8) improve substantially, code generation is essentially unchanged, and knowledge declines slightly (SimpleQA 20.9 to 18.7). The result indicates that synthetic data from public documents can improve the ability to use context without teaching world knowledge, and that the gains are not confined to the public-document types present in the training data. Results come from SFT and SFT+RL comparisons on the same student model, with several public reference models listed (GLM-5.2 20.9, Qwen3.8-2.4T 23.9, Kimi-K3 27.0, Hy4-preview 27.6); RL drops ARC-AGI-1 from 63.8 to 51.8, indicating that RL's additional benefit is concentrated on the target capability.

Analysis shows most of the gain occurs on fictional-context tasks: by the paper's rough classification of 500 CL-bench contexts, 836 tasks are based on public documents and 1,063 on fictional or unclear contexts, and SFT+RL raises the fictional group from 17.5% to 32.9% while the public-document group only rises from 9.0% to 14.1%. This suggests the training data teaches a transferable reasoning mode rather than memorization of specific public-document content, consistent with the design intent of perturbing documents to suppress memorization. The classification is the authors' own rough audit rather than an official split, as the paper explicitly notes; the authors also raise the possibility that public-document tasks are simply harder.

Perspective

The result targets settings where a model must read a document and reason over it, such as question answering over internal documents and manuals, reasoning over retrieved passages, interpreting tool outputs and logs, and checking whether a plan complies with a specification; the intended users are research and engineering teams that want to build long-context supervision from public documents and can use a strong teacher model. The paper covers 16 of CL-bench's 18 subcategories, omitting Workflow Orchestration and Operational Procedures because they are dominated by agent-style, conversation-form task traces that are hard to source from public documents. Data-scaling ablation shows performance rising monotonically from 100 to 9,625 samples with no saturation in the tested range, suggesting the current roughly 10k pool may not exhaust the model's learning capacity.

The paper states the rewriting is not guaranteed to be perfect, only sufficient to make the original and rewritten documents detectably different; the gap check filters 32.5% of samples before retry and its judgment relies on a judge model. SFT outcomes are strongly teacher-dependent: on identical questions, students trained on answers from different teachers span 19.7 to 22.8 on CL-bench, with the strongest teacher Kimi-K3 (27.0) yielding the weakest student (19.7) and the mid-ranked Qwen3.8-2.4T yielding the best (22.8), and the authors leave a more detailed analysis to future work. In the RL stage, reasoning traces and answers grow longer and the model has a higher failure rate (about 2%) on rubrics requiring it to avoid hallucination, which the authors read as possible reward hacking; adding a length penalty drops CL-bench to 20.91, and adding the overall pass rate does not exceed SFT. The reproducibility statement notes that the official CL-bench judge prompt triggers content moderation, so the authors replace full-width brackets with half-width ones and state this may change judge behavior and cannot guarantee identical results. In addition, the public versus fictional classification of contexts is the authors' rough audit of 500 contexts, not an official split.

Sources