Imprint Reader uses SMaRT to read frozen weight updates into natural language, and MetaEdit turns those descriptions into behavioral intervention
Synopsis
The work introduces Imprint Reader: Semantic Mount-and-Read Tuning (SMaRT) mounts a frozen weight update onto the Reader and uses an anchor-free meta-query to elicit a natural-language description, with no-change and random-perturbation controls discouraging unsupported claims; on held-out updates the joint Reader reaches judge-based Pass@100 of 2% for knowledge and 16% for behavior, while the Reader's differentiable proxy for a target behavior transfers back to the original model through MetaEdit, raising measured harmful-prompt refusal from 57.9% to 64.1% at a 0.5% pruning rate and lifting BFCL Overall from 41.69% to 44.60% without target-task training data.
Interpretation
It demonstrates the feasibility of reading natural-language descriptions out of frozen weight updates: under anchor-free meta-queries the Reader describes the factual knowledge or behavioral tendency carried by unseen updates, reaching judge-based Pass@100 of 2% for knowledge and 16% for behavior. Earlier weight-space methods mostly predict coarse attributes such as accuracy or the fine-tuning task, and the few that verbalize weight differences are confined to narrow purpose-built domains and readily fabricate descriptions for uninformative updates; this work extends readout to specific factual propositions and behavioral tendencies and trains the Reader to abstain when an update carries no recoverable semantics. Training data contain 8,592 knowledge items and 8,592 behavior items; updates are built by a 64-step inner-loop LoRA procedure with ranks up to 256, and the Reader is optimized on eight GPUs with 64 episodes per batch (24 knowledge, 24 behavior, 8 no-change, 8 random-perturbation). Evaluation uses 100 held-out knowledge and 100 held-out behavior items with 100 stochastic generations each, judged by Qwen3-30B-A3B-Instruct-2507 under a fixed scoring prompt that requires the cited span to occur in the generated response.
The Reader's target-conditioned likelihood provides a differentiable proxy for the gap between a specified target behavior and a candidate weight update, and its coordinate-aligned gradients transfer back to the original model through MetaEdit for structured intervention. Prior weight readout ends at monitoring and offers no path to acting on what is decoded; here the Reader shares parameter coordinates with its parent, so the readout signal becomes an intervention signal, and the intervention unit is a projection output row rather than an individual scalar. Rows are scored across 2,088,960 output rows, retaining the top 50,000 by contrastive gradient magnitude and then selecting rows with the most negative inner products with the gradients; the safety experiment classifies 1,803 prompts (1,253 harmful) by fixed refusal patterns and reports garbling separately.
Under a safety-maintenance target, Reader-guided row pruning raises measured harmful-prompt refusal from 57.9% to 64.1% at a 0.5% pruning rate, while a refusal-relaxation target moves refusal in the opposite direction, and most baselines do not separate the two targets. The result connects a behavior specified in words directly to a parameter-level intervention, and MetaEdit uses target self-report sentences plus fixed control sentences rather than collections of harmful or harmless training prompts; computing the gradients on the original model instead of the trained Reader also fails to produce the intended separation. Compared against SetDiff, WANDA, ActSVD, and Random at identical row budgets; WANDA and ActSVD produce substantial garbled output after pruning, whereas each MetaEdit condition has a near-zero garbling rate.
Without target-task training data and without inference-time behavioral instructions, MetaEdit improves mathematical reasoning and tool use: GSM8K from 94.77% to 95.00%, MATH-500 from 82.8% to 83.0%, and BFCL Overall from 41.69% to 44.60%, while increasing backtracking and sub-goal expressions. The authors call this description-driven intervention vibe alignment; unlike Direct Prompt, which keeps the behavior description in the system prompt, the edited models need no inference-time behavioral instruction, and a procedure-level self-report outperforms an outcome-level one (MetaEdit (Broad)) on mathematics and BFCL Overall. Mathematical and BFCL coefficients were fixed before evaluation and test scores were not used to select them; BFCL covers 5,106 cases across 22 subsets with the official Overall aggregation; backtracking rises from 0.327 to 0.687 and sub-goal expressions from 5.390 to 6.015 per 1000 generated tokens.
Perspective
The work targets controlled updates: each update corresponds to a single knowledge item or a single behavioral tendency, the Reader and the update builder are both initialized from the same post-trained Qwen3-14B checkpoint, and updates are LoRA-form. It applies to research and engineering settings that want to specify and apply behavioral change in natural language without target-task training data, for example localizing and pruning safety-refusal behavior or adjusting reasoning and tool-use style. The authors state explicitly that the Reader is not evaluated on reconstructing the training examples behind an update, so the results do not constitute a training-data extraction capability.
A careful reader will still watch how free-form readout reliability improves across updates and whether judge-based Pass@100 holds at the same order of magnitude on more models and update types; for safety pruning, benign-prompt garbling and over-refusal results are not fully reported in the main text, and the authors note that this part is left missing rather than set to zero; the mathematics and BFCL gains are not uniform, with MetaEdit (Broad) leading on Multi-turn, Direct Prompt leading on Live and Hallucination, and the unedited model best on Non-live; and what the shared-parent-checkpoint setting between Reader and update builder implies for cross-model transfer remains an open question.
