BiasReducer edits only the reward head, lifting five reward models by 8.3/18.0/6.9 points on average across three bias benchmarks
Synopsis
The work proposes BiasReducer, a lightweight framework that edits only the linear reward head without retraining the reward model: an SAE-style encoder with semantic supervision learns internal representations for predefined attributes such as length, confidence, and sycophancy; the framework then determines the direction and amount by which to adjust the reward head for each attribute and stores these as an edit bank; for a new dataset it ranks attributes by their influence on reward scores and applies the matching edits. Across five reward models, BiasReducer-M improves RM-Bench-Hard, Arena-StyleConflict, and JudgeBiasBench by 8.3, 18.0, and 6.
Interpretation
BiasReducer splits bias mitigation into three stages: an SAE-style encoder learns interpretable coordinates corresponding to predefined attributes, the framework learns how to reduce the reward head's dependence on each attribute (direction and amount), and it then selects which edits to apply for a new dataset. Existing editing methods require the target bias to be specified in advance and apply a fixed correction, while training-based methods need retraining and additional data; this work separates learning how to correct from deciding which attribute to correct, so the same edits can be reused without retraining or pre-specifying the bias. The paper provides a full derivation (including the relation to linear concept erasure, LEACE) and ablations: removing behavioral supervision drops attribute-axis correlation from 0.55 to 0.17, and removing intervention supervision drops sign accuracy from 0.99 to 0.75.
Across five reward models (Qwen3-1.7B, GRM-3B, Mistral-7B, InternLM2-7B, GRM-8B) and three bias benchmarks, BiasReducer-S improves the benchmarks by 6.2, 15.7, and 6.0 percentage points on average, and BiasReducer-M by 8.3, 18.0, and 6.9, beating the original reward model in every model-benchmark combination. Against the stronger training-based baseline RRM, BiasReducer-M adds a further 4.1, 13.2, and 2.3 percentage points on the three benchmarks; against a fixed sycophancy edit or the strongest source-selected fixed edit, dataset-specific selection yields larger gains. The main table reports per-cell numbers for 5 models x 3 benchmarks; Arena-StyleConflict is built from Arena Human Preference 140K and contains 1,000 pairs, and results are averaged over seeds 0 and 1.
Both edit direction and edit amount need to be reward-aware: using the covariance-based direction instead of the SAE decoder vector raises average gains from 0.5/0.4/0.6 to 6.2/15.7/6.0 percentage points, and per-attribute edit strengths beat one globally shared strength (on JudgeBiasBench with Mistral-7B, 8.4 to 11.8 points). This indicates that identifying an interpretable attribute is not sufficient; the edit direction must also reflect how that attribute relates to the reward representation, and different attributes benefit from different edit strengths. Direction and strength comparisons are reported for each of the five reward models with mean rows; attribute selection with 100 examples matches the full-pool decision in 89% of trials, rising to 97% with 200 examples.
The gains transfer downstream: with the edited Mistral-7B reward head in best-of-n selection and GRPO training, Qwen3-4B and Llama-3.2-3B-Instruct show lower verbosity and sycophancy and fewer words, while quality judged independently by Claude Sonnet 5 remains comparable. This connects reward-model-level robustness to policy-level behavior change, suggesting that reducing reliance on superficial attributes need not come at an obvious cost to response quality. Downstream experiments fix the candidate pool or training configuration and change only the reward head; quality, verbosity, and sycophancy are scored separately by an independent evaluator, and training results are evaluated after 80 rounds on 100 held-out prompts.
Perspective
The result targets RLHF and response-selection pipelines that use scalar reward models: editing touches only the linear reward head while all other parameters stay fixed, so it applies to already-trained reward models with a linear head. The method presupposes seven measurable attributes (length, verbosity, formatting, emoji/exclamation, confidence, hedging, sycophancy) and relies on source preference data to learn representations and edit directions and on a source validation split to choose edit amounts; for a new dataset it uses only responses, hidden representations, and original reward scores, not preference labels. For teams wanting to reduce reward-model reliance on superficial attributes without retraining or pre-specifying a bias, this offers a reusable edit bank plus dataset-specific selection.
The attribute set is seven predefined categories, so whether other superficial attributes can be handled by the same edits remains open; edit selection is transductive, requiring the new dataset's responses and original reward scores, and how to operate in a fully online, one-at-a-time setting is still to be examined; downstream quality is judged by Claude Sonnet 5 as an independent evaluator, so whether the evaluator's own preferences shape the conclusions deserves continued attention; in addition, several ablation tables in the appendix appear empty in the loaded text, so per-model breakdowns cannot be checked and only the means and representative results reported in the body can be relied on.
