Public articles linked to the same research event.
arXiv Using Qwen3-1.7B with four RL teachers from the same initialization (mathematics, code, instruction following, science), this study compares gradients, optimizer updates, and task learning curves, with SmolLM3-3B diagnostics: loss averaging implicitly weights longer responses more, Adam's first moment aligns updates across teachers (cosine 0.83 between teachers, 0.96 between averaging rules), BF16 rounding hides about 97% of FP32 master-weight changes and shows only 7-11% of BF16 weights changing, and the top-64 intersection KL gradient closely matches the full-vocabulary gradient yet its mathematics accuracy is 2.6 points above sampled-token policy gradient under response averaging and 2.1 points below under global token averaging.
Using Qwen3-1.7B with four RL teachers from the same initialization (mathematics, code, instruction following, science), this study compares gradients, optimizer updates, and task learning curves, with SmolLM3-3B diagnostics: loss averaging implicitly weights longer responses more, Adam's first moment aligns updates across teachers (cosine 0.83 between teachers, 0.96 between averaging rules), BF16 rounding hides about 97% of FP32 master-weight changes and shows only 7-11% of BF16 weights changing, and the top-64 intersection KL gradient closely matches the full-vocabulary gradient yet its mathematics accuracy is 2.6 points above sampled-token policy gradient under response averaging and 2.1 points below under global token averaging.
Using Qwen3-1.7B with four RL teachers from the same initialization (mathematics, code, instruction following, science), this study compares gradients, optimizer updates, and task learning curves, with SmolLM3-3B diagnostics: loss averaging implicitly weights longer responses more, Adam's first moment aligns updates across teachers (cosine 0.83 between teachers, 0.96 between averaging rules), BF16 rounding hides about 97% of FP32 master-weight changes and shows only 7-11% of BF16 weights changing, and the top-64 intersection KL gradient closely matches the full-vocabulary gradient yet its mathematics accuracy is 2.6 points above sampled-token policy gradient under response averaging and 2.1 points below under global token averaging.
Using Qwen3-1.7B with four RL teachers from the same initialization (mathematics, code, instruction following, science), this study compares gradients, optimizer updates, and task learning curves, with SmolLM3-3B diagnostics: loss averaging implicitly weights longer responses more, Adam's first moment aligns updates across teachers (cosine 0.83 between teachers, 0.96 between averaging rules), BF16 rounding hides about 97% of FP32 master-weight changes and shows only 7-11% of BF16 weights changing, and the top-64 intersection KL gradient closely matches the full-vocabulary gradient yet its mathematics accuracy is 2.6 points above sampled-token policy gradient under response averaging and 2.1 points below under global token averaging.