AdviSD trains a Qwen3-8B advisor to learn selectively from feedback corrections, beating advisor-GRPO by 4.2–6.4 points on BFCL-v3
Synopsis
The work introduces AdviSD, in which a small trainable advisor (Qwen3-8B) steers frozen frontier executors (Gemini 3.7 Flash, Claude Sonnet 4.6) with natural-language advice by pairing outcome-based GRPO with selective feedback-conditioned self-distillation: reflection proposes corrections, and the advisor scores the same recorded executor response with and without its issued advice, using the magnitude of that difference to choose which decisions to supervise, reaching 4.2–6.4 percentage points above advisor-GRPO on BFCL-v3 and 3.9–5.1 score points above it on EnvScaler while beating matched-count random selection and a no-gate variant.
Interpretation
The paper gives a theoretical account of which corrections to retain: in a shared-parameter model, if corrections from persistent failures (execution-insensitive situations) target a weaker preference for useful advice than other corrections, their growing share of supervision as preventable failures become rarer limits learning, and retaining them less often than the rest raises the performance the advisor eventually reaches (Theorems 1–2). Prior feedback-conditioned self-distillation work (such as SDPO and DistIL) focused on constructing teachers or allocating supervision; here the problem is located at the point where the advisor's advice acts through another model, so fitting a teacher and improving execution are distinct objectives (Lemma 1). The evidence is formal: Lemma 1 gives a gradient identity at a fixed prefix with fixed completion and execution laws, and Theorems 1–2 prove in a simplified two-situation, single shared log-odds model that the retained mixture sets the learning limit, with explicit equilibrium expressions. The authors state that this model does not describe the full dynamics of the reflector, the evolving language model, or the optimizer.
AdviSD separates proposing corrections from choosing where to learn: the advisor scores the same recorded executor response under two contexts, one containing its issued advice and one without it, and uses the magnitude of the score difference as the selection signal, so selection needs neither executor likelihoods nor additional executor rollouts; retained decisions are supervised by a feedback-conditioned copy of the pre-update advisor while the rollout batch's GRPO advantages stay unchanged. Compared with trajectory-level or turn-level supervision, selection happens at the advisor–executor interface, and the selection signal comes from paired scoring of an already recorded executor response rather than from a direct measurement of executor behavior or task return. Method details are complete: scoring contexts are built from the request sent to the executor, the target is the serialized and tokenized response including tool-call order but excluding subsequent tool results, and the threshold is calibrated from an empirical quantile of donor-advice contrasts (advice from other tasks, scored but never executed) and frozen during training; teacher and student distributions are renormalized over the pre-update student's top 100 tokens.
With Qwen3-8B advisors for the frozen executors Gemini 3.7 Flash and Claude Sonnet 4.6, AdviSD has the highest in-domain aggregates on BFCL-v3 and EnvScaler, exceeding advisor-GRPO by 4.2–6.4 percentage points and 3.9–5.1 score points; it beats matched-count random selection by 2.5–4.9 points and no gating by 3.2–5.1 points, while the two controls differ by at most 0.9 points and the inverted gate trails GRPO on both benchmarks. These ablations separate choosing which decisions to supervise from merely reducing the amount of supervision: randomly reducing supervision does not recover AdviSD's gains, and prioritizing low-contrast decisions does worse than reward learning alone. Each trained method has three independent training runs, each taking a validation-selected checkpoint and averaging four test evaluations; the BFCL-v3 test set is fixed at 320 tasks (80 per category) and EnvScaler at 200 tasks. The authors note that reported numbers are test results, not checkpoint-selection validation scores.
Without retraining, AdviSD advisors transfer to out-of-domain benchmarks and across executor versions and model families: they lead the four-benchmark macro-average over ACEBench, ToolHop, τ²-bench, and RoTBench, exceeding standalone execution by 2.7 points with Gemini and 3.6 with Claude while GRPO's margin is under one point, and they exceed transferred GRPO by 3.1 percentage points in both cross-family directions. The authors also report that transferred scores remain below those of AdviSD advisors trained for the evaluation executor, so a transferred model does not replace executor-specific training, and on Claude's ACEBench multi-step tasks AdviSD still trails standalone execution and GRPO by 1.7 percentage points. Out-of-domain evaluation keeps each benchmark's native tools, policies, and metrics and reuses the BFCL-trained, validation-selected checkpoints; the transfer experiment fixes the 320 BFCL-v3 test tasks and changes only the executor.
Perspective
The result is aimed at settings where the executor is frozen and can only be influenced through natural-language advice: the advisor decides before each executor response whether to advise or abstain, and advice goes into a temporary request rather than the executor's persistent history. It suits teams that need to adapt frontier API models externally and are willing to train a small advisor; the training side needs a reflector (both executor settings use Gemini 3.7 Flash for reflection) and a threshold calibration based on donor advice, while reflection and scoring are not used at deployment. The theoretical conclusions apply to a simplified model with fixed teachers and stationary retention, and the random-teacher extension additionally requires a stationary conditional law.
The learning analysis assumes fixed teachers and stationary retention, and the authors state that these results do not establish convergence for AdviSD's changing-teacher GRPO–AdamW training; evaluation covers two executor families, one shared reflector, and three runs per trained method, and updates to API models can constrain exact reproducibility. The selection signal is itself the magnitude of a change in the advisor's prediction, which the authors say is neither a measured behavioral effect nor an improvement in task return, and is not identified with the theoretical sensitivity classes; donor calibration also does not remove the asymmetry that the response was observed with advice. In addition, the result tables in the loaded text retain only row labels without numeric values, so specific scores can only be taken from the abstract and prose rather than checked item by item. Training-time statistics show advice becoming less frequent, ordinary gate retention falling from 58.0% early to 32.2% late, and the abstention bypass accounting for about one-fifth of selected decisions in middle and late training, which the authors describe as allocation statistics rather than gradient magnitudes or causal sensitivity measurements.
