The Tasteful Agent: Measuring and Training Long-Horizon Judgment from Trajectory Hindsight
Synopsis
The work defines an agent's taste as its ability to choose the better direction before the outcome is visible, builds Taste-Bench, a 502-question benchmark mined automatically from decision forks in existing agent trajectories, finds that the best model answers only 59.7% correctly while forks whose deciding evidence appears later are much harder and a larger reasoning budget does not help, and shows that distilling the reasoning of a teacher that has seen the outcome into a student improves judgment on unseen tasks by 17.9 percentage points and raises end-to-end success on held-out SWE-bench Pro tasks from 14.6% to 33.7%.
Interpretation
It formalizes taste as a two-choice judgment at a decision fork and labels forks automatically from the outcomes the trajectories themselves record. Existing long-horizon benchmarks report only whether the agent finishes the task and give no measure of decision quality along the way; this work observes that the later part of a trajectory provides hindsight evidence for its earlier decisions and exploits the structure in which parallel attempts at the same task share an equivalent prefix and then diverge, isolating comparisons in which everything except the judgment is the same. The method gives a formal definition and two constructions (parallel trajectories and detour trajectories), with labels taken from recorded outcomes such as passing tests or achieving the experimental objective; a human review of 100 sampled questions found 170 of 172 explicit A/B judgments agreeing with the mined label, a 98.8% agreement rate.
It releases Taste-Bench, 502 taste questions covering software engineering and machine-learning research, together with the accuracy distribution of current frontier models. No prior benchmark specifically measures agent taste; the benchmark keeps 10.8% of 4,657 proposed forks after generator and four-judge filtering, with 390 engineering and 112 research questions, and each question is evaluated twice in both candidate orders so that only questions answered correctly in both orders count as correct. Fourteen contemporary models are evaluated under one interface and token budget; the best model, GPT-5.6 Sol, reaches an Average accuracy of 59.7% and GPT-5.5 reaches 59.5%, while random guessing scores 25% and a model that always prefers the same position scores 0%; per-cell means are 58.1% for parallel engineering, 50.9% for parallel research, 50.8% for detour research, and 35.9% for detour engineering.
The time horizon of a fork drives difficulty: the later the deciding evidence appears, the closer models come to random guessing, and a larger reasoning budget does not improve accuracy. The work annotates every fork with a four-level time horizon (in prefix, inferable, next step, more work), turning long-horizon from an intuition into a measurable quantity, and tests whether the reasoning budget explains the accuracy drop. The mean over 14 models falls from 62.3% at the in-prefix level to 21.0% at the more-work level, near the 25% random score; rerunning two models under three reasoning-effort settings produces six conditions and 6,024 responses, and the change from the lowest to the highest setting is not significant within the reported intervals, while the models produce the most reasoning tokens at the more-work level, which is also the least accurate.
Taste is trainable: distilling the reasoning of a teacher that has seen the outcome into a student transfers judgment to unseen tasks and improves end-to-end success. Beyond measuring taste, the work uses the supervision already contained in each question, two candidate directions plus the outcome that determines the supported one, and distills a privileged teacher into a student that shares the same frozen base model but a different context, using a token-level forward KL to move in-context judgment into the weights. On two task-disjoint folds of the 390 engineering questions, the student reaches 47.9% accuracy on the held-out fold versus 30.0% for the base model, and mean accuracy over the two orders rises from 42.7% to 62.4%, a gain of 17.9 percentage points; on 41 held-out SWE-bench Pro tasks, correct advice raises success from 14.6% to 39.0% as an upper bound, and student advice reaches 33.7%, a gain of 19.1 percentage points.
Perspective
The result is meant for agent systems that already produce many trajectories: whenever the same task has multiple attempts or a single run contains a self-correction, forks can be mined and labeled automatically, so the benchmark scales with trajectory production. It applies to the two settings with existing trajectory pools, software engineering and machine-learning research, and the distillation experiments are validated there; domains without comparable outcome signals, or where outcomes cannot be clearly attributed to a judgment, are not directly covered by this construction.
A careful reader would still watch several things: label quality depends on the generator and four judge models, and pairwise judge agreement on the released parallel-engineering questions is 51.7%, indicating that the judgments themselves are contested; the correlation with SWE-bench Verified is only about 0.2 on the engineering subset with a wide interval, so the ranking relationship between the two benchmarks is not yet stable; the distillation experiment is validated on one base model and engineering questions, with transfer to research questions unreported; and the end-to-end experiment covers 41 tasks and 98 forks, where a gap remains between the correct-advice upper bound and student advice. In addition, this reading is full text, but figures and some appendix tables appear as text, and a few values such as some confidence intervals and correlation coefficients are not fully given in the text, so exact reproduction should consult the original figures and tables.
