Skip to main content
Back to timeline
arXivSource publication:

GraphForge synthesizes 2,169 trajectories from evidence graphs over real files, lifting Qwen3.6-27B from 1380.0 to 1445.7 GDPVal Elo

Synopsis

GraphForge is an evidence-graph framework that starts from O*NET occupation seeds, crawls real files into a workspace, builds an evidence graph over their relations, and compiles that graph into task statements and rubrics, with an initial rollout and a revision agent checking executability; fine-tuning Qwen3.6-27B on 2,169 trajectories brings GDPVal to 1445.7 (+65.7) under OpenHands and Workspace-Bench-Lite and SpreadsheetBench II to 63.7 (+7.7) and 24.0 (+13.7) under Claude Code, with the same data also improving Qwen3.6-35B-A3B.

AI-generated editorial illustration: GraphForge: Training Working Agents with Graph-Anchored Workspace Synthesis

Interpretation

The framework grounds both the task and its verification in real files: seeds come from O*NET occupations and their Detailed Work Activities, keeping only tasks annotated DIGITAL, which leaves 246 occupations across 16 sectors and 43 sub-sectors, 891 DWA task types, and 3,419 valid occupation-task-type pairs; a search agent instantiates each seed into a real-file workspace, over which an evidence graph records files, the facts they provide, and cross-file dependencies. Where EnvCraft generates files with a model and checks workspace state with a Python script, and NexForge builds tasks on real files without task-specific rubrics or verifiers, GraphForge compiles both the task statement and the rubrics from the evidence graph, so each positive criterion carries evidence anchors and a verification procedure. The paper describes five method stages, reports the post-filter seed counts, and gives a construction funnel: of 3,638 materialized tasks, 2,967 (81.6%) reuse unchanged rollouts, 2,153 (59.2%) are admitted, yielding 2,169 validated trajectories over 2,150 unique workspaces.

Training gains appear on all three benchmarks and across scaffolds: SFT improves the 35B base model by 101.7 Elo on GDPVal under OpenHands and 101.4 under Codex, while the 27B SFT model reaches 1445.7 Elo (OpenHands) and 1427.4 Elo (Codex); Workspace-Bench-Lite improves by up to 6.6 and 7.7 points and SpreadsheetBench II by up to 16.5 and 13.7 points. All GraphForge trajectories are rolled out with the Codex scaffold, yet evaluation covers OpenHands, Codex, and Claude Code, and the SFT model beats its base model under every scaffold, suggesting the corpus teaches working skills that transfer across scaffolds rather than habits tied to the rollout scaffold. Evaluation uses pass@1 with a fixed workspace interface and compares models within the same scaffold; GDPVal uses a Bradley-Terry fit over an internal Elo pool anchored by fixing GLM-5.3 (OpenHands) at 1667, with 95% bootstrap intervals reported in the appendix.

Rejection fine-tuning with evidence-anchored rubrics adds a selection signal: from 2,000 newly synthesized queries, up to four SFT-model trajectories are sampled per query, and the anchored arm keeps the rubric-best trajectory when its score exceeds 0.95, leaving 462 after validity filtering; on Workspace-Bench-Lite and SpreadsheetBench II the anchored arm gains most, unanchored selection gains less or turns negative, and random selection is weakest overall. The three RFT arms share the same 462 query IDs, candidate pools, and optimization budget and differ only in the selection rule, so this matched design isolates within-query trajectory selection rather than the full task-admission pipeline. The paper states that GDPVal differences among arms are not statistically resolved at this scale and that Elo differences at 220 tasks are noisy; Appendix D shows all difference intervals include zero, so the authors limit the conclusion to selection quality mattering across benchmarks and evidence anchoring showing its advantage on Workspace-Bench-Lite and SpreadsheetBench II.

The authors audit contamination and judge sensitivity: none of the 39,201 training file hashes coincides with any of the 260 GDPVal file hashes, the top-20 most similar 13-gram text pairs share no substantive content, and only 13 of 44 GDPVal occupations are covered by the training taxonomy; splitting by coverage, the win rate on the 155 uncovered tasks is 0.739 (95% CI [0.671, 0.803]), no lower than 0.692 (95% CI [0.585, 0.800]) on the 65 covered tasks. These audits support reading the improvement as transferable working skills rather than memorization of benchmark content, and a controlled perturbation test further probes whether the judge actually reads the referenced files. Deleting the worksheet a criterion cites drops that criterion's score by 0.377 on average while non-target criteria stay essentially unchanged, but fine-grained corruptions of rows, numbers, and citations move scores by less than 0.03, which the authors attribute to the capability limit of GLM-5.2 as an agentic judge.

Perspective

The method targets training working agents that read diverse files, coordinate tools, and produce deliverables, and it applies where public real files are available and task outcomes can be verified against file contents; because seeds come from O*NET occupations and their official work activities, coverage is bounded by that taxonomy and by which tasks are digitally executable. For users, it offers reusable synthesized data, rubrics, trajectories, and trained checkpoints, plus a pipeline that anchors verification to source files; the paper also states plans to scale GraphForge to more task families and file types and to study how evidence-anchored verification interacts with longer-horizon agent scaffolds.

The corpus holds 2,169 trajectories, and how the benefits scale with larger data budgets is not studied; both the synthesis pipeline and the agent judge are powered by GLM-5.2, so whether stronger frontier models improve data quality and judging reliability remains open; experiments cover two base models from the same family, leaving cross-family transfer unknown. The judge sensitivity experiment shows it reliably detects structural evidence failures but struggles to verify fine-grained content in rows, numbers, and citations, a capability boundary worth watching when rubric scores are used as a selection signal. In addition, the anchored and unanchored judge variants cannot be separated on GDPVal, where Elo differences at 220 tasks are noisy, so those conclusions should be read as directional rather than settled.

Sources