GraphForge: 2,169 Trajectories Push a 27B Open Model Up 65 Points on GDPVal
Office-work agents have a training-data problem. To learn to read a pile of files, juggle tools and hand over a deliverable, they need tasks built on many real files with answers that can be checked. Model-generated files look fake and lack variety. Real files without task-specific checks leave quality unverified. GraphForge (HF 33 upvotes) fixes both sides with one structure: an evidence graph.
The pipeline starts from seeds tied to real occupations to keep the tasks diverse in a controlled way. For each seed it assembles a workspace of real files and builds a graph of how those files relate. Both the task statement and the grading rubric are derived from that graph, so every requirement is backed by something in the workspace, and every rubric item points to the exact files needed to verify it. A first rollout tests whether the task is actually doable, and a revision agent repairs the task and rubric against the original files before any training trajectories are collected.
The payoff is large for a small dataset. Fine-tuning Qwen3.6-27B on just 2,169 GraphForge trajectories raises GDPVal to 1445.7 (+65.7) under OpenHands. Under Claude Code, it lifts Workspace-Bench-Lite to 63.7 (+7.7) and SpreadsheetBench II to 24.0 (+13.7). Rejection fine-tuning, where the evidence-anchored rubrics pick the best of the model's own rollouts, adds more on all three.
The quiet point here is the rubric. A rubric tied to specific files is not just a grader, it's a selection signal you can train on. That is exactly the verifier piece most "AI coworker" products are missing. Rubrics anchored in evidence make the work checkable, and checkable work is trainable work.
Link: arxiv.org/abs/2609.38923
← Back to all articles
The pipeline starts from seeds tied to real occupations to keep the tasks diverse in a controlled way. For each seed it assembles a workspace of real files and builds a graph of how those files relate. Both the task statement and the grading rubric are derived from that graph, so every requirement is backed by something in the workspace, and every rubric item points to the exact files needed to verify it. A first rollout tests whether the task is actually doable, and a revision agent repairs the task and rubric against the original files before any training trajectories are collected.
The payoff is large for a small dataset. Fine-tuning Qwen3.6-27B on just 2,169 GraphForge trajectories raises GDPVal to 1445.7 (+65.7) under OpenHands. Under Claude Code, it lifts Workspace-Bench-Lite to 63.7 (+7.7) and SpreadsheetBench II to 24.0 (+13.7). Rejection fine-tuning, where the evidence-anchored rubrics pick the best of the model's own rollouts, adds more on all three.
The quiet point here is the rubric. A rubric tied to specific files is not just a grader, it's a selection signal you can train on. That is exactly the verifier piece most "AI coworker" products are missing. Rubrics anchored in evidence make the work checkable, and checkable work is trainable work.
Link: arxiv.org/abs/2609.38923
Comments