Skip to main content
Back to timeline
arXivSource publication:

CRAFT fills spreadsheet forms with template-aware reflection and local repair, raising pair accuracy by 8.51 and 23.38 points on FormFillBench

Related research and updates

Synopsis

The authors propose CRAFT, a template-aware agent framework for spreadsheet form filling that couples reflective validation with constrained local repair: it localizes suspicious regions, restores overwritten template content when needed, re-grounds plausible writable slots with a Rectangle-Aware Slot Grounder (RASG), and constrains later edits via label-slot hints and protected regions; they also release FormFillBench with 327 forms across Instruction-Only and Multi-File tracks, where CRAFT improves pair accuracy over the strongest baselines by 8.51 and 23.38 percentage points respectively, with component-removal experiments supporting structural adjudication and slot re-grounding and a second backbone retaining the relative advantage.

Source-provided article image: CRAFT: An Agentic Spreadsheet Form Filling System with Template Awareness
Figure 1 ·

Figure 1: Example of spreadsheet form filling and its key challenges. The agent needs to gather evidence from auxiliary files, satisfy field-level constraints, preserve cross-field consistency, and interpret diverse form layouts.

arXiv

Interpretation

CRAFT turns reflection from a free-form request to regenerate the workbook into constrained repair grounded to specific regions: suspicious ranges carry error types, validation provenance, restoration indicators, and reliable cells, and template-integrity violations trigger restoration before refill. Existing spreadsheet agents already include reflection and validation, such as SheetAgent's iterative reflection and SheetBrain's validation of execution results; the increment here is the template-specific connection between diagnosis and repair, where suspicious ranges carry restoration decisions and protection constraints into slot re-grounding and local refill. Component-removal experiments show that removing Structure-Aware Local Adjudication (SALA) reduces pair accuracy by 6.71 and 41.04 points on the two tracks, and removing Progressive Slot Re-grounding (PSR) reduces it by 9.23 and 41.66 points, with the latter also lowering Multi-File E2E accuracy from 43.31 to 3.15.

The Rectangle-Aware Slot Grounder (RASG) uses a SegFormer-B1 backbone with rectangle logits plus horizontal and vertical boundary cues to reconstruct writable-slot masks, aligned to candidate cells by area-based thresholds, supplying candidate locations for local refill. General-purpose vision-language agents have limited reliability in precise slot localization; RASG explicitly models the rectangular geometry of form slots to narrow the repair space to plausible writable slots, and is trained in three stages from synthetic spreadsheets to CommonForms PDF layouts to a manually filtered, annotated FUSE subset. The paper reports qualitative visualizations of RASG's directional cues and reconstructed slot predictions across diverse layouts including adjacent fields, elongated regions, and irregularly arranged sections, and notes a fallback to guarded local repair when no reliable slot is found.

The authors build FormFillBench with 327 form instances: 200 forms and 5,499 label-value pairs in the Instruction-Only Track, and 127 forms and 2,288 pairs in the Multi-File Track, with roughly five supporting files per case and 11 file types overall. Existing form-filling benchmarks mainly target web-based or document-centered settings, and existing spreadsheet evaluation focuses on spreadsheet understanding and manipulation rather than spreadsheet form filling; this benchmark separates the two tracks to distinguish grounding supplied values into the right locations from consolidating evidence across heterogeneous files before grounding. Templates come from FUSE, VEUSES, Venron2, public online sources, and synthetic templates, with candidate samples undergoing automated consistency checks and manual verification; metrics are pair accuracy and end-to-end accuracy, plus fill precision, recall, F1, and fill coverage.

With the default Gemini 3.1 Flash-Lite-Preview backbone, CRAFT reaches 59.68 pair accuracy and 15.5 E2E accuracy on Instruction-Only, and 66.52 and 43.31 on Multi-File, leading all 11 baselines. Relative to the strongest Instruction-Only baseline VLM+HTML+Image, the gains are 8.51 and 2.00 points; on Multi-File the best baseline differs by metric (OpenAI Agents SDK at 43.14 pair accuracy, CrewAI at 18.90 E2E accuracy), giving gains of 23.38 and 24.41 points. With GPT-5.4 mini as the backbone, CRAFT still leads among the evaluated methods, reaching 41.57 pair accuracy on Instruction-Only and 60.66 on Multi-File, exceeding the strongest corresponding baselines by 7.51 and 10.75 points; absolute performance remains model-dependent, with Multi-File E2E accuracy at 9.45 under that backbone.

Perspective

The work targets spreadsheet form filling: given a template, an instruction, and auxiliary files, it produces entries specifying which value goes into which range while preserving template constraints such as static labels, formatting, and merged regions. It applies to enterprise form settings that require consolidating evidence across heterogeneous sources such as emails, PDFs, and spreadsheets and grounding values precisely, for example reimbursement sheets and compliance questionnaires; the dual-view representation and region-level repair specification also offer a reusable interface for other templated document filling. The evaluated configuration permits at most one reflection round and six patch steps, so the returned workbook may retain unresolved errors.

The paper states that CRAFT is limited by available source parsers and may require human clarification for ambiguous labels or multiple specifications within a cell; the benchmark is curated, and the reported point estimates cover two backbones and an LLM-based semantic judge, establishing performance under this evaluation protocol rather than reliability across arbitrary enterprise workflows. Strict form-level completion remains challenging, especially on Instruction-Only forms and with the second backbone. The two tracks contain different form sets, so their absolute scores are not a controlled comparison of task difficulty, and run-to-run variability is not yet quantified, which is a direction for future work.

Sources