Skip to main content
Back to timeline
arXivSource publication:

Law&Order autoformalizes tax forms and filing instructions into executable programs via a neuro-symbolic framework, reaching 100% cell-level and form-level accuracy on 51 held-out returns

Related research and updates

Synopsis

The work proposes Law&Order, a neuro-symbolic framework that combines large language model synthesis with cell-level verification and iterative localized error repair using human-written OpenTaxSolver tax returns to automatically formalize tax forms and instructions into executable symbolic programs, establishing structural correspondence aligning legal and symbolic components such as cells and schedules and denotational correspondence requiring symbolic components to implement the computations specified by their legal counterparts; evaluated on independently authored, held-out TaxCalcBench returns never exposed during generation or repair, the most advanced LLM achieves only 66% accuracy while Law&Order achieves 100% cell-level and form-level accuracy on 51 held-out returns.

Source-provided article image: Law And Order: Tax Law Autoformalization
Figure 1 ·

Figure 1: The Law&Order generation–verification–repair loop. Starting from an initial set of cell functions F and a repair-set of test cases T, the system repeatedly executes F on T and compares each cell’s model output o against its gold annotation G. Any cell function f whose output disagrees with G is regenerated by an LLM using that cell’s natural-language instruction, and the cycle repeats until every cell matches across all repair-set test cases. The resulting corrected F is then evaluated once on a held-out test set — never seen during synthesis or repair — to produce the final reported accuracy.

arXiv

Interpretation

Introduces Law&Order, a neuro-symbolic framework that automatically formalizes tax forms and instructions into executable symbolic programs handling large computational structures involving arithmetic, branching, recursion, and tabular reasoning. Scalable methods for translating legal texts into accurate symbolic representations remain underdeveloped; this work provides an end-to-end formalization pipeline in the concrete setting of tax law. The abstract describes the framework components and goal without component-level ablations or internal implementation details.

Establishes two forms of correspondence between law and logic: structural correspondence aligning legal and symbolic components such as cells and schedules, and denotational correspondence requiring symbolic components to implement the computations specified by their legal counterparts. Decomposes formalization correctness into two checkable dimensions, structural alignment and semantic alignment, providing explicit criteria for verification. The abstract provides definitional descriptions without reporting independent validation results for each correspondence.

Combines large language model synthesis with cell-level verification and iterative localized error repair, using human-written OpenTaxSolver tax returns as the basis for repair. Closes the loop between LLM generation and symbolic verification rather than relying on a single LLM output. The abstract describes the method combination without reporting repair iteration counts or convergence behavior.

On independently authored, held-out TaxCalcBench returns never exposed during generation or repair, Law&Order achieves 100% cell-level and form-level accuracy on 51 returns, compared with only 66% for the most advanced LLM. Uses a held-out set to contrast LLM synthesis plus symbolic verification against using an LLM alone. The abstract reports the held-out set size (51 returns) and two accuracy figures without confidence intervals or per-return distributions.

Perspective

The result targets tax forms and filing instructions, a rule-dense setting with well-defined computational structure, and is intended for readers who need to turn legal text into executable symbolic programs with cell-level and form-level verification, such as researchers in legal informatics, compliance computation, and formal methods. The evaluation described in the abstract uses independently authored, held-out TaxCalcBench returns never exposed during generation or repair, so its scope is bounded by that benchmark and the tax law domain.

The abstract does not specify the implementation of cell-level verification and iterative localized error repair, the number of repair iterations, or failure modes, nor does it give per-return distributions or statistical uncertainty beyond the 100% and 66% figures; moreover, the held-out set comprises 51 TaxCalcBench returns, so transferability to broader tax provisions and other legal domains remains to be examined in follow-up work.

Sources