Skip to main content
Back to timeline
arXivSource publication:

Handing governance review gates to agents: top model reaches 94.98% strict gate success on DGF-Bench but only 76.92% complete-route success

Synopsis

The work develops a task-substitution framework for Digital Governance Frameworks (DGF) that treats each governance gate as an executable contract, requiring sufficient accessible information, valid decision and authority checks, and a reduction in total human work after exceptions, verification, correction, and maintenance are counted; it derives a residual-work threshold showing why automating most cases can still increase labor, and measures models on DGF-Bench (300 synthetic projects, 899 evaluable model-project runs): Gemini 3.8 Flash, GPT-5.6 Luna, and DeepSeek v4.1 Flash achieve strict gate success of 94.98%, 83.29%, and 74.18%, with complete-route success of 76.92%, 42.33%, and 24.

Source-provided article image: The Last Human Gate: Forward Deployed Engineering for Governance Automation
Figure 5 ·

Figure 5: The generated architecture document in Project Falcon’s actual dossier. The diagram is review evidence, not an independently validated production design. Identity, application, data, monitoring, backup, and recovery elements contextualize the readiness example.

arXiv

Interpretation

The paper models each enterprise governance review gate as an executable contract and proposes a task-substitution framework: human execution should be replaced only when accessible information is sufficient, decision and authority checks are valid, and total human work falls after exceptions, verification, correction, and maintenance are counted. Relative to treating governance automation as simply letting a model decide, the framework states substitution conditions as testable contracts and derives a residual-work threshold, showing that automating most cases can still increase labor. The framework is presented as formal conditions and a threshold derivation and is tested in the controlled DGF-Bench setting; the text is summary-level and does not give the derivation details or threshold values.

DGF-Bench supplies 300 synthetic projects and 899 evaluable model-project runs, measuring strict gate success of 94.98%, 83.29%, and 74.18% for Gemini 3.8 Flash, GPT-5.6 Luna, and DeepSeek v4.1 Flash, with complete-route success of 76.92%, 42.33%, and 24.67%. The benchmark separates getting a single gate right from completing the whole route, exposing a clear gap between the two that earlier discussion of single-point judgment did not capture. Evidence comes from counted evaluable runs on controlled synthetic projects, supported by evidence audits and 135 repeated runs that distinguish correct decisions from reliable execution.

A deterministic control passes all 1,700 gates given the supplied rules and structured facts, locating the model comparison in the execution of a supplied decision kernel. The control attributes task difficulty to execution rather than to the decision rules themselves, providing a reference point for interpreting model success rates. The control passes all 1,700 gates under the same supplied rules and structured facts, a deterministic result under controlled conditions.

A document counterexample establishes an information-sufficiency obstruction, showing that substitution conditions cannot be met when accessible information is insufficient. The counterexample elevates information sufficiency from an engineering detail to a precondition for whether substitution can hold at all. It is given as a single document counterexample, an existence argument rather than a coverage statistic.

Perspective

The work targets the technical feasibility of having agents and software execute already-specified governance review tasks, for governance and engineering teams assessing gate substitution, provided gate rules and structured facts can be supplied. The framework's workforce test is defined by complete human effort at fixed output and quality, while the present measurements cover review performance only, so the results should be read as a test of substitution conditions in a controlled synthetic setting rather than an observation of labor change in real organizations.

The text is summary-level: it does not give the concrete form or value of the residual-work threshold, nor how the 300 synthetic projects were constructed, how gate types are distributed, or how scoring is defined, so the sensitivity of success rates to project difficulty cannot be judged. The deterministic control passing all 1,700 gates was obtained with supplied rules and structured facts, leaving open how models perform when information is incomplete or rules must be inferred. The document counterexample establishes only the existence of an information-sufficiency obstruction, with limited coverage. In addition, the complete human effort required by the workforce test has not been measured, so whether substitution truly reduces total labor remains an open question.

Sources