Skip to main content
Back to timeline
arXivSource publication:

TGMS uses plan and evidence contracts to make database-agent plans rejectable and claims checkable

Synopsis

The work proposes two enforceable contracts for LLM agents over databases: a plan contract that statically checks plan structure, result-field references, and grounded identifiers before execution, and an evidence contract that records belief state, completeness, and provenance so claims can be verified; in TGMS, result-field checking turns invalid references into repairable rejections, completeness propagation detects truncated-count failures in all 15 controlled cases, and claim gating leaves none of 199 emitted answers with an unsupported gated claim at the cost of 21 fewer answers.

Source-provided article image: No Trace, No Claim: Two Contracts for Database Agents
Figure 1 ·

Figure 1: The plan contract checks admissibility before execution. Execution produces annotated results, and the evidence contract checks the resulting claims. The contracts do not guarantee intent interpretation or operator choice.

arXiv · Page 3

Interpretation

It defines and formalizes two contracts: a plan contract specifying what an agent may execute and reference, and an evidence contract recording the belief state, completeness, and provenance under which a result supports a claim, yielding three properties: plan admissibility, execution reproducibility, and claim faithfulness. Prior work largely asks how to obtain useful plans and execute agent workloads efficiently; input schemas constrain only values entering an operation and value grounding only checks numeric and entity consistency. This work places controls at two boundaries, before and after execution, and states that the guarantee envelope excludes task correctness. The contracts are instantiated in TGMS: thirteen temporal operators are each validated against an independent brute-force oracle on 500 randomized cases, and deterministic execution is evidenced by canonical serialization and SHA-256 digest reproduction.

Result-field checking converts the failure of models referencing nonexistent result fields from a runtime error into a structured, repairable pre-execution rejection. In the first live probe run every generated plan passed argument validation yet none executed, because plans referenced fields the producing operators did not return, such as s2.count instead of s2.rows_total; after adding result schemas and inter-step reference checks, execution success on a three-task probe set rose from 0/3 to 3/3, and in the 7B ablation from 0.23 to 0.55. A three-task probe set, 7B and 14B ablations, and a comparison showing grammar-constrained decoding did not address the mismatch; the authors characterize the failure as semantic rather than purely syntactic.

Completeness propagation prevents a count over an incomplete page from supporting a claim about the complete result, and claim gating removes unsupported claims before emission. Value grounding cannot detect this error because the reported entities and derived count are consistent with the visible page; TGMS records rows_total and truncation flags, propagates incompleteness through dependent steps, and incorporates it into claim verification. 15/15 detections in the controlled generator versus 0/15 when disabled, with only three verdicts changing on the natural development split; in the frozen CollegeMsg campaign, 21 of 220 emitted answers contained an unsupported gated claim before gating, and after gating all 199 emitted answers passed, at the cost of 21 fewer answers and a one-percentage-point drop in overall accuracy.

On frozen workloads a fixed operator interface is not a universal accuracy advantage: TGMS reaches 0.408 typed-answer accuracy on CollegeMsg versus 0.064–0.284 for the evaluated baselines, but matches direct SQL over the same bi-temporal store on Bitcoin-OTC. The same-information SQL baseline uses the same bi-temporal DuckDB tables, model, and repair budget, changing only query language, plan representation, and static checks, so it compares two interface designs rather than isolating one mechanism; on correction probes TGMS and bi-temporal SQL answer historical-belief questions while latest-state baselines cannot. Four frozen workloads with test splits of 94, 94, 102, and 94 tasks, primary configuration Qwen2.5-14B-Instruct-AWQ at temperature 0, CollegeMsg aggregating three seeds; the authors note that gains on email-EU and the synthetic workload are positive but inconclusive because their confidence intervals include zero.

Perspective

The design targets settings where past decisions must be reproducible after records are corrected, such as audit and compliance queries, and workloads where plans compose several operators and partial results can change the meaning of an answer. For users, the plan contract makes invalid plans rejectable before execution with structured repair information, and the evidence contract makes gated claims faithful to their cited evidence, at low overhead: plan admission costs 4.1 ms at p50 and claim verification 3.3 ms, together well below 0.01% of end-to-end task time. Prompt size stays independent of raw graph size, remaining between 6,000 and 10,000 tokens per task across the tested scales. The results apply to questions expressible by the current algebra and to the claim types currently gated.

The guarantee envelope remains limited to the claim types currently gated; temporal-pattern claims are checked but not yet withheld, and the verifier is implemented and mutation-tested rather than formally verified. Membership checking cannot establish set completeness: the verifier confirms each reported entity occurs in the evidence but detects none of 100 cases in which valid members are omitted, so complete-set claims require a stronger proof obligation. Because the bi-temporal SQL baseline does not implement the evidence contract, the evaluation does not establish that claim checking requires a fixed operator algebra. The independent question study is small: only 10 of 110 questions were expressible with the current algebra, and only five expressible questions had answer types supported by the current scorer, producing nine CollegeMsg and six Bitcoin-OTC runs across three seeds. Approximate operators and agent write-back are outside the current prototype.

Sources