Skip to main content
Back to timeline
arXivSource publication:

EIO-Agents proposes an evidence-to-decision semantic layer and portable evaluation record, deriving a reproducible BLOCK from 243 claims

Synopsis

EIO-Agents introduces an open specification with two layers: the Evaluation Intelligence Ontology (EIO) defines typed evidence, versioned behavioral predicates with evidence contracts, claims, witness rules, and proof status, while the Portable Evaluation Record (PER) preserves one evaluation's evidence-to-decision chain as a canonical, content-addressed record; in a reference credit-underwriting evaluation, 243 claims project, validate, and explain to a BLOCK decision, whereas an identical capped score in another record yields REVIEW because its proxy observation did not reproduce.

Source-provided article image: EIO-Agents: The Missing Semantic Layer for AI Agent Evaluation
Figure 1 ·

Figure 1: The EIO-Agents architecture. Evaluation systems produce evidence; EIO supplies the semantic contract that turns evidence into claims, proof status, and decisions; PER preserves the resulting evaluation as one content-addressed system of record. Metrics, findings, controls, and release recommendations are views over claims rather than independent facts.

arXiv

Interpretation

The paper proposes EIO, a semantic layer whose core chain is Evidence, Predicate, Claim, Proof, Decision, with computable witness rules, proof rules, recurrence bands, and release semantics. Existing benchmarks, harnesses, LLM judges, and tracing systems each produce scores or labels but do not share a rule for what a given piece of evidence may establish; EIO encodes four separations as machine-checkable rules: Observation versus Evidence, Evidence versus Proof, Judgment versus Proof, and Score versus Decision. The ontology is published as 41 YAML modules pinned by exact version and SHA-256 digest, with 58 predicates, 11 evidence types, 10 source types, and 11 resolvers; the witness rule classifies 110 evidence-type and source-type pairs as 11 witnessing, 3 witnessing only with a paired tool call, 13 compatible but never witnessing, and 83 incompatible.

The paper proposes PER as the canonical record of one evaluation, with exactly 14 required top-level blocks sealed by a content digest, so the record can be re-derived, explained, and verified. A report summarizes an evaluation, whereas PER preserves the state from which that report can be regenerated; the record digest changes when the archive changes, when any byte of the EIO release changes, when the PER schema version changes, or when the conversion code changes its output. The reference library provides project, validate, verify, and explain operations; verification is grouped into schema and release, evidence, claims, views, explanations, decision, record hygiene, and derivation checks; a verifier self-test injects 31 defects into a valid record and requires each to be caught, and 37 build gates check the EIO release with 40 negative vectors.

The reference implementation makes the semantics executable and, across three reference evaluations, shows that the same score can lead to different decisions, that tampering is detected, and that jury consensus is not proof. The contribution is not a new evaluator but a common semantic substrate that lets heterogeneous evaluation output be projected into comparable, exchangeable, auditable records through declared mappings; the producer or adapter author declares what its output means, and EIO does not use fuzzy matching or guess the mapping. The FIN_3 record has 243 claims, 12 findings, readiness 49, and state BLOCK, with canonical bytes identical to the repository golden record; changing BLOCK to PASS by hand fails three independent checks: the schema check, release invariant I-8, and byte-identical re-projection; EXAM_B and FIN_3 share a capped score of 49, but EXAM_B's proxy observation reproduced in 0 of 5 re-runs and its state is REVIEW; across the three records, 618 claims and 58 findings include 47 jury findings, none Proven.

The paper frames the path from open specification to common standard, with a roadmap covering independent implementations, neutral governance, a neutral namespace, reference scoring, signed attestation, and adjudicated cases and reviewed mappings. The author uses 'standard' to mean an open specification designed for common adoption, not an already recognized standards-body standard; the semantic core (the chain, the witness rule, the proof rule, and the release semantics) is fixed and executable today. Normative changes follow a public process: a proposal naming affected identifiers and the effect on records and digests, at least 14 days of public comment, an implementation with regenerated digests and golden records, approval of the specification maintainers, and a new EIO release; framework mappings stay provisional until a legal review process exists.

Perspective

The specification targets teams that need to exchange evaluation results across tools, organizations, or regulatory boundaries, such as evaluation framework authors, auditors, compliance mapping maintainers, and downstream decision systems. It applies to any evaluator that can emit a conforming bundle, including benchmark runners, LLM-judge pipelines, and human review tools, and is designed to map to existing formats such as EARL, PROV, in-toto, and OpenTelemetry rather than replace them. What verification establishes is integrity and derivation: the record is well formed, every derived block agrees with its claims and evidence under the bundled EIO release, the digest is the digest of its canonical bytes, and projecting the given bundle with the same library version produces exactly this record. The specification also states that verification does not establish that the bundle came from a real run, that an agent or a jury would decide the same way again, that a Proven claim is true, or any certification or legal conformity. A control status expresses evidence relevance from one run, and every framework mapping is provisional and flagged for legal review.

Several open questions remain for a careful reader. The three reference archives were produced by an adversarial multi-turn harness whose re-run data come from a version of the harness that is not yet released, so the reproducible scope of the recurrence-band conclusions awaits that release. The fourth bundle comes from a fictional native producer and was written by hand, indicating that native records become fully schema-valid only after the next schema release moves every identifier to a neutral namespace. Framework mappings and all 165 controls are provisional pending a legal review process; a control status is not legal conformity, certification, or attestation. Semantic resolvers are never Proven under the current specification, so a failure resting only on model judgment can raise a release recommendation at most to REVIEW and requires human adjudication to stop a release; this trade-off means some harms detectable only semantically depend on a human process. Verification also does not establish that the bundle came from a real run or that a Proven claim is true, and records may contain personal data and should be handled accordingly.

Sources