Four studies find semantic geometry did not establish pre-action AI judgment: a lexical router recovered every governing and blocking policy, yet the composed pipeline escalated all 2,304 actions
Synopsis
The work represented proposed actions and governing policies as vectors and tested across four studies whether cosine similarity, raw Euclidean distance, and signed raw dot product could support pre-action judgment of an action's relation to policy; in the final synthetic study a lexical router recovered every governing and blocking policy while cutting median policy checks by 97.7%, yet the composed pipeline escalated all 2,304 test actions and supplying every policy to the same downstream mechanism changed no decision, leading the author to propose that judgment requires developing a consequence graph.
Figure 1. From a proposed action to a governed decision.
arXivInterpretation
In the final synthetic study, the selected lexical-overlap router recovered 2,880/2,880 governing-policy instances and 984/984 blocking-policy instances on the sealed surface with no blocking misses, returned a median of 1.5 policies per action, and reduced median policy evaluation by 0.976562500000. Earlier studies entangled retrieval with relation interpretation, leaving the failure unattributed; the fourth study froze separate tests for routing, relation classification, and end-to-end composition so retrieval success could be observed on its own. The result comes from an authored synthetic policy world in which frozen route terms were exposed to the pipeline and the lexical rule operated on authored language with fixed preprocessing; the author states it cannot be promoted into a general claim about natural-language policy retrieval.
In the relation experiment, the full three-measure model reached relation macro-F1 of 0.355163305089 with the first encoder and 0.485257556129 with the second, with Euclidean-only the best singleton in both encoders; adding cosine and signed dot product produced mean gains of 0.002145935796 and 0.005580263897 over the best singleton, below the prespecified 0.03 useful-effect floor, with family-cluster bootstrap intervals crossing zero and Holm-adjusted permutation values of 1.000000000000. The second study had preserved bounded proximity and direction signals and the third tried to compose them into decisions; the fourth tested incremental utility using paired family differences rather than subtracting corpus-level macro-F1 scores, directly asking whether the extra measurements earned their place. The comparison was paired within each of the twelve sealed mission families, which supplied the independent inferential units; only two fixed encoders, one frozen model class, one 0.70 threshold, and one synthetic world were tested.
The composed pipeline returned ESCALATE for all 2,304 sealed actions in each encoder, giving final-disposition accuracy of 840/2,304, or 0.364583333333; the exhaustive semantic comparator sent every policy to the same relation model and aggregation rule and escalated the same 2,304 actions with identical accuracy and escalation rate, a mean accuracy difference of zero and bootstrap intervals of [0.000000000000, 0.000000000000]. The third study showed the composed pipeline failed but could not locate which stage limited it; the fourth used the exhaustive comparator to rule out retrieval scarcity and localize the failure downstream of retrieval in the relation-to-disposition stage. Zero false ACT reflects zero ACT predictions rather than demonstrated safety; the author reports two no-retry replicas completed, clean-root reconstruction agreed, the output-version census reconciled, candidate bytes remained frozen, and independent verification passed.
The author revised the hypothesis to propose that judgment in AI agents requires developing a consequence graph connecting the actor and authority to policy conditions, exceptions, and the changes an action would produce, with follow-on studies planned to compare a consequence-graph representation against a direct language relation reader and a structured policy workflow under fixed retrieval. The original hypothesis held that geometric measurements could supply a basis for judgment; after four studies did not establish that capability, the author reframed the problem as making relationships explicit and states the graph has not yet been tested. This is a follow-on hypothesis and comparison design motivated by the negative results; none of the four studies tested the structure or established its necessity or sufficiency.
Perspective
The work addresses researchers and engineering teams who need agents to act under multiple policies, in settings where policies and actions arrive as natural language and where ACT, HOLD, and ESCALATE must be distinguished. Its reusable output is a stage-separated evaluation structure in which routing, relation interpretation, and final disposition each answer for themselves, and relation representations are compared only after retrieval is fixed. The lexical router's high recall and 97.7% reduction in checks belong to the author's synthetic policy world, where route terms were exposed and preprocessing was fixed. The planned follow-on study would hold retrieval fixed and compare a consequence-graph representation against a direct language relation reader and a structured policy workflow, with relation ground truth, comparison rules, and meaningful error thresholds defined before scoring.
A careful reader would still watch several open questions. The authored policy world's route terms and controlled language may have made localization easier while leaving relation interpretation artificial in ways that do not transfer. Actions and policies entered the relation stage through whole-text vector representations, and the studies did not determine whether those representations discarded the distinctions that mattered or whether a stronger reader could have recovered them. The fourth study tested one deterministic multinomial logistic relation reader, two fixed encoders, a 0.70 threshold, and one fixed aggregation rule; changing any one of those surfaces creates a new scientific question rather than repairing or reinterpreting the frozen result. Final authority remains open, since a machine might make the disposition directly, present structured evidence to a deterministic mechanism, abstain, or preserve a human decision. The author notes that the scientific companion omits separate row predictions for some first-study controls, that fourth-study primitive relation recalculation requires exact gold and ordering joins, that full routing-budget and serial comparisons require original feature inputs not supplied, that model assets are not bundled, and that full experimental replication is therefore not claimed.
