RegLLM diagnostic harness exposes run-to-run variance in two nominally identical GPU pilots: escalation recall 1.0 versus 0.5, with the same adapter's effect flipping direction
Synopsis
The authors propose RegLLM, a diagnostic harness for bounded autonomy in regulated agentic workflows that instruments six trustworthiness signals (citation validity, source grounding, schema compliance, escalation correctness, constitutional alignment, unsafe-action rate) alongside a deterministic runtime supervisor; an offline reference run (n=12) lifts escalation recall from 0 to 0.67 and cuts the unsafe-action rate from 0.33 to 0.08, while two nominally identical single-GPU Qwen2.5-3B LoRA/DPO pilots (n=8, same seed and eval split) show task success of 0.25 versus 0.12 and escalation recall of 1.0 versus 0.
Interpretation
It turns the act-versus-defer decision into a verifiable training signal: each task carries a should_escalate label, the reward checks the agent's terminal decision against that label, and correct escalation contributes to the trajectory score just as a correctly cited provision does. Previously the escalation decision in regulated workflows lived in hand-coded thresholds, and verifiable-reward methods relied on machine-checkable answers as in math and code; this work classifies escalation correctness as a programmatically verifiable signal, making bounded-autonomy behaviour measurable, trainable, and auditable. The paper defines the hybrid reward and four programmatic sub-signals (citation validity, source grounding, schema compliance, escalation correctness), and states that the labels here are author-generated from templates while production requires expert annotation.
One domain constitution serves simultaneously as the soft term of the hybrid reward and as the runtime guardrail: 16 UK FCA principles drive both the AI-judge score and a deterministic supervisor that blocks ungrounded answers, forces escalation on out-of-scope triggers, and writes an append-only audit log keyed by the principle that fired. Prior constitutional-AI work mostly used principles to generate preference data feeding a downstream alignment run, and guardrail systems sat outside the agent; here a single artefact governs evaluation, training, and serving. The paper describes the Constitution interface, the GovernedAgent, and the audit log, and reports the metric changes the supervisor produces in the offline reference run.
The offline reference run shows the deterministic supervisor working on a weak baseline: escalation recall rises from 0 to 0.67, citation validity reaches 1.00, the unsafe-action rate falls from 0.33 to 0.08, and task success rises from 0.08 to 0.25. This isolates the runtime layer's contribution from learned policy behaviour, showing that hard rules directly move bounded-autonomy metrics on a policy that does not learn to defer on its own. Based on 12 held-out tasks in the offline harness, using a deterministic lexical baseline policy and an offline heuristic judge.
The principal empirical finding of the two GPU pilots is the variance itself: under the same code commit, eval split, seed, base model, and hyperparameters, RL-base task success is 0.25 versus 0.12, escalation recall 1.0 versus 0.5, and citation validity 0.50 versus 0.38; the answer-quality adapter moves recall from 1.0 to 0.5 in Run A but from 0.5 to 1.0 in Run B, and the escalation-aware variant designed for Run A's hypothesised failure mode produces no measurable change in Run B. The paper explicitly declines to read either direction as the adapters' true behaviour and instead treats the contrast between runs as evidence that at 8 tasks, 10-20 preferences, and 30 DPO steps, adapter effects cannot be separated from floating-point ordering, kernel selection, and dependency microversions. Two single-GPU Qwen2.5-3B LoRA/DPO pilots, n=8, same seed and eval split, differing only in GPU host; the authors note container image digests and deterministic CUDA flags were not pinned.
Perspective
The framework is for teams shipping an autonomous agent into a domain that has a written rulebook, some way to identify which queries are out of the agent's lane, and a small open base model they can fine-tune. It applies to workflows like financial compliance, where a wrong answer means a regulatory fine, a harmed client, or a missed obligation, and to domains such as data protection, clinical guidance, and taxation that can supply their own principle set. The orchestration layer (rented-GPU lifecycle with guaranteed teardown, S3-staged artefacts, verifiable-reward-driven training) is independently useful for any team running cost-bounded RL pilots. The paper is explicit that this is not a turnkey safety product; the constitution, the corpus, and the escalation labels remain the practitioner's work.
The verifiability of escalation correctness depends on the reliability of the should_escalate labels themselves; here they are author-generated from templates, not annotated by legal-services compliance professionals, and no inter-rater agreement is reported, so the effect of label quality on the conclusions remains an open question. The constitutional-alignment score depends on a model judge; a heuristic judge was used here for reproducibility, and a hosted-API judge would be expected to produce different absolute scores, with judge calibration and bias not yet addressed. The runtime supervisor produced zero interventions across both GPU pilots, because the model cited a provision on every answered turn and missed-escalation trajectories did not reach a terminal answer for the gate to inspect, so judge-based competence-boundary triggers require either a calibrated judge at inference time or richer base-model behaviour. The two runs did not pin container image digests or package microversions and did not enable deterministic CUDA flags, so the sources of variance can only be listed as candidates rather than confirmed. In addition, the full text was loaded here, but some table values appear as placeholders in the body text, so anyone verifying specific numbers item by item should return to the original tables.
