Continuous process-level evaluation of evolving enterprise agent skills: 162 of 175 final-value-correct trials still flagged
Synopsis
The work presents a continuous evaluation framework for enterprise agent skills, applied to two Business Value Determination skills in a real enterprise VAR system, using independently computed live-API ground truth and template test cases with runtime-resolved placeholders to check tool selection, arguments, execution ordering, and database integrity via programmatic checks plus a narrowly scoped LLM judge; across 240 automatic trials, 175 passed all applicable final numerical checks, yet 162 of them (92.6%; Wilson 95% CI 87.7–95.6%) still had at least one evaluator-detected deviation, and under a broader seven-check final-state definition 151/164 (92.1%; 95% CI 86.9–95.3%) still violated a trajectory check, with dependency attribution compressing a mean 6.34 failed checks to 2.65 roots.
Figure 1: Continuous evaluation pipeline. Stages 1 and 5 are new for each run; Stages 2–4 reuse a template suite within each skill across the evaluated specification, model, and harness configurations.
arXiv · Page 5Interpretation
The framework extends evaluation from final values to the process layer: tool selection, argument correctness, execution ordering, and database integrity, referenced against live-API ground truth computed independently of the evaluated trajectory. Compared with outcome-only validation, it detects process deviations even when final values are correct; compared with LLM-generated oracles, it computes ground truth by calling the same real APIs the skill uses. 240 automatic trials over two skills (104 checks for Revenue, 80 for Productivity), two specification variants, two harnesses, and three models, with 10 trials per cell; ground truth is computed by a standalone Python script via an independent execution path.
Correct final values do not establish process conformance: of 175 trials passing all applicable final numerical checks, 162 (92.6%) retained at least one additional deviation; broadening the baseline to seven final-state checks per skill left 151/164 (92.1%) still violating a trajectory check. The result holds after excluding external-error runs (131/144, 91.0%) and unrecovered-error runs (136/147, 92.5%), and the 162 count is unchanged when the no_irrelevant_tools check is removed. All 120 Revenue trials passed the final revenue check and all 120 had at least one additional deviation; root families span data or numerical consistency (106), missing phase or required tool (66), allowlist (55), redundancy or call-count (43), and scope/filter (13).
Template test cases with runtime-resolved placeholders serve as reusable regression artifacts, with one registered suite reused across all 24 model–harness–variant–skill conditions. Placeholders resolve at evaluation time from live API state rather than being hardcoded, so the same checks re-execute after specification revisions, model changes, or tool API changes; across archived materializations the same check IDs cover four Revenue and 16 Productivity expected-value states. One registered suite per skill (104 for Revenue, 80 for Productivity), with six legacy Cloudability checks removed from one archived result so all conditions share a common suite.
Specification-revision sensitivity varies jointly with model and harness: on Revenue, GPT-5.6-Sol is variant-robust on Claude Code (1.2 pp) but most sensitive on Codex (12.9 pp), while Sonnet 4.6 reverses (6.3 vs 0.4 pp). Bootstrap difference-in-differences intervals exclude zero for all three Revenue comparisons and for GPT on Productivity, showing the revision effect can invert across harnesses; the Productivity Opus and Sonnet comparisons remain inconclusive. Ten trials per cell with 20,000-resample nonparametric bootstrap intervals, unadjusted for six comparisons and exploratory; MD and TXT also differ in content, so effects cannot be attributed to formatting alone.
Perspective
The framework targets continuously evolving enterprise agent skills whose outputs are tool calls and persisted database state and can be checked programmatically; it provides reusable regression coverage for two BVD skills (Revenue and Productivity) across two harnesses (Claude Code and Codex), three models, and two specification variants (MD and TXT). It can serve as a quality gate and specification-completeness tool after skill revisions, model upgrades, or harness migrations, helping teams locate process-level regressions before deployment and narrow manual review via root-cause attribution.
Live-API ground truth can drift between execution and deferred evaluation, so immutable response snapshots would strengthen reference integrity; the LLM judge is used only for narrowly designated checks, and a two-author audit agreed on 57 of 60 randomly sampled judged cases (95%), but that audit sampled judgments rather than runs, was not stratified across check types, used authors as reviewers, and the judge model (Claude Sonnet 4.6) is itself one of the evaluated models, so it should not be read as general judge accuracy. The evaluator is normative: in 71/240 trials the irrelevant-tool check is a root failure, sometimes for harness utilities or schema initialization, and ordering and zero-valued EC2 memory/cost checks also encode workflow policy, so domain review must separate harmful behavior from benign alternatives. Traces contain execution noise (83/240 runs record at least one tool or harness error), unrecovered errors concentrate in Codex (38 versus 5), and all 20 Productivity Codex–Sonnet runs contain one, so clean-execution effects for that configuration are not estimable. Experiments compare fixed configurations rather than a longitudinal API or model upgrade, each skill uses one portfolio, one revenue value, and one date window, and repeated stochastic trials do not provide environmental input diversity; a post hoc staleness simulation freezing modal Productivity expected values would place at least one value outside tolerance in 80/120 Productivity runs. The framework diagnoses but does not repair failures, and cost figures are harness-recorded estimates rather than reconciled invoices.
