Skip to main content
Back to timeline
arXivSource publication:

Boundary tests for conversational agents stop guessing: conditions pulled from the agent's own replies lift explicitly defined tests to 62.7%–83.3%

Lead

Boundary tests for conversational agents now come from condition statements extracted from the agent's own replies and paired with transformation instructions; across four domains, 62.7%–83.3% of tests expect behavior explicitly defined by the prompt, tool code, or knowledge base, up from 47.7%–68.1% for AgentEval.

Source-provided article image: AdaT^2: Adaptive Test Transformations for Black-Box Boundary Testing of Conversational Agents
Fig. 1 · arXiv · Page 2

Story

Boundary tests for conversational agents can now take their conditions from the agent's own replies, so the behavior a test expects comes directly from those conditions instead of being supplied by the model writing the test. The earlier approach mined a workflow graph from exploratory conversations and had the generator target the graph's nodes and edges; a node or edge merges replies of the same kind of activity, so one target can hold several boundaries, and because the graph does not state expected behavior, the generator had to infer it. Across four domains, 62.7% to 83.3% of AdaT2's tests are valid boundary tests whose expected behavior is explicitly defined by the agent's prompt, tool code, or knowledge base, against 47.7% to 68.1% for AgentEval's test suites.

Transformed tests reach the other side of a statement's boundary or a different boundary, adding explicitly defined boundaries that single-statement tests miss. A single statement describes only one side of its own boundary, and a test built from it stays on that side. Across the four domains, transformed tests contribute 13 to 46 additional explicit boundaries, and the suites cover 71 to 147 explicit boundaries and 84 to 158 boundary sides.

As regression tests, the suites detect all eight seeded policy faults in the airline domain and four of eight in the retail domain. AgentEval's test suites, with fewer than a third as many tests, detect five and two respectively. Each fault is a predefined policy flip that rewrites one policy passage; a test counts as detecting it when it passes on the original agent, fails on the faulty agent, and the auditor attributes that failure to the fault.

What to watch

Teams auditing policy compliance of conversational agents can use this pipeline under black-box conditions, where only user messages and agent replies are visible, to extract condition statements from exploratory conversations, pair them with transformation instructions, and run the result as a regression suite. The next useful step is to cover both sides of more explicit boundaries and to reduce inconclusive verdicts, since both methods currently cover fewer than a quarter of their explicit boundaries on both sides in every domain.

Each method produced one test suite per domain, and the design-choice comparison used three runs, too few for significance tests, so the differences in rates across domains need more repetitions to confirm. Validity labels, boundary and side counts, and most execution labels come from the auditor model, so absolute rates depend on that model, while comparisons within a domain share one auditor. Seeded faults rewrite only policy text in the agent's prompt, and four retail faults go undetected by all three suites, covering user authentication by a stated user ID, transfer to a human agent, accepted reasons for canceling an order, and exchanging a delivered item only for a different option of the same product. The share of executions labeled inconclusive is high, mostly because the judge returns inconclusive verdicts on plain tests, which lowers suite accuracy.

Sources