EvoOntology: A Self-Evolving Ontology Layer for Data Agents
Synopsis
The work introduces EvoOntology, which encapsulates an ontology as an MCP server with schema, content, and tool layers, builds an initial ontology via a builder agent issuing probe queries, and continuously refines it through attribution-guided typed edits admitted only after backbone-conditional paired evaluation, consistently outperforming ReAct baselines and static semantic-layer baselines on three data-agent benchmarks with six LLM backbones.
Interpretation
It proposes the first autonomous interactive ontology layer for data agents, encapsulated as an MCP server so agents actively query the ontology at runtime rather than passively consuming contextual metadata. Compared with raw-querying methods that let agents explore raw sources directly and semantic-layer methods that inject a manually constructed layer into the prompt, the ontology is split into a schema layer (object types and reference rules), a content layer (Terms, Mappings, Constraints, Evidence nodes with Semantic Relations and Structural References edges), and a tool layer (browse and resolve MCP tools plus a session manifest); only the manifest enters the prompt, while detailed records are retrieved on demand. The paper describes the architecture and reports DDR-Bench comparisons: Baseline + SL even drops on Claude-Sonnet-5, whereas EvoOntology improves on all six backbones; the authors attribute the difference to a static prompt fragment competing with other instructions and being unprunable per turn, while MCP tools retrieve only what the current step needs.
It introduces a builder agent and a self-evolution loop: a workload-guided probing process constructs an evidence-grounded initial ontology, which is then refined through trajectory attribution, single-level localized intervention, and backbone-conditional paired validation. Compared with static, manually maintained semantic layers, ontology maintenance becomes a closed loop driven by agent interaction trajectories, where each candidate edit modifies only one of Content, Tool, or Schema and is accepted only if it beats its parent on the same validation set under identical decoding and interaction budgets. Ablations show removing the gate causes the largest Traj-Wise drop, removing attribution the next largest, and removing diagnose or replacing typed patches with free-form rewrites smaller drops, leading the authors to call the gate and attribution the load-bearing pieces; each editable level alone falls short of the full three-level loop.
It improves consistently across three data-agent benchmarks, with gains coming from the initial ontology plus subsequent evolution. Relative to the no-ontology baseline and the static semantic-layer baseline, EvoOntology improves Trajectory-Wise on all six backbones on DDR-Bench, Overall on InsightBench, and EX and VES on BIRD. On DDR-Bench, GPT-5.5 Traj-Wise rises from 64.2 to 90.9 and GPT-5.6-sol from 68.5 to 93.5; on BIRD, Claude-Opus-4.8 EX rises from 67.5 to 78.3; InsightBench gains are smaller, which the authors explain by Insight being graded against short reference-style findings that saturate once aligned.
The ontology layer lowers total token cost while improving performance, and different backbones evolve different ontologies. Relative to the baseline, the ontology layer increases input tokens per turn but shortens trajectories, reducing total tokens per task; cross-backbone transfer shows performance drops when one backbone's evolved ontology is applied to another. The appendix reports Baseline at 52.6K tokens and 14.6 turns per task versus Evolved at 42.0K tokens and 8.4 turns, with Traj-Wise rising from 69.5 to 89.5; pairwise Jaccard overlap stays below a stated ceiling, the diagonal is the highest entry of each column, and off-diagonal entries drop by at least a stated margin.
Perspective
The result targets data agents that must solve natural-language tasks over heterogeneous sources such as relational databases, semi-structured filings, and unstructured documents, in settings where the ontology can be built and refined automatically from probe queries and interaction trajectories and where a held-out validation set exists for paired evaluation; the paper's validation covers the 10-K scenario of DDR-Bench, InsightBench, and BIRD under the Oracle Knowledge setting, with GPT-5.5, GPT-5.6-sol, Claude-Sonnet-5, Claude-Opus-4.8, DeepSeek-V4-Flash, and Qwen3.5-Flash as backbones.
The paper text is loaded as abstract and section prose, and several numeric values appear as placeholders in the prose, so some specific gain figures can only be read from the tables; the authors also note that identifier overlap alone cannot establish semantic equivalence across backbones. In addition, evolution flattens over the last two rounds, which the authors explain as failure signatures becoming rarer; whether this convergence behavior recurs under other data sources and task distributions remains an open question worth watching.
