Skip to main content
Back to timeline
arXivSource publication:

Four agentic runs produced customer-feedback taxonomies that passed every generic check while 97.7%–100% of leaf names restated an ancestor and branch leakage reached 76.4%

Synopsis

Building six taxonomies over two proprietary customer-feedback corpora (1,940 and 5,000 records) — one production reference plus two repeated runs of the same agentic harness per corpus — the authors found all six pass every generic naming and structure check, yet at least 97.7% of leaf names in every generated tree merely restate an ancestor's name and one tree routes three of four records into multiple top-level categories; they introduce two judge-free whole-tree metric families, structural discriminators and team partitionability.

Source-provided article image: From Plausible Hierarchies to Useful Taxonomies: Evaluating Agentic Harnesses on Customer Feedback
Figure 3 ·

Figure 3: Branch leakage and co-occurrence modularity (Q,

arXiv · Page 6

Interpretation

Generic, individually scoped LLM-judged checks fail to detect whole-tree failure modes: templated-label collapse, generator-imposed branching caps, and cross-team fragmentation. Prior taxonomy evaluation (name-description match, description completeness, hierarchy correctness, sibling distinctness) inspects one node, edge, or sibling pair at a time, while none of these three failures is a property of any single node. All six taxonomies (one production reference plus two repeated runs of the same harness per corpus) pass every generic check: description completeness 1.00 across the board, name-description match 0.97–1.00; hierarchy correctness is 1.00 for C-1 and 0.92 for C-2.

Generated trees build leaf names almost entirely by concatenating an ancestor's name, unlike the references. The authors propose structural discriminators — ancestor containment, token redundancy, type-token ratio, fan-out shape, cross-branch reuse — turning 'does a label add information beyond its ancestors' into computable whole-tree statistics. Ancestor containment is 97.7%–100% for the four generated artifacts versus 13.9% and 2.9% for the two references; type-token ratio collapses 2.5–4.4 on the small sample and 2.0–2.7 on the large; all four generated artifacts have zero cross-branch reuse against 48 and 73 shared nodes in the references.

Generated branches do not partition the record stream into groups teams can own, and this difference is invisible to the generic checks. The authors propose the team-partitionability family (branch leakage, record purity, co-occurrence modularity, dividability), operationalizing the node-level mutual-exclusiveness prescription as a countable record-level violation. Reference L1 leakage is 19.8% and 19.3% versus 34.7%–76.4% for the generated artifacts; C-2's L1 modularity is 0.12 and turns negative at L2, meaning its branch boundaries are worse than a random split; C-1 and C-2 differ by only 0.08 in hierarchy correctness yet by 27 percentage points in leakage with a modularity sign flip.

A deeper coverage check actively rewards these failure modes, since minting a node per documented term is the cheapest way to raise coverage. Product-term coverage is the one check on which generated artifacts beat the references, and the authors show it points in the opposite direction from whole-tree structural quality. Generated artifacts score 0.74–0.86 versus 0.72 on the small sample and 0.90–0.91 versus 0.80 on the large; C-2, the worst on leakage, posts the highest small-sample coverage (0.86).

Perspective

The work targets settings where a customer-feedback taxonomy is a production artifact and top-level branches are assigned to distinct product teams: the metrics need only the taxonomy artifact (structural discriminators) and that artifact's own record classifications (team partitionability), with no LLM judge and no gold taxonomy, so they apply unchanged to any taxonomy artifact meeting the same input contract. The authors note the reference (Live) taxonomies were built by the data owner's in-house pipeline, which was engineered around checks similar in spirit to these metrics, so their favorable profiles are partly by construction; the confound-free evidence is the within-harness contrast among generated runs. Future directions include temporal stability, a second harness family, tuned thresholds, and insertion-time construction gates.

All generated runs use one harness, so harness diversity is untested; the cap-suspect threshold can false-positive on genuinely narrow branching and does once, on the large-sample reference; type-token ratio is size-sensitive across very different scales; leakage, purity, and modularity depend on each artifact's own record classification, so they evaluate the taxonomy-plus-classifier system; neither family is validated against an external usefulness signal such as team velocity; and every metric in Table 1 is LLM-judged without a human-agreement study. The corpora, artifacts, and documentation are proprietary customer data and are not released, so Tables 2 and 3 cannot be independently recomputed, though the formulas and protocol are fully specified for reimplementation.

Sources