TeleAntiFraud 2.0: A Refreshable, Profile-Grounded, and Audio-Based Benchmark for Telecom Fraud Detection
Synopsis
The work presents TeleAntiFraud 2.0, a refreshable Chinese call-audio benchmark organized as monthly frozen snapshots, built with a Mixed-Tree Anti-Fraud Generation Pipeline that turns online fraud-case abstracts into profile-grounded scenarios and expands them into fraud and near-domain lawful sibling dialogues sharing context and diverging only at label-bearing actions, rendered as role-matched speech; each frozen set contains 900 Chinese calls (600 fraud, 300 near-domain non-fraud), controlled text experiments show three classifiers reach perfect Macro-F1 against unrelated or ordinary negatives but drop to 0.65-0.68 with near-domain sibling negatives, and full-set audio and ASR+LLM evaluations further reveal class-prior shortcuts, prediction collapse, and snapshot sensitivity.
Figure 1: Mixed-tree construction and snapshot-level diagnostics, including tree-distance structure, transfer difficulty, prediction collapse, and snapshot variation.
arXivInterpretation
Introduces a versioned, monthly frozen audio benchmark so newly observed scam patterns can be incorporated into later releases while every published snapshot remains immutable, preserving audio, labels and rationales, generation metadata, evaluation prompts, model responses, and provenance records. Compared with fixed test sets such as FDB's shared splits, Fraud-R1's fixed interaction set, and TeleAntiFraud-28k's fixed audio-text set, this design absorbs new cases without overwriting prior evaluation sets, supporting reproducible and auditable comparison over time. Table 1 lists the 'Refresh' column as Monthly and 'Frozen eval.' as Monthly snaps., and the paper states that the two current snapshots support snapshot-sensitivity analysis while longer-term temporal claims require additional releases.
Develops the Mixed-Tree Anti-Fraud Generation Pipeline, converting case abstracts through scenario profiling, mixed-tree expansion, collaborative dialogue realization, and speech rendering into traceable call audio, where fraud and lawful sibling paths share participants, scenario context, opening turns, and early risk language and diverge only after label-bearing actions emerge. Relative to generating fraud and non-fraud calls independently from the same profile, the sibling-path design avoids label-correlated differences in opening context or conversational framing, making the label depend on the completed interaction trajectory rather than topic-level cues. The paper gives a formal definition of the mixed tree (node set, edge set, shared root, node-attribute mapping, path-state mapping, and state-transition function) and states that the state takes ambiguous, fraud, and non-fraud values, updated when an action-labeled edge is traversed.
Controlled text and full-set audio evaluations show that unrelated or ordinary telecom negatives make the task nearly perfectly separable, whereas near-domain sibling negatives reduce Macro-F1 to 0.65-0.68; full-set audio and ASR+LLM evaluations further expose false-positive bias, class-prior shortcuts, prediction collapse, and substantial variation across monthly snapshots. These results indicate that fraud-class F1 alone is insufficient for telecom-fraud evaluation and motivate joint reporting of class-balanced metrics, class-conditional recall, prediction distributions, and collapse behavior. The paper reports the behavior of three classifiers in controlled text experiments and observes these diagnostic behaviors in full-set audio and ASR+LLM evaluations; the specific model list and per-item numbers require consulting the original tables.
Perspective
The benchmark targets Chinese call-audio settings, with each frozen set containing 900 calls (600 fraud, 300 near-domain non-fraud), suited to evaluating discrimination between fraud and near-domain lawful calls under shared context and to snapshot-sensitivity analysis across monthly releases; its pipeline design allows newly collected case abstracts to be incorporated into later monthly snapshots, making it appropriate for teams and evaluators who need to track emerging scam patterns while preserving prior evaluation records and provenance.
The paper notes that only two snapshots currently exist, so longer-term temporal claims require additional releases; moreover, the specific magnitudes of the false-positive bias, class-prior shortcuts, and prediction collapse observed in full-set audio and ASR+LLM evaluations, which model families are involved, and the detailed procedures and agreement results of the 10-expert audit and pilot-plus-audit checks require further verification against the original tables and appendices.
