Skip to main content
Back to timeline
arXivSource publication:

A Refreshable, Near-Domain-Controlled Audio Benchmark for Telecom Fraud Detection: TeleAntiFraud 2.0

Synopsis

The work presents TeleAntiFraud 2.0, a refreshable Chinese audio benchmark for telecom fraud detection organized as monthly frozen snapshots, in which mixed-tree generation makes fraud and lawful near-domain calls share context and diverge only at label-bearing actions, and it reports that three text classifiers reach perfect Macro-F1 against unrelated or ordinary negatives but drop to 0.65-0.68 with near-domain sibling negatives, while full-set audio and ASR+LLM evaluations expose class-prior shortcuts, prediction collapse, and snapshot sensitivity.

AI-generated editorial illustration: TeleAntiFraud 2.0: A Refreshable, Profile-Grounded, and Audio-Based Benchmark for Telecom Fraud Detection

Interpretation

It introduces a refreshable audio benchmark organized as immutable monthly snapshots, each frozen set containing 900 Chinese calls (600 fraud and 300 near-domain non-fraud), with audio, labels, prompts, manifests, and provenance records frozen. Whereas fixed test sets cannot absorb newly observed scam patterns and continually replacing old examples makes results hard to reproduce and compare, this design lets newly collected fraud case abstracts enter later releases through the same pipeline while every published snapshot stays immutable. The paper reports two independently constructed snapshots (June/V1 and July/V2), each with 900 calls and a fixed 2:1 composition, sharing generation, synthesis, and evaluation contracts but independently sampling scenarios, dialogue trees, and voice assignments.

It develops the Mixed-Tree Anti-Fraud Generation Pipeline, which turns online fraud case abstracts into structured scenario profiles, expands them into mixed dialogue trees, and makes fraud and lawful sibling paths share participants, context, opening turns, and early risk language, diverging only after label-bearing actions emerge. Prior non-fraud negatives often came from unrelated topics or different sources, letting models rely on lexical or source-specific shortcuts; shared-context sibling paths bind the label to the completed interaction trajectory rather than topic-level cues. In controlled text experiments, Bigram, LR, SVM, and RoBERTa reach Macro-F1 of 1.000 on unrelated random and ordinary in-domain negatives but fall to 0.650-0.680 on near-domain sibling negatives, with lexical overlap rising across negative difficulty (0.005, 0.161, 0.274).

Full-set audio and ASR+LLM evaluations reveal prediction collapse, class-prior shortcuts, and snapshot sensitivity rather than performance summarizable by a single score. The paper argues that fraud-class F1 alone is insufficient and that Macro-F1, Balanced Accuracy, both recalls, and prediction-class distributions should be reported jointly, with raw predictions preserved. Before prompt-language averaging, 11 of 27 V1 configurations and 15 of 27 deduplicated V2 configurations show an all-FRAUD-like signature, with fraud recall near 1.0, accuracy near the 2:1 prior, and fraud F1 near 0.80; class-prior resampling shows the all-FRAUD baseline's fraud F1 rising from 0.500 to 0.800 as fraud prevalence increases while Balanced Accuracy stays at 0.500 and non-fraud recall at zero.

Construction analyses independently audit tree structure, label traceability, and speech-rendering controls. Beyond detection scores, the paper audits the generated artifacts themselves, including path-level label rationales, expert label review, and voice-pool coverage. BGE-small-zh embeddings over 813 dialogues give mean within-tree distance 0.0188 and cross-tree distance 0.3351 (ratio 17.8); leave-one-tree-out F1 drops from 0.883 in-domain to 0.605; ten experts each reviewed a 100-item packet, and across 1000 item-level judgments corrected-gold agreement averages 0.792 (range 0.700-0.880), with 494 FRAUD and 506 NONFRAUD judgments and 744 marked evidence-sufficient; a single-annotator pilot over 80 dialogues scores 4.60, 4.41, 4.35, and 4.13 on a 1-5 scale for dialogue realism, strategy coherence, victim reaction plausibility, and audio naturalness.

Perspective

The benchmark is intended for defensive research, tracking progress in audio telecom-fraud detection under reproducible and auditable conditions; its setting is Chinese calls organized as monthly frozen snapshots with a 2:1 fraud/non-fraud composition that retains near-domain lawful calls for boundary recognition. The pipeline admits newly collected case abstracts under a fixed schema, so later snapshots can incorporate emerging scam patterns without modifying earlier artifacts or results, and the manifest contract supports balanced 1:1 re-evaluation. The paper explicitly states that it uses synthetic dialogues and rendered speech because real fraud calls often contain private, security-sensitive, and potentially harmful content, positioning the benchmark as a controlled diagnostic instrument.

The paper notes that only two snapshots currently exist, so they support snapshot-sensitivity analysis while longer-term temporal claims require additional releases. In the cross-benchmark comparison, differing sources, audio pipelines, and label definitions mean the analysis measures relative linear separability under a shared classifier rather than serving as a direct benchmark ranking, so the gap cannot be attributed entirely to a single construction factor. On speech rendering, the 35 authorized reference voices include 29 male and 6 female profiles, and the paper lists demographic representativeness as a limitation; removing emotion tokens changes aggregate F1 from 0.920 to 0.960 with a maximum per-model difference of 4.6 points, so the paper treats emotion as a controlled rendering variable without a causal claim. Lower-agreement packets in the label audit point to difficult or ambiguous cases. In addition, clean CER/WER estimates require full-dialogue or segment-level transcription and are stated to remain outside the current scoring contract, which is an open question for readers who need finer speech-recognition error characterization.

Sources