Skip to main content
Back to timeline
arXivSource publication:

Structured Skill Optimization Under Frozen Weights: A New Path for Audio Anti-Fraud Detection

Synopsis

Without changing any parameters of the audio-language model, FRAUDSkill optimizes an external layer of skill programs, routing policies, and decision rules and combines structured output control with validation-guided multi-path inference, raising Macro-F1 on the TeleAntiFraud benchmark from 41.54% for the shared frozen-model baseline to 73.50% while cutting the invalid-output rate from 36.04% to 1.94%.

AI-generated editorial illustration: FRAUDSkill: Structured Frozen-Weight Skill Optimization for Audio Anti-Fraud Detection

Interpretation

It reformulates audio anti-fraud detection as a structured frozen-weight adaptation problem: a frozen audio-language model must complete an ordered decision chain of service-scenario identification, fraud detection, and conditional fraud-type classification, with every output falling inside the official label space. Earlier work typically encodes task knowledge, label constraints, and decision rules into model parameters (fine-tuning) or manually maintained prompts; here the adaptation target is explicitly an editable layer outside the model, with emphasis on the closed-set and sequential dependencies of the decision chain. The paper gives a formal problem setting that distinguishes a required output that is missing or unmappable to an official label from a route that should be skipped under the protocol, and defines the feasible output space accordingly; experiments use the official TeleAntiFraud SFT split (after audio-level deduplication, 10,711 training and 2,677 test samples, of which 1,453 carry fraud-type annotations) under the original sequential decision protocol.

It proposes the FRAUDSkill framework: offline, trajectory-level error diagnosis, critic feedback, and an editor revise external skill programs, with beam search selection on held-out validation; at deployment, label projection, route normalization, and a validation-fitted selector turn open-ended generation into protocol-compliant closed-set predictions. Compared with representing task information as a single editable document (SkillOpt) or a structured skill folder (EvoSkill), this work explicitly models the three-stage decision route, projection onto the official label set, and cross-route consistency, and separates output constraints and final selection from any single textual program. The paper formalizes the full method (root instruction, skill library, route policy, label projection operator, route normalization, selector fitting) and states that all retained programs and search settings are frozen before testing and that test labels are not used for program selection or selector fitting; the appendix gives the search configuration (seed 42, beam width 3, branch factor 3, five search rounds, batches of 64 examples).

Experiments show that rewriting skill text alone yields limited and unstable gains, while structured inference components contribute the main improvement: closed-set projection raises Macro-F1 from 43.66% to 54.47%, route normalization further to 66.21%, complementary multi-path inference to 70.31%, reliability weighting to 70.84%, and class-balanced selection finally to 73.50%. This ablation separates two distinct problems, output validity and class discrimination, indicating that text-layer skill search alone struggles to carry both open-ended generation and structured decision control. The ablation starts from the best textual program and adds structured components cumulatively; FRAUDSkill-Text averages 42.42% Macro-F1 across seeds 42, 43, and 44 with a sample standard deviation of 1.39 points, and the paper itself notes that its mean gain over the baseline is smaller than that variation.

The complete system, with a frozen audio model (Qwen2-Audio-7B-Instruct, deterministic decoding), reaches 73.50% Macro-F1, 79.40% W-F1, 78.72% accuracy, 58.87% joint accuracy, and a 1.94% invalid-output rate; as parameter-updating references, SFT reaches 66.06% and SFT+Memory 75.51%. The paper explicitly lists SFT and SFT+Memory as parameter-training references rather than direct baselines, and notes that these comparisons are protocol-bound and do not imply a task-independent ranking. All directly compared methods share the same frozen audio model and the same label ontology, output schema, audio-evidence guidance, and cross-turn consistency rules, so differences arise from how external task knowledge is represented, optimized, and applied; the evaluation unit is a unique audio recording, which differs from the 7,021 interaction-level records used in the original dataset paper, so results under the two protocols are not directly comparable.

Perspective

The paper's conclusions are scoped to the studied protocol: the evaluation unit is a unique audio recording, which differs from the 7,021 interaction-level records used in the original dataset paper, so results under the two protocols are not directly comparable; SFT and SFT+Memory are parameter-updating references rather than direct baselines, and the comparison does not imply a task-independent ranking. The method suits anti-fraud deployment settings that require closed-set labels and a sequential decision protocol; its external skill programs, label maps, route rules, and selector are frozen before testing and are inspectable and replaceable, making it possible to update the external layer without retraining the audio model when fraud patterns or labeling policies change. For engineering and research teams seeking protocol-compliant predictions without changing model weights, this path offers a reusable way of organizing the problem; a full reproduction requires the TeleAntiFraud audio split, the frozen audio-language model, and the critic/editor model interfaces used for skill search.

Text-level skill search is stochastic: the paper reports a sample standard deviation of 1.39 points for FRAUDSkill-Text across three seeds and notes that its mean gain over the baseline is smaller than that variation, so a single text-search run should be read with care. Route-level error analysis shows a 74.2% scene-route error rate (91.1% of which are missing or off-ontology outputs), a 32.9% fraud-route error rate (60.0% of which are false-normal decisions), and an 83.1% type-route error rate on annotated examples, a distribution that suggests upstream errors can block downstream routes and warrants continued observation in follow-up work. In addition, raw audio files and frozen model weights are not duplicated in the supplement, so a full rerun depends on external data and model interfaces; the paper does not use the original benchmark's LLM-based score for slow-thinking rationales because FRAUDSkill produces closed-set decisions rather than free-form reasoning traces.

Sources