PADBen: A Comprehensive Benchmark for Evaluating AI Text Detectors Against Paraphrase Attacks
Synopsis
Through dual representation space analysis, this work identifies an "intermediate laundering region" mechanism, builds the PADBen benchmark with a five-type text taxonomy and five progressive detection tasks, and evaluates 11 detectors, revealing a critical asymmetry: paraphrase attacks do not universally defeat detection—plagiarism evasion (paraphrasing LLM-generated text) remains detectable (RADAR sentence-pair AUC 0.909), while authorship obfuscation (paraphrasing human-authored text) collapses detection to near-random performance (AUC 0.526–0.748).
Figure 1: Overall pipeline for benchmark curation. Preprocessing details in Appendix A, data generation in Appendix C.3.
· Page 4Interpretation
Proposes the "intermediate laundering region" mechanism: iterative paraphrasing induces semantic displacement while preserving generation patterns, moving texts into an intermediate region that deviates from their origin yet retains generation characteristics. Prior benchmarks largely treat paraphrasing as a single-step perturbation without systematically characterizing how iterative paraphrasing evolves in representation space; this work uses dual representation analysis with BGE-M3 embeddings and Qwen3-4B hidden states to track centroid trajectories over 10 iterations, revealing that this region is universal (reachable from both human and LLM starting points) and stable (reliably reached via iteration). Based on quantitative distance analysis in BGE-M3 embeddings and Qwen3-4B hidden states (e.g., cosine distance between human-authored text and iteratively paraphrased human text rises from 0.085 to 0.134, while distance between LLM-generated text and iteratively paraphrased human text stays stable around 0.698), supplemented by PCA trajectory visualization; Experiment 2 samples 100 texts per category, and the authors note in limitations that sample size and length control could be further extended.
Constructs a five-type text taxonomy and five progressive detection tasks covering the full trajectory from original content to deeply laundered text, evaluating detector robustness under both single-sentence classification and sentence-pair recognition formats. Existing benchmarks such as RAID employ only single-step Dipper-based paraphrasing, and PARAPHRASUS focuses on paraphrase identification rather than adversarial robustness; PADBen is the first to distinguish authorship obfuscation from plagiarism evasion and to set up progressive tasks (Tasks 1–5) for systematic evaluation. The dataset is built from three public corpora—MRPC, HLPC, and PAWS—deduplicated with a cosine similarity threshold of 0.85, yielding 16,233 human-authored texts; generation uses multiple models including Gemini-2.5-Pro, DIPPER, and LLaMA-3-8B, with quality assessed via Jaccard similarity, perplexity, and self-BLEU (e.g., human paraphrase Jaccard 0.798, Type 5-3rd self-BLEU 0.170).
Evaluation of 11 detectors (4 zero-shot, 7 model-based) reveals a critical asymmetry: plagiarism evasion remains detectable, while authorship obfuscation collapses detection. Prior work did not evaluate both paraphrase attack scenarios under a single benchmark; this work empirically confirms through Task 3 vs. Task 5 comparison that although both attacks exploit the intermediate laundering region, they produce different detection signatures. In Task 3, RADAR achieves sentence-pair AUC 0.748 but single-sentence performance drops to 0.50–0.63, with all detectors near random; in Task 5, RADAR reaches sentence-pair AUC 0.909 and exhaustive single-sentence 0.803, while Kimi-K2-Instruct single-sentence AUC is 0.573; in Task 4, all detectors fall between AUC 0.487–0.529, showing universal failure.
Perspective
This benchmark targets researchers and practitioners evaluating AI text detector robustness, applicable to English sentence-level scenarios where paraphrase attacks are the core threat; its five-type taxonomy and five tasks can be used to compare detectors under authorship obfuscation and plagiarism evasion, and to provide an evaluation framework for subsequent detection architecture design.
The mechanism experiment samples 100 texts per category, and text length is not strictly held constant across paraphrasing iterations; the authors note in limitations that these aspects could be further extended. The taxonomy currently covers Type 5 at 1 and 3 iterations, while deeper iterations (e.g., 5–10) and intermediate iteratively-paraphrased human text variants are not yet included. Sentence-pair tasks always include paraphrased text in at least one position, and do not include direct human-original vs. LLM-original pairs. Detectors are evaluated with default configurations without fine-tuning on PADBen. These are directions for future exploration rather than shortcomings of the current work.
