Skip to main content
Back to timeline
arXivSource publication:

EASE adapts logit perturbation and sampling temperature by predictive entropy during generation, cutting average TPR@1%FPR across five detectors on Qwen3-8B from 0.655 to 0.339

Synopsis

Observing that perturbing next-token logits or adjusting sampling temperature reduces AI-generated-text detector performance, the authors propose EASE, a training-free and detector-agnostic framework that computes predictive entropy from the source LLM's next-token distribution and uses it to adapt both logit perturbation strength and sampling temperature, consistently reducing detection performance across three source LLMs and multiple detectors with negligible inference overhead and limited text-quality degradation.

Source-provided article image: EASE: Entropy-Adaptive Distribution Shaping for Evading AI-generated Text Detectors
Figure 1 ·

Figure 1: Effects of decoding changes on AIGT detection and text quality. AUROC and perplexity under (a,b) top- k k logit perturbation with strength s s at τ = 1 \tau=1 and (c,d) varying sampling temperature without logit perturbation. AUROC curves show means over all eight detectors and within two groups: statistical (Likelihood [ 10 ] , Entropy and LogRank [ 4 ] , and LRR [ 25 ] ) and supervised (RoBERTa-Base/Large [ 14 ] , MAGE [ 11 ] , and RADAR [ 7 ] ). Dash-dotted lines mark vanilla baselines.

arXiv

Interpretation

The paper reports a mechanistic observation: adding perturbation to next-token logits or raising sampling temperature lowers detector AUROC while increasing generation perplexity, indicating detector sensitivity to decoding-time distribution changes. Prior evasion work centered on text-level rewriting or generation-process modifications requiring detector feedback, auxiliary models, or fine-tuning; this work isolates changes in the decoding distribution itself as a measurable signal. A small-scale observation study using the Section 4.1 setup with a different random seed, adding Gaussian noise to top-k logits and sweeping perturbation strength, and separately sweeping temperature without logit perturbation, with results shown as AUROC and perplexity curves.

EASE uses predictive entropy to drive two operations: at low entropy (concentrated distributions) it applies larger logit perturbations and higher sampling temperatures to encourage alternative token selections, and both adjustments diminish as the candidate distribution approaches uniformity; logit offsets use a sinusoidal token-dependent modulation to set sign and relative magnitude. Fixed perturbation strength and fixed temperature are replaced by per-step control tied to the concentration of the current candidate distribution, without detector feedback, additional learned parameters, or model fine-tuning. The method is specified in Equations (2)-(5); the ablation (Table 4) shows that removing entropy adaptation yields AUROC 0.902 and TPR@1%FPR 0.361, removing temperature adaptation yields 0.927 / 0.487, and full EASE yields 0.879 / 0.339, with perplexities of 7.576, 5.202, and 7.752 respectively.

On Qwen3-8B, EASE-Plugin lowers the average AUROC across five detectors from 0.952 to 0.879 and average TPR@1%FPR from 0.655 to 0.339 relative to vanilla generation, achieving the lowest values on both RoBERTa detectors; its perplexity of 7.752 is below the compared paraphrasing methods (9.867-15.196). Compared with simple paraphrasing, recursive paraphrasing, and AdvPara baselines, EASE acts during initial generation and does not require post-generation rewriting. Table 1 reports PPL and AUROC / TPR@1%FPR for five detectors on Qwen3-8B; for example, on RoBERTa-Large TPR@1%FPR drops from 0.535 to 0.250, whereas the evaluated paraphrasing baselines range from 0.583 to 0.615.

In cross-model and cross-detector evaluation, EASE reduces AUROC and lowers or maintains TPR@1%FPR across all detector-model pairs; the average TPR@1%FPR over eight detectors falls from 0.425 to 0.203 on Qwen3-8B, from 0.446 to 0.146 on Llama3-8B, and from 0.141 to 0.046 on Ministral-3-8B. The evasion effect is extended from a single model-detector combination to three source LLMs and multiple detectors, with BLEU/ROUGE and perplexity reported alongside to quantify text-quality changes. Table 2 covers eight detectors and three source LLMs; Table 3 shows BLEU-1 and ROUGE-1 remain relatively stable, BLEU-2 and ROUGE-2 show small absolute decreases, and perplexity rises (e.g., from 4.243 to 7.752 on Qwen3-8B).

Perspective

The framework targets a deployer who controls the source LLM's decoding process: it can modify next-token logits and sampling temperature without querying detectors or fine-tuning the model, and applies either during initial generation (EASE-Plugin) or in a subsequent rewriting stage (EASE-Rewrite). Experiments use 2,000 human-written sequences from WikiText-103, taking the first 30 tokens as prompts to generate continuations of up to 200 tokens, with Qwen3-8B, Llama3-8B-Instruct, and Ministral-3-8B-Instruct as source models. For detector developers, this suggests evaluation should cover diverse decoding settings and attend to decoding-induced changes in text statistics.

Effects are uneven across detectors: Likelihood and Binoculars show pronounced reductions in low-false-positive-rate detection, whereas LRR retains substantial discriminative ability; some detectors already perform near or below chance under vanilla generation, so further AUROC reductions in those settings should be interpreted under the fixed score orientation. The relative advantage of EASE-Rewrite versus EASE-Plugin varies with detector and metric, and the AUROC advantage of rewriting does not consistently translate into stronger evasion at a low false-positive rate. On text quality, BLEU-1 and ROUGE-1 remain relatively stable while perplexity rises, indicating a trade-off between quality cost and attack strength. In addition, the observation study is small-scale and uses a different random seed, so generalization to larger scales and more domains remains an open question.

Sources