Skip to main content
Back to timeline
arXivSource publication:

Making machine text sound more human made it easier to detect: a RoBERTa detector tracked statistical complexity and showed a 76.3% false-positive rate on formal human writing

Related research and updates

Synopsis

Using the M4 dataset (N = 10,000) and controlled generations (N = 300), this work perturbs a RoBERTa-based AI-text detector at semantic, structural, and tokenizer levels and finds that asking Mistral-7B-Instruct to make machine text sound more human raised Verb Diversity from 0.77 to 0.92 while making outputs easier to detect, that detection scores appear to track statistical complexity and yield a 76.3% false-positive rate on formal human writing, and, as a control, that event-based Latent Space detection had 87% of its event sequences changed by paraphrasing (Jaccard = 0.067) and 70% of extracted verbs altered by homoglyphs (Jaccard = 0.30), with a best domain AUC of 0.577.

Source-provided article image: Ontological Instability and Statistical Amplification: The Paradox of "Humanizing" LLM-Generated Text
Figure 1 ·

Figure 1: Left: Verb Diversity before (0.766) and after (0.924) Mistral-7B rewriting, with the human baseline (0.572). Right: stability of statistical features (79.4%) and of event sequences (13.1%) under rewriting.

arXiv

Interpretation

Supervised AI-text detectors report high benchmark accuracy, but it is not clear what their decisions are based on; this work probes what a RoBERTa-based detector actually relies on through semantic, structural, and tokenizer-level perturbations. Unlike prior work that reports benchmark accuracy alone, this places the detector under controlled perturbations to observe which features its decisions follow, turning the question from accuracy to evidential basis. Uses the M4 dataset (N = 10,000) and controlled generations (N = 300), with semantic, structural, and tokenizer-level perturbation conditions.

When Mistral-7B-Instruct was asked to make machine text sound more human, Verb Diversity rose from 0.77 to 0.92 and the outputs became easier to detect; detection scores appear to track statistical complexity. This surfaces a counterintuitive effect: rewriting aimed at humanization did not reduce detectability but instead made outputs easier to identify as statistical complexity increased. Based on controlled-generation rewriting experiments (N = 300), reporting a Verb Diversity change from 0.77 to 0.92.

The detector showed a 76.3% false-positive rate on formal human writing. This figure links the detector's reliance on statistical complexity to misclassification of formal human prose, indicating that high benchmark accuracy does not directly translate into reliability on real writing. Measured a 76.3% false-positive rate on formal human writing samples.

As a control, event-based Latent Space detection had 87% of its event sequences changed by paraphrasing (Jaccard = 0.067) and 70% of extracted verbs altered by homoglyphs (Jaccard = 0.30), with a best domain AUC of 0.577; RoBERTa's robustness appears specific to the features it uses, and structural abstraction did not make detection more robust. This control shows another detection paradigm is likewise sensitive to semantic and character-level perturbations, and that the structural-abstraction route did not deliver the expected robustness gain. Reports 87% event-sequence change after paraphrasing with Jaccard = 0.067, 70% extracted-verb change under homoglyphs with Jaccard = 0.30, and a best domain AUC of 0.577.

Perspective

This work addresses researchers and deployers using supervised AI-text detectors, in English-text settings represented by the M4 dataset and controlled generations (N = 300), and under semantic, structural, and tokenizer-level perturbations. It enables follow-up work to pursue the lead that detection scores track statistical complexity, to examine the sources of false positives on formal human writing, and to compare detection paradigms such as event-based Latent Space detection under paraphrasing and homoglyph perturbations.

The phrasing that detection scores 'appear' to track statistical complexity is itself uncertain, and the precise mechanism still needs characterization; the 76.3% false-positive rate comes from the specific genre of formal human writing, and behavior on other genres is unclear; the control detector's best domain AUC of 0.577 indicates limited discrimination in the current setting, and its behavior across domains and perturbation combinations remains an open question; moreover, the loaded text is abstract-level content without figures or full experimental detail, so further judgment about perturbation strength, sample composition, and statistical testing requires consulting the original.

Sources