Public articles linked to the same research event.
arXiv Using the M4 dataset (N = 10,000) and controlled generations (N = 300), this work perturbs a RoBERTa-based AI-text detector at semantic, structural, and tokenizer levels and finds that asking Mistral-7B-Instruct to make machine text sound more human raised Verb Diversity from 0.77 to 0.92 while making outputs easier to detect, that detection scores appear to track statistical complexity and yield a 76.3% false-positive rate on formal human writing, and, as a control, that event-based Latent Space detection had 87% of its event sequences changed by paraphrasing (Jaccard = 0.067) and 70% of extracted verbs altered by homoglyphs (Jaccard = 0.30), with a best domain AUC of 0.577.
Using the M4 dataset (N = 10,000) and controlled generations (N = 300), this work perturbs a RoBERTa-based AI-text detector at semantic, structural, and tokenizer levels and finds that asking Mistral-7B-Instruct to make machine text sound more human raised Verb Diversity from 0.77 to 0.92 while making outputs easier to detect, that detection scores appear to track statistical complexity and yield a 76.3% false-positive rate on formal human writing, and, as a control, that event-based Latent Space detection had 87% of its event sequences changed by paraphrasing (Jaccard = 0.067) and 70% of extracted verbs altered by homoglyphs (Jaccard = 0.30), with a best domain AUC of 0.577.
Using the M4 dataset (N = 10,000) and controlled generations (N = 300), this work perturbs a RoBERTa-based AI-text detector at semantic, structural, and tokenizer levels and finds that asking Mistral-7B-Instruct to make machine text sound more human raised Verb Diversity from 0.77 to 0.92 while making outputs easier to detect, that detection scores appear to track statistical complexity and yield a 76.3% false-positive rate on formal human writing, and, as a control, that event-based Latent Space detection had 87% of its event sequences changed by paraphrasing (Jaccard = 0.067) and 70% of extracted verbs altered by homoglyphs (Jaccard = 0.30), with a best domain AUC of 0.577.
Using the M4 dataset (N = 10,000) and controlled generations (N = 300), this work perturbs a RoBERTa-based AI-text detector at semantic, structural, and tokenizer levels and finds that asking Mistral-7B-Instruct to make machine text sound more human raised Verb Diversity from 0.77 to 0.92 while making outputs easier to detect, that detection scores appear to track statistical complexity and yield a 76.3% false-positive rate on formal human writing, and, as a control, that event-based Latent Space detection had 87% of its event sequences changed by paraphrasing (Jaccard = 0.067) and 70% of extracted verbs altered by homoglyphs (Jaccard = 0.30), with a best domain AUC of 0.577.