Skip to main content
Back to timeline
arXivSource publication:

Auditing 1,813 AmE–BrE variant pairs across the pipeline: six pretraining corpora, 21 post-training datasets and ten checkpoints all lean American English

Synopsis

The study builds a curated resource of 1,813 matched American English–British English variant pairs and introduces DiAlign, a training-free method for estimating regional alignment, then uses them to audit six pretraining corpora, 21 post-training datasets, nine tokenizers, ten model checkpoints and generations under two prompting conditions, finding that American English is systematically favored across data exposure, representation and generation, and that British-English prompting only partially shifts this default.

AI-generated editorial illustration: How Does "English (US)" Become the Default? Triangulating Structural Bias Towards American English Across the LLM Pipeline

Interpretation

The work treats "why English (US) becomes the default" as structural bias spanning the LLM development pipeline, using British English as a controlled reference and tracing the same directional asymmetry through data exposure, representation and generation. Earlier work on English varieties concentrated on single tasks or single stages such as POS tagging, sentiment analysis or retrieval; this study carries one matched variant set from corpus frequencies through tokenization, prediction cost and generation, producing a cross-stage evidence chain. The authors describe it as the first rigorous pipeline-wide study of structural bias across major phases of LLM development; evidence spans six pretraining corpora, 21 post-training datasets, nine tokenizers, ten checkpoints and generation experiments on two QA datasets.

Data exposure: all six pretraining corpora and all 21 post-training datasets favor American English, with orthographic contrasts stronger than vocabulary contrasts. It adds within-English variation to the data-auditing agenda, which has largely tracked language composition, domain and quality. On the pretraining side, AmE orthographic shares often exceed 70% of observed pair frequencies and Wilcoxon signed-rank tests show all six corpora differ significantly from parity; on the post-training side, 11,574,254 samples yield 28,860,119 explicit variant occurrences of which 76.17% are AmE, ranging from 65.29% in Databricks Dolly 15k to 89.72% in WizardLM Evol-Instruct v2.

Representation: AmE variants are generally encoded more compactly by tokenizers, and under a strict equal-sentence-token-count control all ten evaluated checkpoints assign higher per-token prediction cost to the BrE counterfactual. It examines tokenization efficiency, tokenizer provenance and model-assigned probability together, showing that compact segmentation and model preference are separable phenomena. Vocabulary fertility gaps reach roughly 16%–19%; StableLM-2's tokenizer shows 100% token-boundary identity with GPT-4 on both variant sets, with all 100,256 shared base tokens retaining identical rank/id assignments; the prediction-cost analysis contains 85,150 paired-loss records, every paired-bootstrap 95% confidence interval excludes zero, and the loss gap is larger in the post-trained checkpoint of all five model families.

Generation: AmE remains the dominant default under neutral English prompting, and British-English prompting shifts outputs toward BrE without consistently eliminating the AmE default, with the size of the shift varying by model and register. It links upstream data and representation asymmetries to realized output behavior and shows that output-level prompting does not alter the upstream asymmetries already established. Across Natural Questions and ELI5 and ten models, a majority of outputs under neutral English prompting are classified as AmE; for example Llama-3.3 falls to about 30% AmE-classified outputs on ELI5 under British-English prompting; the association between response length and AmE alignment is very small overall.

Perspective

The framework is aimed at corpus builders, tokenizer designers and model evaluators who want to audit or design regional language behavior, and it applies under the setting of paired reference distributions with reliable provenance; the authors state that DiAlign can be extended to other English varieties and beyond English when sufficiently large and comparable reference corpora can be constructed for each variety. British English serves as the controlled reference so that one contrast set can be carried from corpus statistics through tokenization, prediction cost and generation, which is what supports cross-stage comparison.

The authors state that consistency across stages establishes a structural pattern, while causal decomposition requires controlled interventions such as retraining with rebalanced data, tokenizer replacement or matched post-training interventions. The reference variety is currently limited to British English, and broader English and non-English extensions depend on constructing suitable reference corpora. Within the tokenizer panel, Llama-3.3 and StableLM-2 inherit or extend US-origin tokenizer artifacts, and the authors report them separately from independent tokenizers. The prediction-cost experiments are limited to five model families with matched public base and post-trained checkpoints, while DeepSeek-V3 and Velvet-2B enter only the tokenizer and provenance analyses. Generation experiments use two open-domain QA registers under a fixed 50-word target, leaving long-form dialogue, specialized professional domains, mixed-variety inputs and creative generation as open questions.

Sources