Adapting language-dominance probing from cross-lingual to 26 Arabic varieties finds no MSA internal dominance, with dialect representations densely overlapping
Synopsis
Adapting the language-dominance probing framework of Shani and Basirat (2025) from cross-lingual to 25 Arabic dialects plus MSA, the study analyzes hidden representations layer by layer in BLOOM, mGPT and three further models and finds markedly lower normalized mutual information than in the cross-lingual case, a densely overlapping dialect space, and layer-wise dominance scores near the uniform baseline, giving no evidence that MSA or any dialect acts as a dominant internal reference point.
Figure 1: Layer-wise NMI values for cross-lingual vs. cross-dialectal representations for mGPT (left) and BLOOM (right). Dialectal NMI is lower and remains low over a wider middle-layer range. Cross-lingual representations were computed using Parallel Universal Dependencies (PUD) treebanks ( Zeman et al., 2017 ) (following Shani and Basirat (2025) ), cross-dialectal representations were computed using the MADAR corpus ( Bouamor et al., 2018 ) .
arXivInterpretation
Across 26 Arabic varieties, MSA shows no internal representational dominance: layer-wise dominance scores stay close to the uniform baseline of 1/|V|, and the proportion of tokens aligning more strongly with a competing dialect than with their own stays nearly flat at around 30% across layers, rather than forming the middle-layer spike seen for distinct languages. Prior discussion of MSA-biased generation often reads output preference directly as evidence of internal dominance; this work uses a symmetric probe that does not presuppose which variety dominates, separating output preference from internal representational geometry. Built on aligned-token posterior comparisons over the parallel MADAR corpus, pooling every (token, competitor) pair at each layer; the same probe produces the middle-layer spike in the cross-lingual replication, establishing that it detects directional alignment when the effect is strong enough to appear.
Arabic dialects are densely overlapping in hidden space but not erased: average NMI is around 0.1 for mGPT and 0.2 for BLOOM, with the low-NMI region extending 1–3 layers further at both ends of the middle-layer interval, while confusion matrices retain clear diagonals and a supervised linear probe reaches 0.15–0.25 accuracy, roughly 4–6.5 times the 3.84% chance level. Extends language-dominance analysis to the fine-grained setting of a closely related dialect continuum and combines unsupervised clustering with a supervised probe, showing that low separability does not mean the information is absent. 400 parallel sentences per variety for BLOOM and mGPT (four times the per-language sample size of the original framework), with POS filtering, sentence-prefix removal and PCA sensitivity analyses reported in the appendix preserving the main trends.
Where dialect separability changes with depth is architecture-dependent: Meta-Llama-3.1-8B-Instruct maintains relatively stable dialect separability, Qwen2.5-7B shows an early sharp drop followed by a modest recovery toward the final layer, and ArabianGPT-08B-V2 keeps NMI high until later layers before a sharp collapse. The original framework discussed middle-layer collapse mainly for multilingual models; this work contrasts multilingual pretraining with Arabic-native pretraining, suggesting Arabic-focused training does not necessarily make dialects uniformly more separable and may instead shift dialectal compression toward later layers. Three additional models beyond BLOOM and mGPT, with 100 parallel samples per variety; the authors state the model set is designed to contrast multilingual and Arabic-native pretraining rather than to exhaustively benchmark available Arabic LLMs.
MSA-biased generation is not explained by internal MSA dominance: under the applied probing framework MSA behaves as one variety among many rather than as an implicit internal translator, and the authors suggest the bias may instead originate elsewhere in the generation pipeline, such as the disproportionate share of MSA in pretraining and instruction-tuning data, MSA's comparatively standardized orthography versus unstandardized dialectal spelling, or decoding-time preferences shaped by RLHF and higher-frequency token paths. Provides a counterexample-style test of the common inference from output preference to internal routing, and explicitly frames generation bias and representational geometry as two questions requiring separate analysis. The conclusion rests on layer-wise dominance statistics, token-level likelihood-ratio distributions and ablation experiments; the authors note that disentangling the alternative explanations is beyond the scope of the work and that findings should be read as exploratory rather than definitive.
Perspective
The results apply to a probing setup that takes the parallel MADAR travel corpus as input and extracts hidden states layer by layer from BLOOM, mGPT, Qwen2.5-7B, Meta-Llama-3.1-8B-Instruct and ArabianGPT-08B-V2; the conclusions concern the internal representational geometry of these models on this corpus, not a universal statement about all Arabic models or all text domains. For researchers working on dialect NLP, representation probing and multilingual model analysis, this offers a reusable contrast: when generation bias is observed, output preference and internal representational structure should be tested separately rather than treating the former as evidence of the latter. The authors also note that positioning Arabic varieties more precisely on the dialect–language continuum would require computing the same separability curves for closely related language families.
Several open questions remain for a careful reader. The corpus is restricted to the MADAR travel domain, and 400 sentences per variety (BLOOM, mGPT) and 100 sentences (the other models) are moderate samples; the authors note that some results may partly reflect writer-specific idiosyncrasies rather than dialect-level patterns, and that the rule-based POS tagging may introduce inaccuracies. The dominance probe is symmetric and does not presuppose a dominant variety, so distinguishing 'no dominance detected' from 'no dominance present' depends on the probe's sensitivity, which is only indirectly validated in the cross-lingual replication. In addition, the other candidate sources of generation bias (the share of MSA in pretraining and instruction-tuning data, orthographic standardization, RLHF and decoding preferences) are not separately tested here, and the authors explicitly list disentangling them as beyond the scope of the work.
