Across nearly half a million names and 12 tokenizers, whether a name becomes a single token shifts concept accessibility in fellowship, hiring, clinical, and lending judgments
Synopsis
The study introduces NameTrace, measures direct lexical support for names across nearly half a million first names and 12 LLM-associated tokenizers, and, on atomic versus short-fragmented names matched within the same race/ethnicity–gender strata, finds systematically higher task-aligned concept accessibility for atomic names across fellowship, hiring, clinical assessment, and lending; the differences persist in all eight strata, transfer to unseen names, and hidden-state interventions along the measured task directions shift later constrained choices.
Interpretation
Direct single-token access to first names is highly selective and patterned by demographic metadata: among 7,469 higher-frequency, high-confidence names, any-tokenizer atomic access is 49.8% for male-associated names versus 25.7% for female-associated names, and ranges from 17.6% for NH Black-associated names to 47.2% for NH White-associated names, with Asian/PI-associated names at 46.4%. Prior work on name bias largely examined behavior or small-scale vocabulary observations; this work treats lexical comparability as its own variable, quantifies it across nearly half a million names and 12 tokenizers, and reports adjusted odds ratios controlling for frequency and length. Logistic regression shows a one-standard-deviation increase in log frequency corresponds to 2.94 times the odds of atomic access and the same increase in length to 0.51 times; male-associated names have 3.36 times the adjusted odds of female-associated names, Hispanic- and NH Black-associated names 0.47 and 0.42 times the odds of NH White-associated names, and Asian/PI-associated names 1.55 times; intersectional access ranges from 12.1% for NH Black female-associated names to 64.8% for NH White male-associated names.
Among atomic versus short-fragmented names matched within race/ethnicity–gender strata, atomic names show higher task-aligned concept accessibility on all four task axes: weighted gaps of 0.131 for fellowship/promise, 0.072 for hiring/competence, 0.059 for clinical concern, and 0.051 for lending/trustworthiness. Where name-fairness evaluations usually read only final responses, this measures at an intermediate layer using the model's own probability distribution over a task-specific adjective axis with continuous task-aligned weights, so the clinical-concern axis stays positive even when generic sentiment polarity reverses, as with worried and ill. 200 matched pairs (400 names) split into 100 development and 100 unseen evaluation pairs, matched on frequency, character length, demographic-association strength, metadata confidence, and weak orthographic cues; all four axes are evaluated under strong, borderline, and weak evidence conditions with pooled gaps and 95% confidence intervals.
The support-linked differences persist across all eight race/ethnicity–gender strata and extend across model families, while magnitude and direction vary by architecture, depth, and training stage. This shows the pooled result is not driven solely by between-group composition, since comparisons are made within strata, and it treats architecture and training stage as part of the finding rather than reporting a single average effect. Model-by-stratum gaps are positive in 22/24 cells for fellowship, 24/24 for hiring, 23/24 for clinical assessment, and 15/24 for lending; Qwen is positive on all four axes, Llama is positive throughout at smaller scale, and Ministral is positive for hiring and clinical assessment, near zero for fellowship, and negative for lending; in the eight-model extension 23/32 model–task means and 20/32 confidence intervals are positive; base/post-training comparisons show Qwen and Gemma redirect the accessibility profile while Llama largely preserves it.
A support prior estimated from development names predicts gaps on unseen names, and editing name representations along the measured task direction shifts a later constrained two-choice decision. The first shows the support-linked pattern has systematic cross-name structure rather than being name-specific, and the second shows the measured internal direction is not merely readable but has downstream leverage, linking token-level differences to later choices. Subtracting the development prior reduces the pooled held-out gap from 0.131 to 0.004 for fellowship (96.6% reduction), 0.072 to 0.008 for hiring (88.3%), 0.059 to 0.003 for clinical assessment (88.2%), and 0.051 to 0.014 for lending (72.9%); across 36 model–task–evidence cells the prior correlates with held-out gaps at Pearson 0.94 and Spearman 0.93 with 94.4% sign agreement; intervention forward-minus-reverse contrasts are 0.153 for Qwen, 0.155 for Llama, and 0.094 for Ministral, with all 200 Qwen and Llama pairs and 195 of 200 Ministral pairs moving in the expected direction, and Qwen and Llama exceeding unrelated-axis, polarity-shuffled, and random-subspace controls.
Perspective
The work is intended for model evaluation and development: tokenizer designers can see which name surfaces receive direct lexical support, and evaluators can inspect each target name under the relevant tokenizer and treat name-surface support as a separate evaluation variable alongside demographic metadata, frequency, and length, balancing or stratifying by support and running sensitivity analyses when support is uneven. The findings apply to sufficiently frequent, single-word ASCII first names with reliable metadata associations, and they concern how name surfaces are processed as lexical inputs rather than the characteristics of individuals or groups; the task axes can be extended to new application concepts once aligned and opposed adjective poles are defined.
Name-specific pretraining exposure is not directly observed, so the authors frame lexical support as a measurable predictor among names matched on major observed properties rather than as an isolated causal treatment; intermediate accessibility can be preserved, attenuated, or redirected by later computation, with sign reversals for hiring and clinical assessment at the output boundary, so how internal accessibility relates to final open-ended behavior remains an open question; effect magnitude, task coverage, and layer localization vary across architectures, with Ministral negative on lending and, in the eight-model extension, OLMo negative on lending, Gemma-3-4B slightly negative on hiring, and Aya and Ministral negative on clinical assessment; name metadata are aggregate associations of name surfaces rather than identity labels for individuals, and other naming conventions, languages, scripts, and cultural contexts are natural extensions; and if a reader sees only a fast parse without figures and appendices, some stratum-level and layer-localization values would need to be checked against the original text.
