A Taxonomy of Programming Languages for Code Generation
Synopsis
This work presents the first reproducible resource-tier taxonomy for programming languages: the authors merge and deduplicate seven major code corpora, count tokens for 646 programming languages with the StarCoder tokenizer, and split them into four tiers (High, Medium, Low, Scarce) using 100B, 10B, and 1B thresholds, finding that the 12 High-tier languages (1.9% of languages, e.g., Python, JavaScript, Java) supply 74.6% of all 7.63T tokens while the 463 Scarce-tier languages (71.7%) contribute just 1.0%, with Gini coefficients, coefficients of variation, Lorenz curves, and ECDF survival curves characterizing imbalance both across and within tiers.
Figure 1: Resource imbalance across four tiers. (a) : share of languages; (b) : share of tokens. Only 1.9% of languages (High) supply 74.6% of all tokens, while 71.7% of languages (Scarce) contribute just 1.0%.
arXivInterpretation
Introduces the first reproducible resource-tier classification for programming languages, placing 646 languages into four tiers by total token count. Natural-language NLP already had a resource taxonomy from Joshi et al., while the programming-language ecosystem had only anecdotal observations; this work transfers that idea to code with explicit thresholds. Built on 7.63T tokens from seven merged and deduplicated corpora, with published thresholds and pipeline.
Quantifies an extreme long tail: 12 High-tier languages hold 74.6% of tokens, while 463 Scarce-tier languages hold only 1.0%. Turns the 'Python-centric' observation from anecdote into a measurable tier-level phenomenon, with the top-10 languages holding 71.6% and the bottom-10 about 0.00003%. Supported by full token counts and a per-language appendix listing tokens and file counts, showing very large magnitude gaps.
Shows imbalance also exists within tiers: Scarce tier has Gini 0.66 and CV 1.37, High tier has Gini 0.43, and means exceed medians in every tier. Goes beyond between-tier gaps to characterize within-tier distribution using inequality, dispersion, and right-skew statistics. Gini, coefficient of variation, mean-vs-median comparison, Lorenz curves, and ECDF survival curves reinforce one another.
Releases token counts and tier assignments to support dataset curation and tier-aware evaluation. Enables stratified scoring in benchmarks instead of overall means that mask per-tier weaknesses, and lets others swap tokenizers or adjust thresholds. The authors state they release raw counts and tier assignments, publishing only aggregate statistics rather than raw code.
Perspective
The taxonomy targets settings where code LLMs are trained and evaluated on web-derived corpora, and applies to dataset curation, stratified benchmark scoring, and crawling-priority decisions; it is especially relevant to researchers and engineering teams focused on low-resource programming languages, code retrieval, or code security. The authors note that their choices of seven permissively licensed corpora, a single StarCoder tokenizer, extension-based language detection, exact-hash deduplication, and heuristic thresholds are intentional for transparency and reproducibility, and they release raw counts so others can swap tokenizers or adjust thresholds.
Readers may still watch how different tokenizers or finer-grained deduplication would shift absolute counts; how extension-based language detection handles polyglot and template-like files; whether the heuristic thresholds need resetting as corpora keep growing; and how tier membership maps to actual model generation behavior, a link the paper raises with the 'Python default' example but does not test with generation experiments.
