Skip to main content
Back to timeline
Cohere LabsSource publication:

Tiny Aya L2-Thinker lifts in-language reasoning from 12.8% to over 93% via three-pillar data mixing, with accuracy down at most two to three points on five of six benchmarks

Related research and updates

Synopsis

Building on the 32K-context Tiny Aya base model, the work trains Tiny Aya L2-Thinker with a three-pillar mix of English reasoning data, automatically translated multilingual reasoning data of about 5,000 examples per language across roughly 44 languages, and multilingual non-reasoning data, raising the in-language reasoning rate from 12.8% to above 93% across 60 languages and six benchmarks, with accuracy dropping at most two to three points on five benchmarks, a slight gain on the open-ended writing benchmark, and a more noticeable drop on the competition-level math benchmark PolyMath.

AI-generated editorial illustration: Multilingual Bridges: How Data Mixing Unlocks In-Language Reasoning

Interpretation

Trained on English reasoning data alone, the model reasons in the user's language only 12.8% of the time even when overtly instructed to do so; adding about 5,000 automatically translated multilingual reasoning examples per language across roughly 44 languages raises that rate to 86.1% while accuracy also improves. Prior work reported that in-language reasoning often costs accuracy, or avoided that drop only through complex, expensive training; here a comparatively tiny amount of multilingual reasoning data moves language choice and accuracy in the same direction. The comparison is against Tiny Aya En-Thinker, a twin trained on the same base model and underlying data but reasoning in English, with the in-language reasoning rate determined by running a language identification model over each reasoning trace; data sizes and rate changes are stated in the text.

Multilingual non-reasoning data, consisting of questions and answers in the target language without reasoning traces, lets the model carry in-language reasoning over to languages in which it never saw reasoning examples. It separates 'can the model solve the problem' from 'which language does it reason in' as two things learnable from different data sources, and notes that non-reasoning data is more available because it was used in prior multilingual instruction models. The text reports training support for 45 languages and states there is evidence the ability carries over to languages never shown reasoning examples; the detailed transfer evaluation points to the paper.

Across 60 languages and six benchmarks (MGSM, PolyMath, Macaron-MCQ, GlobalPIQA, Marco-Bench-MIF, MIST-OEG), the in-language reasoning rate rises from near zero to above 93%, accuracy drops at most two to three points on five benchmarks, the open-ended writing benchmark MIST-OEG rises slightly, and on English prompts the model still reasons in English and matches the English reasoner on most benchmarks. Prior evaluations concentrated on math and science problems and covered only a handful of languages; this work extends evaluation to cultural and commonsense reasoning, instruction following, and open-ended writing, and examines four tiers of language resourcedness. Direct comparison against the English-reasoning twin trained on the same base model and data, with the in-language reasoning rate judged trace by trace by a language identification model; PolyMath, a competition-level math benchmark, is the exception where accuracy falls more noticeably.

Across four tiers of language resourcedness, Tiny Aya L2-Thinker is the only model that stays strong on all three axes as resource availability declines: its in-language reasoning rate sits near the 99% ceiling on the top two tiers and drops only to 94.3% and 94.5% on the lowest two, averaging under 5,000 thinking tokens, whereas Magistral-Small-24B collapses from 72% to 5% and M-Thinker-7B degrades on every axis with traces roughly doubling in length. It places efficiency, measured in reasoning tokens, and repetition alongside in-language reasoning rate and accuracy rather than reporting accuracy alone. The four resource tiers are grouped by how much web text exists for each language; repetition is measured as how often short sequences of text repeat within a trace and compared against reasoning length in a scatterplot, showing the two rise together.

Perspective

The result applies to multilingual reasoning models trained with data mixing as the central lever on the 32K-context Tiny Aya base model, with training covering 45 languages, evaluation covering 60 languages, and four resource tiers grouped by available web text; developers and researchers who want non-English-speaking users to read and verify reasoning traces can directly reuse the model and multilingual reasoning data released on Hugging Face. Because the method uses data composition as its lever, the conclusions apply most directly to settings with comparable base models and comparable data availability; competition-level math is the exception the text explicitly flags, which the authors attribute to the absence of a reinforcement learning stage and list as a future direction focused on mathematics.

The detailed analysis of each data pillar and the determination of the final data mix point to the paper, so what can be confirmed here is the composition of the three pillars and the overall effect rather than specific mixing proportions. Accuracy falls more noticeably on PolyMath, which the authors explain speculatively by the absence of a reinforcement learning stage and note may also reflect artifacts introduced when translating training traces; that explanation remains to be tested. The co-movement of repetition and reasoning length rests on a scatterplot comparison across the models tested, and the authors state that efficient reasoning is not solved yet. Finally, a reasoning trace being readable by users does not by itself mean it meaningfully influenced the final answer, and the authors list making traces more meaningful, efficient, and interpretable as further work.

Sources