Public articles linked to the same research event.
Cohere Labs Building on the 32K-context Tiny Aya base model, the work trains Tiny Aya L2-Thinker with a three-pillar mix of English reasoning data, automatically translated multilingual reasoning data of about 5,000 examples per language across roughly 44 languages, and multilingual non-reasoning data, raising the in-language reasoning rate from 12.8% to above 93% across 60 languages and six benchmarks, with accuracy dropping at most two to three points on five benchmarks, a slight gain on the open-ended writing benchmark, and a more noticeable drop on the competition-level math benchmark PolyMath.
Building on the 32K-context Tiny Aya base model, the work trains Tiny Aya L2-Thinker with a three-pillar mix of English reasoning data, automatically translated multilingual reasoning data of about 5,000 examples per language across roughly 44 languages, and multilingual non-reasoning data, raising the in-language reasoning rate from 12.8% to above 93% across 60 languages and six benchmarks, with accuracy dropping at most two to three points on five benchmarks, a slight gain on the open-ended writing benchmark, and a more noticeable drop on the competition-level math benchmark PolyMath.
Building on the 32K-context Tiny Aya base model, the work trains Tiny Aya L2-Thinker with a three-pillar mix of English reasoning data, automatically translated multilingual reasoning data of about 5,000 examples per language across roughly 44 languages, and multilingual non-reasoning data, raising the in-language reasoning rate from 12.8% to above 93% across 60 languages and six benchmarks, with accuracy dropping at most two to three points on five benchmarks, a slight gain on the open-ended writing benchmark, and a more noticeable drop on the competition-level math benchmark PolyMath.
Building on the 32K-context Tiny Aya base model, the work trains Tiny Aya L2-Thinker with a three-pillar mix of English reasoning data, automatically translated multilingual reasoning data of about 5,000 examples per language across roughly 44 languages, and multilingual non-reasoning data, raising the in-language reasoning rate from 12.8% to above 93% across 60 languages and six benchmarks, with accuracy dropping at most two to three points on five benchmarks, a slight gain on the open-ended writing benchmark, and a more noticeable drop on the competition-level math benchmark PolyMath.