Faynt controls all 26 Melee characters with one checkpoint, and its 10M model wins 240 of 244 same-character games
Related research and updatesSynopsis
The authors introduce Faynt, a family of 10M- and 75M-parameter Transformer policies for Super Smash Bros. Melee in which a single checkpoint controls all 26 characters; after reinforcement learning, the 10M model wins 240 of 244 same-character games (98.4%) against fourteen specialist and multi-character releases, holds a winning record against every release, and wins all 68 games against a privately supplied zero-delay Slippi-AI model; the pipeline covers pretraining on roughly 840,000 human replays, rank- and outcome-based curricula, 75M-to-10M distillation, and RL restricted to Fox mirror matches, with weights, both benchmark suites, and an automated tournament platform released.
Figure 1: Faynt policy architecture. Dashed connections supply current controller categories directly to the output heads. The main-stick head also conditions on the next joint-controller category, using the recorded target during imitation training and a sample during play. AttnRes denotes Full Attention Residuals.
arXivInterpretation
A single checkpoint covers all 26 characters and wins 240 of 244 same-character games (98.4%), with a winning record against each of fourteen specialist and multi-character releases. Prior releases in this comparison are typically trained per character or per roster, whereas here one checkpoint covers every character and posts a winning record against each release on its supported roster. The abstract reports same-character game counts (240/244) and the win record against each release; opponents retain 21- or 24-frame action delays while Faynt adds none, and the authors state they have not isolated the effect of this difference.
On the initial 152-game benchmark, the supervised 10M wins 69.7% of games versus 45.4% for the pretrained 75M, despite higher overall held-out controller-prediction loss. The result places model scale, supervised loss, and match win rate side by side: the larger model with lower prediction loss does not win more. A win-rate comparison over the 152-game benchmark, with the note that the weighted validation loss used for supervised checkpoint selection agrees with the win-rate ordering of all four pretrained and supervised policies.
After supervised post-training, both models take less damage per minute, build larger early leads, and win more often after losing the first life. It extends the effect of post-training beyond win rate to in-game process measures such as damage taken, early leads, and comebacks from behind. Abstract-level behavioral descriptions without specific values or statistical tests for each measure.
Optimized inference on recorded game states averages 5.2 ms per decision for the 10M and 8.7 ms for the 75M on an NVIDIA T4, excluding emulator execution and communication. It reports measured inference latency for both policy sizes, letting deployment cost be weighed alongside model scale. Average latency measured on recorded game states, explicitly excluding emulator execution and communication overhead.
Perspective
This work targets competitive Super Smash Bros. Melee and is meant for researchers and practitioners who want one checkpoint covering all 26 characters and same-character comparisons against existing releases on their supported rosters. The released weights, both benchmark suites, and the automated tournament platform let follow-up work reproduce the win-rate comparisons under the same conditions, test how distillation and curriculum design transfer, and factor inference latency into deployment trade-offs.
Opponents retain 21- or 24-frame action delays while Faynt adds none, and the authors state they have not isolated the effect of this difference, so how much of the win-rate gap comes from the delay condition remains open. The supervised 10M wins more than the pretrained 75M while showing higher held-out controller-prediction loss, a pattern consistent with the weighted validation-loss ordering but one whose scope needs more settings to confirm. Damage, early-lead, and comeback measures are described only directionally, without specific values. This reading covers the abstract only; figures, hyperparameter details, and statistics in the full text were not included and should be checked against the original.
