Public articles linked to the same research event.
arXiv The authors introduce Faynt, a family of 10M- and 75M-parameter Transformer policies for Super Smash Bros. Melee in which a single checkpoint controls all 26 characters; after reinforcement learning, the 10M model wins 240 of 244 same-character games (98.4%) against fourteen specialist and multi-character releases, holds a winning record against every release, and wins all 68 games against a privately supplied zero-delay Slippi-AI model; the pipeline covers pretraining on roughly 840,000 human replays, rank- and outcome-based curricula, 75M-to-10M distillation, and RL restricted to Fox mirror matches, with weights, both benchmark suites, and an automated tournament platform released.
The authors introduce Faynt, a family of 10M- and 75M-parameter Transformer policies for Super Smash Bros. Melee in which a single checkpoint controls all 26 characters; after reinforcement learning, the 10M model wins 240 of 244 same-character games (98.4%) against fourteen specialist and multi-character releases, holds a winning record against every release, and wins all 68 games against a privately supplied zero-delay Slippi-AI model; the pipeline covers pretraining on roughly 840,000 human replays, rank- and outcome-based curricula, 75M-to-10M distillation, and RL restricted to Fox mirror matches, with weights, both benchmark suites, and an automated tournament platform released.
The authors introduce Faynt, a family of 10M- and 75M-parameter Transformer policies for Super Smash Bros. Melee in which a single checkpoint controls all 26 characters; after reinforcement learning, the 10M model wins 240 of 244 same-character games (98.4%) against fourteen specialist and multi-character releases, holds a winning record against every release, and wins all 68 games against a privately supplied zero-delay Slippi-AI model; the pipeline covers pretraining on roughly 840,000 human replays, rank- and outcome-based curricula, 75M-to-10M distillation, and RL restricted to Fox mirror matches, with weights, both benchmark suites, and an automated tournament platform released.
The authors introduce Faynt, a family of 10M- and 75M-parameter Transformer policies for Super Smash Bros. Melee in which a single checkpoint controls all 26 characters; after reinforcement learning, the 10M model wins 240 of 244 same-character games (98.4%) against fourteen specialist and multi-character releases, holds a winning record against every release, and wins all 68 games against a privately supplied zero-delay Slippi-AI model; the pipeline covers pretraining on roughly 840,000 human replays, rank- and outcome-based curricula, 75M-to-10M distillation, and RL restricted to Fox mirror matches, with weights, both benchmark suites, and an automated tournament platform released.