Controlled Rock–Paper–Scissors and n-gram study finds longer context does not improve strategy identification and second-order dependencies sharply degrade rule execution
Related research and updatesSynopsis
Using controlled two-player Rock–Paper–Scissors games (statistical versus Markov players, 100 to 1000 rounds of context) and a one-player stochastic n-gram continuation task (orders 1 to 8), the authors separate strategy identification, marginal distribution matching, and conditional rule execution, finding that Markov strategies are harder to identify than non-Markov ones, longer context does not improve and often degrades identification, correct recognition does not guarantee faithful simulation, and second-order dependencies markedly lower strict rule match, while neither providing transition rules nor longer prefixes removes high-order degradation and a prefix-only empirical n-gram estimator stays comparatively stable at high order.
Figure 3: Identification accuracy across seven model configurations, averaged across context lengths. Bold model labels and asterisks indicate significant Markov–Non-Markov differences ( p < .001 p<.001 ).
arXivInterpretation
Under a closed candidate pool, Markov (history-dependent) players are identified less accurately on average than non-Markov players, with the difference significant in six of seven model configurations and GPT-5 the sole exception; increasing observed rounds from 100 to 1000 does not improve identification and often degrades it. Prior game-based evaluations often stop at play performance or opponent identification; this work decomposes identification into closed-set recovery of a known generating process and compares Markov versus non-Markov strategy classes separately. Seven model configurations, with 50 statistical-vs-statistical, 50 Markov-as-P1, and 50 Markov-as-P2 pairs per context length; on the same trajectories a non-LLM maximum-likelihood baseline reaches 100% accuracy at every context length and its posterior mass is even more concentrated for Markov players, indicating the trajectories are recoverable.
Correct identification improves rule following, but recognition and simulation are not equivalent: wrong identities are not random, with rule overlap above the random baseline for several models even when the predicted identity is incorrect; on the correct-identity subset Markov MSE is more than 100% higher than non-Markov, versus about 8–9% higher overall. Marginal distribution matching and conditional rule execution are measured separately, showing that distributional similarity can conceal an incorrect generative mechanism. Among 282 incorrectly recovered strategies, 41.8% fall within a TV threshold of the target marginal distribution and 62.1% within a wider threshold; cumulative strict rule match curves show rule-following quality is largely set early in generation and then stabilizes.
Teacher forcing separates immediate rule application from long-horizon maintenance: DeepSeek Reasoner and GPT-5 reach 93.3% and 82.5% next-step accuracy under ground-truth histories (97.2% and 89.5% on the correct-ID subset), whereas DeepSeek Chat and GPT-5-mini reach only 52.5% and 54.2% (66.7% and 70.0%), indicating the former mainly struggle to sustain the rule while the latter also struggle with local rule application. Teacher forcing is used as a diagnostic to attribute free-running errors to two distinct stages rather than recording a single undifferentiated simulation failure. Table 1 reports accuracy for all cases and the correct-ID subset across the four core model configurations.
Structural stress tests point to dependency order rather than joint-state conditioning: adding joint-state information does not reduce rule following, second-order rules drive the main drop, and identity accuracy stays high in that setting; in the one-player stochastic n-gram task, CCLG drops sharply and WJS rises for most models at order 8, neither extending the prefix from 256 to 2048 nor providing transition rules consistently removes the degradation, and a prefix-only empirical n-gram baseline remains comparatively stable at high order. By removing player attribution and Rock–Paper–Scissors semantics, the bottleneck is localized to maintaining and applying higher-order conditional transition structure during generation. Experiment 3 uses deepseek-reasoner only, with 60 Markov-as-P1 and 60 Markov-as-P2 samples per rule family, 1000 generated rounds evaluated in 100-round windows; the one-player task uses CCLG and state-frequency-weighted JS divergence alongside a prefix-only n-gram MLE reference.
Perspective
The framework applies to controlled diagnostic settings where the candidate strategy set is known and the generating process can be specified exactly, for example agent evaluations that must distinguish frequency matching from conditional rule execution. For researchers and engineers using LLMs as behavioral simulators, multi-turn tool users, or opponent models, it supplies a reusable set of separated metrics: identification accuracy, marginal distribution error, strict and cumulative rule match, teacher-forced next-step accuracy, and CCLG and weighted JS divergence for high-order generation. The one-player n-gram experiment further shows that even after removing player attribution and Rock–Paper–Scissors semantics, the same metrics can probe the ability to maintain conditional transition structure.
The strategy space is a predefined closed set, so open-ended strategy discovery is not covered; trajectories are represented symbolically and outputs are evaluated through constrained formats and parsing rules, so the chosen prompts, representations, and decoding settings cannot be guaranteed optimal. The structural analysis in Experiment 3 uses deepseek-reasoner only, leaving open whether the joint-state and second-order findings hold for other models. In the one-player n-gram task, some model-setting combinations show lower valid completion rates under the rule-and-probs setting, so how much metric differences reflect rule-following failure versus output-format failure deserves further separation. In addition, model access dates and API alias resolution are time-specific, so version differences matter for reproduction.
