Skip to main content
Back to timeline
arXivSource publication:

Decision Titan adds test-time training layers to a Decision Transformer and learns long-term dependencies 20x longer than its context window in X-Maze

Synopsis

The work augments a Decision Transformer with Test-Time Training (TTT) layers, forming Decision Titan, and analyzes its performance and memory mechanism in the X-Maze environment built to test sequential memory: the model learns long-term dependencies with ranges 20x longer than the context window and generalizes to lengths 1.7x the training data, but temporal generalization depends on the time embeddings used and the ability to learn long-term dependencies depends on how the relevant information is encoded.

Source-provided article image: Decision Titan: Test-Time Training for Long-Term Memory in Offline Reinforcement Learning
Figure 1 ·

Figure 1 : The Decision Titan architecture. A trajectory of returns-to-go, states, and actions is projected into a shared embedding space before having positional embeddings added. This sequence is then chunked into subsequences of k k timesteps and n n persistent memory tokens are appended to the front of each chunk. Following this, each chunk is then passed through L L Titans MAL blocks, consisting of a TTT layer that is updated every b b timesteps, followed by self-attention with a window size of 3 ​ k + n 3k+n tokens.

arXiv

Interpretation

The authors propose Decision Titan, a Decision Transformer augmented with TTT layers, bringing the TTT framework from natural language processing into offline reinforcement learning, where episodic memory is stored in network parameters rather than vector hidden states. According to the authors, TTT had not yet been applied to reinforcement learning, and no study had analyzed how this memory practically functions; the work addresses both the application and the mechanism analysis. Evidence comes from performance and property analysis in X-Maze, an extension of T-Maze designed to test sequential memory, plus visualization of gate values over time to see how the memory mechanism learns; the abstract gives no specific numbers, sample sizes, or control settings.

Decision Titan can learn long-term dependencies with ranges 20x longer than the context window and generalizes to lengths 1.7x the training data. This gives concrete magnitudes for the temporal span and length extrapolation that TTT memory can reach in an offline RL decision-making task. Measured in the X-Maze environment and reported in the abstract as multipliers; the specific experimental configuration and statistical details are not expanded in the abstract.

Temporal generalization depends on the time embeddings used, and the ability to learn long-term dependencies depends on how the relevant information is encoded. It attributes memory effectiveness to two designable factors, time embeddings and information encoding, rather than to the TTT layers alone. Based on analysis of model properties and visualization of gate values over time; the abstract does not give comparative numbers for each factor.

Perspective

The result targets sequential decision-making in offline reinforcement learning, especially memory-type tasks that require acting correctly across a long history, for which X-Maze is the designed T-Maze extension. For researchers and engineers working on long-term memory architectures, offline RL, or sequence modeling, it suggests storing episodic memory in network parameters updated by gradient descent at both train and test time, and designing temporal spans beyond the context window accordingly. It applies under the premise that task-relevant information can be presented in a form this framework can encode and that the choice of time embeddings matches the task's temporal structure.

The abstract reports only multiplier-form conclusions and does not state the scale of X-Maze, the amount of training data, baseline comparisons, or random seeds, so the conditions under which the 20x and 1.7x magnitudes hold remain open. The separate magnitudes of the effects of time embeddings and information encoding, and the memory-mechanism details revealed by gate-value visualization, also need confirmation in the original text. In addition, conclusions are currently limited to offline reinforcement learning and the X-Maze memory-type environment, and extension to other tasks and online settings remains to be tested.

Sources