Long-Lived Characters, Local Inference: Incremental Memory Maintenance for Game NPCs
Synopsis
This work implements a training-free incremental memory-maintenance runtime for a quantized Qwen hybrid recurrent–attention model that removes superseded attention KV entries, computes replacement records at the true sequence tail, and preserves the continuing recurrent state and unchanged KV; experiments show that independent block composition weakens query-conditioned memory selection, that true-tail updates preserve key current-state and historical bindings across eight scripted maintenance rounds, that slot-preserving alternatives repeat a double-subtraction error, and that attention-distribution proximity alone does not explain these semantic differences.
Figure 1: Rule-governed play makes memory maintenance consequential. Open-ended dialogue expresses intentions; compilation connects them to game-supported activities, while game rules govern authorization and consequences. Experiences revise character records. On a shared local device, repeatedly rebuilding long contexts competes with foreground dialogue. The target is to encode revisions while reusing surviving state. This conceptual illustration motivates the workload; it is not a population-throughput measurement or a claim that the complete envisioned game has been implemented.
arXivInterpretation
It formulates and implements a persistent-memory maintenance runtime for fully local game characters that separates the lifetime of explicit KV records from the continuing recurrent state without changing model weights. Whereas prefix caching and modular reuse focus on reusing shared prefixes or recomputing chunks, this work treats logical memory as editable while its computational history keeps moving forward, deleting superseded KV and appending replacements at the true tail. Implemented on a quantized Qwen3.6-27B hybrid model (48 recurrent layers interleaved with 16 full-attention layers) through a modified local llama.cpp runtime, with multi-update maintenance traces and attention diagnostics.
Independent block composition weakens query-conditioned memory selection without a uniform chunk-initial attention collapse. At the explicit focus-ID probe, five requested memory groups receive 57.4% of memory attention under full refill but 21.5% under independent composition, and the effective number of attended groups rises from 18.3 to 33.2 out of 34, while the first eight tokens of each group do not show the pervasive chunk-start pattern analyzed by EPIC. Paired probes after the same focus-ID instruction, with the five requested groups comprising about 12.16% of memory tokens; these are descriptive attention diagnostics, not generated-answer accuracy scores.
True-tail updates preserve several current-state and historical bindings across eight lossy maintenance rounds, whereas slot-preserving alternatives repeat a double-subtraction error. On the frozen herb-compilation task, dense refill and fresh gapped prefill both compile two herbs, maintaining original slots compiles one, three true-tail reconstructions all compile two, and three tail-compute-then-rotate-back reconstructions all compile one. Three reconstructions are repeatability checks rather than independent scenarios; the final live cache holds 11,894 tokens, with next-position metadata of 261,955 for the original-slot layout and 479,547 for the true-tail layout, and holes consume no placeholder KV.
Attention-distribution proximity does not certify semantic fidelity, and a post-maintenance binding error can be recovered within the retained state by changing task organization. Token-distribution total variation distances from dense to fresh-gapped prefill are 0.095, 0.136, and 0.144, while distances from fresh-gapped prefill to maintained-gapped state are smaller (0.061, 0.079, 0.076), yet the latter is accompanied by mistakes in the rescued-person reference and herb quantity; in the long-horizon case, an understanding-first task yields the correct statement that he carefully picked up and brushed the pages clean, without importing the red cord. Three-way fixed-text diagnostics and frozen task probes with one response per arm; the author explicitly frames this as a behavioral intervention rather than a measured attention-map shift, establishing a recoverable case rather than a general repair rate.
Perspective
The result targets local game characters running on player-owned hardware with rule-governed dialogue and resource accounting, in settings that repeatedly revise memory while reusing unchanged context; it makes change-proportional maintenance an operational option and provides a starting point for later evaluation with multiple characters, longer histories, and adversarial revisions.
A careful reader would still watch how preparation and update budgets compete with foreground dialogue across multiple characters, whether maintained states remain stable under repeated reversals, contradictory reports, and longer histories, how quantization and numerical paths affect conclusions, and the absence of captured attention maps for the true-tail and key-relocation behavior runs, which bears on attributing the semantic differences.
