Making KDA's gate signed lets a single CKDA layer track finite groups and extrapolate periodic waveforms, with 1.3B downstream accuracy on par with KDA
Synopsis
The work introduces Complex KDA (CKDA): extending Kimi Delta Attention by allowing signed gate entries and an extended delta-rule coefficient range so that a single diagonal-plus-rank-one transition can realize a 2D rotation; the authors prove every orthogonal diagonal-plus-rank-one matrix is exactly a CKDA transition, show one CKDA layer tracks every finite group isomorphic to a subgroup of O(2), save one layer relative to comparable diagonal-plus-rank-one linear RNNs, and report experiments on group-word problems, periodic audio continuation, and language modeling.
Interpretation
CKDA realizes a 2D rotation within a single diagonal-plus-rank-one transition: the channel-wise gate supplies one coordinate reflection that composes with the delta-rule Householder reflection, yielding complex-conjugate eigenvalues. Previously DeltaProduct needed two delta-rule transitions per token to model a 2D rotation, raising rank and update cost; standard nonnegative KDA gates keep the spectrum real, and GDN's scalar gate commutes with every matrix and cannot break the symmetry. The paper derives the mechanism: with a scalar gate the transition is a product of two symmetric matrices that commute, so the spectrum is real; with a channel-wise gate the factors need not commute, yet strictly positive gates still make the transition similar to a symmetric matrix, so only signed gates can produce non-real spectra. The 2D expansion gives the negative-discriminant condition, and at the endpoint the transition is a composition of two reflections, i.e. a rotation.
The authors characterize the spectrum of CKDA transitions and prove that every orthogonal diagonal-plus-rank-one matrix is exactly a CKDA transition (signed-Householder form). This elevates CKDA from one parameterization to a complete characterization of the orthogonal DPR1 family, and shows that replacing KDA's structured rank-one term with a non-symmetric one (as in RWKV-7) adds no orthogonal transitions. Theorem 1 gives the signed-Householder form for orthogonal DPR1 matrices; Theorem 8 shows a CKDA matrix has at most one non-real conjugate pair, requiring both β>1 and at least one negative gate entry; Theorem 9 extends the bound to any non-expansive rank-one DPLR transition on the unit circle.
State-tracking expressivity improves: one CKDA layer tracks every finite group isomorphic to a subgroup of O(2), and many tracking results use one fewer layer than comparable diagonal-plus-rank-one linear RNNs. The paper gives both upper and lower bounds: one layer tracks cyclic and dihedral groups (including a three-dimensional realization of the cube rotation group), while Theorem 4 rules out one-layer tracking of S5 for CKDA and DeltaProduct with k≥2 under non-expansive transitions and finite reachability; three CKDA layers solve every finite group-word problem and, allowing β>2, recognize every regular language and compute WFAs in polynomial precision, saving one layer over the four-layer DeltaNet/GDN construction. Constructive proofs give explicit transitions and decoders (including a four-dimensional encoder and quadratic decoder tracking A5 with at most 120 reachable hidden matrices); the lower bound relies on an orthogonalization reduction under finite reachability plus a lemma on two planar rotations.
Experiments show length extrapolation on group-word problems and periodic audio continuation only when both range extensions are combined, while language modeling matches KDA and outperforms Transformers and other baselines. Among the four tested KDA range settings, only CKDA extrapolates well on S3, S5, and periodic waveform continuation; trained 1.3B models do learn negative gates, β>1, and complex eigenvalue pairs, primarily in the first two layers. Single-layer models train up to length 32 and report the best of three seeds; the audio task retains dB-level SNR at length 264 beyond the maximum training length of 136 while a causal Transformer degrades; at 1.3B parameters and 100B tokens of FineWeb-Edu, CKDA's average downstream accuracy is similar to KDA and above Transformer and other linear RNN baselines; kernel implementations retain a fraction of KDA throughput.
Perspective
The result speaks to readers studying linear RNN expressivity and state tracking, and to model designers who need to preserve phase or group-structured information over long sequences; it applies to single-layer or three-layer recurrences under exact arithmetic or a fixed exact datatype, and to language-modeling comparisons at the 1.3B-parameter, 100B-token scale. It enables follow-up work to obtain rotational dynamics through signed channel-wise gates without auxiliary recurrent states or explicit phase gates, while reusing existing KDA kernels.
The paper states that expressivity results do not imply learnability: A5 is representable but was not learned from standard random initialization, and only a smaller model initialized near the quaternion construction achieved extrapolation. The general WFA construction requires β>2, sacrificing the non-expansiveness guarantee, and relies on exact arithmetic over a fixed algebraic number field; error accumulation under floating-point approximation is not analyzed. The gate and β are used during training, but their functional roles remain unclear, which the authors leave to future work. In addition, the loaded text omits some concrete numbers and figure or table identifiers, so several experimental values can only be read through the ranges or relative comparisons given in the text.
