Skip to main content
Back to timeline
arXivSource publication:

ED²: Unleashing LLM Potential for Sequential Recommendation via a Dual Dynamic Index Mechanism

Synopsis

This work proposes the End-to-End Dual Dynamic (ED²) recommender, the first LLM-based sequential recommender to adopt a dual dynamic index mechanism that unifies user/item index generation and sequential recommendation into a single jointly optimized LLM-backbone pipeline, complemented by a multi-grained token regulator (m-GTR) and instruction-tuning tasks for high-order user-item interaction patterns, achieving average improvements of 19.62% in Hit-Rate and 21.11% in NDCG over static-index SOTA LLM recommenders on three public datasets.

Source-provided article image: Unleash LLMs Potential for Sequential Recommendation by Coordinating Dual Dynamic Index Mechanism
Figure 1 ·

Figure 1. Overview of a). User interaction sequence, b). Static item index based sequential recommender systems, and c). Dual dynamic index based sequential recommender system. W/O BP represents for without back-propagation.

arXiv

Interpretation

It proposes ED², the first LLM-based sequential recommender with a dual dynamic index mechanism, assembling the index generator and the sequential recommender into a unified LLM-backbone pipeline optimized end-to-end. Existing LLM-based sequential recommenders mostly use a static index that separates index generation from recommendation and stays frozen during recommender optimization; ED² jointly optimizes index and recommender, fusing semantic and collaborative information into one LLM backbone. The paper provides full method derivations (Formulas 1-11) and experiments on three datasets; Table 1 shows ED² is best on all metrics for Instruments, Games, and Arts with p<0.01 significance, averaging 19.56% Hit-Rate and 21.11% NDCG gains over static-index SOTA LLM recommenders.

It designs the multi-grained token regulator (m-GTR), which builds alignment supervision at both the index level and the token level to improve LLM comprehension of dynamic index tokens. Prior LLM recommenders rely only on sequential-recommendation fine-tuning to understand index tokens, while dynamic indices optimized synchronously with the LLM backbone widen the comprehension gap versus natural-language tokens; m-GTR explicitly aligns dynamic indices with their corresponding textual features. Ablation (Table 2) shows removing m-GTR reduces average performance by 4.13%; parameter sensitivity (Figure 4) shows performance is insensitive to weights β₁ and β₂ (standard deviation below 9.7×10⁻⁴).

It constructs associated user collections and customizes a series of instruction-tuning tasks so the LLM can exploit high-order user-item interaction patterns (user co-purchase and user preference patterns). Most leading LLM sequential recommenders rely only on item-related information (item text and interactive item sequence) and ignore user-related information; ED² adds a user dynamic index branch plus user prediction, comment/query inference, and profile summarization tasks. Ablation (Table 2) shows removing the high-order pattern exploitation tasks reduces average performance by 2.52%, and merely adding user information without corresponding instruction tasks (w/o exploit) yields no gain over fully removing user modules (w/o user); ED² outperforms CLLM4Rec by 72.80% in Hit-Rate and 66.97% in NDCG.

It verifies the general advantage of the dynamic index mechanism over static and LSH indices, and provides engineering solutions for index length and collision handling. Beyond comparing with static indices, the paper introduces dynamic/static LSH index variants and index-length variants of 3 and 5, showing the dynamic index advantage is loosely coupled with the specific indexing method, and uses an extra rehashing token to guarantee index uniqueness. Table 3 shows the Static variant drops 17.57% on average, S-LSH drops 9.73% while D-LSH drops only 4.66%; 3-ED² and 5-ED² drop 3.51% and 2.51% respectively; Appendix H reports an actual collision rate around 3×10⁻⁴.

Perspective

The results target sequential recommendation settings where users and items are described by textual features and interaction sequences are available; experiments are conducted on three categories of the Amazon Product Review dataset (Musical Instruments, Video Games, and Arts, Crafts and Sewing), with roughly 10K-45K items and average sequence lengths around 8-9. The method suits LLM recommender systems that aim to jointly optimize index generation and recommendation while incorporating user-side information; an index length of 4, 256 embeddings per codebook, and 2048 index tokens in total were validated as appropriate at this data scale.

Readers may continue to watch: how index stability and training convergence interact when dynamic indices are optimized synchronously with the LLM backbone; how controllable the collision rate is as data scale, model architecture, and optimization parameters vary; the applicability of user-side instruction tasks when user textual features are sparse or missing; and how the mechanism performs at larger scales, longer sequences, or on non-Amazon datasets. The original is a full paper with figures and appendices loaded, but some implementation details (such as specific instruction templates) are given in appendices and should be read together with them.

Sources