Skip to main content
Back to timeline
arXivSource publication:

From Overlooked to Explored: Recovering Item Relations via Mixture of Perspectives for Sequential Recommendation

Synopsis

Through empirical analysis, this work identifies a "similarity bias" in dot-product self-attention that systematically overlooks heterogeneous item relations carrying causal signal for target prediction, and proposes PRISM, a module that calibrates attention through K Perspective Lenses operating in two complementary views, an Affinity View and a Contrast View, consistently outperforming state-of-the-art baselines on seven real-world benchmarks.

Source-provided article image: From Overlooked to Explored: Recovering Item Relations via Mixture of Perspectives for Sequential Recommendation
Figure 4 ·

Figure 4 . PRISM framework: PRISM performs item-wise multi-perspective analysis followed by relational insight synthesis. All learnable parameters are shared across lenses for efficiency, while the outputs differ solely by relational perspective.

arXiv

Interpretation

The paper identifies and empirically characterizes a "similarity bias" in self-attention: the Spearman rank correlation between attention and an item's causal importance, ρ(A,I), is negative for SASRec on both Toys and Beauty, and Align@1 is similarly low. Prior work had reported erroneous attention distributions, but this paper attributes the cause to dot-product similarity itself and verifies across four transformer baseline categories (vanilla self-attention, attention calibration, intent-based modeling, and MoE within attention heads) that the misalignment persists. On 1,000 test sequences from Amazon Toys and Beauty, ρ(A,I) and Align@1 are computed for five models; causal importance is defined counterfactually as the drop in target probability when a position is masked, I_j.

Counterfactual attention interventions show that overlooked positions do carry useful signal: boosting the bottom 20% of positions by attention (C3) raises target confidence across all baselines and both datasets, and boosting only overlooked-yet-causally-important positions (C4) consistently outperforms blind boosting (C3). This moves the diagnosis from observation to an actionable direction, indicating that useful relations are spread across diverse relational facets rather than concentrated in a single type. Last-layer attention logits are modified at inference with boosting strength β_boost = 5.0, and Δ_confidence(%) measures the relative change in target probability; even random perturbation (C2) helps substantially, indicating baseline attention poorly reflects which items the model relies on.

PRISM uses a Semantic Anchor Router to assign each item a semantic group and a Primary Semantic Anchor, then K Perspective Lenses calibrate attention logits via a Semantic Focus Mask and a Relational Boost Signal, so each item is analyzed under exactly one Affinity View and K−1 Contrast Views. Unlike attention calibration that relies on a learnable matrix without explicit relational guidance, PRISM uses inter-item relations as explicit guidance; unlike MoE applied within attention heads, it recalibrates attention within each perspective, directly addressing similarity bias. The method is specified with full equations (routing logits, semantic composition ratio, mask, boost signal, Perspective Guided Attention, and synthesis), and all learnable parameters are shared across lenses for efficiency.

The training objective combines a sequence-level preference contrastive loss, which builds positive pairs via same-target augmentation and uses Stochastic Perspective Masking to avoid over-alignment, with a collaborative consistency loss that uses symmetric KL divergence to keep lens outputs consistent in item space without collapsing their diversity. It couples item-level relational analysis with sequence-level preference modeling while explicitly preventing multi-perspective representations from collapsing into a single mode. Both losses are given with explicit formulas and masking mechanisms (Bernoulli mask, log-mask operation, temperature τ, batch size B), constituting a design-level argument; the abstract reports consistent outperformance of state-of-the-art baselines on seven real-world benchmarks.

Perspective

The work targets sequential recommendation, predicting the next item from a user's interaction history; its diagnostic experiments are conducted on Amazon Toys and Beauty with 1,000 test sequences, and the method is designed as L−1 modules inserted into an L-layer transformer, so it applies to recommendation models built on self-attention backbones. For researchers and engineering teams seeking to improve how attention selects items rather than simply enlarging models or feature sets, this perspective offers a reusable design idea; the seven real-world benchmarks reported in the abstract point to robustness across varied real-world scenarios.

A careful reader may still watch: whether similarity bias behaves consistently across domains and item-set scales; how the number of semantic groups K and the perspective dropout rate ρ shape the balance between the two views; whether the conclusions from counterfactual interventions with a fixed boosting strength β_boost = 5.0 hold under finer-grained boosting magnitudes; and the specific metrics and ablation details across the seven benchmarks, which require the original figures and tables for confirmation.

Sources