Skip to main content
Back to timeline
arXivSource publication:

Jev holds high NDCG@10 across three Amazon domains while its latency grows far more gradually than pointwise Qwen

Synopsis

On Movies and TV, Video Games, and Books from Amazon Reviews 2023, the study builds controlled hard candidate sets from SASRec's highly ranked negatives (candidate sizes 10 to 100) and compares Jev, a decision-oriented "System One Model" from TypeSafe AI, against SASRec, DCNv2, and pointwise and listwise Qwen2.5 7B Instruct rerankers, finding that Jev maintains strong recommendation effectiveness (NDCG@10, Hit Rate@10, MRR) while its observed serving latency grows substantially more gradually than pointwise Qwen, though it remains far above recommendation-specific models.

AI-generated editorial illustration: Decision-Oriented Recommendation Reranking: An Empirical Study of Jev

Interpretation

Under controlled hard-candidate reranking, Jev maintains consistently strong recommendation effectiveness across all three domains, generally achieving higher NDCG@10 than the other evaluated methods on Video Games and Books, and performing similarly to the best Qwen configuration at smaller candidate sizes on Movies and TV while remaining competitive as the candidate set grows. Prior LLM reranking work largely examined pointwise, pairwise, or listwise prompting formulations; this study places a decision-oriented model as a distinct paradigm in the same candidate sets, same users, and same textual representation against recommendation-specific models and Qwen rerankers. Three Amazon domains with fixed candidate sets shared across all methods; effectiveness measured primarily by NDCG@10, with MRR and Hit Rate@10 showing consistent trends in the appendix; evaluated users number 954 for Movies and TV, 1,000 for Video Games, and 626 for Books.

Candidate-set size is a key experimental dimension that separates reranking paradigms: effectiveness declines for reranking methods other than SASRec as the candidate set grows, but Jev degrades comparatively gradually, pointwise Qwen also degrades relatively smoothly, and listwise Qwen declines substantially more sharply at larger candidate sets. The study treats candidate size from 10 to 100 as an explicit variable and shows that conclusions drawn from a single small candidate set may not generalize to larger reranking problems, a point less visible in single-size comparisons. Repeated evaluation on the same frozen candidate sets across four candidate sizes, cross-checked with MRR and Hit Rate@10 beyond NDCG@10; SASRec's Hit Rate@10 stays unchanged across sizes because the candidate pool derives from its own ranking.

Latency behavior runs opposite to effectiveness: pointwise Qwen latency rises rapidly with candidate-set size, listwise Qwen scales substantially more gradually, and Jev's observed serving latency also grows comparatively gradually, remaining in the low-second range even at candidate size 100 across the three domains and widening its gap over pointwise Qwen as the candidate set expands. The study characterizes effectiveness and latency jointly as operating points rather than reporting a single-dimension winner, yielding distinct operating regimes along the candidate-size dimension. Latency measured under a unified protocol, with local methods executed on a single NVIDIA A800-SXM4-80GB GPU and Jev accessed through the TypeSafe hosted API, so reported values are observed serving latency rather than hardware-normalized intrinsic inference time; the appendix reports medians and quartiles consistent with the averages.

On the effectiveness-latency plane, Jev frequently lies on the empirical non-dominated boundary among the evaluated methods, especially at larger candidate sizes; SASRec and DCNv2 occupy a low-latency but generally lower-quality region, pointwise Qwen achieves stronger effectiveness at substantially greater latency, and listwise Qwen reduces latency but generally occupies lower-quality regions. The authors explicitly bound this boundary to the empirical position relative to the recommendation-specific and Qwen baselines considered here, not a global Pareto frontier for recommender systems, leaving room for testing with larger models, additional decision-oriented models, and real-world environments. Conclusions drawn from joint presentation of effectiveness and latency across three domains and four candidate sizes; the authors also note they do not have access to Jev's architecture, training, or serving implementation, so attribution to the model versus the serving system cannot be established.

Perspective

The study targets the second-stage reranking of a two-stage recommender, applicable when candidates are already produced by a retrieval model and the ranking output is a structured choice among predefined alternatives; evaluation is limited to Movies and TV, Video Games, and Books from Amazon Reviews 2023, with candidate sizes from 10 to 100 and only users whose ground-truth item SASRec retrieved within the top 200. For engineering readers deciding whether to introduce a decision-oriented model under latency constraints, the paper offers an effectiveness and observed-latency comparison under identical candidate sets and textual representations; for research readers, it provides a reusable controlled hard-candidate construction and multi-metric evaluation framework that keeps candidate size as an explicit variable.

Jev is accessed through a hosted API whose underlying hardware, numerical precision, batching, and serving configuration are not disclosed, so latency differences may stem from the model, the serving system, or both, and the authors accordingly frame results as an empirical characterization under the evaluated deployment setting. Evaluation covers only one decision-oriented model and one LLM scale, Qwen2.5 7B Instruct; larger or proprietary models may show different quality and serving costs, so conclusions should not be extended to decision-oriented models or LLM reranking as a whole. Candidate sets are built from SASRec's highly ranked negatives and restricted to users whose target item entered the top 200, a setting that differs from directly reranking the retriever's top items. In addition, some formulas and the specific values of Figures 1, 2, and 3 are not fully rendered in the loaded text, so exact curve shapes for effectiveness and latency still need to be read from the original figures.

Sources