Skip to main content
Back to timeline
arXivSource publication:

RPTune pairs learned catalog curation with post-training to lift conversational product search accuracy by up to 31.4 points across 7 real merchants

Synopsis

For small merchants whose catalogs fit inside a long-context window, RPTune couples an encoder–reorganizer learned catalog curator with LLM post-training using automatically generated, catalog-grounded supervision: across 7 real merchants and 100 complex conversational queries each, context curation raises exact-match accuracy by 13.5 points on average and up to 20 points across 9 frozen LLM backbones, curation alone adds 14.6 points on gemini-3.7-flash, and curation plus post-training lifts gemma-4-E4B-it from 10.7% to 31.0% EM.

AI-generated editorial illustration: RPTune: Learned Context Curation for LLM Catalog Search

Interpretation

The paper establishes in-context catalog search as a distinct problem setting for small merchants: when the whole catalog fits in a long-context window, prompting the LLM with the full catalog is a better fit than multi-stage retrieval. Prior product search methods target large marketplaces with millions of items, dedicated ML teams, and abundant behavioral signals; the authors analyze 56 real Shopify storefronts, find 92.9% of catalogs fit within 1M tokens with a median of 51K tokens, and report that full-catalog LLM search improves over prior retrieval-based approaches by up to 20 percentage points. Based on catalog-size statistics over 56 storefronts and a comparison using gemini-3.7-flash as the shared backbone: full-catalog prompting beats all baselines by at least 9.0 EM points while keeping single-digit latency (8 s).

RPTune's encoder–reorganizer curator ranks, prunes, and positions products guided by downstream LLM feedback, outperforming off-the-shelf dense retrieval and listwise reranking at the same 25% retention budget. Prior context compression and reordering methods are not trained against how a specific downstream LLM actually behaves on a given catalog; RPTune's reorganizer is updated with group-relative-normalized reinforcement learning over the frozen LLM's sampled selections and catalog-wide relevance rewards. Across 9 frozen backbones, EM improves by 13.5 points and FR by 10.1 points on average, with gains in all 63 backbone–merchant combinations; on gemini-3.7-flash, substituting Jina Reranker 3.5 or Gemini Embedding 2 degrades accuracy on 5 and 3 of 7 merchants respectively, whereas RPTune improves all 7; end-to-end latency drops by 28% on average and up to 71%.

Post-training the LLM on curated contexts adds gains beyond inference-time curation, and the context-relative reward is the key design choice. Post-training uses automatically generated, catalog-grounded supervision, requiring no merchant-provided relevance labels or user behavior logs; the reward credits the highest-relevance product available in the curated catalog, preserving a useful learning signal when the globally best variant has been pruned. On gemma-4-E4B-it, EM rises from 10.7% to 16.4% with the encoder, 20.7% with the full curator, and 31.0% with curation plus post-training, while FR rises from 56.7% to 74.4%, increasing monotonically across all 7 merchants at each stage; ablations show EM on Beauty Bakerie drops 10 points with continuous or global top-1 rewards.

RPTune generalizes zero-shot to newly added products and to unseen merchants. Frequent inventory updates make constant retraining impractical; the authors train once on half of the Beauty Bakerie catalog and incrementally inject withheld products, and cross-evaluate each of the 7 merchant-specific models on all catalogs. As the catalog doubles, post-trained RPTune shows only a 14.7% relative EM decline versus the baseline's 29.5%, with FR practically unchanged, and keeps absolute gains of 35.0 EM and 19.2 FR at every injection level; in cross-merchant evaluation all combinations improve both EM and FR over the gemma baseline despite a mean pairwise Jaccard similarity of 0.146 across catalogs.

Perspective

The result targets small merchants whose catalogs fit entirely inside a long-context window: among the 56 Shopify storefronts the authors analyze, 92.9% of catalogs fit within 1M tokens with a median of only 51K tokens, and the 7 evaluated merchants range from 37,506 to 126,958 tokens. Evaluation is limited to single-turn queries; the authors explicitly leave multi-turn evaluation to future work while noting such queries can be decomposed into successive turns and that RPTune can be adapted by running the encoder–reorganizer at each turn. The method needs only the merchant catalog as a data source, with training supervision generated automatically by a frontier LLM, so it suits merchants without relevance labels or user behavior logs; the cross-merchant evaluation and the incremental product-injection protocol indicate the framework is meant for settings with frequently changing inventory.

Evaluation queries were synthesized by gemini-3.1-pro and filtered by consensus plus human audit; the authors themselves flag that reference labels may favor the Gemini family and respond with a non-Gemini judge panel and the observation that Grok's average gain (+19.3 EM) exceeds Gemini's (+12.5 EM). The audit accepts the reference label for 64.9% of queries versus 1.1% for BM25 distractors and 0% for random ones, while an average of 1.85 variants per query are acceptable and pairwise agreement between judges on the best variant ranges only from 36.4% to 46.1%, so a single reference cannot capture every defensible answer. The failure analysis shows remaining errors are close substitutes: error trials have high mean feature reward, and same-product different-variant cases are common, for example on Package Free where 8 of 10 trials return the 100-load pack of the same detergent instead of the 140-load pack. In addition, the New-product group in the catalog-increase experiment contains only 4 to 7 queries at low injection levels, which the authors say should be read as indicative only. A reader who stops at the abstract will miss these label-audit, error-structure, and injection-protocol details.

Sources