ZooWork aligns e-commerce rerankers with cross-family LLM judge labels: the 8B and 4B significantly beat the strongest open baseline on ShopRank-Bench, and the distilled 0.6B matches a 4B base on structured text
Synopsis
ZooWork presents ZooWork-ShopRanker, a family of e-commerce rerankers (0.6B, 4B, 8B) trained on preference labels produced by three reasoning LLM families (Qwen3.5-122B, Gemma-4-31B, DeepSeek-V4-Pro) under a constraint-first, both-orders judging protocol, with the aligned 8B distilling into the 4B and 0.6B; on ShopRank-Bench, built from Gensmo private traffic (10,511 pairs tiered gold/silver/bronze), the 8B and 4B significantly outperform the strongest open baseline Jina-m0, every size significantly beats its own un-aligned base, the 0.6B is statistically indistinguishable from the 4B base on structured text, and alignment costs nothing in general MTEB reranking quality or serving latency.
Interpretation
It builds a family of e-commerce preference rerankers labeled by cross-family LLM judges, where the 8B and 4B significantly beat the strongest open baseline on ShopRank-Bench and every size significantly beats its own base. Prior open rerankers (BGE-Reranker-v2-m3, Jina, Qwen3-Reranker) are trained and evaluated mainly on general retrieval benchmarks such as BEIR and MS MARCO; this work replaces the supervision with constraint-first judge labels and releases open models at 0.6B, 4B, and 8B. On structured text the 8B scores 83.4 overall and the 4B 81.2 against Jina-m0's 79.2; paired McNemar tests support the differences, and a query-clustered bootstrap with Holm-Bonferroni correction leaves every conclusion unchanged.
It releases ShopRank-Bench: 10,511 preference pairs from Gensmo private search traffic, tiered by how many judge families committed into gold (1,843, 3/3), silver (4,445, 2/3), and bronze (4,223, 1/3), in both structured and natural-language formats. Public e-commerce sets such as ESCI-style collections risk pretraining memorization, and a single serialization rewards formatting artifacts; here the pairings and labels have never been public, and the same query, products, and label are held fixed across both formats, with attribute-hierarchy (1,500 pairs) and budget (1,302 pairs) diagnostic tracks alongside. Of 23,000 judged candidates, 10,511 survive (45.7%), with 12,350 unanimous ties and 139 outright conflicts discarded; the training slice and the benchmark share zero query strings and zero exact pairs; a length heuristic wins at most 50.2% and a price heuristic 59.9% on the comparable subset, far below every learned model.
The aligned 8B serves as a distillation teacher so the 0.6B and 4B students combine the teacher's high-volume coverage with the judges' high-precision ordering, and distillation beats aligning the same base directly. The small models are not scaled-down copies of the 8B recipe: they first fit the 8B's soft scores on roughly 135k query-document examples, then are sharpened on judged pairs with the pairwise loss, which makes the 0.6B statistically indistinguishable from the 4B base on structured text at 433 versus 155 docs/s and 4.5 versus 11.2 GB. The alignment-from-base 4B is significantly weaker than the released distilled 4B in both formats (79.0 vs. 81.2 structured, 77.4 vs. 79.5 prose); a panel teacher built from DeepSeek and Gemma grades underperforms at feasible scale, with the 0.6B student reaching 70.7 versus 78.3 on the development hard-structured slice.
Alignment does not trade away general reranking quality or inference cost: with the LoRA update merged, each model matches its own base at median latency, and the 4B and 8B lead the MTEB reranking average. A common concern is that domain alignment causes general-capability regression; this work uses one architecture across three scales with LoRA adapters updating roughly 1% of parameters and reports the single slippage point, SciDocsRR. MTEB averages are 0.799 for both the 4B and 8B against 0.793 for Qwen3-Rnk-8B; median latency is 17.6 vs. 18.6 ms at 0.6B, 25.2 vs. 24.5 ms at 4B, and 25.8 vs. 25.4 ms at 8B; a controlled comparison shows LoRA matches or edges full fine-tuning.
Perspective
The work targets the final reranking stage of an e-commerce search stack, for deployment settings where queries are short and underspecified and candidates come from production retrieval, and it explicitly uses Gensmo private traffic as its evaluation substrate. What a reader can reuse directly is the three open model sizes, the dual-format ShopRank-Bench with its gold/silver/bronze tiers, the attribute-hierarchy and budget diagnostic tracks, and the finding that alignment is free at inference: with the LoRA update merged, each model matches its own base at median latency, and the 0.6B reaches the 4B base level on structured text at roughly a seventh of the size, so quality-bound settings favor the 8B while throughput-bound settings favor the 0.6B or an encoder cross-encoder. The diagnostic tracks further show that a programmatically checkable constraint such as an explicit budget is nearly solved by a little targeted supervision, whereas a hand-designed attribute priority is a poor training target.
Labels come from three LLM families rather than human verification, and the 45-pair fourth-family check is a consistency check with wide intervals; the bronze tier is 40% of the benchmark and largely Gemma-decided, so an aggregate score is dominated by the weakest labels. The panel is asymmetric in commitment rates (Gemma 92.7%, DeepSeek 66.1%, Qwen3.5-122B 18.6%), and two benchmark families also produced training labels; the authors bound this by showing the advantage over Jina-m0 shrinks from about 8 points to about 3 on pairs a training-disjoint family also decided, but run no interaction test. Zero-shot reasoning LLMs still lead the 8B by 8-9 points, so the preference is not fully recovered by a reranker; the natural-language view is model-rendered with varying attribute inclusion, and whether every rendering preserves the attributes its label turned on is not audited; public MTEB comparisons may reflect pretraining contamination. A fast parse without figures would leave the tiering and paired-test details to be checked against the original.
