SEEK externalizes search evaluation criteria into a routable skill bank, improving listwise quality evaluation and attribution diagnosis in Kuaishou short-video search
Synopsis
The work proposes SEEK (Skill-routed Evaluation with Evolvable Knowledge), which externalizes specific search evaluation criteria into a skill bank, dynamically routes relevant skills for each query-result list pair, and employs a task-adapted listwise evaluator to produce page-level judgments and failure mode attribution; a two-stage training pipeline teaches the evaluator to align evaluation criteria with human preferences, while a replay-gated skill bank allows recurring evaluation knowledge gaps to be incorporated without model retraining; experiments on industrial short-video search show that SEEK improves listwise quality evaluation accuracy and achieves significant progress in attribution diagnosis, and SEEK has been deployed at Kuaishou, a platform with over 400 million daily activ
Figure 1. Illustrative example of skill-routed listwise evaluation in SEEK. The router selects task-relevant skills, and the listwise evaluator applies their guidance over the complete result page to produce a page-level judgment and attribution.
arXivInterpretation
SEEK externalizes search evaluation criteria into a skill bank and dynamically routes relevant skills for each query-result list pair, avoiding the irrelevant context and potential criterion interference that come from packing all criteria into a unified prompt. Relative to packing all evaluation criteria into a single prompt, SEEK splits criteria into routable skills and selects relevant ones per query-result list pair. The design is supported by the paper's problem statement about the two existing approaches (unified prompting and post-training internalization) and is validated in experiments on industrial short-video search.
SEEK uses a task-adapted listwise evaluator to produce page-level judgments and failure mode attribution, responding to the reality that users experience search results at the page level while applicable evaluation criteria are multi-dimensional and continuously evolving. Relative to evaluation that yields only an overall score, SEEK outputs both page-level judgments and failure mode attribution, extending evaluation from scoring to diagnosis. The paper reports improved listwise quality evaluation accuracy and significant progress in attribution diagnosis on industrial short-video search.
A two-stage training pipeline teaches the evaluator to align evaluation criteria with human preferences, while a replay-gated skill bank allows recurring evaluation knowledge gaps to be incorporated without model retraining. Relative to internalizing criteria through post-training, which tightly couples rule updates with costly model retraining cycles, SEEK decouples knowledge updates from model training. The paper describes this two-stage training pipeline and the replay-gated skill bank mechanism, and validates their effect in experiments.
SEEK has been deployed at Kuaishou, a platform with over 400 million daily active users, and the deployment significantly improved the scale and quality of online search evaluation. Relative to evaluation methods validated only in offline experimental settings, SEEK is deployed on a real industrial platform and reports improvements in the scale and quality of online evaluation. The paper reports the deployment at Kuaishou and its effect on the scale and quality of online search evaluation.
Perspective
The work targets quality evaluation for industrial search systems, especially page-level, multi-dimensional settings with continuously evolving criteria, and both experiments and deployment center on industrial short-video search. It lets evaluation criteria be updated through the skill bank without retraining the model, making it suitable for search teams that need to adjust evaluation rules frequently while maintaining evaluation scale and quality; for industrial settings seeking to replace manual assessment with scalable automatic evaluation, it offers a reference architecture and training pipeline.
What is available is abstract-level information, lacking specific experimental numbers, sample sizes, control settings, and how attribution diagnosis is evaluated, so the magnitude and statistical robustness of the improvements cannot be judged. The cost of building and maintaining the skill bank, the effect of routing errors on evaluation results, and the specific criteria by which the replay gate identifies knowledge gaps are questions a reader would need to confirm in the full text. In addition, the deployment effect is described in terms of the scale and quality of online search evaluation, and its degree of agreement with human assessment remains to be understood.
