MCRI turns agent-skill evaluation into a pre-execution four-dimensional score, lifting top-1 skill selection by 17.7 to 22.8 percentile points across three benchmarks
Synopsis
The authors propose the four-dimensional MCRI Framework grounded in information gain and behavioral constraint and operationalize it as MCRI-Eval, a large language model-based evaluation method, validated on 63,812 public skills from the OpenClaw skill Hub and 58,275 skill-conditioned model executions across BigCodeBench, BFCL-Fundamental, and Mind2Web: MCRI-Eval scores are positively associated with community popularity signals, achieve the highest downstream ranking agreement among the evaluated methods, and improve top-1 skill selection over each benchmark's strongest baseline by 17.7, 22.8, and 19.6 percentile points, indicating it can serve as a pre-execution signal for prioritizing promising skills.
Figure 1: Constraint strength spectrum from natural language to skills and formal programs.
arXivInterpretation
The paper proposes the four-dimensional MCRI Framework for systematically analyzing agent skills and operationalizes it as MCRI-Eval, a large language model-based evaluation method. The academic community previously lacked a structured framework for systematically analyzing skills; this work grounds the framework in information gain and behavioral constraint, turning skill evaluation from scattered judgment into an operational, dimension-based score. Evidence comes from the framework design and an evaluation setup over 63,812 public skills from the OpenClaw skill Hub; the abstract does not give the concrete definitions or weights of the four dimensions.
MCRI-Eval scores are positively associated with community popularity signals and achieve the highest downstream ranking agreement among the evaluated methods. This offers a skill-quality signal that aligns with community popularity while agreeing more closely with downstream performance ranking, rather than relying on a single source of ordering. Based on 58,275 skill-conditioned model executions covering BigCodeBench, BFCL-Fundamental, and Mind2Web; the abstract does not report correlation coefficients or the specific agreement metric values.
MCRI-Eval improves top-1 skill selection on all three benchmarks, advancing by 17.7, 22.8, and 19.6 percentile points over each benchmark's strongest baseline. The gains are measured in downstream performance rank, indicating that pre-execution scoring can substitute for part of costly execution-based evaluation when prioritizing skills. Gains are reported in percentile points against the strongest baseline on each benchmark, at an evaluation scale of 58,275 executions.
Perspective
The work targets settings that require ranking many candidate skills before execution, such as selecting a top-1 skill for an agent from a skill hub. It applies to public skill ecosystems represented by the OpenClaw skill Hub and reports results on three benchmark families: BigCodeBench, BFCL-Fundamental, and Mind2Web. For teams building skill marketplaces, skill recommenders, or agent capability compositions, MCRI-Eval can serve as a pre-filter before execution-based evaluation, concentrating limited execution budget on skills more likely to be effective.
The abstract does not state the concrete definitions of the four MCRI dimensions, how the dimensions combine into a total score, or the specific values for the association with community popularity signals and the ranking agreement metric. The gains on the three benchmarks are reported in percentile points, but the set of baseline methods and statistical uncertainty are not described. In addition, the evaluation centers on the OpenClaw skill Hub and three benchmarks, so behavior on other skill ecosystems and task domains remains an open question. Because the reading scope here is the abstract only, figures and body details are not included, and the robustness of the reported numbers awaits checking against the full text.
