Public articles linked to the same research event.
arXiv The study investigates whether small language models (SLMs) can serve as task-tool relevance classifiers that evaluate each selected tool independently against the assigned task and return a relevance signal for downstream enforcement, and it introduces a novel dataset of multi-tool tasks whose required tools span distinct Model Context Protocol (MCP) servers, using prompt optimization, supervised fine-tuning, and reinforcement learning through GRPO to optimize and specialize SLMs.
The study investigates whether small language models (SLMs) can serve as task-tool relevance classifiers that evaluate each selected tool independently against the assigned task and return a relevance signal for downstream enforcement, and it introduces a novel dataset of multi-tool tasks whose required tools span distinct Model Context Protocol (MCP) servers, using prompt optimization, supervised fine-tuning, and reinforcement learning through GRPO to optimize and specialize SLMs.
The study investigates whether small language models (SLMs) can serve as task-tool relevance classifiers that evaluate each selected tool independently against the assigned task and return a relevance signal for downstream enforcement, and it introduces a novel dataset of multi-tool tasks whose required tools span distinct Model Context Protocol (MCP) servers, using prompt optimization, supervised fine-tuning, and reinforcement learning through GRPO to optimize and specialize SLMs.
The study investigates whether small language models (SLMs) can serve as task-tool relevance classifiers that evaluate each selected tool independently against the assigned task and return a relevance signal for downstream enforcement, and it introduces a novel dataset of multi-tool tasks whose required tools span distinct Model Context Protocol (MCP) servers, using prompt optimization, supervised fine-tuning, and reinforcement learning through GRPO to optimize and specialize SLMs.