Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

arXiv

Researchers propose using small language models to verify each agent tool call against task intent, with a multi-tool dataset spanning distinct MCP servers

The study investigates whether small language models (SLMs) can serve as task-tool relevance classifiers that evaluate each selected tool independently against the assigned task and return a relevance signal for downstream enforcement, and it introduces a novel dataset of multi-tool tasks whose required tools span distinct Model Context Protocol (MCP) servers, using prompt optimization, supervised fine-tuning, and reinforcement learning through GRPO to optimize and specialize SLMs.