Skip to main content
Back to timeline
arXivSource publication:

A keyword harness passed a 1B model that never calls tools: a strict diagnostic ladder found the missing prior and repaired it in 3.3 GPU-hours

Synopsis

Working on a matched-architecture pair of Spanish security small models (VectraYX-600M at 661.6M parameters and VectraYX-1B at 1,109M, sharing decoder, tokenizer and special-token layout), the work documents a keyword-harness false positive: the two score almost identically on the lenient B4 tool-use metric (0.660 vs. 0.650), yet a verbatim-reproduction check shows the 600M emits valid tool calls with generalized arguments on 6/6 training examples while the 1B does so on 0/4–6 at every pre-remediation checkpoint; a first-token probe localizes the failure to a missing prior of roughly 10^-4–10^-5 on <|tool_call|>, and a targeted SFT recipe (diverse corpus, 5x learning rate, 2,202 steps, about 3.3 GPU-hours) raises the 1B's valid emission from 0.100 to 0.959 and passes 0.

AI-generated editorial illustration: Keyword Harnesses Fail Open: A Cheap Diagnostic Ladder for Tool-Use Claims in Small Language Models

Interpretation

Keyword-matching harnesses can credit a model that merely mentions a tool name in prose and never emits a structured call. The series had already documented the same failure class on the Nano model's B2 classification metric; this work reproduces it on the B4 tool-use metric and supplies a reusable strict replacement check. The two models differ by 0.01 on B4 (0.660 vs. 0.650), while the verbatim-reproduction check under maximally favorable conditions separates them completely (6/6 vs. 0/4–6); the authors therefore treat the 1B's B4 as unreliable and retain the 600M's, where lenient and strict instruments agree.

Whether tool calling arrives by default tracks pretraining composition: a code-heavy single-phase model emits structured calls by default, while a larger web-heavy sibling with a dedicated 6B-token tool-SFT phase does not. The contrast extends the series' earlier finding that bootstrap corpus register dominates downstream behavior from conversational register to structured output, and frames composition as deciding whether the capability is free or must be engineered. The pair shares decoder implementation, 32K tokenizer and special-token layout, but differs in parameter count (1.68x), total tokens and curriculum shape, which the authors explicitly call a natural experiment rather than an ablation; the 600M was frozen at 64% of schedule with no dedicated SFT and still shows the capability by default.

The failure is localizable as an absent prior rather than a noisy one, and repairable with far fewer tokens than the failed phase. The failed dedicated SFT phase spent 6B tokens, whereas the successful targeted recipe used roughly three orders of magnitude fewer tokens over 2,202 steps and about 3.3 GPU-hours, showing token volume predicted nothing and recipe decided everything. The first-token probe reads roughly 10^-4–10^-5 on <|tool_call|>; after repair, valid emission over all 269 corpus rows rises from 0.100 to 0.959 (600M: 0.926), and the repaired model passes 0.536 versus 0.428 on 238 unseen-entity prompts (paired p=0.004).

The repair changed the network's routing, not the trigger token's representation. The check rules out the natural broken-embedding hypothesis and, because input and output embeddings are tied, also covers the geometric target direction the network must aim at. The <|tool_call|> embedding row has cosine similarity around 0.999996 with minimal norm change, comparable to a baseline of 20 ordinary-token rows; 97.7% of the bf16 embedding table is bit-identical while late-layer attention projections show measurable drift as a positive control.

Perspective

The result speaks to teams building sub-1B tool-calling models on small budgets: treat the output format as a first-class pretraining-mix design variable, log first-token probes as near-free training telemetry, and gate any tool-use benchmark number behind a strict structural check. The diagnostic ladder itself costs minutes of CPU time and applies to any model lineage with archived checkpoints; the repair recipe was validated on one 1B checkpoint and re-run on a reproduced checkpoint of the same lineage, in the setting of Spanish-language security MCP tool calling.

The 600M run is permanently frozen at 64% of schedule without its cosine anneal, so its quality is a lower bound rather than a finished product; the matched pair differs in parameter count, total tokens and curriculum shape, which the authors call a natural experiment, leaving composition the leading surviving explanation rather than a demonstrated cause. The strict diagnostics are small-n (six verbatim examples, eight novel prompts), the generalization battery is built by the authors rather than sampled from an external distribution, and it covers neither unseen tools nor multi-call sequences. The repair's original checkpoint and its warm-start parent are lost, so every repaired-1B number comes from an end-to-end reproduction; of the recipe's four bundled ingredients, the warm-start point and the embedding table are eliminated, while corpus diversity and learning rate are separated by a five-arm factorial in which three of four cells were run once and the low-LR diversity advantage on suppression is not reproduced by a second seed, so it remains a hypothesis. Both models over-trigger, suppressing calls poorly on near-miss prompts that mention a CVE or shell command while asking a conceptual question; the repaired 1B largely recites on its own corpus and has not been shown to answer plain conversational questions usefully. In addition, the B4 series paired against the probe comes from a different harness invocation than the main table, and that discrepancy itself is further reason to distrust B4 for this model.

Sources