Skip to main content
Back to timeline
arXivSource publication:

SHarP prunes agent harness modules by single-module ablation saliency: four harnesses cut token cost by at least 10% at comparable performance, and OpenHands on GAIA rises from 29.44% to 35.00%

Synopsis

The work proposes SHarP, which treats tools, instructions, and supporting mechanisms in an agent harness as individually disableable modules, estimates each module's performance and efficiency saliency through single-module ablation, and prunes by an efficiency-oriented or performance-oriented rule; across five settings (OpenHands/GAIA, PaperQA2/LitQA2, LIFE-harness airline and retail, JIT agentfold/DeepPlanning-Shopping), most pruned configurations retain comparable performance, all four harnesses achieve at least a 10% token-cost reduction, and two improve task performance while cutting cost.

Source-provided article image: SHarP: Saliency-based Pruning of Agent Harnesses
Figure 1 ·

Figure 1: Overview of SHarP ’s efficiency-oriented pruning. Single-module ablations estimate performance and efficiency saliency, which determine the protected module set (green) and pruning order, respectively. Pruned configurations and an all-pruned baseline are evaluated for task performance and token cost.

arXiv

Interpretation

The paper formulates harness pruning as an operational procedure: first partition the harness into individually disableable modules along existing implementation boundaries (a registered tool, an instruction block identified by heading or tags, a retrieved skill, a distinct processing mechanism), and define what pruning each module means, for example removing a search tool and its instructions from the prompt and disabling access to it, or pruning a history summarization mechanism by always keeping the full history. Prior work on harness bloat and efficiency largely observed that added mechanisms do not necessarily improve performance or restricted context and tool exposure; it left open how to quantify redundancy across heterogeneous components (tools, instructions, supporting mechanisms) and remove them. This work transfers the neural-network pruning notion of saliency to harnesses and fixes explicit conventions for module partitioning and disabling semantics. The method is specified as four steps (module identification, saliency computation, saliency-guided pruning, pruning evaluation), and Appendix A.1 lists each harness's candidate modules, functions, and disabling notes, making the procedure reproducible.

Saliency is estimated by single-module ablation: disable one module, run the resulting harness, record task score and token cost, compare against the median over all single-module-ablated harnesses on the same task, and test the two difference terms with one-sided paired t-tests whose p-values serve as performance and efficiency saliency; each module requires only one additional evaluation. Compared with exhaustively evaluating combinations of pruned components, this estimation reduces evaluation cost to one run per module, making pruning feasible in evaluation-expensive agent settings; it also folds the harness-specific property that removing a module does not necessarily reduce cost into the same saliency framework. The paper gives the difference terms, the median reference, the one-sided paired t-test with degrees of freedom and p-value computation, and a convention for degenerate paired differences with zero standard deviation; Appendix A.4 reports per-module ablation task scores, token costs, and p-values.

Two pruning rules share the same saliency estimates but swap roles: efficiency-oriented pruning first protects modules crucial for performance, then removes unprotected modules ranked by efficiency saliency from high to low; performance-oriented pruning symmetrically protects modules crucial for efficiency and removes by performance saliency. This introduces a protected-set-plus-ranked-removal budget mechanism to harnesses, so one set of saliency estimates can generate a sequence of pruned configurations under different objectives rather than a single compressed result. On JIT agentfold/DeepPlanning-Shopping and LIFE-harness retail with one trial per validation task, performance-oriented pruning reaches an observed 82.5% match rate at one configuration with a 5.9% mean token-cost reduction, while efficiency-oriented pruning reaches 76.2% match rate at another with a 39.0% token reduction; on retail, efficiency-oriented pruning reaches 87.5% success rate at one configuration, and performance-oriented pruning reaches the lowest token cost at another, reducing mean token cost by 21.2% at 70.0% success rate.

Main results across five settings show substantial harness redundancy: many pruned configurations retain similar observed performance while reducing token cost, all four harnesses achieve at least a 10% token-cost reduction, and two improve task performance while cutting cost; but the extent of redundancy varies by harness and task domain, over-pruning degrades performance, and pruning can also increase cost. Prior work lacked a systematic quantification of how much redundancy exists in a harness; this work provides empirical evidence from a sequence of pruned configurations on held-out validation sets and shows that pruning benefits are not monotonic. OpenHands on GAIA reaches its highest success rate with 24 of 27 modules pruned, rising from 29.44% for the full harness to 35.00%, but falls to 14.44% when all 27 are pruned; PaperQA2 achieves its highest performance with the full harness; LIFE-harness shows a modest decrease on airline and a slight increase on retail in the all-pruned configuration; PaperQA2's 15-module-pruned configuration and JIT agentfold's all-pruned configuration are more expensive than their full harnesses.

Perspective

The results target engineering and research settings that need to reduce agent token cost under a fixed model: a user can partition their own harness into modules, run single-module ablations and saliency estimation, generate a sequence of pruned configurations under the efficiency-oriented or performance-oriented rule, and compare them on a held-out validation set. The evaluation uses Qwen3.5-122B-A10B as the base model throughout and covers OpenHands/GAIA, PaperQA2/LitQA2, LIFE-harness airline and retail, and JIT agentfold/DeepPlanning-Shopping; module partitioning follows existing implementation boundaries, with highly coupled elements such as a tool and its usage instructions merged into a single module. Pruning benefits differ by harness and task domain, and the efficiency-oriented and performance-oriented strategies each lead on DeepPlanning-Shopping and retail respectively, so applicability should be judged by validation results on the target harness and domain.

The pruning-strategy comparison on JIT agentfold and LIFE-harness retail uses one trial per validation task, whereas main results default to three trials, so whether the relative merits of the two strategies hold under different trial budgets is worth rechecking in one's own setting. Module partitioning is done manually along implementation boundaries, and a different granularity could change saliency estimates and the prunable range. The paper reports observed performance and token cost on held-out validation sets; token counts include unsuccessful attempts but do not measure elapsed execution time, and PaperQA2's offline corpus preprocessing and index construction are excluded. Pruning can also remove functionality that helps the agent terminate efficiently or filter evidence, raising cost, so pruned configurations still need validation of termination behavior and evidence-processing paths in the target domain.

Sources