PinePaper-ToolBench: Benchmarking Tool Selection for MCP-Based Design Agents
Abstract
As LLM agent tool catalogs grow into the hundreds, selecting the right tools from a Model Context Protocol (MCP) server becomes a retrieval problem. Yet no public benchmark exists for evaluating tool selection methods in this setting. We introduce PinePaper-ToolBench, a benchmark of 582 test cases across 6 complexity tiers—from single-tool explicit queries to cross-domain ambiguous requests—covering 212 MCP tools for creative design.
Using this benchmark, we find that BM25 over taxonomy-enriched documents is already highly effective, achieving Recall@5 of 0.854. We then ask: does structured knowledge help? We evaluate KG-Hybrid, a two-phase method that augments BM25 with knowledge graph structural re-ranking via method-to-tool linking. KG-Hybrid achieves the highest Recall@5 of 0.868, with the improvement over BM25 statistically significant at the aggregate level (p=0.045), though gains concentrate on compositional and cross-domain tiers. The gains are largest on harder tiers: on compositional queries (T4), KG-Hybrid reaches 0.932 vs. BM25's 0.892 (+4.0 points); on cross-domain ambiguous queries (T6), KG-Hybrid scores 0.608 vs. 0.578 (+3.0 points). Ablation studies confirm that method-to-tool implements edges provide the primary structural signal, while graph propagation (PPR) has negligible effect.
We validate at scale with an extended benchmark (Track B) of 1,000+ cases over 500+ tools using vocabulary-separated synthetic tool generation with full KG integration. All code, benchmark data, and statistical tests are open-sourced to support reproducibility and future work on tool selection for LLM agents.
Cite
@misc{pinepaper2026toolbench,
title={PinePaper-ToolBench: Benchmarking Tool Selection for MCP-Based Design Agents},
author={{PinePaper Research}},
year={2026},
howpublished={\url{https://pinepaper.studio/research/pinepaper-toolbench.pdf}}
}