acceptodds
Under review as a conference paper at ICLR 2027

ToolSAE: Data Selection for Tool Learning via Capability-Relevant Neuron Activations

Abstract

Tool calling is a foundational capability that enables large language model (LLM) agents to interact with the external world. Training such capabilities increasingly relies on large-scale tool-calling traces, making it critical to identify the most useful traces from these extensive datasets. Existing data selection methods typically prioritize alignment with target examples or generic criteria such as quality and diversity, yet these signals often capture only surface features (e.g., schema keywords), while failing to reflect the specific tool-calling capabilities to be strengthened. In this work, we introduce ToolSAE, a capability-guided data selection framework that explicitly identifies tool-capability neurons in the SAE latent space and selects training examples based on their activations. We employ sparse autoencoders to decompose polysemantic model activations into sparse, interpretable representations. Through contrastive neuron profiling and token-level diagnosis, ToolSAE identifies capability-specific neurons while filtering out those responding primarily to surface formatting. It then ranks candidate traces by their weighted activations over the retained neurons, prioritizing examples that engage tool-calling behaviors. Empirical evaluations show that training on just 5% of the tool-calling data selected by ToolSAE achieves tool-calling performance competitive with full-data training while maintaining comparable general-domain performance. Further analysis shows that the identified neurons capture tool-calling semantics and provide effective signals for capability-guided data selection.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.