acceptodds
Under review as a conference paper at ICLR 2027

CountTrigger: What Do LLMs Actually Learn to Count in Cardinality-Conditioned Backdoors?

Abstract

Large language models increasingly operate over structured collections of tools, records, evidence items, and workflow events. We ask whether the cardinality of task-defined structures can become a latent condition for model behavior. We introduce CountTrigger, a family of structural cardinality predicates specified by what to count, where and how counting is applied, and when the resulting count activates behavior. We first give a formal grounding for fixed threshold, exact-count, and band predicates under explicit unit-definability assumptions, separating representability from learnability. Empirically, CountTrigger activation is broadly learnable across 20 model–task settings, while counterfactual controls show that high attack success can coexist with reliance on surrogate structural cues. Matched finite-state supervision makes accumulation states highly decodable, yet multi-seed activation interventions reveal a representation–behavior gap: event-level states rarely flip the final action, whereas late decision-position interventions produce substantially larger effects with recipient-dependent asymmetry. Finally, the same primitive transfers from a controlled setting to native tool registries and banking-event histories, including sharp count boundaries in xLAM and source-disjoint τ-Banking workflows. Together, the results connect formal expressivity, empirical acquisition, internal mechanism, and downstream behavioral control.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.