TAIWO: Structured Test-Time Inference for Open-Ended Semantic Discovery Across Large Language Models
Abstract
Open-ended semantic discovery with large language models (LLMs) is increasingly performed through multi-step prompting, yet comparatively little is known about how reliably a fixed structured workflow transfers across model families or where it fails as concepts become less prevalent. We introduce TAIWO, a six-stage test-time inference pipeline that transforms source documents through familiarization, coding, theme construction, review, definition, and synthesis while retaining source identifiers. We operationalize TAIWO as a frozen benchmark on a deterministic 100-document corpus with nine latent concepts in three prevalence strata. Across 39 model runs, the task-consistent mean score is 0.670 (median 0.737; range 0–0.883); 36 runs complete all structured stages, while three return task-declared zeroes after malformed output. Within the 36 detailed outputs, dominant-concept recall averages 0.907 but rare-concept recall averages 0.481. A provenance audit recovers the exact task source and validates its hash, but also finds two consequential measurement limitations: an exporter treated the three zeroes as missing, and the score's “grounding” term verifies identifier existence rather than evidential support. Exact concept labels paired with deliberately wrong but valid source identifiers can therefore score 1.0. These results provide a transparent cross-model stress test of structured open-ended inference, identify rare concept recovery and schema compliance as concrete failure modes, and show why benchmark conclusions must include scorer-level adversarial audits. In an exploratory three-model compute-matched pilot, TAIWO's mean paired advantage over generic six-call decomposition is 0.0068 (95% model-clustered bootstrap CI ), providing no evidence of a pipeline advantage. A separate transfer pilot on 100 human-labelled Reddit mental-health posts likewise provides no evidence of an advantage: source-grounded category recall is 0.289 for TAIWO and 0.413 for generic decomposition (paired difference , 95% CI ).
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.