acceptodds
Under review as a conference paper at ICLR 2027

Inverse-Scaling Signature in Agentic Cinematic Subtext Detection

Abstract

Agentic pipelines that combine retrieval, decomposition and critic-based refinement are widely assumed to improve as inference-time computation increases. On cinematic subtext detection, which requires pragmatic inference about what characters mean but do not say, the assumption fails. Using LitSubText, 2,606 dialogue instances from 132 film screenplays annotated with 8 intent and 9 theme classes, a single forward pass outperforms every agentic configuration tested. Holding the prompt and the tools fixed and varying only whether the ReAct loop runs, theme macro-F1 falls from 27.3 to 13.0 percent and theme accuracy from 34.2 to 11.7 percent, a significant gap under a screenplay-level cluster bootstrap. The loop breaks 89.0 percent of what the prompt-only configuration answered correctly while repairing 12.0 percent of what it got wrong, so it substitutes predictors rather than partly improving one. Three restraint mechanisms were added to test whether the loss can be recovered, namely Confidence-Gated Retrieval, a Bidirectional Critic that can recommend simplification as well as deepening, and an Interpretation Depth Controller that stops early on critic direction signals. They do not recover it. Instrumented trajectories show that the simplify direction, the only signal expressing restraint, fired zero times across 240 test instances and that the depth controller produced longer trajectories than the unrestrained loop, so two of the three never operated as specified. Probes that hold the input fixed and vary only the candidate locate the reason, because the critic separates wrong candidates from right ones far more sharply than over-detailed ones from concise ones. What it supplies is largely a correctness signal rather than a restraint signal.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.