From Hallucinated Targets to Runaway Repetition: Termination Control in Long Structured Generation
Abstract
Long structured generation requires models to learn both what information to produce and when to terminate. Data curation evaluates item-level correctness but overlooks how supervision content shapes stopping behavior. We study this in long-document relation extraction, where task-relative hallucinated targets contain entities outside an explicit candidate list. Across Qwen3 models at 1.7B, 4B, and 8B parameters, removing these targets reduces continuous pattern repetition by 24.65–34.63 percentage points and improves extraction quality. Replacing 15% of cleaned blocks one for one with real violating blocks raises 4B repetition from 5.52% to 9.71%, preserving inputs, block counts, and termination targets. Process decomposition shows that overall repetition can rise even as local copy gain falls. On identical prefixes at natural stopping points, raw-supervision models close their lists about 32 percentage points less often, yet retain near-certain EOS execution after a forced close. A single activation pulse along the cleaned-minus-raw direction shifts closing decisions bidirectionally, exceeds equal-norm random controls, and transfers to the replacement-trained model. At natural stopping points, it also reduces subsequent repetition in the raw 4B model by 10.94 percentage points. These results trace a specific supervision defect through reuse dynamics to an internal termination control component, establishing learned stopping behavior as a criterion for structured-data quality.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.