LicenseHarness: Preserving Evidence Boundaries in Autonomous Research Writing
Abstract
Large language model research agents always suffer from defensive writing, cluttering generated papers with redundant disclaimers. Existing prompt engineering and rule-based linters fail across long interactions because debugging histories mix factual updates with conversational pressure, inevitably causing defensive relapse. We formalize defensive writing as unlicensed hedges that empirical evidence does not support, proving that removing defensiveness while preserving legitimate scientific boundaries is a unified constrained optimization problem. We introduce **LicenseHarness**, a framework that externalizes evidence into a typed state to block narrative pressure, combined with history invariance distillation and constraint-aware reinforcement learning that internalizes licensing directly into the model weights. Evaluating across 210 empirical machine learning bundles, four base models and production agents, **LicenseHarness** reduces conversational drift by 79 times while maintaining 93.8% epistemic boundary coverage. Double-blind human reviewers consistently prefer its outputs in persuasiveness and acceptance intent.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.