CBR-VERIFY: BUDGETED EVIDENCE REVISIONFOR REUSABLE SCIENTIFIC AGENT WORKFLOW
Abstract
Scientific agents must revise the evidence supporting their conclusions when mea-surements, assumptions or evaluation protocols change. CBR-VERIFY formulatesthis decision as budgeted claim completion. Executable predicates define decisions,version-bound obligations define required evidence, and closure selection allocatescomputation between archive validation and conditional repair. The completion ob-jective captures complementary prerequisites and shared acquisition costs; boundedseed enumeration ranks ready actions by covering density. Across 12 paired GPTtasks, full CBR improves observed utility over Plan–Verify by 9.52 percentagepoints while using 7.85% fewer reported tokens and 12.13% more scientific credits.A broader frozen-task replication, repeated executions, and an additional modelpreserve the main advantage. The replicated mechanism comparison gains 15.48utility points over immediate-completion greedy while saving 4.35 scientific creditsper task. Structural sweeps identify stable pair-expansion gains under high sharing,no gain under low sharing, and reproduction on tasks excluded from tuning. Thesame controller remains beneficial against caching, dependency-aware, cost-awarecontrols without retuning on new test outcomes. Real adapter executions validateevidence invalidation and revision; the real-execution comparison also saves ap-proximately 17% of simulation-tool execution time. Independent replay verifieskey trajectories and outcomes.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.