Demonstration Conditioning Generally Improves Jagged Frontier Model Performance on Digital and Physical Agentic Tasks
Abstract
We measure how much an in-context demonstration improves frontier models, and how much of that improvement comes from the demonstration containing the an- swer. Benchmark averages hide both: they show neither where inside a benchmark an agent is weak nor whether a demonstration taught a procedure or supplied the item’s answer. Holding the prompt fixed, we vary only the demonstration for seven models, including GPT Astra and Claude Opus 5.5, on tasks across web automation, manipulation, assembly, industrial logistics, and self-driving. A demonstration from the same situation raises the score in all five domains and for all seven web models. With the answer removed, the gain survives on driving and manipulation for all three models and disappears on assembly. Across the seven models the gain falls as the baseline rises, narrowing the gap between the best and worst model while the unevenness inside each benchmark remains.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.