Live Skill Evolution for Forecasting Agents with Delayed Outcome Feedback
Abstract
Human forecasters can improve their predictions by reflecting on and learning from past experience. However, for forecasting agents, systematically accumulating such experience and distilling it into reusable forecasting skills remains challenging. To address this challenge, we introduce Future-SE, a live skill-evolution framework that updates an explicit natural-language skill library while keeping the base model and agent harness fixed. The library separates thinking skills for judgment from working skills for evidence acquisition. At prediction time, Future-SE records matched skill-enabled and no-skill trajectory groups. Once outcomes resolve, it uses outcome feedback, prediction-time market baselines, and trace evidence to refine the library. Experiments with DeepSeek-V4-Flash, GLM-5.2, and Kimi-K2.7-Code on two dynamic forecasting tracks and FutureX show that evolved skills improve forecasting performance over matched no-skill agents. Role ablations show larger gains from thinking-only than from working-only configurations. Analysis of the evolved libraries and their revision histories identifies reusable procedures for question interpretation, prior construction, and structured evidence preparation, while revealing conditional evidence weighting as a persistent difficulty.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.