Guiding LLM Agents with Model-Internal Uncertainty for Scientific Decision-Making
Abstract
Scientific experiments can be expensive and slow to return feedback, making uncertainty important for deciding which actions to take under a limited evaluation budget. For large language model agents, however, it remains unclear whether internal representations provide useful uncertainty signals when labeled scientific interactions are scarce. We investigate this question using a frozen Qwen3.5-4B agent in Materials Discovery Environments (MADE) and DiscoveryWorld. Our approach fits a lightweight logistic predictor offline to estimate action-failure risk from public-state and output features augmented with internal activation summaries. These summaries are extracted at the prediction positions of generated action tokens and used to score candidate actions before execution. Four experimental conditions separate single-proposal planning, two-proposal generation, public-feature risk selection, and selection augmented with internal features. Exploratory development comparisons that preserve the public-feature transformation show improved failure ranking: AUROC increases from 0.5467 to 0.5622 in MADE and from 0.9667 to 0.9867 in DiscoveryWorld Proteomics. Online evaluations identify task-specific improvements in DiscoveryWorld, while analyses of matched actions and shared candidate pools distinguish predictive information from changes in action selection. The findings support internal activations as an additional source of failure-risk information in the tested settings, while showing that improved prediction alone does not ensure better scientific outcomes.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.