Regularized Auto-Research with Causal Harness
Abstract
Auto-research agents design, run, and learn from ML experiments on their own, and each experiment costs a training run. We find that such agents overfit to their own experimental history in two ways. On our primary task an unassisted agent locks onto its incumbent, spending half of each wave on one-parameter perturbations of it, and sessions given the same instructions end far apart. Given a correlational dashboard, it reads parameter rankings that its own co-modifications confound, and the rankings degrade as the session proceeds. We introduce CausalHarness, a harness that treats every edit as an intervention and regularizes the search with a Bayesian posterior over causal structure and mechanisms, organized in three layers: an LLM prior sets where the posterior starts, the agent's interventions update it, and the posterior returns to the agent as an information-gain ranking of candidate experiments that the agent is free to ignore. In a study in which agents are not told their condition, with multiple independent sessions of about 300 experiments per condition on up to 64 GPUs, the worst CausalHarness session ends with a lower convergence AUC than the median session of every baseline in all four settings. On the primary task every CausalHarness session converges faster than every session of a matched correlational dashboard (Hedges' , exact ), reaches the target in every session, and beats the unassisted agent. At matched cadence every session also beats every run of a trust-region Bayesian optimizer, and removing any one layer significantly weakens the advantage. The gains hold on two further architectures run with a second agent LLM, and the selected configurations keep their lead under full-length pretraining.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.