AI Research Should Scale Insight, Not Just Compute
Abstract
Scaling compute has driven substantial progress in AI, but performance gains do not necessarily bring a corresponding increase in reusable knowledge about learning. We argue that AI research should scale insight, not just compute. We define insight as reusable, communicable knowledge with a stated scope of applicability. Systematically producing, testing, and accumulating such knowledge could deepen our understanding of learning and guide more effective training, supporting both the science of AI and AI-for-AI research. We explore this approach by actively running experiments and measuring observables to identify learning phenomena, formulate hypotheses, and test interventions. We present OPHIS (Observation–Problem–Hypothesis–Intervention–Speed-up), a conceptual framework for observation-guided research; ComfyResearch, a visual tool for constructing, reproducing, and sharing experiments; and the nanoGPT Observable Library, a dataset of observable trajectories recorded during training. Using the nanoGPT Observable Library, we observe attention-entropy patterns that differ across depth and warmup settings. Applying OPHIS to AutoResearch, we identify interventions that improve validation performance under a fixed training-step budget. We demonstrate ComfyResearch on grokking and edge-of-stability experiments and show how automated experiment generation supports the exploration of learning dynamics. We outline a research agenda for deepening human understanding of learning, accumulating and sharing insights, applying them in AutoResearch, and assessing their value under compute and time constraints.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.