acceptodds
Under review as a conference paper at ICLR 2027

Talking to a Connectome: Language-Model Agents Teach and Interrogate a Whole-Brain Fly Model

Abstract

Whole-brain connectome models make it possible to run behavioural experiments in silico. We let language-model agents experiment on such a model as an experimenter would: presenting odours, pairing them with sugar or shock, and measuring the fly’s approach valence. The model is a leaky integrate-and-fire network of the adult Drosophila brain on the FlyWire connectome (138,639 neurons) with dopamine-gated plasticity at Kenyon-cell output synapses. Manipulating the model shows that an odour recruits its own dopamine through Kenyon-cell-to-dopamineneuron synapses, so that presenting it alone changes its memory, that learning changes this dopamine, and that a memory cannot be undone to within tolerance. In open-loop teaching, one pairing per target meets 0.80 of targets (doing nothing meets 0.71), a zero-shot frontier model writes nearly the same protocols, and teachers optimised on a fast surrogate of the fly exploit the surrogate’s errors (0.99 predicted, 0.75 measured). In closed-loop diagnosis of a hidden training history from twelve measurements, the main difficulty is generalisation between odours that share Kenyon cells. Claude Sonnet 5 scores 0.832 against 0.793 for a thresholding rule because it respects the stated prior that at most two odours were trained; the rule with that prior scores 0.846. Given a one-page summary of the model’s noise and generalisation between odours, and told how answers are scored, Claude reaches 0.894, level with a Bayes player that knows the same (0.912), and gets all seven labels right in 73% of flies instead of 35%. Qwen3-32B cannot use the summary, although thirty steps of multi-turn GRPO raise it from 0.765 to 0.830, Claude’s zero-shot level. Language-model representations of odour names predict

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.