acceptodds
Under review as a conference paper at ICLR 2027

Closing the Perception-Action Loop with Human-Aligned Vision-Language Models

Abstract

Vision-language models (VLMs) have emerged as model systems for studying human visual cognition. However, their visual capabilities are primarily based on processing high-resolution images of full scenes, which is in sharp contrast to how humans process visual information. Human vision is inherently dynamic, relying heavily on eye movements to bring the high-resolution fovea onto different parts of a scene, so each fixation provides only a partial view. Here, we asked whether VLMs under similar partial observation constraints could efficiently choose where to look and use the information they acquired to solve visual question answering (VQA) and visual search (VS) tasks, and whether their behavior would resemble that of human participants. We found that, unlike humans, VLMs made poor use of successive fixations; their performance plateaued after the first few glimpses, and they frequently chose to answer prematurely. Fine-tuning on human fixation trajectories not only helped the models ground their answers in the available visual information across fixations, but also made their internal representations more predictive of neural responses in the primate brain while giving rise to cross-fixation evidence accumulation reminiscent of signals observed in the primate parietal cortex. Altogether, we show that pretrained VLMs largely lack the ability to engage in active vision, but that targeted fine-tuning on human behavior gives them a human-like perception–action loop and brings their representations into closer alignment with the brain. These properties make them strong candidates for studying the neural computations that underpin active vision in humans and non-human primates.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.