Human Modeling Meets AI Metacognition: Toward Coupled Cognitive Alignment for Interactive Agents
Abstract
AI agents increasingly participate in complex human workflows, where failures arise not only from incorrect outputs but also from inappropriate interactions. An agent may answer when it should ask for clarification, personalize despite high uncertainty about the user, or express confidence in ways that distort human trust and reliance. We argue that these failures stem from inaccurate human modeling that aims to infer humans' latent cognitive states, and unreliable AI metacognition which estimates the agent's own uncertainty, capability boundaries, and risk, and more importantly, misalignment between the two in the interaction policy. A primary bottleneck is that the latent states of humans and AI have to be inferred from partial observations. This position paper advocates Coupled Cognitive Alignment (CCA), a framework that unifies human modeling and AI metacognition as a coupled problem of interaction-level decision making under uncertainty. In CCA, agents maintain uncertainty-aware dynamics models of human states while simultaneously reasoning about their own uncertainty and capability limits. These coupled estimates determine not only what the agent outputs, but whether it should answer, ask, explain, retrieve, decompose tasks, defer, or escalate. We formalize it as a partially observable interactive system, identify its key research challenges, and discuss potential strategies for its learning, evaluation, and deployment.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.