acceptodds
Under review as a conference paper at ICLR 2027

Human Modeling Meets AI Metacognition: Toward Coupled Cognitive Alignment for Interactive Agents

Abstract

AI agents increasingly participate in complex human workflows, where failures arise not only from incorrect outputs but also from inappropriate interactions. An agent may answer when it should ask for clarification, personalize despite high uncertainty about the user, or express confidence in ways that distort human trust and reliance. We argue that these failures stem from inaccurate human modeling that aims to infer humans' latent cognitive states, and unreliable AI metacognition which estimates the agent's own uncertainty, capability boundaries, and risk, and more importantly, misalignment between the two in the interaction policy. A primary bottleneck is that the latent states of humans and AI have to be inferred from partial observations. This position paper advocates Coupled Cognitive Alignment (CCA), a framework that unifies human modeling and AI metacognition as a coupled problem of interaction-level decision making under uncertainty. In CCA, agents maintain uncertainty-aware dynamics models of human states while simultaneously reasoning about their own uncertainty and capability limits. These coupled estimates determine not only what the agent outputs, but whether it should answer, ask, explain, retrieve, decompose tasks, defer, or escalate. We formalize it as a partially observable interactive system, identify its key research challenges, and discuss potential strategies for its learning, evaluation, and deployment.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.