Exploring Consciousness in LLMs through IIT-Inspired Reward Learning
Abstract
Consciousness-like processing has been hypothesized to play a role in achieving Artificial General Intelligence (AGI). However, despite numerous studies seeking to assess consciousness in LLMs from the perspectives of various theories of consciousness, comparatively little attention has been devoted to evaluating models after explicitly training them to enhance consciousness-related properties, largely due to the inherent complexity of consciousness. Here, we adopt Integrated Information Theory(IIT) as a quantitative framework for assessing consciousness in a system. This allows maximization of such metrics through Reinforcement Learning(RL). In this theory, a system is viewed as causal evolution of states, and the theory provides a formal, axiom-based mathematical framework for quantifying consciousness. Inspired by its postulates, we formulate three novel reward functions to measure some or all properties associated with consciousness. Our results show that the intrinsic information () reward enables substantially more concise reasoning, reducing response length by up to 70% while maintaining comparable accuracy and improving reasoning faithfulness, consistent with the Information Postulate. On the other hand, training with the -reward, an IIT-based measure that incorporates all IIT postulates, did not yield strong empirical evidence for consciousness. Nevertheless, the proposed framework provides a basis for more concrete empirical investigation of the open question of whether LLMs can exhibit consciousness.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.