Think Ontologically: Internalizing Neuro-Formal State Invariants for Robust Conversational Memory in Large Language Model Agents
Abstract
Large language model (LLM) agents operating in multi-turn dialogues suffer from severe conversational memory degradation: external context compression and prose summarization trigger continuous semantic drift, whereas detached symbolic reasoners impose non-differentiable foreign-function bottlenecks and rigid extraction barriers. In this work, we propose Think Ontologically, a neuro-formal reasoning framework that internalizes conversational state tracking directly into neural policy weights. Rather than offloading memory to external solvers or prose buffers, the agent autoregressively synthesizes first-order relational state invariants (causes, holds_value, has_state, has_goal) grounded in foundational ontologies (UFO and BFO) before generating conversational responses. We train foundation models using a two-stage curriculum: supervised warmup on verified relational tuples followed by Verifiable Reinforcement Fine-Tuning (RLVR) with deterministic multi-tiered reward signals. Investigating multi-reward credit assignment dynamics, we demonstrate that Group reward-Decoupled Normalization Policy Optimization (GDPO) and sequence-level importance sampling (GSPO) effectively eliminate multi-reward gradient masking, enabling models to satisfy deductive constraints while maintaining 100% syntax validity and minimal policy drift (). Across five open-weight model families (Qwen-2.5-14B/7B, Llama-3.1-8B, Gemma-3-4B, and DeepSeek-R1-Distill-7B), our internalized policy consistently outperforms zero-shot and in-context baselines, boosting IFEval strict instruction compliance by up to percentage points, elevating GPQA Diamond scientific reasoning to 29.80%, and achieving an 88.0% goal completion rate on -bench Airline interactions with a >75% reduction in inference latency compared to decoupled neuro-symbolic pipelines.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.