Ontology Concept Bottleneck Models: Towards Human-Aligned Reasoning with Graded Logic
Abstract
As autonomous systems increasingly make consequential decisions on behalf of humans, interpretability alone is insufficient. Their reasoning should not only be explainable but also align with the way humans evaluate evidence and make decisions. This motivates reasoning over intuitive concepts through concise and human-readable compositional processes. We introduce Ontology Concept Bottleneck Models (OCBMs), which jointly learn task-functional concepts and class-specific graded-logic reasoning trees, leveraging the reasoning structure itself as an inductive bias on concept formation. Across standard visual and tabular benchmarks, OCBMs achieve competitive or leading predictive performance while producing compact, auditable reasoning structures with controllable granularity. Moreover, controlled head ablations show that these concepts transfer substantially better across domains than those learned with linear or MLP heads. These results suggest that graded-logic ontologies can serve simultaneously as an interpretable reasoning layer and a structural constraint on concept learning, providing a practical foundation for machine reasoning that can be inspected, intervened upon, and progressively aligned with human reasoning.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.