Mind the Modality Gap: From Constitutional AI to Contextual Policy-Alignment Through Language Bridging
Abstract
Most generative-model alignment methods encode a fixed notion of harmfulness, whereas deployed systems must often satisfy policies specific to a platform, application, or content domain. We introduce AnyCAI, an end-to-end framework that extends Constitutional AI beyond text through a language bridge. Given a natural-language constitution and seed requests, a generator produces an artifact, a modality-capable critic evaluates that artifact without seeing the generating prompt, and a text-only reviser converts the critique into a corrected request. The resulting artifact-grounded pairs adapt an accessible conditioning encoder while the generative backbone remains frozen. AnyCAI generalizes the target from safety to policy compliance: conventional harmlessness is recovered when the constitution specifies harmful content, while other constitutions can express deployer-specific requirements. We pair Policy Compliance Rate with semantic fidelity on policy-compatible requests and evaluate conventional safety, a selective policy relaxation, a marketplace policy, and policy composition across text-to-image, text-to-video, vision-language, and text-to-music settings. The reusable contribution is the complete constitution-to-data-to-alignment pipeline; its supervision pairs, adapters, and evaluation distributions remain policy-specific.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.