acceptodds
Under review as a conference paper at ICLR 2027

Mind the Modality Gap: From Constitutional AI to Contextual Policy-Alignment Through Language Bridging

Abstract

Most generative-model alignment methods encode a fixed notion of harmfulness, whereas deployed systems must often satisfy policies specific to a platform, application, or content domain. We introduce AnyCAI, an end-to-end framework that extends Constitutional AI beyond text through a language bridge. Given a natural-language constitution and seed requests, a generator produces an artifact, a modality-capable critic evaluates that artifact without seeing the generating prompt, and a text-only reviser converts the critique into a corrected request. The resulting artifact-grounded pairs adapt an accessible conditioning encoder while the generative backbone remains frozen. AnyCAI generalizes the target from safety to policy compliance: conventional harmlessness is recovered when the constitution specifies harmful content, while other constitutions can express deployer-specific requirements. We pair Policy Compliance Rate with semantic fidelity on policy-compatible requests and evaluate conventional safety, a selective policy relaxation, a marketplace policy, and policy composition across text-to-image, text-to-video, vision-language, and text-to-music settings. The reusable contribution is the complete constitution-to-data-to-alignment pipeline; its supervision pairs, adapters, and evaluation distributions remain policy-specific.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.