Certifying Culturally Faithful Generation: Feasibility, Ordering, and Release Gates Beyond Calibration
Abstract
Text-to-image systems deployed in cultural heritage settings can produce images that are visually plausible yet culturally unfaithful, substituting generic or neighboring traditions for the requested one. A common mitigation is selective generation: sample several candidates, score them with a verifier, and emit the highestscoring one. But what evidence justifies releasing such a system? Risk-control methods can calibrate a frozen score, yet release has two logically prior failure modes: the candidate pool may contain no culturally acceptable output, or an acceptable output may exist but the verifier ranks another candidate above it. We formalize this distinction with request-level estimands—infeasibility I, mixedpool mass M, and within-mixed-pool selection success γ—which obey the exact decomposition Rsel = I +M(1−γ). The identity is elementary; its value is attributional: the two actionable sources of risk become separately measurable. On top of this attribution layer we define a frozen-policy certification protocol: the operating point is chosen on development data only, the policy is frozen, and release is accepted if and only if an exact one-shot upper bound on the total selected-failure count x = f + o meets the target; Bonferroni component lower bounds only attribute rejections. Common release metrics do not identify the certified quantity: with four candidates, 80.5% pairwise accuracy is compatible with a top-one hit rate from 0% to 100%, and flattening into per-request risk control cannot recover the decomposition. On CulturalFrames, at the prespecified endpoint, the frozen SigLIP verifier’s ordering-error component has a one-sided 95% lower bound of 15.5%—above the full 10% risk budget—and a bounded Qwen2.5-VL repair fails its one-shot held-out gate. LiveCodeBench fails feasibility (I ˆ = 0.292 > 0.10); three benchmark rows pass, including a RewardBench certificate with exact upper risk bound 2.83%. Synthetic validation confirms the gates open on non-degenerate K > 2 pools, and a risk–coverage frontier view certifies coverage ≈ 0.90 at target 0.10. Our cultural audit is a negative, endpoint-conditional diagnosis; we do not claim a benchmark-native positive certificate for non-degenerate cultural pools.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.