acceptodds
Under review as a conference paper at ICLR 2027

Careful Judge: Safe and Efficient Human–AI Collaborative Decision Making

Abstract

In human–AI collaborative decision making, human review can prevent unsafe AI decisions, but each human judgment is costly. Treating human intervention as a one-off fallback misses the opportunity to improve future AI decisions, yet adaptively learning from human feedback changes the model and can invalidate previously calibrated safety guardrails. We approach this challenge with CARE—**c**alibrated **a**daptive **r**ectification and **e**scalation—an end-to-end pipeline that combines AI models and human reviewers to guarantee safe, human-aligned decisions, while continuously learning from human feedback to achieve greater automation with fewer human queries. CARE is principled, general, modular, and works with any black-box AI model. We equip it with a rectification module that improves AI decisions by learning the structure of human–AI misalignment. Our novel adaptive calibration module guarantees risk control at every time step for any rectification module. Experiments on four safety-critical real-world datasets spanning language, vision, driving, and robotics demonstrate that CARE achieves human-aligned decisions while reducing human queries by 25–81% relative to baselines.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.