acceptodds
Under review as a conference paper at ICLR 2027

Refuse Without Abandoning: Circuit Coordination for Safe and Helpful Medical LLMs

Abstract

An ideal medical large language model should refuse harmful requests while preserving useful answers to legitimate needs. Yet existing models tend to achieve only one of these goals at a time. This paper systematically studies the safety-helpfulness imbalance in medical LLMs and identifies two recurring failure modes: models frequently discard information that could be safely provided when refusing harmful content, and safety warnings largely disappear when requests are rewritten as role-play or indirect prompts, with some rewritten requests receiving outright harmful responses. To quantify these behaviors, we propose HELP and CAUTION, two seven-dimensional rubrics that measure useful content and safety guidance respectively. Building on this analysis, we introduce MedCircuit, a method grounded in the observation that models already possess both behaviors but fail to combine them on sensitive inputs. MedCircuit extracts activation difference vectors corresponding to each behavior, constructs two fixed circuits, and trains a lightweight coordinator to jointly control their activation strengths, enabling the model to produce both helpful content and appropriate safety guidance during generation. Experiments across three medical LLMs and two benchmarks show that MedCircuit is the only defense that improves both safety and helpfulness simultaneously across all models, and maintains consistent advantages under role-play and indirect prompt rewriting attacks.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.