acceptodds
Under review as a conference paper at ICLR 2027

Toward Open Weight Models Without Risks: Separating Public and Private Capabilities in LLMs

Abstract

Open-weight Large Language Models (LLMs) enable scientific progress and broad deployment but complicate access control over sensitive capabilities. Current practice either suppresses dangerous capabilities before release or mediates access through closed services. The former is susceptible to jailbreaks while sacrificing capability for all users to mitigate the risks posed by a few, and the latter is incompatible with open-weight release. In this paper, we propose *Tiered Language Models* (TLMs), where a single set of released weights supports distinct capability tiers. In its default public configuration, a TLM behaves as a conventional LLM. A compact secret key specifies a permutation over a small parameter subset, inducing an alternative computation graph over the same weights that exposes additional capabilities. Our training protocol jointly pretrains both configurations from scratch, then fine-tunes the keyed configuration on private data with regularization to preserve the public model's behavior. We evaluate M- and M-parameter TLMs on language acquisition, instruction following, and factual memorization, using these tasks as controlled proxies for private or sensitive capabilities. The keyed configuration successfully learns each task, while the same weights in their public configuration do not exhibit the acquired capabilities. We show that this approach extends naturally to multiple hierarchical tiers. Because authorization operates on the model's weight structure rather than in the input space, the mechanism resists fine-tuning-based extraction and partial key compromise. In general, TLMs take a step toward reconciling open-weight release with selective capability control.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.