acceptodds
Under review as a conference paper at ICLR 2027

LunarZero: A Single-Binary Coding Agent that Runs on Pooled Free-Tier Language Models with Quota-Aware Failover

Abstract

Agentic software engineering systems are commonly evaluated under the assumption that their language model is continuously available. This assumption excludes an important deployment setting: users and institutions that cannot afford sustained access to premium APIs or the infrastructure required to self-host capable models. Such users often depend on free or otherwise rate-limited access, where individual models have limited request and token budgets, availability changes over time, and failures may prevent an agent from completing a task even when alternative models remain accessible. We study this setting empirically. We build LunarZero, a terminal coding agent that treats models from multiple providers as a shared pool and maintains a persistent per-model ledger of request and token usage, failure-class cooldowns, and pre-streaming failover, and we evaluate it on all 300 tasks of SWE-bench Lite using the official instance images and evaluator. Across 886 evaluator-scored runs the pooled configuration resolves of tasks at zero model cost, using 98 distinct models from ten providers. We compare it against a matched control that uses the same agent and protocol pinned to the model ranked highest by the router's quality metric and used most by the pooled system. On the 186 tasks attempted by both, the pooled system resolves against for the control, a paired difference of percentage points ( CI ; exact McNemar, ): under the evaluated configuration we find no statistically detectable difference in task resolution. The configurations nevertheless exhibit substantially different failure modes. The pooled system abandons tasks without producing a patch far more frequently ( pp, ), whereas the control more frequently exhausts its time budget once its provider becomes unavailable ( pp, ); the pooled system also completes tasks faster and sustains far more total model usage than any individual free tier. Ablations show the hand-designed cooldown policy to be conservative: uniform exponential backoff increases resolution by pp on its paired subset () at roughly the request volume. We characterize 47,406 failover events, of which over a quarter arise from catalogue drift rather than rate limiting. These results characterize model availability as a distinct systems constraint: pooling rate-limited models can substantially increase aggregate availability and throughput without necessarily improving task resolution. We specify the routing policy and evaluation protocol in full, and report the accounting behind every claim, so that the study can be replicated independently.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.