acceptodds
Under review as a conference paper at ICLR 2027

ChameleonBench: Evaluating the Cross-Environment Robustness Across Heterogeneous Systems

Abstract

Language models increasingly automate programming and desktop workflows, yet success in one environment leaves open whether the same tasks can be completed under different execution conditions. We use environment broadly to denote the external conditions governing observations, available actions, and execution outcomes, including operating systems, software versions, configurations, permissions, hardware resources, network conditions, and runtime state. We present ChameleonBench, a benchmark of 100 tasks—80 Coding and 20 GUI—constructed around native execution behaviors that vary across environments. We focus on native Linux, macOS, and Windows deployments, which provide differences in filesystem, runtime, and application behavior while supporting equivalent task objectives and success criteria. The tasks form 30 mechanism families, with Coding covering individual mechanisms and their combinations. Using a fixed interaction harness and independent verification, we evaluate six models once per task and deployment, yielding 1,800 runs. With equal weighting across families, the strongest model achieves 75.2% mean success across deployments but 57.9% Robust Success Rate, which requires completing the same task on all three. The other models achieve 7.3–23.0% robust success. Grouping by mechanism family also yields more accurate predictions of performance differences across operating systems, on average, than grouping by interaction type and target OS alone. Our results show why task completion and its preservation across environments should be evaluated together.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.