ChameleonBench: Evaluating the Cross-Environment Robustness Across Heterogeneous Systems
Abstract
Language models increasingly automate programming and desktop workflows, yet success in one environment leaves open whether the same tasks can be completed under different execution conditions. We use environment broadly to denote the external conditions governing observations, available actions, and execution outcomes, including operating systems, software versions, configurations, permissions, hardware resources, network conditions, and runtime state. We present ChameleonBench, a benchmark of 100 tasks—80 Coding and 20 GUI—constructed around native execution behaviors that vary across environments. We focus on native Linux, macOS, and Windows deployments, which provide differences in filesystem, runtime, and application behavior while supporting equivalent task objectives and success criteria. The tasks form 30 mechanism families, with Coding covering individual mechanisms and their combinations. Using a fixed interaction harness and independent verification, we evaluate six models once per task and deployment, yielding 1,800 runs. With equal weighting across families, the strongest model achieves 75.2% mean success across deployments but 57.9% Robust Success Rate, which requires completing the same task on all three. The other models achieve 7.3–23.0% robust success. Grouping by mechanism family also yields more accurate predictions of performance differences across operating systems, on average, than grouping by interaction type and target OS alone. Our results show why task completion and its preservation across environments should be evaluated together.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.