acceptodds
Under review as a conference paper at ICLR 2027

OS-Bench: Evaluating SWE Agents on Cross-OS Code Migration

Abstract

Language-model (LM) agents are increasingly evaluated on repository-level software engineering (SWE) benchmarks derived from GitHub issues and pull requests. Although these benchmarks cover a broad range of development tasks, they provide limited insight into a common yet technically distinct challenge: migrating an existing implementation from one operating system (OS) to another while preserving behavior on both. We introduce OS-Bench, an executable benchmark for cross-OS code migration between Linux and Windows. Motivated by case studies showing that cross-platform features are frequently implemented on one OS before support is added for another, OS-Bench identifies migration opportunities by comparing execution coverage collected from the same repository on both systems. Platform-specific regions are then verified, organized into five migration scenarios, and transformed into repository-level tasks with target-platform tests and source-platform regression tests. The resulting benchmark contains 140 validated tasks from 111 repositories across C, C++, C#, Go, and Rust, comprising 91 Windows-to-Linux and 49 Linux-to-Windows migrations. We evaluate six models with SWE-agent, Claude Code, OpenHands, and Win-agent. The best configurations achieve only 26.4% success on Linux-target tasks and 20.4% on Windows-target tasks while preserving correctness across both systems. These results indicate that cross-OS migration is a challenging and underexplored capability, and OS-Bench is a testbed for improving platform-aware SWE agents.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.