acceptodds
Under review as a conference paper at ICLR 2027

Optimizing Pandas Code for Heterogeneous Backends

Abstract

Data scientists rely on pandas for building data-processing workflows due to its expressive API and rich ecosystem, but Python’s Global Interpreter Lock and single-threaded execution model make pandas a performance bottleneck at scale. We propose PandaX, a system for automatic optimization and heterogeneous scheduling of pandas workflows across CPU and GPU backends. PandaX uses a multi-agent LLM-based rewriting framework to transform inefficient pandas code into backend-targeted, high-performance equivalents, guided by profiling feedback, correctness validation, and formal equivalence checking. It then applies a learned cost model and a dynamic programming scheduler to assign each code block to the backend that minimizes end-to-end execution time, explicitly accounting for computation and data-transfer costs. To make optimization practical, PandaX performs rewriting and scheduling using small data samples and extrapolates performance to full datasets via learned execution and transfer models. Experimental results on 25 real-world data-science workflows and 22 TPC-H queries show that PandaX achieves up to speedup over pandas on real workloads and on TPC-H, significantly outperforming existing GPU-accelerated pandas frameworks and bridging the gap between ease of use and high-performance heterogeneous execution.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.