acceptodds
Under review as a conference paper at ICLR 2027

PhyAI: A Unified Inference Runtime for Robot Policies Across Edge and Cloud

Abstract

Recent breakthroughs in foundational models have driven significant progress in Robot Learning Models (RLMs), spanning Vision-Language-Action (VLA) and World Action Model (WAM) architectures. However, unlike the mature serving ecosystems for LLMs, RLM inference infrastructure remains highly fragmented across diverse deployment scales, from batch-1 real-time device control to massive batch-N edge serving, evaluationa and RL post-training rollouts. We present PhyAI, a unified, high-performance inference runtime for RLMs. By cleanly decoupling model semantics from execution, PhyAI dynamically adapts its parallelism (e.g., unbatched, DP, TP) and kernel dispatch strategies to the target hardware, seamlessly bridging edge and cloud environments. Through system-level optimizations including CUDA graph replay, kernel fusion, and state reuse, PhyAI achieves 1.40x to 4.65x latency speedups over official implementations across 11 benchmark configurations. To investigate how these inference speedups translate to physical control frequencies, we introduce the Control Rate Roofline model, demonstrating that our optimizations either directly boost control rates or unlock time margins for deploying larger models.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.