acceptodds
Under review as a conference paper at ICLR 2027

FlashRT: Agent Harness for Guiding Agents to Deploy Real-Time Multimodal Applications

Abstract

Real-time multimodal applications, such as voice agents and interactive video generation, compose heterogeneous models into pipelines whose efficient deployment requires application-specific decisions about placement, streaming, and intra-model parallelism. Existing serving systems and auto-parallelism compilers commit to limited transformations and fixed workload assumptions, so each new application requires a hand-crafted implementation to achieve high performance. We present FlashRT, an agent harness that guides a generic coding agent to lift a simple developer-written reference implementation into optimized multi-GPU deployments that flexibly weigh target metrics such as latency and throughput. Following a new chain-of-program paradigm, FlashRT directs the agent to convert the reference into an intermediate representation (IR) that captures data dependencies and persistent state, validate the IR with a sequential interpreter, and statically analyze it to identify candidate transformations, which the agent then implements, verifies, and benchmarks in a measurement-gated optimization loop. Across five applications, including video world models and multimodal LLMs, FlashRT delivers up to 70 latency reduction and 2.8 throughput improvement on NVIDIA B200 GPUs and up to 3.6 throughput improvement on AMD MI355X GPUs, and it reduces Qwen3-Omni's response latency by up to 65% relative to the expert-built vLLM-Omni deployment.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.