FlashRT: Agent Harness for Guiding Agents to Deploy Real-Time Multimodal Applications
Abstract
Real-time multimodal applications, such as voice agents and interactive video generation, compose heterogeneous models into pipelines whose efficient deployment requires application-specific decisions about placement, streaming, and intra-model parallelism. Existing serving systems and auto-parallelism compilers commit to limited transformations and fixed workload assumptions, so each new application requires a hand-crafted implementation to achieve high performance. We present FlashRT, an agent harness that guides a generic coding agent to lift a simple developer-written reference implementation into optimized multi-GPU deployments that flexibly weigh target metrics such as latency and throughput. Following a new chain-of-program paradigm, FlashRT directs the agent to convert the reference into an intermediate representation (IR) that captures data dependencies and persistent state, validate the IR with a sequential interpreter, and statically analyze it to identify candidate transformations, which the agent then implements, verifies, and benchmarks in a measurement-gated optimization loop. Across five applications, including video world models and multimodal LLMs, FlashRT delivers up to 70 latency reduction and 2.8 throughput improvement on NVIDIA B200 GPUs and up to 3.6 throughput improvement on AMD MI355X GPUs, and it reduces Qwen3-Omni's response latency by up to 65% relative to the expert-built vLLM-Omni deployment.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.