FlowDVFS: Outperforming Generative Replay for Few-Shot DVFS Control
Abstract
Few-shot reinforcement learning for dynamic voltage and frequency scaling (DVFS) must find, from a few dozen real transitions, the clock giving the highest stable frame rate without throttling. We propose FlowDVFS, a flow-matching replay generator for a few-shot Double DQN: a physics anchor backed up at every clock, plus a flow residual. We judge the generator by a learning curve over 8 to 32 tuples on a throughput reward whose thermal target never binds in the replayed runs. On held-out traces of a Jetson Orin NX and TX2, at MBPO's twenty updates per real transition, FlowDVFS needs fewer real tuples than SynthER, PGR (one update) and MBPO and is above all three at every budget, tested at each against MBPO. Neither replay method cuts the learner's real budget; FlowDVFS does, its anchor's argmax with no learner matching it on development, while the top clock, best on every run, stays above every learned method.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.