OmniZPD: Token Compression via Policy Optimization and Self-Distillation for Omni-Modal Reasoning
Abstract
Omni-modal large language models (OmniLLMs) suffer from high inference costs due to dense audio-visual token redundancy. Existing token compression methods often discard reasoning-critical evidence, as sparse final-answer rewards fail to supervise intermediate thinking processes. We propose OmniZPD, an efficient compression framework integrating policy optimization and self-distillation. OmniZPD features a unified, ratio-conditioned selector that jointly evaluates audio-visual token importance based on queries, compression budgets, and local cross-modal contexts. Keeping the backbone frozen, we optimize the selector using a bounded reward that combines task correctness with self-distillation feedback, enforcing consistent reasoning predictions between full and compressed inputs. At inference, a single trained selector deterministically prunes tokens and flexibly adapts to varying retention budgets without retraining. Experiments across multiple omni-modal reasoning benchmarks show that OmniZPD significantly outperforms existing baselines, achieving a 2.79× prefill speedup while retaining 91.8% of the full-input accuracy with merely 20% of audio-visual tokens.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.