acceptodds
Under review as a conference paper at ICLR 2027

FidelityFlow-PO: Bandit-Guided Allocation of Context Fidelity for Long-Context DPO

Abstract

For long-context Direct Preference Optimization (DPO), the quality of preference pairs depends critically on the evidence available when candidate responses are generated. One adaptive strategy is to use response feedback to revise the context shown across candidate-generation rounds. A recent method, LongMab, instantiates this strategy by partitioning the long context into chunks and using a multi-armed bandit to adapt which chunks are shown during candidate-response generation. However, because decisions are made at the whole-chunk level, each chunk must either be included in full, potentially introducing irrelevant content, or be omitted entirely, potentially discarding critical bridging evidence. This couples source coverage with reading cost. We introduce FidelityFlow-PO, which advances LongMab's chunk-as-arm formulation from binary source selection to feedback-driven control of reading cost. The method preserves a compact, query-conditioned core from each chunk as a starting point for deeper reading, while response feedback adaptively controls how much surrounding context is included. This decouples source coverage from reading depth: access to a source no longer requires full-chunk inclusion, allowing broad coverage to coexist with selective, deeper reading during preference construction. Experiments across the Qwen and Llama model families show that FidelityFlow-PO improves both candidate-response quality and downstream long-context question answering. Across five benchmarks, it increases macro-F1 over the LongMab baselines on Qwen2.5-7B and Llama-3.1-8B. All adaptive reading remains offline, requiring no additional inference-time search.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.