Task-Aware QUBO Allocation for Mixed-Precision Quantization
Abstract
Mixed-precision quantization requires discrete allocation of weight and activation bit-widths, followed by recovery of the selected network. We develop a task-aware quadratic unconstrained binary optimization (QUBO) surrogate with separate weight and activation profiles, a bit-operation (BOP) cost, and selected structural priors, then refine the resulting route through direct validation-based PROTES search. Across three public pretrained SIDD checkpoints with source-image-disjoint calibration and evaluation, the production simulated-annealing QUBO routes are within 0.01 dB of the strongest SFC routes in a retrospective 19-solver comparison. Refinement gives single-run static gains of 0.27–1.01 dB over QUBO and 0.39–1.34 dB over the HAWQ-style allocator at approximately 4.1% routed-layer BOPs, while remaining 0.39–0.87 dB below FP32. On a compact NAFBlock-based denoiser evaluated on a held-out patch split with shared source images, the refined route improves static PSNR over HAWQ-style by 0.21–2.46 dB at nearby costs across the three tightest requested targets. Recovery narrows these differences and can change route ordering, so we evaluate allocation quality separately before and after LSQ+. We report achieved cost, search variation, and optimization expense; BOP reductions are analytical, and the reference inference path remains floating point.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.