BEAR-Med: From a Clinically Grounded Benchmark to Expert-Aligned Reasoning in Medical Image Forensics
Abstract
Clinical image forensics extends beyond binary detection: systems must distinguish local edits from global synthesis, localize affected regions, and justify verdicts with observable evidence. Existing benchmarks remain limited in scale, imaging diversity, forgery coverage, clinical plausibility, or spatial fidelity, while current methods often separate detection from localization, append post-hoc explanations, or derive chain-of-thought supervision from general-purpose multimodal models without explicit forensic evidence. We introduce BEAR-Med-1M, a million-image benchmark spanning four clinical imaging categories and five local-to-global mechanisms: copy-move, splicing, lesion insertion, lesion removal, and full-image synthesis. It targets medically meaningful, anatomically plausible regions and uses bidirectional restoration for lesion insertion and removal, producing spatially faithful local forgeries with low out-of-box spillover. We also introduce BEAR-Med, an expert-aligned model whose four frozen experts capture local traces, image-wide consistency, generation texture, and medical normality. Their evidence is distilled offline into expert-guided CoT targets, while inference uses only the original image. BEAR-Med generates a pre-hoc analyze-localize-decide sequence, placing evidence-based reasoning before localization and verdict rather than explaining predictions post hoc. A segmented objective separately balances reasoning, box, and verdict tokens, followed by GRPO with verifiable task and spatial rewards. BEAR-Med leads dense-localization and multimodal reasoning baselines in accuracy, F1, all reported localization metrics, and authentic-image error control, while retaining strong cross-dataset detection and localization.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.