acceptodds
Under review as a conference paper at ICLR 2027

SARFACT: Benchmarking Task Answers and Visual Evidence for SAR Vision–Language Understanding

Abstract

Evaluating SAR vision–language models only by the correctness of their task outputs leaves a verification gap: a model may produce the expected answer without identifying the visual support needed to inspect it. We introduce SARFACT, a large-scale object-centric SAR vision–language dataset and benchmark that preserves each task's native Task Answer and defines task-specific Visual Evidence: the localized object support needed to inspect the answer. Built from 200K SAR images and 1.11M foreground objects, SARFACT contains 1.42M examples across nine task interfaces. Task Answers and Visual Evidence are derived from shared localized object facts, while task-specific construction controls reduce shortcut-prone patterns and preserve informative evaluation cases. SARFACT separately evaluates Answer, Evidence, and Joint correctness. Experiments with six public VLMs reveal limited zero-shot performance and frequent cases in which correct Answers fail to satisfy the task-defined Visual Evidence requirements. Explicit Evidence supervision substantially improves joint Answer–Evidence performance. Annotations, task records, and evaluation code will be publicly released.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.