GlassSearch-Bench: Task-Adaptive Evaluation and Controlled Rubric Evolution for Smart-Glasses Search Agents
Abstract
Smart-glasses search requests require different evaluation criteria: identifying an object, finding a suitable destination, and giving directions impose different obligations. Task-specific rubrics express these obligations, but regenerating criteria for each evaluation can change the checks applied to an unchanged task. We present GlassSearch-Bench, which constructs task-specific rubrics from versioned, reusable atoms to constrain this source of variation. An answer-blind planner extracts requirements; a constructor selects, parameterizes, and narrows applicable atoms, adding local criteria for uncovered requirements. Repeated product comparisons reuse the resulting frozen rubric. Recurring gaps supply candidate library revisions, with validation and human review governing new versions. In an 18-task development study, candidate expansion reduces local criteria from 32 to 18 relative to the eligibility-repaired library and matches all 15 controlled completion targets from five probes. A 90-case replay study documents changes in requirement records and atom selection across repeated construction. A 500-task archive with four system response sets demonstrates the diagnostic value of task-specific criteria: on 432 tasks assessable for all four systems under both reference-answer conditions, every system scores higher on object identification than on either place discovery or wayfinding. GlassSearch-Bench links task-specific evaluation, explicit control of criterion reconstruction, and cumulative library development to actionable product profiles.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.