acceptodds
Under review as a conference paper at ICLR 2027

PULSEBench: Personalized Understanding of Longitudinal Sensor Evidence for Wearable Health Reasoning

Abstract

Answering open-set health questions from longitudinal wearable records requires identifying relevant signals, time windows, and personal reference points, yet this end-to-end evidence discovery remains insufficiently evaluated. We introduce PULSEBench, a benchmark of 3,000 questions from 120 individuals across three public wearable datasets. Questions are grounded in observed trajectories, while a five-dimensional rubric evaluates free-text answers against question-specific evidence that is withheld from answering systems. Our evaluation reveals a persistent gap in evidence localization and supported attribution: adaptive inspection improves every evaluated backbone over direct inference, but substantial weaknesses remain. Controlled experiments further show that longer inputs and uniformly finer temporal resolution do not reliably improve answer quality, highlighting the importance of temporal organization. Motivated by these findings, we introduce Personalized Hierarchical Memory (PHM), which organizes personal records across temporal scales and retrieves evidence conditioned on the question. PHM improves over direct inference with a single answer-generation pass, although iterative agents remain stronger overall in the shared-model comparison. PULSEBench provides a testbed for studying how models discover and use personal evidence, beyond their ability to reason over preselected observations.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.