acceptodds
Under review as a conference paper at ICLR 2027

Beyond Crowd Labels: Learning Process Supervision from Crowd Reasoning Trajectories

Abstract

Collecting high-quality supervision data is critical to the development of modern AI systems, and crowdsourcing provides a scalable way to obtain such data from multiple workers. Classical learning-from-crowds methods recover reliable labels from noisy worker annotations over predefined response options. Recent advances in large language models have increased demand for supervision that captures not only what the correct answer is but also how it is derived, which motivates the crowdsourced collection of free-form reasoning trajectories. In this paper, we ask whether, when crowd workers provide free-form reasoning trajectories instead of annotations over predefined options, reliable process supervision can still be recovered directly from noisy and heterogeneous worker contributions. To answer this question, we propose a structured inference framework that leverages shared reasoning structure across heterogeneous worker trajectories. We first construct a task-specific latent consensus process and softly align each worker trajectory to this process to establish cross-trajectory correspondences. We then perform structured variational inference over the aligned trajectories and observed final outcomes to recover the validity of individual reasoning steps. Evaluations with both LLM and human crowdworkers show that reliable step-level validity signals can be recovered from heterogeneous crowd reasoning trajectories and verified final outcomes without external step-level evaluators. The recovered signals capture meaningful variation in intermediate reasoning validity and provide useful supervision for downstream LLM fine-tuning across multiple reasoning domains.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.