acceptodds
Under review as a conference paper at ICLR 2027

CATSF: LEARNING ADAPTIVE REASONING FOR MULTIMODAL USER FEEDBACK CLASSIFICATION

Abstract

Android users generate large volumes of multimodal feedback, making manual triage impractical at scale and motivating automated classification. While multimodal large language models exhibit strong reasoning capabilities, high computational costs hinder their deployment, and the scarcity of high-quality task-specific data and reasoning supervision limits the adaptation of compact models. Reasoning-depth selection must also account for both text–image complexity and compact-model capabilities. To address these challenges, we introduce the Android User Feedback (AUF) dataset, a multimodal benchmark containing 3,750 expert-labeled entries, and propose CATSF, a Capability-Aware Teacher–Student Framework that separates learning to execute reasoning paths from learning to select them. In the first fine-tuning stage, a compact student learns basic and complex reasoning from teacher-generated trajectories. In the second stage, the same student learns adaptive path selection from teacher targets selected based on whether its first-stage basic predictions match the training labels. Experiments show that CATSF achieves a relative macro- improvement of up to 63.4% and an accuracy gain of up to 18.3 percentage points over the evaluated baselines. Across teacher–student configurations, CATSF reduces mean output token usage by up to 50.8% relative to matched, separately fine-tuned complex-reasoning baselines.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.