Deep Think with Rehearsal for Low-Latency Team-AI Collaboration
Abstract
The integration of Large Language Models (LLMs) into scientific team meetings presents exciting opportunities to accelerate biomedical discovery, especially through their strong deep thinking capabilities enabled by multi-step reasoning and web search. However, such methods are computationally expensive and introduce substantial latency, limiting their effectiveness in real-time Team-AI communications. In this study, we propose Deep Think with Rehearsal (DTR), a novel framework that decouples deep reasoning from synchronous interaction in the AI4Science context. DTR transfers computationally intensive reasoning into an offline rehearsal phase, allowing the LLM to pre-cognize complex scientific contexts and deliver high-quality, "deep" insights with minimal latency during live interactions. To facilitate this research, we introduce the Scientific Team Meeting Dataset (STMD), a hybrid benchmark comprising authentic transcripts from three real-world biomedical research labs alongside extensive simulated multi-party deliberations synthesized from PubMed literature. Experiments in both simulated and real-world settings demonstrate that DTR consistently improves response quality while reducing inference latency compared to state-of-the-art methods, highlighting the effectiveness of rehearsal in enabling low-latency, high-quality scientific collaboration.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.