acceptodds
Under review as a conference paper at ICLR 2027

Solve, Recover, Reuse: LLM Reasoning Across Workflows with Corrupted Contexts

Abstract

Large language models (LLMs) increasingly operate in complex workflows and noisy contexts. It's important to have a comprehensive understanding of how model performance varies across workflow and corruption types, and whether both harmful and beneficial transitions arise under these challenges. We introduce an Evolutionary Task Generator (ETG) that generates a series of linked tasks across workflows and corruptions, providing a framework to systematically experiment and analyze: 1) model performance changes across workflows and corruptions, 2) the bidirectional transitions under corrupted contexts, and 3) training effects across workflows and corruptions. Across frontier LLMs, we find that model rankings and performance change substantially across workflows, e.g., DeepSeek-V4-Flash drops from 0.75 in direct solving to 0.20 in downstream reuse. Instance-level analysis shows that both harmful and beneficial transitions arise under corrupted contexts. Training results show that direct-solving gains may not transfer across workflows, and may even degrade model performance under complex workflows with corrupted contexts. Our results highlight the importance of examining LLMs reasoning through the lens of complex workflows and corrupted contexts, and suggest a more constructive direction when such noise is unavoidable in practice: resist the harmful and leverage the beneficial.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.