ASM-RISE: Learning to Investigate Assembly under an Evidence Budget
Abstract
Assembly semantics is often distributed across control flow, call relations, and data dependencies rather than contained in the visible instruction sequence. Resolving a question therefore requires not only interpreting assembly, but also deciding which program evidence to inspect next. We introduce ASM-RISE, a framework that turns assembly understanding into budgeted program investigation. A structured environment exposes complementary static views through scoped actions with explicit evidence prices, while trajectory supervised fine-tuning teaches the interaction protocol and agentic reinforcement learning jointly improves evidence acquisition and semantic answering. We evaluate the learned policies on binary code similarity retrieval, tool-free program semantics, semantic answering under four evidence budgets, and general code generation. Across three model families, RL consistently improves retrieval and budgeted reasoning over the corresponding supervised initialization. On Qwen2.5-Coder-7B-Instruct, adaptive RL yields a 52.92% relative improvement in Recall@1 over SFT and outperforms both static-plan and no-tool RL on the evidence-dependent endpoints. Policy rankings differ between evidence-dependent tasks and tool-free semantics, showing that investigation quality and direct-answer proficiency are not interchangeable. ASM-RISE demonstrates that models can learn not only what assembly means, but how to investigate it.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.