acceptodds
Under review as a conference paper at ICLR 2027

FinEvolvingAgent: Self-Evolving Financial Deep Research Agent via Gradient-Guided Group Direct Preference Optimization

Abstract

Financial deep research represents one of the most demanding frontiers for autonomous AI agents, requiring complex multi-hop numerical reasoning, visual understanding of multimodal charts and financial tables, and strategic long-horizon tool invocation across regulatory filings. Existing reinforcement learning methods, such as Proximal Policy Optimization (PPO) and Group Relative Policy Optimization (GRPO), face severe bottlenecks in this domain due to critic instability, high rollout variance, and the absence of continuous preference guidance across complex multi-step reasoning trajectories. To overcome these limitations, we introduce , an autonomous multimodal financial deep research agent built upon the open foundation model Qwen3.6-VL. We first propose , which incorporates continuous differentiable reward signals into policy trajectory rollouts, directing exploratory token generation and tool actions toward mathematically consistent paths and achieving a +5.6% accuracy gain over base Qwen3.6-VL. Furthermore, we develop , a principled framework that generalizes direct preference alignment to multi-candidate group rollouts via a margin-aware Plackett-Luce objective, yielding an additional +3.2% improvement over GRPO. Across six extensive financial benchmarks, including FinQA, TAT-QA, ConvFinQA, FinanceBench, MultiFin-VL, and a newly constructed realistic multi-source benchmark RealFin-DeepResearch, FinEvolvingAgent achieves a new state-of-the-art accuracy of 83.8%, outperforming leading proprietary models including GPT-4o and Claude 3.5 Sonnet while exhibiting superior calibration, numerical precision, and grounded tool execution.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.