DAPF: A Dynamic Agent-Driven Automatic Post-Training Framework
Abstract
Domain-specific fine-tuning of large language models requires repeated decisions about training configurations, data schedules, and recovery from unsuccessful experiments. Automating these decisions requires connecting training feedback to actionable interventions while accounting for limited computational resources. We introduce DAPF, an agent-driven framework for closed-loop post-training of language and vision-language models. Given a target model, task domain, and validation performance target, DAPF iteratively generates hypotheses, designs and executes experiments, diagnoses training outcomes, and selects subsequent interventions. Its attribution module combines loss-pattern analysis with agent-generated diagnoses to propose targeted adjustments and update estimates of intervention effectiveness from experimental feedback. A budget-aware selector prioritizes expected validation improvement per unit cost and chooses the least-cost action predicted to reach the target when available. An append-only evidence graph records experimental decisions and supports recovery of model checkpoints and agent state. We evaluate DAPF on seven benchmarks using Qwen3-1.7B for text tasks and Qwen2.5-VL-3B for multimodal reasoning. Across three independent runs, DAPF achieves higher mean test accuracy than FT-Agent on every benchmark, with an average absolute gain of 1.74 percentage points. Reported average GPU runtime decreases from 1.5 to 1.3 hours, while the average number of effective optimization rounds decreases from 4.9 to 3.8. These results suggest that feedback-guided intervention selection can improve fine-tuning performance and computational efficiency within a bounded search space.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.