Micro-Agent: Mitigating Catastrophic Forgetting in Microscopic Spatial Prediction and Semantic QA through Two-Stage Adaptive Optimization
Abstract
Vision-language models (VLMs) have emerged as powerful foundations for robotic perception systems, yet their adaptation to specialized tasks presents a fundamental challenge: catastrophic forgetting of pre-trained capabilities. This phenomenon is particularly critical in autonomous micro-imaging systems, where precise spatial prediction must coexist with general visual understanding and semantic reasoning. In this paper, we present Micro-Agent, a two-stage optimization framework that enables VLM specialization for spatial prediction tasks while preserving foundational knowledge representations. Our approach addresses the stability-plasticity dilemma through a novel combination of Low-Rank Adaptation (LoRA) and a Mixture-of-Experts (MoE) correction strategy. In the first stage, LoRA-based fine-tuning adapts the pre-trained Qwen-VL model for distance and direction prediction while maintaining parameter-space integrity. The second stage introduces a hybrid expert ensemble comprising the LoRA-adapted model, four handcrafted sharpness features (Entropy, Log-Brenner, Log-Variance, Log-Energy), and a specialized 3-layer CNN with 16 channels, with outputs integrated through a learned fusion MLP. Experimental evaluation on micro-positioning tasks demonstrates that our framework reduces mean positional error from (Stage 1 only) to while improving directional accuracy from to . Critically, image description quality remains comparable to pre-trained levels ( vs ), confirming effective mitigation of catastrophic forgetting. The proposed architecture offers a principled approach for developing specialized vision-language agents that retain broad perceptual capabilities.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.