HatchRCA: Self-Incubated Small Models from LLMs for Root Cause Analysis
Abstract
Root cause analysis (RCA) in cloud systems requires both semantic reasoning over heterogeneous telemetry and quantitative pattern recognition at scale—two capabilities that existing approaches can only combine by manually engineering features and training pipelines for each new dataset, and by admitting empirically summarized rules without quantitative evidence of their reliability. We propose HatchRCA, a framework in which an LLM autonomously designs, trains, and invokes task-specific small models: numerical pattern recognition is delegated to the small models while semantic adjudication remains with the LLM, and both the models and the experience rules are produced by the LLM itself, requiring no human involvement. To make such autonomy safe, every self-generated artifact must meet measured acceptance criteria before deployment—criteria that are data-derived rather than preset, non-degrading, and falsifiable on held-out cases. Models are trained from LLM-written code under an adaptive validation scheme and calibrated before use; those with insufficient accuracy, leakage across splits, or unbounded overfitting are retrained automatically. Rules mined from failure cases are admitted only if they carry sufficient support, correct at least 80% of the cases they fire on, and leave the overall score unchanged or improved, then survive re-verification on held-out cases; artifacts whose benefit cannot be confirmed are never enabled. Under a matched backbone, scorer, and task set, HatchRCA improves the full-score rate by 140% (23.0% to 55.2%) and the mean score by 115% (0.322 to 0.694) over the official RCA agent across 87 tasks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.