BRINE: Matched-State Counterfactual Supervision for Risk-Aware Cognitive Routing in Underwater Robots
Abstract
Long-horizon underwater control requires cognitive computation to be allocated by current risk and capability. Reactive control efficiently handles routine conditions, while degraded sensing and disturbances benefit from adaptive deliberation; learning this allocation requires evidence about the consequences of alternative cognitive branches. We introduce BRINE, a multimodal system that addresses this coupled decision, learning, and execution problem. From observed robot state, RGB images, point clouds, sonar, task commands, and episodic memory, BRINE determines when deliberation is needed and coordinates three role-specialized components: a state refiner, a safety verifier, and a capability planner. A safety-constrained action projection governs execution. To train the router, we construct matched-state counterfactual supervision by evaluating alternative expert combinations from the same simulator state and comparing their task, safety, and computation outcomes. We further formulate risk-shielded contextual-bandit refinement of the router and introduce role-aware mixed-precision execution that retains precision-sensitive semantic and safety computations in FP32. On 72 paired underwater tasks, the standard-precision system succeeds in 62 cases, compared with 48 for reactive control alone and 52 for always-on deliberation. Counterfactual supervision improves held-out expected return by 0.255 (95% confidence interval: 0.139–0.377) over matched risk and teacher supervision. In separate paired 24-task evaluations, contextual-bandit refinement reduces actual expert-forward duty from 0.609 to 0.482 (paired difference −0.127; 95% confidence interval: −0.180 to −0.074; Holm-adjusted p < 0.001), completes 24/24 tasks with zero collisions, and matches the reference route on all 287 high-risk decisions; mixed precision reduces peak temporary allocated memory by 7.13% and matches FP32 with 24/24 successes. These findings support counterfactual learning for selective, safety-conscious deliberation and mixed-precision execution with preserved behavioral semantics.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.