A Composed State-Space World Model for Ultrasound-Guided Spinal Surgery
Abstract
Ultrasound-guided pedicle screw placement is studied as two disconnected problems: bring the probe to a standard imaging plane, and drill a planned corridor once you are there. The second is always evaluated with the probe teleported into place from ground truth, so no published number says whether the pair composes. We solve each problem independently with a state-space world model, then compose the two in sequence. Each is a learned readout of the tool's pose in a physically named frame, propagated by a small transition and planned by CEM-MPC over their respective objectives. The composition replaces the SonoGym simulator's privileged probe placement with a navigated one. In isolation each half is competitive: while learning its dynamics the drilling controller reaches 3.00 mm insertion and 2.80 mm lateral at 90% safety, beating both reported reinforcement-learning baselines on both position axes (their 11.9-12.3 and 5.07-5.42), and exact kinematics improve that again to 2.1/2.0 mm; the probe controller reaches and holds the plane in 0.95 of episodes. Composed, safety falls from 90% to 0% and success to 0.00, and it fails in one place, not everywhere: drill orientation survives a 32 deg probe azimuth error while lateral offset does not. Instead of building a better navigator, we inject known amounts of probe error and measure how much the drilling controller absorbs. The pose estimator turns out not to degrade at the new probe position but to stop responding to its input, reporting the same large error wherever the probe is, which is what an estimator does when both its image and its own output are far outside the distribution it was fitted on. Post-training it at the position where it is actually used repairs this. The composed system then recovers most of what the teleport gave it: safety returns to 85% against the teleported arm's 90% with the exact propagator and to 75% at matched propagator. Auditing the forecasts the planner already computes and discards splits its 13.2 mm error into a propagator term and a perception term: two independent routes put the propagator's share at 9.2 and 9.4 mm, and correcting it leaves a perception-limited 3.8 mm. Two further probes put a hand-designed alignment cap, not the planner, in control inside the corridor. Three lessons about measuring such compositions, demonstrated here rather than established in general: one-channel tolerances do not compose, the joint failing where each channel alone passes; the tightest tolerance is on a probe axis neither task commands, failing by leaving the intervention undone while the safety metric still reads safe; and the controller's own introspective signals invert, the arm with the worst forecast controlling best and the one that fails completely proving the most self-consistent.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.