CARTA: Training-Free LLM-Driven Tensor Mapping for Spatial Accelerators
Abstract
Tensor mapping coordinates parallel execution with data reuse across an accelerator's memory hierarchy. Soter's per-task RL Tuner relearns it for each layer-accelerator pair. In CARTA, a frozen LLM given a short strategy text picks spatial allocations, loop priorities, and tile caps for Soter's deterministic mapping builder, and we ask what transfers across accelerators. On identical problems with an energy-delay product (EDP) objective, the Tuner has 1.85x CARTA's EDP at batch 16 across three independent runs with DeepSeek as the LLM and 1.61x with pinned open-weight Qwen3-14B. However, random search beats this weak incumbent on the main suite by 1.66x at 500 evaluator calls. At 200 evaluator calls, random search has 1.21x DeepSeek CARTA's EDP, while a Gaussian-process (GP) surrogate with a structured encoding and an LLM-free annealer seeded with the same text each have 1.06x. These margins lose support under workload resampling, reach parity at 500 calls, and reverse with Qwen3-14B, which both controls beat (0.92x). Porting changes about 30 lines of strategy text without retraining, but writing them costs at least 3,000 diagnostic evaluator calls per port plus expert reading. Without the text, DeepSeek loses to the Tuner on both held-out ports. With it, Tuner/CARTA EDP is 1.13x on GemminiLike and 1.02x on NVDLALike, a tie, and the annealer has 1.03-1.05x DeepSeek's EDP. Qwen3-14B fails the port prediction registered before the port's agent runs, with 1.45x the annealer's EDP on NVDLALike. The strategy text carries the port, and the model determines the increment, 1.03-1.06x for DeepSeek and negative for Qwen3-14B. We also measure how budget accounting, baseline alignment, and run selection inflate an LLM optimizer's apparent advantage.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.