acceptodds
Under review as a conference paper at ICLR 2027

HeteroEngine: Concurrent Training on Apple Silicon's GPU and Neural Engine

Abstract

Apple Silicon uses a system-on-chip (SoC) architecture in which the CPU, GPU, and Apple Neural Engine (ANE) share unified memory. The two accelerators, the GPU and the ANE, have opposite characteristics: the GPU offers high throughput but consumes more energy, whereas the ANE consumes less energy but offers lower throughput. Moreover, because the two engines share unified memory, they can access the same data and weights without the inter-device transfers required by discrete accelerators with separate device memory; this favors using them together. Standard training frameworks, however, backpropagate on the GPU alone, leaving the ANE idle throughout training. We therefore aim to use both engines concurrently, each compensating for the other's weakness—the GPU's high energy consumption and the ANE's low throughput—and thereby to exceed GPU-only training in both throughput and energy efficiency. In practice, Apple's public software exposes the ANE only through Core ML, a framework for executing trained models that provides no automatic differentiation and thus cannot compute gradients. Prior ANE training work consequently relied on Apple's unpublished internal interfaces. We present HeteroEngine, which derives the backward pass by hand and implements it as a Core ML program, thereby training on the GPU and the ANE concurrently through Apple's official APIs alone, without any such private interface. HeteroEngine partitions each training batch into microbatches and distributes them between the GPU and the ANE. Starting from the same model parameters, both engines concurrently perform the forward and backward passes of the entire model on their assigned microbatches. HeteroEngine then averages the gradients across all microbatches and applies a single optimizer update. When training GPT-2 124M on an M4 Pro, HeteroEngine achieves 1.419× the throughput of GPU-only training at 30.9% less energy per token; relative to ANE-only training, it achieves more than twice the throughput at higher energy per token. Whereas ANE-only training yields a small but statistically significant increase in perplexity, HeteroEngine matches or improves on GPU-only quality: at equal token counts its perplexity is equivalent to GPU-only perplexity on Shakespeare and WikiText-103; on WikiText-103 it is slightly lower after 18.06M tokens at equal token counts and substantially lower at equal training time; and on a Mac mini M2 the same program passes the same test and reproduces the direction of both gains. HeteroEngine thus uses the otherwise idle ANE alongside the GPU, raising both throughput and energy efficiency without loss of quality, and, built on Apple's official APIs alone, runs on a standard macOS installation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.