By chip

Local LLMs on every Apple Silicon chip

Two numbers separate these chips for local inference: how much unified memory they can hold, which decides what runs at all, and how fast that memory is, which decides how quickly it runs. Everything else is secondary. Pick a chip for the full model by model breakdown, or size a specific case in the calculator.

M6

153 GB/s

Up to 32 GB · 26 tok/s on an 8B

Tops out at Qwen 2.5 32B in Q4. In our lineup, on the waitlist.

M5 Ultra

1200 GB/s

Up to 512 GB · 120 tok/s on an 8B

Tops out at DeepSeek V3 in Q4. In our lineup, on the waitlist.

M5 Max (614 GB/s)

614 GB/s

Up to 128 GB · 80 tok/s on an 8B

Tops out at Mistral Large 2 in Q4. In our lineup, on the waitlist.

M5 Pro

307 GB/s

Up to 64 GB · 47.3 tok/s on an 8B

Tops out at Qwen 2.5 72B in Q4. In our lineup, on the waitlist.

M4 Max (40-core GPU)

546 GB/s

Up to 128 GB · 74 tok/s on an 8B

Tops out at Mistral Large 2 in Q4. We rent this chip.

M4 Pro

273 GB/s

Up to 64 GB · 42.9 tok/s on an 8B

Tops out at Qwen 2.5 72B in Q4. We rent this chip.

M4

120 GB/s

Up to 32 GB · 20.8 tok/s on an 8B

Tops out at Qwen 2.5 32B in Q4. We rent this chip.

M3 Ultra

819 GB/s

Up to 512 GB · 97 tok/s on an 8B

Tops out at DeepSeek V3 in Q4. We rent this chip.

M3 Max (40-core GPU)

400 GB/s

Up to 128 GB · 58 tok/s on an 8B

Tops out at Mistral Large 2 in Q4.

M2 Ultra

800 GB/s

Up to 192 GB · 95 tok/s on an 8B

Tops out at Mistral Large 2 in Q4.

M1 Max

400 GB/s

Up to 64 GB · 58 tok/s on an 8B

Tops out at Qwen 2.5 72B in Q4.

Bandwidth and memory figures are Apple’s own published specifications, verified 2026-09-11. Speeds are estimates from a model fitted to benchmarks we measured ourselves; the method is on the calculator page and the raw data on benchmarks.