By chip
Local LLMs on every Apple Silicon chip
Two numbers separate these chips for local inference: how much unified memory they can hold, which decides what runs at all, and how fast that memory is, which decides how quickly it runs. Everything else is secondary. Pick a chip for the full model by model breakdown, or size a specific case in the calculator.
M6
153 GB/sUp to 32 GB · 26 tok/s on an 8B
Tops out at Qwen 2.5 32B in Q4. In our lineup, on the waitlist.
M5 Ultra
1200 GB/sUp to 512 GB · 120 tok/s on an 8B
Tops out at DeepSeek V3 in Q4. In our lineup, on the waitlist.
M5 Max (614 GB/s)
614 GB/sUp to 128 GB · 80 tok/s on an 8B
Tops out at Mistral Large 2 in Q4. In our lineup, on the waitlist.
M5 Pro
307 GB/sUp to 64 GB · 47.3 tok/s on an 8B
Tops out at Qwen 2.5 72B in Q4. In our lineup, on the waitlist.
M4 Max (40-core GPU)
546 GB/sUp to 128 GB · 74 tok/s on an 8B
Tops out at Mistral Large 2 in Q4. We rent this chip.
M4 Pro
273 GB/sUp to 64 GB · 42.9 tok/s on an 8B
Tops out at Qwen 2.5 72B in Q4. We rent this chip.
M4
120 GB/sUp to 32 GB · 20.8 tok/s on an 8B
Tops out at Qwen 2.5 32B in Q4. We rent this chip.
M3 Ultra
819 GB/sUp to 512 GB · 97 tok/s on an 8B
Tops out at DeepSeek V3 in Q4. We rent this chip.
M3 Max (40-core GPU)
400 GB/sUp to 128 GB · 58 tok/s on an 8B
Tops out at Mistral Large 2 in Q4.
M2 Ultra
800 GB/sUp to 192 GB · 95 tok/s on an 8B
Tops out at Mistral Large 2 in Q4.
M1 Max
400 GB/sUp to 64 GB · 58 tok/s on an 8B
Tops out at Qwen 2.5 72B in Q4.
Bandwidth and memory figures are Apple’s own published specifications, verified 2026-09-11. Speeds are estimates from a model fitted to benchmarks we measured ourselves; the method is on the calculator page and the raw data on benchmarks.