M3 generation
Local LLMs on the M3 Ultra
The M3 Ultra moves 819 GB/s and takes up to 512 GB of unified memory. That combination runs everything up to DeepSeek V3 at Q4, at roughly 29.6 tokens/sec. Bandwidth sets the speed, memory sets the ceiling, and the tables below give both for every model.
Worth knowing. 512 GB was the first configuration able to hold a 671B MoE in memory.
Memory bandwidth
819 GB/s
Unified memory
96, 256, 512 GB
8B model at Q4
~97 tok/s
Found in: Mac Studio (2025). Specifications verified 2026-09-11against Apple’s published figures.
What the M3 Ultra runs, at Q4_K_M
Memory figures include a 3 GB reserve for macOS and an 8K context. Speeds come from a model fitted to our own measured benchmarks; the method and its error bars are on the calculator page.
| Model | Needs | 96 GB | 256 GB | 512 GB |
|---|---|---|---|---|
| Llama 3.2 3B | 8 GB | 154 tok/s | 154 tok/s | 154 tok/s |
| Mistral 7B | 16 GB | 103 tok/s | 103 tok/s | 103 tok/s |
| Qwen 2.5 7B | 16 GB | 100 tok/s | 100 tok/s | 100 tok/s |
| Llama 3.1 8B | 16 GB | 97 tok/s | 97 tok/s | 97 tok/s |
| DeepSeek R1 Distill 8B | 16 GB | 97 tok/s | 97 tok/s | 97 tok/s |
| Qwen 2.5 14B | 16 GB | 63 tok/s | 63 tok/s | 63 tok/s |
| Qwen 2.5 32B | 32 GB | 32.9 tok/s | 32.9 tok/s | 32.9 tok/s |
| Llama 3.3 70B | 56 GB | 16.4 tok/s | 16.4 tok/s | 16.4 tok/s |
| Qwen 2.5 72B | 56 GB | 15.9 tok/s | 15.9 tok/s | 15.9 tok/s |
| Mistral Large 2 | 96 GB | 9.7 tok/s | 9.7 tok/s | 9.7 tok/s |
| Llama 3.1 405B | 288 GB | — | — | 3 tok/s |
| DeepSeek V3 | 480 GB | — | — | 29.6 tok/s |
M3 Ultra against the M2 Ultra
Bandwidth went from 800 GB/s to 819 GB/s, which is 1.02x, and that ratio carries almost directly into generation speed at the same model and quantization. The memory ceiling went from 192 GB to 512 GB, which is the part that decides whether a model runs at all rather than how fast. That extra headroom is the stronger reason to upgrade, because speed you can wait out and memory you cannot.
Common questions
What is the largest model an M3 Ultra can run?
At Q4_K_M with 512 GB of unified memory, DeepSeek V3 is the largest of the major open models that fits, needing about 480 GB once macOS is accounted for. Anything larger has to drop to a smaller quantization, run across clustered machines, or move to a chip with a higher memory ceiling.
How fast is the M3 Ultra for local inference?
It has 819 GB/s of memory bandwidth, and generation speed on Apple Silicon is set almost entirely by bandwidth divided by the size of the weights. In practice that works out to roughly 97 tokens/sec on an 8B model at Q4 and 32.9 on a 32B, assuming they fit. That is about 1.02x the M2 Ultra at the same model size.
How much memory should I order with an M3 Ultra?
Memory decides which models you can run at all, and it cannot be upgraded later, so it is the one specification worth overbuying. M3 Ultra ships in 96, 256, 512 GB. The 96 GB configuration caps you at the small end; 512 GB is what opens up the larger models in the table on this page.
Rent an M3 Ultra instead of buying one
We run this exact chip as a dedicated machine, from $571/mo, with Ollama installed and an OpenAI compatible endpoint already exposed. Same silicon and the same numbers as this page, without the hosting, the static IP, or being your own on call.