M5 generation

Local LLMs on the M5 Ultra

The M5 Ultra moves 1200 GB/s and takes up to 512 GB of unified memory. That combination runs everything up to DeepSeek V3 at Q4, at roughly 41.1 tokens/sec. Bandwidth sets the speed, memory sets the ceiling, and the tables below give both for every model.

Worth knowing. 1.2 TB/s. The 512 GB tier exists only on the fully specced part, 36 core CPU and 80 core GPU; the configurations most people see stop at 256 GB.

Memory bandwidth

1200 GB/s

Unified memory

96, 256, 512 GB

8B model at Q4

~120 tok/s

Found in: Mac Studio (2026). Specifications verified 2026-09-11against Apple’s published figures.

What the M5 Ultra runs, at Q4_K_M

Memory figures include a 3 GB reserve for macOS and an 8K context. Speeds come from a model fitted to our own measured benchmarks; the method and its error bars are on the calculator page.

ModelNeeds96 GB256 GB512 GB
Llama 3.2 3B8 GB177 tok/s177 tok/s177 tok/s
Mistral 7B16 GB127 tok/s127 tok/s127 tok/s
Qwen 2.5 7B16 GB124 tok/s124 tok/s124 tok/s
Llama 3.1 8B16 GB120 tok/s120 tok/s120 tok/s
DeepSeek R1 Distill 8B16 GB120 tok/s120 tok/s120 tok/s
Qwen 2.5 14B16 GB83 tok/s83 tok/s83 tok/s
Qwen 2.5 32B32 GB45.5 tok/s45.5 tok/s45.5 tok/s
Llama 3.3 70B56 GB23.3 tok/s23.3 tok/s23.3 tok/s
Qwen 2.5 72B56 GB22.7 tok/s22.7 tok/s22.7 tok/s
Mistral Large 296 GB13.9 tok/s13.9 tok/s13.9 tok/s
Llama 3.1 405B288 GB4.4 tok/s
DeepSeek V3480 GB41.1 tok/s

M5 Ultra against the M3 Ultra

Bandwidth went from 819 GB/s to 1200 GB/s, which is 1.47x, and that ratio carries almost directly into generation speed at the same model and quantization. The memory ceiling went from 512 GB to 512 GB, which is the part that decides whether a model runs at all rather than how fast. The ceiling did not move, so this is a speed upgrade rather than a capability one. If a model did not fit before, it still will not.

Common questions

What is the largest model an M5 Ultra can run?

At Q4_K_M with 512 GB of unified memory, DeepSeek V3 is the largest of the major open models that fits, needing about 480 GB once macOS is accounted for. Anything larger has to drop to a smaller quantization, run across clustered machines, or move to a chip with a higher memory ceiling.

How fast is the M5 Ultra for local inference?

It has 1200 GB/s of memory bandwidth, and generation speed on Apple Silicon is set almost entirely by bandwidth divided by the size of the weights. In practice that works out to roughly 120 tokens/sec on an 8B model at Q4 and 45.5 on a 32B, assuming they fit. That is about 1.47x the M3 Ultra at the same model size.

How much memory should I order with an M5 Ultra?

Memory decides which models you can run at all, and it cannot be upgraded later, so it is the one specification worth overbuying. M5 Ultra ships in 96, 256, 512 GB. The 96 GB configuration caps you at the small end; 512 GB is what opens up the larger models in the table on this page.

The M5 Ultra is in our lineup, not yet in our racks

Straight answer, since you are probably here deciding what to buy. We do not own this hardware yet. It is listed in our configurator at its real price, $786/mo for the entry configuration, and choosing it puts you on the waitlist rather than through checkout. If you want one, that is the honest way to tell us, and it is what decides which machines we buy next. Available today, the closest thing we run is the M3 Ultra at 819 GB/s, up to 256 GB, from $571/mo, which is 68% of the bandwidth in the table above.

Other Apple Silicon