M4 generation

Local LLMs on the M4 Pro

The M4 Pro moves 273 GB/s and takes up to 64 GB of unified memory. That combination runs everything up to Qwen 2.5 72B at Q4, at roughly 5.5 tokens/sec. Bandwidth sets the speed, memory sets the ceiling, and the tables below give both for every model.

Worth knowing. 64 GB is the cheapest sensible way into 70B at Q4.

Memory bandwidth

273 GB/s

Unified memory

24, 48, 64 GB

8B model at Q4

~42.9 tok/s

Found in: MacBook Pro 14/16-inch, Mac mini (2024). Specifications verified 2026-09-11against Apple’s published figures.

What the M4 Pro runs, at Q4_K_M

Memory figures include a 3 GB reserve for macOS and an 8K context. Speeds come from a model fitted to our own measured benchmarks; the method and its error bars are on the calculator page.

ModelNeeds24 GB48 GB64 GB
Llama 3.2 3B8 GB86 tok/s86 tok/s86 tok/s
Mistral 7B16 GB46.8 tok/s46.8 tok/s46.8 tok/s
Qwen 2.5 7B16 GB44.8 tok/s44.8 tok/s44.8 tok/s
Llama 3.1 8B16 GB42.9 tok/s42.9 tok/s42.9 tok/s
DeepSeek R1 Distill 8B16 GB42.9 tok/s42.9 tok/s42.9 tok/s
Qwen 2.5 14B16 GB25.1 tok/s25.1 tok/s25.1 tok/s
Qwen 2.5 32B32 GBtight12 tok/s12 tok/s
Llama 3.3 70B56 GB5.7 tok/s
Qwen 2.5 72B56 GB5.5 tok/s
Mistral Large 296 GB
Llama 3.1 405B288 GB
DeepSeek V3480 GB

What it cannot run

Even at 64 GB, these need more unified memory than the M4 Pro can address at Q4: Mistral Large 2, Llama 3.1 405B, DeepSeek V3. Dropping the quantization buys a little room but not a generation of it, so the real options are a chip with a higher ceiling or several machines with pooled memory.

M4 Pro against the M3 Pro

Bandwidth went from 150 GB/s to 273 GB/s, which is 1.82x, and that ratio carries almost directly into generation speed at the same model and quantization. The memory ceiling went from 36 GB to 64 GB, which is the part that decides whether a model runs at all rather than how fast. That extra headroom is the stronger reason to upgrade, because speed you can wait out and memory you cannot.

Common questions

What is the largest model an M4 Pro can run?

At Q4_K_M with 64 GB of unified memory, Qwen 2.5 72B is the largest of the major open models that fits, needing about 56 GB once macOS is accounted for. Anything larger has to drop to a smaller quantization, run across clustered machines, or move to a chip with a higher memory ceiling.

How fast is the M4 Pro for local inference?

It has 273 GB/s of memory bandwidth, and generation speed on Apple Silicon is set almost entirely by bandwidth divided by the size of the weights. In practice that works out to roughly 42.9 tokens/sec on an 8B model at Q4 and 12 on a 32B, assuming they fit. That is about 1.82x the M3 Pro at the same model size.

How much memory should I order with an M4 Pro?

Memory decides which models you can run at all, and it cannot be upgraded later, so it is the one specification worth overbuying. M4 Pro ships in 24, 48, 64 GB. The 24 GB configuration caps you at the small end; 64 GB is what opens up the larger models in the table on this page.

Rent an M4 Pro instead of buying one

We run this exact chip as a dedicated machine, from $200/mo, with Ollama installed and an OpenAI compatible endpoint already exposed. Same silicon and the same numbers as this page, without the hosting, the static IP, or being your own on call.

Other Apple Silicon