M4 generation
Local LLMs on the M4 Max (40-core GPU)
The M4 Max (40-core GPU) moves 546 GB/s and takes up to 128 GB of unified memory. That combination runs everything up to Mistral Large 2 at Q4, at roughly 6.5 tokens/sec. Bandwidth sets the speed, memory sets the ceiling, and the tables below give both for every model.
Memory bandwidth
546 GB/s
Unified memory
48, 64, 128 GB
8B model at Q4
~74 tok/s
Found in: MacBook Pro 14/16-inch (2024), Mac Studio (2025). Specifications verified 2026-09-11against Apple’s published figures.
What the M4 Max (40-core GPU) runs, at Q4_K_M
Memory figures include a 3 GB reserve for macOS and an 8K context. Speeds come from a model fitted to our own measured benchmarks; the method and its error bars are on the calculator page.
| Model | Needs | 48 GB | 64 GB | 128 GB |
|---|---|---|---|---|
| Llama 3.2 3B | 8 GB | 129 tok/s | 129 tok/s | 129 tok/s |
| Mistral 7B | 16 GB | 79 tok/s | 79 tok/s | 79 tok/s |
| Qwen 2.5 7B | 16 GB | 76 tok/s | 76 tok/s | 76 tok/s |
| Llama 3.1 8B | 16 GB | 74 tok/s | 74 tok/s | 74 tok/s |
| DeepSeek R1 Distill 8B | 16 GB | 74 tok/s | 74 tok/s | 74 tok/s |
| Qwen 2.5 14B | 16 GB | 45.8 tok/s | 45.8 tok/s | 45.8 tok/s |
| Qwen 2.5 32B | 32 GB | 22.9 tok/s | 22.9 tok/s | 22.9 tok/s |
| Llama 3.3 70B | 56 GB | — | 11.2 tok/s | 11.2 tok/s |
| Qwen 2.5 72B | 56 GB | — | 10.9 tok/s | 10.9 tok/s |
| Mistral Large 2 | 96 GB | — | — | 6.5 tok/s |
| Llama 3.1 405B | 288 GB | — | — | — |
| DeepSeek V3 | 480 GB | — | — | — |
What it cannot run
Even at 128 GB, these need more unified memory than the M4 Max (40-core GPU) can address at Q4: Llama 3.1 405B, DeepSeek V3. Dropping the quantization buys a little room but not a generation of it, so the real options are a chip with a higher ceiling or several machines with pooled memory.
M4 Max (40-core GPU) against the M3 Max (40-core GPU)
Bandwidth went from 400 GB/s to 546 GB/s, which is 1.37x, and that ratio carries almost directly into generation speed at the same model and quantization. The memory ceiling went from 128 GB to 128 GB, which is the part that decides whether a model runs at all rather than how fast. The ceiling did not move, so this is a speed upgrade rather than a capability one. If a model did not fit before, it still will not.
Common questions
What is the largest model an M4 Max (40-core GPU) can run?
At Q4_K_M with 128 GB of unified memory, Mistral Large 2 is the largest of the major open models that fits, needing about 96 GB once macOS is accounted for. Anything larger has to drop to a smaller quantization, run across clustered machines, or move to a chip with a higher memory ceiling.
How fast is the M4 Max (40-core GPU) for local inference?
It has 546 GB/s of memory bandwidth, and generation speed on Apple Silicon is set almost entirely by bandwidth divided by the size of the weights. In practice that works out to roughly 74 tokens/sec on an 8B model at Q4 and 22.9 on a 32B, assuming they fit. That is about 1.37x the M3 Max (40-core GPU) at the same model size.
How much memory should I order with an M4 Max (40-core GPU)?
Memory decides which models you can run at all, and it cannot be upgraded later, so it is the one specification worth overbuying. M4 Max (40-core GPU) ships in 48, 64, 128 GB. The 48 GB configuration caps you at the small end; 128 GB is what opens up the larger models in the table on this page.
Rent an M4 Max (40-core GPU) instead of buying one
We run this exact chip as a dedicated machine, from $286/mo, with Ollama installed and an OpenAI compatible endpoint already exposed. Same silicon and the same numbers as this page, without the hosting, the static IP, or being your own on call.