M3 generation

Local LLMs on the M3 Max (40-core GPU)

The M3 Max (40-core GPU) moves 400 GB/s and takes up to 128 GB of unified memory. That combination runs everything up to Mistral Large 2 at Q4, at roughly 4.8 tokens/sec. Bandwidth sets the speed, memory sets the ceiling, and the tables below give both for every model.

Memory bandwidth

400 GB/s

Unified memory

48, 64, 128 GB

8B model at Q4

~58 tok/s

Found in: MacBook Pro 14/16-inch (2023). Specifications verified 2026-09-11against Apple’s published figures.

What the M3 Max (40-core GPU) runs, at Q4_K_M

Memory figures include a 3 GB reserve for macOS and an 8K context. Speeds come from a model fitted to our own measured benchmarks; the method and its error bars are on the calculator page.

ModelNeeds48 GB64 GB128 GB
Llama 3.2 3B8 GB109 tok/s109 tok/s109 tok/s
Mistral 7B16 GB63 tok/s63 tok/s63 tok/s
Qwen 2.5 7B16 GB61 tok/s61 tok/s61 tok/s
Llama 3.1 8B16 GB58 tok/s58 tok/s58 tok/s
DeepSeek R1 Distill 8B16 GB58 tok/s58 tok/s58 tok/s
Qwen 2.5 14B16 GB35.2 tok/s35.2 tok/s35.2 tok/s
Qwen 2.5 32B32 GB17.2 tok/s17.2 tok/s17.2 tok/s
Llama 3.3 70B56 GB8.3 tok/s8.3 tok/s
Qwen 2.5 72B56 GB8 tok/s8 tok/s
Mistral Large 296 GB4.8 tok/s
Llama 3.1 405B288 GB
DeepSeek V3480 GB

What it cannot run

Even at 128 GB, these need more unified memory than the M3 Max (40-core GPU) can address at Q4: Llama 3.1 405B, DeepSeek V3. Dropping the quantization buys a little room but not a generation of it, so the real options are a chip with a higher ceiling or several machines with pooled memory.

M3 Max (40-core GPU) against the M2 Max

Bandwidth went from 400 GB/s to 400 GB/s, which is 1x, and that ratio carries almost directly into generation speed at the same model and quantization. The memory ceiling went from 96 GB to 128 GB, which is the part that decides whether a model runs at all rather than how fast. That extra headroom is the stronger reason to upgrade, because speed you can wait out and memory you cannot.

Common questions

What is the largest model an M3 Max (40-core GPU) can run?

At Q4_K_M with 128 GB of unified memory, Mistral Large 2 is the largest of the major open models that fits, needing about 96 GB once macOS is accounted for. Anything larger has to drop to a smaller quantization, run across clustered machines, or move to a chip with a higher memory ceiling.

How fast is the M3 Max (40-core GPU) for local inference?

It has 400 GB/s of memory bandwidth, and generation speed on Apple Silicon is set almost entirely by bandwidth divided by the size of the weights. In practice that works out to roughly 58 tokens/sec on an 8B model at Q4 and 17.2 on a 32B, assuming they fit. That is about 1x the M2 Max at the same model size.

How much memory should I order with an M3 Max (40-core GPU)?

Memory decides which models you can run at all, and it cannot be upgraded later, so it is the one specification worth overbuying. M3 Max (40-core GPU) ships in 48, 64, 128 GB. The 48 GB configuration caps you at the small end; 128 GB is what opens up the larger models in the table on this page.

We do not rent the M3 Max (40-core GPU)

Straight answer, since you are probably here deciding what to buy: our fleet is M4 generation, so we cannot put an M3 Max (40-core GPU) in front of you today. The closest thing we do run is the M4 Pro at 273 GB/s, which we offer up to 64 GB, from $200/mo. Against the M3 Max (40-core GPU) that is 68% of the bandwidth, so expect roughly that share of the speeds in the table above. If the model you want fits in what we have, renting saves you the upfront cost. If it does not, buy the machine, and the tables on this page will tell you which one.

Other Apple Silicon