M1 generation

Local LLMs on the M1 Max

The M1 Max moves 400 GB/s and takes up to 64 GB of unified memory. That combination runs everything up to Qwen 2.5 72B at Q4, at roughly 8 tokens/sec. Bandwidth sets the speed, memory sets the ceiling, and the tables below give both for every model.

Worth knowing. Still one of the best value-per-GB/s machines on the used market.

Memory bandwidth

400 GB/s

Unified memory

32, 64 GB

8B model at Q4

~58 tok/s

Found in: MacBook Pro 14/16-inch (2021), Mac Studio (2022). Specifications verified 2026-09-11against Apple’s published figures.

Mac Studio with the M1 Max for local LLMs

The Mac Studio is the desktop that carries the M1 Max, and for local inference it is a different class of machine from a Mac mini. The mini stops at the M4 Pro: 64 GB and 273 GB/s. The Studio with the M1 Max goes to 64 GB and 400 GB/s, which is 1.5x the bandwidth and therefore about that much faster on any model that fits both. At 64 GB it holds Qwen 2.5 72B at Q4_K_M.

Two things to check before ordering a Studio for this: memory is fixed at purchase, so size it for the largest model you expect to run, not the one you run today, and the price gap to a used RTX 3090 or a DGX Spark is on the hardware comparison. The calculator gives the memory and tokens per second for any model on this chip.

What the M1 Max runs, at Q4_K_M

Memory figures include a 3 GB reserve for macOS and an 8K context. Speeds come from a model fitted to our own measured benchmarks; the method and its error bars are on the calculator page.

ModelNeeds32 GB64 GB
Llama 3.2 3B8 GB109 tok/s109 tok/s
Mistral 7B16 GB63 tok/s63 tok/s
Qwen 2.5 7B16 GB61 tok/s61 tok/s
Llama 3.1 8B16 GB58 tok/s58 tok/s
DeepSeek R1 Distill 8B16 GB58 tok/s58 tok/s
Qwen 2.5 14B16 GB35.2 tok/s35.2 tok/s
Qwen 2.5 32B32 GB17.2 tok/s17.2 tok/s
Llama 3.3 70B56 GB8.3 tok/s
Qwen 2.5 72B56 GB8 tok/s
Mistral Large 296 GB
Llama 3.1 405B288 GB
DeepSeek V3480 GB

What it cannot run

Even at 64 GB, these need more unified memory than the M1 Max can address at Q4: Mistral Large 2, Llama 3.1 405B, DeepSeek V3. Dropping the quantization buys a little room but not a generation of it, so the real options are a chip with a higher ceiling or several machines with pooled memory.

Common questions

What is the largest model an M1 Max can run?

At Q4_K_M with 64 GB of unified memory, Qwen 2.5 72B is the largest of the major open models that fits, needing about 56 GB once macOS is accounted for. Anything larger has to drop to a smaller quantization, run across clustered machines, or move to a chip with a higher memory ceiling.

How fast is the M1 Max for local inference?

It has 400 GB/s of memory bandwidth, and generation speed on Apple Silicon is set almost entirely by bandwidth divided by the size of the weights. In practice that works out to roughly 58 tokens/sec on an 8B model at Q4 and 17.2 on a 32B, assuming they fit.

Is a Mac Studio with the M1 Max good for local LLMs?

Yes, and it is the reason the Mac Studio exists in this market: it is the only Apple desktop that ships the M1 Max with up to 64 GB of unified memory and 400 GB/s of memory bandwidth. A Mac mini tops out at the M4 Pro with 64 GB and 273 GB/s, so the Studio is the step up when you want 1.5x the generation speed on the same model.

How much memory should I order with an M1 Max?

Memory decides which models you can run at all, and it cannot be upgraded later, so it is the one specification worth overbuying. M1 Max ships in 32, 64 GB. The 32 GB configuration caps you at the small end; 64 GB is what opens up the larger models in the table on this page.

We do not rent the M1 Max

Straight answer, since you are probably here deciding what to buy: our fleet is M4 generation, so we cannot put an M1 Max in front of you today. The closest thing we do run is the M4 Pro at 273 GB/s, which we offer up to 64 GB, from $200/mo. Against the M1 Max that is 68% of the bandwidth, so expect roughly that share of the speeds in the table above. If the model you want fits in what we have, renting saves you the upfront cost. If it does not, buy the machine, and the tables on this page will tell you which one.

Other Apple Silicon