M6 generation

Local LLMs on the M6

The M6 moves 153 GB/s and takes up to 32 GB of unified memory. That combination runs everything up to Qwen 2.5 32B at Q4, at roughly 6.9 tokens/sec. Bandwidth sets the speed, memory sets the ceiling, and the tables below give both for every model.

Worth knowing. Up to 170 GB/s on higher configurations; 153 is the entry figure used here.

Memory bandwidth

153 GB/s

Unified memory

16, 24, 32 GB

8B model at Q4

~26 tok/s

Found in: Mac mini (2026). Specifications verified 2026-09-11against Apple’s published figures.

What the M6 runs, at Q4_K_M

Memory figures include a 3 GB reserve for macOS and an 8K context. Speeds come from a model fitted to our own measured benchmarks; the method and its error bars are on the calculator page.

ModelNeeds16 GB24 GB32 GB
Llama 3.2 3B8 GB56 tok/s56 tok/s56 tok/s
Mistral 7B16 GB28.5 tok/s28.5 tok/s28.5 tok/s
Qwen 2.5 7B16 GB27.2 tok/s27.2 tok/s27.2 tok/s
Llama 3.1 8B16 GB26 tok/s26 tok/s26 tok/s
DeepSeek R1 Distill 8B16 GB26 tok/s26 tok/s26 tok/s
Qwen 2.5 14B16 GB14.7 tok/s14.7 tok/s14.7 tok/s
Qwen 2.5 32B32 GBtight6.9 tok/s
Llama 3.3 70B56 GB
Qwen 2.5 72B56 GB
Mistral Large 296 GB
Llama 3.1 405B288 GB
DeepSeek V3480 GB

What it cannot run

Even at 32 GB, these need more unified memory than the M6 can address at Q4: Llama 3.3 70B, Qwen 2.5 72B, Mistral Large 2, Llama 3.1 405B, DeepSeek V3. Dropping the quantization buys a little room but not a generation of it, so the real options are a chip with a higher ceiling or several machines with pooled memory.

M6 against the M5

Bandwidth went from 153 GB/s to 153 GB/s, which is 1x, and that ratio carries almost directly into generation speed at the same model and quantization. The memory ceiling went from 32 GB to 32 GB, which is the part that decides whether a model runs at all rather than how fast. The ceiling did not move, so this is a speed upgrade rather than a capability one. If a model did not fit before, it still will not.

Common questions

What is the largest model an M6 can run?

At Q4_K_M with 32 GB of unified memory, Qwen 2.5 32B is the largest of the major open models that fits, needing about 32 GB once macOS is accounted for. Anything larger has to drop to a smaller quantization, run across clustered machines, or move to a chip with a higher memory ceiling.

How fast is the M6 for local inference?

It has 153 GB/s of memory bandwidth, and generation speed on Apple Silicon is set almost entirely by bandwidth divided by the size of the weights. In practice that works out to roughly 26 tokens/sec on an 8B model at Q4 and 6.9 on a 32B, assuming they fit. That is about 1x the M5 at the same model size.

How much memory should I order with an M6?

Memory decides which models you can run at all, and it cannot be upgraded later, so it is the one specification worth overbuying. M6 ships in 16, 24, 32 GB. The 16 GB configuration caps you at the small end; 32 GB is what opens up the larger models in the table on this page.

The M6 is in our lineup, not yet in our racks

Straight answer, since you are probably here deciding what to buy. We do not own this hardware yet. It is listed in our configurator at its real price, $128/mo for the entry configuration, and choosing it puts you on the waitlist rather than through checkout. If you want one, that is the honest way to tell us, and it is what decides which machines we buy next. Available today, the closest thing we run is the M4 at 120 GB/s, up to 32 GB, from $99/mo, which is 78% of the bandwidth in the table above.

Other Apple Silicon