M2 generation

Local LLMs on the M2 Ultra

The M2 Ultra moves 800 GB/s and takes up to 192 GB of unified memory. That combination runs everything up to Mistral Large 2 at Q4, at roughly 9.5 tokens/sec. Bandwidth sets the speed, memory sets the ceiling, and the tables below give both for every model.

Worth knowing. 192 GB at 800 GB/s still holds 123B-class models at Q8.

Memory bandwidth

800 GB/s

Unified memory

64, 128, 192 GB

8B model at Q4

~95 tok/s

Found in: Mac Studio, Mac Pro (2023). Specifications verified 2026-09-11against Apple’s published figures.

Mac Studio with the M2 Ultra for local LLMs

The Mac Studio is the desktop that carries the M2 Ultra, and for local inference it is a different class of machine from a Mac mini. The mini stops at the M4 Pro: 64 GB and 273 GB/s. The Studio with the M2 Ultra goes to 192 GB and 800 GB/s, which is 2.9x the bandwidth and therefore about that much faster on any model that fits both. At 192 GB it holds Mistral Large 2 at Q4_K_M, which no mini can.

Two things to check before ordering a Studio for this: memory is fixed at purchase, so size it for the largest model you expect to run, not the one you run today, and the price gap to a used RTX 3090 or a DGX Spark is on the hardware comparison. The calculator gives the memory and tokens per second for any model on this chip.

What the M2 Ultra runs, at Q4_K_M

Memory figures include a 3 GB reserve for macOS and an 8K context. Speeds come from a model fitted to our own measured benchmarks; the method and its error bars are on the calculator page.

ModelNeeds64 GB128 GB192 GB
Llama 3.2 3B8 GB153 tok/s153 tok/s153 tok/s
Mistral 7B16 GB102 tok/s102 tok/s102 tok/s
Qwen 2.5 7B16 GB98 tok/s98 tok/s98 tok/s
Llama 3.1 8B16 GB95 tok/s95 tok/s95 tok/s
DeepSeek R1 Distill 8B16 GB95 tok/s95 tok/s95 tok/s
Qwen 2.5 14B16 GB62 tok/s62 tok/s62 tok/s
Qwen 2.5 32B32 GB32.2 tok/s32.2 tok/s32.2 tok/s
Llama 3.3 70B56 GB16 tok/s16 tok/s16 tok/s
Qwen 2.5 72B56 GB15.6 tok/s15.6 tok/s15.6 tok/s
Mistral Large 296 GB9.5 tok/s9.5 tok/s
Llama 3.1 405B288 GB
DeepSeek V3480 GB

What it cannot run

Even at 192 GB, these need more unified memory than the M2 Ultra can address at Q4: Llama 3.1 405B, DeepSeek V3. Dropping the quantization buys a little room but not a generation of it, so the real options are a chip with a higher ceiling or several machines with pooled memory.

M2 Ultra against the M1 Ultra

Bandwidth went from 800 GB/s to 800 GB/s, which is 1x, and that ratio carries almost directly into generation speed at the same model and quantization. The memory ceiling went from 128 GB to 192 GB, which is the part that decides whether a model runs at all rather than how fast. That extra headroom is the stronger reason to upgrade, because speed you can wait out and memory you cannot.

Common questions

What is the largest model an M2 Ultra can run?

At Q4_K_M with 192 GB of unified memory, Mistral Large 2 is the largest of the major open models that fits, needing about 96 GB once macOS is accounted for. Anything larger has to drop to a smaller quantization, run across clustered machines, or move to a chip with a higher memory ceiling.

How fast is the M2 Ultra for local inference?

It has 800 GB/s of memory bandwidth, and generation speed on Apple Silicon is set almost entirely by bandwidth divided by the size of the weights. In practice that works out to roughly 95 tokens/sec on an 8B model at Q4 and 32.2 on a 32B, assuming they fit. That is about 1x the M1 Ultra at the same model size.

Is a Mac Studio with the M2 Ultra good for local LLMs?

Yes, and it is the reason the Mac Studio exists in this market: it is the only Apple desktop that ships the M2 Ultra with up to 192 GB of unified memory and 800 GB/s of memory bandwidth. A Mac mini tops out at the M4 Pro with 64 GB and 273 GB/s, so the Studio is the step up when a model does not fit in 64 GB or when you want 2.9x the generation speed on the same model.

How much memory should I order with an M2 Ultra?

Memory decides which models you can run at all, and it cannot be upgraded later, so it is the one specification worth overbuying. M2 Ultra ships in 64, 128, 192 GB. The 64 GB configuration caps you at the small end; 192 GB is what opens up the larger models in the table on this page.

We do not rent the M2 Ultra

Straight answer, since you are probably here deciding what to buy: our fleet is M4 generation, so we cannot put an M2 Ultra in front of you today. The closest thing we do run is the M3 Ultra at 819 GB/s, which we offer up to 256 GB, from $571/mo. Against the M2 Ultra that is 102% of the bandwidth, so expect roughly that share of the speeds in the table above. If the model you want fits in what we have, renting saves you the upfront cost. If it does not, buy the machine, and the tables on this page will tell you which one.

Other Apple Silicon