Ollama hosting on a dedicated Mac mini or Mac Studio
Macyou rents dedicated Apple Silicon Macs with Ollama preinstalled and an OpenAI compatible API, from $99/mo for an M4 Mac mini up to an M3 Ultra Mac Studio with 256 GB of unified memory. Each machine is used by one customer, is ready in about 5 minutes, and runs your models around the clock at a fixed monthly price with no per token fees.
Which Mac runs which model
On Apple Silicon the model has to fit in unified memory, and generation speed follows memory bandwidth. Measured figures come from our fleet benchmarks; estimates use the same formula, which stays within about 5% of every machine we have measured.
| Machine | Fits (Q4) | Typical speed | Monthly |
|---|---|---|---|
| M4 Mac mini, 16 GB | 3B to 14B models | Llama 3.1 8B: 21.2 tok/s (measured) | from $99/mo |
| M4 Mac mini, 32 GB | 14B comfortably, 32B at the edge | Qwen 2.5 14B: 11.7 tok/s (measured) | from $171/mo |
| M4 Pro Mac mini, 48 GB | 32B models with long context | 32B: about 12.3 tok/s (estimate) | from $257/mo |
| M4 Pro Mac mini, 64 GB | 70B at Q4 | 70B: about 5.7 tok/s (estimate) | from $286/mo |
| M4 Max Mac Studio, 128 GB | 70B at Q8, 123B at Q4 | 70B: about 11.3 tok/s (estimate) | from $471/mo |
| M3 Ultra Mac Studio, 256 GB | 200B class at Q4, 70B at FP16 | 70B: about 16.6 tok/s (estimate) | from $828/mo |
For any other model, quantization or context length, the Mac LLM calculator works out memory and speed. Yearly billing is cheaper: see pricing.
What comes on the machine
- Ollama with the model of your choice: Llama, Qwen, Mistral, DeepSeek and others
- An OpenAI compatible endpoint with your own API key, served from your machine
- Optional MLX and Jupyter, VS Code Server, and agent frameworks such as CrewAI and LangGraph
- SSH and a browser based remote desktop for anything beyond the API
- A physical Mac nobody else uses, with a hardware encrypted SSD
Calling your model
Every deployment gets its own URL. Point any OpenAI client at it:
from openai import OpenAI
client = OpenAI(
api_key="mcy_live_...",
base_url="https://dep-abc123.macyou.cloud/v1",
)
reply = client.chat.completions.create(
model="llama-3.3-70b",
messages=[{"role": "user", "content": "Hello!"}],
)
print(reply.choices[0].message.content)Full reference in the API docs.
Compared with the usual Ollama picks
The largest M4 Pro each provider sells, against the same machine at Macyou. None of them ships Ollama or an inference API; you install and expose it yourself.
| Provider | Machine | Their price | Same machine at Macyou |
|---|---|---|---|
| MacStadium | M4 Pro, 48 GB, 1 TB | $349/mo | $286/mo, Ollama and API included |
| MacinCloud | M4 Pro, 48 GB, 1 TB | $329/mo | $286/mo, Ollama and API included |
| AWS EC2 Mac | M4 Pro, 48 GB, 1 TB | $1,438/mo | $286/mo, Ollama and API included |
| Scaleway | M4 Pro, 64 GB, 2 TB | $369/mo | $371/mo, Ollama and API included |
| My Remote Mac | M4 Pro, 24 GB, 512 GB | $229/mo | $200/mo, Ollama and API included |
| JUUZ Cloud | M4 Pro, 48 GB, 512 GB | $263/mo | $257/mo, Ollama and API included |
Every provider side by side, with prices and terms: Mac cloud hosting compared.
Frequently asked questions
Can I host Ollama on a Mac in the cloud?
Yes. A Macyou machine is a dedicated Mac mini or Mac Studio that can come with Ollama preinstalled and exposed through an OpenAI compatible API. It is ready about 5 minutes after you deploy, and you can also reach it over SSH or a browser desktop.
Which Mac do I need for Llama 70B?
At Q4 a 70B model needs about 42 GB for weights plus room for context, so 64 GB is the entry point. An M4 Pro with 64 GB generates about 5.7 tokens per second; an M4 Max Mac Studio roughly doubles that, and an M3 Ultra with 256 GB also runs 70B at full precision.
How fast is Ollama on an M4 Mac mini?
On our fleet an M4 Mac mini with 16 GB generates 21.2 tokens per second on Llama 3.1 8B at Q4_K_M, measured with Ollama 0.31.2. Speed follows memory bandwidth, so the Pro, Max and Ultra chips are proportionally faster.
Is the API compatible with the OpenAI SDK?
Yes. The endpoint follows the OpenAI request format (/v1/chat/completions, /v1/models), so existing code works after you change the base URL and the API key. The model itself runs on your machine; no request goes to OpenAI.
Can I use MLX, llama.cpp or LM Studio instead of Ollama?
Yes. MLX and Jupyter come as an optional stack, and because you get SSH you can install llama.cpp, LM Studio's headless server or any other runtime into your own account.
Why a Mac instead of a GPU cloud?
Unified memory lets one machine hold models that would need several consumer GPUs, at a fixed monthly price with no per token or per hour billing. For a model that has to answer around the clock, that usually costs less than renting GPUs by the hour.