On a Mac, the memory question is not a side spec, it is the whole ceiling. Mac unified memory serves as both system RAM and GPU memory at once, which means the number you pick at checkout sets a hard limit on which local AI models will ever run on that machine. Choose 32GB and certain large models are simply off the table forever. The model size you want decides the Mac you buy, not the other way around.
Quick Answer
A base M5 caps at 32GB of unified memory, comfortable for models in the 13B to 30B range. Step up to an M5 Max at 128GB and 70B models run at roughly 12 to 18 tokens per second with room for long context. The rule of thumb: keep model weights to about 60 percent of your unified memory so there is headroom for the context window. Pick your memory tier by the model you intend to run.
Why unified memory sets the ceiling
Apple Silicon routes processing through a single high-bandwidth memory pool shared by the CPU, GPU and Neural Engine alike. There is no separate VRAM to fill, so a model loads once and runs without copying data between memory pools. That is a real strength: a 70B model at 4-bit precision weighs in at roughly 35GB, beyond any single consumer graphics card, yet fits neatly on a Mac with enough unified memory.
The flip side is that the model, plus the overhead of its context cache, must all fit inside that one pool. There is no overflow trick that keeps speed intact. So the memory you buy is the memory you have, and it directly caps how large a model you can load.
The tiers, model by model
Think of it as a staircase, each memory level unlocking a model class.
A 32GB Mac, the ceiling on a base M5, is the practical minimum for 13B to 30B models. It can hold a model of around 28GB and run inference on the GPU directly, which covers a lot of capable everyday models with room for a sensible context window.
A 64GB Mac, available on the M5 Pro tier, opens the door to larger models with extended context and is increasingly treated as the comfortable baseline for serious local AI work.
A 128GB Mac, the M5 Max, runs 70B-class models at around 12 to 18 tokens per second, usable speed for real work, and has the headroom for very long prompts. At 128GB and above, even some very large mixture-of-experts models become practical to run locally.
The guiding principle across all of them is the same: target model weights at no more than about 60 percent of total memory, leaving the rest for the context cache that grows as your prompt lengthens. The AI PCs at Evetech are specced for this kind of local inference work.
Buy for the model, not the badge
The mistake is buying the Mac first and discovering the memory ceiling later. Unified memory is fixed at purchase and cannot be upgraded afterwards, so the decision is permanent. Decide which model tier you actually need, 13B, 30B or 70B, add the context headroom, then choose the memory configuration that comfortably holds it. The wider best-selling desktops at Evetech are worth comparing too if a discrete-GPU build suits your workload better than a unified-memory machine.
Frequently Asked Questions
How much unified memory do I need for a 70B model?
Around 128GB for comfortable headroom. A 70B model at 6-bit quantisation needs roughly 61GB just for weights, and you want context room on top, so an M5 Max class machine is the realistic target.
Is 32GB enough for local AI on a Mac?
Yes, for 13B to 30B models. 32GB is the practical minimum for that range and handles a lot of useful models, but it cannot stretch to 70B-class workloads no matter the quantisation.
Can I upgrade unified memory later?
No. Apple silicon memory is fixed at purchase and cannot be added afterwards. That is why choosing the right tier upfront, based on your target model, matters so much.
Why does unified memory beat a separate graphics card for big models?
Because the whole pool is available to the GPU at once with high bandwidth and no copying. A 70B model that will not fit on a consumer graphics card runs fine on a Mac with enough unified memory.
The model you want to run is the spec that should pick your machine. Browse the AI PC range at Evetech and match the memory to your model tier before you commit, since you only get one chance to choose it.