UK Mac buyback | Apple SiliconGet a free quote
Skip to content

How Much RAM to Run a 70B Model on a Mac?

By The SellMacBooks Team | Published 9 June 2026 | Updated 9 July 2026

Running a 70B model on a Mac needs roughly 40GB of unified memory for the weights at 4-bit, plus headroom for context - so 64GB is the practical floor and 128GB is comfortable. In Mac terms that means a MacBook Pro Max at 64GB or 128GB, or a Mac Studio Ultra for anything larger.

That is the short answer. The rest of this guide shows the arithmetic behind it, what quantisation trades away, how context quietly eats into your memory, and exactly which Macs clear the bar - so you can size a machine with confidence rather than guesswork.

What is the quick memory rule for a 70B model?

Start with one figure and build from there: at 4-bit precision, a model needs about half a gigabyte of memory for every billion parameters. That is the whole rule of thumb. A 70-billion-parameter model therefore lands near 35 to 40GB just to hold its weights in memory, before you account for anything else the machine is doing.

Round that up to 40GB as a working number and the reason for the recommendations becomes obvious. You cannot fit 40GB of weights into a 32GB machine at all, so 32GB is out. A 64GB machine holds the weights with roughly 20GB to spare - enough to run, provided you keep other demands modest. A 128GB machine holds them with room to breathe. The rule scales cleanly, too: an 8B model needs around 4GB, a 34B model around 17GB, and the very largest open models in the hundreds of billions of parameters push past 200GB. Once you know the parameter count, you can estimate the memory in your head.

How does the 0.5GB-per-billion arithmetic work?

It comes down to how many bits each parameter uses. A model's parameters are just numbers, and precision decides how many bits each one occupies. At 16-bit (the format many models ship in), every parameter needs two bytes, so a billion parameters need about 2GB. Quantise down to 4-bit and each parameter uses roughly half a byte instead - a quarter of the size - which lands you at that convenient 0.5GB per billion.

From there it is multiplication. Take the parameter count in billions, multiply by 0.5GB, and you have the approximate weight footprint at 4-bit. The table below applies exactly that method across common sizes. Treat these as indicative planning figures rather than precise measurements - real models vary a little with architecture and the exact quantisation method - but they are close enough to size a purchase.

Model sizeApprox. RAM at 4-bitComfortable Mac tier
7-8B~4-5GB16GB and up
13B~7GB16-24GB
34B~17GB36-48GB
70B~40GB64GB floor, 128GB comfortable
120B~65GB128GB
180B+~100GB+Mac Studio Ultra, 256GB+

The pattern is linear and easy to internalise: double the parameters, double the memory. That is why choosing a Mac for a specific model is really a memory decision dressed up as a chip decision.

What does quantisation trade away?

Quantisation is the lever that makes a 70B model fit a laptop at all, so it is worth understanding the cost. Storing weights at lower precision shrinks the model dramatically - a 70B model needs roughly 140GB at 16-bit, about 70GB at 8-bit, and near 40GB at 4-bit. Each halving of precision roughly halves the memory. That is how a portable machine ends up running a model that would otherwise demand a server.

The trade is a little accuracy. Higher precision preserves more of the model's original behaviour; heavier quantisation nudges its answers slightly, and the effect grows as you push below 4-bit. In practice 4-bit is the sweet spot most people settle on - the quality loss is small on everyday tasks and the memory saving is enormous. The buying lesson is that more memory buys flexibility twice over: you can run larger models, or run the same model at higher precision for sharper answers. If quality is critical, having the memory to step up from 4-bit to 8-bit is a genuine advantage.

How much extra memory does context need?

The weights are only part of the bill. Every model also holds a context window - the running record of your prompt and its reply - and that cache lives in memory too, growing with every token. A short question costs almost nothing, but feed the model a long document, a large codebase, or a lengthy conversation and the context can claim several gigabytes on top of the weights.

This is where the gap between "fits" and "comfortable" appears. On a 64GB Mac, a 70B model's 40GB of weights leaves around 20GB for context, the operating system and any apps you have open - workable for short prompts, but it tightens fast under a long context or a busy desktop. On a 128GB machine the same model leaves ample room for extended context and everything else at once. If your work involves long documents or big code files, budget for context generously and lean toward the higher tier - it is the single most common reason people wish they had bought more memory.

Which Macs actually clear the 70B bar?

Match the memory tier to the machine. For a 70B model, three Mac families qualify, and each suits a different need.

The MacBook Pro M4 or M5 Max at 64GB is the entry point - it holds a 70B model at 4-bit and runs it on your lap, which no ordinary laptop manages. The MacBook Pro Max at 128GB is the comfortable portable, giving you the same models with generous context headroom and the option to run higher-precision weights. Above that sits the Mac Studio M3 Ultra, whose 96GB, 256GB and 512GB tiers run not just 70B but the largest open models available, with desktop cooling built for hours of sustained inference. For a full config-by-config breakdown of what each tier costs and returns, see our guide to high-RAM Macs for AI.

As an indicative 2026 UK resale guide, a 128GB MacBook Pro M5 Max sits near £3,000 in good condition, while a Mac Studio M3 Ultra runs from roughly £2,500 at 96GB to £4,500 at 512GB. Those figures matter because they hold: the memory that lets your Mac run a 70B model is the same memory the next buyer pays a premium for.

Does the memory you buy hold its value?

Unusually well, and that reframes the cost of building a 70B-capable machine. Because unified memory is fixed at purchase and cannot be added later, the used market treats each tier as its own product, and high-memory configurations resell for a strong share of what they cost new. The headroom you pay for to run large models is not money spent - it is largely money parked, recoverable when you move to a bigger rig.

That durability is why buying up the memory ladder rarely stings later. If your workload grows past 70B and you step up to a higher Ultra tier, the machine you replace is worth real money to another builder chasing exactly its memory. When that day comes you can sell your AI Mac through our dedicated route, or if you own the desktop, sell a Mac Studio M3 Ultra directly - with a fixed quote in two working hours, free insured collection, and payment within 24 hours of inspection. Size the memory for the model you want to run, and the resale takes care of itself.

Good to know

Quick answers

Everything sellers ask us before posting their Mac - answered straight.

Can 64GB of unified memory run a 70B model on a Mac?

Yes, at 4-bit quantisation with a modest context window. The weights take around 40GB, leaving roughly 20GB for context, the operating system and other apps. It works, but it is tight - close background apps and keep context short. For long documents or higher precision, 128GB is the safer choice.

Why do you estimate 0.5GB per billion parameters?

At 4-bit precision each parameter uses about half a byte, so a billion parameters need roughly 0.5GB of memory for the weights. Multiply by the model size: a 70B model lands near 40GB. Higher precision doubles or quadruples that, which is why quantisation matters so much on a Mac.

Is 128GB overkill for a 70B model?

Not if you want headroom. A 70B model fits in 64GB, but 128GB lets you run longer context windows, keep other apps open, or step up to higher-precision weights for better answers. It also runs models beyond 70B, so it is the comfortable long-term choice rather than overkill.