📚 More on this topic: VRAM Requirements Guide · Quantization Explained · Multi-GPU Guide · Mac vs PC for Local AI · Used RTX 3090 Buying Guide

Running a 70B model locally used to be the line between “hobby” and “serious local AI” — frontier-class reasoning on your own hardware, no data leaving the box. In mid-2026 that line has blurred. Nothing dense replaced Llama 3.3 70B (Meta went all-MoE with Llama 4, and Qwen dropped its large dense tier entirely), while a single 24GB card now runs Qwen 3.6-27B or an MoE like Llama 4 Scout at 70B-class quality. The dense 70B is still the deepest-reasoning option you can self-host, but it’s no longer the obvious one.

The barrier, when you do want a dense 70B, is still VRAM. A 70B model at full precision needs 141GB of memory. No consumer GPU comes close to that. Quantization brings it down to 43GB at Q4, which still won’t fit on a single RTX 4090 or 3090. You need either two GPUs, a Mac with enough unified memory, or a workstation-class card.

This guide gives you exact VRAM numbers at every quantization level, which hardware setups actually work, realistic speed expectations, and an honest assessment of when a dense 70B is worth the investment versus running a 27B or MoE model instead.


The 70B Math

The formula is simple:

VRAM (GB) = Parameters (billions) × Bytes per parameter

At FP16 (2 bytes per parameter): 70B × 2 = 140GB. That’s the model weights alone. Context and framework overhead are extra.

Quantization compresses those weights:

PrecisionBytes per ParamWeight Size (70B)With Overhead*
FP162.0140 GB~142 GB
Q8_01.070 GB~75 GB
Q6_K0.7552.5 GB~58 GB
Q5_K_M0.62543.75 GB~50 GB
Q4_K_M0.535 GB~43 GB
Q3_K_M0.37526.25 GB~35 GB
Q2_K0.2517.5 GB~27 GB

*Overhead includes KV cache at 4K context, framework memory, and CUDA/Metal context. Real GGUF files are slightly larger than the theoretical minimum due to metadata and mixed-precision layers.

The theoretical calculation gets you in the ballpark. Real file sizes are what matter. See the next section.


Exact VRAM: Llama 3.3 70B and Qwen 2.5 72B

Llama 3.3 70B is still the dense 70B reference — nothing dense replaced it. Qwen 2.5 72B below is the other big dense model people run, but it’s now legacy: Qwen’s 3.5 and 3.6 lines top out at 27B dense and moved everything larger to MoE (covered in its own section below). Numbers from actual GGUF builds on HuggingFace:

Llama 3.3 70B Instruct

QuantizationFile SizeVRAM Needed (4K ctx)VRAM Needed (8K ctx)
FP16141.1 GB~143 GB~148 GB
Q8_075.0 GB~77 GB~82 GB
Q6_K57.9 GB~60 GB~65 GB
Q5_K_M50.0 GB~52 GB~57 GB
Q4_K_M42.5 GB~45 GB~50 GB
Q3_K_M34.3 GB~37 GB~42 GB
Q2_K26.4 GB~29 GB~34 GB

Qwen 2.5 72B Instruct

QuantizationFile SizeVRAM Needed (4K ctx)VRAM Needed (8K ctx)
Q8_077.3 GB~79 GB~84 GB
Q6_K64.4 GB~66 GB~71 GB
Q5_K_M54.5 GB~57 GB~62 GB
Q4_K_M47.4 GB~50 GB~55 GB
Q3_K_M37.7 GB~40 GB~45 GB
Q2_K29.8 GB~32 GB~37 GB

Qwen 2.5 72B is about 10-15% larger than Llama 3.3 70B at the same quantization because it has 72 billion parameters versus 70.6 billion, plus slightly different architectural choices. Both produce similar quality at the same quant level. Starting fresh in 2026, though, Llama 3.3 70B is the better-supported dense pick — and if you specifically want Qwen at scale, you’d reach for the MoE Qwen 3.5-122B-A10B, not the old 72B dense.

For a deeper understanding of what these quantization levels mean and how they affect output quality, see our quantization explainer.

→ Use our Planning Tool to check exact VRAM for your setup.


Context Length Eats Your VRAM

The tables above assume 4K or 8K context. But Llama 3.3 supports 128K tokens and Qwen 2.5 72B supports 128K too. The KV cache (where the model stores attention state for the conversation) grows linearly with context length.

KV Cache VRAM at 70B Scale

Context LengthKV Cache (FP16)KV Cache (Q8)KV Cache (Q4)
4K tokens~2.4 GB~1.2 GB~0.6 GB
8K tokens~4.9 GB~2.4 GB~1.2 GB
16K tokens~9.8 GB~4.9 GB~2.4 GB
32K tokens~14 GB~7 GB~3.5 GB
64K tokens~28 GB~14 GB~7 GB
128K tokens~39 GB~20 GB~11 GB

At 32K context with FP16 KV cache, you’re adding 14GB on top of the model weights. On a dual RTX 3090 setup (48GB total) running Llama 70B Q4_K_M (42.5GB file), that leaves about 5.5GB for everything. With a 14GB KV cache, you’re already over.

This is why most 70B setups run with 4K-8K context and why the 128K advertised context length is mostly theoretical for consumer hardware. You can extend it with quantized KV cache (Ollama and llama.cpp both support this), but even then, 32K+ context on 48GB total VRAM is tight.


Hardware That Can Actually Run 70B

Single Consumer GPUs: Mostly Can’t

GPUVRAMBest 70B QuantContextVerdict
RTX 4060 (8GB)8 GBNoneNot happening
RTX 3060 12GB12 GBNoneNot happening
RTX 4060 Ti 16GB16 GBNoneNot happening
RTX 3090 / 409024 GBNone (Q2_K is 26.4GB)Doesn’t fit even at Q2
RTX 509032 GBQ2_K or Q3_K_M~4K tokensTechnically works. Quality is poor at Q2, marginal at Q3.

The RTX 5090 is the only consumer GPU that can load a 70B model at all. Q3_K_M (34.3GB file) fits with about 4K context, but quality degrades noticeably at Q3 and you have zero headroom. It’s a proof-of-concept, not a daily driver.

Dual GPU Setups

This is where 70B becomes practical on consumer hardware. Two GPUs pool their VRAM.

SetupTotal VRAMBest QuantContextSpeedCost (Aug 2026)
2× RTX 309048 GBQ4_K_M~4-8K16-21 tok/s~$2,200-2,500
2× RTX 409048 GBQ4_K_M~4-8K20-25 tok/s~$4,300-4,800
2× RTX 509064 GBQ4_K_M~16-32K25-30 tok/s~$7,800-8,800

Dual RTX 3090s (~$2,200-2,500 total) is the budget path. 48GB runs Llama 3.3 70B at Q4_K_M with 4-8K context. You get 16-21 tokens per second, which is readable but noticeably slower than the 40+ tok/s you’d get from a 27B model on a single card. Note the price: 3090s appreciated hard through 2026, so this build costs ~$500-800 more than it did a year ago. See our multi-GPU guide for setup instructions.

Dual RTX 5090s (~$7,800-8,800) with 64GB total opens up longer context. Q4_K_M with 16-32K tokens is comfortable, and Q5_K_M becomes viable for better quality. That is close to nine thousand dollars of graphics card, though, and a single 24GB card running gpt-oss 120B or Llama 4 Scout is the smarter buy for most people.

Both setups require a motherboard with two PCIe x16 slots (or at least x16 + x8), a 1000W+ power supply, and good airflow. Two 3090s at full inference draw 700+ watts combined.

Workstation / Datacenter GPUs

GPUVRAMBest QuantContextSpeedPrice
A600048 GBQ4_K_M~4-8K12-16 tok/s~$2,600-3,800 used
A100 80GB80 GBQ5_K_M~16K+19-22 tok/s~$4,000-9,000 used

The A6000 at $2,600-3,800 used gives you the same 48GB as dual 3090s in a single card, no multi-GPU hassle. But it’s slower for inference (lower memory bandwidth) and now costs more than a pair of 3090s. A100 80GB used prices have softened toward the low end ($4,000) as enterprise Ampere leases expire, but it’s still overkill for a single 70B.

Mac (Unified Memory)

This is where Macs win. Unified memory lets the entire RAM pool serve as model memory.

Mac speeds below are community/guidance figures — we bench on a 3090, not Apple silicon — so treat them as the shape of it, not a firsthand measurement.

ConfigUnified MemoryBest QuantContextSpeedPrice (Jul 2026)
Mac Mini M4 Pro 48GB48 GBQ4_K_M~4-8K6-8 tok/s~$2,000
Mac Studio M4 Max 64GB64 GBQ4_K_M~8-16K8-12 tok/s~$2,500-3,000
Mac Studio M3 Ultra 96GB96 GBQ6_K~32K+12-18 tok/s$5,299
M5 Max MacBook Pro 128GB128 GBQ8~32K+10-15 tok/s~$5,800+

Mac speeds are slower than dual NVIDIA GPUs because unified memory bandwidth (546 GB/s on M4 Max, 614 on M5 Max, 819 on M3 Ultra) is lower than GDDR6X (936 GB/s per 3090). But the Mac loads the model at all, which a single 24GB GPU can’t. And it does it silently, at 15 watts idle.

Here’s the part that changed in 2026, and it upends the old advice this guide used to give. The comfortable 70B Mac was the M4 Max Studio at 128GB. That config no longer exists — the DRAM shortage cut the M4 Max Studio to 64GB and the M3 Ultra Studio to 96GB (both 256GB and 512GB gone), and Apple discontinued the 192GB Mac Pro outright. The 96GB M3 Ultra ($5,299, 819 GB/s) is now the roomiest 70B desktop you can buy, and it’s the fastest Mac too. For 128GB you have to go to a laptop — the M5 Max MacBook Pro — because there’s no longer a 128GB Mac desktop at all. A rumored M5 Ultra Studio may land later in 2026, but given the shortage has been removing high-memory configs, don’t buy on the assumption it’ll bring them back. See our Mac vs PC comparison for the full breakdown.


Speed Expectations

70B models are slow. Set your expectations accordingly.

HardwareLlama 3.3 70B Q4_K_MContext
2× RTX 3090 (48GB)16-21 tok/s4-8K
2× RTX 4090 (48GB)20-25 tok/s4-8K
2× RTX 5090 (64GB)25-30 tok/s16K
Mac M3 Ultra 96GB12-18 tok/s32K
Mac M5 Max MBP 128GB10-15 tok/s32K
Mac M4 Max 64GB8-12 tok/s8-16K
A100 80GB19-22 tok/s16K
Single 24GB GPU + CPU offload1-5 tok/s4K

For comparison, Qwen 3.6-27B at Q4_K_M on a single RTX 3090 runs at 35-45 tok/s. A dense 70B model on dual 3090s runs at about half that speed while costing twice the hardware — and an MoE like gpt-oss 120B on that same single 3090 runs faster than either.

CPU offloading (splitting the model between GPU and system RAM) technically works but is painfully slow. The PCIe bus becomes the bottleneck, dropping generation to 1-5 tok/s. At that speed, you’re waiting 10-20 seconds for a single sentence. It’s fine for testing. It’s not usable for daily work.


Quality vs Quantization at 70B

Good news: 70B models tolerate quantization better than smaller models. Research confirms that models above 30B parameters retain ~99% of FP16 accuracy at 4-bit quantization, while 7B models lose 2-5%.

Quality by Quant Level

QuantizationQuality RetentionBest For
Q8_0~99.5% of FP16Maximum quality when VRAM allows
Q6_K~99% of FP16Excellent. Hard to distinguish from Q8 in practice
Q5_K_M~97-99% of FP16Great balance. Most users won’t notice the difference
Q4_K_M~95-97% of FP16The sweet spot. Minor degradation on complex reasoning
Q3_K_M~90-93% of FP16Noticeable. Reasoning and math tasks suffer first
Q2_K~80-85% of FP16Severe. Unpredictable behavior on hard problems. Skip this.

Q4_K_M is the recommendation for almost everyone running 70B locally. The 3-5% quality loss versus FP16 is barely perceptible in normal use. You’d need benchmark suites to measure the difference reliably. The VRAM savings (142GB down to 43GB) make it the only practical option on consumer hardware.

Q3_K_M is where you start noticing. Math problems that Q4 handles cleanly will occasionally fail at Q3. Multi-step reasoning chains break more often. If you’re running on an RTX 5090 and Q3 is your only option, it works. Just know you’re leaving quality on the table.

Q2_K is not worth running. At 70B, even the higher quantization tolerance can’t save Q2 from significant output degradation. If Q2 is your only option, run a 27B model or an MoE at Q4 instead. You’ll get better results.


The MoE shortcut: 70B-class quality without 48GB

Here’s what changed since this guide first ran. The biggest shift in local AI over the past year isn’t a faster dense 70B — it’s that mixture-of-experts (MoE) models now hit 70B-class quality while activating only a fraction of their parameters per token. Big-model quality at small-model VRAM and speed.

An MoE model has a large total parameter count but only activates a few experts per token. Llama 4 Scout is 109B total but 17B active. gpt-oss 120B is 120B total but only ~5B active. You pay VRAM roughly for the total quantized weights, but you pay compute and bandwidth only for the active slice, so they generate fast.

ModelTotal / activeVRAM (Q4 or native)Fits onSpeed
gpt-oss 120B120B / ~5B~14 GB native MXFP4Single 16-24GB cardVery fast
Llama 4 Scout109B / 17B~24 GB at Q4Single 24GB card + 32K ctxFast, like a 17B
Qwen 3.6-35B-A3B35B / 3B~18-22 GB at Q4Single 24GB cardFaster than a dense 9B
Qwen 3.5-122B-A10B122B / 10B~67 GB at Q480GB card, 96GB M3 Ultra, or 128GB M5 Max70B-class, needs the big tier

The practical upshot: for most people asking “how do I run a 70B,” the honest 2026 answer is “run gpt-oss 120B or Llama 4 Scout on the 24GB card you already have.” You get frontier-adjacent quality without the dual-GPU build, the 700W power draw, or the $2,200+ hardware bill. gpt-oss 120B loads in about 14GB with its native MXFP4 weights (ollama pull gpt-oss:120b) and runs faster than a dense 70B ever could.

Dense 70B still wins on the deepest single-shot reasoning. MoE routing occasionally sends a token to the wrong expert, and it shows up on the hardest problems. But for coding, chat, RAG, and general work, the MoE path is faster, cheaper, and fits hardware you can actually buy today.


When 70B Is Worth It

Run 70B For:

Complex reasoning. Multi-step logic problems, mathematical proofs, scientific analysis. The gap between a 27B and a dense 70B is widest here. A 70B model at Q4 catches errors and follows chains of reasoning that a 27B model misses.

Deep research and analysis. Summarizing long documents, comparing multiple sources, identifying inconsistencies. 70B models have broader knowledge and make fewer factual errors.

Nuanced writing. When you need precise tone control, subtle arguments, or professional-grade output. 70B models handle ambiguity and subtext better.

Skip 70B For:

Quick chat and Q&A. A 27B model answers “what’s the capital of France” just as correctly, 3-4x faster.

Simple code generation. For boilerplate, function scaffolding, and straightforward coding tasks, Qwen 3.6-27B is more than sufficient and much faster — it actually beats the old Llama 3.3 70B on SWE-bench Verified. gpt-oss 120B (MoE) is another fast option on a single 24GB card.

Anything speed-sensitive. If you need responses in under 2 seconds, a dense 70B won’t deliver. A 27B model at 40 tok/s starts generating immediately. A dense 70B at 15 tok/s has noticeable latency.

The 27B and MoE alternative

This is the honest question: in 2026, do you need a dense 70B?

Qwen 3.6-27B (Q4_K_M)Llama 3.3 70B (Q4_K_M)
VRAM needed~16 GB~43 GB
HardwareSingle RTX 3090 (~$1,100)Dual RTX 3090 (~$2,200-2,500)
Speed35-45 tok/s16-21 tok/s
Complex reasoningGoodSlightly better
Factual depthGoodBetter
Coding (SWE-bench Verified)77.2% — ahead of the old 70BBehind the 27B

The gap didn’t just narrow — on coding it flipped. Qwen 3.6-27B (dense, Apache 2.0) posts 77.2% on SWE-bench Verified and fits a single 24GB card at Q4 with room for context. Gemma 4 and GLM-5.1 sit in the same tier. And the MoE options from the section above — gpt-oss 120B, Llama 4 Scout — give you 70B-class breadth on that same single card, faster than a dense 70B ever ran.

A dense 70B still wins on the deepest multi-step reasoning and raw factual recall. If that’s your primary use case, the dual-GPU investment is justified. For coding, chat, RAG, and general work, a 27B or an MoE on a single GPU is faster, cheaper, and at least as good.


Bottom Line

Running a dense 70B locally requires either dual GPUs (2× RTX 3090 at ~$2,200-2,500), a Mac with 64GB+ unified memory ($3,000+), or a datacenter card. Q4_K_M is the quantization sweet spot: 43GB for Llama 3.3 70B, excellent quality retention. Below Q4, quality drops noticeably. Below Q3, don’t bother.

But the bigger 2026 answer is to skip the dense 70B for most work. gpt-oss 120B and Llama 4 Scout deliver 70B-class quality on the single 24GB card you probably already have — faster, and with no dual-GPU build. Qwen 3.6-27B fits the same card and beats the old 70B on coding.

If you still want a dense 70B, for the deepest reasoning or maximum factual recall, dual RTX 3090s with Llama 3.3 70B at Q4_K_M is the practical build. You get 16-21 tok/s at 4-8K context. It’s slower than a 27B or an MoE and costs twice the hardware, but on hard reasoning the difference is real.

Not sure? Start with an MoE or Qwen 3.6-27B on a single 24GB GPU. It handles the large majority of tasks. Add the second GPU for a dense 70B only when you consistently hit the quality ceiling on reasoning-heavy work.



Last updated July 17, 2026. Corrected the Mac 70B options for Apple’s 2026 memory cuts: the 128GB M4 Max Studio and 512GB M3 Ultra are gone, the Studio now caps at 96GB ($5,299), and 128GB is laptop-only (M5 Max MacBook Pro). GPU and MoE sections unchanged.