📚 Related: GPU Buying Guide · VRAM Requirements · Used RTX 3090 Buying Guide · Best Local LLMs for Mac (2026)

The $4,000 NVIDIA AI mini-PC isn’t $4,000 anymore. The DGX Spark Founders Edition jumped to $4,699 on February 23, 2026, an 18% hike NVIDIA attributed to memory-supply constraints (NVIDIA Developer Forums; Tom’s Hardware). Hardware unchanged. A lot of comparison posts still show the old price.

The bigger question this guide answers isn’t which GB10 box to buy — it’s whether you should buy into the 128GB-unified tier at all. For a lot of readers, the honest answer is no. Here’s the field check, the three numbers that actually decide it, and an honest used-GPU gut-check before you spend $3,000-5,000.


The Three Numbers That Decide

Before any product table: the local-AI hardware decision pivots on three numbers, in this order.

  1. How big a model do you actually need to fit? Anything 24GB can hold, a discrete consumer GPU wins on speed. Above 24GB at usable quants — 70B FP8, 122B-A10B, 200B at FP4 — you need unified memory, and that’s where the GB10 / Strix Halo / Mac Studio tier exists.
  2. How fast do you need it? Token generation is memory-bandwidth-bound. RTX 3090 has ~936 GB/s. Mac Studio M3 Ultra has ~819 GB/s. GB10 has ~273 GB/s. Strix Halo unified memory is around ~256 GB/s. Bigger memory pool, slower per-token throughput.
  3. What software stack do you need? CUDA (TensorRT-LLM, vLLM, PyTorch CUDA-native), Windows-with-AMD (Strix Halo gives you both), or macOS/MLX. None of these tiers is software-agnostic.

The reason “which box should I buy” has no single answer: those three numbers point different ways for different workloads. The rest of this guide pins them to specific products at July 2026 prices.


The Used-3090 Reality Check

Before sorting which 128GB-unified box, the question worth asking first: what does a used RTX 3090 actually do for $1,000-1,300?

Community-reported benches:

WorkloadUsed RTX 3090 (24GB)GB10 / Strix Halo / Mac (128GB unified)
Models that fit in 24GB (7B-32B dense, 35B-A3B MoE)~50 tok/s on 14B at 16K context (Hardware Corner)~30-60 tok/s on same models
Llama 70B Q5 / FP8Doesn’t fit; needs offload~5 tok/s (Strix Halo dense 70B Q4, community-reported), ~2-3 tok/s (GB10 FP8)
Qwen 3.5-122B-A10B (FP4/Q4)Doesn’t fit~9.5 tok/s (Strix Halo, community-reported)
Memory bandwidth~936 GB/s~256-273 GB/s (GB10/Strix Halo) / ~819 GB/s (M3 Ultra)
Power draw~350W under load~100W (GB10), ~120W (Strix Halo), ~150W (M3 Ultra)
Needs host PCYesNo (standalone)

The 3090 beats every mini-PC on speed for any model that fits in its 24 GB of VRAM, and gets crushed on anything that doesn’t fit. That’s it. That’s the whole trade. For coding agents, chat with 7B-32B, vision models that fit, image generation — the 3090 is faster and cheaper. For Llama 70B FP8, Qwen 3.5-122B-A10B, anything 100B+, the 3090 can’t load it at all.

If your actual workload sits inside 24 GB at usable quantization, the most honest GB10 advice is: don’t buy one. Buy a used 3090 and put the difference toward more memory in your host system or a second card. The 128GB-unified boxes are for workloads that don’t fit in 24GB. That’s the whole story.

If your workload genuinely does need the 128GB tier, the rest of this guide is for you: the GB10 field, the Strix Halo alternative, and where Mac Studio fits.


The GB10 Field, Current Pricing

Every GB10 box runs the same NVIDIA Grace Blackwell GB10 superchip — 20 ARM v9.2 cores (10 performance + 10 efficiency), Blackwell GPU with 6,144 CUDA cores and 192 5th-gen Tensor cores, 128 GB LPDDR5X unified memory at 273 GB/s, ~1 PFLOP sparse FP4 compute, 140W TDP, ConnectX-7 200GbE networking, DGX OS (Ubuntu-based; no Windows). The chip is the chip.

BoxSource / SKUNotableCurrent US price
NVIDIA DGX Spark Founders EditionNVIDIA direct (marketplace)4TB SSD, Gen 5 NVMe, reference design$4,699 (up from $3,999)
ASUS Ascent GX10Amazon, 1TB SKUFront power button; only GB10 with one~$3,099 (1TB) / ~$4,150 (4TB) (TechRadar, IT Pro)
Dell Pro Max (GB10)Dell direct, 2TB SKUMagnetic back panel, enterprise warranty~$3,699 (2TB) / ~$3,999 (4TB) (IT Pro)
MSI EdgeXpert MS-C931MSI, 1TB SKUMost labeled ports; plastic chassis; lineup recently reshuffled~$2,999 (1TB direct) / ~$3,999 (Amazon 4TB)
Acer Veriton GN100Acer direct USStorage Review reports best thermals of the field$3,999 (Acer US MSRP; CDW promo ~$3,702) (TechRadar review)
Gigabyte AI TOP ATOMGigabyte direct / NeweggNow shipping (was “announced”)~$4,662 (4TB)
Lenovo ThinkStation PGXLenovo directNow shipping (was “announced”)~$4,100 (1TB) – $5,079
HP ZGX Nano G1nHP directNewest entrant; priciest GB10 box~$6,030 (4TB)

The performance differences between these boxes are negligible. They run the same SoC, hit the same software-imposed ~100W GPU power cap (the cap John Carmack noticed when reviewing the Spark — same cap is on every OEM build), and bench within margin on the same models. What you’re paying for is chassis quality, NVMe generation, support/warranty, and storage capacity.

The Spark Founders Edition is the only one with Gen 5 NVMe. That means ~13 GB/s sequential read vs ~7 GB/s on the Gen 4 SSDs in every OEM box, which translates to noticeably faster cold model loads. Per OEM reviews on the existing community testing, the gap is roughly a 25% reduction in cold-load time for the same model — meaningful if you swap models frequently, irrelevant if you load one model and run it all day.

The other differentiators (chassis materials, port labeling, thermal behavior under sustained load) come from the published OEM reviews — TechRadar on ASUS, IT Pro on Dell, Storage Review on Acer’s thermals. No GB10 OEM has been independently benched here. Treat the chassis differences as real but small; the performance difference is essentially zero.


The Strix Halo Field — AMD’s GB10 Challenger

This is what shifted in the months since the original DGX Spark launch: AMD’s Ryzen AI Max+ 395 (“Strix Halo”) shipped 128GB unified memory in mini-PCs at prices that undercut every GB10 box, with Windows 11 support that the GB10 lineup doesn’t have. The chip itself is 16 Zen 5 cores at up to 5.1 GHz, Radeon 8060S iGPU (40 compute units), and up to 96 GB of the 128 GB unified pool allocatable as VRAM via AMD’s Variable Graphics Memory.

The GB10’s CUDA premium is real — TensorRT-LLM, vLLM, and PyTorch CUDA-native paths just work on GB10 and don’t on Strix Halo. But for builders who don’t need the CUDA-specific tooling, Strix Halo is in the same memory-capacity category at much lower prices.

BoxNotableCurrent US price
GMKtec EVO-X2Still the cheapest 128GB Strix Halo build~$1,999-2,299 (on sale from $2,199 MSRP; up from ~$1,499 earlier in 2026) (Amazon listing)
Corsair AI Workstation 300Mainstream brand; multiple SSD tiers$2,699 (1TB) / $3,399 (4TB) (TechPowerUp coverage)
AMD Ryzen AI Halo Developer PlatformAMD’s reference build; just launched July 2026 (Micro Center exclusive)$3,999 (Phoronix)
NextNuc AI395Best Buy distribution~$3,699 base (128GB/1TB) / $4,999-5,999 higher-storage tiers (Best Buy listing)
HP Z2 Mini G1aEnterprise-warranted Strix Halo box; HP’s first AI-targeted Z Mini~$3,342-4,781 (128GB, config-dependent)

A few clarifications worth knowing about Strix Halo:

  • The HP Z2 Mini G1a is a Strix Halo box, not a GB10 box. Some 2026 comparison posts have miscategorized it — the Z2 Mini G1a uses AMD Ryzen AI Max+. Confusingly, HP now also ships a separate GB10 box, the ZGX Nano G1n in the table above, so check which HP machine a listing actually means.
  • Strix Halo memory bandwidth is in the ~256 GB/s range for the unified pool, similar to the GB10’s 273 GB/s and far below discrete-GPU bandwidth.
  • Community-reported benches show Strix Halo around 5 tok/s on a dense 70B at Q4 (Level1Techs measured Shisa V2 70B i1-Q4_K_M at 5.0 tok/s; llm-tracker.info). The 8-10 and 30+ tok/s numbers that circulate are MoE models (Llama 4 Scout, Qwen3-30B-class), not dense 70B — a dense 70B is pinned to the ~256 GB/s bandwidth ceiling, and newer ROCm/llama.cpp doesn’t move it much. Usable for batch work, not interactive, the same band as GB10.

If you want CUDA, this isn’t an answer. If you want 128GB unified memory at a lower price with Windows native, this is the answer. The mid-2026 reality.


The Mac Studio Comparison

The third box in this category: Mac Studio — and the story here changed in mid-2026. The RAM shortage walked the whole line down. Apple pulled the 512GB M3 Ultra in March and the 256GB in May, and the 128GB M4 Max config is gone too, so by mid-2026 the entire Mac Studio lineup tops out at 96GB unified memory (the M3 Ultra base). No M5 Mac Studio has shipped, and an M5 Ultra — rumored to scale up to 768GB — isn’t expected until around October 2026, delayed by the same memory supply crunch. The upshot: Mac Studio no longer matches the GB10 and Strix Halo boxes on raw capacity. It’s a step behind at 96GB.

Where it still wins is bandwidth. The M3 Ultra ships at 819 GB/s, roughly 3× the GB10’s 273 GB/s, so anything that fits in 96GB runs far faster than on a GB10 or Strix Halo — a 70B at Q4-Q5 (~40-50GB) is comfortable, and even a 122B-A10B at Q4 (~60GB) fits with room for context. What 96GB can’t reach is the 200B-class at FP4 (~100GB+) that the 128GB boxes can attempt. That capacity ceiling is the new trade.

The other trade is software stack. No CUDA. Inference goes through Apple’s MLX or llama.cpp Metal backend; that’s a different toolchain than Linux GB10 or Windows Strix Halo. For builders inside the Apple ecosystem whose models fit in 96GB, the M3 Ultra is the bandwidth king of this class. For builders who need CUDA or the full 128GB pool, it’s a non-starter.

Apple also raised Mac prices across the board in June 2026: the M3 Ultra Studio base jumped ~$1,300 to $5,299 (96GB / 1TB). That’s more than a Spark Founders and in the range of a loaded OEM GB10 — less memory, but far more bandwidth on what fits. The Best Local LLMs for Mac (2026) guide covers the Apple-side picks in detail.


Real-World Performance Caveats (Mid-2026)

The marketing line on GB10 is “200B parameters on your desk.” The real-world picture is more measured. A few patterns worth knowing before deploying any of these:

  • Token generation is bandwidth-capped. At ~273 GB/s, a GB10 box generates Llama 70B FP8 at roughly 2-3 tok/s and Llama 70B FP4 (via TensorRT-LLM) at around 5 tok/s. That’s usable for batch processing and agent orchestration; it’s painful for interactive chat. Time-to-first-token on a 90B+ model can hit the multi-minute range.
  • Software-imposed 100W power cap. Every GB10 box hits a ~100W GPU power cap, not the 240W power-supply rating. This is by design, not thermal — CPU clocks don’t drop when it engages. NVIDIA could lift it via firmware. Carmack flagged this early; later coverage clarified it’s a software cap, not throttling.
  • Thermal headroom varies by chassis. Storage Review’s testing put the Acer Veriton GN100 as the coolest of the field at 76°C peak; the ASUS Ascent GX10 has the only reports of triggered thermal-slowdown events under sustained stress. None of the boxes throttle under normal LLM inference workloads.
  • Early-buyer reliability reports. Community forums in the first weeks of GB10 shipping flagged sporadic WiFi reset issues, DGX OS update friction, and one-off thermal management oddities. NVIDIA has shipped firmware updates since (a June 2026 DGX Spark software release added multi-node clustering and NVFP4 plus multi-token-prediction gains, up to ~2.6× on Qwen3-35B), so first-batch issues are largely background now. If you’re buying a GB10 box now, you’re buying into a more stable platform than the launch-month cohort.
  • No Windows. Every GB10 box runs DGX OS (Ubuntu-based). For builders whose toolchain is Windows-native, that’s the deal-breaker that pushes the decision to Strix Halo regardless of price.

Decision Tree

Walk this from the top; first match wins.

  1. Your workload fits in 24GB at usable quant (anything 7B-32B dense, 35B-A3B MoE)?Used RTX 3090 ($1,000-1,300). Faster than any mini-PC on these models. Don’t overspend.
  2. You need 70B-200B at usable quants AND CUDA-specific tooling (TensorRT-LLM, vLLM, PyTorch CUDA paths)?GB10 box. Pick DGX Spark Founders ($4,699) for Gen 5 NVMe and the reference design, or a cheaper OEM twin (MSI / ASUS / Dell from $2,999 at 1TB; Acer, Gigabyte, Lenovo and HP’s ZGX Nano run higher) with negligible performance loss.
  3. You need 70B-200B at usable quants AND Windows support OR a tighter budget?Strix Halo (Ryzen AI Max+ 395). $1,999 (GMKtec) to $3,999 (AMD reference / NextNuc). 128GB unified, 96GB allocatable as VRAM, Windows 11. No CUDA.
  4. You’re in the Apple ecosystem, your models fit in 96GB, AND have $5K+ to spend?Mac Studio M3 Ultra 96GB ($5,299). 3× the memory bandwidth of GB10 / Strix Halo, fastest on big-model token throughput in this category — but Apple’s RAM cuts capped the whole line at 96GB, so it can’t reach the 200B-class the 128GB boxes attempt. No CUDA. See Best Local LLMs for Mac 2026.
  5. You want both speed AND capacity? → Dual used 3090s with a real host PC, ~$2,300-2,600 for the GPUs plus a system. Total VRAM 48GB, both at ~936 GB/s. Higher ceiling than any mini-PC on models that fit; see GPU Buying Guide for the build.

The Bottom Line

Three things changed in the five months since the original DGX Spark launch story stabilized:

  1. NVIDIA hiked the Founders Edition price 18% to $4,699 for memory-supply reasons. The hardware didn’t change; the value calculation did. And the memory shortage isn’t easing — DRAM contract prices are still forecast up double digits through Q3 2026.
  2. The field widened, then widened again. AMD’s Strix Halo ships the same 128GB unified-memory category from ~$1,999 with Windows, and on the NVIDIA side the “announced” boxes shipped — Gigabyte AI TOP ATOM, Lenovo ThinkStation PGX, and HP’s new ZGX Nano G1n (the priciest GB10 at ~$6,030) all joined the DGX Spark and its OEM twins. AMD’s own Ryzen AI Halo Developer Platform launched in July at $3,999.
  3. The used-3090 reality check stayed the most important question. For anything that fits in 24GB, a discrete GPU at $1,000-1,300 outruns every box in this guide. The 128GB-unified tier is for workloads that genuinely don’t fit. If yours does, save the money.

If you’ve concluded you need the 128GB-unified tier:

  • For CUDA workflows: DGX Spark Founders (Gen 5 NVMe, reference design, premium price) or MSI / ASUS at $2,999-$3,099 (1TB) for the most cost-effective entry. Performance differences within the GB10 field are negligible; pay for chassis, NVMe, and warranty preferences.
  • For Windows or budget-conscious 128GB: Strix Halo (GMKtec, Corsair, NextNuc, AMD reference). $1,999 to $3,999.
  • For Apple stack with bandwidth headroom (models that fit in 96GB): Mac Studio M3 Ultra 96GB, $5,299. Apple’s RAM cuts pulled the higher-memory tiers, so it trails the 128GB boxes on capacity but triples them on bandwidth.

For everything else, the Used RTX 3090 guide and the Multi-GPU local AI guide are the better reads.