NVIDIA DGX Spark for Local AI: Honest Review, Real Benchmarks, and What to Buy Instead
Want to go deeper than this article?
Free account unlocks the first chapter of all 25 courses — RAG, agents, MCP, voice AI, MLOps, real GitHub repos.
Go from reading about AI to building with AI 25 structured courses. Hands-on projects. Runs on your machine. Start free.
Short answer: don't buy the $4,699 DGX Spark Founders Edition — buy the ASUS Ascent GX10, the same GB10 chip and 128GB of memory for $3,099.99, and only if you run mixture-of-experts models. On GB10 hardware, gpt-oss-120b generates at ~60 tokens/sec (llama.cpp maintainer benchmarks), but dense Llama 3.1 70B decodes at 2.7 tokens/sec (LMSYS). The 273 GB/s memory bandwidth decides everything — this review is mostly the story of that one number.
One honesty note before anything else: we have not bought a DGX Spark. Every performance figure on this page comes from published benchmarks by people who have — the llama.cpp maintainer's own test thread, LMSYS's multi-framework suite, and ServeTheHome's hardware review — and each number is attributed to its source. What we add is the arithmetic that ties the numbers together and the price context from our ongoing hardware market tracking, so you can decide whether this box fits the models you actually want to run.
What the DGX Spark Is
It is a 128GB unified-memory Arm + Blackwell mini-PC — the machine NVIDIA previewed as Project DIGITS at CES 2025, shipped in October 2025 at $3,999, and now sells for $4,699. Everything below is from NVIDIA's official spec sheet:
| Spec | DGX Spark (NVIDIA official) |
|---|---|
| Chip | GB10 Grace Blackwell Superchip |
| CPU | 20 Arm cores: 10x Cortex-X925 + 10x Cortex-A725 |
| Memory | 128GB LPDDR5X, coherent unified (CPU+GPU) |
| Memory bandwidth | 273 GB/s |
| AI compute | Up to 1 PFLOP FP4 (with sparsity) |
| Storage | 4TB NVMe, self-encrypting |
| Networking | ConnectX-7 @ 200Gbps, 10GbE RJ-45, WiFi 7 |
| OS | DGX OS (Ubuntu-based) |
| Size | 150 x 150 x 50.5mm (per ServeTheHome) |
The pitch is simple: a desk-sized box that holds models no consumer GPU can. A 24GB RTX card cannot load gpt-oss-120b at all; the Spark loads it with ~60GB to spare. The two ConnectX-7 200GbE ports support RDMA, and ServeTheHome confirmed you can link two Sparks over copper DAC for a 256GB two-node cluster — a genuinely unusual feature at this size.
The same GB10 chip ships in partner clones — ASUS Ascent GX10, Acer Veriton GN100, plus Dell, HP, Lenovo, MSI, and Gigabyte versions — which matters enormously for what you should actually pay (see pricing).
Reading articles is good. Building is better.
Free account = the first chapter of all 25 courses, with a per-chapter AI tutor. No card.
The Benchmarks That Matter
The headline pair: gpt-oss-120b at 1,956 tok/s prefill / 60.6 tok/s generation, versus dense Llama 3.1 70B at 803 tok/s prefill / 2.7 tok/s decode. Same machine. That 22x generation gap between a 117B model and a 70B model is the entire review in two rows.
All numbers below are published third-party measurements, not ours:
| Model | Quant / framework | Prefill (tok/s) | Generation (tok/s) | Source |
|---|---|---|---|---|
| gpt-oss-120b (117B MoE) | MXFP4, llama.cpp | 1,956 | 60.6 (40.6 @ 32K ctx) | ggerganov (llama.cpp maintainer), Oct 2025 |
| Qwen3 Coder 30B (MoE) | Q8_0, llama.cpp | 1,654 | 44.3 | ggerganov, Oct 2025 |
| gpt-oss-20b (MoE) | MXFP4, Ollama | 2,053 | 49.7 | LMSYS, Oct 2025 |
| Llama 3.1 8B (dense) | FP8, SGLang, batch 1 | 7,991 | 20.5 | LMSYS, Oct 2025 |
| Llama 3.1 8B (dense) | FP8, SGLang, batch 32 | 7,949 | 368 (aggregate) | LMSYS, Oct 2025 |
| Llama 3.1 70B (dense) | FP8, SGLang, batch 1 | 803 | 2.7 | LMSYS, Oct 2025 |
Three things worth pulling out of that table:
- Prefill is genuinely excellent. Nearly 8,000 tok/s prompt processing on an 8B model, and ~2,000 tok/s even on a 117B MoE — that is Blackwell tensor compute doing its job. Long-context RAG ingestion, document summarization, and agent loops with fat prompts all benefit.
- Batch throughput is real. 368 aggregate tok/s on 8B at batch 32 means the Spark can serve a small team on small models even though single-stream feels ordinary.
- Dense 70B decode is not usable for interactive chat. 2.7 tok/s is slower than reading pace. LMSYS's own framing was that 70B/120B-class dense models "are best suited for prototyping and experimentation rather than production" on this box.
A calibration note on early reviews: ServeTheHome's launch-day review measured gpt-oss-120b at just 14.5 tok/s — a quarter of the maintainer's llama.cpp number — and commenters flagged likely configuration issues; the llama.cpp thread also documents model-load times dropping from 104s to 22s with a kernel change. Early Spark numbers scattered widely with software maturity, which is exactly why we only cite sourced figures and name who measured them.
The 273 GB/s Reality Check
Token generation is memory-bandwidth-bound: every generated token reads roughly the model's active weights once, so 273 GB/s divided by weight size is a hard ceiling no software update can beat. This one piece of arithmetic explains every number in the table above.
Run it yourself (this is our arithmetic, and it is checkable):
- Dense Llama 3.1 70B at FP8 is ~70GB of weights. 273 / 70 = ~3.9 tok/s theoretical maximum. LMSYS measured 2.7 — the hardware is already near its ceiling. Not a driver problem. Physics.
- Dense 70B at Q4 (~40GB, the size you will find in our Ollama RAM/VRAM table) tops out around ~6-7 tok/s. Better, still not conversational.
- gpt-oss-120b is a mixture-of-experts model: 117B total parameters but only 5.1B active per token (per OpenAI's model card), stored in MXFP4. Only a few GB of weights move per token — which is how a 117B model hits 60 tok/s on the same 273 GB/s that chokes a 70B dense model at 2.7.
So the Spark is not "fast" or "slow." It is a capacity machine with a bandwidth budget, and MoE models are the models built for exactly that budget. The good news for buyers: the open-model ecosystem swung hard toward MoE — gpt-oss 20b/120b, Qwen3's A3B line, and most frontier-class open releases since. If your model shortlist is MoE, GB10 hardware is well matched. If your shortlist is dense 70B chat models, this is the wrong box and no firmware update will change that.
The compute side ages better: NVIDIA's claim of "up to 2.5x inference gains" from software optimization since launch is a vendor claim, but prefill and batch throughput genuinely are compute-bound, so software wins there are plausible. Decode on dense models is not where the wins will come from.
DGX Spark vs Mac Studio vs Strix Halo
The Spark has the best software stack and the worst bandwidth-per-dollar of the three 128GB-class options: 273 GB/s vs 819 GB/s on a Mac Studio M3 Ultra that launched at a lower price than the Spark now sells for.
| DGX Spark / GB10 clones | Mac Studio (M3 Ultra) | Strix Halo mini-PCs | |
|---|---|---|---|
| Price | $3,099.99 (GX10) - $4,699 (FE) | Launched $3,999 (96GB) | ~$1,999-$3,300 (128GB, volatile) |
| Memory | 128GB LPDDR5X | 96GB base, configurable higher | 128GB LPDDR5X |
| Bandwidth | 273 GB/s | 819 GB/s (Apple spec) | ~256 GB/s |
| Software stack | CUDA (llama.cpp, vLLM, TensorRT) | Metal / MLX | ROCm / Vulkan |
| Wildcard | 200GbE RDMA two-node clustering | Fastest dense-model decode | Cheapest 128GB |
How to actually choose:
- Dense 70B+ models, interactive use: Mac Studio, no contest. 819 GB/s is 3x the Spark's bandwidth, and decode scales almost linearly with it. (The M4 Max tier runs 410 GB/s in the Mac Studio's base configuration and 546 GB/s on the higher-binned chip — up to twice the Spark — and Apple's laptop line now carries 128GB too; our Apple Silicon buying guide maps the configs.)
- CUDA development: Spark-class, and this is its honest niche. If your prototypes deploy to NVIDIA datacenter GPUs, developing on the same stack — same containers, same kernels, TensorRT, NIM — is worth real money. No Mac or AMD box offers that.
- Cheapest ticket to 128GB: Strix Halo. Nearly identical bandwidth (~256 vs 273 GB/s) at half to two-thirds the price, on x86 with a normal Linux ecosystem. The trade is software maturity — ROCm has come a long way but is not CUDA. Our Strix Halo guide covers it at the same depth as this page, and the best mini PC for Ollama roundup has the current box-by-box options.
Note what is not a differentiator: capacity. All three hold 96-128GB. You are choosing on bandwidth, stack, and price — and the Spark only wins on stack.
Run this on your own machine and stop paying every month
Pay once and keep it. No renewal, no per-token bill, and nothing you feed it ever leaves your hardware.
What It Costs — and Why You Should Buy the Clone
NVIDIA raised the Founders Edition from $3,999 to $4,699 in late February 2026 (citing the DRAM/NAND shortage, no hardware changes) while the identical-chip ASUS Ascent GX10 sells for $3,099.99 — so the Founders Edition is currently a $1,600 badge tax.
Prices from our mid-July 2026 market survey; street prices in this category move monthly, so treat them as a snapshot:
| Machine | Chip | Price | Note |
|---|---|---|---|
| DGX Spark Founders Edition | GB10, 128GB, 4TB | $4,699 | Raised +18% from $3,999 launch, Feb 2026 |
| ASUS Ascent GX10 | GB10, 128GB, 1TB | $3,099.99 | Announced at $2,999; in stock at major US retailers |
| ASUS Ascent GX10 (4TB) | GB10, 128GB, 4TB | $4,149.99 | Still undercuts the FE with equal storage |
| Acer Veriton GN100 | GB10, 128GB | ~$3,999 at launch | Reported launch pricing; verify current |
The GX10 runs the same DGX OS-derived software base and the same CUDA stack — the silicon is literally the same GB10. Unless the NVIDIA badge, the FE's 4TB drive, or first-party support specifically matters to you, the GX10 at $3,099.99 is the version of this product we would actually short-list. The wider context — why every box in this category got more expensive during the memory supercycle — is in our GPU price tracker.
Setup: Ollama and llama.cpp on a Spark
Setup is genuinely boring, in the good way: DGX OS is Ubuntu-based, so the standard Linux Ollama installer works, and llama.cpp builds with the ordinary CUDA flags. Commands below are verified against the current Ollama and llama.cpp docs (not run on our own Spark — we don't own one):
Ollama (the path LMSYS used for part of its benchmark suite):
curl -fsSL https://ollama.com/install.sh | sh
ollama run gpt-oss:120b # 65GB download — fits in 128GB with headroom
ollama run gpt-oss:20b # 14GB — the fast everyday option (~50 tok/s per LMSYS)
llama.cpp, built the same way the maintainer's Spark benchmarks were produced (per the project's build docs):
git clone https://github.com/ggml-org/llama.cpp
cd llama.cpp
cmake -B build -DGGML_CUDA=ON
cmake --build build --config Release
Model picks: favor MoE for the bandwidth reasons above — gpt-oss-120b MXFP4 and Qwen3-Coder-30B are the two with published Spark numbers. Our complete Ollama guide covers the day-two stuff (Modelfiles, context settings, serving), and the RAM/VRAM table shows what any given quant actually occupies.
Two practical notes from the llama.cpp benchmark thread worth knowing before day one: model load speed is kernel-sensitive (one user cut gpt-oss-120b load time from 104s to 22s by moving to NVIDIA's 6.17.1 kernel build), and unified-memory settings swung inference results by roughly 3-10%. Also useful: vLLM published an official DGX Spark deployment guide in June 2026, so the serving story has matured well past launch.
Honest Limitations
The short list: bandwidth-capped dense-model decode, an Arm software ecosystem with occasional sharp edges, a raised price, and zero upgrade path.
- The 273 GB/s ceiling is permanent. Dense 70B chat will never be pleasant on this hardware. If model taste shifts back toward dense architectures, the Spark ages badly; if MoE keeps winning, it ages fine. That is a real bet you are making.
- Arm64 means occasional friction. The mainstream stack (Ollama, llama.cpp, vLLM, PyTorch) is solid on aarch64 now, but random tooling — a quantization script, a prebuilt wheel, a Docker image — still sometimes assumes x86. Budget occasional compile-it-yourself time.
- Early benchmark variance was large. A launch-week professional review got a quarter of the maintainer's llama.cpp throughput on the same model. Software has settled since, but when you see Spark numbers anywhere (including here), check the date, framework, and settings before comparing.
- Nothing is upgradeable. Memory is soldered LPDDR5X; 128GB on day one is 128GB forever. The 200-300B-class MoE models at aggressive quants fit today with little room to grow.
- The price went the wrong way. Same hardware, +18% MSRP in February. The GX10 clone blunts this, but the whole category is exposed to the memory shortage, and our July survey pricing may have drifted by the time you read this — check before ordering.
Who Should Buy One (and What to Buy Instead)
Buy a GB10 box — as the $3,099.99 GX10 — if you develop on CUDA for NVIDIA deployment targets or you live on MoE models too big for consumer VRAM. Otherwise there is a better machine for your specific case.
| Your situation | Buy | Why |
|---|---|---|
| CUDA developer, prototypes deploy to NVIDIA infra | ASUS Ascent GX10 ($3,099.99) | The Spark's one unassailable niche, $1,600 cheaper |
| MoE models (gpt-oss-120b class) as daily drivers | GX10 or Strix Halo box | ~60 tok/s is real; Strix is cheaper, CUDA is smoother |
| Dense 70B+ interactive chat | Mac Studio M3 Ultra | 819 GB/s vs 273 — decode is 3x the bandwidth away |
| 128GB on the smallest budget | Strix Halo mini (~$1,999+, volatile) | Same capacity, ~94% of the bandwidth, big savings |
| Your models fit in 32GB | A discrete GPU, not any 128GB box | Far faster per token; see our 32GB VRAM model picks |
| Two-node 200GbE RDMA experiments | 2x Spark/GX10 | Nothing else at this size does it (per ServeTheHome) |
The verdict, compressed: the DGX Spark is a good machine sold at a bad price, sitting next to its own clone at a fair one. The honest reframe from a year of published benchmarks is to stop asking "is it fast?" and ask "are my models MoE?" If yes, GB10 at GX10 pricing is a defensible buy and the CUDA stack is a real amenity. If no, the same money buys visibly better decode on a Mac or the same capacity for much less on Strix Halo. And if you are not sure what your models even need, start from the model side, not the hardware side — pick the models first, check what they occupy in our RAM/VRAM table, and let that number choose the box.
Sources
- NVIDIA DGX Spark official page — specs (GB10, 128GB LPDDR5X, 273 GB/s, 1 PFLOP FP4 sparse, ConnectX-7) and partner list
- llama.cpp DGX Spark benchmark thread — ggerganov (maintainer), Oct 2025: gpt-oss-120b and Qwen3-Coder-30B numbers, kernel/load-time notes
- LMSYS DGX Spark benchmarks — SGLang + Ollama suite: 8B/70B FP8, gpt-oss, batch scaling
- ServeTheHome DGX Spark review — hardware teardown, dimensions, 200GbE RDMA clustering, launch-day performance caveats
- openai/gpt-oss-120b model card — 117B total / 5.1B active parameters, MXFP4
- Ollama gpt-oss library page and llama.cpp build docs — install/build commands and model sizes
- Apple Mac Studio specs — M3 Ultra 819 GB/s (96GB base), M4 Max 410 GB/s in the base Mac Studio configuration
- Our mid-July 2026 market survey (GPU prices & the memory shortage) — FE $4,699 / GX10 $3,099.99 pricing snapshot
FAQ
Go from reading about AI to building with AI
25 structured courses. Hands-on projects. Runs on your machine. Start free.
Liked this? 25 full AI courses are waiting.
From fundamentals to RAG, agents, MCP servers, voice AI, and production deployment with real GitHub repos. First chapter free, every course.
Build Real AI on Your Machine
RAG, agents, NLP, vision, and MLOps - chapters across 25 courses that take you from reading about AI to building AI.
Want the structured version?
Hands-on courses on local AI, from $8.99 a month. The first chapter of each is free.
Keep going
Comments (0)
No comments yet. Be the first to share your thoughts!