★ Reading this for free? Get 25 structured AI courses + per-chapter AI tutor — the first chapter of every course free, no card.Start free in 30 secondsOr own every course: $149 once
Hardware

NVIDIA DGX Spark for Local AI: Honest Review, Real Benchmarks, and What to Buy Instead

August 23, 2026
14 min read
LocalAimaster Research Team

Want to go deeper than this article?

Free account unlocks the first chapter of all 25 courses — RAG, agents, MCP, voice AI, MLOps, real GitHub repos.

📚AI Learning Path

Go from reading about AI to building with AI 25 structured courses. Hands-on projects. Runs on your machine. Start free.

Start free
Or own it for life — Lifetime $149, pay once

Short answer: don't buy the $4,699 DGX Spark Founders Edition — buy the ASUS Ascent GX10, the same GB10 chip and 128GB of memory for $3,099.99, and only if you run mixture-of-experts models. On GB10 hardware, gpt-oss-120b generates at ~60 tokens/sec (llama.cpp maintainer benchmarks), but dense Llama 3.1 70B decodes at 2.7 tokens/sec (LMSYS). The 273 GB/s memory bandwidth decides everything — this review is mostly the story of that one number.

One honesty note before anything else: we have not bought a DGX Spark. Every performance figure on this page comes from published benchmarks by people who have — the llama.cpp maintainer's own test thread, LMSYS's multi-framework suite, and ServeTheHome's hardware review — and each number is attributed to its source. What we add is the arithmetic that ties the numbers together and the price context from our ongoing hardware market tracking, so you can decide whether this box fits the models you actually want to run.


What the DGX Spark Is

It is a 128GB unified-memory Arm + Blackwell mini-PC — the machine NVIDIA previewed as Project DIGITS at CES 2025, shipped in October 2025 at $3,999, and now sells for $4,699. Everything below is from NVIDIA's official spec sheet:

SpecDGX Spark (NVIDIA official)
ChipGB10 Grace Blackwell Superchip
CPU20 Arm cores: 10x Cortex-X925 + 10x Cortex-A725
Memory128GB LPDDR5X, coherent unified (CPU+GPU)
Memory bandwidth273 GB/s
AI computeUp to 1 PFLOP FP4 (with sparsity)
Storage4TB NVMe, self-encrypting
NetworkingConnectX-7 @ 200Gbps, 10GbE RJ-45, WiFi 7
OSDGX OS (Ubuntu-based)
Size150 x 150 x 50.5mm (per ServeTheHome)

The pitch is simple: a desk-sized box that holds models no consumer GPU can. A 24GB RTX card cannot load gpt-oss-120b at all; the Spark loads it with ~60GB to spare. The two ConnectX-7 200GbE ports support RDMA, and ServeTheHome confirmed you can link two Sparks over copper DAC for a 256GB two-node cluster — a genuinely unusual feature at this size.

The same GB10 chip ships in partner clones — ASUS Ascent GX10, Acer Veriton GN100, plus Dell, HP, Lenovo, MSI, and Gigabyte versions — which matters enormously for what you should actually pay (see pricing).


Reading articles is good. Building is better.

Free account = the first chapter of all 25 courses, with a per-chapter AI tutor. No card.

The Benchmarks That Matter

The headline pair: gpt-oss-120b at 1,956 tok/s prefill / 60.6 tok/s generation, versus dense Llama 3.1 70B at 803 tok/s prefill / 2.7 tok/s decode. Same machine. That 22x generation gap between a 117B model and a 70B model is the entire review in two rows.

All numbers below are published third-party measurements, not ours:

ModelQuant / frameworkPrefill (tok/s)Generation (tok/s)Source
gpt-oss-120b (117B MoE)MXFP4, llama.cpp1,95660.6 (40.6 @ 32K ctx)ggerganov (llama.cpp maintainer), Oct 2025
Qwen3 Coder 30B (MoE)Q8_0, llama.cpp1,65444.3ggerganov, Oct 2025
gpt-oss-20b (MoE)MXFP4, Ollama2,05349.7LMSYS, Oct 2025
Llama 3.1 8B (dense)FP8, SGLang, batch 17,99120.5LMSYS, Oct 2025
Llama 3.1 8B (dense)FP8, SGLang, batch 327,949368 (aggregate)LMSYS, Oct 2025
Llama 3.1 70B (dense)FP8, SGLang, batch 18032.7LMSYS, Oct 2025

Three things worth pulling out of that table:

  • Prefill is genuinely excellent. Nearly 8,000 tok/s prompt processing on an 8B model, and ~2,000 tok/s even on a 117B MoE — that is Blackwell tensor compute doing its job. Long-context RAG ingestion, document summarization, and agent loops with fat prompts all benefit.
  • Batch throughput is real. 368 aggregate tok/s on 8B at batch 32 means the Spark can serve a small team on small models even though single-stream feels ordinary.
  • Dense 70B decode is not usable for interactive chat. 2.7 tok/s is slower than reading pace. LMSYS's own framing was that 70B/120B-class dense models "are best suited for prototyping and experimentation rather than production" on this box.

A calibration note on early reviews: ServeTheHome's launch-day review measured gpt-oss-120b at just 14.5 tok/s — a quarter of the maintainer's llama.cpp number — and commenters flagged likely configuration issues; the llama.cpp thread also documents model-load times dropping from 104s to 22s with a kernel change. Early Spark numbers scattered widely with software maturity, which is exactly why we only cite sourced figures and name who measured them.


The 273 GB/s Reality Check

Token generation is memory-bandwidth-bound: every generated token reads roughly the model's active weights once, so 273 GB/s divided by weight size is a hard ceiling no software update can beat. This one piece of arithmetic explains every number in the table above.

Run it yourself (this is our arithmetic, and it is checkable):

  • Dense Llama 3.1 70B at FP8 is ~70GB of weights. 273 / 70 = ~3.9 tok/s theoretical maximum. LMSYS measured 2.7 — the hardware is already near its ceiling. Not a driver problem. Physics.
  • Dense 70B at Q4 (~40GB, the size you will find in our Ollama RAM/VRAM table) tops out around ~6-7 tok/s. Better, still not conversational.
  • gpt-oss-120b is a mixture-of-experts model: 117B total parameters but only 5.1B active per token (per OpenAI's model card), stored in MXFP4. Only a few GB of weights move per token — which is how a 117B model hits 60 tok/s on the same 273 GB/s that chokes a 70B dense model at 2.7.

So the Spark is not "fast" or "slow." It is a capacity machine with a bandwidth budget, and MoE models are the models built for exactly that budget. The good news for buyers: the open-model ecosystem swung hard toward MoE — gpt-oss 20b/120b, Qwen3's A3B line, and most frontier-class open releases since. If your model shortlist is MoE, GB10 hardware is well matched. If your shortlist is dense 70B chat models, this is the wrong box and no firmware update will change that.

The compute side ages better: NVIDIA's claim of "up to 2.5x inference gains" from software optimization since launch is a vendor claim, but prefill and batch throughput genuinely are compute-bound, so software wins there are plausible. Decode on dense models is not where the wins will come from.


DGX Spark vs Mac Studio vs Strix Halo

The Spark has the best software stack and the worst bandwidth-per-dollar of the three 128GB-class options: 273 GB/s vs 819 GB/s on a Mac Studio M3 Ultra that launched at a lower price than the Spark now sells for.

DGX Spark / GB10 clonesMac Studio (M3 Ultra)Strix Halo mini-PCs
Price$3,099.99 (GX10) - $4,699 (FE)Launched $3,999 (96GB)~$1,999-$3,300 (128GB, volatile)
Memory128GB LPDDR5X96GB base, configurable higher128GB LPDDR5X
Bandwidth273 GB/s819 GB/s (Apple spec)~256 GB/s
Software stackCUDA (llama.cpp, vLLM, TensorRT)Metal / MLXROCm / Vulkan
Wildcard200GbE RDMA two-node clusteringFastest dense-model decodeCheapest 128GB

How to actually choose:

  • Dense 70B+ models, interactive use: Mac Studio, no contest. 819 GB/s is 3x the Spark's bandwidth, and decode scales almost linearly with it. (The M4 Max tier runs 410 GB/s in the Mac Studio's base configuration and 546 GB/s on the higher-binned chip — up to twice the Spark — and Apple's laptop line now carries 128GB too; our Apple Silicon buying guide maps the configs.)
  • CUDA development: Spark-class, and this is its honest niche. If your prototypes deploy to NVIDIA datacenter GPUs, developing on the same stack — same containers, same kernels, TensorRT, NIM — is worth real money. No Mac or AMD box offers that.
  • Cheapest ticket to 128GB: Strix Halo. Nearly identical bandwidth (~256 vs 273 GB/s) at half to two-thirds the price, on x86 with a normal Linux ecosystem. The trade is software maturity — ROCm has come a long way but is not CUDA. Our Strix Halo guide covers it at the same depth as this page, and the best mini PC for Ollama roundup has the current box-by-box options.

Note what is not a differentiator: capacity. All three hold 96-128GB. You are choosing on bandwidth, stack, and price — and the Spark only wins on stack.


Own it instead of renting it

Run this on your own machine and stop paying every month

Pay once and keep it. No renewal, no per-token bill, and nothing you feed it ever leaves your hardware.

What It Costs — and Why You Should Buy the Clone

NVIDIA raised the Founders Edition from $3,999 to $4,699 in late February 2026 (citing the DRAM/NAND shortage, no hardware changes) while the identical-chip ASUS Ascent GX10 sells for $3,099.99 — so the Founders Edition is currently a $1,600 badge tax.

Prices from our mid-July 2026 market survey; street prices in this category move monthly, so treat them as a snapshot:

MachineChipPriceNote
DGX Spark Founders EditionGB10, 128GB, 4TB$4,699Raised +18% from $3,999 launch, Feb 2026
ASUS Ascent GX10GB10, 128GB, 1TB$3,099.99Announced at $2,999; in stock at major US retailers
ASUS Ascent GX10 (4TB)GB10, 128GB, 4TB$4,149.99Still undercuts the FE with equal storage
Acer Veriton GN100GB10, 128GB~$3,999 at launchReported launch pricing; verify current

The GX10 runs the same DGX OS-derived software base and the same CUDA stack — the silicon is literally the same GB10. Unless the NVIDIA badge, the FE's 4TB drive, or first-party support specifically matters to you, the GX10 at $3,099.99 is the version of this product we would actually short-list. The wider context — why every box in this category got more expensive during the memory supercycle — is in our GPU price tracker.


Setup: Ollama and llama.cpp on a Spark

Setup is genuinely boring, in the good way: DGX OS is Ubuntu-based, so the standard Linux Ollama installer works, and llama.cpp builds with the ordinary CUDA flags. Commands below are verified against the current Ollama and llama.cpp docs (not run on our own Spark — we don't own one):

Ollama (the path LMSYS used for part of its benchmark suite):

curl -fsSL https://ollama.com/install.sh | sh
ollama run gpt-oss:120b     # 65GB download — fits in 128GB with headroom
ollama run gpt-oss:20b      # 14GB — the fast everyday option (~50 tok/s per LMSYS)

llama.cpp, built the same way the maintainer's Spark benchmarks were produced (per the project's build docs):

git clone https://github.com/ggml-org/llama.cpp
cd llama.cpp
cmake -B build -DGGML_CUDA=ON
cmake --build build --config Release

Model picks: favor MoE for the bandwidth reasons above — gpt-oss-120b MXFP4 and Qwen3-Coder-30B are the two with published Spark numbers. Our complete Ollama guide covers the day-two stuff (Modelfiles, context settings, serving), and the RAM/VRAM table shows what any given quant actually occupies.

Two practical notes from the llama.cpp benchmark thread worth knowing before day one: model load speed is kernel-sensitive (one user cut gpt-oss-120b load time from 104s to 22s by moving to NVIDIA's 6.17.1 kernel build), and unified-memory settings swung inference results by roughly 3-10%. Also useful: vLLM published an official DGX Spark deployment guide in June 2026, so the serving story has matured well past launch.


Honest Limitations

The short list: bandwidth-capped dense-model decode, an Arm software ecosystem with occasional sharp edges, a raised price, and zero upgrade path.

  • The 273 GB/s ceiling is permanent. Dense 70B chat will never be pleasant on this hardware. If model taste shifts back toward dense architectures, the Spark ages badly; if MoE keeps winning, it ages fine. That is a real bet you are making.
  • Arm64 means occasional friction. The mainstream stack (Ollama, llama.cpp, vLLM, PyTorch) is solid on aarch64 now, but random tooling — a quantization script, a prebuilt wheel, a Docker image — still sometimes assumes x86. Budget occasional compile-it-yourself time.
  • Early benchmark variance was large. A launch-week professional review got a quarter of the maintainer's llama.cpp throughput on the same model. Software has settled since, but when you see Spark numbers anywhere (including here), check the date, framework, and settings before comparing.
  • Nothing is upgradeable. Memory is soldered LPDDR5X; 128GB on day one is 128GB forever. The 200-300B-class MoE models at aggressive quants fit today with little room to grow.
  • The price went the wrong way. Same hardware, +18% MSRP in February. The GX10 clone blunts this, but the whole category is exposed to the memory shortage, and our July survey pricing may have drifted by the time you read this — check before ordering.

Who Should Buy One (and What to Buy Instead)

Buy a GB10 box — as the $3,099.99 GX10 — if you develop on CUDA for NVIDIA deployment targets or you live on MoE models too big for consumer VRAM. Otherwise there is a better machine for your specific case.

Your situationBuyWhy
CUDA developer, prototypes deploy to NVIDIA infraASUS Ascent GX10 ($3,099.99)The Spark's one unassailable niche, $1,600 cheaper
MoE models (gpt-oss-120b class) as daily driversGX10 or Strix Halo box~60 tok/s is real; Strix is cheaper, CUDA is smoother
Dense 70B+ interactive chatMac Studio M3 Ultra819 GB/s vs 273 — decode is 3x the bandwidth away
128GB on the smallest budgetStrix Halo mini (~$1,999+, volatile)Same capacity, ~94% of the bandwidth, big savings
Your models fit in 32GBA discrete GPU, not any 128GB boxFar faster per token; see our 32GB VRAM model picks
Two-node 200GbE RDMA experiments2x Spark/GX10Nothing else at this size does it (per ServeTheHome)

The verdict, compressed: the DGX Spark is a good machine sold at a bad price, sitting next to its own clone at a fair one. The honest reframe from a year of published benchmarks is to stop asking "is it fast?" and ask "are my models MoE?" If yes, GB10 at GX10 pricing is a defensible buy and the CUDA stack is a real amenity. If no, the same money buys visibly better decode on a Mac or the same capacity for much less on Strix Halo. And if you are not sure what your models even need, start from the model side, not the hardware side — pick the models first, check what they occupy in our RAM/VRAM table, and let that number choose the box.


Sources


FAQ

🎯
AI Learning Path

Go from reading about AI to building with AI

25 structured courses. Hands-on projects. Runs on your machine. Start free.

Or own it for life — Lifetime $149 $599, pay once

Liked this? 25 full AI courses are waiting.

From fundamentals to RAG, agents, MCP servers, voice AI, and production deployment with real GitHub repos. First chapter free, every course.

Reading now
Join the discussion
TagsDGX SparkGB10NVIDIAUnified MemoryMac StudioStrix HaloLocal AI Hardware

LocalAimaster Research Team

Local AI Master writes hands-on courses and hardware guides for running AI on machines you own. Content is checked against current releases and corrected when readers tell us it is wrong.

Build Real AI on Your Machine

RAG, agents, NLP, vision, and MLOps - chapters across 25 courses that take you from reading about AI to building AI.

Want the structured version?

Hands-on courses on local AI, from $8.99 a month. The first chapter of each is free.

AI Learning Path

Comments (0)

No comments yet. Be the first to share your thoughts!

Is the DGX Spark worth it for local AI in 2026?

Only if you buy the right version and run the right models. The Founders Edition costs $4,699 after NVIDIA's February 2026 price hike, but the ASUS Ascent GX10 has the identical GB10 chip and 128GB of memory for $3,099.99 — about $1,600 less. On any GB10 box, mixture-of-experts models are the sweet spot: gpt-oss-120b runs at roughly 60 tokens/sec (llama.cpp maintainer benchmarks). Dense 70B models are not: Llama 3.1 70B decodes at 2.7 tokens/sec in LMSYS's FP8 SGLang test. If your target models are dense 70B+, look at a Mac Studio's higher bandwidth instead.

How fast is the DGX Spark for local LLMs?

Published numbers, not ours: the llama.cpp maintainer measured gpt-oss-120b (MXFP4) at 1,956 tokens/sec prefill and 60.6 tokens/sec generation (40.6 at 32K context depth), and Qwen3 Coder 30B (Q8_0) at 44.3 tokens/sec generation. LMSYS measured gpt-oss-20b via Ollama at 49.7 tokens/sec decode, Llama 3.1 8B FP8 at 20.5 tokens/sec single-stream (368 aggregate at batch 32), and Llama 3.1 70B FP8 at just 2.7 tokens/sec. The pattern: prefill is strong (Blackwell compute), decode is capped by 273 GB/s memory bandwidth, so sparse MoE models fly and dense large models crawl.

DGX Spark vs Mac Studio — which is better for local AI?

Bandwidth versus ecosystem. The Mac Studio M3 Ultra has 819 GB/s of memory bandwidth — three times the Spark's 273 GB/s — so dense-model decode is much faster on the Mac, and it launched at $3,999 with 96GB (configurable far higher). The Spark counters with CUDA: the entire NVIDIA software stack (llama.cpp CUDA builds, vLLM, TensorRT, NIM containers) runs natively, which matters if you develop for deployment on NVIDIA datacenter GPUs. Rule of thumb: dense 70B+ chat and long-form generation, buy the Mac; CUDA development, fine-tuning experiments, and MoE inference, buy a GB10 box — preferably the $3,099.99 ASUS GX10, not the $4,699 Spark.

Does Ollama run on the DGX Spark?

Yes, out of the box. DGX OS is Ubuntu-based, and the standard Linux installer (curl -fsSL https://ollama.com/install.sh | sh) works on the Spark's arm64 + CUDA setup — LMSYS ran its published Spark benchmarks partly through Ollama. gpt-oss:120b is a 65GB download that fits in the 128GB of unified memory with room to spare, which is exactly the trick a 24GB or even 32GB consumer GPU cannot do.

Is the DGX Spark the same thing as Project DIGITS?

Same machine, renamed. NVIDIA previewed it as Project DIGITS at CES in January 2025, renamed it DGX Spark around GTC in March 2025, and shipped it in October 2025 at $3,999. In late February 2026 NVIDIA raised the Founders Edition MSRP to $4,699, citing the DRAM/NAND shortage — no hardware changes. The partner clones (ASUS Ascent GX10, Dell, HP, Lenovo, MSI, Gigabyte, Acer Veriton GN100) use the same GB10 Grace Blackwell chip and mostly undercut NVIDIA's own price.

Ready to Go Beyond Tutorials?

25 structured courses with hands-on chapters - build RAG chatbots, AI agents, and ML pipelines on your own hardware.

Bonus kit

Ollama Docker Templates

10 one-command Docker stacks for local models — get your new box serving in minutes. Included with paid plans, or free after subscribing to both Local AI Master and Little AI Master on YouTube.

See Plans →

Was this helpful?

📅 Published: August 23, 2026🔄 Last Updated: August 23, 2026✓ Manually Reviewed
LM

Written by the Local AI Master Team

The team behind Local AI Master

We build Local AI Master around practical, testable local AI workflows: model selection, hardware planning, RAG systems, agents, and MLOps. The goal is to turn scattered tutorials into a structured learning path you can follow on your own hardware.

✓ Local AI Curriculum✓ Hands-On Projects✓ Open Source Contributor
📚
Free · no account required

Grab the AI Starter Kit — career roadmap, cheat sheet, setup guide

No spam. Unsubscribe with one click.

🎯
AI Learning Path

Go from reading about AI to building with AI

25 structured courses. Hands-on projects. Runs on your machine. Start free.

Or own it for life — Lifetime $149 $599, pay once
Free Tools & Calculators