A petaFLOP of AI performance in a box the size of a small router sounds like it should be on every developer's desk. The reality is narrower. The NVIDIA DGX Spark is a purpose-built personal AI machine, built around NVIDIA's GB10 chip that pairs a 20-core Arm CPU with a Blackwell GPU and a 128GB pool of shared memory, and it is aimed squarely at people who build and run models rather than people who want a fast gaming or workstation PC. Knowing which group you fall into saves a serious amount of money.

Quick Answer

Buy a DGX Spark if you are an AI developer, researcher, or data scientist who needs to iterate, test, and run inference on large models locally -- a single unit accommodates models in the 200 billion parameter range through its shared memory pool. If you mainly game, render, edit video, or want raw graphics horsepower, look elsewhere: this is an AI development appliance, not a general-purpose desktop. It retails at around 4,699 US dollars, so it is a deliberate investment.

What the DGX Spark Actually Is

The DGX Spark centres on NVIDIA's GB10 chip, which marries a 20-core Arm CPU (ten Cortex-X925 performance cores plus ten Cortex-A725 efficiency cores) with Blackwell-generation GPU compute. NVIDIA quotes up to one petaFLOP of AI performance at FP4 precision with sparsity enabled. The defining feature is a 128GB pool of LPDDR5X shared across CPU and GPU at 273 GB/s, which is what lets it hold genuinely large models resident in memory at once.

It ships preloaded with NVIDIA's AI software stack, so it behaves less like a blank PC you configure and more like a turnkey appliance for model work. Common tools including PyTorch, Jupyter, Ollama and LM Studio are validated to run on it from day one. You can also link two units over high-speed networking to handle models up to around 405 billion parameters, which signals the kind of work it is designed for.

Who Should Actually Buy One

The AI Developer Prototyping Locally

If you are iterating on models daily, cloud GPU rental becomes a recurring cost and a workflow friction point. A local unit that holds a large model in 128GB of unified memory lets you prototype, test, and refine without metering every hour or shipping data off-site. NVIDIA's early 2026 software refresh pushed throughput to roughly 2.5 times the original launch rate, thanks to TensorRT-LLM tuning and speculative-decoding gains baked into the updated stack, which makes the hardware meaningfully more capable than early reviews captured. For this person the DGX Spark earns its keep quickly.

The Researcher or Data Scientist

Academics and research teams who need reproducible local inference, who handle sensitive datasets that cannot leave the building, or who want to experiment with large open-weight models benefit from having the compute on the desk. The preloaded software stack shortens setup time, which matters when the goal is research output, not system administration. The machine runs fully offline, which suits regulated environments where cloud transfer of model training data is not permitted.

The Startup Avoiding Cloud Lock-In

A small team standing up an AI product can use a Spark, or a pair, as on-premises development hardware before committing to ongoing cloud spend. The fixed upfront cost is predictable in a way hourly cloud billing is not. Teams that reach the point where a 70B model is part of their product pipeline often find the per-token cloud cost adds up faster than the hardware price.

Who Should Not Buy One

Gamers get little from it: this is not a graphics card in a desktop, and its strengths do not translate to frame rates. Video editors, 3D artists, and CAD users are better served by a conventional workstation with a strong discrete GPU. Anyone whose models comfortably fit in a normal GPU's memory, or who only occasionally touches AI, will not recoup the price and should rent cloud compute or buy a standard machine. It is also worth noting that on models smaller than around 30B, a consumer GPU with 32GB of VRAM will generate tokens significantly faster than the Spark, because raw bandwidth beats capacity when the model fits. If your interest is local AI but your needs are lighter, compare the AI PC range at Evetech before reaching for an appliance at this tier.

The South African Buying Reality

Evetech stocks the DGX Spark locally, which removes the import friction and adds local warranty cover. The unit is listed at around R70,000, reflecting the US dollar price plus exchange rate and duties. For most South African builders who want serious local AI capability but cannot stretch to that figure, a workstation built on a high-memory consumer or professional GPU is the more flexible starting point. You can weigh those options against the best selling GPUs at Evetech to see where your budget goes furthest.

Frequently Asked Questions

What size models can a single DGX Spark run?

NVIDIA positions a single unit for AI models up to around 200 billion parameters, thanks to its 128GB of unified memory. Linking two units over high-speed networking extends that to roughly 405 billion parameters. In practice, 70B and 120B models run comfortably at 4-bit quantisation; 200B sits at the edge with a tighter context budget.

Is the DGX Spark good for gaming?

No. It is an AI development appliance, not a gaming PC. Its compute is tuned for model inference and training workloads, and it will not deliver the gaming experience a conventional GPU-equipped desktop does.

How is it different from a workstation with a big GPU?

A workstation splits memory between system RAM and GPU VRAM, which limits how large a model the GPU can hold. The Spark's 128GB unified memory pool is shared, so it can keep much larger models entirely in memory, which is its main advantage for AI work. A 70B model at FP16 requires around 140GB and simply will not fit in any single consumer GPU's VRAM.

Should a hobbyist buy one?

Usually not. The price and specialised purpose make it hard to justify for occasional experimentation. A hobbyist is better served by cloud rental or a capable consumer GPU until local model work becomes a daily need. The crossover point is roughly when you are regularly running 70B-and-above models and finding cloud latency or cost is the bottleneck.

Can the DGX Spark fine-tune models?

Yes. A single unit handles targeted fine-tuning on models in the 70B parameter range. Anything larger is loadable for inference but sits beyond what one Spark can practically optimise without splitting the workload across a linked pair of units.

Building serious local AI capability? Browse the AI PC lineup at Evetech and match the hardware to the models you actually intend to run.