AI & HPC

NVIDIA DGX Spark: AI Supercomputing Hosted in Madrid

A petaFLOP of AI performance in a compact box: GB10 Grace Blackwell superchip, 128 GB of unified memory and a 200 Gbps interconnect to cluster nodes. We look at what the DGX Spark can do and why hosting it in a datacenter makes more sense than keeping it in the office.

business EasyDataHost calendar_today July 14, 2026 schedule 8 min read

Until recently, getting a petaFLOP of AI compute meant reserving a full rack, an industrial power contract and a six-figure budget. The NVIDIA DGX Spark changes that equation: it condenses the Grace Blackwell architecture of the large DGX systems into a compact unit that fits in two hands and delivers up to 1 petaFLOP of FP4 performance.

For AI development teams, researchers and companies that want to work with language models without depending on third-party APIs, the DGX Spark offers something very concrete: 128 GB of unified memory capable of loading models that fit in no consumer GPU, running the same software stack as NVIDIA's DGX supercomputers.

In this article we review what the DGX Spark actually is, its technical specifications, its real-world use cases and a question many companies overlook: where it should be installed. At EasyDataHost we offer it hosted in our Madrid datacenter, in 1, 2 or 3-node configurations, and in the following sections we explain why that option beats keeping it on an office desk.

What Is the NVIDIA DGX Spark

The DGX Spark is the most compact member of NVIDIA's DGX family: an AI supercomputing unit built around the GB10 Grace Blackwell superchip, which integrates in a single package a 20-core Arm CPU (Grace architecture) and a Blackwell GPU with fifth-generation Tensor Cores, connected over NVLink-C2C with far more bandwidth than PCIe.

The practical consequence of that integration is unified memory: CPU and GPU share the same 128 GB of LPDDR5x with no need to copy data between system RAM and graphics VRAM. For AI workloads this removes one of the classic bottlenecks and, above all, makes it possible to load much larger models than any consumer GPU with 16 or 24 GB of VRAM can handle.

The unit ships with DGX OS preinstalled, NVIDIA's Ubuntu-based distribution that includes the full NVIDIA AI stack: CUDA drivers, NGC containers, frameworks such as PyTorch and TensorRT-LLM, NIM microservices and the NVIDIA AI Enterprise platform tools. It is the same software environment that runs on datacenter DGX systems, so a project born on a DGX Spark can later scale to larger infrastructure without rewriting anything.

DGX Spark Technical Specifications

These are the unit's key specifications, summarised around the parameters that matter most when sizing AI workloads:

Component Specification
Superchip NVIDIA GB10 Grace Blackwell
CPU 20 Arm cores (Grace architecture)
GPU NVIDIA Blackwell with 5th-generation Tensor Cores
AI performance Up to 1 petaFLOP at FP4
Memory 128 GB unified LPDDR5x (CPU + GPU)
Storage High-performance NVMe SSD
Cluster interconnect NVIDIA ConnectX at 200 Gbps
Operating system DGX OS with the NVIDIA AI stack preinstalled
Form factor Compact desktop-sized unit, low power draw

The key figure: unified memory

A high-end consumer GPU offers 24 GB of VRAM; the DGX Spark puts 128 GB at the model's disposal. That is the difference between running a 7B-parameter model and working with models of up to ~200B at FP4 without splitting layers or offloading to disk.

Use Cases: What a DGX Spark Is For

The DGX Spark does not compete with large-scale training clusters; its territory is the full cycle of development, tuning and local inference of AI models. The scenarios where it performs best:

  • check_circle LLM prototyping and fine-tuning: adapt open models (with techniques such as LoRA or QLoRA) on your own data, iterate fast and validate hypotheses before investing in large-scale training.
  • check_circle Local inference of large models: a single node runs models of up to ~200 billion parameters at FP4; clustering 2 or 3 nodes lets you work with even larger models.
  • check_circle AI development with private data: RAG pipelines, agents and internal assistants where the company's confidential data never leaves a controlled infrastructure — a common requirement in healthcare, banking or the public sector.
  • check_circle Predictable cost versus APIs: for sustained inference workloads, a dedicated unit removes the variable per-token bill and the risk of third-party price changes or usage limits.

If your workload requires intensive training or massive production inference, it probably fits better on our dedicated GPU servers. To understand which configuration suits each phase of an AI project's life cycle, our article on GPU servers for AI and machine learning compares the alternatives in detail.

Scaling by Clustering Nodes: ConnectX at 200 Gbps

One of the features that sets the DGX Spark apart from a plain workstation is its NVIDIA ConnectX interconnect at 200 Gbps, designed to join several nodes into a small AI cluster. Two directly connected nodes add up to 256 GB of unified memory, enough to run inference for models that overflow the capacity of a single unit.

NVIDIA's stack handles the distribution of the model across nodes, so clustering units does not require redesigning the application. For a development team this means a natural growth path: start with one node for prototyping, add a second when the model grows and reach three to serve large models or spread workloads across several simultaneous projects.

That said, the 200 Gbps interconnect introduces a practical requirement: cabling and location matter. Nodes must be physically close, correctly cabled and fed with stable power. It is exactly the kind of installation that is better solved in a datacenter rack than in an office.

Why Host It in a Madrid Datacenter and Not in the Office

The DGX Spark's compact form factor invites you to put it on a desk. But a unit that will run AI workloads for hours or days at a time is, for all practical purposes, production infrastructure, and an office is not the environment that role was designed for: unprotected power cuts, uncontrolled temperature in summer, constant fan noise next to the team, and a device holding valuable data within reach of anyone walking through the room.

That is why at EasyDataHost we offer the NVIDIA DGX Spark hosted in our Madrid datacenter. The advantages over an office installation are direct:

  • check_circle Redundant power: UPS and generator backup guarantee that a 48-hour fine-tuning run is not lost to a micro-outage halfway through.
  • check_circle Controlled cooling: stable temperature and humidity around the clock, which translates into sustained performance without thermal throttling and a longer hardware lifespan.
  • check_circle Low-latency network: datacenter-grade connectivity to move datasets, serve inference to your applications and sync repositories without depending on the office fibre line.
  • check_circle Physical security: access control, video surveillance and fire protection. A unit holding your models and your data should not be within reach of any visitor.
  • check_circle No noise or power consumption in the office: the unit works at full throttle without taking up space, without fans next to the human team and without inflating the company's electricity bill.
  • check_circle Secure remote access: your team works against the DGX Spark over SSH or VPN from anywhere, exactly as if it were sitting next to them, but with the reliability of a professional environment.

There is also a data sovereignty argument: the unit is physically in Madrid, within the GDPR framework and Spanish jurisdiction. For AI projects handling personal or sensitive data, knowing exactly which room houses the hardware that processes them makes compliance work considerably simpler.

DGX Spark at EasyDataHost: 1, 2 and 3-Node Configurations

At EasyDataHost we offer the DGX Spark hosted in our Madrid datacenter in three configurations, designed to accompany the growth of an AI project:

  • arrow_right 1 node: 128 GB of unified memory. Ideal for prototyping, fine-tuning and inference of models up to ~200B parameters at FP4.
  • arrow_right 2 nodes: 256 GB of aggregate unified memory over the 200 Gbps ConnectX interconnect, for larger models or parallel workloads.
  • arrow_right 3 nodes: the full small cluster, with 384 GB of aggregate memory to serve several models or working teams at the same time.

Every configuration includes rack installation, redundant power, connectivity and secure remote access. Up-to-date pricing and terms are available on the NVIDIA DGX Spark product page. And if you are not sure which configuration fits your project, contact our team for a technical assessment with no obligation.

Frequently Asked Questions

What AI models can an NVIDIA DGX Spark run?

With its 128 GB of unified memory, a single node can run inference for models of up to roughly 200 billion parameters quantized to FP4, and fine-tune mid-sized models. By clustering 2 or 3 nodes over 200 Gbps ConnectX you can work with even larger models.

Why host it in a datacenter instead of the office?

In a datacenter the unit runs on redundant power, controlled cooling, a low-latency network and 24/7 physical security, without generating noise, heat or electricity costs in the office. Access is remote and secure from anywhere, and the hardware is protected against power cuts, theft or accidents.

What configurations does EasyDataHost offer and how much do they cost?

We offer the DGX Spark hosted in our Madrid datacenter in 1, 2 or 3-node configurations, with additional nodes interconnected at 200 Gbps. Up-to-date pricing and terms are published on the product page.

Conclusion

The NVIDIA DGX Spark brings AI supercomputing within reach of teams that until now had to choose between insufficient consumer GPUs or cloud infrastructure with variable costs. And its best version is not on a desk, but in a rack:

  • arrow_right The GB10 Grace Blackwell superchip with 128 GB of unified memory makes it possible to work with models of up to ~200B parameters at FP4 in a single compact node.
  • arrow_right The 200 Gbps ConnectX interconnect turns 2 or 3 nodes into a small cluster capable of serving larger models, with a growth path that requires no application rewrite.
  • arrow_right Hosted in a Madrid datacenter, the unit gains redundant power, cooling, physical security and remote access, while the office is spared the noise and the power bill.
  • arrow_right EasyDataHost offers it in 1, 2 or 3-node configurations, with installation, connectivity and support included, and pricing published on the product page.
NVIDIA DGX Spark Grace Blackwell AI HPC LLM Madrid
rocket_launch

AI supercomputing without leaving Madrid

NVIDIA DGX Spark hosted in our datacenter: redundant power, cooling, low-latency network and secure remote access. 1, 2 or 3-node configurations with 24/7 support.