494 tokens per second on four DGX Sparks
A four-node NVIDIA DGX Spark rack ran DeepSeek V4.1-Flash at 494 tokens per second for code with 32 concurrent requests.

494 tokens per second is the peak code output measured on a four-node DGX Spark cluster running DeepSeek V4.1-Flash with 32 concurrent requests. That number does not prove the rig is a general-purpose replacement for a data-center inference node. It measures one model, one workload, one topology, and one burst condition.
The constraint is memory and interconnect, not raw FLOPs. Four 128 GB LPDDR5X pools give about 512 GB of coherent memory, while the switchless QSFP ring avoids a 200 GbE switch but leaves the cluster dependent on a narrow ring path. For a buyer, the claim is that a compact rack can host a 552B-parameter sparse model at practical throughput without optical routing. A fab or platform vendor faces pressure on memory bandwidth and low-latency node links, not on adding more GPU cores.

The article does not state power draw under load, thermal headroom, cost per token, or sustained throughput after the peak. Those gaps matter more than the headline number.
The cluster is constrained by memory bandwidth and ring interconnect, not by GPU FLOPs. The buyer takeaway is that a compact rack can serve a large sparse model, but the document does not prove sustained production value.
The document does not state sustained throughput, power draw under load, thermal headroom, or cost per token.
The 494 tok/s code peak comes from a 552B sparse model with 8B/16B active parameters and 890-byte KV cache, so the bottleneck is memory and node links. The switchless QSFP ring reduces cost but leaves the topology unproven for sustained load.
494 tok/s (code, 32 concurrent · 280 tok/s (32 concurrent) · ~58 tok/s prose, ~96 tok/s code · ~4,764 tok/s prompt processing · ~0.2 s first token idle · 4x NVIDIA DGX Spark units
After Wccftech Hardware. We did not report this. The pictures, if any, are theirs.


