Nvidia Vera Rubin AI Server Shows Significant Performance Gains in Testing
Nvidia's Vera Rubin AI server has undergone performance testing with the DeepSeek-V4-PRO 1.6T model. Results indicate substantial improvements in throughput per megawatt.

Nvidia has evaluated its Vera Rubin AI server using the SemiAnalysis AgentX workload and the DeepSeek-V4-PRO 1.6T model. The tests show that the Vera Rubin system achieves up to 30 times higher throughput per megawatt compared to the previous Grace Blackwell generation.
Throughput per megawatt (Tokens per Megawatt) is a key metric for measuring the energy efficiency of AI data centers. It quantifies the number of AI tokens that can be generated within a fixed power budget.
In tests with the DeepSeek-V4-PRO 1.6T workload, the Grace Blackwell-based GB300 NVL72 server demonstrated 15 times greater throughput per megawatt compared to the H200 NVL8 Hopper configuration. The cost per million tokens for Blackwell servers is approximately one-tenth of the previous generation, allowing for more AI agents within the same power and infrastructure budget or maintaining the same capacity at lower operational costs.
The Vera Rubin system further enhances this advantage. On the DeepSeek V4 Pro model, the Vera Rubin NVL72 system delivers approximately 30 times the throughput per megawatt of the GB300 NVL72. Additionally, the cost per million tokens is reported to be 35 times lower, enabling continuous and large-scale operation of AI agents to meet diverse customer workload demands.