Peter Zhang
Sep 09, 2025 16:44
NVIDIA’s Blackwell Extremely structure units new benchmarks in AI inference efficiency with its debut in MLPerf Inference v5.1, showcasing vital developments in LLM computation.
NVIDIA’s newest technological development, the Blackwell Extremely structure, has made a big impression within the subject of synthetic intelligence by setting new information in AI inference efficiency, in accordance with the official NVIDIA weblog. The debut of Blackwell Extremely within the MLPerf Inference v5.1 benchmark, a number one business commonplace for AI efficiency, highlighted its superior capabilities in dealing with giant language fashions (LLMs).
Benchmark Achievements
The MLPerf Inference v5.1 benchmark consists of a wide range of checks to gauge AI inference efficiency, with NVIDIA’s Blackwell Extremely setting new information throughout a number of newly launched fashions. These embrace DeepSeek-R1, a 671-billion parameter mixture-of-experts mannequin, and the Llama 3.1 sequence fashions. The Blackwell Extremely platform demonstrated distinctive efficiency, surpassing all earlier benchmarks and sustaining per-GPU efficiency information.
Notably, the structure delivered as much as 1.5 occasions greater peak NVFP4 AI compute and doubled the attention-layer compute capabilities in comparison with its predecessors. The introduction of upper HBM3e capability additionally contributed to those developments.
Technological Improvements
The Blackwell Extremely structure incorporates a number of revolutionary applied sciences that improve its efficiency. In depth use of NVFP4 acceleration throughout all DeepSeek-R1 and Llama mannequin submissions performed a vital position in attaining these outcomes. Moreover, the structure’s capability to optimize key-value caches utilizing FP8 precision considerably lowered reminiscence footprint and improved efficiency.
New parallelism methods, similar to skilled parallelism for the MoE portion and information parallelism for the eye mechanism, have been employed to maximise multi-GPU execution. These methods have been complemented by way of CUDA Graphs to scale back CPU overhead throughout inference processes.
Implications for AI Inference
The outcomes from the MLPerf Inference v5.1 benchmark underscore NVIDIA’s continued management in AI inference efficiency. The Blackwell Extremely structure not solely enhances throughput and effectivity but additionally reduces the price per token considerably. That is notably evident within the comparability with Hopper-based programs, the place Blackwell Extremely delivered roughly 5 occasions greater throughput per GPU.
The introduction of disaggregated serving methods additional highlights NVIDIA’s innovation in AI infrastructure. By decoupling context and era throughout separate GPUs or nodes, NVIDIA has optimized useful resource use, notably for big language fashions like Llama 3.1 405B.
Future Prospects
NVIDIA’s developments in AI inference expertise proceed to set new requirements within the business. The Blackwell Extremely structure, with its record-breaking efficiency, positions NVIDIA on the forefront of AI innovation. Because the demand for extra subtle AI fashions grows, NVIDIA’s dedication to increasing its technological capabilities stays evident.
The introduction of Rubin CPX, a processor designed to speed up lengthy context processing, additional exemplifies NVIDIA’s dedication to pushing the boundaries of AI effectivity and efficiency.
Picture supply: Shutterstock


