Nvidia and AMD have both found reasons to celebrate the same set of AI benchmarks. That should make anyone curious about the chip race look beyond the victory headlines.
The new MLPerf Inference 6.1 results, released on 16 September 2026, measure how computer systems run trained AI models. This is the work behind answering prompts and producing other AI outputs. The suite covers multiple tasks and operating conditions; it does not award one universal “best AI chip” trophy. MLCommons announcement
The revealing part is where each company is pressing its advantage.
Nvidia brings its next generation into view
Nvidia’s Vera Rubin NVL72 appears in this round as a preview submission. The company reports up to 3.7 times the throughput of its GB300 NVL72 on Qwen3-VL, and up to 2.5 times on DeepSeek-R1. Those are comparisons with Nvidia’s own earlier platform on particular workloads, rather than a blanket lead over every AMD system. Nvidia’s results
Throughput measures how much work a system completes in a given time. More throughput can mean serving more requests with a rack of equipment, but it does not by itself establish a customer’s electricity bill, purchase price or experience during a busy afternoon.
Nvidia also reports a different achievement on GB300: four racks containing 288 GPUs reached 99% scaling efficiency on the DeepSeek-R1 offline test. In that particular setup, adding hardware produced an almost proportional increase in output. Nvidia’s scaling disclosure
AMD makes the case for hardware already in place
AMD’s most interesting number may be a software result. On the same eight MI355X GPUs, it reports that GPT-OSS-120B throughput rose 28% in the offline scenario and 38% in the server scenario between MLPerf rounds 6.0 and 6.1. The company attributes the gains to improvements in ROCm and the surrounding software. AMD’s submission
That is a compelling argument for buyers: a system can become more useful without replacing the chips.
AMD also highlights wins over selected Nvidia B300 submissions, including 18% higher offline performance on the Wan-2.2 video-generation benchmark. The word selected matters. These results describe particular submitted configurations and workloads. They cannot settle every comparison between the two companies. AMD’s comparison and methodology
The useful contest is bigger than a leaderboard
MLCommons is expanding the tests themselves. This round adds an end-to-end retrieval-augmented generation benchmark, covering the chain from finding relevant information to generating an answer. It also introduces an edge test for multi-step AI workloads. That reflects how useful AI systems often involve several operations, rather than one isolated model response. MLCommons
Our reading is that the competitive picture is becoming more specific. Nvidia is demonstrating a substantial next-generation step, while AMD is making a stronger argument about the useful performance it can extract today.
For customers, the decisive comparison remains their own workload: the same model, accuracy requirements, response-time limits and realistic costs. A benchmark result is evidence to investigate, rather than a substitute for that work.
For everyone watching the rivalry, this is the encouraging development: the fight includes software progress and reproducible results alongside new silicon. That gives competitors more than one way to challenge the leader.
Related reading: AMD’s ROCm software push.
Featured image: Nvidia’s illustration accompanying its MLPerf Inference 6.1 announcement. Credit: NVIDIA.


Leave a Reply