IT Infrastructure

MLPerf Inference v6.1: What the New RAG and Agent Benchmarks Mean for Your AI Costs

A clean isometric 3D render of a neat stack of closed hardcover books linked by a thin copper pipe to a compact matte computer cube, with a brass stopwatch resting on top of the cube, set on a pale gray floor tile, a visual for MLPerf Inference v6.1 and its new benchmarks that time how fast hardware retrieves documents and answers questions

MLPerf Inference v6.1, released by MLCommons on September 16, 2026, added a test that times a full retrieval-augmented generation (RAG) pipeline. That is the same kind of system behind an association knowledge base or an AI site search. A record 30 organizations submitted results.

The round also added an Edge Agentic Inference test for small local machines and listed five new chips. This piece covers what MLPerf measures and leaves out, why the RAG and edge tests matter to AI buyers, and how to read vendor claims. It ends with a simple way to turn benchmark numbers into cost per answer.

The short version: MLPerf tells you how fast a tuned system runs a fixed workload at a set accuracy. It does not tell you what anything costs. The new RAG test is close to the workload many nonprofits and associations run. The edge results show that software can matter as much as hardware. Use the numbers to compare systems, then do the cost math with your own quotes.

What MLPerf Inference v6.1 measured

MLPerf Inference is a set of standard tests run by MLCommons, an industry group. Each test fixes a model, a dataset and an accuracy target. Submitters then run the test on their hardware and report how fast it went.

Datacenter tests cover large servers and cloud systems. Edge tests cover smaller devices. MLCommons describes a "Closed" division, which requires "using the same model as the reference implementation," and an "Open" division, which "allows using a different model or retraining." Closed results are the ones to compare across vendors.

Each result also carries an availability label. "Available" means you can buy or rent the system. "Preview" means it must be submitted as available in the next round. "RDI" covers research, development or internal systems.

According to the MLCommons release, v6.1 added three things:

  • End-to-end RAG: a pipeline of several models answering questions from a document collection.
  • Edge Agentic Inference: a coding agent workload on an edge machine, with accuracy and speed tested separately.
  • Speculative decoding: a speed technique that “predicts and verifies multiple tokens in a single forward pass,” now allowed in two interactive tests.

What MLPerf does not tell you

Know the limits before you quote a number to a board or a client.

No prices. MLPerf reports speed and accuracy. The MLCommons release does not include cost, price or energy figures. Any "cost per token" claim you see comes from a vendor, not from MLCommons.

Tuned systems. Submitters tune their software for the test. Your cloud instance or desktop box will run stock settings unless someone does that work.

Fixed models. The tests use specific models, such as DeepSeek R1 and Qwen3.6-27B. Speed on one model does not transfer cleanly to another.

Best case, not typical case. Many headline numbers are the best result in a category. The median system in the same category may be much slower.

Hardware only. A cloud API adds its own queues, rate limits and pricing tiers. MLPerf Inference measures the machine underneath.

Treat the results as one input to a buying decision. We made a similar point about AI model scores in our post on vendor benchmarks versus leaderboards.

The new RAG test and your knowledge base

RAG is the method behind most "ask our resources" tools. A question comes in, the system finds matching passages in your documents, and a language model writes an answer from them. We explained the basics in What Is Retrieval-Augmented Generation (RAG)?.

Other MLPerf Inference tests time a single model. This one times the whole chain. MLCommons describes it this way: "an embedding model turns a query into a vector, a retriever pulls candidate passages from a vector database, a re-ranker refines the list, and one or more LLMs reason over the data to produce an answer."

The reference version in the MLPerf Inference code repository uses five models:

  • e5-base-v2 to turn text into vectors, stored in a FAISS index
  • ColBERTv2.0 to re-rank the passages it finds
  • GPT-OSS-120B to break a question into smaller searches, over up to five rounds
  • GPT-OSS-20B to write the answer
  • Llama-3.1-8B-Instruct to grade answers for accuracy

The documents are a frozen set of 2,515 Wikipedia pages from the FRAMES dataset. The test reports two tasks: building the vector database from the documents, and answering questions against it.

Why the RAG design matters to cost per answer

One member question can mean several model calls. In the reference pipeline, a large model plans the searches, retrieval can repeat up to five times, and a second model writes the answer.

The cost of one answer is the sum of every step. If your vendor quotes a price per token for the answer model only, the real cost per answer will be higher.

The test also times ingestion separately. That is the work of reading your documents and building the index. For an association with a large library of journals or standards, re-indexing after a site migration is a real cost. Ask a vendor how long it takes.

Two cautions. The test runs in the datacenter category, and its corpus is small next to many association archives. We could not confirm which systems submitted RAG results, so check the MLCommons results table before quoting one.

Our guide to vector databases covers the storage side.

The edge agentic test: numbers for a box on your desk

Watch the Edge Agentic test if you are considering a local AI machine. It uses the Qwen3.6-27B model. According to the repository, it accepts any "OpenAI-compatible endpoint," so the same test can run on different software.

It runs in two phases. The accuracy phase uses the Berkeley Function Calling Leaderboard v4, which checks whether the model picks the right tool with valid inputs. A system must score at least 0.97 times the reference result. The speed phase replays 20 recorded coding sessions, 1,007 turns in total, one request at a time. Context grows to about 23,500 tokens as tool results pile up.

NVIDIA’s developer blog reports these results on one Jetson AGX Thor:

  • 52.33 tokens per second
  • 87.94% accuracy on the function-calling test
  • All 1,007 turns finished in 24 minutes 36 seconds
  • The same workload took 2 hours 37 minutes with llama.cpp, a common open-source runtime

Same device, same model, 6.4 times less time. The llama.cpp run used a different 4-bit format, so software and settings both changed. NVIDIA credits its own 4-bit format, speculative decoding, and reusing earlier context between turns. It says about 96% of prompt tokens came from that reused cache.

StorageReview reports that Atlas Inference, a first-time submitter, ran the test at about 20 tokens per second on both an NVIDIA DGX Spark and an AMD Ryzen AI Max+ 395 machine. We covered that chip in our Ryzen AI Max+ 395 explainer. Those runs used different software and settings from NVIDIA’s, so do not compare them directly.

What the edge numbers mean for local hardware

The main lesson is about software. On the same Jetson board, the runtime changed total time by a factor of 6.4. A local box running default software may be several times slower than its best published result.

Hardware costs the same per hour whether it finishes 400 turns or 2,400. The Jetson figures work out to about 385 turns per hour with llama.cpp and about 2,456 with NVIDIA’s tuned stack. The cost per turn falls by the same factor.

Prices for 128GB local machines rose sharply this year, as we reported in local AI hardware prices in 2026. That makes the software question more important. Ask any vendor of a local AI box which runtime it ships with and whether it has a published MLPerf result on that runtime.

The test also runs one request at a time. If five staff members share the box, you need different numbers.

The 5.7x DeepSeek R1 figure, read carefully

MLCommons states that "the best per-accelerator server scenario result" on DeepSeek R1 "was 5.7x better compared to v5.1 one year ago." For the vision-language test, the best result improved 2.99 times in six months.

"Best" means the top system, not the average. "Per accelerator" means one chip, not a full server. "Server scenario" is one of several test modes. And the gain can come from new hardware, better software, or both.

Your cloud bill for DeepSeek R1 did not fall 5.7 times. Providers set their own prices and may run older hardware. Still, serving costs for a given model can drop fast, so avoid multi-year price locks without a review clause.

For background on how processing power is priced and sold, see What Is Compute?.

New chips in MLPerf Inference v6.1

MLCommons lists five new processors or accelerators. The right-hand column is vendor claims.

Chip Status in v6.1 Detail (attributed)
AMD Ryzen AI Max+ 395 Available AMD’s MLPerf posts do not mention it; StorageReview reports a third-party edge result
AMD Instinct MI350P Available PCIe card with 144 GB of HBM3E; AMD says it posted the top PCIe result in the round
Intel Arc Pro B70 Available Workstation graphics card; Intel also reported results on it in v6.0
NVIDIA Rubin and Vera Rubin NVL72 Preview NVIDIA claims up to 2.5x the DeepSeek R1 throughput of GB300 NVL72

Preview systems are not yet available to buy or rent on normal terms. Our AI accelerator landscape post explains where these chip families fit.

How to read vendor claims built on MLPerf Inference

AMD, NVIDIA and several cloud providers posted on September 16, each leading with its best numbers.

Check the comparison. AMD’s post says its MI355X reached 133% of NVIDIA B200 on one GPT-OSS-120B test. NVIDIA’s Vera Rubin post compares its preview system against its own GB300. Each vendor picks the matchup that suits it.

Check "first place." Nebius, a cloud provider, reports five first-place finishes across 20 available categories. That also means 15 categories without a first place.

Check the category. A preview result is a forecast of a product, not a product you can rent today.

Check for cost language. NVIDIA says each Vera Rubin rack lowers "cost per token" but gives no figure. MLPerf did not measure that.

Check software gains. AMD reports 28% to 38% higher throughput on the same MI355X hardware. Trade press reports Intel claims up to 2.4 times on the same Xeon chips. Ask your provider whether you get those updates.

Turning benchmark numbers into cost per answer

You can do this with two numbers. Cost per answer is the system’s hourly cost divided by the answers it completes per hour at your quality bar.

For a local box, the hourly cost is the purchase price spread over its useful life, plus power and staff time. For cloud capacity, it is the hourly rate on your quote. MLPerf Inference results can help with answers per hour, but only for a matching model and workload.

For a cloud API billed per token, cost per answer equals tokens per answer times the price per token, added up across every model call.

Three steps keep the math honest:

  1. Count the calls. Map every model your RAG or agent tool uses for one answer.
  2. Discount the benchmark. Assume your system runs slower than the best MLPerf Inference result until you test it.
  3. Test on your content. Run a week of real questions, then divide the bill by the answers.

Frequently Asked Questions

What is MLPerf Inference?

Standard tests from MLCommons that measure how fast hardware and software run fixed AI models at a set accuracy. New rounds come out about every six months.

When were the MLPerf Inference v6.1 results released?

September 16, 2026. MLCommons said 30 organizations submitted, a record for the benchmark.

Does MLPerf measure the cost of AI?

No. It reports speed and accuracy. Cost claims built on MLPerf come from vendors, so ask for their math.

What does the new RAG benchmark test?

A full question-answering pipeline: an embedding model, a vector search, a re-ranker and language models that plan searches and write the answer. Indexing and answering are timed separately.

Why does the RAG test matter for an association website?

Member knowledge bases and AI site search use the same pattern. The test shows that one answer can involve several model calls, so the cost per answer is higher than the price of a single prompt.

What is the Edge Agentic Inference benchmark?

A test for small local machines. It replays 20 recorded coding sessions, 1,007 turns in total, with the Qwen3.6-27B model, and separately checks function-calling accuracy against a minimum score.

Was the AMD Ryzen AI Max+ 395 tested?

Yes. MLCommons lists it as new and available. StorageReview reports that Atlas Inference ran the edge agentic test on it at about 20 tokens per second.

Does the 5.7x DeepSeek R1 gain mean cloud AI got 5.7 times cheaper?

No. It compares the best single-chip result in one test with the best one a year earlier. Cloud prices depend on each provider’s choices.

Can I buy NVIDIA Vera Rubin NVL72 now?

MLCommons lists both Rubin systems as preview. Preview systems must be submitted as available in the next round. Ask your provider before planning around them.

Digital Matters

IT Infrastructure Desk