NVFP4 vs MXFP4 Decode Benchmark
If you've been following the latest FP4 formats, you've probably read a bunch of articles comparing NVFP4 and MXFP4: block sizes, scale formats, yada yada.
What I personally haven't seen much of is actual results on actual production inference workloads. Does it matter?!
So, Stas did measure that on a B200, NVFP4 gets ~9% more TFLOPS than MXFP4 on large GEMMs. Well, that's a pretty big difference!
But it's still an isolated GEMM benchmark. Does the perf difference actually translate to production workloads?
Also, what about other non-GEMM-bound workloads, like decode? Also ... why?!
Well, we can test it! Perhaps we can sus it out. (which is the whole point of this post, and probably why you're still reading)
Ok. I tried to keep it simple by picking a realistic-ish dense LLM (Qwen3-32B), served with vLLM, on a B200. Both formats were quantized from scratch from the same BF16 weights, in this way we try to keep the confounders to a minimum.
My hypothesis was that MXFP4 should be faster as we saturate HBM throughput due to MXFP4's slightly smaller bits (4.25 vs 4.5). aka: this is the classic bandwidth-bound workload so why not?
TL;DR

The benchmark Claude mostly built & I ran suggests that if you're using a B200 GPU, you should probably always prefer NVFP4 over MXFP4, by a meaningful margin (up to ~8% faster decode at small batches)!
This holds for low-batch decode, and for GEMM-bound training NVFP4 looks better too (see below), though raw large-GEMM speed depends heavily on the kernel. The difference disappears at larger batch sizes, which at first didn't surprise me: I assumed we were becoming HBM-bandwidth bound as the batch and KV-cache reads grow. I won't spoil it, but that turned out not to be the reason!
Additionally, quantizing from BF16 to the respective formats and measuring the loss indicates that NVFP4 should yield non-trivially better eval perf. I didn't run any evals to confirm this however.
Now, these findings may generalize but as Stas suggests, you should always strive to benchmark your actual production configuration.
Kernel implementations seem to dominate the perf differences
On a B200 (SM100) with vLLM v0.31, the NVFP4/MXFP4 decode perf difference mostly comes down to the kernel implementation, not to the 4.5 vs 4.25 bits difference. Note again that I tested this on only one dense model (Qwen3-32B, w/ BF16 KV cache).
Small-batch decode: NVFP4 is meaningfully faster
+7.1%, +7.6% and +4.2% decode throughput at batch 1, 8 and 32.
Worth highlighting that decode for this workload is neither bandwidth-bound (~40–50% of theoretical HBM bandwidth) nor compute-bound (the GEMMs are tiny), so kernel implementation differences dominate. NVFP4's superior perf here seems to come mostly from the kernels around the GEMMs: activation quantization, fewer kernels per step, etc.

Batch ≥ 64: there's no meaningful difference
So you might as well pick the format based on eval perf, portability or other concerns if you're in this regime.
Batch size dominates the perf difference & it's not due to HBM memory bandwidth saturation
NVFP4's non-GEMM kernels are faster, but its GEMMs fall behind MXFP4's as they grow in size, and from batch 64 up the two differences cancel each other out.
So is it the batch size or the number of tokens that end up dominating? I kept the batch at 1 and grew the context instead, and separately held the tokens constant and grew the batch:

Turns out it's the batch size. Look @ batch 1, NVFP4 keeps saving the same ~0.4 ms per step all the way out to 127k tokens of context (the % only shrinks because each step gets longer). But at the same 128k tokens, bumping the batch to 128 wipes the saving out. More memory traffic didn't make a significant enough difference, but the bigger GEMMs did! Huh?
But why?!
Perhaps a longer context only adds attention work, i.e. more KV cache reads. Remember that I kept the KV cache at BF16 for both formats, so both get slower by the same amount and NVFP4 keeps its lead. So, a bigger batch, on the other hand, means bigger GEMMs, and vLLM autotunes MXFP4's GEMM but (I think?) runs NVFP4's on a fixed heuristic. So as the GEMMs grow, NVFP4's fall further behind MXFP4's, until that slowdown cancels out NVFP4's lead. Maybe!
Now, I'm still not 100% sure it isn't bandwidth: @ batch 128, MXFP4's small edge matches what the byte difference should theoreticaly be. But @ batch 1 w/ a 127k context, with the same memory traffic, and similar bandwidth use, NVFP4 keeps its lead, and the GEMM size is the only thing that changed between the runs.
This confuses me enough that I'll revisit it and read some metrics straight from the GPU to either confirm or rule it out.
Large GEMMs (prefill/training) depend heavily on the kernel implementation
In PyTorch, NVFP4 gets ~9% more TFLOPS (ml-engineering book), but there (at least in the PyTorch 2.13 I checked) NVFP4 runs on cuBLASLt and MXFP4 on a different library (MSLK). In vLLM's kernels (CuTe-DSL), NVFP4's largest GEMMs were from 4% to 20% slower at batch 512.
Additionally, for training, the format itself matters too. NVIDIA's own paper found that an 8B model trained with MXFP4 needed 36% more tokens (1.36T vs 1T) to match NVFP4's loss, while both formats run at the same raw throughput on Blackwell.
You should serve NVFP4 on vLLM's default kernel
That's CuTe-DSL (--linear-backend flashinfer_cutedsl). On cuDNN (flashinfer_cudnn) the same
weights are between 10 and 18% slower and lose to MXFP4 at every batch size. MXFP4 only has a single
W4A4 kernel in vLLM 0.31 (FlashInfer's mm_fp4 with the CuTe-DSL backend).
NVFP4 is more accurate
Its NLL penalty after quantization from BF16 is 2.7× smaller. While no evals were run, the lower NLL (negative log-likelihood) suggests that NVFP4 should yield better eval results.
The code, quantized weights and full write-up are on GitHub and HF (MXFP4, NVFP4).