Intro
The first thing about FP4 that I could not explain to myself was a byte. NVFP4 and MXFP4 use the same 4-bit element, E2M1: one sign bit, two exponent bits, one mantissa bit. A Blackwell tensor core multiplies the same sixteen codes in both formats.
What differs is how a block of values stores its scale: one byte for every 32 values in MXFP4, a different kind of byte for every 16 values in NVFP4.
Every comparison I read puts NVFP4 ahead, sometimes far ahead, and yet OpenAI’s gpt-oss, DeepSeek V4 and Moonshot’s Kimi K3 all shipped in MXFP4, and Moonshot’s previous flagship used neither format.
I wanted the difference in numbers rather than adjectives, and I wanted to know what the silicon does with each format. Almost everything below comes from work that needs no GPU.
We rebuilt the rounding rules of every block format we could find in NumPy and checked them against two independent references: a published table of Gaussian errors, and AMD’s measurements on the activations of two real models.
Then we asked NVIDIA’s own assembler, ptxas 13.4.92 from the PyPI wheels, which FP4 instructions eight targets from Hopper to Rubin accept and what machine code each one becomes. Where measurement was impossible we read the vendor code that ships with the hardware.
At a glance
One byte of scale separates NVFP4 from MXFP4. It is worth 1.6 dB of signal-to-noise on Gaussian data, 4.2 dB on activations and about 9% of perplexity under plain rounding.
Under the open standard’s scale rule, 41.5% of MXFP4 blocks clip their own maximum, and the rule that rescued training is the worse one for inference error.
One constant turns format noise into perplexity, explaining 98.7% of the variance across 21 formats; its value changes from model to model.
FP4 arithmetic needs both operands in FP4. The big open releases ship FP4 weights with 8-bit or 16-bit activations, so they collect the bandwidth half only.
INT8 is gone from the tensor-core instructions of B300 and Rubin. Rubin adds a wider scale format, UE5M3, and 3-bit lookup-table weights.
Sixteen codes
E2M1 has a sign bit, two exponent bits with a bias of 1 and one mantissa bit. Its non-negative values are 0, 0.5, 1, 1.5, 2, 3, 4 and 6. Negative zero is a separate code with the same value, so sixteen codes give fifteen distinct numbers.
The grid is a floating-point format in miniature: neighbours are 0.5 apart up to 2, then 1 apart, then 2 apart. Nothing lies between 4 and 6, so a value that lands on 5 after scaling ends up 20% away from wherever it rounds.
An INT4 grid with the same maximum puts its levels 6/7 apart everywhere. Which grid is better depends entirely on where values land after scaling (Figure 1).
Every block format scales a block so that its largest value lands near the top of the grid, which pushes the other values of a bell-shaped block toward zero, where the FP4 grid is dense: in our simulation 50% of the values of a Gaussian block land below 2, where half of the non-negative levels live.
That is the case for floating point at four bits, and also its weak spot: the few values near the top of each block fall into the coarsest part of the grid.

The yardstick throughout this piece is the signal-to-quantization-noise ratio (SQNR): ten times the base-10 logarithm of signal power over error power. Each extra bit of a well-designed quantizer is worth about 6.02 dB.
The best possible 16-level quantizer for a Gaussian, the Lloyd-Max quantizer, needs no scales because it is told the variance in advance; it reaches 20.19 dB in our simulation, against 20.22 dB in the textbooks.
Real formats must discover the scale from the data, block by block, and they pay for it in bits. MXFP4 spends 8 bits of scale on every 32 values, 4.25 bits per element in total. NVFP4 spends 8 bits on every 16 values plus 32 bits per tensor, 4.5 bits per element.
NVFP4 therefore shrinks an FP8 checkpoint by 1.78 times rather than 2, a detail that matters again when we get to the critical batch.
One byte of scale
MXFP4 comes from the Open Compute Project’s Microscaling specification (OCP MX v1.0, in september 2023). Its scale is E8M0: eight exponent bits and no mantissa, so every scale is a power of two.
The specification sets the shared exponent of a block to floor(log2 amax) - 2, where amax is the largest magnitude in the block and 2 is the exponent of 6, the largest E2M1 value.
Divide the block by that scale and its largest value lands at 4 x 2f, where f is the fractional part of log2 amax. Because f is somewhere in [0, 1), the block maximum lands somewhere in [4, 8), and everything above 6 saturates to 6.
The maximum is clipped whenever 2f exceeds 1.5, that is whenever f exceeds log2 1.5 = 0.585. If block magnitudes are spread evenly on a log scale, which is roughly what happens across the thousands of blocks and dozens of layers of a model, that is 1 - log2 1.5 = 41.5% of all blocks.
Our simulation over 128 tensor scales gives 41.51%. A clipped value can lose up to a quarter of its magnitude: 7.9 becomes 6.
NVFP4, introduced with Blackwell, stores the block scale as E4M3, the FP8 format with three mantissa bits, and computes it as amax / 6 rounded to the nearest E4M3 value (NVIDIA technical blog, June 2025). The block maximum then lands within about 6% of 6 instead of anywhere between 4 and 8, and the small overshoot rounds back to 6.
The price is range. Unsigned E4M3 runs from 2-9 to 448, about 18 binades, which is not enough to hold the scales of an entire model. NVFP4 therefore adds a second level: one FP32 number per tensor, equal to the tensor’s amax / (6 x 448), divides the whole tensor before the block scales are computed.
There is a second way to build an E8M0 scale. Asit Mishra, Dusan Stosic, Simon Layton and Paulius Micikevicius showed that the OCP rule can make MXFP8 pretraining drift away from its BF16 baseline, and fixed it by rounding the scale up, so that amax divided by the scale never exceeds the largest element (Recipes for Pre-training LLMs with MXFP8, arXiv:2506.08027).
Rounding up removes clipping, but the block maximum now lands anywhere in (3, 6], which leaves the top of the grid unused for most blocks (Figure 2). Both rules exist in silicon.
On every Blackwell target we tried, cvt.rz.satfinite.ue8m0x2.f32 and its .rp twin each assemble to a single F2FP.SATFINITE.E8 instruction with a .RZ or .RP modifier, so the choice costs nothing at run time.

For inference the two rules are not equivalent, and the rule that rescued training is the worse one for error. Averaged over tensor scales, the OCP rule gives 18.75 dB on Gaussian data against 18.65 dB for rounding up.
On the activation proxy introduced below the gap is 0.50 dB, again in favour of the OCP rule (16.69 against 16.19 dB). Clipping one value per block costs less squared error than coarsening every other value. Training is hurt by bias, and a clipped maximum is biased; a single forward pass is hurt by squared error.
Clipping one value per block costs less squared error than coarsening every other value in it.
The sawtooth
Power-of-two scales have a property that is easy to miss: the error depends on where the tensor sits relative to powers of two. Multiply a tensor by a constant c and every MX scale shifts with it, but the fractional part of log2 c moves every block maximum within its binade, and with it the clipping fraction.
We multiplied the same 262,144 Gaussian values by 2t for 128 values of t between 0 and 2 (Figure 3). NVFP4 does not move at all: 20.44 dB at every t, because the FP32 tensor scale absorbs the constant exactly.
MXFP4 under the OCP rule moves between 18.58 and 18.96 dB with a period of exactly one binade, and the share of clipped blocks swings between 26.8% and 55.9%. Rounding up moves between 18.52 and 18.79 dB.

The swing is small on Gaussian data, about 0.4 dB from trough to peak, but it means that an MXFP4 model’s accuracy depends on constants nobody chose for this purpose, such as the gain of the normalization layer in front of a projection. It also points at a cheap remedy.
A per-tensor multiplier that moves each tensor to its best phase recovers the distance from the average to the peak, 0.20 dB here.
That remedy is exactly NVFP4’s structure, an FP32 number per tensor above the block scales, and a November 2025 preprint applies the same two-level idea to MXFP8 training, keeping power-of-two block scales under an FP32 tensor scale (arXiv:2511.05811).
Where the error lives
Before trusting any of these numbers we checked the simulator against someone else’s. Jack Cook and colleagues at MIT and NVIDIA publish the mean squared error of five block formats on standard normal data (Adaptive Block-Scaled Data Types, arXiv:2603.28765, Table 1).
Our implementation reproduces all five within 1%: 13.23 against their 13.2 (in units of 10-3) for MXFP4, 9.04 against 9.0 for NVFP4, 7.56 against 7.5 for NVFP4 with Four Over Six, 7.44 against 7.4 for NVINT4 and 6.16 against 6.2 for their IF4 format (Figure 4, left).
The agreement says that the rounding details (ties to even, saturation, scale rounding, the FP32 tensor scale) match theirs, not only the headline numbers.

With the simulator anchored, the next question is where inside each block the error comes from (Figure 5).
Under the OCP rule, the single largest value of a 32-value MXFP4 block carries 23.1% of the block’s squared error on Gaussian data and 39.9% on activations, against the 3.1% it would carry if error were spread evenly.
Its average relative error is 11.9%. In NVFP4 the block maximum carries 2.1% and 5.5% of the error, at or below its even share of 6.25%, with an average relative error of 2.25%.

NVFP4 still has a soft spot near the top of its grid, because nothing exists between 4 and 6. Four Over Six, from Cook, Junxian Guo, Song Han and colleagues (arXiv:2512.02010), scales each block either to 6 or to 4, whichever gives less error.
Scaling to 4 gives up the levels 6 and -6 but makes 3 a level at 75% of the block maximum. In our simulation the choice is worth 0.78 dB on Gaussian data and 0.30 dB on activations.
IF4, from the same group, goes further: each block chooses between the FP4 grid and an INT4 grid rescaled to the same maximum, and records the choice in the sign bit of its E4M3 scale, a bit NVFP4 never uses because scales are positive.
It reaches 22.11 dB on Gaussian data, 1.7 dB above NVFP4 at the same 4.5 bits.
Tails
Gaussian data is the friendly case, and activations are not Gaussian. The input to the down projection of a SwiGLU MLP is silu(a) x b, the product of two roughly Gaussian projections, and products of Gaussian variables have heavy tails.
Our proxy is exactly that product, with independent standard normal a and b. Its kurtosis is about 22, against 3 for a Gaussian.
A proxy that simple needs a reality check, and AMD published the one we needed. In June 2026 its ROCm team measured SQNR on the post-SiLU down_proj inputs of every layer of Llama-3.1-8B and Qwen3.6-27B, using the OCP reference quantizers (Atre, Bao, Tiwari and Sirasao, AMD ROCm blog, 26 June 2026).
Their layer means were 16.8 and 16.4 dB for MXFP4, 29.3 and 28.6 dB for MXFP6 E2M3, and 32.8 and 32.1 dB for FP8 with one scale per token. The proxy gives 16.69, 28.97 and 31.69 dB (Figure 4, right).
The first two fall between AMD’s two models and FP8 is within 1.1 dB. For a distribution defined in one line that is closer than we expected, and it is the activation model we use from here on.
On heavy tails the formats separate (Figure 6). NVFP4 improves, from 20.43 to 20.90 dB, because a floating-point grid suits blocks with one large value and many small ones. MXFP4 falls, from 18.79 to 16.69 dB. The NVFP4 advantage grows from 1.64 dB on Gaussian data to 4.21 dB on activations. Integer grids suffer most.
NVINT4, which beats NVFP4 on Gaussian data at 21.29 dB, drops to 18.55 dB, and INT4 with a scale per 32 values, the grouping of Moonshot’s Kimi K2 Thinking checkpoint (we model the scale as BF16), drops from 20.26 to 16.64 dB.

Weights sit in between. We extracted the 84.9 million linear-layer weights of a real trained transformer, the RoBERTa-base encoder packaged in spaCy’s en_core_web_trf 3.8.0, read them without PyTorch and quantized every matrix along its input dimension.
Their kurtosis is 6.0 and the largest weight lies 27 standard deviations from zero. NVFP4 reaches 20.51 dB, MXFP4 18.58 dB, NVINT4 20.87 dB, Four Over Six 21.23 dB and IF4 21.98 dB.
The NVFP4 advantage over MXFP4 is 1.94 dB pooled, and between 1.67 and 2.55 dB across the 48 matrices. RoBERTa is a 2019 encoder, not a modern decoder, so we read it as a sanity check on the synthetic numbers rather than as a benchmark.
Rotations
The usual cure for heavy tails is a rotation. Multiply blocks of values by a Hadamard matrix and each output becomes a signed average of all inputs, so a single outlier is spread across the block and the distribution moves toward Gaussian.
QuaRot and SpinQuant made this standard practice for INT4. For FP4 it is less obviously useful. Vage Egiazarian, Dan Alistarh and colleagues show that rotations help MXFP4 but hurt NVFP4 under round-to-nearest, and prove that NVFP4’s small 16-value groups already neutralize the outlier mitigation that rotations provide (Bridging the Gap Between Promise and Performance for Microscaling FP4 Quantization, ICLR 2026, arXiv:2509.23202).
Our simulation reproduces the sign of every effect and gives the sizes (Figure 7). A 16-point Hadamard rotation leaves Gaussian data unchanged within 0.03 dB. On activations it costs NVFP4 0.88 dB, gives MXFP4 2.13 dB and gives NVINT4 4.07 dB.
With eight outlier channels at twenty times the scale of the others the changes are -1.03, +2.27 and +3.21 dB. Larger rotations hurt NVFP4 less: a 128-point rotation costs it 0.51 dB on activations.
On RoBERTa’s weights, whose kurtosis is 6 rather than 22, a 16-point rotation changes NVFP4 by -0.09 dB, MXFP4 by +0.16 dB and NVINT4 by +0.42 dB: the same signs, much smaller sizes.

The best 4-bit recipe in our sweep is a rotation followed by NVINT4, integers under an E4M3 scale per 16 values: 22.63 dB on activations, 1.74 dB above plain NVFP4 and 1.36 dB above IF4.
Cook and colleagues found the same ordering in real W4A4 runs on Qwen3.5, where NVINT4 after a Hadamard transform beat NVFP4. The catch comes in the section on the assembler: no tensor-core instruction on a shipping NVIDIA datacenter part multiplies INT4 values under E4M3 block scales.
The block-scaled FP4 instructions accept only E2M1 elements, and the only integer tensor-core instruction left on datacenter Blackwell takes INT8 and is being withdrawn.
From decibels to perplexity
Signal-to-noise ratios matter only if they predict something people care about. Cook and colleagues report the average WikiText-2 perplexity of four Qwen3.5 models (9B, 27B, 35B-A3B and 122B-A10B) for 23 block formats, with the weights and activations of every linear layer quantized by round-to-nearest (arXiv:2603.28765, Table 6).
We simulated 22 of them, all but a block-8 variant of Four Over Six. We computed the noise-to-signal power ratio (NSR) of each format on our activation proxy and fitted the simplest law that could work: the logarithm of the perplexity ratio is proportional to the noise power.
With one free parameter, ln(ppl / pplBF16) = 8.25 x NSR explains 98.7% of the variance across the 21 formats that stay above 12 dB, from 3.5 to 6.5 bits per element, integer and floating point, MX and NV (Figure 8). The rank correlation is 0.986.
Dropping any one format moves the fitted constant between 7.95 and 8.33, and the median error of the fit is 0.6% of perplexity. The same fit on Gaussian noise explains 95.4%, so the heavy-tailed proxy predicts better, as the comparison with AMD suggested.
W4A4 adds weight noise to activation noise, so we also charged each format twice, once on Gaussian data for its weights and once on the proxy for its activations. That doesn’t do better (98.3% with one constant, 98.8% with a constant per term), and the two-term fit puts most of the weight on the activation term.
The two terms rise and fall together across formats, so that split is a hint rather than a measurement. One format breaks the law: MXFP3, at a perplexity of 70. Below roughly 12 dB the models stop degrading gracefully and start to fail.

Read backwards, the fit gives targets. Staying within 1% of BF16 perplexity needs about 29.2 dB on the proxy, within 2% about 26.2 dB, within 5% about 22.3 dB. NVFP4 sits at 20.9 dB, predicted +6.9% against a measured +6.4%; MXFP4 sits at 16.7 dB, predicted +19.3% against a measured +16.0%.
The measured gap between them is 9.1% of perplexity. That is the size of one byte of scale per block when nothing else is done to the model. Quantization-aware training, rotations and better rounding all shrink it;
MR-GPTQ, the method Egiazarian and colleagues propose, brings MXFP4 to within 1 to 2% of NVFP4 accuracy.
The law travels, but its constant does not. AMD’s appendix gives WikiText-2 perplexities for the two models it profiled, quantized W4A4 by round-to-nearest.
For Llama-3.1-8B the law predicts +19.3% for MXFP4 against a measured +20.0% (8.47 against 7.06), and +0.6% for per-token FP8 against a measured +0.7%.
For Qwen3.6-27B it predicts the same +19.3% for MXFP4 against a measured +6.1% (7.53 against 7.10): the larger and newer model tolerates about three times the noise. The constant is a property of the model, not of the format, and 8.25 is an average over a family that runs from 9 to 122 billion parameters.
What the assembler accepts
The numerics are half the story; the other half is which instructions exist. ptxas, nvdisasm and cuobjdump ship as ordinary x86 programs inside NVIDIA’s PyPI wheels, so the question can be asked on any laptop: wrap one instruction in a minimal PTX kernel, assemble it for a target and read the SASS that comes out.
We did this for the FP4-related instructions of PTX ISA 9.4 on eight targets with CUDA 13.4.92 (Figure 9). The targets are sm_90a (H100), sm_100a (B200), sm_100f (the portable Blackwell family), sm_103a (B300), sm_107a (Rubin, compute capability 10.7 according to NVIDIA’s TensorRT 11.3 release notes), sm_110a (Jetson Thor), sm_120a (RTX 50 and RTX PRO) and sm_121a (GB10).

Hopper has no FP4 at all. ptxas rejects conversion to FP4 on sm_90a as a feature the target does not support, rejects the conversion back, rejects conversion to E8M0 scales and has no tcgen05 instructions.
FP4 weights on an H100 must be decoded into BF16 or FP8 with integer tricks or table lookups before they reach a tensor core, and at that point FP4 offers nothing that INT4 does not.
That is the practical background to the Kimi K2 Thinking team choosing INT4 weights with 16-bit activations in November 2025: weight-only INT4 kernels such as Marlin were mature on Hopper (LMSYS blog, January 2026).
On datacenter Blackwell, FP4 reaches its own datapath only through the block-scaled FP4 kinds. tcgen05.mma with .kind::mxf4 or .kind::mxf4nvf4 assembles to UTCOMMA.
The same E2M1 operands sent through .kind::f8f6f4, the kind that also accepts FP8 and FP6, assemble to UTCQMMA, the opcode FP8 uses. The names fit H for half, Q for quarter and O for one eighth of 32 bits, and they fit NVIDIA’s documented rates: FP4 through the mixed kind runs at FP8 speed.
Any model that multiplies FP4 weights by FP8 activations is in this position by construction, because one of its operands is eight bits wide.
NVFP4 with 16-value blocks assembles to UTCOMMA.4X on sm_100a, sm_100f and sm_110a and to UTCOMMA.BLOCK16 on sm_103a and sm_107a. The older spelling of the qualifier, .scale_vec::4X, is rejected for sm_100f, the family target meant to run on both B200 and B300, so portable builds must use .block16.
INT8 is on its way out. tcgen05.mma.kind::i8 assembles to UTCIMMA on sm_100a and on Thor’s sm_110a, and is rejected on sm_100f, sm_103a and sm_107a.
In August 2026 Teng-Ruei Chen traced the same withdrawal for B300 through the spec sheet, the PTX ISA, CUTLASS and the two main open serving engines (Spec Sheets Are Not Kernels, arXiv:2608.11693). The assembler adds two facts: Rubin continues it, and NVIDIA left INT8 out of the portable Blackwell family entirely.
The spec sheets agree. GB300 lists about 0.165 POPS of dense INT8 against 5 PFLOPS of dense FP8, a ratio of 30; a Rubin GPU lists 0.25 POPS against 17.5 PFLOPS, a ratio of 70. On B200 the two were equal.
Consumer Blackwell is a different machine. sm_120a and sm_121a have no tcgen05, but they have warp-level FP4: mma.sync with shape m16n8k64 assembles to OMMA.SF.16864, with UE4M3 scales for NVFP4 through .kind::mxf4nvf4 or E8 scales for MXFP4 through .kind::mxf4.
One more result surprised us: on every datacenter target, Hopper included, the warp-level FP8 instruction mma.sync.m16n8k32 with E4M3 operands assembles to an F2FP unpack to FP16 followed by HMMA.16816, a half-precision multiply.
Only the consumer parts run it natively, as QMMA.16832. On a B200, FP8 tensor throughput is reachable only through tcgen05.
Rubin accepts four instructions that no other target does: conversion to the new UE5M3 scale format (F2FP.SATFINITE.UE5M3), an FP4 conversion that applies a power-of-two scale on the way (F2FP.SATFINITE.E2M1 with a .SCALE_BY_C modifier), FP4 conversion with round-toward-zero, and a matrix multiply that decompresses its weights from a lookup table (UTCQMMA.LUTB). We come back to all four.
K
Spec sheets quote FP4 and FP8 throughput as separate numbers, and the instruction descriptors that CUTLASS builds for each architecture show where the ratio comes from (CUTLASS main branch, October 2026). A dense block-scaled FP4 multiply reduces K = 64 elements per instruction on B200, 96 on B300 and 128 on Rubin.
A dense FP8 multiply reduces 32, 32 and 64. The ratios, 2, 3 and 2, are exactly the ratios of dense FP4 to dense FP8 on the spec sheets: 10 and 5 PFLOPS on GB200, 15 and 5 on GB300, 35 and 17.5 on Rubin (Figure 10).

B300’s extra FP4 throughput is one bit. In the 32-bit Blackwell instruction descriptor for block-scaled FP4, bit 31 selects the K size. A 0 means dense K = 64 or sparse K = 128; a 1 means dense K = 96, with the sparse version marked invalid.
That single comment in mma_sm100_desc.hpp explains why GB300’s dense FP4 rose by half over GB200, from 10 to 15 PFLOPS, while its sparse FP4 stayed at 20: the fast dense mode has no sparse form, so sparse code still runs the K = 128 path both chips share.
Rubin widens the field to two bits, split between bits 3 and 31: dense K of 64, 96 or 128, with sparse forms at 128 and 192 and none for the K = 128 mode. Rubin’s densest sparse mode is therefore 1.5 times its densest dense mode.
NVIDIA’s Rubin architecture overview describes the 50 PFLOPS headline as sparse NVFP4 (as reported by Guru3D, July 2026), and 1.5 times the dense 35 is 52.5, so the descriptor and the spec sheet agree to within 5%. Rubin’s FP4 multiply also requires both operands to be K-major; there is no transposed form.
Making FP4 on the fly
Weights are quantized once, offline. Activations are quantized at every step by the kernel that produces them, and that costs instructions. We wrote the quantization of one block in PTX the way a fused epilogue would: absolute maximum, scale, scale encoding and decoding, multiply, convert, pack.
We assembled it with ptxas -O3 for four targets and counted SASS instructions, excluding loads, stores, address arithmetic and control flow (Figure 11).

NVFP4 costs 2.38 instructions per element on B200, B300 and Rubin alike. For 16 values that is seven three-input FMNMX3 and one FMNMX for the maximum, 18 FMUL, 10 F2FP conversions, one MUFU reciprocal for the decoded scale and one half-precision add. MXFP4 costs 2.16 per element, because its scale is a power of two and its reciprocal takes three integer instructions.
The two scale-rounding rules compile to the same count. Consumer Blackwell pays 2.81 and 2.59, because it has no three-input FMNMX3 and needs 15 comparisons for 16 values.
Rubin’s fused conversion removes the multiplies for MX formats entirely: 1.19 instructions per element. NVFP4 gets no fused path, because a scale that is not a power of two cannot be applied by adjusting an exponent.
Whether 2.4 instructions per element matter depends on how much tensor-core work each quantized value feeds. A quantized activation is reused for every one of the N output columns of the GEMM that consumes it.
If the quantization cannot overlap with other work, its share of the GEMM’s time is about I x F / (256 x S x f x N), where I is instructions per element, F the dense FP4 rate, S the number of SMs and f the clock, since each SM issues 128 lane-instructions per cycle and each multiply-add counts as two FLOPs.
At an assumed 1.9 GHz that is about 330/N for NVFP4 on B200, 460/N on GB300 and 760/N on Rubin: 4.6%, 6.4% and 10.6% at N = 7168, the hidden size of DeepSeek-V3-class models. Rubin’s fused conversion halves its own figure for MX formats.
Tensor cores have grown faster than the general-purpose ALUs around them, and quantization is one of the places where the difference shows. When AMD profiled MXFP4 and MXFP6 serving on MI355X, the standalone activation-quantization step was measurable overhead, and the team wrote a dedicated kernel to remove it (AMD ROCm blog, June 2026).
What Rubin adds
The first addition is a scale format. UE5M3 has no sign bit, five exponent bits with a bias of 15 and three mantissa bits; CUTLASS defines its range as 0 to 114,688, with subnormals and a NaN code (include/cutlass/float8.h). It keeps E4M3’s three mantissa bits, so inside a block it behaves like NVFP4’s scale, but it covers about 34 binades against 18.
In Blackwell’s block-scaled instruction descriptor the scale format is one bit, E4M3 or E8M0; in Rubin’s it is two bits, and the value 2 means UE5M3 (mma_sm107_desc.hpp).
Conversions to UE5M3 assemble only for sm_107a, and we assembled both round-to-nearest and round-up forms.
The extra range matters because of NVFP4’s second level. We quantized Gaussian tensors in blocks of 16 FP4 values with no tensor scale at all, sweeping the standard deviation from 10-8 to 104 (Figure 12).
With UE4M3 block scales the result stays within 0.5 dB of the best only for standard deviations between 0.032 and 562. Weights of large models sit near one over the square root of the hidden size, about 0.011 at a width of 8,192, and a tensor at 0.01 loses 2.4 dB with UE4M3 scales and no tensor scale.
RoBERTa-base, with a width of 768 and a standard deviation of 0.053, happens to sit inside the window. That is why NVFP4 carries its FP32 tensor scale. With UE5M3 the window runs from 10-4 to at least 104, the end of our sweep.
NVIDIA has not said that UE5M3 exists to retire the second scale, but on these numbers it could, and kernels would save a global reduction and a multiply per tensor. The IF4 authors proposed spending the unused sign bit of the E4M3 scale on a choice between FP4 and INT4 grids; UE5M3 spends that bit on range instead.

The second addition is a weight format with no fixed grid. PTX ISA 9.4 adds a .decompress::lut::b qualifier to tcgen05.mma.
In this mode, according to SemiAnalysis’s description, each weight is a 3-bit index into a table of eight E4M3 values, one table per 8 x 64 tile of the weight matrix, so storage is 3 + 64/512 = 3.125 bits per weight; the table lives in tensor memory and the lookup happens inside the multiply.
In the assembler the instruction exists only for sm_107a and becomes UTCQMMA.LUTB, so the multiply itself runs on the FP8 path. A block-scaled variant adds E8M0 scales for every 32 weights.
We fitted eight-entry codebooks by k-means on each group of 512 values, rounded the entries to E4M3 and measured. On Gaussian data the 3.125-bit format reaches 14.76 dB, above NVFP3 at 3.5 bits (14.31 dB) and slightly above an ideal 3-bit quantizer that knows the variance (14.62 dB).
On RoBERTa’s weights, with codebooks shared over 8 x 64 tiles as in the hardware, the plain format reaches 13.87 dB, 0.34 dB below NVFP3 at 3.5 bits (14.21 dB): real weights change scale from row to row, and one table per tile follows that less well than a scale per 16 values.
The block-scaled variant, at 3.375 bits, recovers the loss and reaches 14.23 dB, just above NVFP3 with an eighth of a bit less. Against NVFP4 the plain format gives up 5.7 dB on Gaussian data and 6.6 dB on RoBERTa for 1.375 fewer bits. At 6.02 dB per bit a fair exchange would cost 8.3 dB, so per stored bit the lookup table is still the more efficient weight format.
It would be the wrong tool for activations, where one table cannot follow the local scale (11.73 dB on our proxy, below NVFP3’s 13.86), and the hardware offers it only for the B operand anyway.
The third addition is an operand type. CUTLASS lists an unsigned 8-bit E5M3 type among the A and B operand types of Rubin’s FP8 multiply (the CuTe DSL tcgen05 tables). An unsigned operand only makes sense for values that are never negative.
In a transformer the obvious candidate is the softmax output that multiplies V in attention, which E4M3 represents poorly below 2-9. That reading is ours; the code says only that the type exists.
SemiAnalysis also reports that Rubin’s tensor memory grows from 512 to 576 columns, with the extra columns meant for block scale factors, which on Blackwell compete with accumulators for space.
Who buys which half
FP4 sells two different things. Fewer bytes per weight make decoding faster and models smaller, and that half works whatever the precision of the activations.
Faster arithmetic needs both operands in FP4 and the block-scaled FP4 instructions, and that half needs W4A4. The open models of the past fourteen months bought mostly the first half (Table 1).
FP4 sells two things: fewer bytes and faster arithmetic. Only W4A4 collects both.

OpenAI’s gpt-oss quantizes its MoE weights, more than 90% of the parameters, to MXFP4 at 4.25 bits per parameter (model card, arXiv:2508.10925).
Moonshot moved from INT4 in K2 Thinking to MXFP4 weights with MXFP8 activations in K3, applied with quantization-aware training from the supervised fine-tuning stage onward and a group size of 32 (Kimi K3 model card, July 2026).
On Blackwell that combination runs through .kind::mxf8f6f4 on the FP8 path. On an MI325X, which has no FP4 matrix instructions, vLLM converts K3’s MXFP4 experts to group-32 INT4 at load time and serves them with BF16-by-INT4 kernels (Shakudo, September 2026).
DeepSeek V4 is the most careful case. Its report applies MXFP4 quantization-aware training to the MoE experts and to the query-key path of the attention indexer, then dequantizes the expert weights to FP8 for computation, and notes that this is lossless as long as the ratio between the largest and smallest FP4 sub-block scale inside each 128 x 128 FP8 block stays under a threshold (arXiv:2606.19348). The report does not state the threshold, so we enumerated it.
Every E2M1 value times 2d is exactly representable in E4M3 if and only if d lies between -8 and 6, so the sub-block exponents inside a tile may span 14 binades, a ratio of 16,384, or 11 binades if FP8 subnormals are to be avoided.
Gaussian weights with log-normal row scales span 4.6 binades on average and 6 at most in our simulation, so the condition is loose and V4’s empirical check unsurprising.
NVIDIA’s own NVFP4 checkpoints are the exception. They quantize activations too, by post-training quantization, and NVIDIA reports DeepSeek-R1-0528 within one point of FP8 on six of seven benchmarks, with AIME 2024 two points higher.
AMD’s measurements show what the second half is worth when it is bought without help. On one MI355X, W4A4 MXFP4 delivered 5.1% more total tokens per second than FP8 on Llama-3.1-8B and 10.0% more on Qwen3.6-27B, from twice the peak arithmetic.
If only GEMM time halved, those speedups imply that GEMMs were 9.6% and 18.2% of the FP8 step, which says more about the rest of the step than about FP4.
Accuracy moved the other way: GSM8K on Llama-3.1-8B fell from 80.44 with FP8 to 62.55 with MXFP4. MXFP6 activations with MXFP4 weights recovered 76.42 at 2.8% less throughput than MXFP4, because MI355X runs FP6 at the FP4 rate.
NVIDIA runs FP6 at the FP8 rate, so that particular trade does not exist on Blackwell.
The critical batch
A decode step reads every active weight once and performs two FLOPs per weight for each sequence in the batch. Below a critical batch B* = P x b / (2 x BW), where P is the dense FLOP rate, b the bytes per weight and BW the memory bandwidth, the step is limited by bandwidth; above it, by arithmetic.
Earlier pieces in this publication found B* nearly independent of precision, because P doubles when b halves. Block scales and Blackwell Ultra both break that (Figure 13).
Scales first. NVFP4 weights cost 0.5625 bytes, not 0.5, so on GB200 B* rises from 312.5 with FP8 to 351.6 with W4A4 NVFP4, 12.5% higher purely from scale bytes; MXFP4’s smaller overhead gives 332.0.
Then GB300: dense FP4 is three times dense FP8 there, so W4A4 NVFP4 pushes B* to 527.3, 1.69 times the FP8 value. The FP4-weight, FP8-arithmetic schemes go the other way.
Halving the bytes without speeding up the arithmetic roughly halves B*, to 166 on GB200 and GB300 and 211 on Rubin, and Rubin’s lookup-table weights take it to 155. Those formats make small batches much cheaper and become compute-bound at about half the batch. W4A4 keeps large batches cheap, and on GB300 it is worth three times FP8 arithmetic.
At batch 1, weight traffic alone caps decoding of a 70B dense model at 57 tokens per second in BF16 on 8 TB/s of HBM, 114 in FP8, 203 in NVFP4 and 215 in MXFP4; on Rubin’s 22 TB/s, at 559 in NVFP4 and 805 with lookup-table weights. Attention reads and synchronization come on top of these ceilings.
What to use where
The measurements reduce to a handful of situations (Table 2). The table is our reading of the evidence in this piece, not a benchmark of anyone’s serving stack, and it applies to post-training quantization: models trained with quantization in the loop have already chosen their grid.
Where this could be wrong
Synthetic data. Every SQNR here except RoBERTa’s comes from synthetic distributions. The activation proxy matches AMD’s three numbers, but those are layer means over one tensor type in two models; attention inputs, residual streams and gradients have other shapes.
One family, one recipe. The perplexity law is fitted on Qwen3.5 with round-to-nearest W4A4, where weight and activation noise mix and our single proxy cannot separate them. Models trained with quantization in the loop, which includes gpt-oss, DeepSeek V4 and Kimi K3, learn around their grid, and the law should not be applied to them. The constant also moves between models: AMD’s numbers put it near 8.5 for Llama-3.1-8B and near 2.8 for Qwen3.6-27B. The rank correlation of 0.986 is the more robust number.
Rates from names. We infer that FP4 through .kind::f8f6f4 runs at FP8 speed from the opcode it becomes and from NVIDIA’s documentation, not from a timed kernel. The K values come from CUTLASS source, not from hardware, and a descriptor comment is a strong hint rather than a measurement.
Our codebooks. The lookup-table numbers depend on plain k-means with E4M3 entries. NVIDIA’s own fitting, or quantization-aware training, would do better, so our LUT figures are a floor.
Assumed clocks. The quantization-overhead estimate assumes 1.9 GHz, NVIDIA’s listed SM counts (148 for B200, 160 for B300, 224 for Rubin) and no overlap with other work. Real epilogues hide part of the cost.
Intent. UE5M3 without a tensor scale is our reading of what the format allows, not NVIDIA’s stated purpose, and the softmax use of the unsigned E5M3 operand is a guess from the type’s existence.
Moving spec sheets. Rubin’s numbers come from announcements and partner documents months before broad availability, and NVIDIA’s own figures for B300 range from 13 to 15 PFLOPS of dense FP4 depending on the system.
Predictions
By 30 June 2027, at least one open-weight model above 100 billion parameters will ship its primary checkpoint in a format only Rubin runs natively: FP4 with UE5M3 scales, or 3-bit lookup-table weights.
Through the first PTX ISA release that adds Feynman (sm_140), tcgen05.mma.kind::i8 will not assemble for sm_103a or sm_107a.
By 31 March 2027, vLLM or SGLang will merge a Rubin MoE kernel that reads expert weights through .decompress::lut::b.
Moonshot’s next flagship after K3 will again pair 4-bit weights with 8-bit activations rather than move to W4A4.
By 31 December 2027, a revision of the OCP Microscaling specification will add a non-power-of-two scale type or change the default rounding of E8M0 scales.
Confidence dossier
Every number in the text carries a tier. A: primary source (vendor documentation, archived paper, model card). B: vendor source code or reputable secondary reporting. C: our derivation from A or B sources, with stated assumptions. D: our inference. M: our own measurement.
Sources
Open Compute Project, OCP Microscaling Formats (MX) Specification v1.0, September 2023.
B. Darvish Rouhani et al., Microscaling Data Formats for Deep Learning, arXiv:2310.10537.
E. Alvarez et al., NVIDIA, Introducing NVFP4 for Efficient and Accurate Low-Precision Inference, 24 June 2025.
NVIDIA, Pretraining Large Language Models with NVFP4, arXiv:2509.25149, v2, 4 March 2026.
A. Mishra, D. Stosic, S. Layton, P. Micikevicius, Recipes for Pre-training LLMs with MXFP8, arXiv:2506.08027, v2, 18 August 2025.
V. Egiazarian et al., Bridging the Gap Between Promise and Performance for Microscaling FP4 Quantization, ICLR 2026, arXiv:2509.23202.
J. Cook et al., Four Over Six: More Accurate NVFP4 Quantization with Adaptive Block Scaling, arXiv:2512.02010, v5, 9 May 2026.
J. Cook et al., Adaptive Block-Scaled Data Types, arXiv:2603.28765, 30 March 2026.
S. Atre, B. Bao, S. Tiwari, A. Sirasao, AMD, MXFP6 and MXFP4 Mixed Precision for Accelerating Dense LLMs on AMD Instinct MI355X, 26 June 2026.
T.-R. Chen, Spec Sheets Are Not Kernels: An ISA- and Source-Level Audit of INT8 Availability on NVIDIA Blackwell Ultra, arXiv:2608.11693, August 2026.
NVIDIA, Inside NVIDIA Blackwell Ultra: The Chip Powering the AI Factory Era, NVIDIA technical blog.
NVIDIA, HGX platform specifications (HGX B300 and HGX B200).
Verda, GB300 NVL72 architecture (per-GPU dense FP4, FP8 and INT8 for GB300, GB200, B300, B200).
The Register, Nvidia unpacks its Vera Rubin CPUs and GPUs at CES, 5 January 2026.
NVIDIA, 2026 1H product guide (Vera Rubin NVL72 dense specifications), distributed by Gigabyte.
Guru3D, NVIDIA details Rubin GPU architecture with HBM4, NVLink 6 and 336 billion transistors, July 2026.
SemiAnalysis, Vera Rubin NVL72 vs GB200 NVL72? Inference TCO and Architecture Analysis, 23 July 2026.
NVIDIA, PTX ISA 9.4, CUDA 13.4.
NVIDIA, TensorRT 11.3.0 release notes (Vera Rubin, compute capability 10.7).
NVIDIA, CUTLASS, main branch, October 2026: include/cute/arch/mma_sm100_desc.hpp, mma_sm107_desc.hpp, mma_sm107_umma.hpp, include/cutlass/float8.h, python/CuTeDSL/cutlass/cute/nvgpu/tcgen05/mma.py.
Apache TVM, PTX instruction table and pull request 20271, September 2026.
OpenAI, gpt-oss-120b and gpt-oss-20b Model Card, arXiv:2508.10925.
DeepSeek-AI, DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence, arXiv:2606.19348.
vLLM, DeepSeek V4 model documentation (MXFP4 experts with UE8M0 scales).
Moonshot AI, Kimi K3 model card, as reported by SemiAnalysis InferenceX and the VESSL W4AFP8 card.
LMSYS, INT4 quantization-aware training, 26 January 2026.
Shakudo, How We Deployed Kimi K3 on AMD GPUs, updated 4 September 2026.
heise online, AMD Instinct MI350X and MI355X (FP6 and FP4 at twice the FP8 rate).
S. Lloyd, Least Squares Quantization in PCM, IEEE Transactions on Information Theory, 1982; J. Max, Quantizing for Minimum Distortion, IRE Transactions on Information Theory, 1960.
Two-level MXFP8 training with an FP32 tensor scale, arXiv:2511.05811, November 2025.
Y. Liu et al., RoBERTa, arXiv:1907.11692; Explosion, en_core_web_trf 3.8.0.
S. Ashkboos et al., QuaRot, arXiv:2404.00456; Z. Liu et al., SpinQuant, arXiv:2405.16406.
Corrections
Instruction count. An early count of the NVFP4 quantization cost, 2.31 instructions per element, left out a half-precision add and the byte permutes. The counting rule used here, every instruction except memory access, address arithmetic and control flow, gives 2.38.
Codebook tiles. Our first pass over real weights shared lookup-table codebooks across 512 consecutive values in memory rather than across 8 x 64 tiles of the matrix, as the hardware does. The figures here use tiles.
Weight scale window. A draft said typical weights fall outside the no-tensor-scale window of UE4M3 while quoting RoBERTa’s standard deviation of 0.053, which is inside it. The text now uses the scale of large-model weights, about 0.01, where the loss is 2.4 dB.
Format count. A draft said the perplexity table holds 22 formats and that the fit starts at 3.25 bits. The table holds 23, we simulated 22, and the fit covers 3.5 to 6.5 bits.
Error message. A draft quoted an assembler error message from memory. It was removed; the text now describes the rejection without quoting it.
AMD quantization kernel. The previous version removed a statement that AMD wrote a dedicated activation-quantization kernel, because we could not re-verify it at the time. AMD’s post says exactly that, and the statement is back with its source.
INT8 figures. The previous version cited a secondary page for B300’s INT8 throughput that, on re-reading, lists INT8 at the FP8 rate. The figures now come from Chen’s audit and from per-GPU tables that match NVIDIA’s datasheet.
Kimi K2 Thinking. The previous version gave the scale type of Kimi K2 Thinking’s INT4 weights as BF16. Our sources confirm the grouping of 32 and the 16-bit activations, not the scale type, and the text now says so.
Appendix: reproducing the measurements
The assembler and disassembler come from NVIDIA’s PyPI wheels and run on any x86 Linux machine without a GPU.
pip download --no-deps nvidia-cuda-nvcc nvidia-cuda-nvdisasm # 13.4.92
python3 -m zipfile -e nvidia_cuda_nvcc-13.4.92-*.whl cuda/
python3 -m zipfile -e nvidia_cuda_nvdisasm-13.4.92-*.whl cuda/
chmod +x cuda/nvidia/cu13/bin/*
cuda/nvidia/cu13/bin/ptxas -arch=sm_107a -o k.cubin k.ptx
cuda/nvidia/cu13/bin/nvdisasm -c k.cubin | grep -E "UTC|F2FP|OMMA|QMMA"
A probe is one instruction inside a minimal kernel. This one asks for an NVFP4 block-scaled multiply; changing the target line and the qualifiers gives every cell of Figure 9.
.version 9.4
.target sm_107a
.address_size 64
.visible .entry k(.param .u64 out, .param .u64 in, .param .u32 idesc)
{
.reg .b32 %r<8>; .reg .b64 %rd<8>; .reg .pred %p<2>;
ld.param.u64 %rd1, [out]; ld.param.u64 %rd2, [in]; ld.param.u32 %r1, [idesc];
ld.global.u32 %r2, [%rd2]; ld.global.u32 %r3, [%rd2+4];
ld.global.u64 %rd3, [%rd2+8]; ld.global.u64 %rd4, [%rd2+16];
setp.ne.u32 %p1, %r1, 0;
tcgen05.mma.cta_group::1.kind::mxf4nvf4.block_scale.block16
[%r2], %rd3, %rd4, %r1, [%r3], [%r3], %p1;
ret;
}
The reference semantics of NVFP4 fit in a few lines of NumPy. Every other format in this piece changes the element grid, the block size or the way the scale is rounded.
import numpy as np
E2M1 = np.array([0, .5, 1, 1.5, 2, 3, 4, 6])
E4M3 = np.array(sorted({m / 8 * 2.0**-6 for m in range(8)} |
{(1 + m / 8) * 2.0**(e - 7) for e in range(1, 16) for m in range(8)}))[:-1]
def rne(a, grid): # nearest grid point, ties to even code, saturating
i = np.clip(np.searchsorted(grid, a), 1, len(grid) - 1)
lo, hi = grid[i - 1], grid[i]
up = (hi - a < a - lo) | ((hi - a == a - lo) & (i % 2 == 0))
return np.where(a >= grid[-1], grid[-1], np.where(up, hi, lo))
def nvfp4(x, block=16):
g = np.abs(x).max() / (6 * 448) # FP32 tensor scale
xb = x.reshape(-1, block)
s = rne(np.abs(xb).max(1, keepdims=True) / (6 * g), E4M3) * g # E4M3 block scale
s1 = np.where(s > 0, s, 1.0)
return (np.sign(xb) * rne(np.abs(xb) / s1, E2M1) * s).reshape(x.shape)







