<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[The Software Frontier]]></title><description><![CDATA[Where abstraction ends. Essays on GPU execution, kernel internals, and distributed systems at scale.]]></description><link>https://www.thesoftwarefrontier.com</link><image><url>https://substackcdn.com/image/fetch/$s_!SAY7!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F54550d86-2756-4131-8818-956604f6749d_608x608.png</url><title>The Software Frontier</title><link>https://www.thesoftwarefrontier.com</link></image><generator>Substack</generator><lastBuildDate>Tue, 06 Oct 2026 16:04:53 GMT</lastBuildDate><atom:link href="https://www.thesoftwarefrontier.com/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Lorenzo Bradanini]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[softwarefrontier@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[softwarefrontier@substack.com]]></itunes:email><itunes:name><![CDATA[Lorenzo Bradanini]]></itunes:name></itunes:owner><itunes:author><![CDATA[Lorenzo Bradanini]]></itunes:author><googleplay:owner><![CDATA[softwarefrontier@substack.com]]></googleplay:owner><googleplay:email><![CDATA[softwarefrontier@substack.com]]></googleplay:email><googleplay:author><![CDATA[Lorenzo Bradanini]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[Inside Google’s TPU: How it works, and what it costs]]></title><description><![CDATA[We ran it against five TPU generations, from v4 to Ironwood, then followed the money: a rate card that prices bandwidth, an anchor tenant paying 13% of list, and a ten-gigawatt roadmap.]]></description><link>https://www.thesoftwarefrontier.com/p/inside-googles-tpu-how-it-works-and</link><guid isPermaLink="false">https://www.thesoftwarefrontier.com/p/inside-googles-tpu-how-it-works-and</guid><dc:creator><![CDATA[Lorenzo Bradanini]]></dc:creator><pubDate>Tue, 06 Oct 2026 11:01:40 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!aPVU!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7afb77a7-500b-4766-b6df-716a3696f54f_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!aPVU!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7afb77a7-500b-4766-b6df-716a3696f54f_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!aPVU!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7afb77a7-500b-4766-b6df-716a3696f54f_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!aPVU!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7afb77a7-500b-4766-b6df-716a3696f54f_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!aPVU!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7afb77a7-500b-4766-b6df-716a3696f54f_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!aPVU!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7afb77a7-500b-4766-b6df-716a3696f54f_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!aPVU!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7afb77a7-500b-4766-b6df-716a3696f54f_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/7afb77a7-500b-4766-b6df-716a3696f54f_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:4047069,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/219068497?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7afb77a7-500b-4766-b6df-716a3696f54f_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!aPVU!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7afb77a7-500b-4766-b6df-716a3696f54f_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!aPVU!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7afb77a7-500b-4766-b6df-716a3696f54f_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!aPVU!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7afb77a7-500b-4766-b6df-716a3696f54f_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!aPVU!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7afb77a7-500b-4766-b6df-716a3696f54f_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>Introduction </h2><p>Full disclosure before we start: I&#8217;ve<strong> never run a program on a TPU.</strong> For most of the last decade that was the normal condition of anyone who wrote about accelerators from outside Google. </p><p>The chips lived in <strong>Google&#8217;s data centers</strong>, the papers described them a generation late, and the cloud instances were something you read about more often than something you rented.</p><p>2026 is the year that stopped being true. Anthropic is <strong>bringing up to a million TPUs online</strong>, and Broadcom has booked $21 billion of Ironwood racks for it, sold as complete systems rather than chips. </p><p>Meta signed a multi-year TPU rental in February and is negotiating to buy them for its own buildings next year. <strong>MediaTek&#8217;s TPU business</strong> is limited by how much wafer capacity it can get, not by demand. Last week Google even put four of them in orbit on a prototype satellite, running Gemma for fifteen minutes at a time. The TPU has real <em>tenants</em> now.</p><p>The most interesting part of the TPU, though, is not a chip. <strong>It is a 683 MB file.</strong> The runtime JAX uses to talk to TPUs ships on PyPI as a wheel called libtpu, and it contains the entire TPU compiler. <strong>That compiler will happily compile for a TPU that isn&#8217;t there.</strong> </p><p><em>You describe a topology</em>, say four Ironwood chips in a 2x2x1 slice, and it hands back an executable, a cost estimate, a memory report and, if you ask with the right options, the <strong>VLIW</strong> bundles it would have sent to the chip. </p><p>The feature exists so that people can check whether a model fits on a slice before they pay for one. We used it the way we used ptxas in our pieces on <em>NVIDIA&#8217;s compiler and on Triton</em>: as a microscope pointed at a vendor that has never published an instruction set.</p><p>We compiled the same programs for v4, v5e, v5p, Trillium (v6e) and Ironwood (TPU7x) on a container with no accelerator at all, and read back everything the compiler was willing to say. Then we checked it against the best public description of the hardware, a five-generation retrospective by Google&#8217;s own architects (<strong>Jouppi, Lakshmanamurthy, Young and Patterson</strong>) posted in June and due in IEEE Micro. </p><p>Where the two overlap they mostly agree. The three places they don&#8217;t are <strong>Ironwood&#8217;s HBM bandwidth</strong>, the count of its matrix units, and the compute peaks of v5p and Ironwood.</p><p>A word on what kind of evidence this is. When the compiler tells us which instruction it emitted, where an array lives or which core runs a collective, that is a fact about the program it produced, and <strong>we report it as a measurement</strong>. When it tells us how many cycles something will take, that is the compiler&#8217;s model of the chip, and we label it that way every time. </p><p>The model matters on its own terms, because it is what the compiler optimizes against, but it is <em>not</em> a stopwatch. The public libtpu we used, version 0.0.48, does not know TPU 8t or TPU 8i. For those we have Google&#8217;s April deep dive and arithmetic.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>The very short version</h2><ul><li><p><strong>The compiler is on PyPI.</strong> Google&#8217;s TPU compiler ships inside libtpu, a 683 MB library, and compiles for TPUs that aren&#8217;t attached. We used it on a CPU-only container for v4, v5e, v5p, Trillium (v6e) and Ironwood (TPU7x).</p></li><li><p><strong>The critical batch.</strong> Trillium&#8217;s is 560, the highest of any shipping accelerator we have measured. Ironwood&#8217;s is 313 on Google&#8217;s datasheet and 270 by the compiler&#8217;s own constants, in the range of NVIDIA&#8217;s H100 (295) and B200 (281).</p></li><li><p><strong>The accumulator.</strong> In the final instruction stream for the same 1024&#179; bf16 matmul, Ironwood pops 1,024 results out of its matrix units and Trillium 4,096. Ironwood does no vector adds where Trillium does 3,072, and issues 80% fewer spill loads. Ironwood&#8217;s matrix instructions address a result buffer, <code>mrb[0]</code> to <code>mrb[255]</code>; Trillium&#8217;s results come out of an unaddressed queue.</p></li><li><p><strong>Precision.</strong> By the compiler&#8217;s cycle estimates, Ironwood runs an FP8 matmul in 0.48 times the cycles of bf16 and an int8 matmul in 1.13 times. On v5e, v5p and Trillium, int8 takes 0.47 to 0.62 times and FP8 e4m3 gains little. No chip the public compiler knows runs FP4 faster than its eight-bit path.</p></li><li><p><strong>Batch one.</strong> On v4 through Trillium, the compiler rewrites a batch-1 matmul into a vector multiply and reduction that its own cycle model prices at 2.75 to 3.50 times the matrix-unit version. One option turns the rewrite off. Ironwood doesn&#8217;t apply it.</p></li><li><p><strong>SparseCores.</strong> On Ironwood, every collective we compiled across four or more chips ran on the SparseCores, including the ragged all-to-all that MoE dispatch uses. The one exception was the dense all-to-all. No other generation offloaded anything. TPU 8i replaces those SparseCores with an engine built only for collectives.</p></li><li><p><strong>Storage.</strong> Arrays live in 4 KiB tiles, 128 elements wide. A single row of FP4 stores eight times its logical size. A KV cache with 64-wide heads costs twice its size in the row-major layout an attention kernel wants, and exactly its size in the layout the compiler picks when nobody asks.</p></li><li><p><strong>Prices.</strong> On Google Cloud&#8217;s on-demand rate card, the six accelerators we priced cost $1.40 to $1.65 per TB/s of HBM bandwidth per hour, TPU or GPU. Commitments follow a fixed formula: 70% of on-demand for one year, 45% for three. The estimated anchor rate for Anthropic is 13%.</p></li><li><p><strong>Money.</strong> Broadcom expects Anthropic to go from 1 GW of Ironwood in 2026 to 5 GW of TPU 8i in 2027 and 10 GW in 2028, and guides its AI revenue from $58 billion this fiscal year to $230 billion in fiscal 2028. Alphabet books multi-gigawatt TPU system sales in a $514 billion Cloud backlog and is raising up to $70 billion of equity for its build-out.</p></li></ul><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Compiling for TPUs we don&#8217;t own</h2><p>The whole setup is one line, <code>pip install "jax[tpu]"</code>, which in October 2026 brings jax and jaxlib 0.11.2 and libtpu 0.0.48. JAX can describe a TPU topology that isn&#8217;t attached through <code>jax.experimental.topologies</code>, and any program lowered against that description is compiled by the <strong>real TPU compiler inside libtpu. </strong>Each compile took from under a second to about eight seconds on a single CPU core.</p><p>The topology names are stricter than the marketing names. <code>v4:2x2x1</code>, <code>v5e:2x2</code>, <code>v5p:2x2x1</code>, <code>v6e:2x2</code> and <code>tpu7x:2x2x1</code> are accepted. <code>v7x:2x2x1</code> is rejected with &#8220;Invalid TPU external name: TPU v7x&#8221;, and <code>tpu8t</code>, <code>tpu8i</code> and every other spelling of TPU 8 we tried are rejected the same way. The largest slice we asked for, <code>tpu7x:16x16x16</code>, came back as 8,192 devices.</p><pre><code><code>import numpy as np, jax, jax.numpy as jnp
from jax.experimental import topologies
from jax.sharding import Mesh, NamedSharding, PartitionSpec as P

topo = topologies.get_topology_desc(platform="tpu", topology_name="tpu7x:2x2x1")
mesh = Mesh(np.array(topo.devices[:1]), ("x",))
spec = lambda s: jax.ShapeDtypeStruct(s, jnp.bfloat16, sharding=NamedSharding(mesh, P()))

c = jax.jit(lambda x, w: x @ w).lower(spec((8, 8192)), spec((8192, 28672))).compile()
c.cost_analysis()["optimal_seconds"]   # the compiler's roofline estimate
c.memory_analysis()                    # argument, output, temp and code bytes
c.as_text()                            # optimized HLO, one backend_config per fusion</code></code></pre><p><strong>Five kinds of output came back</strong>, and the rest of this piece is built from them. The cost analysis gives FLOPs, bytes accessed and a time called optimal_seconds. The memory analysis gives argument, output and temporary bytes and the size of the generated code. </p><p><strong>The optimized HLO</strong> attaches a JSON <code>backend_config</code> to every fusion, with the tiling windows the compiler chose, an estimated cycle count, the name of the emitter that generates the code and the on-chip memory it reserves. A dump directory, enabled with <code>XLA_FLAGS</code>, adds 18 entries per module, including one that lists every compiler option with its value. </p><p>And two options passed through <code>LIBTPU_INIT_ARGS</code>, <code>--xla_jf_dump_to</code> and <code>--xla_jf_dump_llo_text=true</code>, make the compiler write every stage of its low-level backend to disk, down to the final bundles, together with a report of how many instructions of each kind a bundle can hold.</p><p>This is <strong>what one fusion looks like</strong> in the optimized HLO for a 4096&#179; bf16 matmul on Ironwood, trimmed to the interesting fields:</p><pre><code><code>ROOT %fusion = bf16[4096,4096]{1,0:T(8,128)(2,1)} fusion(%x.1, %y.1), kind=kOutput,
  backend_config={"window_config":{"kernel_window_bounds":["512","8"],
    "output_window_bounds":["64","8"],"input_window_bounds":["64","32"],
    "estimated_cycles":"324772","iteration_bounds":["4","8","1"],
    "cost_model_type":"COST_MODEL_TYPE_CLASSIC","ml_estimated_microseconds":0,
    "buffering_level":"2"},
  "scoped_memory_configs":[{"memory_space":"1","size":"33554432"}],
  "used_scoped_memory_configs":[{"memory_space":"1","size":"31297536"}],
  "convolution_algorithm_config":{"emitter":"EmitAllBatchInSublanes"}}</code></code></pre><p>The compiler split the output into 32 windows of 512 by 1,024, kept the whole K dimension of each operand window on chip, double-buffered it, reserved 32 MiB of on-chip memory and used 29.8 MiB of it, and predicted 324,772 cycles. <strong>Every number in that line is a decision someone at Google tuned, and none of it appears in the documentation.</strong></p><p>Two fields in that config deserve a second look. <code>cost_model_type</code> says COST_MODEL_TYPE_CLASSIC, and next to it sits <code>ml_estimated_microseconds</code>, set to zero, in every matmul fusion we inspected on Trillium and Ironwood. </p><p>Google published a <strong>learned performance model for TPU programs at MLSys in 2021</strong>, and showed that it beat the production analytical model on exactly the decision this config records, tile-size selection, as well as on fusion. </p><p><strong>TpuGraphs</strong>, a dataset for training such models, followed in 2023. In this release the slot for a learned estimate is present in every matmul fusion, and the analytical model fills in every number we report.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/p/inside-googles-tpu-how-it-works-and?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/p/inside-googles-tpu-how-it-works-and?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><div><hr></div><h2>The 683 MB file</h2><p>libtpu.so in the 0.0.48 wheel is 683,032,200 bytes, with a second 37,636,776-byte library, sdk.so, beside it. </p><p>For scale, ptxas from CUDA 13.3 is 48 MB, tileiras, the closed back end of NVIDIA&#8217;s Tile IR, is 95 MB, and Triton&#8217;s libtriton.so is 180 MB. <strong>libtpu is fourteen times ptxas.</strong></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!SdGi!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F00f7e48d-2585-4425-b2bd-2b77c5440c72_1893x836.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!SdGi!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F00f7e48d-2585-4425-b2bd-2b77c5440c72_1893x836.png 424w, https://substackcdn.com/image/fetch/$s_!SdGi!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F00f7e48d-2585-4425-b2bd-2b77c5440c72_1893x836.png 848w, https://substackcdn.com/image/fetch/$s_!SdGi!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F00f7e48d-2585-4425-b2bd-2b77c5440c72_1893x836.png 1272w, https://substackcdn.com/image/fetch/$s_!SdGi!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F00f7e48d-2585-4425-b2bd-2b77c5440c72_1893x836.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!SdGi!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F00f7e48d-2585-4425-b2bd-2b77c5440c72_1893x836.png" width="1456" height="643" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/00f7e48d-2585-4425-b2bd-2b77c5440c72_1893x836.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:643,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;The TPU compiler is the largest compiler binary we have measured.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="The TPU compiler is the largest compiler binary we have measured." title="The TPU compiler is the largest compiler binary we have measured." srcset="https://substackcdn.com/image/fetch/$s_!SdGi!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F00f7e48d-2585-4425-b2bd-2b77c5440c72_1893x836.png 424w, https://substackcdn.com/image/fetch/$s_!SdGi!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F00f7e48d-2585-4425-b2bd-2b77c5440c72_1893x836.png 848w, https://substackcdn.com/image/fetch/$s_!SdGi!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F00f7e48d-2585-4425-b2bd-2b77c5440c72_1893x836.png 1272w, https://substackcdn.com/image/fetch/$s_!SdGi!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F00f7e48d-2585-4425-b2bd-2b77c5440c72_1893x836.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>The TPU compiler is the largest compiler binary we have measured.</strong> Sizes on disk of the shared libraries and executables as shipped in Python wheels. libtpu 0.0.48 measured for this piece; the NVIDIA and Triton sizes come from the same kind of wheel inventory in our earlier pieces on CUDA binaries, Tile IR and Triton.</figcaption></figure></div><p>The size stops being surprising once you look at what is inside. ptxas is one stage of a pipeline. <strong>libtpu is the pipeline plus the runtime.</strong> Among its 4,237,056 printable strings, the token <code>sparse_core</code> appears 51,866 times, <code>memory_space_assignment</code> 3,100 times and <code>mosaic</code>, the Pallas back end, 828 times. </p><p>Source paths survive as strings: the TPU back end lives under <code>platforms/xla/service/jellyfish</code>, its debugger under <code>platforms/deepsea/jellyfish/xdb</code>.</p><p>The chips are numbered inside the binary. The protocol-buffer descriptor that defines the chip enum is embedded in the file, and decoding it gives <code>TPU_VERSION_INVALID</code> = 0, <code>JELLYFISH</code> = 1, <code>DRAGONFISH</code> = 2, <code>PUFFERFISH</code> = 3, <code>VIPERFISH</code> = 4, <code>GHOSTLITE</code> = 5 and <code>6acc60406</code> = 6, one value per family since TPU v2. </p><p>Ask the compiler which chip it is targeting and the target description it dumps answers: PUFFERFISH for a v4 compile, VIPERFISH for both v5e and v5p, GHOSTLITE for Trillium, and <code>TPU_VERSION_6acc60406</code> for Ironwood. <strong>Every older family kept its codename in the shipped binary. The newest one ships as a string that looks like a hash.</strong></p><p>A second enum, for core types, holds a small piece of history. It has <code>TPU_CORE_TYPE_TENSOR_CORE</code> = 1, <code>TPU_CORE_TYPE_SPARSE_CORE</code> = 3, and value 2 under two names: <code>TPU_CORE_TYPE_BARNA_CORE</code> and <code>TPU_CORE_TYPE_SPARSE_CORE_V0</code>. Google&#8217;s retrospective says the SparseCore was already on the TPU v2 die but was only disclosed in the TPU v4 paper. <strong>The binary still carries the first SparseCore&#8217;s other name, 7,570 times.</strong></p><p>The option surface is just as large. The dump of the compilation environment lists 1,184 top-level options. 861 begin with <code>xla_tpu_</code>, 106 with <code>xla_jf_</code>, and 33 with <code>xla_sc_</code> for the SparseCore. </p><p><strong>Some read like a research agenda</strong>: the scheduler options include a biased random-key genetic algorithm with a generation limit that defaults to 1,200. Most will never matter to anyone outside Google. A handful decide things this piece is about, and we name them where they come up.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>What the TPU compiler thinks each chip is</h2><p>XLA keeps a roofline for every program it compiles. Each instruction gets a time equal to the larger of its FLOPs divided by a peak rate and its bytes divided by a bandwidth, and the sum is reported as optimal_seconds. </p><p>The rates are per-target constants, which means they can be read back. Compile a <strong>matmul large enough to be compute-bound</strong> and divide its FLOPs by its optimal_seconds. Compile an elementwise add large enough to be bandwidth-bound and do the same with its bytes. </p><p><strong>We swept the matmul from 1,024 to 16,384</strong> and the add from 64 MiB to 1 GiB, and both constants plateau exactly.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!RthC!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9df8110d-7c17-4e45-80b2-aad95d6d946a_1928x670.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!RthC!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9df8110d-7c17-4e45-80b2-aad95d6d946a_1928x670.png 424w, https://substackcdn.com/image/fetch/$s_!RthC!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9df8110d-7c17-4e45-80b2-aad95d6d946a_1928x670.png 848w, https://substackcdn.com/image/fetch/$s_!RthC!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9df8110d-7c17-4e45-80b2-aad95d6d946a_1928x670.png 1272w, https://substackcdn.com/image/fetch/$s_!RthC!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9df8110d-7c17-4e45-80b2-aad95d6d946a_1928x670.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!RthC!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9df8110d-7c17-4e45-80b2-aad95d6d946a_1928x670.png" width="1456" height="506" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/9df8110d-7c17-4e45-80b2-aad95d6d946a_1928x670.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:506,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;The five generations, per device, as the compiler describes them.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="The five generations, per device, as the compiler describes them." title="The five generations, per device, as the compiler describes them." srcset="https://substackcdn.com/image/fetch/$s_!RthC!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9df8110d-7c17-4e45-80b2-aad95d6d946a_1928x670.png 424w, https://substackcdn.com/image/fetch/$s_!RthC!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9df8110d-7c17-4e45-80b2-aad95d6d946a_1928x670.png 848w, https://substackcdn.com/image/fetch/$s_!RthC!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9df8110d-7c17-4e45-80b2-aad95d6d946a_1928x670.png 1272w, https://substackcdn.com/image/fetch/$s_!RthC!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9df8110d-7c17-4e45-80b2-aad95d6d946a_1928x670.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>The five generations, per device, as the compiler describes them.</strong> Devices per chip from the topology description; VMEM from Pallas allocation errors; free HBM from the compiler&#8217;s out-of-memory message; peak bf16 rate and HBM bandwidth from the cost-model sweep; B* is the critical decode batch, from the compiler&#8217;s constants and from Google&#8217;s published figures. In these compile-only topologies a v4 TensorCore is one device; in production v4 normally runs both cores as one megacore device.</figcaption></figure></div><p><strong>Two critical things </strong>in that table need explaining before anything else.</p><p>The first is the devices. In these topologies a v4 chip and an Ironwood chip each appear as two devices, with <code>core_on_chip</code> 0 and 1 at the same coordinates, while v5p appears as one device per chip because its two TensorCores are fused into one logical core called <em>megacore</em>. </p><p>Every training TPU since v2 has had <strong>two TensorCores that share only HBM</strong>, according to Google&#8217;s retrospective; megacore is a compiler illusion that started with v4. <strong>Ironwood drops it.</strong> Google&#8217;s TPU7x documentation says frameworks see each Ironwood chip as two devices, one per chiplet, a programming model closer to TPU v3 than to v4 or v5p. </p><p>Each Ironwood device has 94.74 GiB of HBM available, half of the chip&#8217;s 192 GiB minus a reserve, and <strong>64 MiB of VMEM. </strong>Code written for v5p that assumed one device per chip sees twice as many on Ironwood, each with half the memory.</p><p>The second is the ratios between the compiler&#8217;s constants and Google&#8217;s published figures.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!EZpu!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F146980b4-666b-434c-8e11-204430cd9b93_1589x908.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!EZpu!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F146980b4-666b-434c-8e11-204430cd9b93_1589x908.png 424w, https://substackcdn.com/image/fetch/$s_!EZpu!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F146980b4-666b-434c-8e11-204430cd9b93_1589x908.png 848w, https://substackcdn.com/image/fetch/$s_!EZpu!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F146980b4-666b-434c-8e11-204430cd9b93_1589x908.png 1272w, https://substackcdn.com/image/fetch/$s_!EZpu!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F146980b4-666b-434c-8e11-204430cd9b93_1589x908.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!EZpu!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F146980b4-666b-434c-8e11-204430cd9b93_1589x908.png" width="1456" height="832" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/146980b4-666b-434c-8e11-204430cd9b93_1589x908.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:832,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;The compiler's constants against Google's published figures.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="The compiler's constants against Google's published figures." title="The compiler's constants against Google's published figures." srcset="https://substackcdn.com/image/fetch/$s_!EZpu!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F146980b4-666b-434c-8e11-204430cd9b93_1589x908.png 424w, https://substackcdn.com/image/fetch/$s_!EZpu!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F146980b4-666b-434c-8e11-204430cd9b93_1589x908.png 848w, https://substackcdn.com/image/fetch/$s_!EZpu!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F146980b4-666b-434c-8e11-204430cd9b93_1589x908.png 1272w, https://substackcdn.com/image/fetch/$s_!EZpu!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F146980b4-666b-434c-8e11-204430cd9b93_1589x908.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>The compiler&#8217;s constants against Google&#8217;s published figures.</strong> Compiler values per device multiplied by devices per chip, divided by Google&#8217;s per-chip peaks: 275 TFLOP/s and 1,200 GB/s for v4, 197 TFLOP/s and 800 GiB/s for v5e, 459 TFLOP/s and 2,765 GB/s for v5p, 918 TFLOP/s and 1,638 GB/s for Trillium, 2,307 TFLOP/s and 7,380 GB/s for Ironwood. Sources: Google Cloud TPU documentation and Jouppi et al. 2026.</figcaption></figure></div><p>On Trillium the compiler&#8217;s constants are the datasheet: 918 TFLOP/s and 1,638 GB/s. On Ironwood its bandwidth is 3,686 GB/s per core, 7,372 per chip, within 0.11% of the 7,380 GB/s in Google&#8217;s documentation. </p><p>Google&#8217;s own sources don&#8217;t agree with each other here: the architects&#8217; paper lists 7,300. <strong>Its compute is 996 TFLOP/s per core, 1,992 per chip, or 86.3% of the 2,307 Google lists.</strong> v5p carries the same kind of discount on both axes: 394 of 459 TFLOP/s and 2,350 of 2,765 GB/s, 85.8% and 85.0%. v5e and v4 match on compute and are discounted on bandwidth, to 85.9% and 81.9%. </p><p>One more cross-check falls out of the memory report: v5p&#8217;s 95.73 GiB of free HBM per device fits the 96 GiB in the retrospective better than the 95 GiB in the documentation table.</p><p>The binary doesn&#8217;t say whether the compute discounts are clock assumptions or margins, but it gives a hint. The <strong>low-level backend reports how many instructions of each kind fit in one bundle</strong>, and the MXU count changes between generations: four MXU slots per bundle on v4, v5e and v5p, two on Trillium and Ironwood.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!sraB!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97e5b2b9-3b50-4ffe-b2a9-9947505e737e_1654x892.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!sraB!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97e5b2b9-3b50-4ffe-b2a9-9947505e737e_1654x892.png 424w, https://substackcdn.com/image/fetch/$s_!sraB!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97e5b2b9-3b50-4ffe-b2a9-9947505e737e_1654x892.png 848w, https://substackcdn.com/image/fetch/$s_!sraB!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97e5b2b9-3b50-4ffe-b2a9-9947505e737e_1654x892.png 1272w, https://substackcdn.com/image/fetch/$s_!sraB!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97e5b2b9-3b50-4ffe-b2a9-9947505e737e_1654x892.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!sraB!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97e5b2b9-3b50-4ffe-b2a9-9947505e737e_1654x892.png" width="1456" height="785" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/97e5b2b9-3b50-4ffe-b2a9-9947505e737e_1654x892.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:785,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Issue slots per VLIW bundle, by generation.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Issue slots per VLIW bundle, by generation." title="Issue slots per VLIW bundle, by generation." srcset="https://substackcdn.com/image/fetch/$s_!sraB!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97e5b2b9-3b50-4ffe-b2a9-9947505e737e_1654x892.png 424w, https://substackcdn.com/image/fetch/$s_!sraB!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97e5b2b9-3b50-4ffe-b2a9-9947505e737e_1654x892.png 848w, https://substackcdn.com/image/fetch/$s_!sraB!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97e5b2b9-3b50-4ffe-b2a9-9947505e737e_1654x892.png 1272w, https://substackcdn.com/image/fetch/$s_!sraB!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97e5b2b9-3b50-4ffe-b2a9-9947505e737e_1654x892.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Issue slots per VLIW bundle, by generation.</strong> Capacities as printed at the top of the compiler&#8217;s static per-bundle utilization report for a matmul fusion. Red marks a change from the generation to the left. The vector ALU column matches Google&#8217;s description of the VPU growing from two restricted ALUs per lane to four general-purpose ones. At the slot level an Ironwood TensorCore is identical to a Trillium TensorCore.</figcaption></figure></div><p>Divide the compiler&#8217;s cycle estimates for large matmuls by the number of MXU slots and you get the work one slot does per cycle. On the four-slot chips it is 15,350 to 15,860 bf16 multiply-adds, close to the 16,384 of one 128 by 128 systolic array. </p><p>That matches Google&#8217;s counts:<strong> four MXUs per TensorCore on v5e, eight 128 by 128 MXUs per chip on v4 and v5p.</strong> On Trillium a slot reaches 113,900 multiply-adds per cycle and on Ironwood 108,800, which is impossible for one 256 by 256 array (65,536) and consistent with two arrays per slot, or one array retiring two rows per cycle.</p><p>Combine slot rates with the published peaks and you get clocks the datasheets don&#8217;t print: 1.05 GHz for v4, which matches the 1,050 MHz in Google&#8217;s TPU v4 paper, 1.50 GHz for v5e, 1.75 GHz for v5p and Trillium, and 2.20 GHz for Ironwood. <strong>The retrospective describes Ironwood&#8217;s matrix units</strong> as four 256 by 256 arrays for bf16. </p><p>Counted per chip, reaching 2,307 TFLOP/s with four such arrays would need a 4.4 GHz clock; counted per TensorCore, it needs 2.2 GHz, which is what the compiler&#8217;s slot rates imply. </p><p><strong>We read the paper&#8217;s four arrays as per TensorCore</strong>. Run the same arithmetic on the compiler&#8217;s constants instead of the datasheet and v5p drops to 1.50 GHz and Ironwood to 1.90 GHz. </p><p>One consistent reading is that the compiler models those two chips at a lower clock than their peak figures assume. Another is that it simply derates them. <strong>Either way, the compiler&#8217;s own roofline for Ironwood uses 996 TFLOP/s per core, not 1,153.</strong></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>The critical decode batch, v4 to TPU 8i</h2><p>In earlier pieces we kept coming back to one number: the decode batch at which a chip stops waiting for its weights. A decode step has to stream every weight from HBM at least once. </p><p>At small batch that stream is the whole cost, and adding sequences is nearly free until the arithmetic catches up.<strong> The crossover is B* = P&#183;b / (2&#183;B_hbm)</strong>, where P is the FLOP rate at the precision you use, b the bytes per weight and B_hbm the bandwidth. </p><p>Because P usually doubles whenever b halves, B* barely moves with precision. On an H100 it is 295 whether you run bf16 or FP8, on a B200 about 281.</p><p><strong>Between</strong> <strong>TPU generations it moves a lot.</strong></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!hCVO!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe3e3f98f-9632-4361-9443-4fc51a0bf1a5_1764x1191.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!hCVO!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe3e3f98f-9632-4361-9443-4fc51a0bf1a5_1764x1191.png 424w, https://substackcdn.com/image/fetch/$s_!hCVO!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe3e3f98f-9632-4361-9443-4fc51a0bf1a5_1764x1191.png 848w, https://substackcdn.com/image/fetch/$s_!hCVO!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe3e3f98f-9632-4361-9443-4fc51a0bf1a5_1764x1191.png 1272w, https://substackcdn.com/image/fetch/$s_!hCVO!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe3e3f98f-9632-4361-9443-4fc51a0bf1a5_1764x1191.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!hCVO!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe3e3f98f-9632-4361-9443-4fc51a0bf1a5_1764x1191.png" width="1456" height="983" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e3e3f98f-9632-4361-9443-4fc51a0bf1a5_1764x1191.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:983,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Critical decode batch per chip.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Critical decode batch per chip." title="Critical decode batch per chip." srcset="https://substackcdn.com/image/fetch/$s_!hCVO!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe3e3f98f-9632-4361-9443-4fc51a0bf1a5_1764x1191.png 424w, https://substackcdn.com/image/fetch/$s_!hCVO!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe3e3f98f-9632-4361-9443-4fc51a0bf1a5_1764x1191.png 848w, https://substackcdn.com/image/fetch/$s_!hCVO!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe3e3f98f-9632-4361-9443-4fc51a0bf1a5_1764x1191.png 1272w, https://substackcdn.com/image/fetch/$s_!hCVO!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe3e3f98f-9632-4361-9443-4fc51a0bf1a5_1764x1191.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Critical decode batch per chip.</strong> Hatched bars use Google&#8217;s published figures, solid bars the compiler&#8217;s own constants. TPU 8 values use Google&#8217;s April figures and assume, as Google states for 8t, that FP4 doubles FP8 throughput; the 8i FP8 rate is derived from the 11.6 EFLOP/s pod figure. NVIDIA values from published dense peaks: H100 SXM 989 TFLOP/s at 3.35 TB/s, H200 989 at 4.8, HGX B200 2,250 at 8.0.</figcaption></figure></div><p><strong>Trillium&#8217;s critical batch is 560</strong>, by the datasheet and by the compiler, which agree exactly. That is the highest of any shipping accelerator in this piece, NVIDIA&#8217;s included. </p><p>Trillium raised compute 4.7 times over v5e, and the slot table shows how: four times the multiply-adds per cycle, from half as many slots each eight times larger, at 1.75 GHz instead of 1.50. <strong>Bandwidth only grew 1.9 times</strong>, so Trillium needs about twice as many concurrent sequences per chip as an H100 before its FLOPs start to pay. </p><p>That&#8217;s an odd profile for a chip Google&#8217;s architects describe, in a footnote, as focused on inference. <strong>It suits batched, throughput-bound serving, and it punishes latency-bound decode.</strong></p><p><strong>Ironwood undid it.</strong> It raised bandwidth 4.5 times over Trillium, to 7,380 GB/s, while raising compute 2.5 times, which puts its critical batch at 313 on paper and 270 by the compiler&#8217;s constants, close to the H100&#8217;s 295 and the B200&#8217;s 281. </p><p>Google introduced Ironwood as its first TPU built for inference, and its architects list it as the fifth of their training supercomputers. In roofline terms both descriptions fit: <strong>it is the generation that brought the critical batch back down.</strong></p><p>The compiler&#8217;s estimates agree when you compile real shapes. We compiled one 8192 by 28672 projection, the shape of a large dense MLP layer, at every batch from 2 to 2,048 and plotted the estimated time against batch 2.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!6iBv!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff72a3607-ea7d-4bd4-8a56-ebcabbaae46b_1546x975.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!6iBv!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff72a3607-ea7d-4bd4-8a56-ebcabbaae46b_1546x975.png 424w, https://substackcdn.com/image/fetch/$s_!6iBv!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff72a3607-ea7d-4bd4-8a56-ebcabbaae46b_1546x975.png 848w, https://substackcdn.com/image/fetch/$s_!6iBv!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff72a3607-ea7d-4bd4-8a56-ebcabbaae46b_1546x975.png 1272w, https://substackcdn.com/image/fetch/$s_!6iBv!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff72a3607-ea7d-4bd4-8a56-ebcabbaae46b_1546x975.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!6iBv!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff72a3607-ea7d-4bd4-8a56-ebcabbaae46b_1546x975.png" width="1456" height="918" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f72a3607-ea7d-4bd4-8a56-ebcabbaae46b_1546x975.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:918,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Where each generation leaves the memory wall.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Where each generation leaves the memory wall." title="Where each generation leaves the memory wall." srcset="https://substackcdn.com/image/fetch/$s_!6iBv!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff72a3607-ea7d-4bd4-8a56-ebcabbaae46b_1546x975.png 424w, https://substackcdn.com/image/fetch/$s_!6iBv!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff72a3607-ea7d-4bd4-8a56-ebcabbaae46b_1546x975.png 848w, https://substackcdn.com/image/fetch/$s_!6iBv!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff72a3607-ea7d-4bd4-8a56-ebcabbaae46b_1546x975.png 1272w, https://substackcdn.com/image/fetch/$s_!6iBv!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff72a3607-ea7d-4bd4-8a56-ebcabbaae46b_1546x975.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Where each generation leaves the memory wall.</strong> Compiler-estimated time (optimal_seconds) for one bf16 projection of 8,192 by 28,672, per device, relative to batch 2. Triangles mark each generation&#8217;s critical batch from its compiler constants. Every curve stays flat until its triangle.</figcaption></figure></div><p>On Ironwood the estimate at batch 2 is 127.5 &#181;s. It is 7.5% higher at batch 256 and 1.97 times higher at 512. </p><p>On Trillium it is 287 &#181;s at batch 2, only 15.1% higher at batch 512 and 1.97 times higher at 1,024. v5p, the chip with the most bandwidth per FLOP, bends first: 17% above its floor at batch 192 and 56% above at 256. In absolute terms Ironwood is still the fastest at every batch, 251 &#181;s at batch 512 against Trillium&#8217;s 330. <strong>What changes is where each chip stops being free.</strong></p><p>The eighth generation splits the difference in an instructive way. TPU 8t keeps HBM bandwidth modest at 6,528 GB/s while doubling MXU throughput with native FP4, so its critical batch is about 482 at FP8 or FP4.</p><p> That&#8217;s a training chip. TPU 8i has more bandwidth, 8,601 GB/s, and by Google&#8217;s own numbers the same throughput at FP4 as at FP8: <strong>10.1 PFLOP/s at FP4 per chip</strong>, and <strong>11.6 EFLOP/s at FP8 across a 1,152-chip pod, </strong>which is 10.07 per chip. If those figures mean what they appear to mean, 8i&#8217;s critical batch is 585 with FP8 weights and 294 with FP4 weights.</p><p>That breaks the precision invariance our earlier pieces relied on, and it breaks it in a useful direction. On 8i, FP4 is <em>not</em> a way to get more FLOPs. It is a way to halve the bytes each decode step has to read, which halves the batch you need before the chip is busy. </p><p><strong>It is a compute format on the training chip and a bandwidth format on the inference chip.</strong> One caveat: the per-chip FP8 rate comes from a pod total. If Google counts only the 1,024 active chips its Boardfly description mentions, FP8 would be 11.3 PFLOP/s per chip and the two formats would no longer match.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>A batch of one</h2><p>Below the critical batch every extra row is nearly free, so a batch of one should cost what a batch of eight costs. <strong>On Ironwood it does. On every older TPU, by the compiler&#8217;s own estimate, it costs about three times as much.</strong></p><p>The reason is visible in the optimized HLO. Compile a [1, 8192] by [8192, 28672] product for v4, v5e, v5p or Trillium and there is no matrix multiply left in it. The compiler has rewritten the dot into an elementwise multiply followed by a reduction, which runs on the vector unit. </p><p>Give it two rows and you get a convolution, generated by the MXU emitter <code>EmitAllBatchInSublanes</code>. The rewrite comes from a pass whose source file name survives in the binary, <code>tpu_dot_strength_reducer.cc</code>, and it is controlled by the option <code>xla_tpu_enable_dot_strength_reduction</code>, which defaults to true.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!H5sY!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9752c977-1b0a-478b-83ea-26a9be78d41a_1551x908.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!H5sY!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9752c977-1b0a-478b-83ea-26a9be78d41a_1551x908.png 424w, https://substackcdn.com/image/fetch/$s_!H5sY!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9752c977-1b0a-478b-83ea-26a9be78d41a_1551x908.png 848w, https://substackcdn.com/image/fetch/$s_!H5sY!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9752c977-1b0a-478b-83ea-26a9be78d41a_1551x908.png 1272w, https://substackcdn.com/image/fetch/$s_!H5sY!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9752c977-1b0a-478b-83ea-26a9be78d41a_1551x908.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!H5sY!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9752c977-1b0a-478b-83ea-26a9be78d41a_1551x908.png" width="1456" height="852" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/9752c977-1b0a-478b-83ea-26a9be78d41a_1551x908.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:852,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;A batch of one, priced by the compiler's own cycle model.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="A batch of one, priced by the compiler's own cycle model." title="A batch of one, priced by the compiler's own cycle model." srcset="https://substackcdn.com/image/fetch/$s_!H5sY!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9752c977-1b0a-478b-83ea-26a9be78d41a_1551x908.png 424w, https://substackcdn.com/image/fetch/$s_!H5sY!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9752c977-1b0a-478b-83ea-26a9be78d41a_1551x908.png 848w, https://substackcdn.com/image/fetch/$s_!H5sY!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9752c977-1b0a-478b-83ea-26a9be78d41a_1551x908.png 1272w, https://substackcdn.com/image/fetch/$s_!H5sY!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9752c977-1b0a-478b-83ea-26a9be78d41a_1551x908.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>A batch of one, priced by the compiler&#8217;s own cycle model.</strong> Estimated cycles for a [1, 8192] by [8192, 28672] bf16 product, relative to the same product at batch 2. Red: as compiled by default. Blue: with xla_tpu_enable_dot_strength_reduction=false. On Ironwood the rewrite never fires for this shape.</figcaption></figure></div><p>The cycle estimates attached to each program price the decision. On v4 the rewritten product takes 3,154,024 cycles against 1,077,235 on the MXU, 2.93 times as many. </p><p>On v5e it is 3,280,393 against 1,147,598 (2.86 times), on v5p 1,292,915 against 470,757 (2.75 times), and on Trillium 1,986,785 against 568,013 (3.50 times). <strong>Turn the option off with </strong><code>LIBTPU_INIT_ARGS=--xla_tpu_enable_dot_strength_reduction=false</code><strong> and batch one costs what batch two costs.</strong> </p><p>On Ironwood the pass doesn&#8217;t fire for this shape at all, and batch one stays on the MXU at 352,341 cycles either way.</p><p>The final bundles make the difference concrete. With the rewrite, Trillium&#8217;s program for a [1, 8192] by [8192, 1024] product contains no matrix instruction. <strong>It&#8217;s 855 vector multiplies, 506 vector adds, 966 unpack instructions </strong>that widen bf16 to f32, and 244 cross-lane broadcasts, permutes and rotates. </p><p>Without the rewrite, the same product is 128 matmul issues fed by 2,048 weight pushes, the same instruction stream the compiler emits for batch eight.</p><p><strong>We don&#8217;t know why the rewrite is on</strong>. A plausible reason is that it was tuned for small matrices, where pushing weights into the MXU costs more than the arithmetic they feed. For an 8,192-wide projection the compiler&#8217;s cost model disagrees with its own pattern pass, and the pattern pass runs first. </p><p>In production this rarely bites, because serving engines usually pad the decode batch to a few fixed sizes so that each size compiles once, and the smallest size is usually larger than one. <strong>It does bite in hand-written single-stream decode loops</strong>, in draft models for speculative decoding that run at batch one, and in mixture-of-experts code that issues one dot per expert when an expert receives a single token. </p><p>If you run any of those on v5e or Trillium, the option is worth a test on real hardware. <em>We can only tell you what the compiler predicts.</em></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>The tile and the KV cache</h2><p>Every TPU array lives in tiles, and the layout string in the optimized HLO says which. A bf16 matrix is <code>bf16[4096,4096]{1,0:T(8,128)(2,1)}</code>: the last dimension varies fastest, the array is cut into tiles of 8 by 128 32-bit words, and (2,1) means two bf16 rows are packed into each word. f32 is <code>T(8,128)</code>. </p><p>Eight-bit types are <code>T(32,128)(4,1)</code> and four-bit types <code>T(64,128)(8,1)</code>. Every one of these tiles is 4 KiB. On the training chips from TPU v2 through v5p that is exactly one vector register, 8 sublanes by 128 lanes of 32 bits. </p><p>Ironwood&#8217;s registers are 16 by 256 according to the retrospective, four times larger, but its memory layouts use the same 8 by 128 tiles as every other generation we compiled for. <strong>Narrow types don&#8217;t get smaller tiles. They get more rows per word.</strong></p><p>The compiler is careful with small arrays. For an array with one row it picks a <code>T(1,128)</code> tile instead of padding to eight rows, and for two or four rows <code>T(2,128)</code> or <code>T(4,128)</code>. </p><p>What it can&#8217;t do is store less than one 32-bit word per lane. <strong>So on all five generations a single row of bf16 occupies twice its logical size, a single row of FP8 four times and a single row of FP4 eight times.</strong> FP4 activations need eight rows before they stop paying for air; FP8 needs four.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!siWb!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37891f82-84d5-4ef7-957f-6fec17818291_1918x992.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!siWb!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37891f82-84d5-4ef7-957f-6fec17818291_1918x992.png 424w, https://substackcdn.com/image/fetch/$s_!siWb!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37891f82-84d5-4ef7-957f-6fec17818291_1918x992.png 848w, https://substackcdn.com/image/fetch/$s_!siWb!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37891f82-84d5-4ef7-957f-6fec17818291_1918x992.png 1272w, https://substackcdn.com/image/fetch/$s_!siWb!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37891f82-84d5-4ef7-957f-6fec17818291_1918x992.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!siWb!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37891f82-84d5-4ef7-957f-6fec17818291_1918x992.png" width="1456" height="753" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/37891f82-84d5-4ef7-957f-6fec17818291_1918x992.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:753,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Bytes stored per logical byte on Ironwood.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Bytes stored per logical byte on Ironwood." title="Bytes stored per logical byte on Ironwood." srcset="https://substackcdn.com/image/fetch/$s_!siWb!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37891f82-84d5-4ef7-957f-6fec17818291_1918x992.png 424w, https://substackcdn.com/image/fetch/$s_!siWb!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37891f82-84d5-4ef7-957f-6fec17818291_1918x992.png 848w, https://substackcdn.com/image/fetch/$s_!siWb!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37891f82-84d5-4ef7-957f-6fec17818291_1918x992.png 1272w, https://substackcdn.com/image/fetch/$s_!siWb!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37891f82-84d5-4ef7-957f-6fec17818291_1918x992.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Bytes stored per logical byte on Ironwood.</strong> Left: a [rows, 4096] array at four widths, from the compiler&#8217;s memory report; v4, v5e, v5p and Trillium give identical results. Right: KV-cache arrays in bf16. With no dimension of 128 or more, [64,16,8,64] pays 2x. A realistic [4096,16,8,64] cache pays nothing in the layout the compiler picks, which puts the 4,096 pages innermost, and 2x when forced into the row-major layout an attention kernel reads. Folding heads into the minor dimension, [64,16,512], removes the padding in row-major order.</figcaption></figure></div><p><strong>The minor dimension is less forgiving</strong>. Storage tiles are 128 elements wide on every generation, and a minor dimension narrower than that is padded unless the compiler can move a larger dimension into its place. When nobody constrains the layout, it does. </p><p>A [4096, 64] array is stored transposed, layout <code>{0,1}</code>, with the 4,096 dimension innermost, and costs exactly its logical size. A KV cache shaped [4096 pages, 16 tokens, 8 heads, 64 dims] gets layout <code>{0,3,2,1}</code>, with the page index innermost, and also costs exactly its size, as do the same cache with 16,384 pages, 16 heads or 32-token pages.</p><p><strong>That default is a storage decision, not an access decision.</strong> It scatters each token&#8217;s 64-wide head across 64 sublane rows and interleaves 128 different pages in every vector, which is the opposite of what an attention kernel wants. <strong>Force the same cache into row-major order, layout </strong><code>{3,2,1,0}</code><strong>, and it costs 2.0 times its logical size, in bf16 and FP8 alike.</strong> </p><p>A small block such as [64, 16, 8, 64], which has no dimension of 128 or more to move inward, pays 2.0 times even in the default layout. Fold the heads into the minor dimension, [64, 16, 512], and row-major order pays nothing. With 128-wide heads the problem doesn&#8217;t exist. </p><p>Several open models use 64-wide heads, and <strong>on a TPU the layout of their KV cache decides whether half of it is padding</strong>. Kernels fix their own layouts, so this is a decision the kernel author makes once for everyone who uses the kernel.</p><p>This is the TPU version of a constraint we met on Blackwell, where the smallest UMMA tile has 64 rows. On NVIDIA the floor is in the math unit. On the TPU it is in the storage format, which means it costs HBM capacity and bandwidth, not just idle multipliers.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share The Software Frontier&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share The Software Frontier</span></a></p><div><hr></div><h2>Precision: bf16, int8, FP8 and FP4 per generation</h2><p>Datasheets list one number per format, when they list it at all. The compiler has to decide how to run every format on every chip, and its cycle estimates show what it decided. </p><p>We compiled <strong>one 8192 by 8192 by 8192 matmul</strong> in seven input types on all five generations and divided each estimate by the bf16 one.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Wtq4!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc2a270d7-e691-4c5e-a6e0-664b4aa60fd6_1984x889.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Wtq4!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc2a270d7-e691-4c5e-a6e0-664b4aa60fd6_1984x889.png 424w, https://substackcdn.com/image/fetch/$s_!Wtq4!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc2a270d7-e691-4c5e-a6e0-664b4aa60fd6_1984x889.png 848w, https://substackcdn.com/image/fetch/$s_!Wtq4!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc2a270d7-e691-4c5e-a6e0-664b4aa60fd6_1984x889.png 1272w, https://substackcdn.com/image/fetch/$s_!Wtq4!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc2a270d7-e691-4c5e-a6e0-664b4aa60fd6_1984x889.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Wtq4!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc2a270d7-e691-4c5e-a6e0-664b4aa60fd6_1984x889.png" width="1456" height="652" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c2a270d7-e691-4c5e-a6e0-664b4aa60fd6_1984x889.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:652,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;The precision matrix, as the compiler prices it.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="The precision matrix, as the compiler prices it." title="The precision matrix, as the compiler prices it." srcset="https://substackcdn.com/image/fetch/$s_!Wtq4!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc2a270d7-e691-4c5e-a6e0-664b4aa60fd6_1984x889.png 424w, https://substackcdn.com/image/fetch/$s_!Wtq4!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc2a270d7-e691-4c5e-a6e0-664b4aa60fd6_1984x889.png 848w, https://substackcdn.com/image/fetch/$s_!Wtq4!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc2a270d7-e691-4c5e-a6e0-664b4aa60fd6_1984x889.png 1272w, https://substackcdn.com/image/fetch/$s_!Wtq4!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc2a270d7-e691-4c5e-a6e0-664b4aa60fd6_1984x889.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>The precision matrix, as the compiler prices it.</strong> Estimated cycles for an 8192&#179; matmul in each input type, relative to bf16 on the same chip, with f32 accumulation for floating types and int32 for integer types. Below 1.0 is faster. No convert instructions appear in the optimized HLO for any cell; the emitters handle every type directly.</figcaption></figure></div><p>Ironwood runs FP8 in 0.48 times the cycles of bf16, which is the 2x Google advertises: 4,614 TFLOP/s against 2,307. <strong>It runs int8 in 1.13 times the cycles of bf16, slower than bf16</strong>, and int4 and FP4 in 1.15 times. Nothing in the optimized HLO converts the inputs. The emitter handles each narrow type itself, and on Ironwood only FP8 has a fast path in the cost model.</p><p>The older chips are the mirror image. int8 runs in 0.47 times the bf16 cycles on v5e, 0.53 on v5p and 0.62 on Trillium, consistent with the int8 rate Google publishes for v5e, 393 TOPS, twice its bf16 peak. </p><p><strong>FP8 e4m3 gets little or nothing</strong>: 0.94, 1.03 and 0.86, in line with Google&#8217;s own table, which lists v5p and Trillium at the same FP8 rate as bf16. One result we can&#8217;t explain: FP8 e5m2 runs exactly as fast as int8 on v5e, v5p and Trillium, in the same number of cycles to the unit. </p><p>Either the cost model files e5m2 under the eight-bit integer path, or the hardware has a fast path for it that no datasheet mentions. Without a chip we can&#8217;t tell which. v4 accelerates nothing: every narrow type costs 1.08 to 1.17 times bf16.</p><p>Two practical consequences follow. If you serve int8-quantized weights on v5e or Trillium, those checkpoints lose their speed on Ironwood and <strong>should be re-quantized to FP8 before the migration, not after</strong>. And no chip the public compiler knows runs FP4 faster than its eight-bit path. </p><p>TPU 8t is the first TPU with native FP4, and on 8i, as above, FP4 appears to run at the FP8 rate. Our <strong>FP4 piece</strong> <strong>compared NVIDIA&#8217;s NVFP4 with the OCP MXFP4 format that AMD uses.</strong> Google hasn&#8217;t said, in the material we read, which FP4 the eighth generation uses or how it scales its blocks.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Inside one matmul on Trillium and Ironwood</h2><p><strong>The deepest thing libtpu will show you is its low-level IR</strong>, which it calls LLO, all the way to the final bundles. For a single 1024 by 1024 by 1024 bf16 matmul, the two dump options write 72 stages for Ironwood and 73 for Trillium. </p><p>The names read like a compiler textbook written for a VLIW machine: pre-auto-mxu-assigner, bf16-coalescing, x8-coalescing, MXU-assigner, vmac-transform, critical-path-scheduler, vliw-packed-bundles, vliw-bundle-scheduler, register-pressure, post-ra, delay-converter and, last, final_bundles. Trillium has one stage Ironwood doesn&#8217;t, post-allocation-offset-shuffling.</p><p>A bundle is the set of instructions issued together. TPU v2&#8217;s scalar unit fetched bundles of 322 bits, according to the retrospective, and Ironwood&#8217;s are more than 50% wider. </p><p>The slot table above says what one can hold. <strong>On both Trillium and Ironwood it is two MXU instructions</strong>, four vector ALU operations, two cross-lane operations, two result pops, one transcendental, three vector loads, two vector stores and two scalar operations, plus dedicated slots for spill traffic. </p><p>At the level of the instruction format the two TensorCores are identical. <strong>The programs they get are not.</strong></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!JgXa!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98f582f7-0c52-49e3-a1e8-73511d84c509_1763x1067.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!JgXa!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98f582f7-0c52-49e3-a1e8-73511d84c509_1763x1067.png 424w, https://substackcdn.com/image/fetch/$s_!JgXa!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98f582f7-0c52-49e3-a1e8-73511d84c509_1763x1067.png 848w, https://substackcdn.com/image/fetch/$s_!JgXa!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98f582f7-0c52-49e3-a1e8-73511d84c509_1763x1067.png 1272w, https://substackcdn.com/image/fetch/$s_!JgXa!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98f582f7-0c52-49e3-a1e8-73511d84c509_1763x1067.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!JgXa!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98f582f7-0c52-49e3-a1e8-73511d84c509_1763x1067.png" width="1456" height="881" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/98f582f7-0c52-49e3-a1e8-73511d84c509_1763x1067.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:881,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;One 1024&#179; bf16 matmul, counted instruction by instruction.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="One 1024&#179; bf16 matmul, counted instruction by instruction." title="One 1024&#179; bf16 matmul, counted instruction by instruction." srcset="https://substackcdn.com/image/fetch/$s_!JgXa!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98f582f7-0c52-49e3-a1e8-73511d84c509_1763x1067.png 424w, https://substackcdn.com/image/fetch/$s_!JgXa!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98f582f7-0c52-49e3-a1e8-73511d84c509_1763x1067.png 848w, https://substackcdn.com/image/fetch/$s_!JgXa!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98f582f7-0c52-49e3-a1e8-73511d84c509_1763x1067.png 1272w, https://substackcdn.com/image/fetch/$s_!JgXa!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98f582f7-0c52-49e3-a1e8-73511d84c509_1763x1067.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>One 1024&#179; bf16 matmul, counted instruction by instruction.</strong> Final VLIW bundles from the compiler&#8217;s LLO dump, grouped by kind; spill loads and stores are vector loads and stores to compiler-allocated spill buffers, shown separately. The two chips load the same 2,048 vectors and store the same 512. The difference is accumulator traffic.</figcaption></figure></div><p><strong>Trillium&#8217;s program is 4,342 bundles and 17,639 instruction</strong>s. Ironwood&#8217;s is 3,094 bundles and 8,677 instructions. Strip out the spills and both load the same 2,048 vectors and store the same 512 results, and both issue the same 1,024 matrix multiplies after the same 256 weight pushes. </p><p>Everything else is the cost of getting partial sums out of the matrix unit. <strong>Trillium pops 4,096 results and Ironwood 1,024. Trillium executes 3,072 vector adds and Ironwood none.</strong> Trillium issues 2,384 spill loads and 1,396 spill stores, Ironwood 476 of each.</p><p>Here are three consecutive bundles from the middle of Trillium&#8217;s program, with registers renamed and <strong>VMEM addresses shortened:</strong></p><pre><code><code>vmatmul.mubr.bf16.gmra.mxu0 v51 ;; v21 = vpop.f32.mrf.mxu0 ;; vmatprep.mubr.bf16.mxu1 v61 ;; v30 = vld [buf+0xc30]
vmatprep.mubr.bf16.mxu0 v44 ;; v24 = vpack.c.bf16 v41, v23 ;; v3 = vadd.f32 v19, v5 ;; v57 = vadd.f32 v21, v22
    ;; v1 = vpop.f32.mrf.mxu1 ;; vmatpush1.bf16.msra.mxu1 v45 ;; v41 = vld [spill] ;; v23 = vld [spill] ;; v21 = vld [spill]
v31 = vpop.f32.mrf.mxu0 ;; v45 = vld [spill] ;; v22 = vld [spill]</code></code></pre><p>and four from Ironwood&#8217;s:</p><pre><code><code>v21 = vpop.f32.mrb[225].mxu0 ;; v43 = vpack.c.bf16 v37, v19 ;; v0 = vpop.f32.mrb[226].mxu1
v27 = vpack.c.bf16 v21, v62 ;; v12 = vpop.f32.mrb[226].mxu0 ;; v40 = vpop.f32.mrb[227].mxu1 ;; v21 = vld [spill]
v33 = vpop.f32.mrb[227].mxu0 ;; vst [buf+0xe08] v43 ;; v41 = vpack.c.bf16 v40, v0
vst [buf+0xe00] v27 ;; v63 = vpack.c.bf16 v33, v12 ;; vmatmul.mubr.bf16.gmra.mrb[76].mxu1 v2 ;; v27 = vld [spill]</code></code></pre><p>On Trillium, results leave the matrix unit through something called mrf, with no address, and are added together on the vector unit while partial sums wait in registers and spill to VMEM. <strong>On Ironwood the matrix instruction names a destination, </strong><code>mrb[76]</code><strong>, and each pop names a source, </strong><code>mrb[225]</code><strong>, and nothing is added afterwards.</strong> Across the whole program the indices run from <code>mrb[0]</code> to <code>mrb[255]</code>.</p><p>The counts fit one reading exactly. The 1024 by 1024 f32 output is 1,024 vector registers&#8217; worth of 8 by 128 tiles. Trillium pops each of them four times, because <strong>one 256-deep pass through the matrix unit</strong> covers a quarter of K = 1024, and sums the four partials with three adds: 4,096 pops and 3,072 adds. Ironwood pops each one once. </p><p>A decode-shaped product, [8, 8192] by [8192, 1024], has 8 output tiles and a K that takes 32 passes. Trillium pops 256 times and adds 248 times; Ironwood pops 8 times and adds nothing.</p><p>We read mrf as <em>matrix result FIFO</em> and mrb as <em>matrix result buffer</em>: <strong>on Ironwood, partial sums accumulate inside the matrix unit, in addressed entries, for as many passes as K needs</strong>. <em>That interpretation is ours, not Google&#8217;s.</em> It is the same move NVIDIA made between Hopper and Blackwell, when the tensor core&#8217;s accumulator left the register file for tensor memory. </p><p>In our TMEM piece the reason was that Blackwell&#8217;s largest accumulator no longer fit in 255 registers. Here the measurable effect is that the vector unit stops doing the matrix unit&#8217;s bookkeeping. The <strong>result-pop slots are busy 16.5% of the time</strong> on Ironwood against 47.2% on Trillium, the vector ALUs 12.4% against 26.5%, and the spill-store slots 7.7% against 16.1%, while the MXU slots stay above 90% on both.</p><p>Those idle slots now have a second use. The retrospective describes a hardware replay unit in Ironwood&#8217;s vector unit that samples vector bundles at random and <strong>re-executes them in idle VLIW slots</strong>, replaying odd-lane operations on even lanes to catch silent data corruption without slowing the program. A matmul that leaves seven in eight vector ALU slots empty gives it plenty of room.</p><p>A caution about where the accumulator matters. In the decode-shaped program both chips spend about 4,100 bundles, dominated by roughly 8,200 vector loads and 2,048 weight pushes. <strong>When a program is a stream of weights, an accumulator saves little.</strong> </p><p>It matters for prefill, for training and for any matmul deep enough in K that partial sums pile up, which is exactly the work that turns into spills on Trillium.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/p/inside-googles-tpu-how-it-works-and/comments&quot;,&quot;text&quot;:&quot;Leave a comment&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/p/inside-googles-tpu-how-it-works-and/comments"><span>Leave a comment</span></a></p><div><hr></div><h2>SparseCores, collectives and MoE dispatch</h2><p>SparseCores are small dataflow processors built for embedding lookups. Google&#8217;s retrospective counts two per chip on TPU v2 and v3 and four on v4, v5p and Ironwood; <strong>Google&#8217;s documentation gives Trillium two.</strong> Each has 16 compute tiles; together they take about 5% of the die area and of the power, and Ironwood&#8217;s are 2.4 times faster than v5p&#8217;s. </p><p>The retrospective also says they began serving as offload engines for collectives such as AllReduce, AllGather, ReduceScatter and Broadcast once Transformers took over the workload. <strong>In the public compiler, with default options, that only happens on Ironwood.</strong></p><p>Compile an all-reduce across four Ironwood chips, eight devices, and the optimized HLO contains no all-reduce running on the TensorCore. </p><p>It contains an asynchronous call on an execution thread named <code>sparsecore</code>, with a configuration that says <code>"device_type":"DEVICE_TYPE_SPARSECORE"</code> and <code>"offload":"OFFLOAD_COLLECTIVE"</code>, and a ring description in which one phase is marked <code>"across_cores_on_chip":true</code>, the hop between the two TensorCores of a chip. <strong>The TensorCore starts the call and carries on. The SparseCores move the data.</strong></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!pfEp!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F364e6c53-caad-4cd9-820f-b1e97646a237_1700x953.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!pfEp!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F364e6c53-caad-4cd9-820f-b1e97646a237_1700x953.png 424w, https://substackcdn.com/image/fetch/$s_!pfEp!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F364e6c53-caad-4cd9-820f-b1e97646a237_1700x953.png 848w, https://substackcdn.com/image/fetch/$s_!pfEp!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F364e6c53-caad-4cd9-820f-b1e97646a237_1700x953.png 1272w, https://substackcdn.com/image/fetch/$s_!pfEp!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F364e6c53-caad-4cd9-820f-b1e97646a237_1700x953.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!pfEp!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F364e6c53-caad-4cd9-820f-b1e97646a237_1700x953.png" width="1456" height="816" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/364e6c53-caad-4cd9-820f-b1e97646a237_1700x953.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:816,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Which core runs each collective.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Which core runs each collective." title="Which core runs each collective." srcset="https://substackcdn.com/image/fetch/$s_!pfEp!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F364e6c53-caad-4cd9-820f-b1e97646a237_1700x953.png 424w, https://substackcdn.com/image/fetch/$s_!pfEp!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F364e6c53-caad-4cd9-820f-b1e97646a237_1700x953.png 848w, https://substackcdn.com/image/fetch/$s_!pfEp!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F364e6c53-caad-4cd9-820f-b1e97646a237_1700x953.png 1272w, https://substackcdn.com/image/fetch/$s_!pfEp!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F364e6c53-caad-4cd9-820f-b1e97646a237_1700x953.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Which core runs each collective.</strong> Engine named in the optimized HLO for each collective at 1 MiB per device, on slices of each generation; the same pattern held at 64 KiB and 16 MiB. The ragged all-to-all was compiled once per slice with 512 rows of 1,024 bf16 values per device. Two-device collectives on Ironwood, within a chip or between neighboring chips, stay on the TensorCore.</figcaption></figure></div><p>We compiled all-reduce, all-gather, reduce-scatter and the dense all-to-all at 64 KiB, 1 MiB and 16 MiB per device on Ironwood slices of 4, 8 and 64 chips and on v5e, v5p and Trillium slices, plus the ragged <strong>all-to-all that JAX uses</strong> for mixture-of-experts dispatch. </p><p>On Ironwood, all-reduce, all-gather and reduce-scatter were offloaded to the SparseCores in every case that compiled, and so was the ragged all-to-all, on 8, 16 and 128 devices. <strong>The dense all-to-all never was, on any generation.</strong> <strong>On v5e, v5p and Trillium nothing was offloaded, although v5p has four SparseCores per chip.</strong> </p><p>Their collectives run on the TensorCore, with a short-message all-reduce emitter and a ring emitter above it; on the 2- to 4-device slices we bisected, the switch happens <em>just under 1 MiB per device</em>, and around 2 MiB on v5p. A two-device collective on Ironwood, whether between the two cores of one chip or between neighboring chips, also stays on the TensorCore.</p><p>The library behind these choices is large. libtpu&#8217;s strings name 146 distinct identifiers ending in Emitter or Strategy, among them a <code>D2DUniDirRingStrategy</code> for the die-to-die hop inside a chip, an <code>InferenceShortRingSumEmitter</code>, a <code>RaggedAllToAllEmitter</code> and a <code>RotatedPincerQuantizedEmitter</code>, whose name suggests collectives on quantized data. <em>Our compiles used a handful of them</em>, and each choice is recorded in the HLO.</p><p>The switches are in the compilation environment. <code>xla_tpu_enable_sparse_core_collective_offload_all_reduce</code>, <code>_all_gather</code> and <code>_reduce_scatter</code> are set to AUTO, and <code>xla_tpu_sparse_core_all_reduce_offload_min_size_in_bytes</code> is 65,536, so very small all-reduces stay where they are.</p><p>Read that next to Google&#8217;s description of TPU 8i. Each 8i chip has two TensorCores on core dies and one <em>Collectives Acceleration Engine</em> on the chiplet die, &#8220;<em>replacing four SparseCores</em>&#8221; of Ironwood, with <strong>five times lower on-chip collective latency</strong>, plus a new topology, <em>Boardfly</em>, that cuts the worst path across a 1,024-chip pod from 16 hops to 7 to speed up all-to-all traffic. </p><p><strong>On Ironwood the SparseCores are already the TPU&#8217;s collectives engine, MoE dispatch included.</strong> TPU 8i replaces them with hardware built only for that job and attacks the remaining cost, hop count, through the network. TPU 8t keeps its SparseCores, and Google describes them as offloading data-dependent all-gathers and other collectives too.</p><p>For serving, the reason to care is tensor parallelism. A Megatron-style transformer layer performs<strong> two all-reduces per layer in decode</strong>, each of batch times hidden times two bytes: 1 MiB at batch 64 and a hidden size of 8,192, right where the older chips switch emitters. </p><p>On Ironwood that traffic, and the MoE dispatch next to it, leaves the TensorCore, which can keep multiplying while it moves.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>VMEM and TPU 8i&#8217;s 384 MB</h2><p><strong>VMEM is the TPU&#8217;s software-managed on-chip memory</strong>, where a Pallas kernel stages its tiles. The compiler enforces its size and says so when you ask for too much. </p><p>We compiled a <strong>Pallas kernel</strong> with an ever larger scratch buffer until it failed, which on Ironwood reads &#8220;<em>Allocation (size=68157440) would exceed memory (size=67108864)</em>&#8221;. The limits per device: v4 16 MiB, v5e 128 MiB, v5p 64 MiB, Trillium 128 MiB, Ironwood 64 MiB per core. </p><p>Per chip that is <strong>32 MiB on v4 and 128 MiB on v5e</strong>, v5p, Trillium and Ironwood, which matches the retrospective&#8217;s table for the training chips; on v5p a megacore program is still limited to one core&#8217;s 64 MiB. </p><p>XLA&#8217;s own fusions stay well below that: on Ironwood the large matmul fusions carried a 32 MiB scoped budget, and the largest fusion allocations we saw were about 15 MB on v4, v5e and v5p and up to 33 MB on <strong>Trillium and Ironwood</strong>. A kernel that wants the rest has to ask for it.</p><p><strong>The same trick gives HBM. </strong>Compile something larger than memory and the error reports the space available per device: 15.25 GiB on v4, 15.75 on v5e, 95.73 on v5p, 31.24 on Trillium and 94.74 per Ironwood core.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!WgEL!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe4ae0c4-2f03-4a11-af5a-6bcd613ad02f_1911x899.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!WgEL!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe4ae0c4-2f03-4a11-af5a-6bcd613ad02f_1911x899.png 424w, https://substackcdn.com/image/fetch/$s_!WgEL!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe4ae0c4-2f03-4a11-af5a-6bcd613ad02f_1911x899.png 848w, https://substackcdn.com/image/fetch/$s_!WgEL!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe4ae0c4-2f03-4a11-af5a-6bcd613ad02f_1911x899.png 1272w, https://substackcdn.com/image/fetch/$s_!WgEL!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe4ae0c4-2f03-4a11-af5a-6bcd613ad02f_1911x899.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!WgEL!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe4ae0c4-2f03-4a11-af5a-6bcd613ad02f_1911x899.png" width="1456" height="685" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/fe4ae0c4-2f03-4a11-af5a-6bcd613ad02f_1911x899.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:685,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;On-chip memory per chip, and how much KV cache it could hold.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="On-chip memory per chip, and how much KV cache it could hold." title="On-chip memory per chip, and how much KV cache it could hold." srcset="https://substackcdn.com/image/fetch/$s_!WgEL!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe4ae0c4-2f03-4a11-af5a-6bcd613ad02f_1911x899.png 424w, https://substackcdn.com/image/fetch/$s_!WgEL!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe4ae0c4-2f03-4a11-af5a-6bcd613ad02f_1911x899.png 848w, https://substackcdn.com/image/fetch/$s_!WgEL!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe4ae0c4-2f03-4a11-af5a-6bcd613ad02f_1911x899.png 1272w, https://substackcdn.com/image/fetch/$s_!WgEL!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe4ae0c4-2f03-4a11-af5a-6bcd613ad02f_1911x899.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>On-chip memory per chip, and how much KV cache it could hold.</strong> VMEM measured per device and multiplied by TensorCores per chip; TPU 8t and 8i from Google&#8217;s April figures. Per-chip VMEM has been 128 MiB since v5e and v5p. Token counts assume the whole SRAM holds FP8 KV for a 70B-class model with grouped-query attention, 80 layers and 8 KV heads of 128, at 163,840 bytes per token. Nothing else would fit.</figcaption></figure></div><p><strong>On-chip memory per chip has been 128 MiB on every TPU since v5e and v5p.</strong> The retrospective&#8217;s explanation is that SRAM density grew more slowly than logic density, so VMEM only quadrupled from TPU v2 to Ironwood while the die grew much more. </p><p>TPU 8i is the first chip in this lineage to break the plateau, with 384 MB of on-chip SRAM, three times Ironwood, and Google says it can host a larger KV cache entirely on silicon. It is worth putting a number on larger. </p><p>A 70B-class model with grouped-query attention, 80 layers and 8 KV heads of 128, stores 163,840 bytes of KV per token in FP8. <strong>384 MB holds about 2,300 of those tokens, and only if nothing else lived in VMEM.</strong> Ironwood&#8217;s 128 MiB holds about 820. A model with DeepSeek&#8217;s compressed latent attention, 576 values per token per layer over 61 layers, needs 35,136 bytes per token, so 8i holds about 10,900 tokens of it.</p><p>Per chip that is one modest conversation. Per pod it is another matter: 1,152 chips carry 442 GB of SRAM, about 2.7 million tokens of the 70B-class cache. <strong>On-chip KV is a pod-scale idea.</strong> </p><p>It works when the cache is sharded across hundreds of chips that can reach each other quickly, which is what the collectives engine and Boardfly are for. <strong>The SRAM, the CAE and the topology</strong> only make sense together.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>The bill: TPU vs GPU, per byte and per FLOP</h2><p>As of September, Google&#8217;s pricing page lists Ironwood at $12.00 per chip-hour on demand in Iowa, $8.40 on a one-year commitment and $5.40 on three years. </p><p><strong>Trillium lists at $2.70 on demand and $1.22 on three years</strong>, v5p at $4.20 and $1.89, v5e at $1.20 and $0.54. SemiAnalysis estimates that Anthropic pays about $1.60 per Ironwood chip-hour on Google Cloud. For GPUs on the same cloud, a September snapshot of public prices put an H100 at $5.38 and a B200 at $11.28 per GPU-hour on demand. </p><p>The transacted neocloud market, as measured by Ornn&#8217;s index on 28 September, was $2.56 for an H100 and $8.11 for a B200.</p><p><strong>Every hourly rate hides two prices.</strong> One is what you pay per unit of HBM bandwidth, which is what a decode step below the critical batch consumes. The other is what you pay per unit of compute, which is what you consume above it.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!W83c!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F27911563-a7b4-4522-b454-e4f9e61738fe_1997x1037.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!W83c!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F27911563-a7b4-4522-b454-e4f9e61738fe_1997x1037.png 424w, https://substackcdn.com/image/fetch/$s_!W83c!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F27911563-a7b4-4522-b454-e4f9e61738fe_1997x1037.png 848w, https://substackcdn.com/image/fetch/$s_!W83c!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F27911563-a7b4-4522-b454-e4f9e61738fe_1997x1037.png 1272w, https://substackcdn.com/image/fetch/$s_!W83c!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F27911563-a7b4-4522-b454-e4f9e61738fe_1997x1037.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!W83c!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F27911563-a7b4-4522-b454-e4f9e61738fe_1997x1037.png" width="1456" height="756" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/27911563-a7b4-4522-b454-e4f9e61738fe_1997x1037.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:756,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Two prices for every chip.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Two prices for every chip." title="Two prices for every chip." srcset="https://substackcdn.com/image/fetch/$s_!W83c!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F27911563-a7b4-4522-b454-e4f9e61738fe_1997x1037.png 424w, https://substackcdn.com/image/fetch/$s_!W83c!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F27911563-a7b4-4522-b454-e4f9e61738fe_1997x1037.png 848w, https://substackcdn.com/image/fetch/$s_!W83c!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F27911563-a7b4-4522-b454-e4f9e61738fe_1997x1037.png 1272w, https://substackcdn.com/image/fetch/$s_!W83c!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F27911563-a7b4-4522-b454-e4f9e61738fe_1997x1037.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Two prices for every chip.</strong> Hourly price divided by HBM bandwidth (horizontal) and by dense 8-bit compute on each chip&#8217;s native 8-bit path (vertical): int8 on v5e (393 TOPS), v5p (918) and Trillium (1,836), FP8 on Ironwood (4,614), H100 (1,979) and B200 (4,500). Filled points are on-demand prices; hollow points are three-year commitments, the market index and the analyst estimate for Anthropic. The shaded band spans $1.40 to $1.65 per TB/s-hour.</figcaption></figure></div><p>On demand, the first price is almost the same for all six accelerators we priced: $1.40 per TB/s-hour for v5e, $1.41 for the B200, $1.52 for v5p, $1.61 for the H100, $1.63 for Ironwood and $1.65 for Trillium. </p><p>The second ranges more than threefold, from $1.47 per dense 8-bit PFLOP/s-hour on Trillium to $4.58 on v5p. <strong>Whether by design or by convergence, Google&#8217;s on-demand rate card prices bandwidth and lets FLOPs fall where they may.</strong></p><p>That has a direct consequence for tokens. Take a 70B-class dense model with 8-bit weights and short contexts, and assume a perfect roofline: no KV traffic, no communication, no idle time, and per-chip throughput that doesn&#8217;t depend on how many chips the model is split across. <em>These are floors, not forecasts.</em></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!bEV3!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb24bd74f-ec46-4ff8-8d9d-25f68c17b681_1928x1246.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!bEV3!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb24bd74f-ec46-4ff8-8d9d-25f68c17b681_1928x1246.png 424w, https://substackcdn.com/image/fetch/$s_!bEV3!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb24bd74f-ec46-4ff8-8d9d-25f68c17b681_1928x1246.png 848w, https://substackcdn.com/image/fetch/$s_!bEV3!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb24bd74f-ec46-4ff8-8d9d-25f68c17b681_1928x1246.png 1272w, https://substackcdn.com/image/fetch/$s_!bEV3!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb24bd74f-ec46-4ff8-8d9d-25f68c17b681_1928x1246.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!bEV3!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb24bd74f-ec46-4ff8-8d9d-25f68c17b681_1928x1246.png" width="1456" height="941" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b24bd74f-ec46-4ff8-8d9d-25f68c17b681_1928x1246.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:941,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Price per chip-hour and the floor cost of a 70B-class model's output tokens.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Price per chip-hour and the floor cost of a 70B-class model's output tokens." title="Price per chip-hour and the floor cost of a 70B-class model's output tokens." srcset="https://substackcdn.com/image/fetch/$s_!bEV3!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb24bd74f-ec46-4ff8-8d9d-25f68c17b681_1928x1246.png 424w, https://substackcdn.com/image/fetch/$s_!bEV3!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb24bd74f-ec46-4ff8-8d9d-25f68c17b681_1928x1246.png 848w, https://substackcdn.com/image/fetch/$s_!bEV3!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb24bd74f-ec46-4ff8-8d9d-25f68c17b681_1928x1246.png 1272w, https://substackcdn.com/image/fetch/$s_!bEV3!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb24bd74f-ec46-4ff8-8d9d-25f68c17b681_1928x1246.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Price per chip-hour and the floor cost of a 70B-class model&#8217;s output tokens.</strong> Per-chip throughput is 64 times HBM bandwidth over 70 GB of weights at batch 64, and dense 8-bit FLOP/s over 140 GFLOP per token at the critical batch. Real systems reach a fraction of either. Sources for prices in the dossier.</figcaption></figure></div><p>At batch 64 every chip in the table is bandwidth-bound, and the on-demand floor per million output tokens lands between $0.42 and $0.50 on all of them. <strong>The chip doesn&#8217;t matter, because the rate card has already priced the bandwidth.</strong> </p><p>At each chip&#8217;s critical batch the floor falls to between $0.057 on Trillium and $0.18 on v5p, and now the chip matters a great deal. <strong>Trillium is the cheapest FLOP Google sells, as long as you can keep 560 sequences per chip in flight.</strong> </p><p>Ironwood&#8217;s floor at its critical batch is $0.10 on demand, $0.045 on a three-year commitment, and $0.013 at the estimated Anthropic rate, about a quarter of the transacted H100 market&#8217;s $0.050 and a fifth of the B200&#8217;s $0.070.</p><p>Long contexts push every chip back toward the bandwidth price, whatever the batch. With an<strong> FP8 KV cache, the same model reads 163,840 bytes per token </strong>for every position of context, so at 32K tokens each generated token reads 5.4 GB of cache no matter how many sequences share the weights. </p><p>That caps utilization at 4.2% of peak on Ironwood, 2.3% on Trillium, 4.4% on an H100 and 4.6% on a B200, and it puts the on-demand floor at between $2.08 and $2.46 per million output tokens on every chip in the table, about five times the short-context floor at batch 64. </p><p>For agentic workloads with long contexts, <strong>the rate card&#8217;s bandwidth price is the price of a token</strong>, which is why so much attention research in the last year has been about reading fewer bytes of cache.</p><p>The comparison at the critical batch is the economic story of the TPU in 2026. At list price Ironwood is priced like a B200 on Google Cloud, $12.00 against $11.28, and the same rate card puts every chip near the same bandwidth price. </p><p>At the price a tenant with a gigawatt-scale contract reportedly pays, <strong>Ironwood is the cheapest accelerator in this piece on both axes, by 1.9 to 2.9 times over the next cheapest option</strong>. The scale of those contracts is public now. Broadcom booked $10 billion and then another $11 billion of Ironwood racks for Anthropic, describing them as complete system sales, and in September put Anthropic&#8217;s deployment at 1 GW of Ironwood in 2026, 5 GW of TPU 8i in 2027 and 10 GW in 2028. </p><p>Meta signed a multi-year TPU rental in February and is negotiating to buy chips for its own data centers from 2027.</p><p>Tenants also change what the compiler is. A chip that only Google ran could keep its codenames, its 1,184 options and its rewrites to itself. A chip rented by Anthropic and Meta has customers who will read the same dumps we read, find the same options and tune them. <strong>For now, the most detailed description of the TPU&#8217;s internals that anyone outside Google can read is a 683 MB binary.</strong></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Economic and financial implications</h2><p>Everything above is a property of a chip or of a compiler. Money turns those properties into decisions, and in 2026 the decisions are large enough to show up in the earnings of three companies. Start with what the findings are worth to a buyer, then follow the money up the stack.</p><h3>What the compiler&#8217;s findings are worth to a buyer</h3><p>Four results in this piece change the cost of a token directly, using the same floor model as the table above.</p><p><strong>Precision on Ironwood.</strong> The compiler prices an int8 matmul at 1.13 times the bf16 cycles and an FP8 matmul at 0.48 times. Above the critical batch, where compute sets the price, <strong>an int8 checkpoint costs about 2.3 times as much per token as the same model in FP8.</strong> Below it, both read one byte per weight and cost the same. The penalty lands on the high-batch traffic that pays for a deployment.</p><p><strong>The KV layout.</strong> At 32K tokens of context the cache dominates every byte a decode step reads. For a model with 64-wide heads and the same cache size per token as our reference model, a row-major cache stores and moves twice its logical size, which <strong>doubles the long-context floor, from $2.08 to $2.46 per million output tokens on demand to $4.17 to $4.92.</strong></p><p><strong>Batch one.</strong> For single-stream decode on v4 through Trillium, the multiply-and-reduce rewrite costs 2.75 to 3.50 times the cycles of the matrix-unit path in the compiler&#8217;s model. A latency-bound product that runs one sequence per replica pays that multiple on every token, unless someone changes one option.</p><p><strong>Which generation.</strong> At three-year list prices Trillium&#8217;s dense 8-bit compute costs $0.66 per PFLOP/s-hour, the cheapest in this piece, and its bandwidth $0.74 per TB/s-hour, the same as Ironwood&#8217;s $0.73. Above 560 concurrent sequences per chip Trillium&#8217;s floor is $0.026 per million tokens against Ironwood&#8217;s $0.045. </p><p>Below that, <strong>the two cost the same per byte.</strong> <em>Offline work that can keep hundreds of sequences in flight, such as evaluations, synthetic data and reinforcement-learning rollouts, is Trillium&#8217;s economic home. Interactive serving is Ironwood&#8217;s.</em></p><h3>The rate card is a formula</h3><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Ep9g!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd532b61-dfb1-4cdd-aad4-07ae238c208b_1928x770.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Ep9g!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd532b61-dfb1-4cdd-aad4-07ae238c208b_1928x770.png 424w, https://substackcdn.com/image/fetch/$s_!Ep9g!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd532b61-dfb1-4cdd-aad4-07ae238c208b_1928x770.png 848w, https://substackcdn.com/image/fetch/$s_!Ep9g!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd532b61-dfb1-4cdd-aad4-07ae238c208b_1928x770.png 1272w, https://substackcdn.com/image/fetch/$s_!Ep9g!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd532b61-dfb1-4cdd-aad4-07ae238c208b_1928x770.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Ep9g!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd532b61-dfb1-4cdd-aad4-07ae238c208b_1928x770.png" width="1456" height="581" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/dd532b61-dfb1-4cdd-aad4-07ae238c208b_1928x770.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:581,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Google's TPU rate card, per chip-hour.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Google's TPU rate card, per chip-hour." title="Google's TPU rate card, per chip-hour." srcset="https://substackcdn.com/image/fetch/$s_!Ep9g!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd532b61-dfb1-4cdd-aad4-07ae238c208b_1928x770.png 424w, https://substackcdn.com/image/fetch/$s_!Ep9g!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd532b61-dfb1-4cdd-aad4-07ae238c208b_1928x770.png 848w, https://substackcdn.com/image/fetch/$s_!Ep9g!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd532b61-dfb1-4cdd-aad4-07ae238c208b_1928x770.png 1272w, https://substackcdn.com/image/fetch/$s_!Ep9g!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd532b61-dfb1-4cdd-aad4-07ae238c208b_1928x770.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Google&#8217;s TPU rate card, per chip-hour.</strong> On-demand, one-year, three-year and Flex-start prices from Google&#8217;s pricing pages and September 2026 reports of them; the anchor rate is SemiAnalysis&#8217;s estimate for Anthropic. Every generation with a published price follows the same ratios.</figcaption></figure></div><p><strong>Google&#8217;s TPU prices follow a fixed schedule</strong>. For every generation with a published price, the one-year commitment is 70% of the on-demand rate, the three-year commitment 45% and Flex-start 50%: $0.84, $0.54 and $0.60 against $1.20 for v5e; $2.94, $1.89 and $2.10 against $4.20 for v5p; $1.89, $1.22 and $1.35 against $2.70 for Trillium; $8.40, $5.40 and $6.00 against $12.00 for Ironwood. </p><p><strong>The estimated rate for Anthropic, $1.60 per Ironwood chip-hour, is 13.3% of on-demand.</strong> It is not on the card, and it is a different kind of price: a tenant price for gigawatts rather than a list price for hours.</p><h3>What a gigawatt costs and earns</h3><p>Broadcom booked $21 billion of Ironwood racks for Anthropic and describes Anthropic&#8217;s 2026 deployment as 1 GW of Ironwood: <strong>roughly $21 billion of rack hardware per gigawatt</strong>, before buildings, power and the network outside the racks. </p><p>On the revenue side, one Ironwood chip rented for 90% of the hours in a year earns $12,600 at the estimated anchor rate, $42,600 at the three-year list rate, $66,200 at the one-year rate and $94,600 on demand.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!YJWK!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b58e29b-b288-4b2b-994e-dc2eaf2ccdf9_1833x955.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!YJWK!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b58e29b-b288-4b2b-994e-dc2eaf2ccdf9_1833x955.png 424w, https://substackcdn.com/image/fetch/$s_!YJWK!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b58e29b-b288-4b2b-994e-dc2eaf2ccdf9_1833x955.png 848w, https://substackcdn.com/image/fetch/$s_!YJWK!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b58e29b-b288-4b2b-994e-dc2eaf2ccdf9_1833x955.png 1272w, https://substackcdn.com/image/fetch/$s_!YJWK!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b58e29b-b288-4b2b-994e-dc2eaf2ccdf9_1833x955.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!YJWK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b58e29b-b288-4b2b-994e-dc2eaf2ccdf9_1833x955.png" width="1456" height="759" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0b58e29b-b288-4b2b-994e-dc2eaf2ccdf9_1833x955.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:759,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;What an Ironwood chip earns, and how much power it can draw for the racks to pay back.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="What an Ironwood chip earns, and how much power it can draw for the racks to pay back." title="What an Ironwood chip earns, and how much power it can draw for the racks to pay back." srcset="https://substackcdn.com/image/fetch/$s_!YJWK!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b58e29b-b288-4b2b-994e-dc2eaf2ccdf9_1833x955.png 424w, https://substackcdn.com/image/fetch/$s_!YJWK!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b58e29b-b288-4b2b-994e-dc2eaf2ccdf9_1833x955.png 848w, https://substackcdn.com/image/fetch/$s_!YJWK!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b58e29b-b288-4b2b-994e-dc2eaf2ccdf9_1833x955.png 1272w, https://substackcdn.com/image/fetch/$s_!YJWK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b58e29b-b288-4b2b-994e-dc2eaf2ccdf9_1833x955.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>What an Ironwood chip earns, and how much power it can draw for the racks to pay back.</strong> Left: revenue per chip-year at 90% utilization at each price. Right: the largest all-in power per chip, meaning chip plus its share of host, network and cooling, at which that revenue repays $21 billion of racks per gigawatt within four years. Energy and buildings are excluded, so these are ceilings.</figcaption></figure></div><p>Google doesn&#8217;t publish Ironwood&#8217;s power per chip, so the two numbers can&#8217;t be joined directly. They can be joined conditionally. For the anchor rate to repay $21 billion of racks in four years at 90% utilization, a gigawatt has to hold about 416,000 chips, which means each chip, with its share of host, network and cooling, can draw no more than about 2.4 kW. </p><p>At the three-year list rate the same payback allows 8.1 kW per chip, and on demand 18 kW. <strong>At any published price the racks pay for themselves in well under four years. At the anchor price the answer depends on watts per chip, a number Google hasn&#8217;t published.</strong> Energy and buildings come on top of all of these.</p><h3>Who captures the margin</h3><p>A<strong> GPU buyer pays NVIDIA&#8217;s gross margin: 75.0%</strong> in the quarter to July, guided to 74.0% for the next and, according to the company&#8217;s call, heading for a trough of 71% to 72% in its fiscal fourth quarter as memory costs rise. </p><p>A TPU buyer pays Google, which pays Broadcom, whose company-wide gross margin was also 75% in its quarter to August but fell 210 basis points in a single quarter because AI chips, 73% of them custom accelerators, made up more of its revenue. </p><p><strong>The gap between those two margin stacks is the room Google has to price anchor tenants far below its own list.</strong> It is also how the B200 rental index could rise 79% in a quarter while Ironwood capacity went to Anthropic, by SemiAnalysis&#8217;s estimate, at about a fifth of the B200&#8217;s market rate.</p><h3>Alphabet</h3><p>Alphabet&#8217;s filings now describe the TPU as something it sells. Its first-quarter 10-Q says Google Cloud has agreements to supply multiple gigawatts of TPU hardware to customers who run their own infrastructure, with the revenue counted in the Cloud backlog. </p><p>That backlog reached <strong>$514 billion</strong> at the end of June, up more than $50 billion in a quarter. Google Cloud revenue grew 82% to $24.8 billion in the second quarter, with a 35.6% operating margin, and management said growth accelerated <em>even after excluding TPU system sales</em>.</p><p><strong>The bill is on the other side of the ledger</strong>. Alphabet raised its 2026 capital-expenditure guidance to $195 billion to $205 billion, spent $44.9 billion in the second quarter alone and reported negative free cash flow of $5.9 billion for the quarter. </p><p>On 1 June it announced <strong>up to $70 billion of new equity</strong>: $15 billion of mandatory convertible preferred stock, $15 billion of common stock and a $40 billion at-the-market program, to fund AI infrastructure. It has also reportedly agreed a joint venture with a large investment firm to lease TPUs to other customers. </p><p>A company that funds part of its build-out with equity has to care about the <strong>utilization and the price of every chip</strong>, which is the trade-off an 87% discount to list makes visible.</p><h3>Broadcom</h3><p>Broadcom&#8217;s AI semiconductor revenue was $16.7 billion in its quarter to August, up 221% and 56% of total revenue, with custom accelerators 73% of it. </p><p>It shipped Ironwood in high volume to Anthropic and to Google, began production shipments of TPU 8i for Google, and guided AI revenue to $58 billion for fiscal 2026, <strong>$115 billion for fiscal 2027 and $230 billion for fiscal 2028</strong>. On the same call it gave Anthropic&#8217;s roadmap: <strong>1 GW of Ironwood in 2026, 5 GW of TPU 8i in 2027 and 10 GW in 2028.</strong></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!ewWg!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a97c992-797c-448d-92bb-be51cfc3e1ee_1810x966.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!ewWg!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a97c992-797c-448d-92bb-be51cfc3e1ee_1810x966.png 424w, https://substackcdn.com/image/fetch/$s_!ewWg!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a97c992-797c-448d-92bb-be51cfc3e1ee_1810x966.png 848w, https://substackcdn.com/image/fetch/$s_!ewWg!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a97c992-797c-448d-92bb-be51cfc3e1ee_1810x966.png 1272w, https://substackcdn.com/image/fetch/$s_!ewWg!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a97c992-797c-448d-92bb-be51cfc3e1ee_1810x966.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!ewWg!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a97c992-797c-448d-92bb-be51cfc3e1ee_1810x966.png" width="1456" height="777" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/7a97c992-797c-448d-92bb-be51cfc3e1ee_1810x966.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:777,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Anthropic's TPU roadmap and Broadcom's AI revenue guidance.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Anthropic's TPU roadmap and Broadcom's AI revenue guidance." title="Anthropic's TPU roadmap and Broadcom's AI revenue guidance." srcset="https://substackcdn.com/image/fetch/$s_!ewWg!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a97c992-797c-448d-92bb-be51cfc3e1ee_1810x966.png 424w, https://substackcdn.com/image/fetch/$s_!ewWg!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a97c992-797c-448d-92bb-be51cfc3e1ee_1810x966.png 848w, https://substackcdn.com/image/fetch/$s_!ewWg!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a97c992-797c-448d-92bb-be51cfc3e1ee_1810x966.png 1272w, https://substackcdn.com/image/fetch/$s_!ewWg!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a97c992-797c-448d-92bb-be51cfc3e1ee_1810x966.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Anthropic&#8217;s TPU roadmap and Broadcom&#8217;s AI revenue guidance.</strong> Gigawatts and fiscal-year AI revenue guidance as stated on Broadcom&#8217;s September 2026 call. The hatched bar applies Ironwood&#8217;s roughly $21 billion of racks per gigawatt to 5 GW of TPU 8i; it is arithmetic, not a forecast.</figcaption></figure></div><p><strong>Put the two together and something has to give.</strong> If TPU 8i racks cost what Ironwood racks cost per gigawatt, Anthropic&#8217;s 2027 deployment alone would be about $105 billion of hardware, close to Broadcom&#8217;s entire $115 billion fiscal-2027 AI guidance, which also has to cover Google&#8217;s own TPUs, other custom-chip customers and networking. </p><p><strong>Either TPU 8i is much cheaper per gigawatt than Ironwood, or a large share of the value flows through Google&#8217;s TPU system sales rather than through Broadcom&#8217;s revenue, or both.</strong> </p><p>Reports conflict on which partner designs which eighth-generation chip, and Google hasn&#8217;t confirmed any of them; Broadcom&#8217;s own call places TPU 8i in its shipments.</p><h3>MediaTek and the inputs</h3><p>MediaTek is Google&#8217;s second TPU partner. It reportedly holds orders for Google&#8217;s v7e and v8e, asked TSMC for a sevenfold increase in CoWoS packaging capacity for its Google projects by 2027, and expects more than $1 billion of ASIC revenue in 2026 and several billion in 2027. </p><p>Digitimes reports that its TPU orders are now limited by capacity. <strong>The binding inputs for every TPU</strong>, as for every GPU,<strong> are advanced packaging and HBM.</strong> NVIDIA attributes its margin step-down to memory costs, and each new TPU carries more HBM than the last: 192 GiB on Ironwood, 216 GB on TPU 8t and 288 GB on TPU 8i. </p><p>A rate card that is flat in dollars per unit of bandwidth is consistent with HBM being the input that sets the cost.</p><h3>Nvidia</h3><p><em>NVIDIA&#8217;s own numbers show no damage yet</em>: revenue of $96.2 billion in the quarter to July, up 106%, data-center revenue of $89.0 billion, guidance of $108 billion for the next quarter and about 70% growth for fiscal 2028. </p><p>Demand still exceeds supply. What TPUs change is the price at the margin. The largest buyers outside Google now have a second supplier: Anthropic is scaling toward 10 GW of TPUs by 2028, Meta rents TPUs alongside its NVIDIA and AMD contracts, and Google sells TPU systems for customers&#8217; own buildings. <strong>At list price Google matches NVIDIA&#8217;s cost per byte. In contracts it doesn&#8217;t have to.</strong></p><h3>Compute markets</h3><p>GPU rental indices, from Ornn&#8217;s transacted index to Silicon Data&#8217;s H100 index, price GPU-hours. <strong>TPU hours have no index</strong>, and comparing a TPU-hour with a GPU-hour is a category error unless you normalize. </p><p>The normalization this piece suggests is bandwidth for inference below the critical batch and dense compute above it. On that basis Ironwood&#8217;s list price is a B200&#8217;s, its three-year price is below the B200 market and its anchor price has no GPU equivalent. </p><p><strong>A market that wants to hedge inference capacity across vendors would need to quote it per TB/s-hour, not per chip-hour.</strong></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Notes for anyone serving on TPUs</h2><ul><li><p><strong>Size the batch to the generation.</strong> Trillium needs about 560 concurrent sequences per chip before compute is the limit, Ironwood 270 to 313, v5p about 168. Below those numbers you are paying the bandwidth price and extra sequences are almost free. Quantizing weights from bf16 to FP8 does not change the number on Ironwood; FP4 weights on TPU 8i would halve it.</p></li><li><p><strong>Count devices, not chips, on Ironwood.</strong> Each chip is two devices with 94.74 GiB of usable HBM and 64 MiB of VMEM each. A sharding plan written for v5p&#8217;s megacore sees twice the devices with half the memory, and a tensor-parallel degree of two can live inside one chip.</p></li><li><p><strong>Move int8 checkpoints to FP8 before moving to Ironwood.</strong> In the compiler&#8217;s model int8 takes 1.13 times the bf16 cycles there, and FP8 0.48 times.</p></li><li><p><strong>Fix the KV layout yourself.</strong> With 64-wide heads, a row-major cache costs twice its size. Fold heads into the minor dimension, or accept the compiler&#8217;s default only if no kernel needs row-major order.</p></li><li><p><strong>Watch for batch one on v4 through Trillium.</strong> Hand-written decode loops, draft models and per-expert dots can hit the multiply-and-reduce rewrite. Pad to two rows, or test <code>xla_tpu_enable_dot_strength_reduction=false</code> on real hardware.</p></li><li><p><strong>Ask Pallas for VMEM explicitly.</strong> XLA&#8217;s fusions budget 16 to 32 MiB; the hardware has 64 MiB per Ironwood core and 128 MiB on v5e and Trillium.</p></li><li><p><strong>Expect collectives to change hands at 64 KiB on Ironwood.</strong> Above it, all-reduce, all-gather, reduce-scatter and ragged all-to-all across four or more chips run on the SparseCores while the TensorCore computes. Only the dense all-to-all stays on the TensorCore.</p></li><li><p><strong>Buy bandwidth below the critical batch and FLOPs above it.</strong> Compare chips in dollars per TB/s-hour for interactive decode and long contexts, and in dollars per PFLOP/s-hour for batch work above the critical batch. The published commitment discounts are fixed at 30% and 55% off on-demand; anything better is a contract, not a list price.</p></li></ul><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>What the compiler can&#8217;t tell us</h2><p><strong>Every time in this piece is a model output.</strong> The cycle estimates and optimal_seconds are what the compiler believes, and the compiler is optimizing against them, but real programs also pay for DMA contention, ICI congestion, host stalls and everything else a static model leaves out. </p><p>The relative results, such as batch one against batch two or int8 against bf16 on the same chip, are much safer than any absolute time.</p><p><strong>We can&#8217;t tell whether the compiler&#8217;s lower compute constants</strong> for v5p and Ironwood reflect a clock, a margin or something else, and the implied clocks depend on our reading of the slot rates and of the retrospective&#8217;s MXU count. </p><p>We can&#8217;t tell <strong>whether FP8 e5m2 really runs at the int8 rate on v5e</strong>, v5p and Trillium or is only filed that way by the cost model. The names matrix result FIFO and matrix result buffer are our reading of mrf and mrb, supported by the counts but not confirmed by Google.</p><p>Layouts in this piece are the ones the compiler chooses for a program&#8217;s parameters when nothing else constrains them, plus one forced row-major case. <strong>Inside a real model,</strong> kernels and neighboring operations constrain layouts, so the padding you pay depends on the code around the array. </p><p>Coverage is narrow in places: matmuls go up to 16,384 on a side, <strong>collective tests up to 64 chips and three message sizes</strong>, and everything uses the compiler&#8217;s default options except where we say otherwise. All of it comes from one release, libtpu 0.0.48, as of October 2026, and a later release can change any compiler decision described here.</p><p>The TPU 8 numbers come from Google&#8217;s April deep dive and from reporting on it, and the 8i FP8 rate per chip is derived from a pod total. <strong>The Ironwood list prices reach us through two reports quoting Google&#8217;s pricing page</strong>, the GPU prices on Google Cloud through a third-party snapshot, and the Anthropic rate is an analyst estimate. </p><p>GPU rental prices move weekly: SemiAnalysis&#8217;s index has the B200 up 79% in the three months to late September. <strong>Company figures come from filings and earnings calls</strong>; the gigawatt arithmetic combines numbers from different companies and periods, and the power per chip that would join cost and revenue is unpublished.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Five predictions</h2><ol><li><p><strong>By 30 June 2027</strong>, a public libtpu release on PyPI will accept a TPU 8t or TPU 8i topology name for compile-only use.</p></li><li><p>When it does, an all-reduce compiled for four or more TPU 8i chips will name an engine other than the <strong>TensorCore and the SparseCore,</strong> and the ragged all-to-all will run on the same engine.</p></li><li><p><strong>Through 31 December 2027,</strong> no public libtpu release will estimate an int8 matmul on Ironwood as cheaper than the same matmul in bf16.</p></li><li><p><strong>By 31 December 2027</strong>, vLLM&#8217;s TPU documentation will recommend FP8 rather than int8 weights for serving on Ironwood.</p></li><li><p>When TPU 8i gets a<strong> public on-demand price</strong>, it will be within 15% of $1.63 per TB/s-hour of HBM bandwidth, which means between $11.92 and $16.12 per chip-hour.</p></li></ol><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share The Software Frontier&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share The Software Frontier</span></a></p><div><hr></div><h2>Doing this yourself</h2><p>Everything in this piece came from one install and a handful of options. The topology call and the compile are in the method section above. The rest is environment variables:</p><pre><code><code>XLA_FLAGS="--xla_dump_to=/tmp/hlo --xla_dump_hlo_as_text"                    # HLO stages, compiler options
LIBTPU_INIT_ARGS="--xla_jf_dump_to=/tmp/llo --xla_jf_dump_llo_text=true"      # every LLO stage, final bundles
LIBTPU_INIT_ARGS="--xla_tpu_enable_dot_strength_reduction=false"             # keep batch-1 dots on the MXU</code></code></pre><p>libtpu prints warnings about TPU worker hostnames and accelerator types when no TPU is attached. For compile-only use they are harmless. The VMEM probe is a Pallas kernel with a <code>pltpu.VMEM</code> scratch shape that grows until the compile fails, and the HBM probe is any program whose temporaries exceed memory. </p><p>The chip enum comes from the protocol-buffer descriptor inside libtpu.so: find the string <code>TPU_VERSION_JELLYFISH</code>, and the byte after each name is a field tag followed by the value. </p><p>If you have access to real TPUs, the most useful thing you could do with this piece is time the batch-one option and the int8 against FP8 result on Ironwood and tell us what you see. We will publish corrections with credit.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Confidence dossier</h2><p>Tier A is a primary document: Google&#8217;s documentation, Google&#8217;s blog, Google&#8217;s architects&#8217; retrospective, a company&#8217;s own statement. Tier B is reputable reporting, a published index or a figure from vendor material we did not re-read. Tier C is our arithmetic on A or B. </p><p>Tier D is our interpretation. Tier M is our own measurement with libtpu 0.0.48 and JAX 0.11.2 in compile-only mode, with no TPU attached. Every claim in the piece has a row.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!26WH!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd96b498-b9ac-4153-94ea-6f40956bac33_1900x1570.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!26WH!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd96b498-b9ac-4153-94ea-6f40956bac33_1900x1570.png 424w, https://substackcdn.com/image/fetch/$s_!26WH!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd96b498-b9ac-4153-94ea-6f40956bac33_1900x1570.png 848w, https://substackcdn.com/image/fetch/$s_!26WH!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd96b498-b9ac-4153-94ea-6f40956bac33_1900x1570.png 1272w, https://substackcdn.com/image/fetch/$s_!26WH!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd96b498-b9ac-4153-94ea-6f40956bac33_1900x1570.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!26WH!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd96b498-b9ac-4153-94ea-6f40956bac33_1900x1570.png" width="1456" height="1203" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/cd96b498-b9ac-4153-94ea-6f40956bac33_1900x1570.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1203,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Confidence dossier, part 1 of 9&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Confidence dossier, part 1 of 9" title="Confidence dossier, part 1 of 9" srcset="https://substackcdn.com/image/fetch/$s_!26WH!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd96b498-b9ac-4153-94ea-6f40956bac33_1900x1570.png 424w, https://substackcdn.com/image/fetch/$s_!26WH!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd96b498-b9ac-4153-94ea-6f40956bac33_1900x1570.png 848w, https://substackcdn.com/image/fetch/$s_!26WH!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd96b498-b9ac-4153-94ea-6f40956bac33_1900x1570.png 1272w, https://substackcdn.com/image/fetch/$s_!26WH!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd96b498-b9ac-4153-94ea-6f40956bac33_1900x1570.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Sb03!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe36cf6fa-17b8-4771-9c29-468dc203ce3e_1900x1570.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Sb03!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe36cf6fa-17b8-4771-9c29-468dc203ce3e_1900x1570.png 424w, https://substackcdn.com/image/fetch/$s_!Sb03!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe36cf6fa-17b8-4771-9c29-468dc203ce3e_1900x1570.png 848w, https://substackcdn.com/image/fetch/$s_!Sb03!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe36cf6fa-17b8-4771-9c29-468dc203ce3e_1900x1570.png 1272w, https://substackcdn.com/image/fetch/$s_!Sb03!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe36cf6fa-17b8-4771-9c29-468dc203ce3e_1900x1570.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Sb03!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe36cf6fa-17b8-4771-9c29-468dc203ce3e_1900x1570.png" width="1456" height="1203" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e36cf6fa-17b8-4771-9c29-468dc203ce3e_1900x1570.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1203,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Confidence dossier, part 2 of 9&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Confidence dossier, part 2 of 9" title="Confidence dossier, part 2 of 9" srcset="https://substackcdn.com/image/fetch/$s_!Sb03!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe36cf6fa-17b8-4771-9c29-468dc203ce3e_1900x1570.png 424w, https://substackcdn.com/image/fetch/$s_!Sb03!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe36cf6fa-17b8-4771-9c29-468dc203ce3e_1900x1570.png 848w, https://substackcdn.com/image/fetch/$s_!Sb03!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe36cf6fa-17b8-4771-9c29-468dc203ce3e_1900x1570.png 1272w, https://substackcdn.com/image/fetch/$s_!Sb03!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe36cf6fa-17b8-4771-9c29-468dc203ce3e_1900x1570.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!SaVN!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3a718435-5890-4c5c-b067-053dbdf0831c_1900x1570.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!SaVN!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3a718435-5890-4c5c-b067-053dbdf0831c_1900x1570.png 424w, https://substackcdn.com/image/fetch/$s_!SaVN!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3a718435-5890-4c5c-b067-053dbdf0831c_1900x1570.png 848w, https://substackcdn.com/image/fetch/$s_!SaVN!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3a718435-5890-4c5c-b067-053dbdf0831c_1900x1570.png 1272w, https://substackcdn.com/image/fetch/$s_!SaVN!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3a718435-5890-4c5c-b067-053dbdf0831c_1900x1570.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!SaVN!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3a718435-5890-4c5c-b067-053dbdf0831c_1900x1570.png" width="1456" height="1203" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/3a718435-5890-4c5c-b067-053dbdf0831c_1900x1570.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1203,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Confidence dossier, part 3 of 9&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Confidence dossier, part 3 of 9" title="Confidence dossier, part 3 of 9" srcset="https://substackcdn.com/image/fetch/$s_!SaVN!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3a718435-5890-4c5c-b067-053dbdf0831c_1900x1570.png 424w, https://substackcdn.com/image/fetch/$s_!SaVN!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3a718435-5890-4c5c-b067-053dbdf0831c_1900x1570.png 848w, https://substackcdn.com/image/fetch/$s_!SaVN!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3a718435-5890-4c5c-b067-053dbdf0831c_1900x1570.png 1272w, https://substackcdn.com/image/fetch/$s_!SaVN!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3a718435-5890-4c5c-b067-053dbdf0831c_1900x1570.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!PQpv!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89d5a95c-4944-47d5-8b8e-1c12b1f6da10_1900x1570.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!PQpv!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89d5a95c-4944-47d5-8b8e-1c12b1f6da10_1900x1570.png 424w, https://substackcdn.com/image/fetch/$s_!PQpv!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89d5a95c-4944-47d5-8b8e-1c12b1f6da10_1900x1570.png 848w, https://substackcdn.com/image/fetch/$s_!PQpv!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89d5a95c-4944-47d5-8b8e-1c12b1f6da10_1900x1570.png 1272w, https://substackcdn.com/image/fetch/$s_!PQpv!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89d5a95c-4944-47d5-8b8e-1c12b1f6da10_1900x1570.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!PQpv!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89d5a95c-4944-47d5-8b8e-1c12b1f6da10_1900x1570.png" width="1456" height="1203" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/89d5a95c-4944-47d5-8b8e-1c12b1f6da10_1900x1570.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1203,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Confidence dossier, part 4 of 9&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Confidence dossier, part 4 of 9" title="Confidence dossier, part 4 of 9" srcset="https://substackcdn.com/image/fetch/$s_!PQpv!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89d5a95c-4944-47d5-8b8e-1c12b1f6da10_1900x1570.png 424w, https://substackcdn.com/image/fetch/$s_!PQpv!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89d5a95c-4944-47d5-8b8e-1c12b1f6da10_1900x1570.png 848w, https://substackcdn.com/image/fetch/$s_!PQpv!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89d5a95c-4944-47d5-8b8e-1c12b1f6da10_1900x1570.png 1272w, https://substackcdn.com/image/fetch/$s_!PQpv!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89d5a95c-4944-47d5-8b8e-1c12b1f6da10_1900x1570.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!WT0Q!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a9b00e6-e9d1-4013-87cf-71915d79b1c4_1900x1570.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!WT0Q!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a9b00e6-e9d1-4013-87cf-71915d79b1c4_1900x1570.png 424w, https://substackcdn.com/image/fetch/$s_!WT0Q!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a9b00e6-e9d1-4013-87cf-71915d79b1c4_1900x1570.png 848w, https://substackcdn.com/image/fetch/$s_!WT0Q!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a9b00e6-e9d1-4013-87cf-71915d79b1c4_1900x1570.png 1272w, https://substackcdn.com/image/fetch/$s_!WT0Q!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a9b00e6-e9d1-4013-87cf-71915d79b1c4_1900x1570.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!WT0Q!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a9b00e6-e9d1-4013-87cf-71915d79b1c4_1900x1570.png" width="1456" height="1203" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/9a9b00e6-e9d1-4013-87cf-71915d79b1c4_1900x1570.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1203,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Confidence dossier, part 5 of 9&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Confidence dossier, part 5 of 9" title="Confidence dossier, part 5 of 9" srcset="https://substackcdn.com/image/fetch/$s_!WT0Q!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a9b00e6-e9d1-4013-87cf-71915d79b1c4_1900x1570.png 424w, https://substackcdn.com/image/fetch/$s_!WT0Q!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a9b00e6-e9d1-4013-87cf-71915d79b1c4_1900x1570.png 848w, https://substackcdn.com/image/fetch/$s_!WT0Q!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a9b00e6-e9d1-4013-87cf-71915d79b1c4_1900x1570.png 1272w, https://substackcdn.com/image/fetch/$s_!WT0Q!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a9b00e6-e9d1-4013-87cf-71915d79b1c4_1900x1570.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!PZBB!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe1607cb5-79bc-479f-becc-3a5f2a16fb11_1900x1570.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!PZBB!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe1607cb5-79bc-479f-becc-3a5f2a16fb11_1900x1570.png 424w, https://substackcdn.com/image/fetch/$s_!PZBB!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe1607cb5-79bc-479f-becc-3a5f2a16fb11_1900x1570.png 848w, https://substackcdn.com/image/fetch/$s_!PZBB!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe1607cb5-79bc-479f-becc-3a5f2a16fb11_1900x1570.png 1272w, https://substackcdn.com/image/fetch/$s_!PZBB!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe1607cb5-79bc-479f-becc-3a5f2a16fb11_1900x1570.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!PZBB!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe1607cb5-79bc-479f-becc-3a5f2a16fb11_1900x1570.png" width="1456" height="1203" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e1607cb5-79bc-479f-becc-3a5f2a16fb11_1900x1570.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1203,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Confidence dossier, part 6 of 9&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Confidence dossier, part 6 of 9" title="Confidence dossier, part 6 of 9" srcset="https://substackcdn.com/image/fetch/$s_!PZBB!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe1607cb5-79bc-479f-becc-3a5f2a16fb11_1900x1570.png 424w, https://substackcdn.com/image/fetch/$s_!PZBB!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe1607cb5-79bc-479f-becc-3a5f2a16fb11_1900x1570.png 848w, https://substackcdn.com/image/fetch/$s_!PZBB!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe1607cb5-79bc-479f-becc-3a5f2a16fb11_1900x1570.png 1272w, https://substackcdn.com/image/fetch/$s_!PZBB!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe1607cb5-79bc-479f-becc-3a5f2a16fb11_1900x1570.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!75On!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbe49a97b-d69e-47e8-875e-45c185a9a485_1900x1570.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!75On!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbe49a97b-d69e-47e8-875e-45c185a9a485_1900x1570.png 424w, https://substackcdn.com/image/fetch/$s_!75On!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbe49a97b-d69e-47e8-875e-45c185a9a485_1900x1570.png 848w, https://substackcdn.com/image/fetch/$s_!75On!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbe49a97b-d69e-47e8-875e-45c185a9a485_1900x1570.png 1272w, https://substackcdn.com/image/fetch/$s_!75On!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbe49a97b-d69e-47e8-875e-45c185a9a485_1900x1570.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!75On!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbe49a97b-d69e-47e8-875e-45c185a9a485_1900x1570.png" width="1456" height="1203" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/be49a97b-d69e-47e8-875e-45c185a9a485_1900x1570.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1203,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Confidence dossier, part 7 of 9&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Confidence dossier, part 7 of 9" title="Confidence dossier, part 7 of 9" srcset="https://substackcdn.com/image/fetch/$s_!75On!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbe49a97b-d69e-47e8-875e-45c185a9a485_1900x1570.png 424w, https://substackcdn.com/image/fetch/$s_!75On!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbe49a97b-d69e-47e8-875e-45c185a9a485_1900x1570.png 848w, https://substackcdn.com/image/fetch/$s_!75On!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbe49a97b-d69e-47e8-875e-45c185a9a485_1900x1570.png 1272w, https://substackcdn.com/image/fetch/$s_!75On!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbe49a97b-d69e-47e8-875e-45c185a9a485_1900x1570.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Z-Di!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc197ade7-e4fb-4bdc-9b74-e57244ce8df7_1900x1570.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Z-Di!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc197ade7-e4fb-4bdc-9b74-e57244ce8df7_1900x1570.png 424w, https://substackcdn.com/image/fetch/$s_!Z-Di!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc197ade7-e4fb-4bdc-9b74-e57244ce8df7_1900x1570.png 848w, https://substackcdn.com/image/fetch/$s_!Z-Di!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc197ade7-e4fb-4bdc-9b74-e57244ce8df7_1900x1570.png 1272w, https://substackcdn.com/image/fetch/$s_!Z-Di!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc197ade7-e4fb-4bdc-9b74-e57244ce8df7_1900x1570.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Z-Di!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc197ade7-e4fb-4bdc-9b74-e57244ce8df7_1900x1570.png" width="1456" height="1203" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c197ade7-e4fb-4bdc-9b74-e57244ce8df7_1900x1570.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1203,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Confidence dossier, part 8 of 9&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Confidence dossier, part 8 of 9" title="Confidence dossier, part 8 of 9" srcset="https://substackcdn.com/image/fetch/$s_!Z-Di!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc197ade7-e4fb-4bdc-9b74-e57244ce8df7_1900x1570.png 424w, https://substackcdn.com/image/fetch/$s_!Z-Di!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc197ade7-e4fb-4bdc-9b74-e57244ce8df7_1900x1570.png 848w, https://substackcdn.com/image/fetch/$s_!Z-Di!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc197ade7-e4fb-4bdc-9b74-e57244ce8df7_1900x1570.png 1272w, https://substackcdn.com/image/fetch/$s_!Z-Di!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc197ade7-e4fb-4bdc-9b74-e57244ce8df7_1900x1570.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!qas2!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8600193e-3e89-4cd0-9524-44efdfe1b57e_1900x1026.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!qas2!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8600193e-3e89-4cd0-9524-44efdfe1b57e_1900x1026.png 424w, https://substackcdn.com/image/fetch/$s_!qas2!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8600193e-3e89-4cd0-9524-44efdfe1b57e_1900x1026.png 848w, https://substackcdn.com/image/fetch/$s_!qas2!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8600193e-3e89-4cd0-9524-44efdfe1b57e_1900x1026.png 1272w, https://substackcdn.com/image/fetch/$s_!qas2!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8600193e-3e89-4cd0-9524-44efdfe1b57e_1900x1026.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!qas2!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8600193e-3e89-4cd0-9524-44efdfe1b57e_1900x1026.png" width="1456" height="786" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/8600193e-3e89-4cd0-9524-44efdfe1b57e_1900x1026.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:786,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Confidence dossier, part 9 of 9&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Confidence dossier, part 9 of 9" title="Confidence dossier, part 9 of 9" srcset="https://substackcdn.com/image/fetch/$s_!qas2!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8600193e-3e89-4cd0-9524-44efdfe1b57e_1900x1026.png 424w, https://substackcdn.com/image/fetch/$s_!qas2!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8600193e-3e89-4cd0-9524-44efdfe1b57e_1900x1026.png 848w, https://substackcdn.com/image/fetch/$s_!qas2!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8600193e-3e89-4cd0-9524-44efdfe1b57e_1900x1026.png 1272w, https://substackcdn.com/image/fetch/$s_!qas2!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8600193e-3e89-4cd0-9524-44efdfe1b57e_1900x1026.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Sources</h2><ol><li><p><em>N. P. Jouppi, S. Lakshmanamurthy, C. Young and D. Patterson, &#8220;Google&#8217;s Training Supercomputers from TPU v2 to Ironwood,&#8221; to appear in IEEE Micro, July/August 2026. <a href="https://arxiv.org/abs/2606.15870">arxiv.org/abs/2606.15870</a></em></p></li><li><p><em>Google Cloud, &#8220;TPU7x (Ironwood),&#8221; Cloud TPU documentation. <a href="https://docs.cloud.google.com/tpu/docs/tpu7x">docs.cloud.google.com/tpu/docs/tpu7x</a></em></p></li><li><p><em>Google Cloud, &#8220;TPU v5e,&#8221; Cloud TPU documentation, last updated 11 August 2026. <a href="https://docs.cloud.google.com/tpu/docs/v5e">docs.cloud.google.com/tpu/docs/v5e</a></em></p></li><li><p><em>D. Gupta and S. Mugazambi, &#8220;Inside the eighth-generation TPU: An architecture deep dive,&#8221; Google Cloud Blog, 22 April 2026. <a href="https://cloud.google.com/blog/products/compute/tpu-8t-and-tpu-8i-technical-deep-dive">cloud.google.com/blog</a></em></p></li><li><p><em>HPCwire, &#8220;Google Bolsters AI Hypercomputer with New TPU Chips, Virgo Interconnect, Speedier Lustre,&#8221; 22 April 2026. <a href="https://www.hpcwire.com/2026/04/22/google-bolsters-ai-hypercomputer-with-new-tpu-chips-virgo-interconnect-speedier-lustre/">hpcwire.com</a></em></p></li><li><p><em>DatacenterDynamics, &#8220;Google unveils eighth-generation TPUs, two dedicated training and inference chips.&#8221; <a href="https://www.datacenterdynamics.com/en/news/google-unveils-eighth-generation-tpus-two-dedicated-training-and-inference-chips/">datacenterdynamics.com</a></em></p></li><li><p><em>Moor Insights and Strategy, &#8220;Research note: Google TPU 8: Architecture, Context, and Enterprise Relevance.&#8221; <a href="https://moorinsightsstrategy.com/research-notes/google-tpu-8-architecture-context-and-enterprise-relevance/">moorinsightsstrategy.com</a></em></p></li><li><p><em>N. Jouppi et al., &#8220;TPU v4: An Optically Reconfigurable Supercomputer for Machine Learning with Hardware Support for Embeddings,&#8221; ISCA 2023. <a href="https://arxiv.org/abs/2304.01433">arxiv.org/abs/2304.01433</a></em></p></li><li><p><em>S. J. Kaufman et al., &#8220;A Learned Performance Model for Tensor Processing Units,&#8221; MLSys 2021. <a href="https://arxiv.org/abs/2008.01040">arxiv.org/abs/2008.01040</a></em></p></li><li><p><em>P. M. Phothilimthana et al., &#8220;TpuGraphs: A Performance Prediction Dataset on Large Tensor Computational Graphs,&#8221; 2023. <a href="https://arxiv.org/abs/2308.13490">arxiv.org/abs/2308.13490</a></em></p></li><li><p><em>Google Cloud, &#8220;Cloud TPU pricing.&#8221; <a href="https://cloud.google.com/tpu/pricing">cloud.google.com/tpu/pricing</a></em></p></li><li><p><em>Google Cloud, &#8220;Dynamic Workload Scheduler pricing.&#8221; <a href="https://cloud.google.com/products/dws/pricing">cloud.google.com/products/dws/pricing</a></em></p></li><li><p><em>shattered.io, &#8220;Google Puts $12/Hr Price on Ironwood TPU vs Nvidia,&#8221; September 2026. <a href="https://shattered.io/google-ironwood-tpu-pricing-vs-nvidia-2026/">shattered.io</a></em></p></li><li><p><em>Spheron, &#8220;Google TPU v7 Ironwood vs NVIDIA B200: Inference Cost (2026),&#8221; citing a SemiAnalysis estimate of Anthropic&#8217;s rate. <a href="https://www.spheron.network/blog/google-tpu-v7-ironwood-vs-nvidia-b200-inference-cost/">spheron.network</a></em></p></li><li><p><em>Ornn, Compute Price Index, settlement of 28 September 2026. <a href="https://data.ornn.com/markets">data.ornn.com/markets</a></em></p></li><li><p><em>Silicon Data, H100 Rental Price Index (SDH100RT). <a href="https://www.silicondata.com/products/silicon-index/h100">silicondata.com</a></em></p></li><li><p><em>fastgpu, &#8220;What it costs to rent an H100, B200 or RTX 4090 in September 2026,&#8221; DEV Community. <a href="https://dev.to/fastgpu/what-it-costs-to-rent-an-h100-b200-or-rtx-4090-in-september-2026-live-prices-from-28-gpu-clouds-12n3">dev.to</a></em></p></li><li><p><em>shattered.io, &#8220;Nvidia B200 Cloud Price Hits $8.01/Hr,&#8221; on the SemiAnalysis GPU pricing index, September 2026. <a href="https://shattered.io/nvidia-b200-cloud-price-79-percent-surge-2026/">shattered.io</a></em></p></li><li><p><em>SemiAnalysis, &#8220;TPUv7: Google Takes a Swing at the King.&#8221; <a href="https://newsletter.semianalysis.com/p/tpuv7-google-takes-a-swing-at-the">newsletter.semianalysis.com</a></em></p></li><li><p><em>TrendForce, &#8220;Anthropic Emerges as Broadcom&#8217;s Mega-Client; Margin Challenges Ahead,&#8221; 12 December 2025. <a href="https://www.trendforce.com/news/2025/12/12/news-anthropic-emerges-as-broadcoms-mega-client-margin-challenges-ahead/">trendforce.com</a></em></p></li><li><p><em>DatacenterDynamics, &#8220;Broadcom bulges with Anthropic AI orders.&#8221; <a href="https://www.datacenterdynamics.com/en/news/broadcom-bulges-with-anthropic-ai-orders/">datacenterdynamics.com</a></em></p></li><li><p><em>SDxCentral, coverage of Broadcom&#8217;s 2026 earnings call. <a href="https://www.sdxcentral.com/news/broadcom-gulps-ai-infused-silicon-revenues-touts-superior-supply-chain-planning/">sdxcentral.com</a></em></p></li><li><p><em>Dataconomy, &#8220;Meta signs multibillion-dollar deal to rent Google TPUs for AI training,&#8221; 27 February 2026, reporting The Information. <a href="https://dataconomy.com/2026/02/27/meta-signs-multibillion-dollar-deal-to-rent-google-tpus-for-ai-training/">dataconomy.com</a></em></p></li><li><p><em>Digitimes, &#8220;MediaTek TPU orders rise as chip capacity war intensifies,&#8221; 1 October 2026. <a href="https://digitimes.com/news/a20261001PD216/mediatek-tpu-capacity-google-revenue.html">digitimes.com</a></em></p></li><li><p><em>NPR, &#8220;Google launches Project Suncatcher, a step towards AI data centers in space,&#8221; 1 October 2026. <a href="https://npr.org/2026/10/01/nx-s1-5983697/project-suncatcher-google-ai-data-center-space">npr.org</a></em></p></li><li><p><em>J. Austin et al., &#8220;How to Scale Your Model,&#8221; Google DeepMind. <a href="https://jax-ml.github.io/scaling-book/">jax-ml.github.io/scaling-book</a></em></p></li><li><p><em>JAX documentation, &#8220;Ahead-of-time lowering and compilation.&#8221; <a href="https://docs.jax.dev/en/latest/aot.html">docs.jax.dev</a></em></p></li><li><p><em>L. Bradanini and L. Tettamanti, &#8220;How Blackwell&#8217;s Tensor Memory Actually Works,&#8221; The Software Frontier, 3 August 2026. <a href="https://www.thesoftwarefrontier.com/p/how-blackwells-tensor-memory-actually">thesoftwarefrontier.com</a></em></p></li><li><p><em>Alphabet, &#8220;Alphabet Announces Second Quarter 2026 Results,&#8221; Form 8-K exhibit 99.1. <a href="https://www.sec.gov/Archives/edgar/data/0001652044/000165204426000066/googexhibit991q22026.htm">sec.gov</a></em></p></li><li><p><em>Alphabet, Q2 2026 earnings call. <a href="https://abc.xyz/investor/events/event-details/2026/2026-Q2-Earnings-Call-2026-GgTAq7Is0z/default.aspx">abc.xyz</a></em></p></li><li><p><em>Alphabet, Form 10-Q for the quarter ended 31 March 2026. <a href="https://www.sec.gov/Archives/edgar/data/0001652044/000165204426000048/goog-20260331.htm">sec.gov</a></em></p></li><li><p><em>Alphabet, Form 8-K of 1 June 2026 on equity offerings. <a href="https://www.sec.gov/Archives/edgar/data/0001652044/000119312526257724/d83560dex991.htm">sec.gov</a></em></p></li><li><p><em>Yahoo Finance, &#8220;Alphabet Q2 2026 earnings: revenue up 24%, Cloud surges 82%.&#8221; <a href="https://finance.yahoo.com/markets/stocks/articles/alphabet-q2-2026-earnings-revenue-203058727.html">finance.yahoo.com</a></em></p></li><li><p><em>Broadcom, &#8220;Broadcom Inc. Announces Third Quarter Fiscal Year 2026 Financial Results,&#8221; 2 September 2026. <a href="https://investors.broadcom.com/news-releases/news-release-details/broadcom-inc-announces-third-quarter-fiscal-year-2026-financial">investors.broadcom.com</a></em></p></li><li><p><em>The Motley Fool, Broadcom Q3 2026 earnings call transcript. <a href="https://www.fool.com/earnings/call-transcripts/2026/09/09/broadcom-avgo-q3-2026-earnings-call-transcript/">fool.com</a></em></p></li><li><p><em>NVIDIA, &#8220;NVIDIA Announces Financial Results for Second Quarter Fiscal 2027,&#8221; 26 August 2026. <a href="https://www.sec.gov/Archives/edgar/data/0001045810/000104581026000073/q2fy27pr.htm">sec.gov</a></em></p></li><li><p><em>NVIDIA Q2 fiscal 2027 earnings call summary, on margin and fiscal 2028 outlook. <a href="https://transcripts.platformaeronaut.com/summaries/NVDA-2Q27-AI-Summary">transcripts.platformaeronaut.com</a></em></p></li><li><p><em>TrendForce, &#8220;MediaTek Reportedly Secures Google v7e, v8e TPU Orders, Requests 7-Fold CoWoS Increase from TSMC,&#8221; 15 December 2025. <a href="https://www.trendforce.com/news/2025/12/15/news-mediatek-reportedly-secures-google-v7e-v8e-tpu-orders-requests-7-fold-cowos-increase-from-tsmc/">trendforce.com</a></em></p></li><li><p><em>TrendForce, &#8220;MediaTek Forecasts $1B in ASIC Sales for 2026.&#8221; <a href="https://www.trendforce.com/news/?p=52932">trendforce.com</a></em></p></li><li><p><em>tech-insider, &#8220;Google TPU 8t and 8i: 121 Exaflops, $21B Nvidia Challenge,&#8221; with one account of the design-partner split. <a href="https://tech-insider.org/google-tpu-8t-8i-broadcom-mediatek-nvidia-2026/">tech-insider.org</a></em></p></li><li><p><em>Midas, summary of a Commercial Times report with a conflicting account of the split. <a href="https://www.getmidas.com/midasin-kulaklari/google-yeni-nesil-tpu-cipleri-icin-mediatek-ile-calisacak-p-414278">getmidas.com</a></em></p></li><li><p><em>Storyboard18, &#8220;Meta signs multi-billion dollar AI chip rental deal with Google,&#8221; reporting The Information on a TPU leasing joint venture. <a href="https://storyboard18.com/digital/meta-signs-multi-billion-dollar-ai-chip-rental-deal-with-google-90907.htm">storyboard18.com</a></em></p></li><li><p><em>pump.co, &#8220;Google Cloud TPU: Pricing, Performance and Use Cases,&#8221; on commitment prices. <a href="https://www.pump.co/blog/google-cloud-tpu/">pump.co</a></em></p></li><li><p><em>dataknobs, &#8220;Google TPU Pricing 2026,&#8221; on Ironwood Flex-start and commitment prices. <a href="https://www.dataknobs.com/generativeai/8-tpu-gpu/tpu-gpu-cost.html">dataknobs.com</a></em></p></li></ol>]]></content:encoded></item><item><title><![CDATA[A wonderful dive into how FP4 works]]></title><description><![CDATA[NVFP4 and MXFP4 multiply the same sixteen codes. The difference is one byte of scale per block, and it is worth 1.6 dB on Gaussian data, 4.2 dB on activations and about 9% of perplexity.]]></description><link>https://www.thesoftwarefrontier.com/p/a-wonderful-dive-into-how-fp4-works</link><guid isPermaLink="false">https://www.thesoftwarefrontier.com/p/a-wonderful-dive-into-how-fp4-works</guid><dc:creator><![CDATA[Lorenzo Bradanini]]></dc:creator><pubDate>Sun, 04 Oct 2026 08:21:01 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!A8Uw!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd504476-42c6-4428-b186-05722e724d1e_2600x2600.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!A8Uw!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd504476-42c6-4428-b186-05722e724d1e_2600x2600.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!A8Uw!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd504476-42c6-4428-b186-05722e724d1e_2600x2600.jpeg 424w, https://substackcdn.com/image/fetch/$s_!A8Uw!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd504476-42c6-4428-b186-05722e724d1e_2600x2600.jpeg 848w, https://substackcdn.com/image/fetch/$s_!A8Uw!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd504476-42c6-4428-b186-05722e724d1e_2600x2600.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!A8Uw!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd504476-42c6-4428-b186-05722e724d1e_2600x2600.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!A8Uw!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd504476-42c6-4428-b186-05722e724d1e_2600x2600.jpeg" width="1456" height="1456" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/cd504476-42c6-4428-b186-05722e724d1e_2600x2600.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1456,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:582521,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/210569862?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd504476-42c6-4428-b186-05722e724d1e_2600x2600.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!A8Uw!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd504476-42c6-4428-b186-05722e724d1e_2600x2600.jpeg 424w, https://substackcdn.com/image/fetch/$s_!A8Uw!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd504476-42c6-4428-b186-05722e724d1e_2600x2600.jpeg 848w, https://substackcdn.com/image/fetch/$s_!A8Uw!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd504476-42c6-4428-b186-05722e724d1e_2600x2600.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!A8Uw!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd504476-42c6-4428-b186-05722e724d1e_2600x2600.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>Intro</h2><p>The first thing about <strong>FP4</strong> that I could not explain to myself was a byte. <strong>NVFP4</strong> and <strong>MXFP4</strong> use the same 4-bit element, <em>E2M1</em>: one sign bit, two exponent bits, one mantissa bit. A Blackwell tensor core multiplies <strong>the same sixteen codes</strong> in both formats. </p><p>What differs is how a block of values stores its <strong>scale</strong>: one byte for every <strong>32 values</strong> in MXFP4, a different kind of byte for every <strong>16 values</strong> in NVFP4.</p><p> Every comparison I read puts NVFP4 ahead, sometimes far ahead, and yet OpenAI&#8217;s <em>gpt-oss</em>, <em>DeepSeek V4</em> and Moonshot&#8217;s <em>Kimi K3</em> all shipped in <strong>MXFP4</strong>, and Moonshot&#8217;s previous flagship used neither format.</p><p>I wanted the difference <strong>in numbers rather than adjectives</strong>, and I wanted to know what the silicon does with each format. Almost everything below comes from work that needs <strong>no GPU</strong>. </p><p>We rebuilt the rounding rules of every block format we could find in <em>NumPy</em> and checked them against <strong>two independent references</strong>: a published table of Gaussian errors, and AMD&#8217;s measurements on the activations of two real models. </p><p>Then we asked NVIDIA&#8217;s own assembler, <strong>ptxas 13.4.92</strong> from the PyPI wheels, which FP4 instructions <strong>eight targets from Hopper to Rubin</strong> accept and what machine code each one becomes. Where measurement was impossible we read <em>the vendor code that ships with the hardware</em>.</p><h3>At a glance</h3><ul><li><p><em><strong>One byte of scale</strong> separates NVFP4 from MXFP4. It is worth <strong>1.6 dB</strong> of signal-to-noise on Gaussian data, <strong>4.2 dB</strong> on activations and about <strong>9% of perplexity</strong> under plain rounding.</em></p></li><li><p><em>Under the open standard&#8217;s scale rule, <strong>41.5% of MXFP4 blocks clip their own maximum</strong>, and the rule that rescued training is the worse one for inference error.</em></p></li><li><p><em><strong>One constant turns format noise into perplexity</strong>, explaining <strong>98.7%</strong> of the variance across 21 formats; its value changes from model to model.</em></p></li><li><p><em><strong>FP4 arithmetic needs both operands in FP4.</strong> The big open releases ship FP4 weights with 8-bit or 16-bit activations, so they collect <strong>the bandwidth half only</strong>.</em></p></li><li><p><em><strong>INT8 is gone</strong> from the tensor-core instructions of B300 and Rubin. <strong>Rubin adds a wider scale format, UE5M3, and 3-bit lookup-table weights.</strong></em></p></li></ul><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Sixteen codes</h2><p><strong>E2M1</strong> has a sign bit, two exponent bits with a <em>bias of 1</em> and one mantissa bit. Its non-negative values are <strong>0, 0.5, 1, 1.5, 2, 3, 4 and 6</strong>. <em>Negative zero</em> is a separate code with the same value, so sixteen codes give <strong>fifteen distinct numbers</strong>. </p><p>The grid is a floating-point format in miniature: neighbours are <strong>0.5 apart up to 2</strong>, then 1 apart, then <strong>2 apart</strong>. Nothing lies between 4 and 6, so a value that lands on 5 after scaling ends up <strong>20% away</strong> from wherever it rounds.</p><p>An <strong>INT4</strong> grid with the same maximum puts its levels <em>6/7 apart everywhere</em>. Which grid is better depends entirely on <strong>where values land after scaling</strong> (Figure 1). </p><p>Every block format scales a block so that its largest value lands near the top of the grid, which pushes the other values of a bell-shaped block toward zero, where the FP4 grid is dense: in our simulation <strong>50% of the values of a Gaussian block land below 2</strong>, where half of the non-negative levels live. </p><p>That is the case for floating point at four bits, and also its weak spot: the few values <strong>near the top of each block fall into the coarsest part of the grid</strong>.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!1hIF!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc9fccc75-52a1-448d-9a9c-a0c5a130ba66_1760x1140.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!1hIF!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc9fccc75-52a1-448d-9a9c-a0c5a130ba66_1760x1140.png 424w, https://substackcdn.com/image/fetch/$s_!1hIF!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc9fccc75-52a1-448d-9a9c-a0c5a130ba66_1760x1140.png 848w, https://substackcdn.com/image/fetch/$s_!1hIF!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc9fccc75-52a1-448d-9a9c-a0c5a130ba66_1760x1140.png 1272w, https://substackcdn.com/image/fetch/$s_!1hIF!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc9fccc75-52a1-448d-9a9c-a0c5a130ba66_1760x1140.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!1hIF!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc9fccc75-52a1-448d-9a9c-a0c5a130ba66_1760x1140.png" width="1456" height="943" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c9fccc75-52a1-448d-9a9c-a0c5a130ba66_1760x1140.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:943,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Figure 1. The E2M1 grid against INT4 at the same maximum, and where the values of a Gaussian block land once its maximum is scaled to 6.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 1. The E2M1 grid against INT4 at the same maximum, and where the values of a Gaussian block land once its maximum is scaled to 6." title="Figure 1. The E2M1 grid against INT4 at the same maximum, and where the values of a Gaussian block land once its maximum is scaled to 6." srcset="https://substackcdn.com/image/fetch/$s_!1hIF!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc9fccc75-52a1-448d-9a9c-a0c5a130ba66_1760x1140.png 424w, https://substackcdn.com/image/fetch/$s_!1hIF!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc9fccc75-52a1-448d-9a9c-a0c5a130ba66_1760x1140.png 848w, https://substackcdn.com/image/fetch/$s_!1hIF!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc9fccc75-52a1-448d-9a9c-a0c5a130ba66_1760x1140.png 1272w, https://substackcdn.com/image/fetch/$s_!1hIF!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc9fccc75-52a1-448d-9a9c-a0c5a130ba66_1760x1140.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Figure 1.</strong> The E2M1 grid against INT4 at the same maximum, and where the values of a Gaussian block land once its maximum is scaled to 6.</figcaption></figure></div><p>The yardstick throughout this piece is the <strong>signal-to-quantization-noise ratio (SQNR)</strong>: ten times the base-10 logarithm of signal power over error power. Each extra bit of a well-designed quantizer is worth about <strong>6.02 dB</strong>. </p><p>The best possible 16-level quantizer for a Gaussian, the <em>Lloyd-Max quantizer</em>, needs no scales because it is told the variance in advance; it reaches <strong>20.19 dB</strong> in our simulation, against 20.22 dB in the textbooks. </p><p>Real formats must <strong>discover the scale from the data, block by block</strong>, and they pay for it in bits. MXFP4 spends 8 bits of scale on every 32 values, <strong>4.25 bits per element</strong> in total. NVFP4 spends 8 bits on every 16 values plus 32 bits per tensor, <strong>4.5 bits per element</strong>. </p><p>NVFP4 therefore shrinks an FP8 checkpoint by <strong>1.78 times rather than 2</strong>, a detail that matters again when we get to the <em>critical batch</em>.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>One byte of scale</h2><p>MXFP4 comes from the Open Compute Project&#8217;s <em>Microscaling specification</em> (OCP MX v1.0, in september 2023). Its scale is <strong>E8M0</strong>: eight exponent bits and no mantissa, so <strong>every scale is a power of two</strong>. </p><p>The specification sets the shared exponent of a block to <strong>floor(log<sub>2</sub> amax) - 2</strong>, where <em>amax</em> is the largest magnitude in the block and 2 is the exponent of 6, the largest E2M1 value. </p><p>Divide the block by that scale and its largest value lands at 4 x 2<sup>f</sup>, where <em>f</em> is the fractional part of log<sub>2</sub> amax. Because f is somewhere in [0, 1), <strong>the block maximum lands somewhere in [4, 8)</strong>, and everything above 6 <strong>saturates to 6</strong>.</p><p>The maximum is clipped whenever 2<sup>f</sup> exceeds 1.5, that is whenever f exceeds log<sub>2</sub> 1.5 = 0.585. If block magnitudes are spread evenly on a log scale, which is roughly what happens across the thousands of blocks and dozens of layers of a model, that is <strong>1 - log<sub>2</sub> 1.5 = 41.5% of all blocks</strong>. </p><p>Our simulation over 128 tensor scales gives <strong>41.51%</strong>. A clipped value can lose <strong>up to a quarter of its magnitude</strong>: 7.9 becomes 6.</p><p>NVFP4, introduced with <strong>Blackwell</strong>, stores the block scale as <strong>E4M3</strong>, the FP8 format with <em>three mantissa bits</em>, and computes it as <strong>amax / 6</strong> rounded to the nearest E4M3 value (NVIDIA technical blog, June 2025). The block maximum then lands <strong>within about 6% of 6</strong> instead of anywhere between 4 and 8, and the small overshoot rounds back to 6. </p><p>The price is <strong>range</strong>. Unsigned E4M3 runs from 2<sup>-9</sup> to 448, about <strong>18 binades</strong>, which is not enough to hold the scales of an entire model. NVFP4 therefore adds <strong>a second level</strong>: one <strong>FP32 number per tensor</strong>, equal to the tensor&#8217;s amax / (6 x 448), divides the whole tensor before the block scales are computed.</p><p>There is a second way to build an E8M0 scale. Asit Mishra, Dusan Stosic, Simon Layton and <strong>Paulius Micikevicius </strong>showed that the OCP rule can make <em>MXFP8 pretraining</em> drift away from its BF16 baseline, and fixed it by <strong>rounding the scale up</strong>, so that amax divided by the scale never exceeds the largest element (<em>Recipes for Pre-training LLMs with MXFP8</em>, arXiv:2506.08027). </p><p>Rounding up <strong>removes clipping</strong>, but the block maximum now lands anywhere in (3, 6], which <strong>leaves the top of the grid unused</strong> for most blocks (Figure 2). <strong>Both rules exist in silicon.</strong> </p><p>On every Blackwell target we tried, cvt.rz.satfinite.ue8m0x2.f32 and its .rp twin each assemble to <strong>a single F2FP.SATFINITE.E8 instruction</strong> with a <em>.RZ</em> or <em>.RP</em> modifier, so the choice costs nothing at run time.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!HT61!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F474a2ed5-0188-4ad2-b248-3440434e5a13_1760x1200.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!HT61!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F474a2ed5-0188-4ad2-b248-3440434e5a13_1760x1200.png 424w, https://substackcdn.com/image/fetch/$s_!HT61!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F474a2ed5-0188-4ad2-b248-3440434e5a13_1760x1200.png 848w, https://substackcdn.com/image/fetch/$s_!HT61!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F474a2ed5-0188-4ad2-b248-3440434e5a13_1760x1200.png 1272w, https://substackcdn.com/image/fetch/$s_!HT61!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F474a2ed5-0188-4ad2-b248-3440434e5a13_1760x1200.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!HT61!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F474a2ed5-0188-4ad2-b248-3440434e5a13_1760x1200.png" width="1456" height="993" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/474a2ed5-0188-4ad2-b248-3440434e5a13_1760x1200.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:993,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Figure 2. Where the block maximum lands after division by its scale, for 65,536 Gaussian blocks with random tensor scales, under the two MX rules and under NVFP4.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 2. Where the block maximum lands after division by its scale, for 65,536 Gaussian blocks with random tensor scales, under the two MX rules and under NVFP4." title="Figure 2. Where the block maximum lands after division by its scale, for 65,536 Gaussian blocks with random tensor scales, under the two MX rules and under NVFP4." srcset="https://substackcdn.com/image/fetch/$s_!HT61!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F474a2ed5-0188-4ad2-b248-3440434e5a13_1760x1200.png 424w, https://substackcdn.com/image/fetch/$s_!HT61!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F474a2ed5-0188-4ad2-b248-3440434e5a13_1760x1200.png 848w, https://substackcdn.com/image/fetch/$s_!HT61!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F474a2ed5-0188-4ad2-b248-3440434e5a13_1760x1200.png 1272w, https://substackcdn.com/image/fetch/$s_!HT61!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F474a2ed5-0188-4ad2-b248-3440434e5a13_1760x1200.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Figure 2.</strong> Where the block maximum lands after division by its scale, for 65,536 Gaussian blocks with random tensor scales, under the two MX rules and under NVFP4.</figcaption></figure></div><p>For inference the two rules are not equivalent, and <strong>the rule that rescued training is the worse one for error</strong>. Averaged over tensor scales, the OCP rule gives <strong>18.75 dB</strong> on Gaussian data against <strong>18.65 dB</strong> for rounding up. </p><p>On the activation proxy introduced below the gap is <strong>0.50 dB</strong>, again in favour of the OCP rule (16.69 against 16.19 dB). Clipping one value per block costs less squared error than coarsening every other value. <em>Training is hurt by bias</em>, and a clipped maximum is biased; <em>a single forward pass is hurt by squared error</em>.</p><blockquote><p><em>Clipping one value per block costs less squared error than coarsening every other value in it.</em></p></blockquote><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>The sawtooth</h2><p><strong>Power-of-two scales</strong> have a property that is easy to miss: <strong>the error depends on where the tensor sits relative to powers of two</strong>. Multiply a tensor by a constant c and every MX scale shifts with it, but the fractional part of log<sub>2</sub> c moves every block maximum <em>within its binade</em>, and with it the clipping fraction. </p><p>We multiplied the same 262,144 Gaussian values by 2<sup>t</sup> for 128 values of t between 0 and 2 (<em>Figure 3</em>). <strong>NVFP4 does not move at all</strong>: 20.44 dB at every t, because the FP32 tensor scale absorbs the constant exactly. </p><p><strong>MXFP4 under the OCP rule moves between 18.58 and 18.96 dB</strong> with a period of exactly one binade, and the share of clipped blocks swings between <strong>26.8% and 55.9%</strong>. Rounding up moves between 18.52 and 18.79 dB.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!iMXS!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb0b974cf-f773-4774-b2f5-bc355f50c118_1760x1160.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!iMXS!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb0b974cf-f773-4774-b2f5-bc355f50c118_1760x1160.png 424w, https://substackcdn.com/image/fetch/$s_!iMXS!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb0b974cf-f773-4774-b2f5-bc355f50c118_1760x1160.png 848w, https://substackcdn.com/image/fetch/$s_!iMXS!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb0b974cf-f773-4774-b2f5-bc355f50c118_1760x1160.png 1272w, https://substackcdn.com/image/fetch/$s_!iMXS!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb0b974cf-f773-4774-b2f5-bc355f50c118_1760x1160.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!iMXS!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb0b974cf-f773-4774-b2f5-bc355f50c118_1760x1160.png" width="1456" height="960" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b0b974cf-f773-4774-b2f5-bc355f50c118_1760x1160.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:960,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Figure 3. SQNR as the same tensor is multiplied by 2^t (top), and the share of MXFP4 blocks whose maximum saturates under the OCP rule (bottom).&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 3. SQNR as the same tensor is multiplied by 2^t (top), and the share of MXFP4 blocks whose maximum saturates under the OCP rule (bottom)." title="Figure 3. SQNR as the same tensor is multiplied by 2^t (top), and the share of MXFP4 blocks whose maximum saturates under the OCP rule (bottom)." srcset="https://substackcdn.com/image/fetch/$s_!iMXS!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb0b974cf-f773-4774-b2f5-bc355f50c118_1760x1160.png 424w, https://substackcdn.com/image/fetch/$s_!iMXS!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb0b974cf-f773-4774-b2f5-bc355f50c118_1760x1160.png 848w, https://substackcdn.com/image/fetch/$s_!iMXS!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb0b974cf-f773-4774-b2f5-bc355f50c118_1760x1160.png 1272w, https://substackcdn.com/image/fetch/$s_!iMXS!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb0b974cf-f773-4774-b2f5-bc355f50c118_1760x1160.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Figure 3.</strong> SQNR as the same tensor is multiplied by 2^t (top), and the share of MXFP4 blocks whose maximum saturates under the OCP rule (bottom).</figcaption></figure></div><p>The swing is small on Gaussian data, <strong>about 0.4 dB</strong> from trough to peak, but it means that an MXFP4 model&#8217;s accuracy depends on <strong>constants nobody chose for this purpose</strong>, such as the <em>gain of the normalization layer</em> in front of a projection. It also points at a cheap remedy. </p><p><strong>A per-tensor multiplier</strong> that moves each tensor to its best phase recovers the distance from the average to the peak, <strong>0.20 dB</strong> here. </p><p>That remedy is exactly NVFP4&#8217;s structure, an FP32 number per tensor above the block scales, and a November 2025 preprint applies the same <em>two-level idea</em> to <strong>MXFP8 training</strong>, keeping power-of-two block scales under an FP32 tensor scale (arXiv:2511.05811).</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Where the error lives</h2><p>Before trusting any of these numbers we <strong>checked the simulator against someone else&#8217;s</strong>. Jack Cook and colleagues at MIT and NVIDIA publish the <em>mean squared error</em> of five block formats on standard normal data (<em>Adaptive Block-Scaled Data Types</em>, arXiv:2603.28765, Table 1). </p><p>Our implementation <strong>reproduces all five within 1%</strong>: 13.23 against their 13.2 (in units of 10<sup>-3</sup>) for MXFP4, 9.04 against 9.0 for NVFP4, 7.56 against 7.5 for NVFP4 with Four Over Six, 7.44 against 7.4 for NVINT4 and 6.16 against 6.2 for their IF4 format (Figure 4, left). </p><p>The agreement says that the rounding details (<em>ties to even</em>, <em>saturation</em>, <em>scale rounding</em>, the FP32 tensor scale) match theirs, not only the headline numbers.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!s3k3!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a68d503-46d6-412c-81ea-501379668593_1760x1140.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!s3k3!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a68d503-46d6-412c-81ea-501379668593_1760x1140.png 424w, https://substackcdn.com/image/fetch/$s_!s3k3!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a68d503-46d6-412c-81ea-501379668593_1760x1140.png 848w, https://substackcdn.com/image/fetch/$s_!s3k3!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a68d503-46d6-412c-81ea-501379668593_1760x1140.png 1272w, https://substackcdn.com/image/fetch/$s_!s3k3!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a68d503-46d6-412c-81ea-501379668593_1760x1140.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!s3k3!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a68d503-46d6-412c-81ea-501379668593_1760x1140.png" width="1456" height="943" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/5a68d503-46d6-412c-81ea-501379668593_1760x1140.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:943,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Figure 4. Left: our simulator against Cook et al., Table 1. Right: our activation proxy against AMD's per-layer measurements on two real models.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 4. Left: our simulator against Cook et al., Table 1. Right: our activation proxy against AMD's per-layer measurements on two real models." title="Figure 4. Left: our simulator against Cook et al., Table 1. Right: our activation proxy against AMD's per-layer measurements on two real models." srcset="https://substackcdn.com/image/fetch/$s_!s3k3!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a68d503-46d6-412c-81ea-501379668593_1760x1140.png 424w, https://substackcdn.com/image/fetch/$s_!s3k3!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a68d503-46d6-412c-81ea-501379668593_1760x1140.png 848w, https://substackcdn.com/image/fetch/$s_!s3k3!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a68d503-46d6-412c-81ea-501379668593_1760x1140.png 1272w, https://substackcdn.com/image/fetch/$s_!s3k3!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a68d503-46d6-412c-81ea-501379668593_1760x1140.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Figure 4.</strong> Left: our simulator against Cook et al., Table 1. Right: our activation proxy against AMD&#8217;s per-layer measurements on two real models.</figcaption></figure></div><p>With the simulator anchored, the next question is <strong>where inside each block the error comes from</strong> (Figure 5). </p><p>Under the OCP rule, <strong>the single largest value of a 32-value MXFP4 block carries 23.1% of the block&#8217;s squared error</strong> on Gaussian data and <strong>39.9% on activations</strong>, against the 3.1% it would carry if error were spread evenly. </p><p>Its average relative error is <strong>11.9%</strong>. In NVFP4 the block maximum carries <strong>2.1% and 5.5%</strong> of the error, at or below its even share of 6.25%, with an average relative error of <strong>2.25%</strong>.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!3UPT!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a4ae409-522a-4108-8746-f1e4e2dcb3ce_1760x1100.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!3UPT!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a4ae409-522a-4108-8746-f1e4e2dcb3ce_1760x1100.png 424w, https://substackcdn.com/image/fetch/$s_!3UPT!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a4ae409-522a-4108-8746-f1e4e2dcb3ce_1760x1100.png 848w, https://substackcdn.com/image/fetch/$s_!3UPT!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a4ae409-522a-4108-8746-f1e4e2dcb3ce_1760x1100.png 1272w, https://substackcdn.com/image/fetch/$s_!3UPT!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a4ae409-522a-4108-8746-f1e4e2dcb3ce_1760x1100.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!3UPT!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a4ae409-522a-4108-8746-f1e4e2dcb3ce_1760x1100.png" width="1456" height="910" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/5a4ae409-522a-4108-8746-f1e4e2dcb3ce_1760x1100.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:910,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Figure 5. Share of each block's squared error carried by its largest value, on Gaussian data and on activations.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 5. Share of each block's squared error carried by its largest value, on Gaussian data and on activations." title="Figure 5. Share of each block's squared error carried by its largest value, on Gaussian data and on activations." srcset="https://substackcdn.com/image/fetch/$s_!3UPT!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a4ae409-522a-4108-8746-f1e4e2dcb3ce_1760x1100.png 424w, https://substackcdn.com/image/fetch/$s_!3UPT!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a4ae409-522a-4108-8746-f1e4e2dcb3ce_1760x1100.png 848w, https://substackcdn.com/image/fetch/$s_!3UPT!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a4ae409-522a-4108-8746-f1e4e2dcb3ce_1760x1100.png 1272w, https://substackcdn.com/image/fetch/$s_!3UPT!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a4ae409-522a-4108-8746-f1e4e2dcb3ce_1760x1100.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Figure 5.</strong> Share of each block&#8217;s squared error carried by its largest value, on Gaussian data and on activations.</figcaption></figure></div><p>NVFP4 still has <strong>a soft spot near the top of its grid</strong>, because nothing exists between 4 and 6. <strong>Four Over Six</strong>, from Cook, Junxian Guo, Song Han and colleagues (arXiv:2512.02010), scales each block <strong>either to 6 or to 4</strong>, whichever gives less error. </p><p>Scaling to 4 gives up the levels 6 and -6 but makes 3 a level at 75% of the block maximum. In our simulation the choice is worth <strong>0.78 dB on Gaussian data and 0.30 dB on activations</strong>. </p><p><strong>IF4</strong>, from the same group, goes further: each block chooses between the FP4 grid and an <em>INT4 grid rescaled to the same maximum</em>, and records the choice in <strong>the sign bit of its E4M3 scale</strong>, a bit NVFP4 never uses because scales are positive. </p><p>It reaches <strong>22.11 dB</strong> on Gaussian data, <strong>1.7 dB above NVFP4 at the same 4.5 bits</strong>.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Tails</h2><p>Gaussian data is the friendly case, and <strong>activations are not Gaussian</strong>. The input to the <em>down projection</em> of a <strong>SwiGLU</strong> MLP is silu(a) x b, the product of two roughly Gaussian projections, and <strong>products of Gaussian variables have heavy tails</strong>. </p><p>Our proxy is exactly that product, with independent standard normal a and b. Its <em>kurtosis</em> is <strong>about 22</strong>, against 3 for a Gaussian.</p><p>A proxy that simple needs a reality check, and AMD published the one we needed. In June 2026 its ROCm team measured SQNR on the <em>post-SiLU down_proj inputs</em> of every layer of <strong>Llama-3.1-8B</strong> and <strong>Qwen3.6-27B</strong>, using the <em>OCP reference quantizers</em> (Atre, Bao, Tiwari and Sirasao, AMD ROCm blog, 26 June 2026). </p><p>Their layer means were <strong>16.8 and 16.4 dB for MXFP4</strong>, <strong>29.3 and 28.6 dB for MXFP6 E2M3</strong>, and <strong>32.8 and 32.1 dB for FP8</strong> with one scale per token. The proxy gives <strong>16.69, 28.97 and 31.69 dB</strong> (Figure 4, right). </p><p>The first two fall between AMD&#8217;s two models and FP8 is within 1.1 dB. For a distribution defined in one line that is closer than we expected, and it is <strong>the activation model we use from here on</strong>.</p><p>On heavy tails <strong>the formats separate</strong> (Figure 6). <strong>NVFP4 improves</strong>, from 20.43 to <strong>20.90 dB</strong>, because a floating-point grid suits blocks with one large value and many small ones. <strong>MXFP4 falls</strong>, from 18.79 to <strong>16.69 dB</strong>. The NVFP4 advantage grows from <strong>1.64 dB on Gaussian data to 4.21 dB on activations</strong>. <strong>Integer grids suffer most.</strong> </p><p>NVINT4, which beats NVFP4 on Gaussian data at 21.29 dB, drops to <strong>18.55 dB</strong>, and INT4 with a scale per 32 values, the grouping of Moonshot&#8217;s <em>Kimi K2 Thinking</em> checkpoint (we model the scale as BF16), drops from 20.26 to <strong>16.64 dB</strong>.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!yL5Y!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee334d81-9f74-45d4-9100-5251c5834ec3_1760x1220.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!yL5Y!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee334d81-9f74-45d4-9100-5251c5834ec3_1760x1220.png 424w, https://substackcdn.com/image/fetch/$s_!yL5Y!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee334d81-9f74-45d4-9100-5251c5834ec3_1760x1220.png 848w, https://substackcdn.com/image/fetch/$s_!yL5Y!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee334d81-9f74-45d4-9100-5251c5834ec3_1760x1220.png 1272w, https://substackcdn.com/image/fetch/$s_!yL5Y!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee334d81-9f74-45d4-9100-5251c5834ec3_1760x1220.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!yL5Y!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee334d81-9f74-45d4-9100-5251c5834ec3_1760x1220.png" width="1456" height="1009" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ee334d81-9f74-45d4-9100-5251c5834ec3_1760x1220.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1009,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Figure 6. Every format we simulated, by storage cost and SQNR (left), and the 4-bit formats side by side on Gaussian data, activations and real weights (right).&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 6. Every format we simulated, by storage cost and SQNR (left), and the 4-bit formats side by side on Gaussian data, activations and real weights (right)." title="Figure 6. Every format we simulated, by storage cost and SQNR (left), and the 4-bit formats side by side on Gaussian data, activations and real weights (right)." srcset="https://substackcdn.com/image/fetch/$s_!yL5Y!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee334d81-9f74-45d4-9100-5251c5834ec3_1760x1220.png 424w, https://substackcdn.com/image/fetch/$s_!yL5Y!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee334d81-9f74-45d4-9100-5251c5834ec3_1760x1220.png 848w, https://substackcdn.com/image/fetch/$s_!yL5Y!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee334d81-9f74-45d4-9100-5251c5834ec3_1760x1220.png 1272w, https://substackcdn.com/image/fetch/$s_!yL5Y!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee334d81-9f74-45d4-9100-5251c5834ec3_1760x1220.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Figure 6.</strong> Every format we simulated, by storage cost and SQNR (left), and the 4-bit formats side by side on Gaussian data, activations and real weights (right).</figcaption></figure></div><p><strong>Weights sit in between.</strong> We extracted the <strong>84.9 million linear-layer weights</strong> of a real trained transformer, the <em>RoBERTa-base</em> encoder packaged in spaCy&#8217;s en_core_web_trf 3.8.0, read them <em>without PyTorch</em> and quantized every matrix along its input dimension. </p><p>Their kurtosis is <strong>6.0</strong> and the largest weight lies <strong>27 standard deviations</strong> from zero. NVFP4 reaches <strong>20.51 dB</strong>, MXFP4 <strong>18.58 dB</strong>, NVINT4 20.87 dB, Four Over Six 21.23 dB and IF4 <strong>21.98 dB</strong>. </p><p>The NVFP4 advantage over MXFP4 is <strong>1.94 dB</strong> pooled, and between 1.67 and 2.55 dB across the 48 matrices. RoBERTa is <em>a 2019 encoder, not a modern decoder</em>, so we read it as <strong>a sanity check</strong> on the synthetic numbers rather than as a benchmark.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share The Software Frontier&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share The Software Frontier</span></a></p><div><hr></div><h2>Rotations</h2><p>The usual cure for heavy tails is <strong>a rotation</strong>. Multiply blocks of values by a <em>Hadamard matrix</em> and each output becomes a signed average of all inputs, so a single outlier is spread across the block and the distribution moves toward Gaussian. </p><p><em>QuaRot</em> and <em>SpinQuant</em> made this standard practice for INT4. For FP4 it is <strong>less obviously useful</strong>. Vage Egiazarian, Dan Alistarh and colleagues show that <strong>rotations help MXFP4 but hurt NVFP4 under round-to-nearest</strong>, and prove that NVFP4&#8217;s <em>small 16-value groups</em> already neutralize the outlier mitigation that rotations provide (<em>Bridging the Gap Between Promise and Performance for Microscaling FP4 Quantization</em>, ICLR 2026, arXiv:2509.23202).</p><p>Our simulation <strong>reproduces the sign of every effect and gives the sizes</strong> (Figure 7). A 16-point Hadamard rotation leaves Gaussian data unchanged within 0.03 dB. On activations it <strong>costs NVFP4 0.88 dB</strong>, <strong>gives MXFP4 2.13 dB</strong> and <strong>gives NVINT4 4.07 dB</strong>. </p><p>With eight <em>outlier channels</em> at twenty times the scale of the others the changes are -1.03, +2.27 and +3.21 dB. Larger rotations hurt NVFP4 less: a 128-point rotation costs it 0.51 dB on activations. </p><p>On RoBERTa&#8217;s weights, whose kurtosis is 6 rather than 22, a 16-point rotation changes NVFP4 by -0.09 dB, MXFP4 by +0.16 dB and NVINT4 by +0.42 dB: <strong>the same signs, much smaller sizes</strong>.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!rLVL!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbda7a876-014a-4806-86df-d94ae318ac91_1760x1100.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!rLVL!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbda7a876-014a-4806-86df-d94ae318ac91_1760x1100.png 424w, https://substackcdn.com/image/fetch/$s_!rLVL!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbda7a876-014a-4806-86df-d94ae318ac91_1760x1100.png 848w, https://substackcdn.com/image/fetch/$s_!rLVL!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbda7a876-014a-4806-86df-d94ae318ac91_1760x1100.png 1272w, https://substackcdn.com/image/fetch/$s_!rLVL!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbda7a876-014a-4806-86df-d94ae318ac91_1760x1100.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!rLVL!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbda7a876-014a-4806-86df-d94ae318ac91_1760x1100.png" width="1456" height="910" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/bda7a876-014a-4806-86df-d94ae318ac91_1760x1100.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:910,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Figure 7. Change in SQNR when a 16-point Hadamard rotation is applied before quantization, by distribution.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 7. Change in SQNR when a 16-point Hadamard rotation is applied before quantization, by distribution." title="Figure 7. Change in SQNR when a 16-point Hadamard rotation is applied before quantization, by distribution." srcset="https://substackcdn.com/image/fetch/$s_!rLVL!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbda7a876-014a-4806-86df-d94ae318ac91_1760x1100.png 424w, https://substackcdn.com/image/fetch/$s_!rLVL!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbda7a876-014a-4806-86df-d94ae318ac91_1760x1100.png 848w, https://substackcdn.com/image/fetch/$s_!rLVL!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbda7a876-014a-4806-86df-d94ae318ac91_1760x1100.png 1272w, https://substackcdn.com/image/fetch/$s_!rLVL!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbda7a876-014a-4806-86df-d94ae318ac91_1760x1100.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Figure 7.</strong> Change in SQNR when a 16-point Hadamard rotation is applied before quantization, by distribution.</figcaption></figure></div><p><strong>The best 4-bit recipe in our sweep is a rotation followed by NVINT4</strong>, integers under an E4M3 scale per 16 values: <strong>22.63 dB on activations</strong>, 1.74 dB above plain NVFP4 and 1.36 dB above IF4. </p><p>Cook and colleagues found the same ordering in real W4A4 runs on Qwen3.5, where NVINT4 after a Hadamard transform beat NVFP4. The catch comes in the section on the assembler: <strong>no tensor-core instruction on a shipping NVIDIA datacenter part multiplies INT4 values under E4M3 block scales</strong>. </p><p>The block-scaled FP4 instructions accept only E2M1 elements, and the only integer tensor-core instruction left on datacenter Blackwell takes INT8 and is <em>being withdrawn</em>.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>From decibels to perplexity</h2><p>Signal-to-noise ratios matter only if they <strong>predict something people care about</strong>. Cook and colleagues report the average <em>WikiText-2 perplexity</em> of four <strong>Qwen3.5</strong> models (9B, 27B, 35B-A3B and 122B-A10B) for <strong>23 block formats</strong>, with the weights and activations of every linear layer quantized by <em>round-to-nearest</em> (arXiv:2603.28765, Table 6). </p><p>We simulated 22 of them, all but a block-8 variant of Four Over Six. We computed the <strong>noise-to-signal power ratio (NSR)</strong> of each format on our activation proxy and fitted the simplest law that could work: <strong>the logarithm of the perplexity ratio is proportional to the noise power</strong>.</p><p>With one free parameter, <strong>ln(ppl / ppl<sub>BF16</sub>) = 8.25 x NSR explains 98.7% of the variance</strong> across the 21 formats that stay above 12 dB, from 3.5 to 6.5 bits per element, integer and floating point, MX and NV (Figure 8). The <em>rank correlation</em> is <strong>0.986</strong>. </p><p>Dropping any one format moves the fitted constant between 7.95 and 8.33, and the <em>median error</em> of the fit is <strong>0.6% of perplexity</strong>. The same fit on Gaussian noise explains 95.4%, so <strong>the heavy-tailed proxy predicts better</strong>, as the comparison with AMD suggested. </p><p>W4A4 adds weight noise to activation noise, so we also charged each format twice, once on <strong>Gaussian data for its weights </strong>and once on the proxy for its activations. That doesn&#8217;t do better (98.3% with one constant, 98.8% with a constant per term), and the two-term fit puts most of the weight on the activation term. </p><p>The two terms rise and fall together across formats, so that split is <em>a hint rather than a measurement</em>. <strong>One format breaks the law: MXFP3</strong>, at a perplexity of 70. Below roughly 12 dB the models <strong>stop degrading gracefully and start to fail</strong>.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!ZCK_!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8013e9aa-d3ec-41df-9ef5-5b7e352a014f_1760x1220.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!ZCK_!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8013e9aa-d3ec-41df-9ef5-5b7e352a014f_1760x1220.png 424w, https://substackcdn.com/image/fetch/$s_!ZCK_!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8013e9aa-d3ec-41df-9ef5-5b7e352a014f_1760x1220.png 848w, https://substackcdn.com/image/fetch/$s_!ZCK_!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8013e9aa-d3ec-41df-9ef5-5b7e352a014f_1760x1220.png 1272w, https://substackcdn.com/image/fetch/$s_!ZCK_!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8013e9aa-d3ec-41df-9ef5-5b7e352a014f_1760x1220.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!ZCK_!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8013e9aa-d3ec-41df-9ef5-5b7e352a014f_1760x1220.png" width="1456" height="1009" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/8013e9aa-d3ec-41df-9ef5-5b7e352a014f_1760x1220.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1009,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Figure 8. Perplexity increase against format noise for 21 formats, on log scales, with four out-of-sample points from AMD's measurements.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 8. Perplexity increase against format noise for 21 formats, on log scales, with four out-of-sample points from AMD's measurements." title="Figure 8. Perplexity increase against format noise for 21 formats, on log scales, with four out-of-sample points from AMD's measurements." srcset="https://substackcdn.com/image/fetch/$s_!ZCK_!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8013e9aa-d3ec-41df-9ef5-5b7e352a014f_1760x1220.png 424w, https://substackcdn.com/image/fetch/$s_!ZCK_!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8013e9aa-d3ec-41df-9ef5-5b7e352a014f_1760x1220.png 848w, https://substackcdn.com/image/fetch/$s_!ZCK_!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8013e9aa-d3ec-41df-9ef5-5b7e352a014f_1760x1220.png 1272w, https://substackcdn.com/image/fetch/$s_!ZCK_!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8013e9aa-d3ec-41df-9ef5-5b7e352a014f_1760x1220.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Figure 8.</strong> Perplexity increase against format noise for 21 formats, on log scales, with four out-of-sample points from AMD&#8217;s measurements.</figcaption></figure></div><p>Read backwards, the fit gives <strong>targets</strong>. Staying within <strong>1%</strong> of BF16 perplexity needs about <strong>29.2 dB</strong> on the proxy, within 2% about 26.2 dB, within <strong>5%</strong> about <strong>22.3 dB</strong>. NVFP4 sits at 20.9 dB, predicted <strong>+6.9%</strong> against a measured +6.4%; MXFP4 sits at 16.7 dB, predicted <strong>+19.3%</strong> against a measured +16.0%. </p><p>The measured gap between them is <strong>9.1% of perplexity</strong>. That is <strong>the size of one byte of scale per block</strong> when nothing else is done to the model. <em>Quantization-aware training</em>, rotations and better rounding all shrink it; </p><p><strong>MR-GPTQ</strong>, the method Egiazarian and colleagues propose, brings MXFP4 to <strong>within 1 to 2% of NVFP4 accuracy</strong>.</p><p><strong>The law travels, but its constant does not.</strong> AMD&#8217;s appendix gives WikiText-2 perplexities for the two models it profiled, quantized W4A4 by round-to-nearest. </p><p>For <strong>Llama-3.1-8B</strong> the law predicts <strong>+19.3%</strong> for MXFP4 against a measured <strong>+20.0%</strong> (8.47 against 7.06), and +0.6% for per-token FP8 against a measured +0.7%. </p><p>For <strong>Qwen3.6-27B</strong> it predicts the same +19.3% for MXFP4 against a measured <strong>+6.1%</strong> (7.53 against 7.10): the larger and newer model <strong>tolerates about three times the noise</strong>. The constant is <em>a property of the model, not of the format</em>, and 8.25 is an average over a family that runs from 9 to 122 billion parameters.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/p/a-wonderful-dive-into-how-fp4-works/comments&quot;,&quot;text&quot;:&quot;Leave a comment&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/p/a-wonderful-dive-into-how-fp4-works/comments"><span>Leave a comment</span></a></p><div><hr></div><h2>What the assembler accepts</h2><p>The numerics are half the story; <strong>the other half is which instructions exist</strong>. ptxas, nvdisasm and cuobjdump ship as <em>ordinary x86 programs</em> inside NVIDIA&#8217;s PyPI wheels, so the question can be asked <strong>on any laptop</strong>: wrap one instruction in a minimal PTX kernel, assemble it for a target and read the <em>SASS</em> that comes out. </p><p>We did this for the FP4-related instructions of <strong>PTX ISA 9.4</strong> on <strong>eight targets</strong> with <strong>CUDA 13.4.92</strong> (Figure 9). The targets are sm_90a (<em>H100</em>), sm_100a (<em>B200</em>), sm_100f (<em>the portable Blackwell family</em>), sm_103a (<em>B300</em>), sm_107a (<em>Rubin</em>, compute capability 10.7 according to NVIDIA&#8217;s TensorRT 11.3 release notes), sm_110a (<em>Jetson Thor</em>), sm_120a (<em>RTX 50 and RTX PRO</em>) and sm_121a (<em>GB10</em>).</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Oqhg!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa5b7c7d5-2b39-46a0-bd81-ce1d4f35ffa7_1760x1460.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Oqhg!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa5b7c7d5-2b39-46a0-bd81-ce1d4f35ffa7_1760x1460.png 424w, https://substackcdn.com/image/fetch/$s_!Oqhg!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa5b7c7d5-2b39-46a0-bd81-ce1d4f35ffa7_1760x1460.png 848w, https://substackcdn.com/image/fetch/$s_!Oqhg!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa5b7c7d5-2b39-46a0-bd81-ce1d4f35ffa7_1760x1460.png 1272w, https://substackcdn.com/image/fetch/$s_!Oqhg!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa5b7c7d5-2b39-46a0-bd81-ce1d4f35ffa7_1760x1460.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Oqhg!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa5b7c7d5-2b39-46a0-bd81-ce1d4f35ffa7_1760x1460.png" width="1456" height="1208" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a5b7c7d5-2b39-46a0-bd81-ce1d4f35ffa7_1760x1460.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1208,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Figure 9. Which FP4-related instructions assemble on which target, and the SASS opcode each becomes.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 9. Which FP4-related instructions assemble on which target, and the SASS opcode each becomes." title="Figure 9. Which FP4-related instructions assemble on which target, and the SASS opcode each becomes." srcset="https://substackcdn.com/image/fetch/$s_!Oqhg!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa5b7c7d5-2b39-46a0-bd81-ce1d4f35ffa7_1760x1460.png 424w, https://substackcdn.com/image/fetch/$s_!Oqhg!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa5b7c7d5-2b39-46a0-bd81-ce1d4f35ffa7_1760x1460.png 848w, https://substackcdn.com/image/fetch/$s_!Oqhg!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa5b7c7d5-2b39-46a0-bd81-ce1d4f35ffa7_1760x1460.png 1272w, https://substackcdn.com/image/fetch/$s_!Oqhg!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa5b7c7d5-2b39-46a0-bd81-ce1d4f35ffa7_1760x1460.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Figure 9.</strong> Which FP4-related instructions assemble on which target, and the SASS opcode each becomes.</figcaption></figure></div><p><strong>Hopper has no FP4 at all.</strong> ptxas rejects conversion to FP4 on sm_90a as a feature the target does not support, rejects the conversion back, rejects conversion to E8M0 scales and has <strong>no tcgen05 instructions</strong>. </p><p>FP4 weights on an H100 must be <strong>decoded into BF16 or FP8</strong> with integer tricks or table lookups before they reach a tensor core, and at that point <strong>FP4 offers nothing that INT4 does not</strong>. </p><p>That is the practical background to the <em>Kimi K2 Thinking</em> team choosing <strong>INT4 weights with 16-bit activations</strong> in November 2025: weight-only INT4 kernels such as <em>Marlin</em> were mature on Hopper (LMSYS blog, January 2026).</p><p>On datacenter Blackwell, <strong>FP4 reaches its own datapath only through the block-scaled FP4 kinds</strong>. tcgen05.mma with .kind::mxf4 or .kind::mxf4nvf4 assembles to <strong>UTCOMMA</strong>. </p><p>The same E2M1 operands sent through .kind::f8f6f4, the kind that also accepts FP8 and FP6, assemble to <strong>UTCQMMA, the opcode FP8 uses</strong>. The names fit <em>H for half, Q for quarter and O for one eighth</em> of 32 bits, and they fit NVIDIA&#8217;s documented rates: <strong>FP4 through the mixed kind runs at FP8 speed</strong>. </p><p>Any model that multiplies FP4 weights by FP8 activations is in this position <strong>by construction</strong>, because one of its operands is eight bits wide.</p><p>NVFP4 with 16-value blocks assembles to <strong>UTCOMMA.4X</strong> on sm_100a, sm_100f and sm_110a and to <strong>UTCOMMA.BLOCK16</strong> on sm_103a and sm_107a. The older spelling of the qualifier, .scale_vec::4X, is <strong>rejected for sm_100f</strong>, the family target meant to run on both B200 and B300, so <em>portable builds must use .block16</em>.</p><p><strong>INT8 is on its way out.</strong> tcgen05.mma.kind::i8 assembles to UTCIMMA on sm_100a and on Thor&#8217;s sm_110a, and is <strong>rejected on sm_100f, sm_103a and sm_107a</strong>. </p><p>In August 2026 Teng-Ruei Chen traced the same withdrawal for B300 through the spec sheet, the PTX ISA, CUTLASS and the two main open serving engines (<em>Spec Sheets Are Not Kernels</em>, arXiv:2608.11693). The assembler adds two facts: <strong>Rubin continues it</strong>, and NVIDIA <strong>left INT8 out of the portable Blackwell family</strong> entirely. </p><p>The spec sheets agree. GB300 lists about <strong>0.165 POPS of dense INT8 against 5 PFLOPS of dense FP8</strong>, a ratio of <strong>30</strong>; a Rubin GPU lists <strong>0.25 POPS against 17.5 PFLOPS</strong>, a ratio of <strong>70</strong>. On B200 the two were equal.</p><p><strong>Consumer Blackwell is a different machine.</strong> sm_120a and sm_121a have no tcgen05, but they have <em>warp-level FP4</em>: mma.sync with shape m16n8k64 assembles to <strong>OMMA.SF.16864</strong>, with UE4M3 scales for NVFP4 through .kind::mxf4nvf4 or E8 scales for MXFP4 through .kind::mxf4. </p><p>One more result surprised us: on every datacenter target, Hopper included, <strong>the warp-level FP8 instruction mma.sync.m16n8k32 with E4M3 operands assembles to an F2FP unpack to FP16 followed by HMMA.16816</strong>, a half-precision multiply. </p><p>Only the consumer parts run it natively, as QMMA.16832. On a B200, <strong>FP8 tensor throughput is reachable only through tcgen05</strong>.</p><p><strong>Rubin accepts four instructions that no other target does</strong>: conversion to the new <strong>UE5M3</strong> scale format (F2FP.SATFINITE.UE5M3), an FP4 conversion that applies a power-of-two scale on the way (F2FP.SATFINITE.E2M1 with a <em>.SCALE_BY_C</em> modifier), FP4 conversion with <em>round-toward-zero</em>, and a matrix multiply that <strong>decompresses its weights from a lookup table</strong> (UTCQMMA.LUTB). We come back to all four.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/p/a-wonderful-dive-into-how-fp4-works?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/p/a-wonderful-dive-into-how-fp4-works?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><div><hr></div><h2>K</h2><p>Spec sheets quote FP4 and FP8 throughput as separate numbers, and <strong>the instruction descriptors that CUTLASS builds for each architecture show where the ratio comes from</strong> (CUTLASS main branch, October 2026). A dense block-scaled FP4 multiply reduces <strong>K = 64</strong> elements per instruction on B200, <strong>96 on B300</strong> and <strong>128 on Rubin</strong>. </p><p>A dense FP8 multiply reduces 32, 32 and 64. The ratios, <strong>2, 3 and 2</strong>, are <strong>exactly the ratios of dense FP4 to dense FP8 on the spec sheets</strong>: 10 and 5 PFLOPS on GB200, 15 and 5 on GB300, 35 and 17.5 on Rubin (Figure 10).</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!B7Bb!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F595210c0-8555-405e-aaf1-f981a9c97f01_1760x1060.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!B7Bb!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F595210c0-8555-405e-aaf1-f981a9c97f01_1760x1060.png 424w, https://substackcdn.com/image/fetch/$s_!B7Bb!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F595210c0-8555-405e-aaf1-f981a9c97f01_1760x1060.png 848w, https://substackcdn.com/image/fetch/$s_!B7Bb!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F595210c0-8555-405e-aaf1-f981a9c97f01_1760x1060.png 1272w, https://substackcdn.com/image/fetch/$s_!B7Bb!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F595210c0-8555-405e-aaf1-f981a9c97f01_1760x1060.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!B7Bb!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F595210c0-8555-405e-aaf1-f981a9c97f01_1760x1060.png" width="1456" height="877" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/595210c0-8555-405e-aaf1-f981a9c97f01_1760x1060.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:877,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Figure 10. Elements reduced per dense MMA instruction for FP4 and FP8, by architecture, against the ratio of dense spec-sheet throughputs.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 10. Elements reduced per dense MMA instruction for FP4 and FP8, by architecture, against the ratio of dense spec-sheet throughputs." title="Figure 10. Elements reduced per dense MMA instruction for FP4 and FP8, by architecture, against the ratio of dense spec-sheet throughputs." srcset="https://substackcdn.com/image/fetch/$s_!B7Bb!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F595210c0-8555-405e-aaf1-f981a9c97f01_1760x1060.png 424w, https://substackcdn.com/image/fetch/$s_!B7Bb!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F595210c0-8555-405e-aaf1-f981a9c97f01_1760x1060.png 848w, https://substackcdn.com/image/fetch/$s_!B7Bb!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F595210c0-8555-405e-aaf1-f981a9c97f01_1760x1060.png 1272w, https://substackcdn.com/image/fetch/$s_!B7Bb!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F595210c0-8555-405e-aaf1-f981a9c97f01_1760x1060.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Figure 10.</strong> Elements reduced per dense MMA instruction for FP4 and FP8, by architecture, against the ratio of dense spec-sheet throughputs. <a href="https://www.claudeusercontent.com/?domain=claude.ai&amp;parentOrigin=https%3A%2F%2Fclaude.ai&amp;errorReportingMode=parent&amp;formattedSpreadsheets=true#">save PNG</a></figcaption></figure></div><p><strong>B300&#8217;s extra FP4 throughput is one bit.</strong> In the 32-bit Blackwell <em>instruction descriptor</em> for block-scaled FP4, <strong>bit 31 selects the K size</strong>. A 0 means dense K = 64 or sparse K = 128; a 1 means <strong>dense K = 96, with the sparse version marked invalid</strong>. </p><p>That single comment in mma_sm100_desc.hpp explains why GB300&#8217;s <strong>dense FP4 rose by half</strong> over GB200, from 10 to 15 PFLOPS, while <strong>its sparse FP4 stayed at 20</strong>: the fast dense mode has no sparse form, so sparse code still runs the K = 128 path both chips share. </p><p>Rubin widens the field to <strong>two bits</strong>, split between bits 3 and 31: dense K of 64, 96 or 128, with sparse forms at 128 and 192 and none for the K = 128 mode. Rubin&#8217;s densest sparse mode is therefore <strong>1.5 times its densest dense mode</strong>. </p><p>NVIDIA&#8217;s Rubin architecture overview describes the <strong>50 PFLOPS</strong> headline as <em>sparse</em> NVFP4 (as reported by Guru3D, July 2026), and 1.5 times the dense 35 is 52.5, so <strong>the descriptor and the spec sheet agree to within 5%</strong>. Rubin&#8217;s FP4 multiply also requires <strong>both operands to be K-major</strong>; there is no transposed form.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Making FP4 on the fly</h2><p><strong>Weights are quantized once, offline. Activations are quantized at every step</strong> by the kernel that produces them, and that costs instructions. We wrote the quantization of one block in PTX the way a <em>fused epilogue</em> would: absolute maximum, scale, scale encoding and decoding, multiply, convert, pack. </p><p>We assembled it with ptxas -O3 for four targets and counted SASS instructions, excluding loads, stores, address arithmetic and control flow (Figure 11).</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!NlGq!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa6785055-0883-437a-a300-482e8f1752cf_1760x1140.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!NlGq!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa6785055-0883-437a-a300-482e8f1752cf_1760x1140.png 424w, https://substackcdn.com/image/fetch/$s_!NlGq!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa6785055-0883-437a-a300-482e8f1752cf_1760x1140.png 848w, https://substackcdn.com/image/fetch/$s_!NlGq!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa6785055-0883-437a-a300-482e8f1752cf_1760x1140.png 1272w, https://substackcdn.com/image/fetch/$s_!NlGq!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa6785055-0883-437a-a300-482e8f1752cf_1760x1140.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!NlGq!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa6785055-0883-437a-a300-482e8f1752cf_1760x1140.png" width="1456" height="943" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a6785055-0883-437a-a300-482e8f1752cf_1760x1140.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:943,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Figure 11. SASS instructions per element to quantize one block of activations, split by instruction class, by format and target.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 11. SASS instructions per element to quantize one block of activations, split by instruction class, by format and target." title="Figure 11. SASS instructions per element to quantize one block of activations, split by instruction class, by format and target." srcset="https://substackcdn.com/image/fetch/$s_!NlGq!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa6785055-0883-437a-a300-482e8f1752cf_1760x1140.png 424w, https://substackcdn.com/image/fetch/$s_!NlGq!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa6785055-0883-437a-a300-482e8f1752cf_1760x1140.png 848w, https://substackcdn.com/image/fetch/$s_!NlGq!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa6785055-0883-437a-a300-482e8f1752cf_1760x1140.png 1272w, https://substackcdn.com/image/fetch/$s_!NlGq!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa6785055-0883-437a-a300-482e8f1752cf_1760x1140.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Figure 11.</strong> SASS instructions per element to quantize one block of activations, split by instruction class, by format and target.</figcaption></figure></div><p><strong>NVFP4 costs 2.38 instructions per element</strong> on B200, B300 and Rubin alike. For 16 values that is seven three-input <em>FMNMX3</em> and one FMNMX for the maximum, <strong>18 FMUL</strong>, <strong>10 F2FP conversions</strong>, one <em>MUFU</em> reciprocal for the decoded scale and one half-precision add. <strong>MXFP4 costs 2.16 per element</strong>, because its scale is a power of two and its reciprocal takes three integer instructions. </p><p><strong>The two scale-rounding rules compile to the same count.</strong> Consumer Blackwell pays <strong>2.81 and 2.59</strong>, because it has no three-input FMNMX3 and needs 15 comparisons for 16 values. </p><p><strong>Rubin&#8217;s fused conversion removes the multiplies for MX formats</strong> entirely: <strong>1.19 instructions per element</strong>. NVFP4 gets <strong>no fused path</strong>, because a scale that is not a power of two cannot be applied by adjusting an exponent.</p><p>Whether 2.4 instructions per element matter depends on <strong>how much tensor-core work each quantized value feeds</strong>. A quantized activation is reused for every one of the <strong>N output columns</strong> of the GEMM that consumes it. </p><p>If the quantization cannot overlap with other work, its share of the GEMM&#8217;s time is about <strong>I x F / (256 x S x f x N)</strong>, where I is instructions per element, F the dense FP4 rate, S the number of SMs and f the clock, since each SM issues <em>128 lane-instructions per cycle</em> and each multiply-add counts as two FLOPs. </p><p>At an assumed <strong>1.9 GHz</strong> that is about <strong>330/N</strong> for NVFP4 on B200, <strong>460/N</strong> on GB300 and <strong>760/N</strong> on Rubin: <strong>4.6%, 6.4% and 10.6% at N = 7168</strong>, the hidden size of DeepSeek-V3-class models. Rubin&#8217;s fused conversion halves its own figure for MX formats. </p><p><strong>Tensor cores have grown faster than the general-purpose ALUs around them</strong>, and quantization is one of the places where the difference shows. When AMD profiled MXFP4 and MXFP6 serving on MI355X, the standalone activation-quantization step was <strong>measurable overhead</strong>, and the team wrote <strong>a dedicated kernel</strong> to remove it (AMD ROCm blog, June 2026).</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>What Rubin adds</h2><p><strong>The first addition is a scale format.</strong> <strong>UE5M3</strong> has no sign bit, <em>five exponent bits with a bias of 15</em> and three mantissa bits; CUTLASS defines its range as 0 to <strong>114,688</strong>, with subnormals and a NaN code (include/cutlass/float8.h). It <strong>keeps E4M3&#8217;s three mantissa bits</strong>, so inside a block it behaves like NVFP4&#8217;s scale, but it covers <strong>about 34 binades against 18</strong>. </p><p>In Blackwell&#8217;s block-scaled instruction descriptor the scale format is <strong>one bit</strong>, E4M3 or E8M0; in Rubin&#8217;s it is <strong>two bits</strong>, and the value 2 means UE5M3 (mma_sm107_desc.hpp). </p><p>Conversions to UE5M3 <strong>assemble only for sm_107a</strong>, and we assembled both round-to-nearest and round-up forms.</p><p>The extra range matters because of <strong>NVFP4&#8217;s second level</strong>. We quantized Gaussian tensors in blocks of 16 FP4 values <strong>with no tensor scale at all</strong>, sweeping the standard deviation from 10<sup>-8</sup> to 10<sup>4</sup> (Figure 12). </p><p><strong>With UE4M3 block scales the result stays within 0.5 dB of the best only for standard deviations between 0.032 and 562.</strong> Weights of large models sit near <em>one over the square root of the hidden size</em>, about 0.011 at a width of 8,192, and a tensor at 0.01 <strong>loses 2.4 dB</strong> with UE4M3 scales and no tensor scale. </p><p>RoBERTa-base, with a width of 768 and a standard deviation of 0.053, happens to sit inside the window. That is why NVFP4 carries its FP32 tensor scale. <strong>With UE5M3 the window runs from 10<sup>-4</sup> to at least 10<sup>4</sup></strong>, the end of our sweep. </p><p>NVIDIA has not said that UE5M3 exists to retire the second scale, but on these numbers <strong>it could</strong>, and kernels would save a global reduction and a multiply per tensor. The IF4 authors proposed spending the unused sign bit of the E4M3 scale on a choice between FP4 and INT4 grids; <strong>UE5M3 spends that bit on range</strong> instead.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!8m5M!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa754a7e3-cc24-4f22-bc44-b09e2de78d62_1760x1140.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!8m5M!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa754a7e3-cc24-4f22-bc44-b09e2de78d62_1760x1140.png 424w, https://substackcdn.com/image/fetch/$s_!8m5M!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa754a7e3-cc24-4f22-bc44-b09e2de78d62_1760x1140.png 848w, https://substackcdn.com/image/fetch/$s_!8m5M!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa754a7e3-cc24-4f22-bc44-b09e2de78d62_1760x1140.png 1272w, https://substackcdn.com/image/fetch/$s_!8m5M!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa754a7e3-cc24-4f22-bc44-b09e2de78d62_1760x1140.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!8m5M!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa754a7e3-cc24-4f22-bc44-b09e2de78d62_1760x1140.png" width="1456" height="943" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a754a7e3-cc24-4f22-bc44-b09e2de78d62_1760x1140.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:943,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Figure 12. SQNR of FP4 in blocks of 16 with no tensor-level scale, as the tensor's standard deviation varies, for three scale formats.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 12. SQNR of FP4 in blocks of 16 with no tensor-level scale, as the tensor's standard deviation varies, for three scale formats." title="Figure 12. SQNR of FP4 in blocks of 16 with no tensor-level scale, as the tensor's standard deviation varies, for three scale formats." srcset="https://substackcdn.com/image/fetch/$s_!8m5M!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa754a7e3-cc24-4f22-bc44-b09e2de78d62_1760x1140.png 424w, https://substackcdn.com/image/fetch/$s_!8m5M!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa754a7e3-cc24-4f22-bc44-b09e2de78d62_1760x1140.png 848w, https://substackcdn.com/image/fetch/$s_!8m5M!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa754a7e3-cc24-4f22-bc44-b09e2de78d62_1760x1140.png 1272w, https://substackcdn.com/image/fetch/$s_!8m5M!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa754a7e3-cc24-4f22-bc44-b09e2de78d62_1760x1140.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Figure 12.</strong> SQNR of FP4 in blocks of 16 with no tensor-level scale, as the tensor&#8217;s standard deviation varies, for three scale formats.</figcaption></figure></div><p><strong>The second addition is a weight format with no fixed grid.</strong> PTX ISA 9.4 adds a <strong>.decompress::lut::b</strong> qualifier to tcgen05.mma. </p><p>In this mode, according to SemiAnalysis&#8217;s description, each weight is <strong>a 3-bit index into a table of eight E4M3 values</strong>, one table per <strong>8 x 64 tile</strong> of the weight matrix, so storage is 3 + 64/512 = <strong>3.125 bits per weight</strong>; the table lives in <em>tensor memory</em> and the lookup happens inside the multiply. </p><p>In the assembler the instruction <strong>exists only for sm_107a</strong> and becomes <strong>UTCQMMA.LUTB</strong>, so the multiply itself <strong>runs on the FP8 path</strong>. A <em>block-scaled variant</em> adds E8M0 scales for every 32 weights.</p><p>We fitted eight-entry codebooks by <em>k-means</em> on each group of 512 values, rounded the entries to E4M3 and measured. On Gaussian data the 3.125-bit format reaches <strong>14.76 dB</strong>, above NVFP3 at 3.5 bits (14.31 dB) and slightly above an ideal 3-bit quantizer that knows the variance (14.62 dB). </p><p>On RoBERTa&#8217;s weights, with codebooks shared over 8 x 64 tiles as in the hardware, the plain format reaches <strong>13.87 dB</strong>, 0.34 dB below NVFP3 at 3.5 bits (14.21 dB): <em>real weights change scale from row to row</em>, and one table per tile follows that less well than a scale per 16 values. </p><p><strong>The block-scaled variant, at 3.375 bits, recovers the loss and reaches 14.23 dB</strong>, just above NVFP3 with an eighth of a bit less. Against NVFP4 the plain format gives up 5.7 dB on Gaussian data and 6.6 dB on RoBERTa for 1.375 fewer bits. At 6.02 dB per bit a fair exchange would cost 8.3 dB, so <strong>per stored bit the lookup table is still the more efficient weight format</strong>. </p><p>It would be <strong>the wrong tool for activations</strong>, where one table cannot follow the local scale (11.73 dB on our proxy, below NVFP3&#8217;s 13.86), and the hardware offers it only for the B operand anyway.</p><p><strong>The third addition is an operand type.</strong> CUTLASS lists <strong>an unsigned 8-bit E5M3 type</strong> among the A and B operand types of Rubin&#8217;s FP8 multiply (the CuTe DSL tcgen05 tables). An unsigned operand only makes sense for values that are never negative. </p><p>In a transformer the obvious candidate is <strong>the softmax output that multiplies V in attention</strong>, which E4M3 represents poorly below 2<sup>-9</sup>. <em>That reading is ours</em>; the code says only that the type exists. </p><p>SemiAnalysis also reports that Rubin&#8217;s <em>tensor memory</em> grows from <strong>512 to 576 columns</strong>, with the extra columns meant for <strong>block scale factors</strong>, which on Blackwell compete with accumulators for space.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Who buys which half</h2><p><strong>FP4 sells two different things.</strong> <strong>Fewer bytes per weight</strong> make decoding faster and models smaller, and that half works whatever the precision of the activations. </p><p><strong>Faster arithmetic</strong> needs both operands in FP4 and the block-scaled FP4 instructions, and that half needs <strong>W4A4</strong>. The open models of the past fourteen months <strong>bought mostly the first half</strong> (Table 1).</p><blockquote><p><em>FP4 sells two things: fewer bytes and faster arithmetic. Only W4A4 collects both.</em></p></blockquote><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!7zi9!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65d7e6e1-efa1-4fee-8779-915f51caad4d_1760x847.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!7zi9!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65d7e6e1-efa1-4fee-8779-915f51caad4d_1760x847.png 424w, https://substackcdn.com/image/fetch/$s_!7zi9!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65d7e6e1-efa1-4fee-8779-915f51caad4d_1760x847.png 848w, https://substackcdn.com/image/fetch/$s_!7zi9!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65d7e6e1-efa1-4fee-8779-915f51caad4d_1760x847.png 1272w, https://substackcdn.com/image/fetch/$s_!7zi9!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65d7e6e1-efa1-4fee-8779-915f51caad4d_1760x847.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!7zi9!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65d7e6e1-efa1-4fee-8779-915f51caad4d_1760x847.png" width="1456" height="701" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/65d7e6e1-efa1-4fee-8779-915f51caad4d_1760x847.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:701,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Table 1. Weight and activation formats of the main FP4 releases, and what each collects on Blackwell tensor cores.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Table 1. Weight and activation formats of the main FP4 releases, and what each collects on Blackwell tensor cores." title="Table 1. Weight and activation formats of the main FP4 releases, and what each collects on Blackwell tensor cores." srcset="https://substackcdn.com/image/fetch/$s_!7zi9!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65d7e6e1-efa1-4fee-8779-915f51caad4d_1760x847.png 424w, https://substackcdn.com/image/fetch/$s_!7zi9!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65d7e6e1-efa1-4fee-8779-915f51caad4d_1760x847.png 848w, https://substackcdn.com/image/fetch/$s_!7zi9!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65d7e6e1-efa1-4fee-8779-915f51caad4d_1760x847.png 1272w, https://substackcdn.com/image/fetch/$s_!7zi9!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65d7e6e1-efa1-4fee-8779-915f51caad4d_1760x847.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Table 1.</strong> Weight and activation formats of the main FP4 releases, and what each collects on Blackwell tensor cores.</figcaption></figure></div><p>OpenAI&#8217;s <em>gpt-oss</em> quantizes its MoE weights, <strong>more than 90% of the parameters</strong>, to MXFP4 at <strong>4.25 bits per parameter</strong> (model card, arXiv:2508.10925). </p><p>Moonshot moved from <strong>INT4 in K2 Thinking</strong> to <strong>MXFP4 weights with MXFP8 activations in K3</strong>, applied with <em>quantization-aware training</em> from the supervised fine-tuning stage onward and a group size of 32 (Kimi K3 model card, July 2026).</p><p>On Blackwell that combination runs through .kind::mxf8f6f4 <strong>on the FP8 path</strong>. On an <strong>MI325X</strong>, which has no FP4 matrix instructions, vLLM <strong>converts K3&#8217;s MXFP4 experts to group-32 INT4 at load time</strong> and serves them with BF16-by-INT4 kernels (Shakudo, September 2026).</p><p><strong>DeepSeek V4 is the most careful case.</strong> Its report applies <strong>MXFP4 quantization-aware training</strong> to the MoE experts and to the query-key path of the attention indexer, then <strong>dequantizes the expert weights to FP8</strong> for computation, and notes that this is <strong>lossless</strong> as long as the ratio between the largest and smallest FP4 sub-block scale inside each 128 x 128 FP8 block stays under a threshold (arXiv:2606.19348). The report does not state the threshold, so <strong>we enumerated it</strong>. </p><p>Every E2M1 value times 2<sup>d</sup> is exactly representable in E4M3 if and only if d lies between -8 and 6, so the sub-block exponents inside a tile may span <strong>14 binades, a ratio of 16,384</strong>, or 11 binades if FP8 subnormals are to be avoided. </p><p>Gaussian weights with log-normal row scales span <strong>4.6 binades on average</strong> and 6 at most in our simulation, so <strong>the condition is loose</strong> and V4&#8217;s empirical check unsurprising.</p><p><strong>NVIDIA&#8217;s own NVFP4 checkpoints are the exception.</strong> They quantize activations too, by <em>post-training quantization</em>, and NVIDIA reports <strong>DeepSeek-R1-0528 within one point of FP8 on six of seven benchmarks</strong>, with AIME 2024 two points higher.</p><p>AMD&#8217;s measurements show <strong>what the second half is worth when it is bought without help</strong>. On one MI355X, W4A4 MXFP4 delivered <strong>5.1% more total tokens per second</strong> than FP8 on Llama-3.1-8B and <strong>10.0% more</strong> on Qwen3.6-27B, from <strong>twice the peak arithmetic</strong>. </p><p>If only GEMM time halved, those speedups imply that GEMMs were <strong>9.6% and 18.2% of the FP8 step</strong>, which says more about the rest of the step than about FP4. </p><p>Accuracy moved the other way: GSM8K on Llama-3.1-8B fell from <strong>80.44 with FP8 to 62.55 with MXFP4</strong>. <strong>MXFP6 activations with MXFP4 weights recovered 76.42</strong> at 2.8% less throughput than MXFP4, because <strong>MI355X runs FP6 at the FP4 rate</strong>. </p><p>NVIDIA runs FP6 at the FP8 rate, so that particular trade does not exist on Blackwell.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>The critical batch</h2><p>A decode step reads every active weight once and performs two FLOPs per weight for each sequence in the batch. Below a <strong>critical batch B* = P x b / (2 x BW)</strong>, where <em>P</em> is the dense FLOP rate, <em>b</em> the bytes per weight and <em>BW</em> the memory bandwidth, the step is <strong>limited by bandwidth</strong>; above it, <strong>by arithmetic</strong>. </p><p>Earlier pieces in this publication found B* nearly <strong>independent of precision</strong>, because P doubles when b halves. <strong>Block scales and Blackwell Ultra both break that</strong> (Figure 13).</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!qbXw!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5efc8980-7426-423a-af40-fce99d2a987d_1760x1260.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!qbXw!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5efc8980-7426-423a-af40-fce99d2a987d_1760x1260.png 424w, https://substackcdn.com/image/fetch/$s_!qbXw!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5efc8980-7426-423a-af40-fce99d2a987d_1760x1260.png 848w, https://substackcdn.com/image/fetch/$s_!qbXw!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5efc8980-7426-423a-af40-fce99d2a987d_1760x1260.png 1272w, https://substackcdn.com/image/fetch/$s_!qbXw!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5efc8980-7426-423a-af40-fce99d2a987d_1760x1260.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!qbXw!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5efc8980-7426-423a-af40-fce99d2a987d_1760x1260.png" width="1456" height="1042" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/5efc8980-7426-423a-af40-fce99d2a987d_1760x1260.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1042,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Figure 13. Critical batch by GPU and scheme, from dense spec-sheet throughput and HBM bandwidth.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 13. Critical batch by GPU and scheme, from dense spec-sheet throughput and HBM bandwidth." title="Figure 13. Critical batch by GPU and scheme, from dense spec-sheet throughput and HBM bandwidth." srcset="https://substackcdn.com/image/fetch/$s_!qbXw!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5efc8980-7426-423a-af40-fce99d2a987d_1760x1260.png 424w, https://substackcdn.com/image/fetch/$s_!qbXw!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5efc8980-7426-423a-af40-fce99d2a987d_1760x1260.png 848w, https://substackcdn.com/image/fetch/$s_!qbXw!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5efc8980-7426-423a-af40-fce99d2a987d_1760x1260.png 1272w, https://substackcdn.com/image/fetch/$s_!qbXw!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5efc8980-7426-423a-af40-fce99d2a987d_1760x1260.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Figure 13.</strong> Critical batch by GPU and scheme, from dense spec-sheet throughput and HBM bandwidth.</figcaption></figure></div><p><strong>Scales first.</strong> NVFP4 weights cost <strong>0.5625 bytes, not 0.5</strong>, so on GB200 B* rises from 312.5 with FP8 to <strong>351.6 with W4A4 NVFP4</strong>, 12.5% higher purely from scale bytes; MXFP4&#8217;s smaller overhead gives 332.0. </p><p><strong>Then GB300</strong>: dense FP4 is three times dense FP8 there, so W4A4 NVFP4 pushes B* to <strong>527.3, 1.69 times the FP8 value</strong>. <strong>The FP4-weight, FP8-arithmetic schemes go the other way.</strong> </p><p>Halving the bytes without speeding up the arithmetic roughly <strong>halves B*</strong>, to <strong>166</strong> on GB200 and GB300 and 211 on Rubin, and Rubin&#8217;s lookup-table weights take it to <strong>155</strong>. Those formats make <strong>small batches much cheaper</strong> and become compute-bound at about half the batch. <strong>W4A4 keeps large batches cheap</strong>, and on GB300 it is worth three times FP8 arithmetic.</p><p>At batch 1, weight traffic alone caps decoding of a <strong>70B dense model</strong> at <strong>57 tokens per second in BF16</strong> on 8 TB/s of HBM, <strong>114 in FP8</strong>, <strong>203 in NVFP4</strong> and <strong>215 in MXFP4</strong>; on Rubin&#8217;s 22 TB/s, at <strong>559 in NVFP4</strong> and <strong>805 with lookup-table weights</strong>. <em>Attention reads and synchronization come on top of these ceilings.</em></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>What to use where</h2><p>The measurements reduce to a handful of situations (Table 2). The table is <em>our reading of the evidence in this piece</em>, not a benchmark of anyone&#8217;s serving stack, and it applies to <strong>post-training quantization</strong>: models trained with quantization in the loop have already chosen their grid.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!G-H2!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2c7855d5-356f-45bf-b19b-3275591aa8db_1760x1094.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!G-H2!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2c7855d5-356f-45bf-b19b-3275591aa8db_1760x1094.png 424w, https://substackcdn.com/image/fetch/$s_!G-H2!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2c7855d5-356f-45bf-b19b-3275591aa8db_1760x1094.png 848w, https://substackcdn.com/image/fetch/$s_!G-H2!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2c7855d5-356f-45bf-b19b-3275591aa8db_1760x1094.png 1272w, https://substackcdn.com/image/fetch/$s_!G-H2!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2c7855d5-356f-45bf-b19b-3275591aa8db_1760x1094.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!G-H2!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2c7855d5-356f-45bf-b19b-3275591aa8db_1760x1094.png" width="1456" height="905" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2c7855d5-356f-45bf-b19b-3275591aa8db_1760x1094.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:905,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Table 2. What the measurements favour, by hardware and workload, for post-training quantization.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Table 2. What the measurements favour, by hardware and workload, for post-training quantization." title="Table 2. What the measurements favour, by hardware and workload, for post-training quantization." srcset="https://substackcdn.com/image/fetch/$s_!G-H2!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2c7855d5-356f-45bf-b19b-3275591aa8db_1760x1094.png 424w, https://substackcdn.com/image/fetch/$s_!G-H2!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2c7855d5-356f-45bf-b19b-3275591aa8db_1760x1094.png 848w, https://substackcdn.com/image/fetch/$s_!G-H2!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2c7855d5-356f-45bf-b19b-3275591aa8db_1760x1094.png 1272w, https://substackcdn.com/image/fetch/$s_!G-H2!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2c7855d5-356f-45bf-b19b-3275591aa8db_1760x1094.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Table 2.</strong> What the measurements favour, by hardware and workload, for post-training quantization.</figcaption></figure></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Where this could be wrong</h2><ul><li><p><strong>Synthetic data.</strong> Every SQNR here except RoBERTa&#8217;s comes from <em>synthetic distributions</em>. The activation proxy matches AMD&#8217;s three numbers, but those are layer means over <strong>one tensor type in two models</strong>; attention inputs, residual streams and gradients have other shapes.</p></li><li><p><strong>One family, one recipe.</strong> The perplexity law is fitted on <strong>Qwen3.5 with round-to-nearest W4A4</strong>, where weight and activation noise mix and our single proxy cannot separate them. Models trained with quantization in the loop, which includes gpt-oss, DeepSeek V4 and Kimi K3, <strong>learn around their grid</strong>, and the law should not be applied to them. The constant also <strong>moves between models</strong>: AMD&#8217;s numbers put it near 8.5 for Llama-3.1-8B and near 2.8 for Qwen3.6-27B. <em>The rank correlation of 0.986 is the more robust number.</em></p></li><li><p><strong>Rates from names.</strong> We infer that FP4 through .kind::f8f6f4 runs at FP8 speed <strong>from the opcode it becomes</strong> and from NVIDIA&#8217;s documentation, <em>not from a timed kernel</em>. The K values come from <strong>CUTLASS source, not from hardware</strong>, and a descriptor comment is a strong hint rather than a measurement.</p></li><li><p><strong>Our codebooks.</strong> The lookup-table numbers depend on <em>plain k-means with E4M3 entries</em>. NVIDIA&#8217;s own fitting, or quantization-aware training, would do better, so <strong>our LUT figures are a floor</strong>.</p></li><li><p><strong>Assumed clocks.</strong> The quantization-overhead estimate assumes <strong>1.9 GHz</strong>, NVIDIA&#8217;s listed SM counts (148 for B200, 160 for B300, 224 for Rubin) and <strong>no overlap</strong> with other work. Real epilogues hide part of the cost.</p></li><li><p><strong>Intent.</strong> UE5M3 without a tensor scale is <em>our reading of what the format allows</em>, not NVIDIA&#8217;s stated purpose, and the softmax use of the unsigned E5M3 operand is <em>a guess from the type&#8217;s existence</em>.</p></li><li><p><strong>Moving spec sheets.</strong> Rubin&#8217;s numbers come from announcements and partner documents <strong>months before broad availability</strong>, and NVIDIA&#8217;s own figures for B300 range from 13 to 15 PFLOPS of dense FP4 depending on the system.</p></li></ul><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/p/a-wonderful-dive-into-how-fp4-works?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/p/a-wonderful-dive-into-how-fp4-works?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><div><hr></div><h2>Predictions</h2><ol><li><p>By <strong>30 June 2027</strong>, at least one open-weight model above 100 billion parameters will ship its primary checkpoint in <strong>a format only Rubin runs natively</strong>: FP4 with UE5M3 scales, or 3-bit lookup-table weights.</p></li><li><p>Through the first PTX ISA release that adds <strong>Feynman (sm_140)</strong>, tcgen05.mma.kind::i8 <strong>will not assemble</strong> for sm_103a or sm_107a.</p></li><li><p>By <strong>31 March 2027</strong>, vLLM or SGLang will merge <strong>a Rubin MoE kernel</strong> that reads expert weights through .decompress::lut::b.</p></li><li><p>Moonshot&#8217;s next flagship after K3 will again pair <strong>4-bit weights with 8-bit activations</strong> rather than move to W4A4.</p></li><li><p>By <strong>31 December 2027</strong>, a revision of the <strong>OCP Microscaling specification</strong> will add a non-power-of-two scale type or change the default rounding of E8M0 scales.</p></li></ol><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Confidence dossier</h2><p>Every number in the text carries a tier. <strong>A</strong>: <em>primary source</em> (vendor documentation, archived paper, model card). <strong>B</strong>: <em>vendor source code or reputable secondary reporting</em>. <strong>C</strong>: <em>our derivation</em> from A or B sources, with stated assumptions. <strong>D</strong>: <em>our inference</em>. <strong>M</strong>: <em>our own measurement</em>.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!ChfQ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcb70a6cb-273b-47ba-b59b-1a6d49bf3782_1760x2246.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!ChfQ!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcb70a6cb-273b-47ba-b59b-1a6d49bf3782_1760x2246.png 424w, https://substackcdn.com/image/fetch/$s_!ChfQ!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcb70a6cb-273b-47ba-b59b-1a6d49bf3782_1760x2246.png 848w, https://substackcdn.com/image/fetch/$s_!ChfQ!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcb70a6cb-273b-47ba-b59b-1a6d49bf3782_1760x2246.png 1272w, https://substackcdn.com/image/fetch/$s_!ChfQ!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcb70a6cb-273b-47ba-b59b-1a6d49bf3782_1760x2246.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!ChfQ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcb70a6cb-273b-47ba-b59b-1a6d49bf3782_1760x2246.png" width="1456" height="1858" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/cb70a6cb-273b-47ba-b59b-1a6d49bf3782_1760x2246.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1858,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Table 3. Confidence dossier, claims 1 to 23.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Table 3. Confidence dossier, claims 1 to 23." title="Table 3. Confidence dossier, claims 1 to 23." srcset="https://substackcdn.com/image/fetch/$s_!ChfQ!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcb70a6cb-273b-47ba-b59b-1a6d49bf3782_1760x2246.png 424w, https://substackcdn.com/image/fetch/$s_!ChfQ!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcb70a6cb-273b-47ba-b59b-1a6d49bf3782_1760x2246.png 848w, https://substackcdn.com/image/fetch/$s_!ChfQ!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcb70a6cb-273b-47ba-b59b-1a6d49bf3782_1760x2246.png 1272w, https://substackcdn.com/image/fetch/$s_!ChfQ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcb70a6cb-273b-47ba-b59b-1a6d49bf3782_1760x2246.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Table 3.</strong> Confidence dossier, claims 1 to 23.</figcaption></figure></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!3zh9!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F24bcb52c-6a39-4e1a-8c3d-36201ef333f5_1760x2280.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!3zh9!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F24bcb52c-6a39-4e1a-8c3d-36201ef333f5_1760x2280.png 424w, https://substackcdn.com/image/fetch/$s_!3zh9!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F24bcb52c-6a39-4e1a-8c3d-36201ef333f5_1760x2280.png 848w, https://substackcdn.com/image/fetch/$s_!3zh9!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F24bcb52c-6a39-4e1a-8c3d-36201ef333f5_1760x2280.png 1272w, https://substackcdn.com/image/fetch/$s_!3zh9!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F24bcb52c-6a39-4e1a-8c3d-36201ef333f5_1760x2280.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!3zh9!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F24bcb52c-6a39-4e1a-8c3d-36201ef333f5_1760x2280.png" width="1456" height="1886" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/24bcb52c-6a39-4e1a-8c3d-36201ef333f5_1760x2280.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1886,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Table 4. Confidence dossier, claims 24 to 46.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Table 4. Confidence dossier, claims 24 to 46." title="Table 4. Confidence dossier, claims 24 to 46." srcset="https://substackcdn.com/image/fetch/$s_!3zh9!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F24bcb52c-6a39-4e1a-8c3d-36201ef333f5_1760x2280.png 424w, https://substackcdn.com/image/fetch/$s_!3zh9!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F24bcb52c-6a39-4e1a-8c3d-36201ef333f5_1760x2280.png 848w, https://substackcdn.com/image/fetch/$s_!3zh9!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F24bcb52c-6a39-4e1a-8c3d-36201ef333f5_1760x2280.png 1272w, https://substackcdn.com/image/fetch/$s_!3zh9!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F24bcb52c-6a39-4e1a-8c3d-36201ef333f5_1760x2280.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Table 4.</strong> Confidence dossier, claims 24 to 46.</figcaption></figure></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!9TZk!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff6eeec3a-0e0b-4db2-8397-0b14c85b6c75_1760x2314.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!9TZk!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff6eeec3a-0e0b-4db2-8397-0b14c85b6c75_1760x2314.png 424w, https://substackcdn.com/image/fetch/$s_!9TZk!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff6eeec3a-0e0b-4db2-8397-0b14c85b6c75_1760x2314.png 848w, https://substackcdn.com/image/fetch/$s_!9TZk!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff6eeec3a-0e0b-4db2-8397-0b14c85b6c75_1760x2314.png 1272w, https://substackcdn.com/image/fetch/$s_!9TZk!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff6eeec3a-0e0b-4db2-8397-0b14c85b6c75_1760x2314.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!9TZk!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff6eeec3a-0e0b-4db2-8397-0b14c85b6c75_1760x2314.png" width="1456" height="1914" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f6eeec3a-0e0b-4db2-8397-0b14c85b6c75_1760x2314.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1914,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Table 5. Confidence dossier, claims 47 to 69.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Table 5. Confidence dossier, claims 47 to 69." title="Table 5. Confidence dossier, claims 47 to 69." srcset="https://substackcdn.com/image/fetch/$s_!9TZk!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff6eeec3a-0e0b-4db2-8397-0b14c85b6c75_1760x2314.png 424w, https://substackcdn.com/image/fetch/$s_!9TZk!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff6eeec3a-0e0b-4db2-8397-0b14c85b6c75_1760x2314.png 848w, https://substackcdn.com/image/fetch/$s_!9TZk!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff6eeec3a-0e0b-4db2-8397-0b14c85b6c75_1760x2314.png 1272w, https://substackcdn.com/image/fetch/$s_!9TZk!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff6eeec3a-0e0b-4db2-8397-0b14c85b6c75_1760x2314.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Table 5.</strong> Confidence dossier, claims 47 to 69.</figcaption></figure></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/p/a-wonderful-dive-into-how-fp4-works?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/p/a-wonderful-dive-into-how-fp4-works?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><div><hr></div><h2>Sources</h2><ol><li><p>Open Compute Project, <em>OCP Microscaling Formats (MX) Specification v1.0</em>, September 2023.</p></li><li><p>B. Darvish Rouhani et al., <em>Microscaling Data Formats for Deep Learning</em>, <a href="https://arxiv.org/abs/2310.10537">arXiv:2310.10537</a>.</p></li><li><p>E. Alvarez et al., NVIDIA, <em><a href="https://developer.nvidia.com/blog/introducing-nvfp4-for-efficient-and-accurate-low-precision-inference/">Introducing NVFP4 for Efficient and Accurate Low-Precision Inference</a></em>, 24 June 2025.</p></li><li><p>NVIDIA, <em>Pretraining Large Language Models with NVFP4</em>, <a href="https://arxiv.org/abs/2509.25149">arXiv:2509.25149</a>, v2, 4 March 2026.</p></li><li><p>A. Mishra, D. Stosic, S. Layton, P. Micikevicius, <em>Recipes for Pre-training LLMs with MXFP8</em>, <a href="https://arxiv.org/abs/2506.08027">arXiv:2506.08027</a>, v2, 18 August 2025.</p></li><li><p>V. Egiazarian et al., <em>Bridging the Gap Between Promise and Performance for Microscaling FP4 Quantization</em>, ICLR 2026, <a href="https://arxiv.org/abs/2509.23202">arXiv:2509.23202</a>.</p></li><li><p>J. Cook et al., <em>Four Over Six: More Accurate NVFP4 Quantization with Adaptive Block Scaling</em>, <a href="https://arxiv.org/abs/2512.02010">arXiv:2512.02010</a>, v5, 9 May 2026.</p></li><li><p>J. Cook et al., <em>Adaptive Block-Scaled Data Types</em>, <a href="https://arxiv.org/abs/2603.28765">arXiv:2603.28765</a>, 30 March 2026.</p></li><li><p>S. Atre, B. Bao, S. Tiwari, A. Sirasao, AMD, <em><a href="https://rocm.blogs.amd.com/artificial-intelligence/w4a6-quant-mm/README.html">MXFP6 and MXFP4 Mixed Precision for Accelerating Dense LLMs on AMD Instinct MI355X</a></em>, 26 June 2026.</p></li><li><p>T.-R. Chen, <em>Spec Sheets Are Not Kernels: An ISA- and Source-Level Audit of INT8 Availability on NVIDIA Blackwell Ultra</em>, <a href="https://arxiv.org/abs/2608.11693">arXiv:2608.11693</a>, August 2026.</p></li><li><p>NVIDIA, <em>Inside NVIDIA Blackwell Ultra: The Chip Powering the AI Factory Era</em>, NVIDIA technical blog.</p></li><li><p>NVIDIA, <a href="https://www.nvidia.com/en-us/data-center/hgx/">HGX platform specifications</a> (HGX B300 and HGX B200).</p></li><li><p>Verda, <em><a href="https://verda.com/blog/gb300-nvl72-architecture">GB300 NVL72 architecture</a></em> (per-GPU dense FP4, FP8 and INT8 for GB300, GB200, B300, B200).</p></li><li><p>The Register, <em><a href="https://www.theregister.com/2026/01/05/ces_rubin_nvidia/">Nvidia unpacks its Vera Rubin CPUs and GPUs at CES</a></em>, 5 January 2026.</p></li><li><p>NVIDIA, <em><a href="https://www.gigabyte.com/FileUpload/Global/WebPage/1052/NVIDIA_2026_1H_V2.pdf">2026 1H product guide</a></em> (Vera Rubin NVL72 dense specifications), distributed by Gigabyte.</p></li><li><p>Guru3D, <em><a href="https://www.guru3d.com/story/nvidia-details-rubin-gpu-architecture-with-hbm4-nvlink-6-and-336-billion-transistors/">NVIDIA details Rubin GPU architecture with HBM4, NVLink 6 and 336 billion transistors</a></em>, July 2026.</p></li><li><p>SemiAnalysis, <em><a href="https://newsletter.semianalysis.com/p/vera-rubin-nvl72-vs-gb200-nvl72-inference">Vera Rubin NVL72 vs GB200 NVL72? Inference TCO and Architecture Analysis</a></em>, 23 July 2026.</p></li><li><p>NVIDIA, <em><a href="https://docs.nvidia.com/cuda/developer-preview/13.4/parallel-thread-execution/index.html">PTX ISA 9.4</a></em>, CUDA 13.4.</p></li><li><p>NVIDIA, <em><a href="https://docs.nvidia.com/deeplearning/tensorrt/latest/getting-started/release-notes-11/11.3.0.html">TensorRT 11.3.0 release notes</a></em> (Vera Rubin, compute capability 10.7).</p></li><li><p>NVIDIA, <a href="https://github.com/NVIDIA/cutlass">CUTLASS</a>, main branch, October 2026: include/cute/arch/mma_sm100_desc.hpp, mma_sm107_desc.hpp, mma_sm107_umma.hpp, include/cutlass/float8.h, python/CuTeDSL/cutlass/cute/nvgpu/tcgen05/mma.py.</p></li><li><p>Apache TVM, PTX instruction table and <a href="https://github.com/apache/tvm/pull/20271">pull request 20271</a>, September 2026.</p></li><li><p>OpenAI, <em>gpt-oss-120b and gpt-oss-20b Model Card</em>, <a href="https://arxiv.org/abs/2508.10925">arXiv:2508.10925</a>.</p></li><li><p>DeepSeek-AI, <em>DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence</em>, <a href="https://arxiv.org/abs/2606.19348">arXiv:2606.19348</a>.</p></li><li><p>vLLM, <em><a href="https://docs.vllm.ai/en/latest/api/vllm/model_executor/models/deepseek_v4/">DeepSeek V4 model documentation</a></em> (MXFP4 experts with UE8M0 scales).</p></li><li><p>Moonshot AI, <em>Kimi K3 model card</em>, as reported by <a href="https://inferencex.semianalysis.com/model/kimi-k3">SemiAnalysis InferenceX</a> and the <a href="https://huggingface.co/vessl/Kimi-K3-W4AFP8">VESSL W4AFP8 card</a>.</p></li><li><p>LMSYS, <em><a href="https://www.lmsys.org/blog/2026-01-26-int4-qat">INT4 quantization-aware training</a></em>, 26 January 2026.</p></li><li><p>Shakudo, <em><a href="https://www.shakudo.io/blog/kimi-k3-amd-gpu-deployment">How We Deployed Kimi K3 on AMD GPUs</a></em>, updated 4 September 2026.</p></li><li><p>heise online, <em><a href="https://heise.de/-10443422">AMD Instinct MI350X and MI355X</a></em> (FP6 and FP4 at twice the FP8 rate).</p></li><li><p>S. Lloyd, <em>Least Squares Quantization in PCM</em>, IEEE Transactions on Information Theory, 1982; J. Max, <em>Quantizing for Minimum Distortion</em>, IRE Transactions on Information Theory, 1960.</p></li><li><p>Two-level MXFP8 training with an FP32 tensor scale, <a href="https://arxiv.org/abs/2511.05811">arXiv:2511.05811</a>, November 2025.</p></li><li><p>Y. Liu et al., <em>RoBERTa</em>, <a href="https://arxiv.org/abs/1907.11692">arXiv:1907.11692</a>; Explosion, en_core_web_trf 3.8.0.</p></li><li><p>S. Ashkboos et al., <em>QuaRot</em>, <a href="https://arxiv.org/abs/2404.00456">arXiv:2404.00456</a>; Z. Liu et al., <em>SpinQuant</em>, <a href="https://arxiv.org/abs/2405.16406">arXiv:2405.16406</a>.</p></li></ol><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Corrections</h2><p><strong>Instruction count.</strong> An early count of the NVFP4 quantization cost, 2.31 instructions per element, left out a half-precision add and the byte permutes. The counting rule used here, every instruction except memory access, address arithmetic and control flow, gives 2.38.</p><p><strong>Codebook tiles.</strong> Our first pass over real weights shared lookup-table codebooks across 512 consecutive values in memory rather than across 8 x 64 tiles of the matrix, as the hardware does. The figures here use tiles.</p><p><strong>Weight scale window.</strong> A draft said typical weights fall outside the no-tensor-scale window of UE4M3 while quoting RoBERTa&#8217;s standard deviation of 0.053, which is inside it. The text now uses the scale of large-model weights, about 0.01, where the loss is 2.4 dB.</p><p><strong>Format count.</strong> A draft said the perplexity table holds 22 formats and that the fit starts at 3.25 bits. The table holds 23, we simulated 22, and the fit covers 3.5 to 6.5 bits.</p><p><strong>Error message.</strong> A draft quoted an assembler error message from memory. It was removed; the text now describes the rejection without quoting it.</p><p><strong>AMD quantization kernel.</strong> The previous version removed a statement that AMD wrote a dedicated activation-quantization kernel, because we could not re-verify it at the time. AMD&#8217;s post says exactly that, and the statement is back with its source.</p><p><strong>INT8 figures.</strong> The previous version cited a secondary page for B300&#8217;s INT8 throughput that, on re-reading, lists INT8 at the FP8 rate. The figures now come from Chen&#8217;s audit and from per-GPU tables that match NVIDIA&#8217;s datasheet.</p><p><strong>Kimi K2 Thinking.</strong> The previous version gave the scale type of Kimi K2 Thinking&#8217;s INT4 weights as BF16. Our sources confirm the grouping of 32 and the 16-bit activations, not the scale type, and the text now says so.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Appendix: reproducing the measurements</h2><p>The assembler and disassembler come from <strong>NVIDIA&#8217;s PyPI </strong>wheels and run on any x86 Linux machine without a GPU.</p><pre><code><code>pip download --no-deps nvidia-cuda-nvcc nvidia-cuda-nvdisasm      # 13.4.92
python3 -m zipfile -e nvidia_cuda_nvcc-13.4.92-*.whl cuda/
python3 -m zipfile -e nvidia_cuda_nvdisasm-13.4.92-*.whl cuda/
chmod +x cuda/nvidia/cu13/bin/*
cuda/nvidia/cu13/bin/ptxas -arch=sm_107a -o k.cubin k.ptx
cuda/nvidia/cu13/bin/nvdisasm -c k.cubin | grep -E "UTC|F2FP|OMMA|QMMA"
</code></code></pre><p>A probe is one instruction inside a minimal kernel. This one asks for an <strong>NVFP4 block-scaled multiply</strong>; changing the target line and the qualifiers gives every cell of Figure 9.</p><pre><code><code>.version 9.4
.target sm_107a
.address_size 64
.visible .entry k(.param .u64 out, .param .u64 in, .param .u32 idesc)
{
  .reg .b32 %r&lt;8&gt;; .reg .b64 %rd&lt;8&gt;; .reg .pred %p&lt;2&gt;;
  ld.param.u64 %rd1, [out]; ld.param.u64 %rd2, [in]; ld.param.u32 %r1, [idesc];
  ld.global.u32 %r2, [%rd2]; ld.global.u32 %r3, [%rd2+4];
  ld.global.u64 %rd3, [%rd2+8]; ld.global.u64 %rd4, [%rd2+16];
  setp.ne.u32 %p1, %r1, 0;
  tcgen05.mma.cta_group::1.kind::mxf4nvf4.block_scale.block16
      [%r2], %rd3, %rd4, %r1, [%r3], [%r3], %p1;
  ret;
}
</code></code></pre><p>The<strong> reference semantics</strong> of NVFP4 fit in a few lines of NumPy. Every other format in this piece changes the element grid, the block size or the way the scale is rounded.</p><pre><code><code>import numpy as np
E2M1 = np.array([0, .5, 1, 1.5, 2, 3, 4, 6])
E4M3 = np.array(sorted({m / 8 * 2.0**-6 for m in range(8)} |
                       {(1 + m / 8) * 2.0**(e - 7) for e in range(1, 16) for m in range(8)}))[:-1]

def rne(a, grid):                      # nearest grid point, ties to even code, saturating
    i = np.clip(np.searchsorted(grid, a), 1, len(grid) - 1)
    lo, hi = grid[i - 1], grid[i]
    up = (hi - a &lt; a - lo) | ((hi - a == a - lo) &amp; (i % 2 == 0))
    return np.where(a &gt;= grid[-1], grid[-1], np.where(up, hi, lo))

def nvfp4(x, block=16):
    g = np.abs(x).max() / (6 * 448)                                   # FP32 tensor scale
    xb = x.reshape(-1, block)
    s = rne(np.abs(xb).max(1, keepdims=True) / (6 * g), E4M3) * g     # E4M3 block scale
    s1 = np.where(s &gt; 0, s, 1.0)
    return (np.sign(xb) * rne(np.abs(xb) / s1, E2M1) * s).reshape(x.shape)</code></code></pre><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p>]]></content:encoded></item><item><title><![CDATA[Why Cached Tokens Cost 10% and Vanish in 5 Minutes]]></title><description><![CDATA[Every agent resends its whole context at every step, and every provider bills the repeat at a discount with a timer attached. The discount and the timer both come from one single ratio.]]></description><link>https://www.thesoftwarefrontier.com/p/why-cached-tokens-cost-10-and-vanish</link><guid isPermaLink="false">https://www.thesoftwarefrontier.com/p/why-cached-tokens-cost-10-and-vanish</guid><dc:creator><![CDATA[Lorenzo Bradanini]]></dc:creator><pubDate>Tue, 29 Sep 2026 14:48:41 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!T0Ee!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F28f8d841-4a52-4691-a7cb-4d64addcfbdc_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!T0Ee!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F28f8d841-4a52-4691-a7cb-4d64addcfbdc_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!T0Ee!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F28f8d841-4a52-4691-a7cb-4d64addcfbdc_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!T0Ee!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F28f8d841-4a52-4691-a7cb-4d64addcfbdc_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!T0Ee!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F28f8d841-4a52-4691-a7cb-4d64addcfbdc_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!T0Ee!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F28f8d841-4a52-4691-a7cb-4d64addcfbdc_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!T0Ee!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F28f8d841-4a52-4691-a7cb-4d64addcfbdc_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/28f8d841-4a52-4691-a7cb-4d64addcfbdc_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2402261,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/209898768?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F28f8d841-4a52-4691-a7cb-4d64addcfbdc_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!T0Ee!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F28f8d841-4a52-4691-a7cb-4d64addcfbdc_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!T0Ee!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F28f8d841-4a52-4691-a7cb-4d64addcfbdc_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!T0Ee!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F28f8d841-4a52-4691-a7cb-4d64addcfbdc_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!T0Ee!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F28f8d841-4a52-4691-a7cb-4d64addcfbdc_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>Introduction</h2><p>Every coding agent works the same way underneath. At each step it sends the model everything it has so far, the system prompt, the tool definitions, every file it opened and<strong> every command it ran</strong>, followed by one new tool result, and the model reads all of it to choose one small next action. </p><p>In the session we price later in this piece, the <strong>context starts at 12,000 tokens and ends at 85,500</strong> after fifty steps, and the model is sent 2.44 million input tokens in order to write 20,000.</p><p>Almost <em>no one of that is billed at full price</em>. The provider has seen those tokens before, has kept what it computed the first time, and charges the repeat at a fraction of the normal rate. On Anthropic&#8217;s current price list the <strong>fraction is 10 percent for most models</strong>, 5 percent for Opus 5.5 and 2.5 percent for Fable 5.1 and Mythos 5.1. On DeepSeek it is 2 percent. </p><p>The discount comes with a clock: five minutes, thirty minutes, an hour or a day, depending on the provider and on what you pay.</p><p>I wanted to know where those tokens go between two calls, who pays to keep them, and why the numbers are what they are. Why a tenth and not a half. <strong>Why five minutes and not five hours. </strong>Why DeepSeek can charge a fiftieth while Google charges rent by the hour. </p><p>We ended up with <strong>one ratio that answers most of these questions</strong>, a replay of public production traces that answers most of the rest, and a view of where inference hardware is heading that neither of us expected when we started. It also produced a handful of rules that, in the agent session we model, cut the bill in half.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>The short version</h2><ol><li><p><strong>What is cached.</strong> Not your text: the key and value tensors every attention layer computes for every token. Depending on the architecture that is 890 bytes to 512 KiB per token, so a 100,000-token agent context occupies anywhere from 89 MB to 52 GB.</p></li><li><p><strong>Why a read costs a tenth.</strong> A hit costs the provider a reload and some rent, not a recompute, and on an H100 even a single SSD returns cache faster than the GPU can regenerate it for every model here released since 2024. How cheap a read can be depends on how much compute each stored byte saves, which varies 700-fold across the models we measured; DeepSeek&#8217;s V4.1-Flash saves the most per byte and charges 2 percent.</p></li><li><p><strong>Why five minutes.</strong> GPU memory is the most expensive place to keep anything: an H100 fills its own HBM with fresh cache in about a minute on a 70B-class model, host memory pays for roughly an hour and SSD for days. Five minutes is also where most reuse happens; in Kimi&#8217;s production traces it captures 79 percent of the possible reuse for chat and 93 percent for agents.</p></li><li><p><strong>Why writes cost extra.</strong> Keeping bytes is rent, but only for entries nobody reads again, because every read refreshes the lifetime for free. Converted to dollars per gigabyte-hour, most providers price that rent like GPU memory, far above what DRAM costs.</p></li><li><p><strong>What to do about it.</strong> Caching everything pays only if more than 22 percent of the tokens that pass through the cache are hits (53 percent with the one-hour lifetime). Gaps of 5 to 36 minutes are cheapest to cross with a zero-output ping every four and a half minutes, longer ones with the one-hour lifetime. In our fifty-step agent session with nine long tool calls, that is the difference between $1.88 and $0.97.</p></li><li><p><strong>Where it is going.</strong> A Rubin GPU fills its own memory with fresh cache 2.4 times faster than a B200. That is why NVIDIA is adding petabytes of flash to every pod and why DeepSeek shrank its cache 437-fold in three years.</p></li></ol><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>The price list in September 2026</h2><p><em>Four providers sell the same physical thing in three different ways.</em></p><p>Anthropic charges a premium to write the cache and a small fee to read it. <strong>A write costs 1.25 times the base input price</strong> if the entry should live five minutes and 2 times if it should live an hour. A read costs 0.1 times the input price on most models, 0.05 on Opus 5.5 and 0.025 on Fable 5.1 and Mythos 5.1<sup>. </sup></p><p>Every read refreshes the lifetime at no charge, and the clock starts when the request that wrote or read the entry begins, not when its response finishes, so <strong>a</strong> <strong>response that streams for four minutes leaves about one minute</strong> for the follow-up. Anthropic also states that the cached key and value tensors are held in memory only and never stored at rest.</p><p>OpenAI has moved to the same structure. From GPT-5.6 onward a write costs 1.25 times the uncached input rate, a read 0.1 times, and an entry stays eligible for at least thirty minutes after its latest write or reuse<sup>. </sup></p><p>Earlier models charged nothing to write and kept entries in what OpenAI describes as volatile GPU memory for roughly five to ten minutes of inactivity; an <strong>extended policy kept them for up to 24 hours by offloading the key and value tensors</strong> to storage local to the GPU once memory filled up. The two companies differ in one detail that matters for agents at scale: Anthropic does not count cache hits against rate limits, and OpenAI does.</p><p>Google separates reading from keeping. Reading an explicit Gemini cache costs a tenth of the input price, and <strong>keeping the cache alive costs rent:</strong> $4.50 per million tokens per hour on Gemini 3.1 Pro and $0.50 on Gemini 3.8 Flash, a rate the price page schedules to double on January 1, 2027. </p><p>Gemini also caches without being asked: according to Google&#8217;s caching guide, implicit caching is on by default for Gemini 2.5 and newer models, discounting repeated prefixes of at least 4,096 tokens on the current 3.x models, with no storage fee and <strong>no guarantee of a hit.</strong></p><p>DeepSeek charges for neither writing nor keeping. A hit on its <strong>current flash model costs $0.003 per million tokens off-peak </strong>against $0.15 for a miss, both doubling during two weekday windows<sup>. </sup>DeepSeek has offered context caching since August 2024 and serves its hits from a cache on disk.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!oZso!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F56264f62-c4e2-4ab7-9b56-aa769d023e90_1634x950.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!oZso!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F56264f62-c4e2-4ab7-9b56-aa769d023e90_1634x950.png 424w, https://substackcdn.com/image/fetch/$s_!oZso!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F56264f62-c4e2-4ab7-9b56-aa769d023e90_1634x950.png 848w, https://substackcdn.com/image/fetch/$s_!oZso!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F56264f62-c4e2-4ab7-9b56-aa769d023e90_1634x950.png 1272w, https://substackcdn.com/image/fetch/$s_!oZso!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F56264f62-c4e2-4ab7-9b56-aa769d023e90_1634x950.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!oZso!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F56264f62-c4e2-4ab7-9b56-aa769d023e90_1634x950.png" width="1456" height="847" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/56264f62-c4e2-4ab7-9b56-aa769d023e90_1634x950.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:847,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Horizontal bar chart of cache-read prices: DeepSeek V4.1-Flash 2 percent, Claude Fable 5.1 and Mythos 5.1 2.5 percent, Claude Opus 5.5 5 percent, Claude Sonnet 5, GPT-5.6 and Gemini 10 percent, DeepSeek R1 in 2025 25.5 percent.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Horizontal bar chart of cache-read prices: DeepSeek V4.1-Flash 2 percent, Claude Fable 5.1 and Mythos 5.1 2.5 percent, Claude Opus 5.5 5 percent, Claude Sonnet 5, GPT-5.6 and Gemini 10 percent, DeepSeek R1 in 2025 25.5 percent." title="Horizontal bar chart of cache-read prices: DeepSeek V4.1-Flash 2 percent, Claude Fable 5.1 and Mythos 5.1 2.5 percent, Claude Opus 5.5 5 percent, Claude Sonnet 5, GPT-5.6 and Gemini 10 percent, DeepSeek R1 in 2025 25.5 percent." srcset="https://substackcdn.com/image/fetch/$s_!oZso!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F56264f62-c4e2-4ab7-9b56-aa769d023e90_1634x950.png 424w, https://substackcdn.com/image/fetch/$s_!oZso!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F56264f62-c4e2-4ab7-9b56-aa769d023e90_1634x950.png 848w, https://substackcdn.com/image/fetch/$s_!oZso!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F56264f62-c4e2-4ab7-9b56-aa769d023e90_1634x950.png 1272w, https://substackcdn.com/image/fetch/$s_!oZso!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F56264f62-c4e2-4ab7-9b56-aa769d023e90_1634x950.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Figure 1.</strong> Cache-read price as a share of the fresh input price at list prices in September 2026. The line under each name gives the write charge or the rent. DeepSeek R1 in February 2025 is shown for reference.</figcaption></figure></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!_dBp!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcc41a763-242c-4cbb-8edc-28bdc8e8c290_1643x1075.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!_dBp!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcc41a763-242c-4cbb-8edc-28bdc8e8c290_1643x1075.png 424w, https://substackcdn.com/image/fetch/$s_!_dBp!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcc41a763-242c-4cbb-8edc-28bdc8e8c290_1643x1075.png 848w, https://substackcdn.com/image/fetch/$s_!_dBp!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcc41a763-242c-4cbb-8edc-28bdc8e8c290_1643x1075.png 1272w, https://substackcdn.com/image/fetch/$s_!_dBp!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcc41a763-242c-4cbb-8edc-28bdc8e8c290_1643x1075.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!_dBp!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcc41a763-242c-4cbb-8edc-28bdc8e8c290_1643x1075.png" width="1456" height="953" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/cc41a763-242c-4cbb-8edc-28bdc8e8c290_1643x1075.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:953,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Table of read multipliers, write charges, lifetimes and storage charges for Anthropic, OpenAI, Google and DeepSeek.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Table of read multipliers, write charges, lifetimes and storage charges for Anthropic, OpenAI, Google and DeepSeek." title="Table of read multipliers, write charges, lifetimes and storage charges for Anthropic, OpenAI, Google and DeepSeek." srcset="https://substackcdn.com/image/fetch/$s_!_dBp!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcc41a763-242c-4cbb-8edc-28bdc8e8c290_1643x1075.png 424w, https://substackcdn.com/image/fetch/$s_!_dBp!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcc41a763-242c-4cbb-8edc-28bdc8e8c290_1643x1075.png 848w, https://substackcdn.com/image/fetch/$s_!_dBp!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcc41a763-242c-4cbb-8edc-28bdc8e8c290_1643x1075.png 1272w, https://substackcdn.com/image/fetch/$s_!_dBp!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcc41a763-242c-4cbb-8edc-28bdc8e8c290_1643x1075.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Table 1.</strong> How each provider bills its cache in September 2026.</figcaption></figure></div><p>Read one way, the market has converged: <em>a read at a tenth of the input price is the default at Anthropic, OpenAI and Google.</em> Read another way, it has already broken below that point, with the newest models from Anthropic and DeepSeek at 5, 2.5 and 2 percent. <strong>The charges for writing and keeping have not converged at all.</strong> Both patterns come from what the provider is physically holding.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>What the provider keeps</h2><p>A cached prompt is not text. When a transformer processes a token, every attention layer computes a key vector and a value vector for it, and every later token in the sequence reads them. <strong>Keeping those vectors lets the next request with the same beginning</strong> skip computing them again. OpenAI&#8217;s documentation says it directly: the cache holds key and value tensors, not the tokens themselves.</p><p>How many bytes that is depends on the attention design, and the range is wide. With grouped-query attention, the design used by Llama 3, Qwen3, GLM and MiniMax, every token costs two vectors per layer per key-value head:</p><pre><code>bytes per token = 2 &#215; layers &#215; kv_heads &#215; head_dim &#215; bytes per element
Llama-3.1-70B:    2 &#215; 80 &#215; 8 &#215; 128 &#215; 2 (BF16) = 327,680 bytes = 320 KiB</code></pre><p>A 100,000-token agent context for that model is 32.8 GB,<strong> two fifths of an H100&#8217;s memory</strong> for a single conversation. Multi-head latent attention, which DeepSeek introduced in V2<sup> </sup>and Moonshot reused in Kimi K2, stores one compressed latent per layer instead of per-head keys and values: 512 values plus 64 for position. </p><p>In BF16 that is 1,152 bytes per layer and 70,272 bytes per token across DeepSeek-V3&#8217;s 61 layers, so the same 100,000 tokens take 7.0 GB. DeepSeek-V3.2 adds a <strong>small FP8 key per token for its sparse-attention indexer </strong>but stores the latent itself in FP8, which by our layout accounting comes to 48,068 bytes.</p><p>The newest designs compress across tokens as well. DeepSeek&#8217;s V4 family keeps compressed and heavily compressed attention caches, <strong>about 3,600 bytes per token for V4-Flash</strong> by our reconstruction in August, and DeepSeek&#8217;s own report puts V4-Flash at 7 percent of V3.2&#8217;s cache at a million tokens.</p><p><strong>V4.1-Flash, released on September 10, goes further</strong>: its decoder layers take their global cache from a projection of the final encoder state instead of computing their own, and the cache is stored in FP4 with one scale per 16 channels. The model card gives 890 bytes per token. A 100,000-token context is 89 MB.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!XEVx!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F51915238-b395-495c-964f-e278b8f601f5_1634x1140.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!XEVx!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F51915238-b395-495c-964f-e278b8f601f5_1634x1140.png 424w, https://substackcdn.com/image/fetch/$s_!XEVx!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F51915238-b395-495c-964f-e278b8f601f5_1634x1140.png 848w, https://substackcdn.com/image/fetch/$s_!XEVx!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F51915238-b395-495c-964f-e278b8f601f5_1634x1140.png 1272w, https://substackcdn.com/image/fetch/$s_!XEVx!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F51915238-b395-495c-964f-e278b8f601f5_1634x1140.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!XEVx!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F51915238-b395-495c-964f-e278b8f601f5_1634x1140.png" width="1456" height="1016" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/51915238-b395-495c-964f-e278b8f601f5_1634x1140.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1016,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Log-scale bar chart of cache bytes per token, from Llama-2-7B at 524,288 bytes to DeepSeek-V4.1-Flash at 890 bytes.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Log-scale bar chart of cache bytes per token, from Llama-2-7B at 524,288 bytes to DeepSeek-V4.1-Flash at 890 bytes." title="Log-scale bar chart of cache bytes per token, from Llama-2-7B at 524,288 bytes to DeepSeek-V4.1-Flash at 890 bytes." srcset="https://substackcdn.com/image/fetch/$s_!XEVx!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F51915238-b395-495c-964f-e278b8f601f5_1634x1140.png 424w, https://substackcdn.com/image/fetch/$s_!XEVx!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F51915238-b395-495c-964f-e278b8f601f5_1634x1140.png 848w, https://substackcdn.com/image/fetch/$s_!XEVx!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F51915238-b395-495c-964f-e278b8f601f5_1634x1140.png 1272w, https://substackcdn.com/image/fetch/$s_!XEVx!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F51915238-b395-495c-964f-e278b8f601f5_1634x1140.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Figure 2.</strong> Cache bytes per token for fifteen open models, computed from their published configurations, with the size of a 100,000-token context. Sliding-window and linear-attention layers also keep a fixed per-sequence state that is not included.</figcaption></figure></div><p><strong>Two things in that chart are easy to miss</strong>. The first is that cache size does not follow model size. Llama-2-7B, with plain multi-head attention, keeps 512 KiB per token, more than Llama-3.1-70B; MiniMax-M2, with 10 billion active parameters, keeps 248 KiB because all 62 of its layers use full attention with eight key-value heads; GLM-4.6 keeps 368 KiB across 92 layers. </p><p>The second is how far one lab has pushed. DeepSeek&#8217;s first model, the 67B release of late 2023, kept 389,120 bytes per token: 95 layers, eight key-value heads of dimension 128. </p><p>Divide by 890 and the result is 437.2, which is the 437-fold reduction DeepSeek reports in the V4.1 model card. <strong>We computed the V1 figure from its architecture</strong> and got the same ratio, a useful check on both numbers.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!7ij3!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb1312a9c-38d6-41ee-9c85-cce00f6217fe_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!7ij3!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb1312a9c-38d6-41ee-9c85-cce00f6217fe_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!7ij3!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb1312a9c-38d6-41ee-9c85-cce00f6217fe_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!7ij3!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb1312a9c-38d6-41ee-9c85-cce00f6217fe_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!7ij3!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb1312a9c-38d6-41ee-9c85-cce00f6217fe_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!7ij3!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb1312a9c-38d6-41ee-9c85-cce00f6217fe_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b1312a9c-38d6-41ee-9c85-cce00f6217fe_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Line chart of DeepSeek cache bytes per token: 389,120 in 2023, about 70,000 in 2024, 48,068 in 2025, 3,621 in April 2026 and 890 in September 2026.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Line chart of DeepSeek cache bytes per token: 389,120 in 2023, about 70,000 in 2024, 48,068 in 2025, 3,621 in April 2026 and 890 in September 2026." title="Line chart of DeepSeek cache bytes per token: 389,120 in 2023, about 70,000 in 2024, 48,068 in 2025, 3,621 in April 2026 and 890 in September 2026." srcset="https://substackcdn.com/image/fetch/$s_!7ij3!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb1312a9c-38d6-41ee-9c85-cce00f6217fe_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!7ij3!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb1312a9c-38d6-41ee-9c85-cce00f6217fe_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!7ij3!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb1312a9c-38d6-41ee-9c85-cce00f6217fe_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!7ij3!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb1312a9c-38d6-41ee-9c85-cce00f6217fe_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Figure 3.</strong> Cache per token across DeepSeek generations at the precision each was released with. V4-Flash is our reconstruction and V4.1-Flash comes from the model card. The first and last points differ by a factor of 437.</figcaption></figure></div><p>Sliding windows and linear attention trade differently. gpt-oss-120b keeps full history in only 18 of its 36 layers and a 128-token window in the rest; <strong>Qwen3-Next keeps full attention in 12 of 48 layers</strong> and a fixed-size recurrent state in the others. </p><p><em>Both shrink the per-token cache</em>, and both <strong>add a per-sequence state</strong> that can only be reused at the exact position where it was saved, which makes prefix matching harder than it is for a paged cache.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Compute saved per byte</h2><p>The size alone does not say whether keeping the cache is worth anything. What matters is the work the bytes save. <strong>Recomputing a token&#8217;s cache means running prefill for it again</strong>, which costs at least two floating-point operations per active parameter. Call that F and divide it by the bytes kept:</p><pre><code>&#961; = F / k        prefill FLOPs avoided per byte of cache kept
F = 2 &#215; active parameters in prefill      (a floor: attention adds more at long context)</code></pre><p>For <strong>Llama-3.1-70B, F is 141 GFLOP per token and k is 327,680 bytes,</strong> so &#961; is 431 thousand: every byte kept saves 431 thousand operations. DeepSeek-V3 is at 1.05 million, V4-Flash at 7.2 million, and V4.1-Flash, whose prefill runs through only 8 billion active parameters, at 18 million. </p><p>At the bottom are Llama-2-7B at 26 thousand and MiniMax-M2 at 79 thousand. Across the fifteen models in this piece &#961; spans a factor of about 700.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!nLlT!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F52c9048f-1594-4bec-bc8a-f876554a34d2_1634x1064.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!nLlT!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F52c9048f-1594-4bec-bc8a-f876554a34d2_1634x1064.png 424w, https://substackcdn.com/image/fetch/$s_!nLlT!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F52c9048f-1594-4bec-bc8a-f876554a34d2_1634x1064.png 848w, https://substackcdn.com/image/fetch/$s_!nLlT!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F52c9048f-1594-4bec-bc8a-f876554a34d2_1634x1064.png 1272w, https://substackcdn.com/image/fetch/$s_!nLlT!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F52c9048f-1594-4bec-bc8a-f876554a34d2_1634x1064.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!nLlT!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F52c9048f-1594-4bec-bc8a-f876554a34d2_1634x1064.png" width="1456" height="948" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/52c9048f-1594-4bec-bc8a-f876554a34d2_1634x1064.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:948,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Log-log scatter of prefill GFLOP per token against cache bytes per token for fifteen models, with diagonal lines of constant compute per byte.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Log-log scatter of prefill GFLOP per token against cache bytes per token for fifteen models, with diagonal lines of constant compute per byte." title="Log-log scatter of prefill GFLOP per token against cache bytes per token for fifteen models, with diagonal lines of constant compute per byte." srcset="https://substackcdn.com/image/fetch/$s_!nLlT!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F52c9048f-1594-4bec-bc8a-f876554a34d2_1634x1064.png 424w, https://substackcdn.com/image/fetch/$s_!nLlT!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F52c9048f-1594-4bec-bc8a-f876554a34d2_1634x1064.png 848w, https://substackcdn.com/image/fetch/$s_!nLlT!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F52c9048f-1594-4bec-bc8a-f876554a34d2_1634x1064.png 1272w, https://substackcdn.com/image/fetch/$s_!nLlT!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F52c9048f-1594-4bec-bc8a-f876554a34d2_1634x1064.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Figure 4.</strong> Prefill compute per token against cache bytes per token. Diagonal lines have constant &#961;, the compute saved per byte kept; models toward the upper left save more work for every byte they hold.</figcaption></figure></div><p>Mixture-of-experts made caching relatively less valuable before DeepSeek made it more valuable.<strong> Sparsity cuts compute per token without touching attention</strong>, so a model like Qwen3-235B-A22B or MiniMax-M2 carries the cache of a large dense model and the prefill cost of a small one. </p><p>DeepSeek compressed attention in the same generations in which it sparsified the experts, and every release since V2 sits further toward the upper left of the chart.</p><p>Our F is a floor. With dense attention, recomputing a long prefix also pays for attention itself, which grows with position. Averaged over a 64,000-token prefix, <strong>attention adds about 84 GFLOP per token to Llama-3.1-70B&#8217;s 141</strong>, and about 160 GFLOP to DeepSeek-V3&#8217;s 74 when V3 runs prefill without absorbing its projections. </p><p>At agent context lengths the real &#961; of dense-attention models is 1.6 to 3.2 times the floors we use, so every conclusion below gets stronger when those terms are added back.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>How fast a GPU really writes cache</h2><p><strong>Put &#961; next to a GPU </strong>and you get the quantity that ties the rest of the piece together. A GPU running prefill at peak throughput P and utilization u produces fresh cache at the rate</p><pre><code>G = P &#215; u / &#961;        bytes of new cache per second</code></pre><p>We use dense FP8 peak for P, 1,979 TFLOPS on an H100 and 4.5 PFLOPS on a B200, and u = 0.30. That is at the optimistic end of what has been published, so it is worth calibrating. <strong>SGLang&#8217;s open reproduction of DeepSeek&#8217;s serving system</strong> processed 52,300 input tokens per second per eight-GPU H100 node on 2,000-token prompts, which is about 0.26 of the FP8 peak once attention is counted. </p><p>On GB200 the same team reached 18,471 input tokens per second per GPU with FP8 experts, 0.29 of dense FP8 peak, and 26,156 with NVFP4 experts. <strong>DeepSeek&#8217;s own production prefill nodes</strong>, averaged over a day of real traffic, processed 73,700 input tokens per second per node including cache hits; removing the 56.3 percent that were hits leaves about 4,000 computed tokens per second per H800, roughly 0.16 of peak. </p><p>We keep 0.30 as the central value and carry 0.15 to 0.5 in every range below. <strong>On an H100 at our assumption, Llama-3.1-70B writes 1.38 GB of cache per second</strong>, DeepSeek-V3 0.56 GB, V4-Flash 83 MB and V4.1-Flash 33 MB.</p><p>G has one property worth stating separately. Move a model&#8217;s arithmetic from BF16 to FP8 and P doubles; store its cache in FP8 as well and k halves; G does not move. </p><p>Quantizing both sides together leaves the rate at which a GPU fills memory where it was, the <strong>same kind of precision invariance we actually found for the critical batch size in August</strong>, seen from the memory side. Quantizing only the arithmetic doubles it.</p><p>G answers three different questions. It is the break-even bandwidth for reloading, because any link faster than G delivers cached bytes faster than the GPU could regenerate them. </p><p>It sets the capacity bill for a lifetime, because holding <strong>everything a GPU produces for &#964; seconds takes G &#215; &#964; bytes</strong>. And it sets the clock for the GPU&#8217;s own memory, because HBM capacity divided by G is how long the GPU takes to fill its HBM with fresh cache.</p><h3>Reload or recompute</h3><p>The first question has a lopsided answer. At 1.38 GB/s, a single Gen5 NVMe drive, which reads at about 14 GB/s, already returns Llama-3.1-70B&#8217;s cache faster than an H100 can recompute it, and every link above it in the hierarchy is faster still. </p><p>For the <strong>MLA and compressed models</strong> the margin over a single drive is 25 to 420 times. Only models at the bottom of the &#961; scale need more than a drive: MiniMax-M2 produces 7.5 GB/s of cache on an H100 and about 67 GB/s on a Rubin GPU, a full PCIe Gen5 x16 link.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!xORI!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2b1e90d4-7433-4232-b189-d99c33987153_1634x988.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!xORI!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2b1e90d4-7433-4232-b189-d99c33987153_1634x988.png 424w, https://substackcdn.com/image/fetch/$s_!xORI!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2b1e90d4-7433-4232-b189-d99c33987153_1634x988.png 848w, https://substackcdn.com/image/fetch/$s_!xORI!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2b1e90d4-7433-4232-b189-d99c33987153_1634x988.png 1272w, https://substackcdn.com/image/fetch/$s_!xORI!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2b1e90d4-7433-4232-b189-d99c33987153_1634x988.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!xORI!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2b1e90d4-7433-4232-b189-d99c33987153_1634x988.png" width="1456" height="880" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2b1e90d4-7433-4232-b189-d99c33987153_1634x988.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:880,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Horizontal log-scale bar chart of cache production rate per GPU for seven models on H100 and Rubin, with vertical lines for NVMe, PCIe Gen5, an 800G NIC and NVLink-C2C.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Horizontal log-scale bar chart of cache production rate per GPU for seven models on H100 and Rubin, with vertical lines for NVMe, PCIe Gen5, an 800G NIC and NVLink-C2C." title="Horizontal log-scale bar chart of cache production rate per GPU for seven models on H100 and Rubin, with vertical lines for NVMe, PCIe Gen5, an 800G NIC and NVLink-C2C." srcset="https://substackcdn.com/image/fetch/$s_!xORI!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2b1e90d4-7433-4232-b189-d99c33987153_1634x988.png 424w, https://substackcdn.com/image/fetch/$s_!xORI!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2b1e90d4-7433-4232-b189-d99c33987153_1634x988.png 848w, https://substackcdn.com/image/fetch/$s_!xORI!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2b1e90d4-7433-4232-b189-d99c33987153_1634x988.png 1272w, https://substackcdn.com/image/fetch/$s_!xORI!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2b1e90d4-7433-4232-b189-d99c33987153_1634x988.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Figure 5.</strong> Fresh cache produced per second by one GPU at 30 percent of dense FP8 peak, against common link bandwidths. A link to the right of a bar returns that model&#8217;s cache faster than the GPU could recompute it. Rubin uses NVIDIA&#8217;s published specifications.</figcaption></figure></div><p>The latency tells the same story. Reloading a 100,000-token Llama-3.1-70B context, 32.8 GB, over a Gen5 x16 link at an effective 55 GB/s <strong>takes circa 0.6 seconds</strong>; recomputing it on eight H100s at our utilization takes about 3 seconds before attention is counted. </p><p>For V4.1-Flash the context is 89 MB and the reload takes a couple of milliseconds. Recomputation wins only where compute per byte is low and prefixes are short. A September characterization of SSD-backed caching for vLLM put the <strong>break-even at about 6,200 tokens for Qwen3-4B on its H100 setup.</strong> </p><h3>Hiding the reload</h3><p>A reload does not have to be waited on. The engine can load layer l + 1 of the cached prefix while <strong>it computes layer l of the new tokens</strong>, which is how the layer-wise pipelines in Mooncake and the systems that followed it work. </p><p>The reload disappears behind the computation when the link moves the cached bytes in no more time than the GPU spends on the new ones:</p><pre><code>L_cached &#215; k / W  &#8804;  L_new &#215; F / (P &#215; u)      which gives      W &#8805; (L_cached / L_new) &#215; G</code></pre><p>This is where agents are the hard case. A late agent step carries a long cached prefix and a short new suffix: <strong>84,000 tokens cached and 1,500 new</strong> in the last step of our session, a ratio of 56. </p><p>Hiding that reload for Llama-3.1-70B on an H100 needs 77 GB/s, more than a Gen5 x16 link delivers, although <strong>attention over the long prefix adds compute</strong> to the new tokens and cuts the requirement by more than half. </p><p>DeepSeek-V3 needs 32 GB/s and V4.1-Flash under 2 GB/s. The ratio of cached to new tokens is exactly what an agent pushes up, which is why coherent CPU links and small caches matter more for agents than for chat.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share The Software Frontier&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share The Software Frontier</span></a></p><div><hr></div><h2>The clock</h2><p>Reloading is relatively cheap. <strong>Keeping is certainly not free, and it is paid for by the second.</strong> An entry is worth keeping as long as the expected saving from a future hit exceeds the rent. Recomputing one byte of cache costs the GPU&#8217;s price per second divided by G, and holding one byte costs the tier&#8217;s price per byte-second. </p><p>The holding time at which the two are equal is the break-even:</p><pre><code>&#964;* = (GPU cost per second) / (G &#215; s)       s = holding cost of the tier per byte-second
&#964;*_HBM = H / G                             when HBM is priced at its share of the GPU</code></pre><p><strong>The HBM line needs one assumption</strong>. When serving is limited by memory, the normal state of a decode-heavy fleet, <em>every gigabyte of HBM holding an idle cache entry is a gigabyte not holding an active request</em>, so its price is its share of the GPU&#8217;s rent. </p><p>With that price the break-even reduces to something physical: the time the GPU takes to fill its own HBM with fresh cache. On an H100 serving Llama-3.1-70B that is 58 seconds. <strong>H/G is an upper bound, because part of HBM holds the weights</strong>: about 11 percent for Llama-3.1-70B spread over eight H100s, about 60 percent for a 671-billion-parameter FP8 model on eight H200s. The break-even shrinks in the same proportion.</p><p>For the other tiers we priced memory at 2026 levels, which are not those of a year ago. <strong>TrendForce expected conventional DRAM contract prices to rise 55 to 60 percent</strong> in the first quarter alone; by late August a 64 GB DDR5 server module cost 2.5 to 3 times its January price and a 3.84 TB enterprise NVMe drive 2.3 to 2.8 times, and TrendForce still expected both to rise in the third quarter. </p><p>We tend to assume <strong>$12 per GB of DRAM over four years</strong>, with 50 percent added for power, space and the CPU socket it hangs from, and $0.25 per GB of SSD over five years, doubled for the servers and network around it.</p><p> For the GPU we use an H100 at $2.50 an hour, inside the range specialist clouds charge in 2026.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!l-Nz!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0314bcb7-cb5d-498e-bc56-6405b759d1c8_1643x670.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!l-Nz!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0314bcb7-cb5d-498e-bc56-6405b759d1c8_1643x670.png 424w, https://substackcdn.com/image/fetch/$s_!l-Nz!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0314bcb7-cb5d-498e-bc56-6405b759d1c8_1643x670.png 848w, https://substackcdn.com/image/fetch/$s_!l-Nz!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0314bcb7-cb5d-498e-bc56-6405b759d1c8_1643x670.png 1272w, https://substackcdn.com/image/fetch/$s_!l-Nz!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0314bcb7-cb5d-498e-bc56-6405b759d1c8_1643x670.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!l-Nz!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0314bcb7-cb5d-498e-bc56-6405b759d1c8_1643x670.png" width="1456" height="594" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0314bcb7-cb5d-498e-bc56-6405b759d1c8_1643x670.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:594,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Table of holding costs: HBM 0.031 dollars per GB-hour, DRAM 0.00051, SSD 0.000011.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Table of holding costs: HBM 0.031 dollars per GB-hour, DRAM 0.00051, SSD 0.000011." title="Table of holding costs: HBM 0.031 dollars per GB-hour, DRAM 0.00051, SSD 0.000011." srcset="https://substackcdn.com/image/fetch/$s_!l-Nz!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0314bcb7-cb5d-498e-bc56-6405b759d1c8_1643x670.png 424w, https://substackcdn.com/image/fetch/$s_!l-Nz!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0314bcb7-cb5d-498e-bc56-6405b759d1c8_1643x670.png 848w, https://substackcdn.com/image/fetch/$s_!l-Nz!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0314bcb7-cb5d-498e-bc56-6405b759d1c8_1643x670.png 1272w, https://substackcdn.com/image/fetch/$s_!l-Nz!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0314bcb7-cb5d-498e-bc56-6405b759d1c8_1643x670.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Table 2.</strong> Holding costs per tier used throughout the piece.</figcaption></figure></div><p>Per byte-hour, <strong>HBM, DRAM and SSD come out at roughly 2,700 to 45 to 1.</strong> Each step down the ladder buys 45 to 60 times more holding time for the same money.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!8MmR!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61120904-2b26-4f8b-a1b3-858158b89b31_1634x1026.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!8MmR!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61120904-2b26-4f8b-a1b3-858158b89b31_1634x1026.png 424w, https://substackcdn.com/image/fetch/$s_!8MmR!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61120904-2b26-4f8b-a1b3-858158b89b31_1634x1026.png 848w, https://substackcdn.com/image/fetch/$s_!8MmR!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61120904-2b26-4f8b-a1b3-858158b89b31_1634x1026.png 1272w, https://substackcdn.com/image/fetch/$s_!8MmR!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61120904-2b26-4f8b-a1b3-858158b89b31_1634x1026.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!8MmR!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61120904-2b26-4f8b-a1b3-858158b89b31_1634x1026.png" width="1456" height="914" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/61120904-2b26-4f8b-a1b3-858158b89b31_1634x1026.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:914,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Log-time chart of break-even holding times in HBM, host DRAM and SSD for six models, from 11 seconds in HBM for MiniMax-M2 to 77 days on SSD for DeepSeek-V4.1-Flash.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Log-time chart of break-even holding times in HBM, host DRAM and SSD for six models, from 11 seconds in HBM for MiniMax-M2 to 77 days on SSD for DeepSeek-V4.1-Flash." title="Log-time chart of break-even holding times in HBM, host DRAM and SSD for six models, from 11 seconds in HBM for MiniMax-M2 to 77 days on SSD for DeepSeek-V4.1-Flash." srcset="https://substackcdn.com/image/fetch/$s_!8MmR!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61120904-2b26-4f8b-a1b3-858158b89b31_1634x1026.png 424w, https://substackcdn.com/image/fetch/$s_!8MmR!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61120904-2b26-4f8b-a1b3-858158b89b31_1634x1026.png 848w, https://substackcdn.com/image/fetch/$s_!8MmR!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61120904-2b26-4f8b-a1b3-858158b89b31_1634x1026.png 1272w, https://substackcdn.com/image/fetch/$s_!8MmR!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61120904-2b26-4f8b-a1b3-858158b89b31_1634x1026.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Figure 6.</strong> Longest holding time at which keeping a cache entry costs less than recomputing it, per tier, on an H100 at $2.50 an hour and 30 percent utilization. Markers are central estimates and whiskers span our assumption ranges; dashed lines mark lifetimes the providers sell. With a reuse probability below one, every point moves left in proportion.</figcaption></figure></div><p>For Llama-3.1-70B the ladder reads about a minute in HBM, about an hour in DRAM and about two days on SSD. <strong>For DeepSeek-V3 it is two and a half minutes</strong>, two and a half hours and four and a half days. For V4.1-Flash it is 40 minutes, 41 hours and 77 days. MiniMax-M2 gets 11 seconds, 11 minutes and 8 hours. </p><p>Not a single one of these numbers is totally accurate. Across the ranges we consider plausible (<em>utilization from 0.15, what DeepSeek&#8217;s production fleet achieves, to 0.5, an H100 from $1.50 to $4.00 an hour, the low and high ends of 2026 memory quotes, and 60 to 100 percent of HBM available to cache</em>), the <strong>HBM break-evens move by a factor of 0.36 to 2.0, the DRAM ones by 0.2 to 5.8 and the SSD ones by 0.17 to 7.1. </strong>The whiskers in the chart show those ranges; the order of the tiers and of the models never changes.</p><p>Now lay the product menu over it. Five minutes is past the HBM break-even of every grouped-query model in our set even at the top of its range, and inside the DRAM break-even of all of them at our central assumptions, which says <strong>a five-minute cache for those models belongs in host memory</strong> rather than on the GPU. </p><p>One hour sits at the DRAM break-even of a 70B-class grouped-query model: 59 minutes at the center, between 12 minutes and almost 6 hours across the ranges. At the center, <strong>twenty-four hours is past the DRAM break-even </strong>of every model here except V4.1-Flash, and pays on SSD only for models above roughly 230 thousand FLOPs per byte, which leaves out most of the sparse grouped-query models. </p><p>The providers&#8217; own descriptions match the ladder: OpenAI keeps short-lived entries in GPU memory and moves to local storage for the 24-hour policy, <strong>DeepSeek serves hits from disk</strong>, and Google charges rent for anything kept on request.</p><p>These are break-evens for an entry that is certain to be read again. With a probability p of reuse, each one shrinks by p. The probabilities are what the traces are for.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/p/why-cached-tokens-cost-10-and-vanish?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/p/why-cached-tokens-cost-10-and-vanish?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><div><hr></div><h2>What the traces say</h2><p>Moonshot published request traces from Kimi&#8217;s serving system with the Mooncake paper, which <strong>won the best paper award at FAST 2025.</strong> Each request carries an arrival time, its input and output lengths, and one hash per 512-token block, remapped so that equal hashes mean reusable cache. </p><p>The conversation trace holds 12,031 requests and the tool-and-agent trace 23,608, both sampled from one hour of production traffic. Average inputs are 12,035 and 8,596 tokens and average outputs 343 and 182, ratios of 35 and 47 to one.</p><p>We replayed both traces block by block against a cache with unlimited space whose entries expire a fixed time after their last use. With no expiry at all, 37.4 percent of the conversation trace&#8217;s input tokens and 57.1 percent of the agent trace&#8217;s could have come from cache. <strong>Those ceilings sit where others have found them. </strong></p><p>Alibaba&#8217;s study of its production traces found ideal hit ratios of 62 and 54 percent over a day, and in its API trace single-turn requests produced 97 percent of the hits. <strong>DeepSeek reported that 342 billion of the 608 billion input tokens</strong> it served in one day of February 2025, 56.3 percent, hit its on-disk cache. </p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!oEkz!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feeb4ee80-7326-462a-9863-26ea3b22cc61_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!oEkz!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feeb4ee80-7326-462a-9863-26ea3b22cc61_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!oEkz!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feeb4ee80-7326-462a-9863-26ea3b22cc61_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!oEkz!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feeb4ee80-7326-462a-9863-26ea3b22cc61_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!oEkz!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feeb4ee80-7326-462a-9863-26ea3b22cc61_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!oEkz!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feeb4ee80-7326-462a-9863-26ea3b22cc61_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/eeb4ee80-7326-462a-9863-26ea3b22cc61_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Line chart: at a five-minute lifetime the conversation trace captures 79 percent of ideal reuse and the agent trace 93 percent; both approach 100 percent by thirty minutes.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Line chart: at a five-minute lifetime the conversation trace captures 79 percent of ideal reuse and the agent trace 93 percent; both approach 100 percent by thirty minutes." title="Line chart: at a five-minute lifetime the conversation trace captures 79 percent of ideal reuse and the agent trace 93 percent; both approach 100 percent by thirty minutes." srcset="https://substackcdn.com/image/fetch/$s_!oEkz!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feeb4ee80-7326-462a-9863-26ea3b22cc61_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!oEkz!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feeb4ee80-7326-462a-9863-26ea3b22cc61_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!oEkz!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feeb4ee80-7326-462a-9863-26ea3b22cc61_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!oEkz!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feeb4ee80-7326-462a-9863-26ea3b22cc61_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Figure 7.</strong> Share of the ideal reuse captured by a sliding cache lifetime, replaying Kimi&#8217;s public conversation and tool-and-agent traces with unlimited capacity.</figcaption></figure></div><p><strong>A five-minute lifetime captures 79 percent of the ideal reuse in the conversation</strong> trace and 93 percent in the agent trace. Ten minutes captures 94 and 99 percent. Thirty minutes captures essentially all of it. Most of the value is in the first five minutes, and by thirty there is little left to buy.</p><p>The gaps between reuses show where the two workloads differ. In the agent trace <strong>about half of all reused tokens were reused </strong>within one second of their previous use (<em>52 percent over the whole trace, 50 percent in the edge-corrected sample below</em>): requests arriving together and sharing a long prefix, the signature of an agent fanning out parallel calls. </p><p>In the conversation trace that <strong>share is about 10 percent</strong>, and the median gap is about two minutes, roughly the time a person takes to read an answer and type the next question.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Na7p!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8f49bd73-b8e7-4462-b0f7-a06684c892f2_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Na7p!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8f49bd73-b8e7-4462-b0f7-a06684c892f2_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!Na7p!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8f49bd73-b8e7-4462-b0f7-a06684c892f2_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!Na7p!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8f49bd73-b8e7-4462-b0f7-a06684c892f2_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!Na7p!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8f49bd73-b8e7-4462-b0f7-a06684c892f2_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Na7p!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8f49bd73-b8e7-4462-b0f7-a06684c892f2_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/8f49bd73-b8e7-4462-b0f7-a06684c892f2_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Cumulative distribution of reuse gaps: half of agent reuse happens within one second; 75 percent of conversation reuse and 91 percent of agent reuse happens within five minutes.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Cumulative distribution of reuse gaps: half of agent reuse happens within one second; 75 percent of conversation reuse and 91 percent of agent reuse happens within five minutes." title="Cumulative distribution of reuse gaps: half of agent reuse happens within one second; 75 percent of conversation reuse and 91 percent of agent reuse happens within five minutes." srcset="https://substackcdn.com/image/fetch/$s_!Na7p!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8f49bd73-b8e7-4462-b0f7-a06684c892f2_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!Na7p!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8f49bd73-b8e7-4462-b0f7-a06684c892f2_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!Na7p!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8f49bd73-b8e7-4462-b0f7-a06684c892f2_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!Na7p!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8f49bd73-b8e7-4462-b0f7-a06684c892f2_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Figure 8.</strong> Cumulative share of reused tokens by time since the same block was last used, counting only blocks whose previous use came at least thirty minutes before the end of the one-hour trace.</figcaption></figure></div><p>Fan-out is also why Anthropic&#8217;s documentation notes that <strong>an entry becomes available</strong> only after the first response starts, and suggests waiting for it before sending parallel requests. Ten identical prefixes sent in the same second are ten cache writes, not one write and nine reads.</p><p>The trace covers one hour, which biases every gap statistic toward short gaps, so <strong>for the chart above we counted only blocks whose previous use </strong>came at least thirty minutes before the end. Among those, 75 percent of the conversation reuse and 91 percent of the agent reuse that happens within half an hour happens within the first five minutes.</p><p>Capacity is the other half of the question. An <em>LRU cache sized to hold ten minutes of the trace&#8217;s unique traffic captures 89 percent of the ideal reuse </em>for conversations and<strong> 98 percent for agents</strong>, and twenty minutes captures 98 percent of the conversation ideal. </p><p>The <strong>unique traffic arriving at a prefill GPU is exactly what that GPU computes</strong>, which is G. So the capacity a ten-minute window needs per GPU, with prefill running flat out, is G &#215; 600 seconds: 827 GB for Llama-3.1-70B on an H100, 338 GB for DeepSeek-V3 and 20 GB for V4.1-Flash.</p><p>An H100 has 80 GB of HBM, and an eight-GPU DGX H100 has 2 TB of system memory, about 256 GB per GPU. A ten-minute window for a 70B grouped-query model fits in neither, and has to live on SSD or in a pool shared across machines. <strong>For DeepSeek-V3 it spills just past DRAM</strong>. For V4.1-Flash it fits in HBM next to the weights. That comparison goes a long way toward explaining why DeepSeek can nearly give reads away and charge nothing for keeping.</p><p>It also <strong>explains why nobody lets you write everything for free.</strong> If every miss in the conversation trace were written at 1.25 times the input price and every hit read at 0.1, a five-minute cache would cut the input bill by 9 percent; on the agent trace the cut would be 36 percent. </p><p><strong>The margin is in what gets written</strong>, which is why both Anthropic and OpenAI let callers place breakpoints, and why OpenAI&#8217;s new explicit mode charges no write for anything after the last breakpoint. </p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/p/why-cached-tokens-cost-10-and-vanish/comments&quot;,&quot;text&quot;:&quot;Leave a comment&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/p/why-cached-tokens-cost-10-and-vanish/comments"><span>Leave a comment</span></a></p><div><hr></div><h2>Where the bytes live</h2><p>NVIDIA describes the <strong>hierarchy its Dynamo framework </strong>manages in four levels: G1 is GPU memory for the cache of requests in flight, G2 is host DRAM for staging, G3 is SSD inside a node and G4 is shared storage for what has to survive. </p><p>At CES in January it added a level between the last two. The CMX context memory storage platform, first shown as ICMS, puts <strong>Ethernet-attached flash managed by BlueField-4 processors</strong> into each pod, with petabytes of shared capacity, and NVIDIA claims five times the tokens per second and five times the power efficiency of general-purpose storage for this job. Its own description of the data is the best argument for a separate tier: transient, derived, and recomputable if lost.</p><p>The rest of the rack is moving the same way. A GB200 superchip pairs two Blackwell GPUs with a <strong>Grace CPU carrying up to 480 GB of LPDDR5X</strong>. A Vera Rubin NVL72 rack carries 54 TB of LPDDR5X, 1.5 TB per Vera CPU and 750 GB per GPU, reachable over a coherent NVLink-C2C link of 1.8 TB/s. </p><p>The two labs that price hits lowest built their own versions of this years ago. Mooncake pools the DRAM and SSD of Kimi&#8217;s GPU servers into one distributed cache; the paper reports<strong> more than 100 billion tokens a day</strong> on thousands of nodes, and 115 and 107 percent more requests on A800 and H800 clusters than the previous system. </p><p>DeepSeek&#8217;s <strong>3FS file system, built on SSDs and RDMA</strong>, reached 6.6 TiB/s of aggregate reads on a 180-node cluster, and its README shows cache reads for inference peaking at 40 GiB/s.</p><p>Flash has one limit that bandwidth arithmetic hides, which is wear. A prefill GPU that persisted everything it computed would write G bytes per second to flash, all day. For<strong> Llama-3.1-70B on an H100 that is 119 TB a day</strong>; a 30 TB drive rated for one full write per day absorbs 30. Endurance alone would call for four drives per GPU before capacity or bandwidth enter the picture. </p><p><strong>For DeepSeek-V3 the figure is 49 TB a day</strong>, for V4.1-Flash under 3 TB. At these write rates, admission control, writing only the blocks likely to be reused, decides whether the tier outlives its warranty. </p><p>The V4.1 model card makes the same point from the model side: its bounded replay of the sliding-window layers exists so that <strong>their cache never has to be persisted to SSD</strong>, and it cuts the persistent footprint to about an eighth of V4-Flash&#8217;s.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share The Software Frontier&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share The Software Frontier</span></a></p><div><hr></div><h2>Rubin and the context tier</h2><p>The part we did not expect is what <strong>the next GPU does to all of this. </strong>The HBM break-even is H/G, and G is proportional to peak compute, so the clock for GPU memory is set by a single hardware constant: HBM capacity per unit of compute. </p><p>In megabytes of <strong>HBM per dense FP8 TFLOPS it is 40.4 on an H100</strong>, 71.2 on an H200, 40.0 on a B200 and 37.2 on a GB200. On Rubin, with 288 GB against 17.5 dense FP8 PFLOPS on NVIDIA&#8217;s specification page, it is 16.5. NVIDIA revised that page on September 21, lowering HBM bandwidth to 19.2 TB/s and NVLink to 3 TB/s per GPU; neither enters this ratio, and capacity and compute are unchanged.</p><p>That one constant moves the whole ladder. For Llama-3.1-70B the time to fill HBM with fresh cache falls from 58 seconds on an H100 to 24 on Rubin, for DeepSeek-V3 from 142 to 58, for V4.1-Flash from 40 minutes to 16. </p><p>The capacity a ten-minute window needs per GPU grows with compute instead: <strong>7.3 TB for Llama-3.1-70B, 3.0 TB for DeepSeek-V3 </strong>and 175 GB for V4.1-Flash, against 288 GB of HBM and 750 GB of LPDDR5X per Rubin GPU.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!04Da!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F07576417-f773-44f2-9784-44b200dc093f_1634x931.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!04Da!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F07576417-f773-44f2-9784-44b200dc093f_1634x931.png 424w, https://substackcdn.com/image/fetch/$s_!04Da!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F07576417-f773-44f2-9784-44b200dc093f_1634x931.png 848w, https://substackcdn.com/image/fetch/$s_!04Da!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F07576417-f773-44f2-9784-44b200dc093f_1634x931.png 1272w, https://substackcdn.com/image/fetch/$s_!04Da!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F07576417-f773-44f2-9784-44b200dc093f_1634x931.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!04Da!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F07576417-f773-44f2-9784-44b200dc093f_1634x931.png" width="1456" height="830" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/07576417-f773-44f2-9784-44b200dc093f_1634x931.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:830,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Grouped bar chart of HBM fill time for three models on five GPUs; Rubin is fastest to fill at 24 seconds for Llama-3.1-70B.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Grouped bar chart of HBM fill time for three models on five GPUs; Rubin is fastest to fill at 24 seconds for Llama-3.1-70B." title="Grouped bar chart of HBM fill time for three models on five GPUs; Rubin is fastest to fill at 24 seconds for Llama-3.1-70B." srcset="https://substackcdn.com/image/fetch/$s_!04Da!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F07576417-f773-44f2-9784-44b200dc093f_1634x931.png 424w, https://substackcdn.com/image/fetch/$s_!04Da!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F07576417-f773-44f2-9784-44b200dc093f_1634x931.png 848w, https://substackcdn.com/image/fetch/$s_!04Da!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F07576417-f773-44f2-9784-44b200dc093f_1634x931.png 1272w, https://substackcdn.com/image/fetch/$s_!04Da!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F07576417-f773-44f2-9784-44b200dc093f_1634x931.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Figure 9.</strong> Time for a GPU running prefill at 30 percent of dense FP8 peak to fill its HBM with fresh cache, which is also the HBM break-even. The line under each GPU gives megabytes of HBM per dense FP8 TFLOPS.</figcaption></figure></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!2U9y!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e8c2b10-dde9-4e17-a39a-ed2fca6a77b0_1634x931.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!2U9y!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e8c2b10-dde9-4e17-a39a-ed2fca6a77b0_1634x931.png 424w, https://substackcdn.com/image/fetch/$s_!2U9y!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e8c2b10-dde9-4e17-a39a-ed2fca6a77b0_1634x931.png 848w, https://substackcdn.com/image/fetch/$s_!2U9y!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e8c2b10-dde9-4e17-a39a-ed2fca6a77b0_1634x931.png 1272w, https://substackcdn.com/image/fetch/$s_!2U9y!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e8c2b10-dde9-4e17-a39a-ed2fca6a77b0_1634x931.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!2U9y!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e8c2b10-dde9-4e17-a39a-ed2fca6a77b0_1634x931.png" width="1456" height="830" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2e8c2b10-dde9-4e17-a39a-ed2fca6a77b0_1634x931.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:830,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Two-panel bar chart: on H100 a ten-minute window needs 827 GB for Llama-3.1-70B, 338 GB for DeepSeek-V3 and 20 GB for V4.1-Flash; on Rubin 7.3 TB, 3.0 TB and 175 GB.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Two-panel bar chart: on H100 a ten-minute window needs 827 GB for Llama-3.1-70B, 338 GB for DeepSeek-V3 and 20 GB for V4.1-Flash; on Rubin 7.3 TB, 3.0 TB and 175 GB." title="Two-panel bar chart: on H100 a ten-minute window needs 827 GB for Llama-3.1-70B, 338 GB for DeepSeek-V3 and 20 GB for V4.1-Flash; on Rubin 7.3 TB, 3.0 TB and 175 GB." srcset="https://substackcdn.com/image/fetch/$s_!2U9y!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e8c2b10-dde9-4e17-a39a-ed2fca6a77b0_1634x931.png 424w, https://substackcdn.com/image/fetch/$s_!2U9y!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e8c2b10-dde9-4e17-a39a-ed2fca6a77b0_1634x931.png 848w, https://substackcdn.com/image/fetch/$s_!2U9y!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e8c2b10-dde9-4e17-a39a-ed2fca6a77b0_1634x931.png 1272w, https://substackcdn.com/image/fetch/$s_!2U9y!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e8c2b10-dde9-4e17-a39a-ed2fca6a77b0_1634x931.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Figure 10.</strong> Cache needed per GPU to hold ten minutes of fresh prefill output at 30 percent of dense FP8 peak, against HBM and host memory per GPU.</figcaption></figure></div><p><strong>A Rubin GPU produces cache faster than anything</strong> in its own rack can hold for more than a minute or two, unless the model&#8217;s cache is small. Its share of the rack&#8217;s host memory holds about a minute of Llama-3.1-70B&#8217;s output and<em> about two and a half minutes of DeepSeek-V3&#8217;s</em>; for V4.1-Flash it holds 43 minutes. </p><p>The flash tier in the pod is the hardware answer to that gap, and a 437-fold smaller cache is the model answer. <strong>They are substitutes</strong>, and different companies are pursuing them.</p><p>The same arithmetic casts a different light on a product that disappeared. In September 2025 <strong>NVIDIA announced Rubin CPX</strong>, a GPU with 128 GB of GDDR7 built for the compute-heavy prefill of long contexts and due at the end of 2026. It was missing from the roadmap at GTC in March. </p><p>Commentators have pointed to <strong>tight 3nm capacity and the price of GDDR7 during the memory shortage</strong>. Our reading, which is an interpretation and nothing more, is that caching also shrank the market the chip was designed for. </p><p>When half or more of the input tokens in agent traffic are cache hits, the prefill that remains is the new part of each request, and a pool of flash that keeps the rest of the context close does more for time to first token than a pool of extra prefill compute.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>How the lookup works</h2><p>No one of this helps unless a request can find its cache, and finding it is stricter than it looks. <strong>Attention is causal:</strong> the cache for position i depends on every token before i. Change one token near the start of the prompt and every block after it has a different cache, even where the text afterward is identical. </p><p>So every system identifies cache by prefix. <strong>vLLM hashes each block of tokens together</strong> with the hash of the block before it, adding extra keys for images, adapters and a per-tenant salt, and a lookup walks that chain until the first miss.</p><p>The API products expose the same chain at a coarser grain. Anthropic writes an entry only at a breakpoint, as a cumulative hash of everything before it, and a <strong>read walks back at most 20 blocks</strong> looking for an entry an earlier request wrote; it will not discover stable content behind a changing block unless something wrote an entry there. </p><p>Minimum cacheable lengths run from 512 tokens on its <em>newest models to 4,096 on Haiku 4.5</em>. OpenAI&#8217;s GPT-5.6 has a 1,024-token minimum, places an implicit breakpoint at the end of the latest eligible message, allows up to four writes per request, and <strong>checks the first two and the latest fifty explicit breakpoints on a lookup.</strong> Gemini&#8217;s implicit cache starts at 4,096 tokens on its current 3.x models. </p><p>The minimums reflect bookkeeping. Every entry carries a hash, an index record and a routing decision, and below a few hundred tokens that overhead is comparable to recomputing. <strong>OpenAI&#8217;s guide works through the trade from the customer&#8217;s side </strong>and shows that padding a short shared prefix up to the 1,024-token minimum pays off once the prefix is longer than about 102 tokens and reused often enough.</p><p>Then the request has to arrive where the cache is. OpenAI&#8217;s cache lives on individual machines, and requests are routed by a hash of the initial tokens after its <strong>hidden system content plus an optional key</strong>; above about 15 requests per minute for one prefix and key, some requests overflow to machines that do not hold the entry. </p><p><strong>Open-source stacks solve the same problem</strong> with cache-aware routers that trade some load balance for hits. For agents the difficult moment is the pause: while a tool runs, the engine is tempted to evict the conversation&#8217;s cache to make room for someone else. </p><p><em>Continuum, from Berkeley, pins the cache in GPU memory for a time-to-live predicted</em> from the tool&#8217;s expected duration, and reports large reductions in job completion time on SWE-Bench and BFCL agent workloads. </p><p>Shared caches also leak. A hit is faster than a miss, and <strong>in audits run in September and October 2024</strong>, Stanford researchers detected prompt caching at 8 of 17 API providers and sharing across users at 7 of them, OpenAI among them. </p><p>The large providers now isolate: OpenAI by organization and processing region, Anthropic by workspace on its own API. OpenAI still recommends a separate cache key per end user inside one organization, which stops one customer from probing for another&#8217;s cached prefixes.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>What the prices imply</h2><p>With the physics in place, the <strong>price lists can be read as statements </strong>about cost. Start with the question a developer can answer from them directly: </p><blockquote><p><em>how many times must a cached prefix be read before caching beats resending it? </em></p></blockquote><p>With a write multiplier w and a read multiplier r, writing once and reading n times costs w + n&#183;r input-equivalents, and not caching costs 1 + n.</p><pre><code>break-even reads:  n &gt; (w - 1) / (1 - r)
5-minute write  (w = 1.25, r = 0.1):   n &gt; 0.28    one read wins:  1.35 against 2.00
1-hour write    (w = 2.00, r = 0.1):   n &gt; 1.11    one read loses: 2.10 against 2.00</code></pre><p>A five-minute write pays for itself with one read and an hour-long write needs two. The <strong>second read is the price of the longer clock</strong>, and the hourly cost of holding can be backed out of each price list. Anthropic&#8217;s five-minute premium is a quarter of the input price for five minutes, three times the input price per hour; its one-hour premium is one times the input price per hour. </p><p>OpenAI&#8217;s is a quarter for at least thirty minutes, at most half the input price per hour. <strong>Google&#8217;s rent is 2.25 times the input price per hour</strong> on Gemini 3.1 Pro and two thirds of it on Gemini 3.8 Flash. DeepSeek&#8217;s is zero.</p><p>Those hourly figures describe an entry that is written and never read again. <em>An entry that keeps being read costs nothing more to keep,</em> because at both Anthropic and OpenAI every read refreshes the lifetime without a new write charge. </p><p>For an active agent the write premium is a deposit, not rent. <strong>The provider charges for the chance that you walk away</strong> and gives the holding away to anyone who comes back, which is the pricing you would design if idle bytes in fast memory were the cost you worried about.</p><p>Converting those charges to dollars per gigabyte-hour requires the <strong>cache size per token of the model behind each price</strong>, and none of these providers publishes it for its closed models. So we plotted every charge against every plausible size.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!agdj!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F62ff4f79-8f25-486d-8087-110029963a6e_1634x988.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!agdj!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F62ff4f79-8f25-486d-8087-110029963a6e_1634x988.png 424w, https://substackcdn.com/image/fetch/$s_!agdj!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F62ff4f79-8f25-486d-8087-110029963a6e_1634x988.png 848w, https://substackcdn.com/image/fetch/$s_!agdj!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F62ff4f79-8f25-486d-8087-110029963a6e_1634x988.png 1272w, https://substackcdn.com/image/fetch/$s_!agdj!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F62ff4f79-8f25-486d-8087-110029963a6e_1634x988.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!agdj!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F62ff4f79-8f25-486d-8087-110029963a6e_1634x988.png" width="1456" height="880" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/62ff4f79-8f25-486d-8087-110029963a6e_1634x988.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:880,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Log-log chart of implied dollars per gigabyte-hour against cache bytes per token for four storage charges, with bands for HBM, DRAM and SSD costs.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Log-log chart of implied dollars per gigabyte-hour against cache bytes per token for four storage charges, with bands for HBM, DRAM and SSD costs." title="Log-log chart of implied dollars per gigabyte-hour against cache bytes per token for four storage charges, with bands for HBM, DRAM and SSD costs." srcset="https://substackcdn.com/image/fetch/$s_!agdj!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F62ff4f79-8f25-486d-8087-110029963a6e_1634x988.png 424w, https://substackcdn.com/image/fetch/$s_!agdj!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F62ff4f79-8f25-486d-8087-110029963a6e_1634x988.png 848w, https://substackcdn.com/image/fetch/$s_!agdj!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F62ff4f79-8f25-486d-8087-110029963a6e_1634x988.png 1272w, https://substackcdn.com/image/fetch/$s_!agdj!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F62ff4f79-8f25-486d-8087-110029963a6e_1634x988.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Figure 11.</strong> Hourly storage charge per gigabyte implied by each price, as a function of the undisclosed cache size per token of the model behind it. Bands show our holding costs for HBM, DRAM and SSD; the shaded range spans the modern open models in this piece.</figcaption></figure></div><p>Across the cache sizes of the modern open models in this piece, from V4.1-Flash&#8217;s 890 bytes to GLM-4.6&#8217;s 368 KiB, every one of these charges is <strong>at least two and a half times our DRAM cost</strong>, and most are more than ten times it. For caches under 64 KB per token, all of them except Gemini Flash&#8217;s are at or above what HBM costs to rent. </p><p><strong>Sonnet 5&#8217;s hourly rate, $2 per million tokens,</strong> equals H100 HBM rent at 64 KB per token and exceeds DRAM cost for any cache under 3.9 MB per token, a size no production model approaches. None of the providers prices its clock at DRAM cost, whichever tier the bytes actually sit in.</p><p>DeepSeek is the one provider for which both sides of the calculation are public. At 890 bytes per token a million cached tokens occupy 0.89 GB, and <strong>holding them on SSD for a full day costs about $0.0002 at our rates</strong>. Recomputing them costs at least 16 PFLOP, 27 seconds of an H800 at our utilization, about $0.015 at $2 an hour. </p><p>The off-peak read price, $0.003, is a fifth of that recompute floor and about twelve times a day of SSD rent; the miss price, $0.15, is ten times the recompute floor. <strong>Nineteen months earlier DeepSeek charged $0.14 for an R1 hit against $0.55 for a miss</strong>: 25 percent, on a model that kept 79 times more cache per token.</p><p>We cannot see inside Anthropic&#8217;s models, so we cannot say why a read on Fable 5.1 costs a quarter of what it costs on Sonnet 5 as a share of the input price. <strong>The arithmetic says a read gets cheaper to serve as the cache per token shrinks</strong> relative to compute, and a factor of four is the size of step DeepSeek&#8217;s generations have taken. </p><p>It is also what a company would do to win agent workloads, where reads are most of the bill. The two explanations are not exclusive.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>An agent session, priced</h2><p>Here is what the clock does to a real bill. Take a coding agent that starts with a <strong>12,000-token system prompt</strong> and tool catalog and takes fifty steps. Each step appends 1,100 tokens of tool output and 400 tokens of model output. </p><p>The context ends at 85,500 tokens; the session sends 2.44 million input tokens and receives 20,000 output tokens, a ratio of 122 to one. Manus, which runs this kind of loop at scale, reports a ratio around 100 to one and names cache hit rate as the metric it watches above all others.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!qRbi!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F85106de1-2f46-4ad3-a868-9aa53a8c58e3_1634x950.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!qRbi!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F85106de1-2f46-4ad3-a868-9aa53a8c58e3_1634x950.png 424w, https://substackcdn.com/image/fetch/$s_!qRbi!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F85106de1-2f46-4ad3-a868-9aa53a8c58e3_1634x950.png 848w, https://substackcdn.com/image/fetch/$s_!qRbi!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F85106de1-2f46-4ad3-a868-9aa53a8c58e3_1634x950.png 1272w, https://substackcdn.com/image/fetch/$s_!qRbi!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F85106de1-2f46-4ad3-a868-9aa53a8c58e3_1634x950.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!qRbi!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F85106de1-2f46-4ad3-a868-9aa53a8c58e3_1634x950.png" width="1456" height="847" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/85106de1-2f46-4ad3-a868-9aa53a8c58e3_1634x950.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:847,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Stacked bar chart of session cost: $5.08 without cache, $0.88 with a five-minute cache, $1.88 with nine expiries, $0.97 with keepalive pings, $1.01 with a one-hour cache.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Stacked bar chart of session cost: $5.08 without cache, $0.88 with a five-minute cache, $1.88 with nine expiries, $0.97 with keepalive pings, $1.01 with a one-hour cache." title="Stacked bar chart of session cost: $5.08 without cache, $0.88 with a five-minute cache, $1.88 with nine expiries, $0.97 with keepalive pings, $1.01 with a one-hour cache." srcset="https://substackcdn.com/image/fetch/$s_!qRbi!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F85106de1-2f46-4ad3-a868-9aa53a8c58e3_1634x950.png 424w, https://substackcdn.com/image/fetch/$s_!qRbi!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F85106de1-2f46-4ad3-a868-9aa53a8c58e3_1634x950.png 848w, https://substackcdn.com/image/fetch/$s_!qRbi!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F85106de1-2f46-4ad3-a868-9aa53a8c58e3_1634x950.png 1272w, https://substackcdn.com/image/fetch/$s_!qRbi!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F85106de1-2f46-4ad3-a868-9aa53a8c58e3_1634x950.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Figure 12.</strong> Cost of one fifty-step agent session at Claude Sonnet 5 list prices. In the last three scenarios, nine of the forty-nine gaps between steps last eight minutes.</figcaption></figure></div><p><strong>Without caching the session costs $5.08, almost all of it input</strong>. With a five-minute cache and every tool call finishing inside the window it costs $0.88: the session reads 96.5 percent of its input tokens from cache, its effective input price falls to 14 percent of list, and 69 percent of what remains is reads. </p><p>Now let nine of the forty-nine gaps last eight minutes: the test suite, the build, the person who went for coffee. Each time the cache is gone, and <strong>the entire context is written again at 1.25 times the input price. </strong>The session costs $1.88, more than double. Writing everything with the one-hour lifetime brings it to $1.01; keeping the five-minute entries alive with one zero-output ping per gap brings it to $0.97.</p><p>The <strong>same session on Opus 5.5, with reads at 5 percent</strong>, costs $1.30 against $10.15 uncached. On DeepSeek&#8217;s V4.1-Flash off-peak it costs 3.2 cents against 38 cents. Because reads dominate an agent&#8217;s bill, <em>the read multiplier matters more than any other number on the price list</em>: halving it from 0.1 to 0.05 at Sonnet 5&#8217;s base price would cut the input side of this session by 34 percent.</p><p>Price lists also differ in structure, so the same session lands at very different fractions of each provider&#8217;s own input price. <strong>With a warm cache it runs at 0.14 of list on Sonnet 5 </strong>and GPT-5.6, 0.13 on Gemini&#8217;s implicit cache if every repeat hits, 0.09 on Opus 5.5, 0.07 on Fable 5.1 and 0.05 on DeepSeek&#8217;s V4.1-Flash. </p><p><strong>The eight-minute gaps separate them further. </strong>GPT-5.6&#8217;s thirty-minute minimum absorbs them, and DeepSeek publishes no lifetime but serves hits from disk, so we actually assume its cache survives them too. Anthropic&#8217;s five-minute models need pings or the one-hour write to stay near their warm figures; left to expire, Sonnet 5 climbs to 0.34, Opus 5.5 to 0.30 and Fable 5.1 to 0.29. A lower read price does little for a session whose cache keeps expiring.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Blia!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F180c311a-fa46-4cf0-8b79-8bc01bb815a0_1634x950.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Blia!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F180c311a-fa46-4cf0-8b79-8bc01bb815a0_1634x950.png 424w, https://substackcdn.com/image/fetch/$s_!Blia!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F180c311a-fa46-4cf0-8b79-8bc01bb815a0_1634x950.png 848w, https://substackcdn.com/image/fetch/$s_!Blia!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F180c311a-fa46-4cf0-8b79-8bc01bb815a0_1634x950.png 1272w, https://substackcdn.com/image/fetch/$s_!Blia!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F180c311a-fa46-4cf0-8b79-8bc01bb815a0_1634x950.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Blia!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F180c311a-fa46-4cf0-8b79-8bc01bb815a0_1634x950.png" width="1456" height="847" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/180c311a-fa46-4cf0-8b79-8bc01bb815a0_1634x950.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:847,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Horizontal bar chart of effective input multipliers: DeepSeek 0.05, Fable 5.1 0.07, Opus 5.5 0.09, Gemini implicit 0.13, Sonnet 5 and GPT-5.6 0.14, Sonnet 5 one-hour 0.17; expiries raise Anthropic five-minute models to about 0.3.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Horizontal bar chart of effective input multipliers: DeepSeek 0.05, Fable 5.1 0.07, Opus 5.5 0.09, Gemini implicit 0.13, Sonnet 5 and GPT-5.6 0.14, Sonnet 5 one-hour 0.17; expiries raise Anthropic five-minute models to about 0.3." title="Horizontal bar chart of effective input multipliers: DeepSeek 0.05, Fable 5.1 0.07, Opus 5.5 0.09, Gemini implicit 0.13, Sonnet 5 and GPT-5.6 0.14, Sonnet 5 one-hour 0.17; expiries raise Anthropic five-minute models to about 0.3." srcset="https://substackcdn.com/image/fetch/$s_!Blia!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F180c311a-fa46-4cf0-8b79-8bc01bb815a0_1634x950.png 424w, https://substackcdn.com/image/fetch/$s_!Blia!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F180c311a-fa46-4cf0-8b79-8bc01bb815a0_1634x950.png 848w, https://substackcdn.com/image/fetch/$s_!Blia!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F180c311a-fa46-4cf0-8b79-8bc01bb815a0_1634x950.png 1272w, https://substackcdn.com/image/fetch/$s_!Blia!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F180c311a-fa46-4cf0-8b79-8bc01bb815a0_1634x950.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Figure 13.</strong> The same fifty-step session as a fraction of each provider&#8217;s own list input price. Bars assume a warm cache; crosses let the nine eight-minute gaps expire; circles bridge them with pings.</figcaption></figure></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>A field guide</h2><p>The arithmetic above turns into rules that can be applied without redoing it. Each comes with the number behind it, so it can be rechecked when a price list changes.</p><h3>When caching pays at all</h3><p>If every input token either writes to the cache or reads from it, the input bill as a multiple of the uncached bill is (1 - h)&#183;w + h&#183;r, where h is the share of those tokens that are hits. Caching pays when that is below one:</p><pre><code>break-even hit rate:  h* = (w - 1) / (w - r)
5-minute write  (w = 1.25, r = 0.10):   h* = 21.7%
1-hour write    (w = 2.00, r = 0.10):   h* = 52.6%
Opus 5.5, 5-minute (r = 0.05):          h* = 20.8%
DeepSeek (no write premium):            any hit rate pays</code></pre><p>Kimi&#8217;s chat trace at a five-minute lifetime has a hit rate of 29 percent, just above the line, which is why caching everything there saves only 9 percent. <strong>The agent trace sits at 53 percent and our modeled session at 96.5 percent,</strong> where the input bill falls to 14 percent of list. For chat-like traffic, mark what is shared across requests, such as the system prompt and the tool definitions, and let the rest go uncached; for agents, cache everything.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!DpUk!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F41aabbab-0688-44c8-9e91-9a051ac58e82_1634x931.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!DpUk!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F41aabbab-0688-44c8-9e91-9a051ac58e82_1634x931.png 424w, https://substackcdn.com/image/fetch/$s_!DpUk!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F41aabbab-0688-44c8-9e91-9a051ac58e82_1634x931.png 848w, https://substackcdn.com/image/fetch/$s_!DpUk!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F41aabbab-0688-44c8-9e91-9a051ac58e82_1634x931.png 1272w, https://substackcdn.com/image/fetch/$s_!DpUk!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F41aabbab-0688-44c8-9e91-9a051ac58e82_1634x931.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!DpUk!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F41aabbab-0688-44c8-9e91-9a051ac58e82_1634x931.png" width="1456" height="830" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/41aabbab-0688-44c8-9e91-9a051ac58e82_1634x931.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:830,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Line chart of cost multiplier against hit rate for four pricing schemes; break-even at 21.7 percent for 5-minute writes and 52.6 percent for 1-hour writes.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Line chart of cost multiplier against hit rate for four pricing schemes; break-even at 21.7 percent for 5-minute writes and 52.6 percent for 1-hour writes." title="Line chart of cost multiplier against hit rate for four pricing schemes; break-even at 21.7 percent for 5-minute writes and 52.6 percent for 1-hour writes." srcset="https://substackcdn.com/image/fetch/$s_!DpUk!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F41aabbab-0688-44c8-9e91-9a051ac58e82_1634x931.png 424w, https://substackcdn.com/image/fetch/$s_!DpUk!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F41aabbab-0688-44c8-9e91-9a051ac58e82_1634x931.png 848w, https://substackcdn.com/image/fetch/$s_!DpUk!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F41aabbab-0688-44c8-9e91-9a051ac58e82_1634x931.png 1272w, https://substackcdn.com/image/fetch/$s_!DpUk!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F41aabbab-0688-44c8-9e91-9a051ac58e82_1634x931.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Figure 14.</strong> Input bill relative to no caching when every input token is either written to or read from the cache. Caching pays below the dashed line; points show Kimi&#8217;s traces at a five-minute lifetime and our modeled agent session.</figcaption></figure></div><h3>Crossing a long tool call</h3><p>The expensive event in an agent session is an expiry in the middle of a long context, because it turns the cheapest tokens in the session into the most expensive ones. <strong>There are three ways across a gap longer than five minutes. </strong>Let the entry expire and rewrite the context at 1.25 times the input price. </p><p>Buy the one-hour lifetime and pay 2 times instead of 1.25 on the write. Or keep the five-minute entry alive with a request that repeats the prefix every four and a half minutes. <strong>Anthropic documents this as pre-warming</strong>: with max_tokens set to 0 the request reads the cache, refreshes its lifetime and generates nothing. Each ping costs one read, a tenth of the context&#8217;s input price.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!sUBh!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d1c20b0-cc25-4963-9a94-6dfbcafe0c13_1634x931.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!sUBh!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d1c20b0-cc25-4963-9a94-6dfbcafe0c13_1634x931.png 424w, https://substackcdn.com/image/fetch/$s_!sUBh!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d1c20b0-cc25-4963-9a94-6dfbcafe0c13_1634x931.png 848w, https://substackcdn.com/image/fetch/$s_!sUBh!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d1c20b0-cc25-4963-9a94-6dfbcafe0c13_1634x931.png 1272w, https://substackcdn.com/image/fetch/$s_!sUBh!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d1c20b0-cc25-4963-9a94-6dfbcafe0c13_1634x931.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!sUBh!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d1c20b0-cc25-4963-9a94-6dfbcafe0c13_1634x931.png" width="1456" height="830" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0d1c20b0-cc25-4963-9a94-6dfbcafe0c13_1634x931.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:830,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Line chart: letting the cache expire costs 1.25x after five minutes, pings cost 0.1x per 4.5 minutes, the one-hour cache costs 0.85x including its premium; pings win up to 36 minutes.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Line chart: letting the cache expire costs 1.25x after five minutes, pings cost 0.1x per 4.5 minutes, the one-hour cache costs 0.85x including its premium; pings win up to 36 minutes." title="Line chart: letting the cache expire costs 1.25x after five minutes, pings cost 0.1x per 4.5 minutes, the one-hour cache costs 0.85x including its premium; pings win up to 36 minutes." srcset="https://substackcdn.com/image/fetch/$s_!sUBh!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d1c20b0-cc25-4963-9a94-6dfbcafe0c13_1634x931.png 424w, https://substackcdn.com/image/fetch/$s_!sUBh!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d1c20b0-cc25-4963-9a94-6dfbcafe0c13_1634x931.png 848w, https://substackcdn.com/image/fetch/$s_!sUBh!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d1c20b0-cc25-4963-9a94-6dfbcafe0c13_1634x931.png 1272w, https://substackcdn.com/image/fetch/$s_!sUBh!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d1c20b0-cc25-4963-9a94-6dfbcafe0c13_1634x931.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Figure 15.</strong> Cost to resume a context after a gap, in multiples of its uncached input price, for three ways of crossing the gap at Anthropic&#8217;s multipliers and for OpenAI&#8217;s thirty-minute minimum.</figcaption></figure></div><p>Charging the whole one-hour premium to a single gap, pings are cheaper than the one-hour write for <strong>gaps up to about 36 minutes</strong> and cheaper than a rewrite up to about 54 minutes. When a session has several long gaps the one-hour premium is spread across them and wins sooner. </p><p>Pings have to reproduce the prefix and the settings exactly, thinking and effort configuration included, and <strong>Anthropic rejects the zero-output form </strong>when it is combined with streaming, extended thinking, structured outputs or a forced tool choice. </p><p><strong>OpenAI&#8217;s GPT-5.6 needs none of this</strong> for gaps under half an hour, because its minimum lifetime is already thirty minutes. In our session, one ping per eight-minute gap brings the bill from $1.88 to $0.97.</p><h3>Keeping the prefix identical</h3><p>One changed token invalidates every block after it, so the rules are mechanical. <strong>Put what never changes first</strong>: tool definitions, then instructions, then reference material, then history. Append tool results and messages instead of editing or reordering earlier turns. </p><p><strong>Serialize JSON with a fixed key order</strong>. Keep timestamps and per-request data at the end or in later messages. Do not toggle settings that are rendered into the prompt: Anthropic&#8217;s list includes tool definitions, web search and citation toggles, speed mode, and thinking or effort settings;</p><p> OpenAI&#8217;s includes tools, parallel tool calls, output format, reasoning effort and verbosity. <strong>Manus describes arriving at the first three rules after several rewrites</strong> of its agent framework.</p><p>Compaction is the deliberate exception. Replacing a context of C tokens with a summary of c tokens forfeits the cache and costs a summary generated at output prices. At Sonnet 5 prices, where output costs five times input, it pays back after about ((w - r + 5)&#183;c + r&#183;C) / (r&#183;(C - c)) further steps. <strong>Compacting 80,000 tokens into a 5,000-token summary</strong> pays back in about five steps; into a 20,000-token summary, in about 22.</p><h3>Measuring what you pay</h3><p>Both APIs report the split. Anthropic returns cache_creation_input_tokens, cache_read_input_tokens and input_tokens, where input_tokens counts only tokens after the last breakpoint; OpenAI returns cached_tokens and cache_write_tokens inside input_tokens_details, and its input_tokens includes both.</p><p>The number to watch is the effective input multiplier:</p><pre><code>effective input multiplier = f_uncached + w &#215; f_written + r &#215; f_read</code></pre><p>where the f are shares of all input tokens. Our session runs at 0.14 with a warm cache and 0.34 with nine expiries. A rising multiplier with a steady hit rate usually means more writes: a <strong>breakpoint on content that changes</strong>, or a conversation that grows more than 20 blocks past its last write, beyond the reach of Anthropic&#8217;s lookback.</p><h3>Sizing a cache for your own fleet</h3><p>For a self-hosted deployment the same formulas become a sizing recipe. Per prefill GPU, a cache that holds &#964; seconds of fresh output needs G &#215; &#964; bytes, and in Kimi&#8217;s traces ten minutes of unique traffic captured 89 to 98 percent of the achievable reuse. </p><p>On an H100 that is 827 GB for Llama-3.1-70B, 338 GB for DeepSeek-V3 and 20 GB for V4.1-Flash, which picks the tier for you: GPU or host memory for compressed caches, SSD or a shared pool for the rest. <strong>A flash tier has to be sized for writes as well as capacity, G &#215; 86,400 bytes per GPU per day</strong> if everything is persisted, which is why admission control belongs in the design from the start. </p><p>And if reloads must hide behind prefill, the link needs (cached tokens / new tokens) &#215; G: 77 GB/s for the last step of our session on Llama-3.1-70B, which a PCIe Gen5 x16 link cannot deliver and a coherent CPU link can.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!JScv!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7acaa57b-5f5c-4924-a432-e78a4607e95a_1634x1204.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!JScv!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7acaa57b-5f5c-4924-a432-e78a4607e95a_1634x1204.png 424w, https://substackcdn.com/image/fetch/$s_!JScv!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7acaa57b-5f5c-4924-a432-e78a4607e95a_1634x1204.png 848w, https://substackcdn.com/image/fetch/$s_!JScv!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7acaa57b-5f5c-4924-a432-e78a4607e95a_1634x1204.png 1272w, https://substackcdn.com/image/fetch/$s_!JScv!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7acaa57b-5f5c-4924-a432-e78a4607e95a_1634x1204.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!JScv!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7acaa57b-5f5c-4924-a432-e78a4607e95a_1634x1204.png" width="1456" height="1073" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/7acaa57b-5f5c-4924-a432-e78a4607e95a_1634x1204.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1073,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Table of eight caching rules with the number behind each, from the break-even hit rate to cache sizing for self-hosted fleets.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Table of eight caching rules with the number behind each, from the break-even hit rate to cache sizing for self-hosted fleets." title="Table of eight caching rules with the number behind each, from the break-even hit rate to cache sizing for self-hosted fleets." srcset="https://substackcdn.com/image/fetch/$s_!JScv!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7acaa57b-5f5c-4924-a432-e78a4607e95a_1634x1204.png 424w, https://substackcdn.com/image/fetch/$s_!JScv!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7acaa57b-5f5c-4924-a432-e78a4607e95a_1634x1204.png 848w, https://substackcdn.com/image/fetch/$s_!JScv!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7acaa57b-5f5c-4924-a432-e78a4607e95a_1634x1204.png 1272w, https://substackcdn.com/image/fetch/$s_!JScv!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7acaa57b-5f5c-4924-a432-e78a4607e95a_1634x1204.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Table 3.</strong> The field guide in one table, at September 2026 list prices.</figcaption></figure></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Where this could be wrong</h2><p><strong>Utilization and prices move G and the clock linearly.</strong> We used 30 percent of dense FP8 peak and an H100 at $2.50 an hour. Double the utilization and every break-even time halves; halve the GPU price and the DRAM and SSD break-evens halve. The ordering of tiers and models does not change, and neither do the log-scale charts, but individual minutes and hours can be off by a factor of two. Published measurements put real prefill between 0.16 of peak in DeepSeek&#8217;s production fleet and 0.29 in SGLang&#8217;s GB200 benchmark, so our central value is on the optimistic side: at production utilization every HBM break-even roughly doubles and every capacity figure halves.</p><p><strong>The HBM price is an opportunity cost.</strong> Pricing idle HBM at its full share of GPU rent is right when serving is memory-bound. A compute-bound prefill fleet values HBM less at the margin, and its HBM break-even is longer than H/G.</p><p><strong>Closed models are closed.</strong> Everything we say about Anthropic, OpenAI and Google infrastructure is inference from prices and documentation. If their caches per token are far smaller than the open models&#8217; caches, which DeepSeek has shown is possible, the implied storage charges sit even further above HBM rent; if far larger, closer to it.</p><p><strong>The traces are one sampled hour of one service.</strong> Kimi&#8217;s traces come from chat and tool traffic, not from long-running coding agents with multi-minute tool calls. Agent sessions in 2026 probably have longer gaps and higher reuse, which would favor longer lifetimes than our replay suggests. Our correction for the one-hour window removes the truncation bias only for gaps up to thirty minutes. Alibaba has published two-hour traces from its Bailian service in a compatible block-hash format, including a coding trace and a reasoning trace; we have not replayed them here, and they are the obvious next check.</p><p><strong>Prices are strategy as much as cost.</strong> A provider can price reads below cost to win agent workloads, and rent above cost because explicit caches are bought by customers with few alternatives. Reading prices as cost signals assumes competition does some of the work.</p><p><strong>F is a floor, and we left latency out.</strong> Attention makes recomputation more expensive at long context, and a hit also buys time to first token, which has a value we did not price. Both make caching more attractive than our numbers show.</p><p><strong>The rules are only as durable as the price lists.</strong> Pings work because reads refresh lifetimes for free and zero-output requests are documented, and the crossovers in the field guide move with every multiplier. Either could change with the next pricing update, which is why every rule above is written as a formula that can be rerun.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Predictions</h2><p>Each of these can be checked against a public price list, model card or announcement.</p><ol><li><p><em>By December 31, 2027, at least two of Anthropic, OpenAI and Google will list a cache-read price at or below 5 percent of the fresh input price for their flagship model. Today only Anthropic does.</em></p></li><li><p><em>By December 31, 2027, an open-weight model from a lab other than DeepSeek, competitive with frontier models on agentic coding benchmarks, will ship with a cache under 4 KB per token at its reference precision.</em></p></li><li><p><em>By December 31, 2027, DeepSeek will list a cache-hit price at or below 1 percent of its cache-miss price.</em></p></li><li><p><em>By December 31, 2027, at least two of the five largest GPU clouds will have announced a production flash tier dedicated to KV cache, built on CMX or an equivalent design.</em></p></li></ol><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Confidence dossier</h2><p>Every number in this piece is listed below with a tier. <strong>A means a primary source </strong>we read directly: documentation, a model card, a paper or a company post. <strong>B means secondary reporting</strong>, a preliminary specification or market data. <strong>C means our own derivation</strong> or model from stated assumptions. <strong>D means our interpretation</strong> of closed systems, strategy or vendor claims. <strong>M means our own measurement </strong>on public data, either the Mooncake traces or published model configurations.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!BK1G!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1011db06-31d0-4270-bdc9-d7d07a36a5cd_1643x1776.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!BK1G!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1011db06-31d0-4270-bdc9-d7d07a36a5cd_1643x1776.png 424w, https://substackcdn.com/image/fetch/$s_!BK1G!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1011db06-31d0-4270-bdc9-d7d07a36a5cd_1643x1776.png 848w, https://substackcdn.com/image/fetch/$s_!BK1G!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1011db06-31d0-4270-bdc9-d7d07a36a5cd_1643x1776.png 1272w, https://substackcdn.com/image/fetch/$s_!BK1G!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1011db06-31d0-4270-bdc9-d7d07a36a5cd_1643x1776.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!BK1G!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1011db06-31d0-4270-bdc9-d7d07a36a5cd_1643x1776.png" width="1456" height="1574" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/1011db06-31d0-4270-bdc9-d7d07a36a5cd_1643x1776.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1574,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Table of claims with confidence tiers A, B, C, D and M, part 1.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Table of claims with confidence tiers A, B, C, D and M, part 1." title="Table of claims with confidence tiers A, B, C, D and M, part 1." srcset="https://substackcdn.com/image/fetch/$s_!BK1G!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1011db06-31d0-4270-bdc9-d7d07a36a5cd_1643x1776.png 424w, https://substackcdn.com/image/fetch/$s_!BK1G!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1011db06-31d0-4270-bdc9-d7d07a36a5cd_1643x1776.png 848w, https://substackcdn.com/image/fetch/$s_!BK1G!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1011db06-31d0-4270-bdc9-d7d07a36a5cd_1643x1776.png 1272w, https://substackcdn.com/image/fetch/$s_!BK1G!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1011db06-31d0-4270-bdc9-d7d07a36a5cd_1643x1776.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Table 4.</strong> Confidence dossier.</figcaption></figure></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!ouDN!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F45972c36-537b-4267-a1ae-ab26fe50f60b_1643x1776.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!ouDN!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F45972c36-537b-4267-a1ae-ab26fe50f60b_1643x1776.png 424w, https://substackcdn.com/image/fetch/$s_!ouDN!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F45972c36-537b-4267-a1ae-ab26fe50f60b_1643x1776.png 848w, https://substackcdn.com/image/fetch/$s_!ouDN!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F45972c36-537b-4267-a1ae-ab26fe50f60b_1643x1776.png 1272w, https://substackcdn.com/image/fetch/$s_!ouDN!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F45972c36-537b-4267-a1ae-ab26fe50f60b_1643x1776.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!ouDN!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F45972c36-537b-4267-a1ae-ab26fe50f60b_1643x1776.png" width="1456" height="1574" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/45972c36-537b-4267-a1ae-ab26fe50f60b_1643x1776.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1574,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Table of claims with confidence tiers, part 2.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Table of claims with confidence tiers, part 2." title="Table of claims with confidence tiers, part 2." srcset="https://substackcdn.com/image/fetch/$s_!ouDN!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F45972c36-537b-4267-a1ae-ab26fe50f60b_1643x1776.png 424w, https://substackcdn.com/image/fetch/$s_!ouDN!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F45972c36-537b-4267-a1ae-ab26fe50f60b_1643x1776.png 848w, https://substackcdn.com/image/fetch/$s_!ouDN!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F45972c36-537b-4267-a1ae-ab26fe50f60b_1643x1776.png 1272w, https://substackcdn.com/image/fetch/$s_!ouDN!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F45972c36-537b-4267-a1ae-ab26fe50f60b_1643x1776.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Table 5.</strong> Confidence dossier.</figcaption></figure></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!T3k7!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbce5f3a5-ad13-4505-8006-87d61dc9a06d_1643x1776.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!T3k7!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbce5f3a5-ad13-4505-8006-87d61dc9a06d_1643x1776.png 424w, https://substackcdn.com/image/fetch/$s_!T3k7!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbce5f3a5-ad13-4505-8006-87d61dc9a06d_1643x1776.png 848w, https://substackcdn.com/image/fetch/$s_!T3k7!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbce5f3a5-ad13-4505-8006-87d61dc9a06d_1643x1776.png 1272w, https://substackcdn.com/image/fetch/$s_!T3k7!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbce5f3a5-ad13-4505-8006-87d61dc9a06d_1643x1776.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!T3k7!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbce5f3a5-ad13-4505-8006-87d61dc9a06d_1643x1776.png" width="1456" height="1574" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/bce5f3a5-ad13-4505-8006-87d61dc9a06d_1643x1776.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1574,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Table of claims with confidence tiers, part 3.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Table of claims with confidence tiers, part 3." title="Table of claims with confidence tiers, part 3." srcset="https://substackcdn.com/image/fetch/$s_!T3k7!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbce5f3a5-ad13-4505-8006-87d61dc9a06d_1643x1776.png 424w, https://substackcdn.com/image/fetch/$s_!T3k7!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbce5f3a5-ad13-4505-8006-87d61dc9a06d_1643x1776.png 848w, https://substackcdn.com/image/fetch/$s_!T3k7!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbce5f3a5-ad13-4505-8006-87d61dc9a06d_1643x1776.png 1272w, https://substackcdn.com/image/fetch/$s_!T3k7!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbce5f3a5-ad13-4505-8006-87d61dc9a06d_1643x1776.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Table 6.</strong> Confidence dossier.</figcaption></figure></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!JoTq!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcdb152bc-8cf8-4417-a4c6-bb4a46a9261c_1643x1779.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!JoTq!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcdb152bc-8cf8-4417-a4c6-bb4a46a9261c_1643x1779.png 424w, https://substackcdn.com/image/fetch/$s_!JoTq!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcdb152bc-8cf8-4417-a4c6-bb4a46a9261c_1643x1779.png 848w, https://substackcdn.com/image/fetch/$s_!JoTq!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcdb152bc-8cf8-4417-a4c6-bb4a46a9261c_1643x1779.png 1272w, https://substackcdn.com/image/fetch/$s_!JoTq!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcdb152bc-8cf8-4417-a4c6-bb4a46a9261c_1643x1779.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!JoTq!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcdb152bc-8cf8-4417-a4c6-bb4a46a9261c_1643x1779.png" width="1456" height="1577" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/cdb152bc-8cf8-4417-a4c6-bb4a46a9261c_1643x1779.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1577,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Table of claims with confidence tiers, part 4.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Table of claims with confidence tiers, part 4." title="Table of claims with confidence tiers, part 4." srcset="https://substackcdn.com/image/fetch/$s_!JoTq!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcdb152bc-8cf8-4417-a4c6-bb4a46a9261c_1643x1779.png 424w, https://substackcdn.com/image/fetch/$s_!JoTq!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcdb152bc-8cf8-4417-a4c6-bb4a46a9261c_1643x1779.png 848w, https://substackcdn.com/image/fetch/$s_!JoTq!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcdb152bc-8cf8-4417-a4c6-bb4a46a9261c_1643x1779.png 1272w, https://substackcdn.com/image/fetch/$s_!JoTq!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcdb152bc-8cf8-4417-a4c6-bb4a46a9261c_1643x1779.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Table 7.</strong> Confidence dossier.</figcaption></figure></div><h2>Method</h2><p>Cache sizes come from each model&#8217;s published configuration: layer counts, <strong>key-value heads and head dimensions</strong> for grouped-query models, latent and positional dimensions for MLA models, and the model card for V4.1-Flash. </p><p><strong>The V3.2 layout counts 512 FP8 latent values</strong>, 16 bytes of scales and 64 BF16 positional values per layer, plus a 128-value FP8 indexer key with a 4-byte scale. The V4-Flash figure comes from our August reconstruction of its compressed attention. F counts two operations per active parameter and nothing else.</p><p>G, the break-even times and the capacities follow from the formulas in the text, with P at dense FP8 peak, u at 0.30, an H100 at $2.50 an hour, a B200 at $4.50 and the tier costs in the table. Rubin figures use NVIDIA&#8217;s specification page as revised on September 21, 2026.</p><p>The trace replay sorts requests by arrival time and treats each hash as one 512-token block. A request hits its longest prefix of blocks present in the cache and stops at the first miss, and every block it touches has its last-use time set to the request&#8217;s arrival. <strong>Tokens are counted with the final partial block capped at the request&#8217;s input length.</strong> Lifetimes are sliding: an entry expires a fixed time after its last use. The capacity replay evicts the least recently used block once the cache is full. For the gap statistics we counted only reuses whose previous use came at least 1,800 seconds before the end of the trace.</p><p><strong>The session model has a 12,000-token starting prefix</strong>, fifty steps, 1,100 tokens of tool output and 400 of model output per step, and Claude Sonnet 5 list prices. A step after a long gap pays a full write for its whole context. </p><p>The whiskers in the clock chart vary utilization from 0.15 to 0.5, the H100 price from $1.50 to $4.00 an hour, <strong>DRAM from $8 to $16 per GB with 25 to 100 percent overhead,</strong> SSD from $0.15 to $0.35 per GB with 50 to 200 percent overhead, and the share of HBM available to cache from 60 to 100 percent. </p><p>The gap model sends a ping every 4.5 minutes after the first five, charges each ping one read of the context, and charges the whole one-hour premium to the gap it protects; in the session with pings, each of the nine long gaps lasts eight minutes and takes one ping. No number in this piece required a GPU.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Sources</h2><ol><li><p>Anthropic. Prompt caching. Claude Platform Docs, accessed September 2026. <a href="https://platform.claude.com/docs/en/build-with-claude/prompt-caching">https://platform.claude.com/docs/en/build-with-claude/prompt-caching</a></p></li><li><p>OpenAI. Prompt caching. OpenAI API documentation, accessed September 2026. <a href="https://developers.openai.com/api/docs/guides/prompt-caching">https://developers.openai.com/api/docs/guides/prompt-caching</a></p></li><li><p>Google. Gemini Developer API pricing, accessed September 2026. <a href="https://ai.google.dev/gemini-api/docs/pricing">https://ai.google.dev/gemini-api/docs/pricing</a></p></li><li><p>Google. Context caching. Gemini API documentation. <a href="https://ai.google.dev/gemini-api/docs/caching">https://ai.google.dev/gemini-api/docs/caching</a></p></li><li><p>DeepSeek. Models and Pricing. DeepSeek API Docs, accessed September 2026. <a href="https://api-docs.deepseek.com/quick_start/pricing">https://api-docs.deepseek.com/quick_start/pricing</a></p></li><li><p>DeepSeek. DeepSeek-V4.1-Flash: Smarter, Faster, More Efficient. Release note, September 10, 2026; the news index dates Context Caching to August 2, 2024. <a href="https://api-docs.deepseek.com/news/news260910">https://api-docs.deepseek.com/news/news260910</a></p></li><li><p>DeepSeek. Day 6: One More Thing, DeepSeek-V3/R1 Inference System Overview. Open Infra Index, February 2025. <a href="https://github.com/deepseek-ai/open-infra-index/blob/main/202502OpenSourceWeek/day_6_one_more_thing_deepseekV3R1_inference_system_overview.md">https://github.com/deepseek-ai/open-infra-index/blob/main/202502OpenSourceWeek/day_6_one_more_thing_deepseekV3R1_inference_system_overview.md</a></p></li><li><p>Grattafiori et al. The Llama 3 Herd of Models. arXiv:2407.21783, 2024. <a href="https://arxiv.org/abs/2407.21783">https://arxiv.org/abs/2407.21783</a></p></li><li><p>DeepSeek-AI. DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model. arXiv:2405.04434, 2024. <a href="https://arxiv.org/abs/2405.04434">https://arxiv.org/abs/2405.04434</a></p></li><li><p>Moonshot AI. Kimi K2 model summary, README. <a href="https://github.com/MoonshotAI/Kimi-K2">https://github.com/MoonshotAI/Kimi-K2</a></p></li><li><p>DeepSeek-AI. DeepSeek-V3 inference configuration, config_671B.json. <a href="https://github.com/deepseek-ai/DeepSeek-V3/blob/main/inference/configs/config_671B.json">https://github.com/deepseek-ai/DeepSeek-V3/blob/main/inference/configs/config_671B.json</a></p></li><li><p>DeepSeek-AI. DeepSeek-V3.2-Exp inference configuration, config_671B_v3.2.json. <a href="https://github.com/deepseek-ai/DeepSeek-V3.2-Exp/blob/main/inference/config_671B_v3.2.json">https://github.com/deepseek-ai/DeepSeek-V3.2-Exp/blob/main/inference/config_671B_v3.2.json</a></p></li><li><p>Bradanini and Tettamanti. DeepSeek V4-Flash: The Cost of Deciding What to Read. August 2026.</p></li><li><p>Simon Willison. DeepSeek V4, almost on the frontier, a fraction of the price. April 24, 2026 (quotes the V4 report on cache size at 1M tokens). <a href="https://simonwillison.net/2026/apr/24/deepseek-v4">https://simonwillison.net/2026/apr/24/deepseek-v4</a></p></li><li><p>DeepSeek-AI. DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression. Model card, September 2026. <a href="https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash">https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash</a></p></li><li><p>Touvron et al. Llama 2: Open Foundation and Fine-Tuned Chat Models. arXiv:2307.09288, 2023. <a href="https://arxiv.org/abs/2307.09288">https://arxiv.org/abs/2307.09288</a></p></li><li><p>MiniMax. MiniMax-M2 configuration, config.json. <a href="https://huggingface.co/MiniMaxAI/MiniMax-M2/resolve/main/config.json">https://huggingface.co/MiniMaxAI/MiniMax-M2/resolve/main/config.json</a></p></li><li><p>GLM-4.6 configuration, as published in RedHatAI&#8217;s NVFP4 release of zai-org/GLM-4.6, config.json. <a href="https://huggingface.co/RedHatAI/GLM-4.6-NVFP4/blob/main/config.json">https://huggingface.co/RedHatAI/GLM-4.6-NVFP4/blob/main/config.json</a></p></li><li><p>DeepSeek-AI. DeepSeek LLM: Scaling Open-Source Language Models with Longtermism. arXiv:2401.02954, 2024. <a href="https://arxiv.org/abs/2401.02954">https://arxiv.org/abs/2401.02954</a></p></li><li><p>OpenAI. gpt-oss reference implementation, gpt_oss/torch/model.py. <a href="https://github.com/openai/gpt-oss/blob/main/gpt_oss/torch/model.py">https://github.com/openai/gpt-oss/blob/main/gpt_oss/torch/model.py</a></p></li><li><p>Hugging Face Transformers v4.57.1. Qwen3-Next configuration defaults. <a href="https://github.com/huggingface/transformers/blob/v4.57.1/src/transformers/models/qwen3_next/configuration_qwen3_next.py">https://github.com/huggingface/transformers/blob/v4.57.1/src/transformers/models/qwen3_next/configuration_qwen3_next.py</a></p></li><li><p>Qwen Team. Qwen3 Technical Report. arXiv:2505.09388, 2025. <a href="https://arxiv.org/abs/2505.09388">https://arxiv.org/abs/2505.09388</a></p></li><li><p>The SGLang Team. Deploying DeepSeek with PD Disaggregation and Large-Scale Expert Parallelism on 96 H100 GPUs. LMSYS blog, May 5, 2025. <a href="https://lmsys.org/blog/2025-05-05-large-scale-ep/">https://lmsys.org/blog/2025-05-05-large-scale-ep/</a></p></li><li><p>The SGLang Team. Deploying DeepSeek on GB200 NVL72 with PD and Large Scale EP (Part II): 3.8x Prefill, 4.8x Decode Throughput. LMSYS blog, September 25, 2025. <a href="https://www.lmsys.org/blog/2025-09-25-gb200-part-2">https://www.lmsys.org/blog/2025-09-25-gb200-part-2</a></p></li><li><p>Building py-kvcache: A Performance Characterization of External KV Caching for vLLM with NVMe SSDs. arXiv:2609.11744, September 2026. <a href="https://arxiv.org/abs/2609.11744">https://arxiv.org/abs/2609.11744</a></p></li><li><p>Qin et al. Mooncake: Trading More Storage for Less Computation, a KVCache-centric Architecture for Serving LLM Chatbot. USENIX FAST 2025, Best Paper Award. <a href="https://www.usenix.org/conference/fast25/presentation/qin">https://www.usenix.org/conference/fast25/presentation/qin</a></p></li><li><p>TrendForce. Memory Makers Prioritize Server Applications, Driving Across-the-Board Price Increases in 1Q26. January 5, 2026. <a href="https://www.trendforce.com/presscenter/news/20260105-12860.html">https://www.trendforce.com/presscenter/news/20260105-12860.html</a></p></li><li><p>ITLDC. Server RAM and SSD market update, 2026 (supplier prices tracked from January 1 to August 21). <a href="https://itldc.com/es/blog/server-ram-ssd-market-update-2026">https://itldc.com/es/blog/server-ram-ssd-market-update-2026</a></p></li><li><p>igor&#8217;sLAB. DRAM and NAND remain more expensive: TrendForce sees slowing but still rising memory prices in Q3 2026. <a href="https://www.igorslab.de/en/dram-and-nand-remain-more-expensive-trendforce-sees-slowing-but-still-rising-memory-prices-q3-2026/">https://www.igorslab.de/en/dram-and-nand-remain-more-expensive-trendforce-sees-slowing-but-still-rising-memory-prices-q3-2026/</a></p></li><li><p>IntuitionLabs. Data Center GPU Pricing 2026: The Full AI Pricing Index. July 20, 2026. <a href="https://intuitionlabs.ai/articles/data-center-gpu-pricing-2026">https://intuitionlabs.ai/articles/data-center-gpu-pricing-2026</a></p></li><li><p>Moonshot AI and Tsinghua University. Mooncake FAST&#8217;25 trace release: conversation, tool and agent, and synthetic traces. <a href="https://github.com/kvcache-ai/Mooncake/tree/main/FAST25-release">https://github.com/kvcache-ai/Mooncake/tree/main/FAST25-release</a></p></li><li><p>Wang et al. KVCache Cache in the Wild: Characterizing and Optimizing KVCache Cache at a Large Cloud Provider. USENIX ATC 2025. <a href="https://arxiv.org/abs/2506.02634">https://arxiv.org/abs/2506.02634</a></p></li><li><p>NVIDIA Technical Blog. Introducing NVIDIA BlueField-4-Powered CMX Context Memory Storage Platform for the Next Frontier of AI. March 16, 2026. <a href="https://developer.nvidia.com/blog/introducing-nvidia-bluefield-4-powered-inference-context-memory-storage-platform-for-the-next-frontier-of-ai/">https://developer.nvidia.com/blog/introducing-nvidia-bluefield-4-powered-inference-context-memory-storage-platform-for-the-next-frontier-of-ai/</a></p></li><li><p>NVIDIA. Vera Rubin NVL72, product page, preliminary specifications. <a href="https://www.nvidia.com/en-us/data-center/vera-rubin-nvl72">https://www.nvidia.com/en-us/data-center/vera-rubin-nvl72</a></p></li><li><p>Tom&#8217;s Hardware. Nvidia launches Vera Rubin NVL72 AI supercomputer at CES. January 2026. <a href="https://www.tomshardware.com/pc-components/gpus/nvidia-launches-vera-rubin-nvl72-ai-supercomputer-at-ces-promises-up-to-5x-greater-inference-performance-and-10x-lower-cost-per-token-than-blackwell-coming-2h-2026">https://www.tomshardware.com/pc-components/gpus/nvidia-launches-vera-rubin-nvl72-ai-supercomputer-at-ces-promises-up-to-5x-greater-inference-performance-and-10x-lower-cost-per-token-than-blackwell-coming-2h-2026</a></p></li><li><p>DeepSeek-AI. Fire-Flyer File System (3FS), README, performance section. <a href="https://github.com/deepseek-ai/3FS">https://github.com/deepseek-ai/3FS</a></p></li><li><p>NVIDIA. NVIDIA Unveils Rubin CPX: A New Class of GPU Designed for Massive-Context Inference. Press release, September 9, 2025, via AIwire. <a href="https://dev.aiwire.net/2025/09/09/nvidia-unveils-rubin-cpx-a-new-class-of-gpu-designed-for-massive-context-inference/">https://dev.aiwire.net/2025/09/09/nvidia-unveils-rubin-cpx-a-new-class-of-gpu-designed-for-massive-context-inference/</a></p></li><li><p>Tom&#8217;s Hardware. Nvidia removes Rubin CPX accelerators from its roadmap. March 2026. <a href="https://tomshardware.com/pc-components/gpus/nvidia-removes-rubin-cpx-accelerators-from-its-roadmap-groq-3-lpus-take-center-stage-as-cpx-is-removed">https://tomshardware.com/pc-components/gpus/nvidia-removes-rubin-cpx-accelerators-from-its-roadmap-groq-3-lpus-take-center-stage-as-cpx-is-removed</a></p></li><li><p>ASCII weekly. Analysis of the 2026 NVIDIA roadmap and the missing Rubin CPX, in Japanese. <a href="https://weekly.ascii.jp/elem/000/004/387/4387523/3/">https://weekly.ascii.jp/elem/000/004/387/4387523/3/</a></p></li><li><p>vLLM. Automatic Prefix Caching, design document. <a href="https://docs.vllm.ai/en/latest/design/prefix_caching/">https://docs.vllm.ai/en/latest/design/prefix_caching/</a></p></li><li><p>Li et al. Continuum: Efficient and Robust Multi-Turn LLM Agent Scheduling with KV Cache Time-to-Live. arXiv:2511.02230. <a href="https://arxiv.org/abs/2511.02230">https://arxiv.org/abs/2511.02230</a></p></li><li><p>Gu, Li, Kuditipudi, Liang and Hashimoto. Auditing Prompt Caching in Language Model APIs. ICML 2025. <a href="https://proceedings.mlr.press/v267/gu25b.html">https://proceedings.mlr.press/v267/gu25b.html</a></p></li><li><p>Yichao Ji. Context Engineering for AI Agents: Lessons from Building Manus. July 18, 2025. <a href="https://manus.im/blog/Context-Engineering-for-AI-Agents-Lessons-from-Building-Manus">https://manus.im/blog/Context-Engineering-for-AI-Agents-Lessons-from-Building-Manus</a></p></li><li><p>Alibaba. Qwen-Bailian anonymized usage traces: to-C, to-B, thinking and coder workloads, README. <a href="https://github.com/alibaba-edu/qwen-bailian-usagetraces-anon">https://github.com/alibaba-edu/qwen-bailian-usagetraces-anon</a></p></li></ol>]]></content:encoded></item><item><title><![CDATA[How Mojo Actually Compiles]]></title><description><![CDATA[Ask Mojo 1.1 for sm_90 and it quietly returns sm_90a, hinting that portability no longer lives in the binary. An H100 executable can carry kernels NVIDIA's own assembler refuses to build for Blackwell]]></description><link>https://www.thesoftwarefrontier.com/p/how-mojo-actually-compiles</link><guid isPermaLink="false">https://www.thesoftwarefrontier.com/p/how-mojo-actually-compiles</guid><dc:creator><![CDATA[Lorenzo Bradanini]]></dc:creator><pubDate>Thu, 24 Sep 2026 15:55:18 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!eMX6!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff96f0e96-98ee-49ef-b1fe-a957769047a4_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!eMX6!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff96f0e96-98ee-49ef-b1fe-a957769047a4_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!eMX6!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff96f0e96-98ee-49ef-b1fe-a957769047a4_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!eMX6!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff96f0e96-98ee-49ef-b1fe-a957769047a4_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!eMX6!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff96f0e96-98ee-49ef-b1fe-a957769047a4_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!eMX6!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff96f0e96-98ee-49ef-b1fe-a957769047a4_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!eMX6!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff96f0e96-98ee-49ef-b1fe-a957769047a4_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f96f0e96-98ee-49ef-b1fe-a957769047a4_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2151246,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/217227968?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff96f0e96-98ee-49ef-b1fe-a957769047a4_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!eMX6!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff96f0e96-98ee-49ef-b1fe-a957769047a4_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!eMX6!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff96f0e96-98ee-49ef-b1fe-a957769047a4_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!eMX6!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff96f0e96-98ee-49ef-b1fe-a957769047a4_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!eMX6!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff96f0e96-98ee-49ef-b1fe-a957769047a4_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>Introduction</h2><p><em>Four days after Modular shipped Mojo 1.1</em>, I did what I just always do with a brand new compiler. I installed it in a container with one CPU core and no GPU, <strong>wrote a kernel that doubles a vector</strong>, and asked for PTX for an H100. I asked for <code>sm_90</code>. The first two directives of what came back were these:</p><pre><code><code>.version 8.5
.target sm_90a</code></code></pre><p>That one letter is the subject of this piece. <code>sm_90</code> is Hopper. <code>sm_90a</code> is Hopper with its architecture-specific instructions unlocked, among them the warpgroup matrix multiply and the warpgroup register reallocation that the <strong>fastest Hopper kernels </strong>are built on. </p><p>The price of the suffix is written into <strong>NVIDIA&#8217;s own assembler.</strong> PTX that declares <code>.target sm_90a</code> can only be compiled for <code>sm_90a</code>. It cannot be compiled for plain <code>sm_90</code>, and it cannot be carried forward to the next generation. I didn&#8217;t ask for that trade. </p><p>The compiler made it for me, without a warning, and it makes the same trade for every <strong>Hopper and Blackwell</strong> name it accepts but one.</p><p>I spent the following days pulling on that thread, and it runs through the whole stack. I installed Mojo 1.1.0 and MAX&#8217;s <code>max-core</code> 26.6.0 from PyPI, cloned the compiler <strong>Modular open-sourced in August</strong>, took NVIDIA&#8217;s <code>ptxas</code> 13.4 from PyPI to check what Mojo emits, and compiled well over a hundred small probe programs. </p><p>Nothing here ran on a GPU. Everything is about what the compiler decides <strong>before a GPU is involved. </strong>I&#8217;m not a compiler engineer, I&#8217;m someone who likes taking tools apart, and Mojo is now a tool you can take apart all the way down. Where I infer instead of measure, I say so, and the dossier at the end grades every claim.</p><p>The short version is basically this: <strong>Mojo&#8217;s portability is real</strong>, and it lives entirely in the source. The compiler treats a target name as a product rather than an instruction set and compiles for the exact chip it is told about. When the compiler travels with the program, as it does inside MAX&#8217;s runtime library, that costs nothing. </p><p>When it does not, as in an executable made with <code>mojo build</code>, the kernels inside are locked to one chip: an executable built for an H100 carries <strong>PTX that NVIDIA&#8217;s own assembler </strong>refuses to build for any Blackwell part. Either way, every NVIDIA kernel is finished by NVIDIA&#8217;s closed assembler, which MAX ships inside its wheel. </p><p>The design is, at least in my own view, pretty coherent. In several places the implementation is not: a device table that quietly compiles a Jetson Orin as an A100, a Blackwell guard that reads the machine you compile on instead of the machine you compile for, and a <strong>Jetson Thor entry</strong> that cannot reach the tensor memory its hardware has.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>What changed since the Mojo you remember</h2><p><strong>The Mojo of 2023 </strong>was (<em>probably</em>) pitched as a superset of Python with a split personality. <code>def</code> behaved like Python: dynamic, permissive, allowed to raise. <code>fn</code> was the strict twin: typed, non-raising unless declared, arguments immutable by default. </p><p>That split was the idea people remembered. It is gone. Modular said in August 2025 that Mojo may or may not grow into a full Python superset and that this was acceptable, which is how <strong>Simon Willison summarized the change</strong> when the compiler went open. </p><p>The repository&#8217;s own release notes date the rest. <code>fn</code> appears in the November 2022 notes as a new, stricter declaration. It is <strong>deprecated in 0.26.2</strong>, warned on by the compiler from 1.0.0b1 in May 2026, and finally removed in 1.1.0, released on 17 September 2026, together with <code>alias</code>, the <code>__comptime_assert</code> keyword, and the <code>@parameter if</code> and <code>@parameter for</code> forms.</p><p>The 1.0 announcement promised that changes during 1.x would be primarily additive, with breaking changes managed carefully in the way <em>mature languages like C++ </em>manage them. The <strong>first minor release after 1.0 deleted three keywords</strong> and<strong> two syntax forms. </strong></p><p>To be fair, all of them were deprecated during the 1.0 cycle, so code that compiled without deprecation warnings under 1.0 should still compile under 1.1. The more interesting thing is what the removals leave behind. <strong>We wrote a batch of small programs</strong> to test what 1.1 actually enforces, and ran each one through the compiler:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!RzwC!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc13df972-330f-40fe-83a6-ad22443b649a_2400x1821.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!RzwC!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc13df972-330f-40fe-83a6-ad22443b649a_2400x1821.png 424w, https://substackcdn.com/image/fetch/$s_!RzwC!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc13df972-330f-40fe-83a6-ad22443b649a_2400x1821.png 848w, https://substackcdn.com/image/fetch/$s_!RzwC!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc13df972-330f-40fe-83a6-ad22443b649a_2400x1821.png 1272w, https://substackcdn.com/image/fetch/$s_!RzwC!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc13df972-330f-40fe-83a6-ad22443b649a_2400x1821.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!RzwC!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc13df972-330f-40fe-83a6-ad22443b649a_2400x1821.png" width="1456" height="1105" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c13df972-330f-40fe-83a6-ad22443b649a_2400x1821.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1105,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:561270,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/217227968?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc13df972-330f-40fe-83a6-ad22443b649a_2400x1821.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!RzwC!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc13df972-330f-40fe-83a6-ad22443b649a_2400x1821.png 424w, https://substackcdn.com/image/fetch/$s_!RzwC!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc13df972-330f-40fe-83a6-ad22443b649a_2400x1821.png 848w, https://substackcdn.com/image/fetch/$s_!RzwC!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc13df972-330f-40fe-83a6-ad22443b649a_2400x1821.png 1272w, https://substackcdn.com/image/fetch/$s_!RzwC!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc13df972-330f-40fe-83a6-ad22443b649a_2400x1821.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><code>def</code> now means what <code>fn</code> used to mean. It does not raise unless it says <code>raises</code>, every variable is introduced with <code>var</code>, and arguments default to the immutable <code>imm</code> convention. </p><p>The memory model underneath is affine by default: a value moved out with <code>^</code> is uninitialized, and using it is a compile error. It is<strong> linear on request</strong>: a type declared with <code>Deinitable where False</code> never conforms to <code>Deinitable</code>, so it must be explicitly consumed, and the compiler rejects any path that drops it. </p><p>Rust has affine types and<strong> no native linear ones</strong>; Mojo has both in one type system. Exclusivity is enforced at call sites, and since 1.0 the lifetime checker, experimentally, understands the inside of containers, so a reference into a <code>List</code> held across an <code>append</code> is rejected instead of dangling after reallocation. Strings refuse to have a length until you say which length you mean.</p><p>What is still missing is also clear. Pattern matching exists on the main branch as an experimental <code>__match</code> statement, and the nightly changelog notes that exhaustiveness is not yet checked for enums or <code>Bool</code>. An async model and unions are on the roadmap. </p><p><strong>So Mojo 1.1 is not Python with types anymore</strong>. It reads like a Rust whose surface borrowed Python&#8217;s indentation, attached to a compile-time metaprogramming system more capable than either. We have argued before that strict semantics and one way to do things suit a world where models write a growing share of the code. </p><p>Mojo 1.1 reads like a language that took that argument seriously, whether or not anyone at Modular would phrase it that way.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>What pip actually installs</h2><p><code>pip install mojo</code> pulls five packages: <code>mojo</code> itself (<em>a 48.9 MB wheel holding the language server, the debugger and the REPL entry point</em>), <code>mojo-compiler</code> (<em>an 80.7 MB wheel</em>) and its companion <code>mojo-compiler-mojo-libs</code>, <code>mblack</code> 26.6.0 for formatting, and <code>mojo-lldb-libs</code>. </p><p>The directory it installs is 675 MB. The <strong>compiler driver is a single binary of 142,215,032 bytes.</strong> It carries its own linker, an <code>lld</code> of 115,567,904 bytes, which is 81% of the size of the compiler that calls it. The standard library ships as one precompiled container, <code>std.mojoc</code>, 3,226,898 bytes, with the magic bytes <code>MPKG</code> and an entropy of 7.999 bits per byte. </p><p><em>Its entropy says it is compressed,</em> or encrypted, and either way it is opaque. The source of that library is <strong>on</strong> <strong>GitHub under Apache 2.0</strong>; the file the compiler actually loads is not something you can read.</p><p>A &#8220;<em>hello world</em>&#8221; builds in <em>2.33 seconds cold and 0.60 seconds from a warm cache</em> in our first session, and in 2.97 to 3.03 and 0.74 to 0.76 seconds across three later runs on the same, busier shared core. The ratio held: a warm build costs a quarter of a cold one. </p><p>The result is a 17,944-byte executable. The <strong>executable is not standalone.</strong> It links <code>libKGENCompilerRTShared.so</code> through an absolute <code>RUNPATH</code> that points into the Python environment Mojo was installed into, <code>/usr/local/lib/python3.12/dist-packages/modular/lib</code> in our case. Move the binary to a machine without that directory and the loader will not find its runtime unless you point it there. That is a <strong>normal trade for a young toolchain</strong>, and worth knowing before you ship anything built this way.</p><p>The bigger surprise is what the wheel does not contain. Mojo is sold as the language for GPUs, and after <code>pip install mojo</code> there is no GPU package. <code>from gpu.id import block_idx</code> fails with <code>unable to locate module 'gpu'</code>. The 1.0 changelog explains it: the accelerator APIs moved into a new <code>max</code> Mojo package and the <code>layout</code> package was <strong>bundled with MAX instead of Mojo. </strong></p><p>Install <code>max-core</code> 26.6.0, an 85.2 MB wheel, and a <code>max</code> package appears with <code>max.gpu</code> inside it, one of 31 precompiled packages totalling 24.4 MB: <code>nn</code> at 6.8 MB, <code>linalg</code> at 4.8 MB, <code>layout</code>, <code>kv_cache</code>, <code>quantization</code>, <code>comm</code>, <code>shmem</code>, <code>state_space</code>, and five vendor binding packages we come back to later. </p><p>It also brings <code>libmax.so</code> at 162.6 MB, the GPU runtime <code>libMGPRT.so</code> at 44.7 MB, and a 37.7 MB file called <code>libNVPTX.so</code> that turns out to be NVIDIA&#8217;s assembler. With both installed, the tree is 982 MB.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!wqAt!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb8354e25-71dd-48b8-8df4-02111497ac81_1806x987.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!wqAt!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb8354e25-71dd-48b8-8df4-02111497ac81_1806x987.png 424w, https://substackcdn.com/image/fetch/$s_!wqAt!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb8354e25-71dd-48b8-8df4-02111497ac81_1806x987.png 848w, https://substackcdn.com/image/fetch/$s_!wqAt!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb8354e25-71dd-48b8-8df4-02111497ac81_1806x987.png 1272w, https://substackcdn.com/image/fetch/$s_!wqAt!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb8354e25-71dd-48b8-8df4-02111497ac81_1806x987.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!wqAt!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb8354e25-71dd-48b8-8df4-02111497ac81_1806x987.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b8354e25-71dd-48b8-8df4-02111497ac81_1806x987.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Everything Mojo 1.1.0 and max-core 26.6.0 put on disk, by file. The GPU programming model, MAX's kernel packages and NVIDIA's PTX compiler all arrive with max-core, not with the Mojo wheel.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Everything Mojo 1.1.0 and max-core 26.6.0 put on disk, by file. The GPU programming model, MAX's kernel packages and NVIDIA's PTX compiler all arrive with max-core, not with the Mojo wheel." title="Everything Mojo 1.1.0 and max-core 26.6.0 put on disk, by file. The GPU programming model, MAX's kernel packages and NVIDIA's PTX compiler all arrive with max-core, not with the Mojo wheel." srcset="https://substackcdn.com/image/fetch/$s_!wqAt!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb8354e25-71dd-48b8-8df4-02111497ac81_1806x987.png 424w, https://substackcdn.com/image/fetch/$s_!wqAt!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb8354e25-71dd-48b8-8df4-02111497ac81_1806x987.png 848w, https://substackcdn.com/image/fetch/$s_!wqAt!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb8354e25-71dd-48b8-8df4-02111497ac81_1806x987.png 1272w, https://substackcdn.com/image/fetch/$s_!wqAt!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb8354e25-71dd-48b8-8df4-02111497ac81_1806x987.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 1. Everything Mojo 1.1.0 and max-core 26.6.0 put on disk, by file. The GPU programming model, MAX&#8217;s kernel packages and NVIDIA&#8217;s PTX compiler all arrive with max-core, not with the Mojo wheel.</figcaption></figure></div><p>The licenses deserve <strong>one careful paragraph,</strong> because the headline from August was that Mojo is now open source, and that is true of the source. The PyPI metadata for <code>mojo</code> 1.1.0 and for <code>max</code> 26.6.0 both declare <code>LicenseRef-MAX-Platform-Software-License</code>. </p><p>The <strong>Python shim</strong> that provides the wheel&#8217;s entry points opens with a header calling itself Modular proprietary. The repository&#8217;s <code>LICENSE</code> is Apache 2.0 with LLVM exceptions. </p><p>Reporting from <strong>ModCon by NAND Research </strong>and RuntimeWire says Modular also rewrote the MAX license there to remove device restrictions and is moving MAX toward a source-available model. We are not lawyers and read none of this as a problem. </p><p>What is measurable is narrower: the tree you clone and the wheel you install carry different license identifiers, and if you need the<strong> Apache terms end to end</strong>, Modular&#8217;s own instructions build the compiler from source and run a file with <code>./bazelw run --config=build-mojo KGEN:mojo -- run hello.mojo</code>. </p><p>The same post notes that a prebuilt compiler is still required today if you customize MAX kernels or models.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>The compiler, now that it can be read</h2><p>We cloned <code>modular/modular</code> at commit <code>26cfe94</code>, dated 21 September 2026, and fetched the <code>mojo/v1.1.0</code> tag (commit <code>c6fa49f</code>, 17 September) to compare against what actually shipped. The working tree is 11,054 files. </p><p>The compiler proper, the C++ and TableGen under <code>Mojo/lib</code>, <code>Mojo/include</code> and <code>Mojo/tools</code>, is 314,179 lines. The standard library, written in Mojo, is 185,284 lines. The compiler&#8217;s tests are another 149,354. </p><p>Then there is MAX: <strong>1,048,124 lines of Mojo across kernels,</strong> models and tests, and 880,903 lines of Python. The Mojo code inside MAX alone is 3.3 times the size of the compiler. The language is the smaller asset in its own repository.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!7ATP!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc0aa5afc-e01f-4932-a4b7-11f740163d6d_1806x987.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!7ATP!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc0aa5afc-e01f-4932-a4b7-11f740163d6d_1806x987.png 424w, https://substackcdn.com/image/fetch/$s_!7ATP!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc0aa5afc-e01f-4932-a4b7-11f740163d6d_1806x987.png 848w, https://substackcdn.com/image/fetch/$s_!7ATP!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc0aa5afc-e01f-4932-a4b7-11f740163d6d_1806x987.png 1272w, https://substackcdn.com/image/fetch/$s_!7ATP!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc0aa5afc-e01f-4932-a4b7-11f740163d6d_1806x987.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!7ATP!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc0aa5afc-e01f-4932-a4b7-11f740163d6d_1806x987.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c0aa5afc-e01f-4932-a4b7-11f740163d6d_1806x987.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Lines of code in the open repository by component, and the TableGen census of Mojo's five MLIR dialects. kgen declares more types and attributes than operations because Mojo's compile-time values live in MLIR's attribute system.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Lines of code in the open repository by component, and the TableGen census of Mojo's five MLIR dialects. kgen declares more types and attributes than operations because Mojo's compile-time values live in MLIR's attribute system." title="Lines of code in the open repository by component, and the TableGen census of Mojo's five MLIR dialects. kgen declares more types and attributes than operations because Mojo's compile-time values live in MLIR's attribute system." srcset="https://substackcdn.com/image/fetch/$s_!7ATP!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc0aa5afc-e01f-4932-a4b7-11f740163d6d_1806x987.png 424w, https://substackcdn.com/image/fetch/$s_!7ATP!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc0aa5afc-e01f-4932-a4b7-11f740163d6d_1806x987.png 848w, https://substackcdn.com/image/fetch/$s_!7ATP!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc0aa5afc-e01f-4932-a4b7-11f740163d6d_1806x987.png 1272w, https://substackcdn.com/image/fetch/$s_!7ATP!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc0aa5afc-e01f-4932-a4b7-11f740163d6d_1806x987.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 2. Lines of code in the open repository by component, and the TableGen census of Mojo&#8217;s five MLIR dialects. kgen declares more types and attributes than operations because Mojo&#8217;s compile-time values live in MLIR&#8217;s attribute system.</figcaption></figure></div><p>Under <code>Mojo/include/Mojo</code> there are five MLIR dialects, and their own descriptions say what each one is for. <code>lit</code> holds the language-level abstractions and lowers to <code>kgen</code>. <code>pop</code> defines parametric operations on top of<strong> MLIR&#8217;s LLVM dialect</strong>. <code>kgen</code> holds the top-level and support operations of what its description still calls the kernel generation framework. <code>hlcf</code> is a structured, <strong>high-level control-flow dialect</strong>, and <code>co</code> models async functions as coroutines with arbitrary suspension points. </p><p>Counting the TableGen definitions gives 196 operations, 114 types and 174 attributes across the five, and <strong>45 passes declared next to them</strong>, plus 2 more in the shared Support library. The shape of <code>kgen</code> is the telling one: 39 operations, 64 types and 98 attributes, in 9,238 lines of TableGen, more than <code>lit</code> and <code>pop</code> combined. </p><p><strong>Compile-time values in Mojo</strong> are carried as MLIR attributes; the device records we meet below are literally <code>#kgen.target&lt;...&gt;</code> attributes. That is why the dialect that carries a program&#8217;s compile-time structure is mostly types and attributes.</p><p>The <strong>name KGEN is everywhere</strong> once you look for it. Modular&#8217;s documented build target for the compiler is <code>KGEN:mojo</code>. The runtime every Mojo executable links is <code>libKGENCompilerRTShared.so</code>. The tools directory contains <code>kgen</code>, <code>kgen-doc</code> and <code>kgen-reduce</code>, and part of the test suite lives under <code>test/kgen</code>. </p><p>In our <strong>July piece on MLIR we described KGEN</strong>, from the public material available then, as a proprietary kernel generator that sat as a layer above Mojo. With the source open, that picture was wrong. KGEN is the compiler&#8217;s historical name and one of its dialects, sitting inside Mojo, not above it. </p><p>The directory list also has an <code>Interpreter</code> and an <code>Elaborator</code>. By their names and the passes they feed, the first runs <code>comptime</code> code inside the compiler and the second instantiates parametric code, which is where a large share of every build&#8217;s time goes.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Where the compile time goes</h2><p>Mojo can time its own pipeline. <code>--mlir-timing</code> reports every MLIR pass, <code>--llvm-timing</code> every LLVM pass, and <code>--timing-json</code> writes both as one object. <em>Mojo caches compiled modules </em>in <code>MODULAR_CACHE_DIR</code> and a warm cache skips most of the pipeline, so every measurement here points it at an empty directory first.</p><p>For the &#8220;hello world&#8221; example, 46 distinct passes run, plus five analyses. The four largest are importing the Mojo modules at 16.6% of pipeline time, <code>VerifyParameters</code> at 16.1%, <code>LowerLIT</code> at 13.1% and <code>InlineParametric</code> at 12.8%, followed by <code>CheckLifetimes</code>, the borrow checker, at 6.8%. </p><p>For a program that compiles the vector kernel for <code>sm_90a</code> as an offload target, the order shifts: imports 19.3%, <code>ElaborateGenerators</code> 15.4%, <code>VerifyParameters</code> 12.7%, <code>InlineParametric</code> 9.4%, <code>LowerLIT</code> 8.6%, <code>CheckLifetimes</code> 4.7%. </p><p>The three passes that exist to instantiate and check parametric code take 31.7% of the hello-world pipeline and 37.5% of the GPU one. The tree view shows the elaborator working in rounds: <code>InlineParametric</code> followed by <code>VerifyParameters</code>, repeated, with dead-symbol elimination and canonicalization in between.</p><p>Then there is LLVM. In the timing run for the <code>sm_90a</code> program, which the <strong>LLVM timers force onto one thread</strong>, the MLIR pipeline took 5.92 seconds and every LLVM pass on both the x86 host and the NVPTX offload added up to 15.6 milliseconds, or 0.26% of the build. Instruction selection for the whole GPU kernel took less than half a millisecond. </p><p>In our piece on Triton, <strong>Triton&#8217;s own MLIR pipeline was 81% of compile time</strong> and <code>ptxas</code> the other 19%. In Mojo the assembler is not in the build at all, because PTX is turned into machine code when the program loads, and LLVM is a rounding error. Nearly everything the compiler spends its time on is deciding what your program means.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!UU6n!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F308356c7-bee1-4f15-b745-72176f4d9bba_1806x987.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!UU6n!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F308356c7-bee1-4f15-b745-72176f4d9bba_1806x987.png 424w, https://substackcdn.com/image/fetch/$s_!UU6n!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F308356c7-bee1-4f15-b745-72176f4d9bba_1806x987.png 848w, https://substackcdn.com/image/fetch/$s_!UU6n!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F308356c7-bee1-4f15-b745-72176f4d9bba_1806x987.png 1272w, https://substackcdn.com/image/fetch/$s_!UU6n!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F308356c7-bee1-4f15-b745-72176f4d9bba_1806x987.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!UU6n!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F308356c7-bee1-4f15-b745-72176f4d9bba_1806x987.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/308356c7-bee1-4f15-b745-72176f4d9bba_1806x987.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Share of the Mojo MLIR pipeline by pass, cold builds on one core. The parameter passes (VerifyParameters, InlineParametric, ElaborateGenerators) grow from 31.7% to 37.5% when a GPU offload target is added.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Share of the Mojo MLIR pipeline by pass, cold builds on one core. The parameter passes (VerifyParameters, InlineParametric, ElaborateGenerators) grow from 31.7% to 37.5% when a GPU offload target is added." title="Share of the Mojo MLIR pipeline by pass, cold builds on one core. The parameter passes (VerifyParameters, InlineParametric, ElaborateGenerators) grow from 31.7% to 37.5% when a GPU offload target is added." srcset="https://substackcdn.com/image/fetch/$s_!UU6n!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F308356c7-bee1-4f15-b745-72176f4d9bba_1806x987.png 424w, https://substackcdn.com/image/fetch/$s_!UU6n!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F308356c7-bee1-4f15-b745-72176f4d9bba_1806x987.png 848w, https://substackcdn.com/image/fetch/$s_!UU6n!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F308356c7-bee1-4f15-b745-72176f4d9bba_1806x987.png 1272w, https://substackcdn.com/image/fetch/$s_!UU6n!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F308356c7-bee1-4f15-b745-72176f4d9bba_1806x987.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 3. Share of the Mojo MLIR pipeline by pass, cold builds on one core. The parameter passes (VerifyParameters, InlineParametric, ElaborateGenerators) grow from 31.7% to 37.5% when a GPU offload target is added.</figcaption></figure></div><p>That made us expect compile time to grow with the number of specializations, since <strong>generating kernels instead of enumerating them</strong> is the whole bet behind the language. </p><p>So we built programs that instantiate a small parametric function <code>work[N]</code> once, 8, 32, 128 and 256 times through a <code>comptime for</code>, each from an empty cache. <strong>We ran each three times</strong>. The medians were 3.08, 3.09, 3.08, 3.18 and 3.21 seconds, and a least-squares line through them gives about 0.56 milliseconds per specialization on a floor of about 3.1 seconds: two hundred and fifty-six specializations add roughly 0.14 seconds. </p><p>At this size an extra specialization costs almost nothing. The floor is paid before any user code. An empty <code>def main(): pass</code> builds in 2.92 seconds cold, median of three, and its pipeline looks like hello world&#8217;s: importing modules 18.1%, <code>VerifyParameters</code> 16.7%. <strong>With no user code at all,</strong> that is the standard library being imported and its parametric code verified. Adding <code>print(1)</code> costs about 0.13 seconds more, <code>print("hello")</code> about 0.06. </p><p><strong>The honest caveat is that our bodies were quite small. </strong>A kernel that unrolls deeply at compile time will cost more, and the compiler has a <code>--loop-unrolling-warn-threshold</code> flag, defaulting to 1024, for when it does.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!ZDEs!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F04e0fe7b-71a6-4baf-812d-13a97cd49734_1806x987.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!ZDEs!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F04e0fe7b-71a6-4baf-812d-13a97cd49734_1806x987.png 424w, https://substackcdn.com/image/fetch/$s_!ZDEs!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F04e0fe7b-71a6-4baf-812d-13a97cd49734_1806x987.png 848w, https://substackcdn.com/image/fetch/$s_!ZDEs!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F04e0fe7b-71a6-4baf-812d-13a97cd49734_1806x987.png 1272w, https://substackcdn.com/image/fetch/$s_!ZDEs!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F04e0fe7b-71a6-4baf-812d-13a97cd49734_1806x987.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!ZDEs!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F04e0fe7b-71a6-4baf-812d-13a97cd49734_1806x987.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/04e0fe7b-71a6-4baf-812d-13a97cd49734_1806x987.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Cold build time against the number of distinct comptime specializations in one program, median of three runs with the full range. Almost all of the cost is a fixed floor; the slope is about half a millisecond per specialization for small bodies.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Cold build time against the number of distinct comptime specializations in one program, median of three runs with the full range. Almost all of the cost is a fixed floor; the slope is about half a millisecond per specialization for small bodies." title="Cold build time against the number of distinct comptime specializations in one program, median of three runs with the full range. Almost all of the cost is a fixed floor; the slope is about half a millisecond per specialization for small bodies." srcset="https://substackcdn.com/image/fetch/$s_!ZDEs!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F04e0fe7b-71a6-4baf-812d-13a97cd49734_1806x987.png 424w, https://substackcdn.com/image/fetch/$s_!ZDEs!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F04e0fe7b-71a6-4baf-812d-13a97cd49734_1806x987.png 848w, https://substackcdn.com/image/fetch/$s_!ZDEs!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F04e0fe7b-71a6-4baf-812d-13a97cd49734_1806x987.png 1272w, https://substackcdn.com/image/fetch/$s_!ZDEs!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F04e0fe7b-71a6-4baf-812d-13a97cd49734_1806x987.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 4. Cold build time against the number of distinct comptime specializations in one program, median of three runs with the full range. Almost all of the cost is a fixed floor; the slope is about half a millisecond per specialization for small bodies.</figcaption></figure></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>The target sweep</h2><p><code>mojo build --print-supported-accelerators</code> lists every GPU the compiler will target. In 1.1.0 that is fourteen AMD architectures from <code>gfx90a</code> (MI250X) through <strong>CDNA3 and CDNA4 to RDNA2, RDNA3, RDNA3.5 and RDNA4 consumer parts</strong>, plus the <code>mi300a</code> APU name and two aliases. </p><p>Ten Apple entries, M1 through M5 each with and without Metal 4; and twenty NVIDIA names from <code>sm_52</code> (Maxwell) to <code>sm_121a</code> (DGX Spark). <code>--print-supported-targets</code> lists the host CPU backends: the AArch64 family, 32 and 64-bit RISC-V in both endiannesses, and 32 and 64-bit x86. There is no Hexagon, no Qualcomm accelerator, no TPU and no Trainium in either list, which will matter at the end.</p><p><strong>Modular&#8217;s GPU package </strong>exposes a function, <code>_compile_code</code>, that compiles a kernel for a named target at compile time and returns the assembly as a string. You don&#8217;t need to own any particular device. </p><p>We compiled the same small kernel for all twenty NVIDIA names and fed every result to NVIDIA&#8217;s <code>ptxas</code> 13.4.92, taken from the <code>nvidia-cuda-nvcc</code> wheel on PyPI.</p><pre><code><code>from max.gpu.host.compile import _compile_code, get_gpu_target
from max.gpu import thread_idx, block_idx, block_dim

def vadd(c: Pointer[Float32, MutAnyOrigin], a: Pointer[Float32, ImmutAnyOrigin], n: Int):
    var i = Int(block_idx.x * block_dim.x + thread_idx.x)
    if i &lt; n:
        c[unsafe_offset=i] = a[unsafe_offset=i] * 2.0

def main() raises:
    print(_compile_code[vadd, target = get_gpu_target["sm_90"]()]())</code></code></pre><p>Five names come back with a suffix nobody asked for: <code>sm_90</code> becomes <code>sm_90a</code>, <code>sm_100</code> becomes <code>sm_100a</code>, <code>sm_103</code> becomes <code>sm_103a</code>, <code>sm_120</code> becomes <code>sm_120a</code> and <code>sm_121</code> becomes <code>sm_121a</code>. </p><p>One name goes the other way: <code>sm_87</code>, Jetson Orin, comes back as <code>.target sm_80</code>. And <code>sm_110a</code>, Jetson Thor, loses the suffix it asked for and comes back as <code>sm_110</code>. </p><p>The<strong> PTX ISA version follows the family</strong>: 5.0 for Maxwell and Pascal, 6.3 for Turing, 8.1 for Ampere and Ada, 8.5 for Hopper, 8.7 for <code>sm_120</code>, 8.8 for <code>sm_100</code>, <code>sm_103</code> and <code>sm_121</code>, and 9.0 for Thor. The path that real programs use behaves the same way. </p><p>Built with <code>mojo build --target-accelerator sm_90 --emit asm</code>, a program that compiles the kernel through <code>DeviceContext.compile_function</code> writes a PTX file next to its assembly that says <code>.target sm_90a</code>, and the same build for <code>sm_100</code> and <code>sm_87</code> says <code>sm_100a</code> and <code>sm_80</code>.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!fo5i!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F77466de3-2b27-482c-a4c7-a261a8d14614_1806x987.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!fo5i!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F77466de3-2b27-482c-a4c7-a261a8d14614_1806x987.png 424w, https://substackcdn.com/image/fetch/$s_!fo5i!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F77466de3-2b27-482c-a4c7-a261a8d14614_1806x987.png 848w, https://substackcdn.com/image/fetch/$s_!fo5i!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F77466de3-2b27-482c-a4c7-a261a8d14614_1806x987.png 1272w, https://substackcdn.com/image/fetch/$s_!fo5i!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F77466de3-2b27-482c-a4c7-a261a8d14614_1806x987.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!fo5i!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F77466de3-2b27-482c-a4c7-a261a8d14614_1806x987.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/77466de3-2b27-482c-a4c7-a261a8d14614_1806x987.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;The same kernel compiled for every NVIDIA name Mojo 1.1 accepts. Red: silently upgraded to the architecture-specific target. Gold: downgraded or stripped of its suffix. Slate: emitted, but rejected by the CUDA 13.4 assembler.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="The same kernel compiled for every NVIDIA name Mojo 1.1 accepts. Red: silently upgraded to the architecture-specific target. Gold: downgraded or stripped of its suffix. Slate: emitted, but rejected by the CUDA 13.4 assembler." title="The same kernel compiled for every NVIDIA name Mojo 1.1 accepts. Red: silently upgraded to the architecture-specific target. Gold: downgraded or stripped of its suffix. Slate: emitted, but rejected by the CUDA 13.4 assembler." srcset="https://substackcdn.com/image/fetch/$s_!fo5i!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F77466de3-2b27-482c-a4c7-a261a8d14614_1806x987.png 424w, https://substackcdn.com/image/fetch/$s_!fo5i!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F77466de3-2b27-482c-a4c7-a261a8d14614_1806x987.png 848w, https://substackcdn.com/image/fetch/$s_!fo5i!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F77466de3-2b27-482c-a4c7-a261a8d14614_1806x987.png 1272w, https://substackcdn.com/image/fetch/$s_!fo5i!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F77466de3-2b27-482c-a4c7-a261a8d14614_1806x987.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 5. The same kernel compiled for every NVIDIA name Mojo 1.1 accepts. Red: silently upgraded to the architecture-specific target. Gold: downgraded or stripped of its suffix. Slate: emitted, but rejected by the CUDA 13.4 assembler.</figcaption></figure></div><p>NVIDIA&#8217;s assembler makes the consequences concrete, and its own help text states the rules. PTX for <code>sm_XY</code> compiles to any target at or above <code>XY</code>, suffixed or not. PTX for the family target <code>sm_XYf</code> compiles to members of the same family at or above it. </p><p>PTX for <code>sm_XYa</code> compiles to <code>sm_XYa</code> and nothing else. Every one of Mojo&#8217;s upgraded outputs therefore fails when handed to <code>ptxas</code> with the architecture that was requested, for example:</p><pre><code><code>$ ptxas -arch sm_90 vadd_for_sm90.ptx
ptxas fatal   : PTX with .target 'sm_90a' cannot be compiled for architecture 'sm_90'
$ ptxas -arch sm_100 vadd_for_sm90.ptx
ptxas fatal   : Program with .target 'sm_90a' cannot be compiled to future architecture</code></code></pre><p>Change nothing but the directive back to <code>.target sm_90</code> and the same body assembles for <code>sm_100</code> and for <code>sm_120</code>. <strong>The same holds one generation later. </strong>Mojo&#8217;s output for a B200, <code>sm_100a</code>, is refused for the B300&#8217;s <code>sm_103</code>. Rewritten as the family target <code>sm_100f</code>, it assembles for <code>sm_103</code> and is refused for <code>sm_120</code>, which belongs to a different family. </p><p>Of the twenty names Mojo accepts, twelve produce PTX that <code>ptxas</code> 13.4 will assemble for the architecture that was asked for. The three Maxwell and Pascal names fail earlier: <code>ptxas</code> 13.4 does not recognize <code>sm_52</code>, <code>sm_60</code> or <code>sm_61</code> as architectures at all. </p><p><strong>Mojo lists them, and emits PTX 5.0 for them</strong>, but the current CUDA assembler cannot finish the job. Modular knows. A string inside the GPU runtime tells users of older hardware to point <code>MODULAR_NVPTX_COMPILER_PATH</code> at an external <code>ptxas</code>.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>A target is a product, not an instruction set</h2><p>The suffix is not produced by a transformation anywhere in the compiler. It comes from a table. In the shipped 1.1.0 the table lives in <code>info.mojo</code>, where the H100 record already says <code>sm_90a</code>. </p><p>On the main branch it has moved to <code>Mojo/stdlib/std/_gpu/host/_builtin_targets.mojo</code>, and the first thing that file does is map command-line names to products:</p><pre><code><code>        A10       ._with_cli_values[["sm_86"]],
        A100      ._with_cli_values[["sm_80"]],
        OrinNano  ._with_cli_values[["sm_87"]],
        L4        ._with_cli_values[["sm_89"]],
        RTX4090   ._with_cli_values[[]       ], # Could be ["sm_89"], but `sm_89` resolves to L4.
        B100      ._with_cli_values[[]                   ], # Could be ["sm_100", "sm_100a"], but both resolve to B200.
        B200      ._with_cli_values[["sm_100", "sm_100a"]],
        H100      ._with_cli_values[["sm_90" , "sm_90a" ]],</code></code></pre><p>Each product carries a compilation target written as a <code>kgen</code> attribute, and the H100&#8217;s is where the letter comes from:</p><pre><code><code>comptime _h100_target = CompilationTarget[
    _mlir_value=__mlir_attr[
        `#kgen.target&lt;triple = "nvptx64-nvidia-cuda", `,
        `stdlib_plugin = "cuda", `,
        `arch = "sm_90a", `,
        `features = "+ptx85,+sm_90a", `,
        `tune_cpu = "sm_90a", `,
        ...</code></code></pre><p>So <code>sm_90</code> is not an instruction set to Mojo. It is a lookup key that resolves to &#8220;an H100&#8221;, and an H100 is compiled as <code>sm_90a</code> with PTX 8.5, 132 streaming multiprocessors in its record. </p><p>We parsed the whole table: 44 device records and 43 target records, because the B100 and the B200 share one. <code>sm_89</code> resolves to an L4, so an RTX 4090 cannot be selected by name. <code>sm_86</code> resolves to an A10. </p><p>On main, <strong>the lookup itself is a compile-time Mojo program that walks a type list</strong> of target collections, and the list has two members: <code>BuiltinTargets</code> and one called <code>ADDITIONAL_TARGETS</code>. The device registry is pluggable at compile time, which we come back to at the end.</p><p>Whether the suffix is a bug depends on where the compiler is when the kernel gets built. MAX carries it at runtime. <code>libmax.so</code> contains the pipeline we timed above: the pass names <code>LowerLIT</code>, <code>ElaborateGenerators</code>, <code>InlineParametric</code>, <code>VerifyParameters</code>, <code>CheckLifetimes</code> and <code>LowerKGENToLLVM</code> are all in it, <strong>next to 255 strings that mention NVPTX and 466 that mention AMDGPU</strong>. When MAX compiles a kernel for the GPU in the machine, <code>sm_90a</code> unlocks every instruction the H100 has and costs nothing.</p><p>An executable made with <code>mojo build</code> is a different object. We built one for <code>sm_90</code>. It is 44,512 bytes, links only <code>libKGENCompilerRTShared.so</code>, <code>libAsyncRTMojoBindings.so</code> and the C library, and contains no pass names at all. What it does contain is <strong>one 1,354-byte PTX module</strong> that says <code>.target sm_90a</code>. NVIDIA&#8217;s assembler builds that module for <code>sm_90a</code> and refuses it for <code>sm_100</code>, <code>sm_103</code> and <code>sm_120</code> as a future architecture. </p><p>Change the one directive to <code>sm_90</code> and it builds for all three. So the <strong>kernels of an H100 executable cannot be carried to Blackwell</strong> by NVIDIA&#8217;s compiler, where a plain target would have let the driver compile them forward, and nothing in the build said so.</p><p>Triton makes the same choice as Mojo: in our Triton piece, its <code>sm_arch_from_capability</code> appends the suffix to every capability from 9.0 upward, with a live TODO next to it. NVIDIA&#8217;s own Tile IR makes the opposite one. </p><p>Its <code>tileiras</code> backend accepts eleven targets and none with the suffix, and buys forward compatibility by excluding the instructions that would need it. <strong>Three serious projects, two good answers.</strong> Mojo and Triton recompile for the exact chip. cuTile narrows what a kernel can say. </p><p>The family targets sit between them, and the NVIDIA assembler bundled in <code>max-core</code> knows <code>sm_100f</code>, <code>sm_103f</code>, <code>sm_110f</code>, <code>sm_120f</code> and <code>sm_121f</code>. Mojo&#8217;s table asks for none of them.</p><p>So the cost of the choice depends on where portability has to live. In MAX it lives in the source plus a runtime that carries a full compiler inside the 162.6 MB <code>libmax.so</code>, with <strong>NVIDIA&#8217;s assembler beside it in the wheel. </strong></p><p>In a <code>mojo build</code> executable it ends at build time, and the executable is exactly as portable as a CUDA binary compiled only for <code>sm_90a</code>.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Where the table is wrong</h2><p>A device-centric design is only as good as the device table, and a <strong>hand-maintained table of 44 products </strong>accumulates small errors. We found four. They matter less individually than as a pattern.</p><p>The first explains Jetson Orin. The <code>OrinNano</code> record on main says <code>arch = "sm_87"</code>, yet 1.1.0 emitted <code>sm_80</code>. The emitted LLVM IR already carried <code>"target-cpu"="sm_80"</code> with <code>+ptx81,+sm_80</code>, so the substitution happens in Mojo&#8217;s own code, <strong>before LLVM sees the kernel.</strong> The answer is in the 1.1.0 tag, where the table had a different structure. </p><p><code>Mojo/stdlib/std/_gpu/host/info.mojo</code> maps <code>"sm_87"</code> to <code>OrinNano</code> correctly, then a method named <code>target()</code> turns a device into its compilation target through a chain of name comparisons that ends like this:</p><pre><code><code>        if self.name == "A100":
            return _get_a100_target()
        ...
        if self.name == "Jetson Thor":
            return _get_jetson_thor_target()
        ...
        if self.name == "":
            return _get_empty_target()
        return _get_a100_target()</code></code></pre><p>The chain has 43 name comparisons, and every product has a branch except one. There is no <code>"Orin Nano"</code> case, so the Orin falls through to the default, which is the A100. </p><p>Code built this way still runs on an Orin, because <strong>Ampere binaries are compatible within the family</strong>, so this is not a crash. It is a silent substitution, and the failure mode it represents is the one to worry about: an unrecognized device becoming an A100 instead of becoming an error. </p><p>Main has since restructured the table so that each device is declared together with its target record, as <code>TargetAccelerator[GPUInfo, target]</code>, which removes this failure mode by construction. The shipped 1.1.0 still has it.</p><p><strong>The second is an RTX 3090 record</strong>, unreachable from the command line because <code>sm_86</code> resolves to the A10, whose features read <code>+ptx63,+sm_86</code>. PTX 6.3 predates <code>sm_86</code>. Compiling our kernel against that record, the way the standard library itself would, stops LLVM outright:</p><pre><code><code>LLVM ERROR: PTX version 6.3 does not support target 'sm_86'. Minimum required
PTX version is 7.1. Either remove the PTX version to use the default, or increase
it to at least 7.1.</code></code></pre><p>Like the GTX 1080 Ti, GTX 1060, GTX 970 and Tesla P100 records, its name is<strong> NVIDIA&#8217;s full marketing string</strong>, <code>"NVIDIA GeForce RTX 3090"</code>, which suggests these entries exist to match physical cards by name at runtime. We could not test whether a real 3090 selects it. The entry is identical at the 1.1.0 tag and on main.</p><p>The third is cosmetic: the RTX 4090 and 4090 mobile records compile for <code>sm_89</code> but set <code>tune_cpu</code> to <code>sm_90a</code>, Hopper&#8217;s tuning, in both versions. In practice the NVPTX backend does little architecture-specific scheduling, and <code>ptxas</code> does the real work, so we expect no measurable effect. </p><p>The <strong>fourth is Jetson Thor</strong>, whose record asks for <code>sm_110</code> without the suffix. That one does have consequences, and they show up with the tensor cores.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Where the tensor cores are</h2><p>MAX exposes a <strong>warp-level matrix multiply-accumulate </strong>as one generic function, <code>mma(d, a, b, c)</code> in <code>max.gpu.compute.mma</code>, taking four SIMD values. Its body is a compile-time dispatch on the target: <code>_mma_nvidia</code> if the target is an NVIDIA GPU, <code>_mma_amd</code> if it is AMD, and <code>_mma_apple</code> only if the target is an Apple M5 with the <code>metal4_0</code> feature. </p><p>We compiled a one-call kernel against every family and read what came out.</p><pre><code><code>def k(o: Pointer[Float32, MutAnyOrigin]):
    var a = SIMD[DType.bfloat16, 8](1)
    var b = SIMD[DType.bfloat16, 4](1)
    var c = SIMD[DType.float32, 4](0)
    var d = SIMD[DType.float32, 4](0)
    mma(d, a, b, c)
    o[unsafe_offset=0] = d[0] + d[3]</code></code></pre><p>With those operand widths, <code>sm_80</code>, <code>sm_90a</code>, <code>sm_100a</code> and <code>sm_120a</code> all produce the same PTX instruction, <code>mma.sync.aligned.m16n8k16.row.col.f32.bf16.bf16.f32</code>. On Turing, <code>sm_75</code>, which has no bfloat16 tensor cores, compilation does not fail with a Mojo diagnostic. LLVM aborts: <code>Cannot select: intrinsic %llvm.nvvm.mma.m16n8k16.row.col.bf16</code>. </p><p>Change the fragments to four lanes each and the AMD targets compile: <code>gfx90a</code> emits <code>v_mfma_f32_16x16x16bf16_1k</code>, while <code>gfx942</code> and <code>gfx950</code> emit <code>v_mfma_f32_16x16x16_bf16</code>. RDNA3 rejects four-lane fragments inside its own implementation file and accepts sixteen-lane bfloat16 inputs with eight accumulators, emitting <code>v_wmma_f32_16x16x16_bf16</code>. </p><p><strong>The M5 under Metal 4 wants eight lanes of fp16 in and eight of fp32 out</strong>, and emits the AIR intrinsic <code>air.simdgroup_matrix_16x16x16_multiply_accumulate</code>. Hand it four lanes and a constraint explains that Apple MMA requires eight-element fragments. </p><p>On an M4, a <strong>clean compile-time error</strong> says the target does not support the operation at all.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!ODaB!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb10a311a-9dde-4bd5-bee5-3d287f729bb1_1806x987.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!ODaB!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb10a311a-9dde-4bd5-bee5-3d287f729bb1_1806x987.png 424w, https://substackcdn.com/image/fetch/$s_!ODaB!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb10a311a-9dde-4bd5-bee5-3d287f729bb1_1806x987.png 848w, https://substackcdn.com/image/fetch/$s_!ODaB!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb10a311a-9dde-4bd5-bee5-3d287f729bb1_1806x987.png 1272w, https://substackcdn.com/image/fetch/$s_!ODaB!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb10a311a-9dde-4bd5-bee5-3d287f729bb1_1806x987.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!ODaB!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb10a311a-9dde-4bd5-bee5-3d287f729bb1_1806x987.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b10a311a-9dde-4bd5-bee5-3d287f729bb1_1806x987.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;One call to the generic mma() across NVIDIA, AMD and Apple targets. Five distinct instructions from one function name, but four different operand shapes: the widths of the SIMD arguments encode each vendor's fragment layout.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="One call to the generic mma() across NVIDIA, AMD and Apple targets. Five distinct instructions from one function name, but four different operand shapes: the widths of the SIMD arguments encode each vendor's fragment layout." title="One call to the generic mma() across NVIDIA, AMD and Apple targets. Five distinct instructions from one function name, but four different operand shapes: the widths of the SIMD arguments encode each vendor's fragment layout." srcset="https://substackcdn.com/image/fetch/$s_!ODaB!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb10a311a-9dde-4bd5-bee5-3d287f729bb1_1806x987.png 424w, https://substackcdn.com/image/fetch/$s_!ODaB!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb10a311a-9dde-4bd5-bee5-3d287f729bb1_1806x987.png 848w, https://substackcdn.com/image/fetch/$s_!ODaB!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb10a311a-9dde-4bd5-bee5-3d287f729bb1_1806x987.png 1272w, https://substackcdn.com/image/fetch/$s_!ODaB!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb10a311a-9dde-4bd5-bee5-3d287f729bb1_1806x987.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 6. One call to the generic mma() across NVIDIA, AMD and Apple targets. Five distinct instructions from one function name, but four different operand shapes: the widths of the SIMD arguments encode each vendor&#8217;s fragment layout.</figcaption></figure></div><p>So one function name really does reach five distinct matrix instructions in four families of matrix hardware. It does not do so from one piece of code. </p><p>The operand widths are the vendor&#8217;s register fragment layout, and they differ on every family: eight, four and <strong>four lanes on NVIDIA</strong> for this shape, four, four and four on CDNA, sixteen, sixteen and eight on RDNA3, eight, eight and eight on Apple. </p><p>The abstraction that actually hides this is the <code>layout</code> package. Its <code>TensorCore</code> type is parameterized on output type, input type, an MMA shape and a transpose flag, and computes the fragment shapes for the target at compile time. </p><p>This is the <strong>precise sense in which Mojo is portable</strong>: the portability is in parameters evaluated at compile time, not in the operation. It is a good design. It also means that the thing a kernel author writes against is <code>layout</code>, which ships with MAX, not with Mojo.</p><p>The Hopper warpgroup multiply shows the other edge. <code>wgmma_async</code> compiled for <code>sm_90a</code> emits <code>wgmma.mma_async.sync.aligned.m64n64k16.f32.bf16.bf16</code>, and NVIDIA&#8217;s assembler accepts it. Compiled for <code>sm_100a</code>, <code>sm_120a</code> or <code>sm_80</code>, it produces no Mojo diagnostic. LLVM aborts on <code>%llvm.nvvm.wgmma.commit_group.sync.aligned</code>. The function&#8217;s compile-time asserts check shapes and scale factors, and we found none that checks the target. </p><p><strong>It&#8217;s a third, independent confirmation</strong> of what we measured in our piece on Blackwell&#8217;s tensor memory, that <code>wgmma</code> exists on no Blackwell target at all. It also shows where Mojo&#8217;s knowledge of the hardware ends and LLVM&#8217;s begins.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>The tcgen05 guard</h2><p><strong>Blackwell&#8217;s datacenter tensor cores</strong> are programmed through the <code>tcgen05</code> family and a fifth memory space, tensor memory, which has to be allocated before use. MAX guards these instructions. </p><p>The file&#8217;s ten public <code>tcgen05</code> functions make eleven calls to one constraint, identical at the 1.1.0 tag and on main:</p><pre><code><code># max/mojo/max/gpu/compute/arch/tcgen05.mojo
def check_blackwell_constraint():
    comptime assert _has_blackwell_tcgen05(), (
        "The tcgen05 instructions are only applicable on nVidia Blackwell"
        " (sm_100a, sm_101a) hardware."
    )

# Mojo/stdlib/std/sys/info.mojo
comptime _SM_101X_ARCHS: List[StaticString] = ["sm_101", "sm_101a"]

def _has_blackwell_tcgen05() -&gt; Bool:
    return _has_nvidia_gpu_any[
        _SM_100X_ARCHS + _SM_101X_ARCHS + _SM_103X_ARCHS
    ]()

def _has_nvidia_gpu_any[archs: List[StaticString]]() -&gt; Bool:
    comptime if not has_nvidia_gpu_accelerator():
        return False
    comptime for arch in archs:
        comptime if arch.removeprefix("sm_") in _accelerator_arch():
            return True
    return False</code></code></pre><p>Two details in that code decide everything. The first is the choice of <code>_has_nvidia_gpu_any</code>. It asks whether the build has an <strong>NVIDIA accelerator configured,</strong> from the <code>--target-accelerator</code> flag or the detected GPU, and compares against that accelerator&#8217;s name. It does not ask what the kernel is being compiled for. </p><p>The same file has a sibling, <code>_is_nvidia_gpu_any</code>, that checks the compilation target. The second detail is the list: <code>sm_101</code> was the name of <strong>Jetson Thor&#8217;s architecture</strong> before CUDA 13 renamed it <code>sm_110</code>, and Mojo&#8217;s own device table uses the new name. NVIDIA&#8217;s current assembler no longer accepts <code>sm_101a</code> at all.</p><p>We compiled a kernel that allocates tensor memory with every combination that matters, then gave each result to <code>ptxas</code> 13.4:</p><p><strong>Kernel targetBuild acceleratorMojo 1.1ptxas 13.4</strong><code>sm_100a</code>none (no flag, no GPU)refused by the guardnot reached<code>sm_100asm_100a</code>emits <code>tcgen05.alloc.cta_group::1.sync.aligned.shared::cta.b32</code>accepted<code>sm_103asm_103a</code>emits the sameaccepted<code>sm_100asm_90a</code>refused by the guardnot reached<code>sm_90asm_100a</code><strong>emits tcgen05 into sm_90a PTX</strong>rejected: tcgen05.alloc needs PTX ISA 8.6<code>sm_110a</code> (Thor)<code>sm_110a</code>refused by the guardaccepts tcgen05 on <code>sm_110a</code> and <code>sm_110f</code>, refuses it on <code>sm_110sm_120a</code>, <code>sm_121a</code>same as the targetrefused by the guardrefuses tcgen05 on both</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!k83t!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F119f4ced-250f-47d5-845b-d9c2ee179004_1806x987.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!k83t!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F119f4ced-250f-47d5-845b-d9c2ee179004_1806x987.png 424w, https://substackcdn.com/image/fetch/$s_!k83t!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F119f4ced-250f-47d5-845b-d9c2ee179004_1806x987.png 848w, https://substackcdn.com/image/fetch/$s_!k83t!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F119f4ced-250f-47d5-845b-d9c2ee179004_1806x987.png 1272w, https://substackcdn.com/image/fetch/$s_!k83t!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F119f4ced-250f-47d5-845b-d9c2ee179004_1806x987.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!k83t!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F119f4ced-250f-47d5-845b-d9c2ee179004_1806x987.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/119f4ced-250f-47d5-845b-d9c2ee179004_1806x987.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Where tcgen05 is legal according to NVIDIA's assembler, against what Mojo 1.1 will emit. The two disagree on Jetson Thor, and the guard can be satisfied by a build flag while the kernel targets Hopper.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Where tcgen05 is legal according to NVIDIA's assembler, against what Mojo 1.1 will emit. The two disagree on Jetson Thor, and the guard can be satisfied by a build flag while the kernel targets Hopper." title="Where tcgen05 is legal according to NVIDIA's assembler, against what Mojo 1.1 will emit. The two disagree on Jetson Thor, and the guard can be satisfied by a build flag while the kernel targets Hopper." srcset="https://substackcdn.com/image/fetch/$s_!k83t!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F119f4ced-250f-47d5-845b-d9c2ee179004_1806x987.png 424w, https://substackcdn.com/image/fetch/$s_!k83t!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F119f4ced-250f-47d5-845b-d9c2ee179004_1806x987.png 848w, https://substackcdn.com/image/fetch/$s_!k83t!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F119f4ced-250f-47d5-845b-d9c2ee179004_1806x987.png 1272w, https://substackcdn.com/image/fetch/$s_!k83t!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F119f4ced-250f-47d5-845b-d9c2ee179004_1806x987.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 7. Where tcgen05 is legal according to NVIDIA&#8217;s assembler, against what Mojo 1.1 will emit. The two disagree on Jetson Thor, and the guard can be satisfied by a build flag while the kernel targets Hopper.</figcaption></figure></div><p>On a GPU-less build machine, a common kind of CI runner, public <code>tcgen05</code> code does not compile for a B200 unless someone passes <code>--target-accelerator sm_100a</code>. On a machine configured for an H100, a correct B200 kernel is refused. On a machine configured for a B200, the guard lets <code>tcgen05</code> into <strong>Hopper PTX</strong> that NVIDIA&#8217;s assembler then rejects. </p><p>It&#8217;s possible that<strong> reading the build host is intentional</strong>, a way of instantiating Blackwell paths only when building for a Blackwell system, and that MAX&#8217;s own launch path always configures the accelerator to match. </p><p>The <em>measured behaviour </em>stands either way: in a language whose whole premise is compiling for hardware you are not running on, a hardware guard that consults the build host is the wrong primitive.</p><p>Jetson Thor gets both problems at once. Its architecture is missing from the guard&#8217;s list, which still says <code>sm_101</code>, and its device record compiles for <code>sm_110</code> without the suffix, where <code>ptxas</code> refuses <code>tcgen05</code> even when it is emitted. <strong>NVIDIA&#8217;s assembler</strong> accepts the identical instruction on <code>sm_110a</code> and <code>sm_110f</code>. As shipped, Mojo 1.1 cannot put Thor&#8217;s tensor memory to use, and fixing either blocker alone would not be enough.</p><p>The source explains why the Hopper and Blackwell failures looked so different: <code>wgmma</code> aborted inside LLVM, while <code>tcgen05</code> sailed through LLVM into Hopper PTX. </p><p>They reach the instruction set by different roads. The <strong>warpgroup commit reaches LLVM as an NVVM intrinsic</strong>, the one named in LLVM&#8217;s error, and LLVM&#8217;s instruction selector checks intrinsics against the target. <code>tcgen05.mojo</code> contains 17 inline-assembly strings and no LLVM intrinsic calls: <code>tcgen05.alloc</code> is written out as text, and LLVM passes inline assembly through without reading it. </p><p>For <code>tcgen05</code>, the Mojo guard is the only check between the source and NVIDIA&#8217;s assembler. It is not an isolated choice. Across the <code>max.gpu</code> package we count 128 inline-assembly call sites against <strong>76 LLVM-intrinsic call sites</strong>, alongside 296 compile-time asserts; across the kernel sources, 55 against 85, alongside 4,413 asserts.</p><p>Step back and there are three places where a portable stack can keep its knowledge of the hardware. The first is <strong>compile-time constraints written in the language,</strong> and when Mojo uses them the diagnostics are good: the M4&#8217;s clean refusal, the M5&#8217;s fragment-size message, the intent of the <code>tcgen05</code> guard. </p><p>The second is<strong> LLVM&#8217;s instruction selector,</strong> which knows what each target can encode but can only say so by aborting the process. The third is the vendor&#8217;s assembler, which knows everything and speaks last. Inline assembly is a road around the middle layer: it goes from a Mojo string to the vendor&#8217;s assembler with nothing in between except whatever constraint the library author remembered to write. </p><p>A portable language is only as good as the first layer. In Mojo 1.1, a good part of the hardware knowledge still lives in the second and third.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>The last mile belongs to the vendor</h2><p>On AMD, Mojo owns the path to machine code. The <strong>AMDGPU backend runs in process,</strong> the assembly it prints is the final GCN or RDNA instruction stream, and the runtime loads the result through <code>hipModuleLoadData</code> from <code>libamdhip64.so</code>. </p><p>We found no reference to AMD&#8217;s <code>comgr</code> library in the runtime. On Apple, Mojo stops earlier and deliberately: its output for an M-series target is LLVM IR with the <code>air64-apple-macosx</code> triple and Metal feature flags, handed to Apple&#8217;s toolchain to finish.</p><p>On NVIDIA, Mojo stops at PTX, and the rest is NVIDIA&#8217;s. The GPU runtime, <code>libMGPRT.so</code>, and <code>libmax.so</code> both reference <code>cuModuleLoadDataEx</code> from <code>libcuda.so.1</code>, the driver&#8217;s own just-in-time path, and the entry points of NVIDIA&#8217;s PTX compiler library, <code>nvPTXCompilerCreate</code> and <code>nvPTXCompilerCompile</code>. The library they look for, named in the runtime&#8217;s own strings, is the 37.7 MB <code>libNVPTX.so</code> in <code>max-core</code>, and despite the name it has nothing to do with LLVM&#8217;s NVPTX backend. </p><p>It exports 45,501 dynamic symbols, and 45,474 of them are prefixed <code>libnvptxcompiler_static_</code>. The <strong>other 27 are NVIDIA&#8217;s public API:</strong> fifteen <code>nvPTXCompiler</code> functions, eleven <code>nvLinker</code> functions for linking device code, and one JIT entry point. Its version string reads <code>Cuda compilation tools, release 13.3, V13.3.33</code>, and it depends on nothing but the C and C++ runtimes. </p><p>We&#8217;re talking about <code>ptxas</code> as a library, NVIDIA&#8217;s static PTX compiler repackaged as a shared object and redistributed inside the MAX wheel. The runtime&#8217;s error strings name the two ways around it: <code>MODULAR_NVPTX_LIBRARY_PATH</code> for a different copy of the library, and <code>MODULAR_NVPTX_COMPILER_PATH</code> for an external <code>ptxas</code> binary, recommended for older drivers and older hardware.</p><p>This is the part of the stack we called the<strong> load-bearing moat</strong> in our piece on NVIDIA&#8217;s compiler: since Kepler and Volta, the dependency interlocks for fixed-latency instructions live in control bits that <code>ptxas</code> writes, so the assembler is correctness-critical and not merely an optimizer. </p><p>Mojo replaces CUDA&#8217;s source language and a growing share of its library layer, and on NVIDIA it is finished by the same closed assembler as a <strong>CUDA C++ kernel.</strong> That&#8217;s not a criticism of Modular, to be fair. It is the most direct measurement we have of how far a language can reach on NVIDIA hardware, and the answer is up to PTX, one step short of the instructions the chip executes.</p><p>We measured nothing about speed, so the best evidence on what that reach buys comes from outside. Godoy and colleagues at <strong>Oak Ridge National Laboratory</strong>, in a paper at the SC&#8217;25 workshops, ported four scientific workloads to Mojo, a seven-point stencil, BabelStream, miniBUDE and Hartree-Fock, and compared them with CUDA on an H100 and HIP on an MI300A. </p><p>They found Mojo competitive with the vendor baselines on the memory-bound kernels, and <strong>gaps on AMD for atomic operations</strong> and for fast-math compute-bound kernels. The study predates 1.0 by almost a year, and it fits the picture here: the language reaches each vendor&#8217;s backend, and what remains is how well each backend is driven.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share The Software Frontier&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share The Software Frontier</span></a></p><div><hr></div><h2>The vendor libraries never left</h2><p>MAX is usually described, including by us in July, as a stack that needs no vendor libraries: its own kernels, generated by its own compiler. <code>max-core</code> contains five precompiled binding packages, <code>_cublas</code>, <code>_cudnn</code>, <code>_cufft</code>, <code>_rocblas</code> and <code>_miopen</code>, 1.25 MB together. </p><p><strong>Seven source files in MAX&#8217;s linear-algebra kernels</strong> reference vendor BLAS, including the Hopper and Blackwell matmul dispatchers. The Hopper dispatcher&#8217;s own comment says that on a miss, an unsupported configuration or no tuning for the shape, it returns a miss so the caller can fall back to vendor BLAS. </p><p>A helper named <code>_vendor_blas_fallback_disabled</code> turns that fallback off, and its documentation says it returns true only when the fallback has been disabled globally or a benchmark asks specifically for the Mojo kernel. The fallback is on unless someone turns it off.</p><p>The tuning is where the work is. The matmul tree contains 188 Hopper and 148 Blackwell tuning-config entries, most of them in lists keyed by exact problem shapes. One list is <strong>headed as the GEMM shapes of Llama 405B in FP8</strong>, with entries such as M=64, N=16384, K=2048. </p><p>That is the same structure we found in AMD&#8217;s hipBLASLt, where 16,177 kernels cover 2.9 million shape mappings: the moat is not the code generator but the <strong>campaign of measurements </strong>that decides which generated kernel to use for which shape. </p><p>The phrase &#8220;<em>no vendor library</em>&#8221; describes the tuned path. The default dependency graph still contains cuBLAS. We repeated Modular&#8217;s framing uncritically in our July piece on MLIR, and this is the correction.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share The Software Frontier&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share The Software Frontier</span></a></p><div><hr></div><h2>How portable the kernels are</h2><p>The language reaches<strong> four families of matrix hardware. </strong>The question left is how much of MAX&#8217;s kernel library does. We classified every Mojo file under <code>max/kernels/src</code>, 544,039 lines, by its path. </p><p>Files that live in a directory or carry a name for one vendor or architecture, such as <code>sm90</code>, <code>sm100</code>, <code>amd_structured</code> or <code>apple</code>, hold 156,621 NVIDIA-specific lines (28.8%), 93,998 AMD-specific lines (17.3%) and 11,903 Apple-specific lines (2.2%). </p><p>The remaining 281,517 lines (51.7%) sit on generic paths, and those<strong> still branch on the target at compile time.</strong> At least 48.3% of the kernel library is written for one vendor or one architecture.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Gajk!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb41c64a4-093c-4e25-9a7c-dd2a23b62cae_1806x987.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Gajk!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb41c64a4-093c-4e25-9a7c-dd2a23b62cae_1806x987.png 424w, https://substackcdn.com/image/fetch/$s_!Gajk!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb41c64a4-093c-4e25-9a7c-dd2a23b62cae_1806x987.png 848w, https://substackcdn.com/image/fetch/$s_!Gajk!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb41c64a4-093c-4e25-9a7c-dd2a23b62cae_1806x987.png 1272w, https://substackcdn.com/image/fetch/$s_!Gajk!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb41c64a4-093c-4e25-9a7c-dd2a23b62cae_1806x987.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Gajk!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb41c64a4-093c-4e25-9a7c-dd2a23b62cae_1806x987.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b41c64a4-093c-4e25-9a7c-dd2a23b62cae_1806x987.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Left: MAX's kernel sources by the vendor or architecture their file path names; a lower bound on specificity. Right: how GPU operations reach the instruction set, as inline assembly or as LLVM intrinsics.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Left: MAX's kernel sources by the vendor or architecture their file path names; a lower bound on specificity. Right: how GPU operations reach the instruction set, as inline assembly or as LLVM intrinsics." title="Left: MAX's kernel sources by the vendor or architecture their file path names; a lower bound on specificity. Right: how GPU operations reach the instruction set, as inline assembly or as LLVM intrinsics." srcset="https://substackcdn.com/image/fetch/$s_!Gajk!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb41c64a4-093c-4e25-9a7c-dd2a23b62cae_1806x987.png 424w, https://substackcdn.com/image/fetch/$s_!Gajk!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb41c64a4-093c-4e25-9a7c-dd2a23b62cae_1806x987.png 848w, https://substackcdn.com/image/fetch/$s_!Gajk!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb41c64a4-093c-4e25-9a7c-dd2a23b62cae_1806x987.png 1272w, https://substackcdn.com/image/fetch/$s_!Gajk!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb41c64a4-093c-4e25-9a7c-dd2a23b62cae_1806x987.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 8. Left: MAX&#8217;s kernel sources by the vendor or architecture their file path names; a lower bound on specificity. Right: how GPU operations reach the instruction set, as inline assembly or as LLVM intrinsics.</figcaption></figure></div><p>That is not a failure of the language. It is what fast kernels look like, and CUTLASS keeps separate kernel families per architecture generation for the same reason. </p><p>What Mojo changes is that the<strong> NVIDIA, AMD and Apple slices are written in one language,</strong> against one layout algebra and one set of compile-time abstractions, and share the half of the library that is generic. It also gives the ModCon claim a yardstick. </p><p>If adding an accelerator really takes ten times less effort than it used to, the next vendor&#8217;s slice should look more like Apple&#8217;s 11,903 lines than AMD&#8217;s 93,998. That is our inference, and it becomes checkable the day that slice appears.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>What Qualcomm bought, as far as the binary can tell</h2><p>The price is public. <strong>Modular raised a $250 million round in September 2025 at a $1.6 billion valuation</strong>, bringing total capital to $380 million, according to RuntimeWire. </p><p>Qualcomm announced the acquisition on 24 June 2026 in an all-stock deal of about $3.9 billion and closed it on 29 July, with <strong>Chris Lattner named executive vice president</strong> of advanced AI software and platforms. </p><p>That is 10.3 times the capital raised and 2.4 times the last private valuation, the best outcome in the census of GPU software exits we assembled in August. Then <strong>came Mojo 1.0 on 11 August,</strong> the open compiler on 18 August, and Mojo 1.1 on 17 September, which also opened the compiler to outside contributions.</p><p>At <strong>ModCon</strong>, according to the <em>NAND Research&#8217;s account</em>, Modular announced platform support for AWS Trainium, Google TPUs and Qualcomm&#8217;s Cloud AI100 and Dragonfly accelerators, said that integrating a new accelerator now takes more than<strong> ten times less engineering effort than before</strong>, and announced an alliance program for hardware vendors, model providers and clouds. </p><p>None of those accelerators appears in what we measured. Mojo 1.1.0 lists NVIDIA, AMD and Apple GPUs and nothing else. <strong>No source file in the open repository mentions Qualcomm</strong>; the word occurs only inside four test-data files, two GGUF model files and two tokenizer vocabularies. </p><p>Hexagon appears in a comment in a C++ support file, in a test fixture, and in a documentation page whose sample target listing includes it next to <code>amdgcn</code>. The <strong>shipped binary lists neither</strong>: its host backends are AArch64, RISC-V and x86.</p><p>The mechanism for adding them is visible, though. On main, the target lookup walks two collections of devices, the builtin one and <code>ADDITIONAL_TARGETS</code>. A vendor can supply its own collection of device records at compile time without editing the table everyone else uses. </p><p>Our reading, which is inference and not measurement, is that <strong>the announced backends live outside the open tree, </strong>at least for now, or reach customers in a form we cannot see, and that the plumbing for plugging them in is already public. </p><p>The claim of ten times less effort is testable in exactly one way: the day a Qualcomm target appears in <code>--print-supported-accelerators</code> or in the open tree, the size of the change that added it will say how much effort it took.</p><p>What this says about the deal is the more interesting question. Nearly everything we counted is open: the compiler, the standard library and MAX&#8217;s kernels, <strong>all 523 kernel source files carrying the same Apache 2.0 with LLVM exceptions header</strong> as the compiler. </p><p>What stays closed is the runtime that compiles and runs models, the 162.6 MB <code>libmax.so</code> and the GPU runtime beside it, shipped under the MAX license, along with whatever target collections live outside the tree. </p><p>Opening the compiler commoditizes the frontend, much as <strong>NVIDIA let its own become a commodity years ago</strong>, building NVVM on LLVM and contributing the NVPTX backend upstream. </p><p>Qualcomm is not buying a way to make kernels portable; roughly half of them are written per architecture anyway. It is buying the team, the closed runtime and <strong>the plumbing that turns its own accelerators </strong>into one more target collection, next to a kernel library that was tuned on other people&#8217;s hardware first. </p><p>The repository&#8217;s root already carries governance for that future: an AI tool policy that asks contributors to label assisted work with an <code>Assisted-by</code> trailer, to keep pull requests small because <strong>AI lowers the cost of generating code </strong>and not of reviewing it, and to keep a human in the loop, alongside <code>AGENTS.md</code> and <code>CLAUDE.md</code> files written for coding agents.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/p/how-mojo-actually-compiles?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/p/how-mojo-actually-compiles?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><div><hr></div><h2>If you build on Mojo today</h2><p>Build one executable per GPU architecture. A <code>mojo build</code> for <code>sm_90</code> embeds <code>sm_90a</code> PTX that<strong> NVIDIA&#8217;s compiler will not build for Blackwell</strong>, and the Hopper outputs and nearly all the Blackwell ones are locked the same way. The family targets such as <code>sm_100f</code> would carry a B200 build to a B300, and Mojo does not ask for them. </p><p>The <strong>path that follows the hardware is MAX&#8217;s</strong>, which compiles on the machine the model loads on. On a CI runner without a GPU, pass <code>--target-accelerator</code> with the architecture your kernels target, or public <code>tcgen05</code> code will not compile, and never build Hopper kernels on a runner configured for Blackwell. </p><p>On Jetson Orin, 1.1.0 compiles as <code>sm_80</code>; on Jetson Thor it cannot use tensor memory; on Maxwell and Pascal, point <code>MODULAR_NVPTX_COMPILER_PATH</code> at a <code>ptxas</code> old enough to know them. </p><p>If you need a guarantee that no vendor library is in the path, build with <code>-D MODULAR_DISABLE_VENDOR_FALLBACK=true</code>, the define the matmul dispatch reads to turn off the cuBLAS and rocBLAS fallback. </p><p>Executables built from a pip install carry an absolute <code>RUNPATH</code> into that Python environment. And<strong> don&#8217;t contort code to avoid Mojo&#8217;s 64-bit</strong> <code>Int</code> in index math: on our vector kernel, <code>Int32</code> index math saved 2 of 20 PTX instructions and 4 of 11 64-bit operations, <code>ptxas</code> allocated 10 registers either way, and the cubin shrank by 128 bytes.</p><p>Three of the problems above, <strong>the Orin fallthrough</strong>, the RTX 3090 PTX version and the two Thor blockers, are small patches. Since 17 September the compiler accepts outside contributions, which may be the most useful thing the open-sourcing did.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Where I might be wrong</h2><p>The <code>tcgen05</code> guard may read the build host on purpose.<strong> If MAX&#8217;s own launch path always sets the build accelerator</strong> to the device a kernel targets, the mismatch we produced can only happen when someone compiles by hand the way we did. </p><p>The consequence we measured, <code>tcgen05</code> inside Hopper PTX, stands, but its practical reach may be small.</p><p>We compiled through <code>_compile_code</code> and, to check target selection, through <code>DeviceContext.compile_function</code>, and both picked the same targets. We never launched a kernel. <strong>Launching applies rules we never exercised</strong>, such as the requirement since 1.0 that kernel arguments be fixed-width types rather than <code>Int</code>. PTX that assembles is not code that is correct or fast. </p><p>Nor did we load a <code>mojo build</code> executable on a Blackwell GPU: the refusal of its embedded PTX is the verdict of NVIDIA&#8217;s assembler, which the runtime would also have to get past, not something we watched happen.</p><p>The <strong>kernel census classifies files by path. </strong>It is a lower bound on architecture-specific code, and a path that names a vendor can still hold helpers other vendors use.</p><p>Our timings come from a one-core container and small programs. Absolute numbers will differ on real machines. <strong>We expect the ratios to hold</strong>, the LLVM share especially, but the specialization curve is a result about small bodies only.</p><p>A <strong>PyPI metadata field</strong> is not a license text. We did not read the MAX license in full, and we are not lawyers.</p><p><strong>The RTX 3090 record may be unreachable in practice</strong>. And backends absent from the open tree may exist and ship through channels we cannot see, such as MAX containers or partner builds. The claim is &#8220;<em>not in the open tree</em>&#8221;, not &#8220;<em>does not exist</em>&#8221;.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/p/how-mojo-actually-compiles?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/p/how-mojo-actually-compiles?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><div><hr></div><h2>Five dated predictions</h2><p>First: by <strong>31 March 2027</strong>, a stable Mojo release removes at least one of the two Jetson Thor blockers, either by adding <code>sm_110</code> names to the <code>tcgen05</code> guard or by compiling Thor as <code>sm_110a</code> or <code>sm_110f</code>.</p><p>Second: through <strong>31 December 2027,</strong> the default target records for the H100, B200 and B300 stay architecture-specific, with the <code>a</code> suffix. MAX&#8217;s runtime compiler has no reason to give those instructions up, and we expect <code>mojo build</code> users to be told to build per architecture rather than see the default change.</p><p>Third: by <strong>30 June 2027</strong>, a Qualcomm accelerator appears in <code>--print-supported-accelerators</code> of a public Mojo release or in the open repository.</p><p>Fourth: through <strong>31 December 2027</strong>, every shipped configuration of Mojo and MAX that runs on NVIDIA hardware still finishes code with NVIDIA&#8217;s closed PTX compiler, whether bundled, invoked through the driver&#8217;s JIT, or as an external <code>ptxas</code>.</p><p>Fifth: <code>fn</code> does not return in any 1.x release.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Confidence dossier</h2><p>Every claim that carries weight in the piece, graded. M means we measured it and <code>verify.py</code> re-checks it; A means we read it in a primary source, usually the code at a named revision; B means it was reported or claimed and we did not reproduce it; C is our inference. 42 rows are M, 1 reuse our own earlier measurements, 15 are A, 4 are B and 6 are C.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!XY1s!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a770ab8-259a-4920-923b-61d108b257bc_2400x6832.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!XY1s!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a770ab8-259a-4920-923b-61d108b257bc_2400x6832.png 424w, https://substackcdn.com/image/fetch/$s_!XY1s!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a770ab8-259a-4920-923b-61d108b257bc_2400x6832.png 848w, https://substackcdn.com/image/fetch/$s_!XY1s!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a770ab8-259a-4920-923b-61d108b257bc_2400x6832.png 1272w, https://substackcdn.com/image/fetch/$s_!XY1s!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a770ab8-259a-4920-923b-61d108b257bc_2400x6832.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!XY1s!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a770ab8-259a-4920-923b-61d108b257bc_2400x6832.png" width="1456" height="4145" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6a770ab8-259a-4920-923b-61d108b257bc_2400x6832.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:4145,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:3271485,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/217227968?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a770ab8-259a-4920-923b-61d108b257bc_2400x6832.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!XY1s!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a770ab8-259a-4920-923b-61d108b257bc_2400x6832.png 424w, https://substackcdn.com/image/fetch/$s_!XY1s!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a770ab8-259a-4920-923b-61d108b257bc_2400x6832.png 848w, https://substackcdn.com/image/fetch/$s_!XY1s!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a770ab8-259a-4920-923b-61d108b257bc_2400x6832.png 1272w, https://substackcdn.com/image/fetch/$s_!XY1s!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a770ab8-259a-4920-923b-61d108b257bc_2400x6832.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/p/how-mojo-actually-compiles?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/p/how-mojo-actually-compiles?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><div><hr></div><h2>Reproduce it</h2><p>Everything above ran in a<strong> one-core x86-64 container </strong>with 3 GB of memory and no GPU, on 21 and 22 September 2026. The probe scripts and the harness are published with this piece as <code>mojo-compile-probe</code>, MIT licensed. <code>verify.py</code> holds 187 checks and exits non-zero at the first number that no longer matches.</p><pre><code><code>pip install mojo==1.1.0 max-core==26.6.0      # Mojo 1.1.0 (8189361e), MAX 26.6.0
pip download nvidia-cuda-nvcc==13.4.92 --no-deps  # ptxas V13.4.92, unzip and use bin/ptxas
git clone --depth 1 https://github.com/modular/modular.git   # measured at 26cfe94
git -C modular fetch --depth 1 origin tag mojo/v1.1.0        # c6fa49f, what shipped

python3 m2_target_sweep.py      # 30 targets, emitted .target and PTX version
python3 m3_ptxas_roundtrip.py   # every NVIDIA output through ptxas, forward-compat tests
python3 m4_tensor_cores.py      # mma, wgmma, tcgen05 per target and build flag
python3 m5_language.py          # the language probes
python3 m6_compile_time.py      # specialization scaling and timing census
mojo build aot.mojo --target-accelerator sm_90 -o aot_sm_90   # then extract the PTX and run ptxas
mojo build m7_launch_path.mojo --target-accelerator sm_90 --emit asm   # launch-path check
python3 verify.py --live        # re-checks every number quoted; exits non-zero on failure</code></code></pre><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Corrections before publication</h2><p>Writing the checks before the prose caught these, in the order they happened. We list them because the method only works if its failures are visible.</p><ol><li><p><em>While probing imports, a shell redirection error made it look as if wildcard imports of missing modules were silently accepted. They are not; the compiler reports them correctly. Caught by re-running the probe with separate output files.</em></p></li><li><p><em>The family targets sm_100f through sm_121f were first attributed to Mojo&#8217;s LLVM NVPTX backend. The strings came from libNVPTX.so, which is NVIDIA&#8217;s nvptxcompiler, not LLVM.</em></p></li><li><p><em>The harness caught a conflation in the symbol count: libNVPTX.so exports 45,501 dynamic symbols, of which 45,474 carry the libnvptxcompiler_static_ prefix.</em></p></li><li><p><em>A chart subtitle said MAX&#8217;s code was five times the compiler. The measured ratio for MAX&#8217;s Mojo code is 3.3.</em></p></li><li><p><em>The Jetson Orin downgrade was first diagnosed against main, whose table is structured differently. The diagnosis in the text is against the v1.1.0 tag that actually shipped.</em></p></li><li><p><em>A draft said the RTX 3090 record was the only one named with a full marketing string. The GTX 1080 Ti, GTX 1060, GTX 970 and Tesla P100 records are too.</em></p></li><li><p><em>The first specialization-scaling run used one measurement per point and called the curve flat. Three runs per point show a small but real slope, about 0.56 ms per specialization; the text and Figure 4 use the repeated data.</em></p></li><li><p><em>Version 2 said the suffix costs nothing because Mojo compiles at load time and never ships PTX. That holds for MAX, whose runtime library carries the compiler. It is false for mojo build executables, which embed build-time PTX and no compiler; the targets section now separates the two.</em></p></li><li><p><em>Tuning-config counts included the struct definitions: the call sites are 188 and 148, not 189 and 149.</em></p></li><li><p><em>The pass count mixed 2 passes from the shared Support library into the Mojo compiler&#8217;s 45.</em></p></li><li><p><em>Hexagon also appears in a documentation page&#8217;s sample target list, not only in a comment and a test fixture.</em></p></li><li><p><em>Version 2 explained the compile floor as the standard library&#8217;s cost without measuring it; the empty-program census now does.</em></p></li><li><p><em>Final read-through: the opening said sm_90a unlocks the tensor memory accelerator. ptxas 13.4 accepts TMA copies on plain sm_90; the suffix is what wgmma and setmaxnreg need, and the text now names those.</em></p></li><li><p><em>Final read-through: version 3 said 1.1 deleted two keywords. The release notes list three keywords (fn, alias, __comptime_assert) and two syntax forms.</em></p></li><li><p><em>Final read-through: version 3 contrasted an open language with a kernel library it implied was not. All 523 kernel source files carry the same Apache 2.0 header; the closed part is the runtime, and the Qualcomm section now says so.</em></p></li><li><p><em>Final read-through: a sentence left over from version 2 still called the specialization curve flat.</em></p></li><li><p><em>Two framings from our July piece on MLIR are corrected in the text: KGEN is a dialect inside Mojo rather than a layer above it, and MAX&#8217;s matmul path keeps a vendor BLAS fallback.</em></p></li></ol><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Sources</h2><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!b_id!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0defe443-54d8-4a27-87de-d5c598665e3e_2400x2000.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!b_id!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0defe443-54d8-4a27-87de-d5c598665e3e_2400x2000.png 424w, https://substackcdn.com/image/fetch/$s_!b_id!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0defe443-54d8-4a27-87de-d5c598665e3e_2400x2000.png 848w, https://substackcdn.com/image/fetch/$s_!b_id!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0defe443-54d8-4a27-87de-d5c598665e3e_2400x2000.png 1272w, https://substackcdn.com/image/fetch/$s_!b_id!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0defe443-54d8-4a27-87de-d5c598665e3e_2400x2000.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!b_id!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0defe443-54d8-4a27-87de-d5c598665e3e_2400x2000.png" width="1456" height="1213" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0defe443-54d8-4a27-87de-d5c598665e3e_2400x2000.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1213,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:996736,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/217227968?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0defe443-54d8-4a27-87de-d5c598665e3e_2400x2000.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!b_id!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0defe443-54d8-4a27-87de-d5c598665e3e_2400x2000.png 424w, https://substackcdn.com/image/fetch/$s_!b_id!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0defe443-54d8-4a27-87de-d5c598665e3e_2400x2000.png 848w, https://substackcdn.com/image/fetch/$s_!b_id!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0defe443-54d8-4a27-87de-d5c598665e3e_2400x2000.png 1272w, https://substackcdn.com/image/fetch/$s_!b_id!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0defe443-54d8-4a27-87de-d5c598665e3e_2400x2000.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Measurements, analysis and conclusions are ours. Source excerpts from modular/modular are <strong>Apache 2.0 with LLVM exceptions</strong>; compiler and assembler messages are output we generated. Charts and harness: mojo-compile-probe.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/p/how-mojo-actually-compiles?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/p/how-mojo-actually-compiles?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><p></p>]]></content:encoded></item><item><title><![CDATA[How NVIDIA's Data Center Business Actually Works in 2026]]></title><description><![CDATA[NVIDIA booked $89.0 billion of data center revenue in its latest quarter. The company now reports about $530 billion of commitments and maximum guarantee exposure, against $91 billion of liabilities.]]></description><link>https://www.thesoftwarefrontier.com/p/how-nvidias-data-center-business</link><guid isPermaLink="false">https://www.thesoftwarefrontier.com/p/how-nvidias-data-center-business</guid><dc:creator><![CDATA[Lorenzo Bradanini]]></dc:creator><pubDate>Sun, 20 Sep 2026 17:29:24 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!FM_D!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf5493ce-2fda-48e5-9b07-58c83ba79c94_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!FM_D!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf5493ce-2fda-48e5-9b07-58c83ba79c94_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!FM_D!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf5493ce-2fda-48e5-9b07-58c83ba79c94_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!FM_D!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf5493ce-2fda-48e5-9b07-58c83ba79c94_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!FM_D!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf5493ce-2fda-48e5-9b07-58c83ba79c94_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!FM_D!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf5493ce-2fda-48e5-9b07-58c83ba79c94_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!FM_D!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf5493ce-2fda-48e5-9b07-58c83ba79c94_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/cf5493ce-2fda-48e5-9b07-58c83ba79c94_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2914693,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/216005940?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf5493ce-2fda-48e5-9b07-58c83ba79c94_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!FM_D!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf5493ce-2fda-48e5-9b07-58c83ba79c94_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!FM_D!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf5493ce-2fda-48e5-9b07-58c83ba79c94_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!FM_D!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf5493ce-2fda-48e5-9b07-58c83ba79c94_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!FM_D!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf5493ce-2fda-48e5-9b07-58c83ba79c94_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>A very short version </h2><p><strong>NVIDIA&#8217;s data center revenue</strong> more than doubled in a year, to $89.0 billion a quarter. The company now sells whole racks, supplies reference designs for the buildings around them, and increasingly backs the financing as well. </p><p>The <strong>Vera Rubin rack</strong> started shipping in August, and the version of its product page updated on September 11 lists lower memory and NVLink bandwidth than the August version did, without explanation, which moves the batch size at which these machines stop being limited by memory. </p><p>Power, cooling and memory are the constraints that matter now, and the balance sheet shows it: <strong>$279 billion of supply commitments</strong>, mostly memory, up to $105 billion of guarantees behind one Ohio campus, and operating cash flow that fell to 40% of net income, partly because some customers were given longer to pay.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>The real introduction</h2><p>We&#8217;ve alredy explored the fact that <strong>NVIDIA&#8217;s second quarter of fiscal 2027 closed on July 26</strong>, 2026, and the company reported it on August 26. On it we saw that revenue was $96.2 billion, up 106% from a year earlier, and $89.0 billion of it came from the data center.</p><p>But as many of you know, those are just the mere &#8220;<em>headline numbers</em>&#8221;. The part of the filing that says more about where the data center business is going towards sits further down the CFO commentary, in a <strong>section about commitments and guarantees.</strong></p><p>That small section lists <strong>$279 billion of supply and capacity commitments, </strong>up from $119 billion just three months earlier, an increase NVIDIA itself  attributes mainly to memory. </p><p>It lists<strong> $36 billion under agreements with AI clouds</strong>, a structure the CFO described on the call as a take-or-pay commitment on part of a facility&#8217;s capacity in exchange for a share of its revenue, and $20 billion of data center leases that NVIDIA signed and expects to reassign to third parties.</p><p>What sounds obvious is that neither of those two lines appears in the first quarter&#8217;s commentary. And it lists guarantees with a <strong>maximum gross exposure of $108.5 billion</strong>, of which $105 billion backs the land, power and shell for 4.25 gigawatts at a campus in Ohio that OpenAI will lease for 20 years.</p><p>Read next to the product news since the end of May, those tables describe a company whose unit of sale has <em>moved from the chip to the rack and now to the gigawatt.</em> At that scale the scarce inputs are memory, power, buildings and credit, and <strong>NVIDIA is now putting its own balance sheet behind all four.</strong></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>The quarter</h2><p>In the filing&#8217;s own terms, second quarter revenue grew 18% sequentially, and Data Center grew 18% sequentially and 117% year over year. That put Data Center at 92.5% of the total, against 92.2% in the first quarter. </p><p>Everything else now sits in a platform called<strong> Edge Computing</strong>: $7.2 billion in the quarter, up 13% sequentially and 27% from a year ago. NVIDIA adopted this two-platform framework in the first quarter. </p><p>Data Center is split into <strong>Hyperscale and ACIE</strong>, and Edge Computing covers PCs, game consoles, workstations, AI-RAN base stations, robotics and automotive. <em>GAAP gross margin was 75.0%,</em> GAAP operating income $63.7 billion and GAAP net income $59.7 billion, which includes $7.8 billion of pre-tax net gains on equity securities. </p><p>The CFO pointed out on the call that <strong>growth accelerated for the fourth quarter in a row. </strong><em>Year over year, revenue grew 56% in the second quarter of fiscal 2026,<sup> </sup>then 62%, 73%, 85% and now 106%.</em> The third quarter guide of $108.0 billion, plus or minus 2%, means 89% growth at the midpoint and 93% at the top of the range, so extending the streak would take a result above $117.3 billion, 8.7% over the midpoint. </p><p><strong>NVIDIA beat its own midpoint by 4.6%</strong> in the first quarter and by 5.7% in the second. The comparison base is also harder: revenue in the third quarter of fiscal 2026 was 22% above the quarter before it. </p><p>For fiscal 2028, which roughly corresponds to calendar 2027, NVIDIA gave a <strong>preliminary outlook of about 70% revenue growth</strong> and described it as supply-constrained. Huang said the company had never guided a full year ahead before. </p><p>The CFO said customer forecasts point to growth doubling, and that supply should remain a bottleneck at least through the end of fiscal 2028. </p><p>The third quarter outlook assumes no Data Center compute revenue from China. <strong>Shipments of Hopper products to China were below 1%</strong> of Data Center revenue in the quarter; in the first quarter of fiscal 2026, before the license requirement, NVIDIA had sold $4.6 billion of H20. </p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!kwox!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F13d8c17a-9cc2-489e-9a11-c9b0e138e68f_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!kwox!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F13d8c17a-9cc2-489e-9a11-c9b0e138e68f_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!kwox!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F13d8c17a-9cc2-489e-9a11-c9b0e138e68f_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!kwox!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F13d8c17a-9cc2-489e-9a11-c9b0e138e68f_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!kwox!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F13d8c17a-9cc2-489e-9a11-c9b0e138e68f_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!kwox!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F13d8c17a-9cc2-489e-9a11-c9b0e138e68f_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/13d8c17a-9cc2-489e-9a11-c9b0e138e68f_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Bar chart of NVIDIA revenue by fiscal quarter from Q3 FY25 to Q2 FY27 with the Data Center portion, the Q3 FY27 guide of $108 billion, and a line of year over year growth ranging from 56% to 106%&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Bar chart of NVIDIA revenue by fiscal quarter from Q3 FY25 to Q2 FY27 with the Data Center portion, the Q3 FY27 guide of $108 billion, and a line of year over year growth ranging from 56% to 106%" title="Bar chart of NVIDIA revenue by fiscal quarter from Q3 FY25 to Q2 FY27 with the Data Center portion, the Q3 FY27 guide of $108 billion, and a line of year over year growth ranging from 56% to 106%" srcset="https://substackcdn.com/image/fetch/$s_!kwox!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F13d8c17a-9cc2-489e-9a11-c9b0e138e68f_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!kwox!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F13d8c17a-9cc2-489e-9a11-c9b0e138e68f_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!kwox!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F13d8c17a-9cc2-489e-9a11-c9b0e138e68f_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!kwox!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F13d8c17a-9cc2-489e-9a11-c9b0e138e68f_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 1. Revenue by fiscal quarter with Data Center labeled, and year over year growth. The hatched bar and the hollow marker are the third quarter guide.</figcaption></figure></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Who is buying</h2><p>Hyperscale, which NVIDIA defines as the public clouds and the largest consumer internet companies, <strong>brought in $48.7 billion, up 13% sequentially and 102% from a year ago. </strong>ACIE, short for AI clouds, industrial and enterprise, brought in $40.3 billion, up 25% sequentially and 138% from a year ago. </p><p>On the call the CFO tied ACIE growth to capacity that neoclouds added for enterprises, AI startups and sovereign customers, and to hyperscalers buying capacity from AI clouds to supplement their own buildouts. </p><p><strong>NVIDIA expects its neocloud partners to finish 2026 with 8 gigawatts installed</strong>, against roughly 3 gigawatts at the end of 2025, and said sovereign revenue, mostly booked through regional neoclouds, grew 35% sequentially and more than tripled year over year. </p><p>The <strong>second quarter commentary</strong> also contains a line that matters for anyone tracking this split: a company was moved from ACIE to Hyperscale because its business model changed, and prior periods were recast. </p><p>The size of the move can be read off the two commentaries. The first quarter was originally reported as $37.87 billion of Hyperscale and $37.38 billion of ACIE, so ACIE was 49.7% of Data Center. </p><p>In the recast <strong>the same quarter shows $43.05 billion and $32.20 billion</strong>, and ACIE drops to 42.8%. The difference between the two versions, $5.18 billion, is the first quarter revenue of that one customer, which NVIDIA does not name, and 6.9% of the whole Data Center line. </p><p>On the new basis, ACIE was 45.3% of Data Center in the second quarter. When NVIDIA describes the <strong>non-hyperscale business as roughly half of Data Center</strong>, as both executives did on the call, it helps to remember that a single classification decision moved that share by about seven points. </p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!InkI!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F498f519c-9be7-4a41-9e41-3622f3bf3061_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!InkI!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F498f519c-9be7-4a41-9e41-3622f3bf3061_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!InkI!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F498f519c-9be7-4a41-9e41-3622f3bf3061_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!InkI!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F498f519c-9be7-4a41-9e41-3622f3bf3061_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!InkI!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F498f519c-9be7-4a41-9e41-3622f3bf3061_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!InkI!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F498f519c-9be7-4a41-9e41-3622f3bf3061_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/498f519c-9be7-4a41-9e41-3622f3bf3061_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Stacked bars of Hyperscale and ACIE revenue for Q1 FY27 as first reported, Q1 FY27 recast, and Q2 FY27&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Stacked bars of Hyperscale and ACIE revenue for Q1 FY27 as first reported, Q1 FY27 recast, and Q2 FY27" title="Stacked bars of Hyperscale and ACIE revenue for Q1 FY27 as first reported, Q1 FY27 recast, and Q2 FY27" srcset="https://substackcdn.com/image/fetch/$s_!InkI!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F498f519c-9be7-4a41-9e41-3622f3bf3061_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!InkI!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F498f519c-9be7-4a41-9e41-3622f3bf3061_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!InkI!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F498f519c-9be7-4a41-9e41-3622f3bf3061_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!InkI!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F498f519c-9be7-4a41-9e41-3622f3bf3061_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 2. The same quarter before and after the recast, next to the latest quarter.</figcaption></figure></div><p><strong>Two other statements</strong> from the call frame the customer mix. The CFO said the AI labs whose buildouts NVIDIA expects to support with its balance sheet should account for roughly a quarter of NVIDIA&#8217;s business next year. </p><p>On the hyperscale side, <strong>NVIDIA put cloud industry backlog above $2 trillion</strong> and capital expenditure by the five largest hyperscalers at nearly $800 billion in 2026 and $1.3 trillion in 2027. Those are NVIDIA&#8217;s figures, not a sum rebuilt here from each company&#8217;s guidance. </p><p>The same call announced that AWS will deploy an additional 2 million GPUs between this quarter and the second quarter of fiscal 2029. </p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Timeline</h2><p>The data center news of the last few months, in order, with the sources used for each entry:</p><ul><li><p><strong>Jan 5. </strong>CES keynote: Vera Rubin in full production; Huang says 45&#176;C water needs no chillers. </p></li><li><p><strong>May 18.</strong> Quarterly dividend raised from $0.01 to $0.25 a share; $80.0 billion added to the buyback authorization. </p></li><li><p><strong>May 31.</strong> Vera Rubin ramps into full production; DSX announced; Spectrum-X Ethernet Photonics in production. </p></li><li><p><strong>Early July.</strong> SemiAnalysis reports a Kyber delay to 2028; NVIDIA says its roadmap is intact. </p></li><li><p><strong>Aug 10.</strong> Financing platforms with six asset managers and banks, memorandums of understanding for more than $500 billion.</p></li><li><p><strong>Aug 17. </strong>Guarantees for SB Energy&#8217;s PORTS-Pike campus; 8-K filed. </p></li><li><p><strong>Aug 26. </strong>Second quarter results: $96.2 billion of revenue; Vera Rubin production shipments under way; Groq 3 LPX in full production. </p></li><li><p><strong>Sep 9.</strong> Australia: up to 2 GW with eight partners by 2027.</p></li><li><p><strong>Sep 10. </strong>Huang at a Goldman Sachs conference repeats $3 trillion to $4 trillion by 2030.</p></li><li><p><strong>Sep 11. </strong>Vera Rubin NVL72 product page updated; memory and NVLink bandwidth lower than in the August version.</p></li><li><p><strong>Oct 21. </strong>GTC Berlin keynote.</p></li><li><p><strong>Nov 17. </strong>Third quarter results. </p></li></ul><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share The Software Frontier&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share The Software Frontier</span></a></p><div><hr></div><h2>The rack</h2><p><strong>Production shipments of Vera Rubin</strong> started precisely in August. NVIDIA says it already holds purchase orders from every major hyperscaler, AI cloud and system OEM, and it expects Vera Rubin to be about 20% of Data Center revenue in the third quarter. </p><p><strong>If Data Center keeps the 92.5% share of revenue </strong>it had in the second quarter, it would be about $100 billion of the $108.0 billion guide, and 20% of that is <em>roughly $20 billion of Vera Rubin</em> in a single quarter.</p><p>NVIDIA&#8217;s specification table for the Vera Rubin NVL72 rack gives the shape of the machine; the figures here are from the version of the page updated on September 11.</p><p>The rack holds <strong>72 Rubin GPUs and 36 Vera CPUs.</strong> Each GPU carries 288 GB of HBM4 with 19.2 TB/s of memory bandwidth, and the rack is listed at 20.7 TB and 1,400 TB/s. NVLink 6 gives every GPU 3 TB/s of scale-up bandwidth, 216 TB/s for the rack, and NVLink-C2C links each Vera to its GPUs at 1.8 TB/s. </p><p><strong>Each Vera has 88 custom Olympus cores and 176 threads</strong>, 3,168 cores and 6,336 threads in the rack, with up to 1.5 TB of LPDDR5X per CPU and up to 54 TB per rack. Scale-out networking is listed at 0.45 TB/s per GPU and 32.4 TB/s per rack, both bidirectional, and the inlet temperature at 45&#176;C. </p><p><strong>NVIDIA&#8217;s own technical blog</strong>, which I tried to read carefully before writing all of this post, adds that HBM4 doubles the interface width of HBM3e and that Rubin nearly triples memory bandwidth over Blackwell.</p><p>The same table rates the rack at 3,600 PFLOPS for NVFP4 inference and 288 PFLOPS at BF16, a factor of 12.5 on identical silicon. Its footnotes also mark the <strong>NVFP4 inference row</strong> as a sparse specification and the training, BF16 and TF32 rows as dense ones, so the headline number is not only in the most compressed format in the table but also on a different basis from the rows below it. </p><p>A comparison between generations only means something when both sides use the same precision and the same basis.</p><p>A second ratio from the table: <strong>each GPU&#8217;s NVLink bandwidth is 16% of its memory bandwidth, 3 TB/s against 19.2 TB/s. </strong>Work that has to cross GPUs inside the rack, such as tensor parallel collectives or the all-to-all traffic of mixture-of-experts models, moves over about a sixth of the bandwidth each GPU has for reading its own weights and KV cache.</p><p>NVIDIA&#8217;s performance claims for Vera Rubin use at least three baselines. The product page says the <strong>NVL72 delivers up to 10x more tokens per megawatt and one tenth the cost per million tokens</strong> compared with GB200 NVL72, both measured on Kimi-K2-Thinking with 32K input and 8K output tokens, and that it trains a 10 trillion parameter mixture-of-experts model on 100 trillion tokens in a month with a quarter of the GPUs, in NVIDIA&#8217;s projection. </p><p>The <em>August call claimed 30x higher throughput per megawatt and 35x lower token cost</em> against Grace Blackwell Ultra, the GB300 generation that followed GB200. The May release claimed 10x agent throughput at scale against Grace Blackwell. </p><p>If GB300 is at least as efficient per megawatt as GB200, a<strong> 30x gain over it and a gain of at most 10x over GB200 </strong>cannot both hold on the same workload and configuration; the product page itself attaches a 35x per-megawatt figure to Vera Rubin paired with LPX. All of these remain vendor figures until independent benchmarks run on shipping racks.</p><p>Around the NVL72, NVIDIA now sells what it calls a POD of five rack types working as one system: the NVL72 itself, Vera CPU racks, Groq 3 LPX, Vera BlueField-4 STX storage and Spectrum-6 SPX Ethernet. </p><p><strong>Groq 3 LPX, NVIDIA&#8217;s first rack-scale LPU system, is in full production</strong>, with volume shipments to early adopters expected later this quarter and Nebius first in line. </p><p>Jensen Huang described it as an <strong>add-on for services that will pay for very high interactivity</strong>, with lower throughput and a higher cost per token, and said most data centers will simply run Vera Rubin NVL72. </p><p>The Vera CPU is also sold on its own. Grace revenue passed $5 billion over the trailing twelve months, <em>Vera is already shipping to lead partners</em> with AWS starting this quarter, and NVIDIA expects CPU revenue to more than double in fiscal 2028. </p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>The revised table</h2><p>The<strong> table above differs from the one quoted before September</strong>, and several write-ups published in the last month still carry the older numbers.</p><p>Spheron, which retrieved NVIDIA&#8217;s table on August 16, recorded 22 TB/s of HBM4 bandwidth per GPU, 1,580 TB/s per rack and 260 TB/s of NVLink 6, and NVIDIA&#8217;s own January technical blog gives NVLink 6 as 3.6 TB/s per GPU. </p><p>The <strong>page NVIDIA updated on September 11 </strong>lists 19.2 TB/s, 1,400 TB/s and 216 TB/s instead. The rack&#8217;s scale-out figure moved the other way, from the 28.8 TB/s reported in July to 32.4 TB/s. The page unfortunately does not say why. </p><p>The changes are not a <em>single unit conversion</em>, because the ratios differ, and the compute ratings and memory capacities match the January figures. </p><p>The new numbers can be checked against the rest of the page. The 100 MW factory table only reproduces with them: <strong>40,000 GPUs at 19.2 TB/s is 768 PB/s,</strong> which the page rounds to 800, while 22 TB/s would give 880. </p><p>The <em>rack total no longer equals 72 times the per-GPU figure</em> exactly: that product is 1,382.4 TB/s against the 1,400 listed, a 1.3% rounding. Anything calculated from the old bandwidth, including the critical batch below, has to be redone.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Ratings by precision</h2><p>The ratings in the table span three orders of magnitude on the same silicon. Per GPU, <strong>NVIDIA lists 50 PFLOPS for sparse NVFP4 inference</strong>, 35 for dense NVFP4 training, 17.5 for FP8 or FP6 training, 4 at FP16 or BF16 and 2 at TF32, then 130 TFLOPS at FP32 and 33 TFLOPS at FP64.</p><p>The top of that list is 1,515 times the bottom. The steps between rows are simple ratios: <strong>dense NVFP4 is exactly twice FP8</strong>, FP8 is 4.375 times BF16, BF16 is twice TF32, and TF32 is 15.4 times the plain FP32 rating. The sparse inference row sits 1.43 times above dense NVFP4 training.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!RGoP!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fef572899-d86a-4141-acff-19ad1e8909b0_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!RGoP!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fef572899-d86a-4141-acff-19ad1e8909b0_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!RGoP!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fef572899-d86a-4141-acff-19ad1e8909b0_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!RGoP!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fef572899-d86a-4141-acff-19ad1e8909b0_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!RGoP!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fef572899-d86a-4141-acff-19ad1e8909b0_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!RGoP!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fef572899-d86a-4141-acff-19ad1e8909b0_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ef572899-d86a-4141-acff-19ad1e8909b0_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Bar chart on a log scale of the seven per-GPU ratings in NVIDIA's Vera Rubin table, from 50 PFLOPS for sparse NVFP4 inference down to 33 TFLOPS at FP64&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Bar chart on a log scale of the seven per-GPU ratings in NVIDIA's Vera Rubin table, from 50 PFLOPS for sparse NVFP4 inference down to 33 TFLOPS at FP64" title="Bar chart on a log scale of the seven per-GPU ratings in NVIDIA's Vera Rubin table, from 50 PFLOPS for sparse NVFP4 inference down to 33 TFLOPS at FP64" srcset="https://substackcdn.com/image/fetch/$s_!RGoP!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fef572899-d86a-4141-acff-19ad1e8909b0_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!RGoP!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fef572899-d86a-4141-acff-19ad1e8909b0_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!RGoP!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fef572899-d86a-4141-acff-19ad1e8909b0_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!RGoP!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fef572899-d86a-4141-acff-19ad1e8909b0_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 3. The seven ratings NVIDIA lists for one Rubin GPU, marked by the basis the page footnotes give.</figcaption></figure></div><p>Each row answers a different question. The training rows describe dense matrix math, the <strong>sparse inference row</strong> folds in whatever sparsity NVIDIA assumed, and the FP64 row is the one that matters for classical simulation. </p><p>Ratios between a sparse row and a dense one, or between generations rated on different bases, look precise and mean little.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>The critical batch</h2><p>A decode step reads every active weight once and does roughly two floating point operations per active parameter for every sequence in the batch. With N active parameters stored at b bytes each, a batch of B needs N &#215; b bytes of memory traffic and <strong>2 &#215; N &#215; B floating point operations per step. </strong></p><p>The step is limited by memory bandwidth while N &#215; b divided by the bandwidth is larger than 2 &#215; N &#215; B divided by the compute rating P, and by compute beyond that. N cancels, which leaves a property of the hardware alone:</p><p style="text-align: center;"><code>B* = P &#215; b / (2 &#215; B_mem)</code></p><p><strong>B* is the batch per model replica</strong> above which adding sequences stops being free. It counts weights only, at zero context. Longer contexts add KV cache reads for every sequence, which on their own push the real crossover higher, and at long enough contexts no batch makes the step compute-bound at all; attention compute, which also grows with context, pulls the other way. </p><p>B* also only makes sense with <strong>dense ratings, </strong>because the sparse row already assumes less work per parameter than the formula does. <em>The model behind this post enforces that</em>: each rating carries a basis field, and the function refuses a sparse one.</p><p>With the September 11 bandwidth of 19.2 TB/s, the dense rows give <strong>B* of 456 for NVFP4 and FP8 training and 208 for BF16 and TF32. </strong>The pairs match exactly because each rating scales with the inverse of its bytes per parameter: 35 &#215; 0.5 and 17.5 &#215; 1 give the same product, and so do 4 &#215; 2 and 2 &#215; 4. </p><p>With the old 22 TB/s, the same rows gave 398 and 182, so the revision moves the crossover 14.6% higher: a <strong>Rubin GPU stays memory-bound </strong>at larger batches than the earlier table implied. Plugging in the sparse inference row anyway would give 651, a number that describes nothing.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!28sr!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F48b7d1b3-d158-4905-af16-28416ddd5bf4_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!28sr!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F48b7d1b3-d158-4905-af16-28416ddd5bf4_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!28sr!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F48b7d1b3-d158-4905-af16-28416ddd5bf4_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!28sr!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F48b7d1b3-d158-4905-af16-28416ddd5bf4_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!28sr!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F48b7d1b3-d158-4905-af16-28416ddd5bf4_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!28sr!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F48b7d1b3-d158-4905-af16-28416ddd5bf4_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/48b7d1b3-d158-4905-af16-28416ddd5bf4_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Grouped bars of the critical batch B* for four dense ratings before and after the September 11 revision, 456 and 208 after, 398 and 182 before, with the sparse row shown as refused&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Grouped bars of the critical batch B* for four dense ratings before and after the September 11 revision, 456 and 208 after, 398 and 182 before, with the sparse row shown as refused" title="Grouped bars of the critical batch B* for four dense ratings before and after the September 11 revision, 456 and 208 after, 398 and 182 before, with the sparse row shown as refused" srcset="https://substackcdn.com/image/fetch/$s_!28sr!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F48b7d1b3-d158-4905-af16-28416ddd5bf4_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!28sr!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F48b7d1b3-d158-4905-af16-28416ddd5bf4_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!28sr!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F48b7d1b3-d158-4905-af16-28416ddd5bf4_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!28sr!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F48b7d1b3-d158-4905-af16-28416ddd5bf4_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 4. The weights-only critical batch for each dense rating, before and after the bandwidth revision.</figcaption></figure></div><h2>The memory ladder</h2><p>Bandwidth falls sharply at each step away from the GPU&#8217;s own memory. Per GPU, <strong>NVIDIA lists 19.2 TB/s for HBM4</strong>, 3 TB/s for NVLink 6 and 0.45 TB/s for scale-out networking; each Vera connects at 1.8 TB/s over NVLink-C2C; and BlueField-4 runs at up to 800 Gb/s, which is 0.1 TB/s NVIDIA&#8217;s Rubin platform page gives Vera&#8217;s own LPDDR5X band.</p><p>width as up to 1.2 TB/s, the <strong>figure VideoCardz also reported from CES</strong>. On the figures as listed, local HBM4 is 6.4 times NVLink, 10.7 times NVLink-C2C, 16 times Vera&#8217;s memory and 42.7 times the scale-out figure.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!355B!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F14ea6b3b-79e9-420a-8f91-e5ae43843855_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!355B!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F14ea6b3b-79e9-420a-8f91-e5ae43843855_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!355B!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F14ea6b3b-79e9-420a-8f91-e5ae43843855_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!355B!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F14ea6b3b-79e9-420a-8f91-e5ae43843855_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!355B!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F14ea6b3b-79e9-420a-8f91-e5ae43843855_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!355B!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F14ea6b3b-79e9-420a-8f91-e5ae43843855_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/14ea6b3b-79e9-420a-8f91-e5ae43843855_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Horizontal bars on a log scale: HBM4 19.2 TB/s, NVLink 6 3 TB/s, NVLink-C2C 1.8 TB/s, Vera LPDDR5X 1.2 TB/s, scale-out 0.45 TB/s, BlueField-4 0.1 TB/s&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Horizontal bars on a log scale: HBM4 19.2 TB/s, NVLink 6 3 TB/s, NVLink-C2C 1.8 TB/s, Vera LPDDR5X 1.2 TB/s, scale-out 0.45 TB/s, BlueField-4 0.1 TB/s" title="Horizontal bars on a log scale: HBM4 19.2 TB/s, NVLink 6 3 TB/s, NVLink-C2C 1.8 TB/s, Vera LPDDR5X 1.2 TB/s, scale-out 0.45 TB/s, BlueField-4 0.1 TB/s" srcset="https://substackcdn.com/image/fetch/$s_!355B!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F14ea6b3b-79e9-420a-8f91-e5ae43843855_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!355B!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F14ea6b3b-79e9-420a-8f91-e5ae43843855_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!355B!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F14ea6b3b-79e9-420a-8f91-e5ae43843855_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!355B!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F14ea6b3b-79e9-420a-8f91-e5ae43843855_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 5. Bandwidth at each step away from a Rubin GPU&#8217;s own memory, as NVIDIA lists it.</figcaption></figure></div><p><strong>Those ratios turn into time</strong> when data has to move. <em>Assuming each interconnect figure counts both directions</em>, as the product page states for scale-out and NVIDIA&#8217;s NVLink page states for NVLink, one direction carries half. </p><p>On that basis, and <strong>taking BlueField-4&#8217;s 800 Gb/s as a single direction</strong>, 100 GB takes about 67 milliseconds over NVLink 6, 111 milliseconds over NVLink-C2C, 444 milliseconds over the scale-out network and 1 second through BlueField-4, while a GPU reads 100 GB from its own HBM4 in about 5.2 milliseconds.</p><p>The rack holds 20.7 TB of HBM4 and up to 54 TB of LPDDR5X behind NVLink-C2C, 74.7 TB in all, and <strong>NVIDIA&#8217;s 100 MW factory table counts 12 PB of HBM4</strong> and up to 30 PB of LPDDR5X across 40,000 GPUs, which it calls 42 PB of fast memory. </p><p>Whatever does not fit in that tier, KV cache for long agent sessions included, sits behind links that are several times slower.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/p/how-nvidias-data-center-business?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/p/how-nvidias-data-center-business?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><div><hr></div><h2>SRAM next to HBM</h2><p>The <strong>Groq 3 LPX rack makes the opposite trade. </strong>NVIDIA&#8217;s page lists 256 LPUs per rack, each with 500 MB of SRAM, 150 TB/s of SRAM bandwidth and 2.5 TB/s of scale-up bandwidth, which gives 128 GB of SRAM, about 40 PB/s of SRAM bandwidth and 640 TB/s of chip-to-chip bandwidth per rack, plus 12 TB of DDR5 for larger models.</p><p>Against a Vera Rubin NVL72, an <em>LPX rack has about 162 times less accelerator memory</em>, SRAM against HBM4, and about 28 times more accelerator memory bandwidth.</p><p>A compact way to compare the two is the sweep rate, bandwidth divided by capacity: how many times per second a chip could read its entire memory. <strong>A Rubin GPU sweeps its 288 GB about 67 times a second.</strong> A Groq 3 LPU sweeps its 500 MB 300,000 times a second, 4,500 times as often. </p><p>If a chip&#8217;s memory were full of weights that every token has to read, the sweep rate would be its ceiling on tokens per second at batch one. </p><p>The limit is capacity: at 500 MB per chip, a trillion-parameter model cannot live in SRAM alone, which is <strong>why NVIDIA positions LPX as an accelerator for Vera Rubin</strong> and says the GPUs and LPUs jointly compute every layer for every output token.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!GPiH!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fba9d8f2e-71e0-49fa-8075-4de1016bbaec_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!GPiH!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fba9d8f2e-71e0-49fa-8075-4de1016bbaec_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!GPiH!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fba9d8f2e-71e0-49fa-8075-4de1016bbaec_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!GPiH!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fba9d8f2e-71e0-49fa-8075-4de1016bbaec_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!GPiH!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fba9d8f2e-71e0-49fa-8075-4de1016bbaec_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!GPiH!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fba9d8f2e-71e0-49fa-8075-4de1016bbaec_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ba9d8f2e-71e0-49fa-8075-4de1016bbaec_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Log-log scatter of memory capacity against memory bandwidth for the Groq 3 LPU, the LPX rack, the Rubin GPU, the NVL72 rack and the Vera CPU, with lines of constant sweep rate&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Log-log scatter of memory capacity against memory bandwidth for the Groq 3 LPU, the LPX rack, the Rubin GPU, the NVL72 rack and the Vera CPU, with lines of constant sweep rate" title="Log-log scatter of memory capacity against memory bandwidth for the Groq 3 LPU, the LPX rack, the Rubin GPU, the NVL72 rack and the Vera CPU, with lines of constant sweep rate" srcset="https://substackcdn.com/image/fetch/$s_!GPiH!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fba9d8f2e-71e0-49fa-8075-4de1016bbaec_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!GPiH!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fba9d8f2e-71e0-49fa-8075-4de1016bbaec_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!GPiH!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fba9d8f2e-71e0-49fa-8075-4de1016bbaec_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!GPiH!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fba9d8f2e-71e0-49fa-8075-4de1016bbaec_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 6. SRAM and HBM4 on the same axes. Points on the same dashed line can read their whole memory the same number of times per second.</figcaption></figure></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Networking</h2><p><strong>Networking is the part of the rack that grew fastest.</strong> Data Center networking revenue went from $3.0 billion in the fourth quarter of fiscal 2025 to $14.8 billion in the first quarter of fiscal 2027, 4.9 times in five quarters, and from 8.5% to 19.7% of Data Center revenue. </p><p>For the whole of fiscal 2026 it was $31.4 billion, up 142%. The fourth quarter commentary credited the NVLink compute fabric of GB200 and GB300 systems along with the growth of Ethernet and InfiniBand. </p><p>The second quarter commentary no longer splits Data Center into compute and networking. On the call <strong>NVIDIA said networking revenue grew 18% sequentially</strong> and that Spectrum-X Ethernet grew 2.6 times year over year. </p><p>If that 18% applies to the same line as the first quarter&#8217;s $14.8 billion, second quarter networking was about $17.5 billion. That is an estimate made here, not a reported number.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!LaID!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fddd5fa90-e10f-436f-85aa-9d66833dced6_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!LaID!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fddd5fa90-e10f-436f-85aa-9d66833dced6_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!LaID!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fddd5fa90-e10f-436f-85aa-9d66833dced6_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!LaID!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fddd5fa90-e10f-436f-85aa-9d66833dced6_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!LaID!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fddd5fa90-e10f-436f-85aa-9d66833dced6_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!LaID!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fddd5fa90-e10f-436f-85aa-9d66833dced6_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ddd5fa90-e10f-436f-85aa-9d66833dced6_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Bar chart of Data Center networking revenue from $3.1 billion to $14.8 billion with a line of its share of Data Center revenue going from 10.2% to 19.7%&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Bar chart of Data Center networking revenue from $3.1 billion to $14.8 billion with a line of its share of Data Center revenue going from 10.2% to 19.7%" title="Bar chart of Data Center networking revenue from $3.1 billion to $14.8 billion with a line of its share of Data Center revenue going from 10.2% to 19.7%" srcset="https://substackcdn.com/image/fetch/$s_!LaID!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fddd5fa90-e10f-436f-85aa-9d66833dced6_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!LaID!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fddd5fa90-e10f-436f-85aa-9d66833dced6_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!LaID!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fddd5fa90-e10f-436f-85aa-9d66833dced6_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!LaID!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fddd5fa90-e10f-436f-85aa-9d66833dced6_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 7. Networking revenue and its share of Data Center, through the last quarter with a disclosed split.</figcaption></figure></div><p><strong>Spectrum-X Ethernet Photonics</strong>, which NVIDIA describes as the first co-packaged optics switches with 200 Gb/s SerDes, is in production, and NVIDIA claims<strong> 5x better power efficiency</strong>, 5x longer uptime and 1.3x faster deployment than networks built on traditional transceivers. </p><p>In the second quarter NVIDIA said Spectrum-6 switch systems supporting both pluggable and co-packaged optics are arriving in gigascale AI factories. <strong>NVIDIA&#8217;s case for co-packaged optics</strong> is made at the facility level: power the optics no longer burn is power the operator can give to compute. </p><p>Per GPU, the <strong>September 11 table puts scale-up at 3 TB/s and scale-out at 0.45 TB/s,</strong> as NVIDIA lists them, a ratio of 6.7. The same page names the rack&#8217;s scale-out fabrics, Quantum-X800 InfiniBand and Spectrum-X Ethernet, alongside the ConnectX-9 SuperNICs and BlueField-4 DPUs inside it. </p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Kyber</h2><p>The rack after this one is under more pressure. Jen Huang said at GTC in March 2025 that each Kyber rack for Rubin Ultra would draw 600 kW. </p><p>In early July, SemiAnalysis reported that<strong> Kyber had slipped to 2028</strong> because its printed circuit midplane is hard to manufacture, and that a stopgap design joining two current racks back to back had been dropped after large customers objected. </p><p>Tom&#8217;s Hardware, reporting those claims, cites trade analyses describing a 78-layer board and notes that<strong> Kyber holds 144 GPU packages against 72 in the current Oberon rack. </strong></p><p>NVIDIA&#8217;s reply to the publication was four words: &#8220;<em>Our roadmap is intact</em>.&#8221; The reported delay does not touch Vera Rubin, which uses the current rack. </p><h2>The gigawatt</h2><p>On the call NVIDIA offered a way to measure its share of a buildout: revenue opportunity per gigawatt. It put <em>Hopper at about $18 billion, Grace Blackwell at about $25 billion and Vera Rubin at about $40 billion</em>, with the <strong>Vera Rubin figure</strong> covering the Vera CPU, the Rubin GPU, NVLink, InfiniBand or Ethernet, and the Groq LPU. </p><p>That is 2.2 times the Hopper figure in two generations. Huang also said the total investment in a gigawatt of data center has gone from about $30 billion five years ago to <strong>about $60 billion today</strong>. If both statements describe the same gigawatt, NVIDIA&#8217;s content is about two thirds of it. </p><p>The <strong>Ohio disclosure </strong>gives a second way to check the $40 billion figure. NVIDIA says each generation of its infrastructure deployed at PORTS-Pike could represent about 1.5 million GPUs, or $150 billion to $200 billion of NVIDIA revenue. </p><p>On the call that statement follows the description of the initial 4.25 gigawatts, and reading it that way gives $35.3 billion to $47.1 billion per gigawatt; <strong>$40 billion times 4.25 is $170 billion</strong>, inside the range, which supports that reading. </p><p>The same numbers give about 353,000 GPUs per gigawatt, or 2.83 kW of IT load per GPU once everything else inside the IT envelope, from CPUs to switches to storage, is divided among the GPUs, and $100,000 to $133,000 of NVIDIA revenue per GPU.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!5lsy!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe7fed3ec-9804-40ea-be18-dca9072616ee_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!5lsy!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe7fed3ec-9804-40ea-be18-dca9072616ee_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!5lsy!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe7fed3ec-9804-40ea-be18-dca9072616ee_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!5lsy!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe7fed3ec-9804-40ea-be18-dca9072616ee_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!5lsy!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe7fed3ec-9804-40ea-be18-dca9072616ee_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!5lsy!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe7fed3ec-9804-40ea-be18-dca9072616ee_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e7fed3ec-9804-40ea-be18-dca9072616ee_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Bar chart of NVIDIA revenue opportunity per gigawatt: Hopper 18, Grace Blackwell 25, Vera Rubin 40 billion dollars, with the PORTS-Pike implied range of 35.3 to 47.1 and a guarantee cap of 24.7&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Bar chart of NVIDIA revenue opportunity per gigawatt: Hopper 18, Grace Blackwell 25, Vera Rubin 40 billion dollars, with the PORTS-Pike implied range of 35.3 to 47.1 and a guarantee cap of 24.7" title="Bar chart of NVIDIA revenue opportunity per gigawatt: Hopper 18, Grace Blackwell 25, Vera Rubin 40 billion dollars, with the PORTS-Pike implied range of 35.3 to 47.1 and a guarantee cap of 24.7" srcset="https://substackcdn.com/image/fetch/$s_!5lsy!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe7fed3ec-9804-40ea-be18-dca9072616ee_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!5lsy!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe7fed3ec-9804-40ea-be18-dca9072616ee_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!5lsy!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe7fed3ec-9804-40ea-be18-dca9072616ee_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!5lsy!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe7fed3ec-9804-40ea-be18-dca9072616ee_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 8. NVIDIA&#8217;s per-gigawatt figures from the call, checked against the numbers it gave for the Ohio site.</figcaption></figure></div><p>The product page gives a third data point. NVIDIA&#8217;s 100 MW factory table, built on<strong> DSX with MaxLPS</strong>, counts 40,000 Rubin GPUs, which is 2.5 kW per GPU or 400,000 GPUs per gigawatt. </p><p>If the MaxLPS gain of up to 40% applied in full, the same site without it would hold fewer GPUs at 2.5 kW &#215; 1.4, or 3.5 kW each. The Ohio figures, 2.83 kW per GPU and 88.2% of the MaxLPS density, fall between the two. </p><p><em>The bases are not identical</em>, since the Ohio numbers are per IT gigawatt and the factory table does not say whether its 100 MW is IT load, so the comparison shows consistency rather than proof.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!ITLz!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe48269cd-d32c-4f0f-b0d1-a0b514356783_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!ITLz!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe48269cd-d32c-4f0f-b0d1-a0b514356783_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!ITLz!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe48269cd-d32c-4f0f-b0d1-a0b514356783_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!ITLz!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe48269cd-d32c-4f0f-b0d1-a0b514356783_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!ITLz!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe48269cd-d32c-4f0f-b0d1-a0b514356783_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!ITLz!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe48269cd-d32c-4f0f-b0d1-a0b514356783_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e48269cd-d32c-4f0f-b0d1-a0b514356783_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Horizontal bars of kW per GPU: 2.50 for NVIDIA's 100 MW factory with MaxLPS, 2.83 implied for PORTS-Pike, 3.50 for the same factory without the full MaxLPS gain&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Horizontal bars of kW per GPU: 2.50 for NVIDIA's 100 MW factory with MaxLPS, 2.83 implied for PORTS-Pike, 3.50 for the same factory without the full MaxLPS gain" title="Horizontal bars of kW per GPU: 2.50 for NVIDIA's 100 MW factory with MaxLPS, 2.83 implied for PORTS-Pike, 3.50 for the same factory without the full MaxLPS gain" srcset="https://substackcdn.com/image/fetch/$s_!ITLz!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe48269cd-d32c-4f0f-b0d1-a0b514356783_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!ITLz!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe48269cd-d32c-4f0f-b0d1-a0b514356783_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!ITLz!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe48269cd-d32c-4f0f-b0d1-a0b514356783_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!ITLz!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe48269cd-d32c-4f0f-b0d1-a0b514356783_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 9. Three ways to put a number on watts per GPU at facility scale.</figcaption></figure></div><h2>Power</h2><p>Most of what changes at the facility level comes from power. NVIDIA&#8217;s own description of the problem is that racks today distribute 54 V DC over copper busbars, that<strong> GB200 and GB300 NVL72 racks use up to eight power shelves</strong>, that a Kyber rack at megawatt scale would need up to 64U of power shelves on 54 V, and that a single 1 MW rack on 54 V would need up to 200 kg of copper busbar. </p><p>Its answer is 800 V DC distribution. NVIDIA says row-level 800 V DC busways can move<strong> 85% more power than 415 V AC through the same conductor size and reduce copper requirements by 45%</strong>, and that the architecture improves end-to-end efficiency by up to 5%. It has also said full-scale production of 800 V DC data centers will coincide with Kyber. </p><p>The arithmetic behind the change is Ohm&#8217;s law. A 600 kW rack at 54 V draws <strong>about 11,100 amps</strong>; at 800 V it draws 750 amps, 14.8 times less. Resistive loss scales with the square of the current, so the same conductor would dissipate about 219 times less heat, and that margin is what lets a designer use far less copper instead.</p><p>The calculation leaves out conversion stages and distances, which <strong>differ between the two designs</strong>, but it shows why the voltage had to change. It also ties the facility to the rack: if Kyber slips, so does the point at which NVIDIA expects 800 V DC to reach full scale.</p><p>Cooling moved in the same direction. The September 11 table lists a 45&#176;C inlet for the rack, and at CES in January Huang said<strong> the power of Vera Rubin is twice that of Grace Blackwell </strong>while the airflow is about the same and the water still enters at 45&#176;C, and that at that temperature data centers need no water chillers. The product page does not list a rack power figure. </p><p>The physics of warm water is simple: heat carried equals mass flow times specific heat times the temperature rise, so <strong>removing 100 kW with a 10 K rise takes about 145 liters of water per minute</strong>, and doubling the rise to 20 K halves the flow to about 72 liters. Whenever the outdoor air is cooler than the water coming back from the racks, the heat can be rejected by dry coolers with fans and no compressor, and a warmer loop makes that true for more hours of the year than a colder one.</p><p>Assembly is the other constraint NVIDIA keeps pointing at. Its January technical blog says the<strong> modular compute and switch trays enable up to 18x faster assembly</strong>, its Rubin platform page says the comparison is against Blackwell, and Thunder Compute describes assembly time going from over an hour and a half to about five minutes, a ratio that matches. The product page adds that the rack uses cable-free modular trays and is <strong>supported by more than 80 MGX ecosystem partners.</strong> </p><p>The DSX platform is NVIDIA&#8217;s attempt to sell the facility layer as a reference design. It bundles validated designs that cover compute, networking, storage, power and cooling, a simulation layer called DSX Sim, grid integration called DSX Flex,<strong> open source operations software called DSX OS</strong>, and<strong> DSX MaxLPS</strong>, which NVIDIA says combines 45&#176;C liquid cooling with in-rack technologies so an operator can run up to 40% more GPUs at their most efficient operating point inside a fixed power budget.</p><p>The trade is easy to state: 1.4 times the GPUs comes out ahead as long as each GPU gives up less than 28.6% of its throughput. NVIDIA says the impact on workload performance is minimal, and the announcement does not include the curve. <strong>DSX Flex is running a multi-megawatt pilot </strong>with Emerald AI and Silicon Valley Power that adjusts AI factory load to grid signals. </p><p>The sites are getting larger and more specific. PORTS-Pike is being developed on the grounds of the decommissioned Portsmouth Gaseous Diffusion Plant in Pike County, Ohio, with capacity coming online in phases from 2028.<strong> SB Energy and SoftBank plan at least 10 GW of new generation</strong> to support 8 GW of IT capacity, all of it for OpenAI, plus at least $4.2 billion of regional grid investment, and NVIDIA is investing $1.5 billion in SB Energy.</p><p>The ratio of planned generation to IT capacity there is 1.25. In Australia, NVIDIA announced on September 9 that <em>eight cloud and data center partners there are expanding land</em>, power and shell to host its DSX factories, with up to 2 GW by 2027. </p><p>At a Goldman Sachs conference the next day, Huang repeated his estimate of <strong>$3 trillion to $4 trillion of AI infrastructure spending by 2030</strong> and acknowledged that land and power could slow deployment. </p><h2>Memory</h2><p><strong>Supply-related commitments were $50.3 billion</strong> at the end of the third quarter of fiscal 2026, $95.2 billion a quarter later, $119.0 billion at the end of the first quarter of fiscal 2027, and $279 billion now. </p><p>That is <strong>5.5 times in three quarters</strong>, with $160 billion added in the last quarter alone and attributed mainly to memory. By fiscal year, $92 billion falls due in the rest of fiscal 2027, $87 billion in fiscal 2028 and $88 billion in fiscal 2029, so $267 billion, or 96%, is due by the end of fiscal 2029. </p><p>The <strong>10-Q says the commitments cover data center infrastructure systems</strong>, primarily memory and manufacturing facilities, and that some of the underlying agreements can be cancelled, rescheduled or adjusted before firm orders are placed, possibly at additional cost. </p><p>For scale: at the third quarter guide, <strong>one quarter of cost of revenue is about $28.1 billion</strong>, which is $108.0 billion times one minus the 74.0% margin. The <em>$92 billion due in the rest of fiscal 2027 is about 3.3 quarters of that</em>, and two quarters at that rate come to about $56.2 billion. Inventory rose from $19.8 billion to $31.6 billion over the same three quarters, and NVIDIA ties the latest increase to the Vera Rubin introduction. </p><p>Whatever is not consumed by shipments in those two quarters has to show up as inventory, as prepayments, or as capacity paid for ahead of fiscal 2028, when NVIDIA plans to grow about 70%. <strong>The filing does not break the $92 billion down.</strong></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!1_xs!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F461f72c8-de6f-43eb-aa65-75b0e953c0c5_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!1_xs!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F461f72c8-de6f-43eb-aa65-75b0e953c0c5_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!1_xs!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F461f72c8-de6f-43eb-aa65-75b0e953c0c5_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!1_xs!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F461f72c8-de6f-43eb-aa65-75b0e953c0c5_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!1_xs!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F461f72c8-de6f-43eb-aa65-75b0e953c0c5_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!1_xs!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F461f72c8-de6f-43eb-aa65-75b0e953c0c5_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/461f72c8-de6f-43eb-aa65-75b0e953c0c5_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Grouped bars of supply-related commitments and inventory from Q3 FY26 to Q2 FY27&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Grouped bars of supply-related commitments and inventory from Q3 FY26 to Q2 FY27" title="Grouped bars of supply-related commitments and inventory from Q3 FY26 to Q2 FY27" srcset="https://substackcdn.com/image/fetch/$s_!1_xs!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F461f72c8-de6f-43eb-aa65-75b0e953c0c5_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!1_xs!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F461f72c8-de6f-43eb-aa65-75b0e953c0c5_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!1_xs!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F461f72c8-de6f-43eb-aa65-75b0e953c0c5_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!1_xs!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F461f72c8-de6f-43eb-aa65-75b0e953c0c5_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 10. Supply-related commitments against inventory at the end of each quarter.</figcaption></figure></div><p>The margin guide points the same way. The CFO described &#8220;<em>extreme pricing conditions in memory,</em>&#8221; said the increases had exceeded NVIDIA&#8217;s expectations and would go higher next year, and reset the outlook: <strong>74.0% gross margin in the third quarter, plus or minus 50 basis points</strong>, a trough of 71% to 72% in the fourth quarter, and 72% to 73% in fiscal 2028 as price increases NVIDIA has already executed take effect in its first quarter. </p><p>On a GAAP basis the second quarter was 75.0%, the same as the fourth quarter of fiscal 2026. The only large dip in the eight quarters charted was the 60.5% of the first quarter of fiscal 2026, when NVIDIA took a $4.5 billion charge on H20. </p><p><strong>Put per $100 of revenue, cost of revenue is $25.00</strong> at the second quarter&#8217;s 75.0% margin, $26.00 at the third quarter guide, $28.50 at the midpoint of the fourth quarter trough and $27.50 at the midpoint of the fiscal 2028 range: 4%, 14% and 10% more cost for each dollar of sales than in the second quarter.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!A7f1!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4396d810-43ca-40ad-bf4d-19d8f0065de8_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!A7f1!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4396d810-43ca-40ad-bf4d-19d8f0065de8_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!A7f1!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4396d810-43ca-40ad-bf4d-19d8f0065de8_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!A7f1!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4396d810-43ca-40ad-bf4d-19d8f0065de8_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!A7f1!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4396d810-43ca-40ad-bf4d-19d8f0065de8_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!A7f1!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4396d810-43ca-40ad-bf4d-19d8f0065de8_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/4396d810-43ca-40ad-bf4d-19d8f0065de8_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Line chart of GAAP gross margin from Q3 FY25 to Q2 FY27 with guidance for Q3 FY27, the Q4 FY27 trough and fiscal 2028&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Line chart of GAAP gross margin from Q3 FY25 to Q2 FY27 with guidance for Q3 FY27, the Q4 FY27 trough and fiscal 2028" title="Line chart of GAAP gross margin from Q3 FY25 to Q2 FY27 with guidance for Q3 FY27, the Q4 FY27 trough and fiscal 2028" srcset="https://substackcdn.com/image/fetch/$s_!A7f1!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4396d810-43ca-40ad-bf4d-19d8f0065de8_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!A7f1!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4396d810-43ca-40ad-bf4d-19d8f0065de8_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!A7f1!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4396d810-43ca-40ad-bf4d-19d8f0065de8_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!A7f1!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4396d810-43ca-40ad-bf4d-19d8f0065de8_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 11. GAAP gross margin, actual and guided.</figcaption></figure></div><p>Rubin is built around HBM4, 288 GB per GPU in NVIDIA&#8217;s specification. On the call <strong>NVIDIA said it works with all three major memory suppliers</strong> to add the capacity its roadmap requires, and in the second quarter it announced a multiyear technology partnership with SK hynix. </p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Credit</h2><p><strong>The full list, as of July 26, 2026, is in the CFO commentary</strong>.<sup> </sup>Commitments total $366 billion: $279 billion for supply and capacity; $29 billion of cloud service agreements that support NVIDIA&#8217;s own research, open models and autonomous vehicle software; <strong>$25 billion of data center leases </strong>that have not started, with terms of up to 20 years beginning between the third quarter of fiscal 2027 and fiscal 2033; $25 billion of equity investments in AI model makers, infrastructure financiers and other private companies; and $8 billion of capital expenditure. </p><p>Additional commitments total $56 billion: $36 billion of AI cloud agreements, and $20 billion of leases with terms of about 15 years, starting in fiscal 2028 or 2029, that NVIDIA expects to reassign. </p><p><strong>Guarantees add $108.5 billion.</strong> The sum, $530.5 billion, is 5.8 times the $91.3 billion of total liabilities on the balance sheet and 1.75 times the $303.0 billion of revenue over the last four quarters. The sum mixes firm purchase obligations with a contingent maximum exposure, so it is a ceiling rather than a bill.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!sYbF!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33128ace-b7ac-4c11-b077-47394b82c45f_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!sYbF!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33128ace-b7ac-4c11-b077-47394b82c45f_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!sYbF!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33128ace-b7ac-4c11-b077-47394b82c45f_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!sYbF!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33128ace-b7ac-4c11-b077-47394b82c45f_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!sYbF!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33128ace-b7ac-4c11-b077-47394b82c45f_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!sYbF!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33128ace-b7ac-4c11-b077-47394b82c45f_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/33128ace-b7ac-4c11-b077-47394b82c45f_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Horizontal stacked bar of $530.5 billion of commitments and guarantees next to $91.3 billion of total liabilities&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Horizontal stacked bar of $530.5 billion of commitments and guarantees next to $91.3 billion of total liabilities" title="Horizontal stacked bar of $530.5 billion of commitments and guarantees next to $91.3 billion of total liabilities" srcset="https://substackcdn.com/image/fetch/$s_!sYbF!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33128ace-b7ac-4c11-b077-47394b82c45f_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!sYbF!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33128ace-b7ac-4c11-b077-47394b82c45f_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!sYbF!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33128ace-b7ac-4c11-b077-47394b82c45f_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!sYbF!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33128ace-b7ac-4c11-b077-47394b82c45f_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 12. Everything NVIDIA lists under commitments, additional commitments and guarantees, next to the liabilities on its balance sheet.</figcaption></figure></div><p>The <strong>AI cloud agreements</strong> are the new commercial structure. NVIDIA earns revenue on the upfront sale of the infrastructure and, if certain criteria are met, takes part of the revenue the AI cloud earns from its own customers. </p><p>On the call the CFO explained the mechanics. NVIDIA commits to take or pay for part of a facility&#8217;s capacity, which gives lenders a minimum revenue to underwrite, and in exchange shares in the neocloud&#8217;s revenue above that floor.<strong> NVIDIA says independent capital still underwrites each deal</strong>, that it is not making loans, and that it gets paid twice, once for the hardware and again from rental income. </p><p>The Ohio guarantees are residual value guaranties, and the 8-K describes how they work. There are several agreements tied to the leases, covering about 4.25 GW of IT load in total. <strong>Each generally becomes effective when its lease commences</strong>, total payment obligations are capped at $105 billion, and payment is conditional, among other things, on the lessor meeting ready-for-service conditions, expected from 2028. </p><p>The trigger events are an OpenAI insolvency that causes a lease default, or OpenAI failing to pay under a lease. In either case NVIDIA pays roughly the <strong>shortfall between a guaranteed minimum value of the lease and whatever is recovered</strong> by reletting or selling, and it can choose to assume the lease, have the lessor try to relet, start a sale, let the lease terminate, or defer for up to a year while covering specified costs. </p><p>The<strong> CFO commentary</strong> adds that the exposure declines as OpenAI pays rent, and that NVIDIA has the option to support about 3.8 GW more. Data Center Frontier reports that OpenAI has agreed to reimburse NVIDIA for amounts it actually pays. </p><p>In effect, <strong>NVIDIA is guaranteeing the longest-lived layer of the site</strong> (<em>land, power and shell under 20-year leases</em>) in order to sell the shortest-lived one (<em>compute that NVIDIA expects to upgrade several times over those 20 years</em>). </p><p>The cap works out to $24.7 billion per gigawatt, and to between 52.5% and 70% of the revenue NVIDIA attaches to a single generation of hardware at the site. What NVIDIA would actually lose in a default is not the rent itself but the <strong>gap between the guaranteed value and what a powered shell of that size would fetch</strong> from a new tenant or a buyer at that moment.</p><p>The rest of the lab exposure came out on the call. NVIDIA has invested nearly $50 billion in frontier AI labs. For another lab, which it did not name, it will provide <strong>selective credit enhancement for nearly 2 GW of compute.</strong> OpenAI&#8217;s existing and planned commitments add up to about 12 GW of NVIDIA compute. </p><p>The financing platforms NVIDIA announced on August 10 with six large asset managers and banks aim to <strong>mobilize more than $500 billion of third-party capital</strong>, but they rest on memorandums of understanding and remain subject to final agreements.</p><p>The CFO anticipated the obvious objection, that this is circular financing, and answered that NVIDIA compute is fungible and can be redeployed to other customers, which in NVIDIA&#8217;s view limits the risk. </p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Cash</h2><p>GAAP net income was $59.7 billion in the quarter, operating cash flow $24.1 billion and free cash flow $21.3 billion, against $58.3 billion, $50.3 billion and $48.6 billion in the first quarter.<strong> Operating cash flow was 40% of net income</strong>, down from 86%. The CFO commentary attributes the sequential decline to higher working capital and cash taxes.</p><p>The <strong>cash flow statement </strong>shows the working capital part: receivables absorbed $22.3 billion, inventories $5.8 billion, and prepaid expenses and other assets $5.5 billion, and net income included $7.8 billion of non-cash, pre-tax gains on equity securities. </p><p><strong>Days sales outstanding rose from 45 to 60</strong>, which NVIDIA attributes to extended payment terms on large multi-quarter agreements with certain investment-grade customers. </p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!nfGk!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fecf190bb-4c7c-4ffb-8273-33ca27e6461b_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!nfGk!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fecf190bb-4c7c-4ffb-8273-33ca27e6461b_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!nfGk!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fecf190bb-4c7c-4ffb-8273-33ca27e6461b_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!nfGk!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fecf190bb-4c7c-4ffb-8273-33ca27e6461b_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!nfGk!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fecf190bb-4c7c-4ffb-8273-33ca27e6461b_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!nfGk!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fecf190bb-4c7c-4ffb-8273-33ca27e6461b_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ecf190bb-4c7c-4ffb-8273-33ca27e6461b_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Grouped bars of net income, operating cash flow and free cash flow for four quarters with a line of days sales outstanding&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Grouped bars of net income, operating cash flow and free cash flow for four quarters with a line of days sales outstanding" title="Grouped bars of net income, operating cash flow and free cash flow for four quarters with a line of days sales outstanding" srcset="https://substackcdn.com/image/fetch/$s_!nfGk!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fecf190bb-4c7c-4ffb-8273-33ca27e6461b_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!nfGk!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fecf190bb-4c7c-4ffb-8273-33ca27e6461b_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!nfGk!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fecf190bb-4c7c-4ffb-8273-33ca27e6461b_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!nfGk!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fecf190bb-4c7c-4ffb-8273-33ca27e6461b_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 13. Earnings against cash, with days sales outstanding.</figcaption></figure></div><p>The balance sheet has changed shape as well. <strong>Marketable equity securities of $42.8 billion and non-marketable securities of $51.2 billion </strong>total $93.9 billion, up from $35.1 billion at January 25 and now larger than the <em>$56.6 billion of cash and marketable debt securities. </em>In the first half of the fiscal year NVIDIA spent $42.4 billion on purchases of equity securities and booked $23.7 billion of net gains on them. </p><p><strong>Total debt is $33.4 billion </strong>after $25.0 billion of senior unsecured notes issued in the quarter, against $8.5 billion at January 25, and financing activities include a $2.9 billion outflow labeled Groq, Inc. </p><p>Cash paid for buybacks and dividends was $25.8 billion <em>($19.7 billion and $6.0 billion)</em>, which NVIDIA describes as approximately $26.0 billion returned. The <strong>quarterly dividend went from $0.01 to $0.25 a share </strong>after a board decision on May 18, and $99.0 billion of buyback authorization remains. </p><div><hr></div><h2>November 17</h2><p><strong>NVIDIA reports its third quarter on November 17</strong>. The numbers above give a short list of things to check in that release: revenue against the $108.0 billion guide, and whether it clears the $117.3 billion that would keep growth accelerating;</p><p><strong>Vera Rubin&#8217;s share of Data Center against the 20% NVIDIA expects</strong>; gross margin at 74.0% on the way to the fourth quarter trough; days sales outstanding against this quarter&#8217;s 60; and the supply line against inventory, which shows whether the commitments are turning into goods or into capacity reserved for later. Before that, Huang gives a keynote at GTC Berlin on October 21.</p><h2>Reproduce the numbers</h2><p>Every value marked M in the dossier comes from arithmetic on the sources, and the core of it fits in a short Python script with no dependencies. It prints the values used above.</p><pre><code><code># Inputs come from the NVIDIA filings and product pages listed under Sources.
rev = {"Q2FY26": 46743, "Q3FY26": 57006, "Q2FY27": 96221}   # $ millions

# Q3 FY27 revenue needed to keep year-over-year growth accelerating ($B)
print(round(rev["Q3FY26"] * rev["Q2FY27"] / rev["Q2FY26"] / 1000, 1))

# Q1 FY27 revenue moved from ACIE to Hyperscale by the recast ($B)
print((43050 - 37869) / 1000)

# Critical batch, weights only, zero context: B* = P * b / (2 * B_mem)
def b_star(pflops, bytes_per_param, bandwidth_tbs):
    return pflops * 1e15 * bytes_per_param / (2 * bandwidth_tbs * 1e12)

print(round(b_star(17.5, 1.0, 19.2)), round(b_star(4, 2.0, 19.2)))   # Sep 11 table
print(round(b_star(17.5, 1.0, 22.0)), round(b_star(4, 2.0, 22.0)))   # earlier table

# Sweep rate: bandwidth over capacity, per second
rubin = 19.2e3 / 288
lpu = 150e3 / 0.5
print(round(rubin, 1), round(lpu), round(lpu / rubin))

# Current for a 600 kW rack at 54 V and 800 V, and the resistive loss ratio
i54, i800 = 600e3 / 54, 600e3 / 800
print(round(i54), round(i800), round((i54 / i800) ** 2))

# Ohio: NVIDIA revenue per gigawatt ($B) and IT kW per GPU
print(round(150 / 4.25, 1), round(200 / 4.25, 1), round(4.25e6 / 1.5e6, 2))

# Water to remove 100 kW with a 10 K rise at 45 C (liters per minute)
print(round(100 / (4.18 * 10) / 0.990 * 60))
</code></code></pre><h2>Glossary</h2><h3>Notes on the numbers</h3><p>NVIDIA&#8217;s fiscal 2027 began on January 26, 2026, and its second quarter ended on July 26, 2026. Fiscal 2028 roughly corresponds to calendar 2027. Margins are GAAP throughout, because since the <strong>first quarter of fiscal 2027 NVIDIA&#8217;s non-GAAP measures </strong>include stock-based compensation and the historical non-GAAP figures were restated.</p><p>The fourth quarter trough and the fiscal 2028 range were given on the call without specifying GAAP or non-GAAP; the two measures were within 0.1 points of each other in each of the last two quarters.</p><p>The compute and networking split of Data Center was last given for the first quarter of fiscal 2027, rounded to $60.4 billion and $14.8 billion. The second quarter networking figure above is an estimate made here. <strong>Hyperscale and ACIE figures for the second quarter and the recast first quarter</strong> come from the second quarter commentary, and the as-reported first quarter figures come from the first quarter commentary. </p><p>Supply-related commitments are charted for the four quarters in which the commentaries used here report a total under that label. The first quarter of fiscal 2026 used a different definition, $29.8 billion of purchase commitments and obligations for inventory and manufacturing capacity, and is not charted. </p><p><strong>Vera Rubin specifications are taken from NVIDIA&#8217;s product page</strong> as updated on September 11, 2026. The earlier values come from Spheron&#8217;s record of the page on August 16 and NVIDIA&#8217;s January technical blog, and the earlier scale-out figure from SiliconReport. </p><p>Transfer times assume each interconnect figure counts both directions and treat <strong>BlueField-4&#8217;s 800 Gb/s as one direction.</strong> Vera&#8217;s memory bandwidth comes from NVIDIA&#8217;s Rubin platform page, not from the NVL72 table. </p><p>Per-gigawatt revenue, the performance multipliers, the neocloud gigawatt figures and the hyperscaler capital expenditure figures are statements by NVIDIA, not audited numbers. The Kyber delay is a third-party report that NVIDIA has not confirmed. <strong>Values marked M in the dossier below are arithmetic on filed numbers</strong>, with the formula given.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Sources</h2><ol><li><p>NVIDIA, CFO Commentary on Second Quarter Fiscal 2027 Results (8-K exhibit), Aug 26, 2026. <a href="https://www.sec.gov/Archives/edgar/data/1045810/000104581026000073/q2fy27cfocommentary.htm">https://www.sec.gov/Archives/edgar/data/1045810/000104581026000073/q2fy27cfocommentary.htm</a></p></li><li><p>NVIDIA, press release: Financial Results for Second Quarter Fiscal 2027, Aug 26, 2026. <a href="https://nvidianews.nvidia.com/news/nvidia-announces-financial-results-for-second-quarter-fiscal-2027">https://nvidianews.nvidia.com/news/nvidia-announces-financial-results-for-second-quarter-fiscal-2027</a></p></li><li><p>NVIDIA Q2 FY2027 earnings call, corrected transcript (FactSet, hosted by NVIDIA IR), Aug 26, 2026. <a href="https://s201.q4cdn.com/141608511/files/content_files/TRANSCRIPT_-NVIDIA-Corp-NVDA-US-Q2-2027-Earnings-Call-26-August-2026-5_00-PM-ET.pdf">https://s201.q4cdn.com/141608511/files/content_files/TRANSCRIPT_-NVIDIA-Corp-NVDA-US-Q2-2027-Earnings-Call-26-August-2026-5_00-PM-ET.pdf</a></p></li><li><p>NVIDIA, CFO Commentary on First Quarter Fiscal 2027 Results, May 2026. <a href="https://s201.q4cdn.com/141608511/files/doc_financials/2027/Q127/Q1FY27-CFO-Commentary.pdf">https://s201.q4cdn.com/141608511/files/doc_financials/2027/Q127/Q1FY27-CFO-Commentary.pdf</a></p></li><li><p>NVIDIA, CFO Commentary on Fourth Quarter and Fiscal 2026 Results (8-K exhibit), Feb 25, 2026. <a href="https://www.sec.gov/Archives/edgar/data/1045810/000104581026000019/q4fy26cfocommentary.htm">https://www.sec.gov/Archives/edgar/data/1045810/000104581026000019/q4fy26cfocommentary.htm</a></p></li><li><p>NVIDIA, CFO Commentary on Third Quarter Fiscal 2026 Results (8-K exhibit), Nov 19, 2025. <a href="https://www.sec.gov/Archives/edgar/data/1045810/000104581025000228/q3fy26cfocommentary.htm">https://www.sec.gov/Archives/edgar/data/1045810/000104581025000228/q3fy26cfocommentary.htm</a></p></li><li><p>NVIDIA, CFO Commentary on First Quarter Fiscal 2026 Results, May 2025. <a href="https://s201.q4cdn.com/141608511/files/doc_financials/2026/Q126/Q1FY26-CFO-Commentary.pdf">https://s201.q4cdn.com/141608511/files/doc_financials/2026/Q126/Q1FY26-CFO-Commentary.pdf</a></p></li><li><p>NVIDIA, press release: Financial Results for Second Quarter Fiscal 2026 (8-K exhibit), Aug 27, 2025. <a href="https://www.sec.gov/Archives/edgar/data/1045810/000104581025000207/q2fy26pr.htm">https://www.sec.gov/Archives/edgar/data/1045810/000104581025000207/q2fy26pr.htm</a></p></li><li><p>NVIDIA, Form 10-Q for the quarter ended July 26, 2026, Aug 2026. <a href="https://www.sec.gov/Archives/edgar/data/0001045810/000104581026000075/nvda-20260726.htm">https://www.sec.gov/Archives/edgar/data/0001045810/000104581026000075/nvda-20260726.htm</a></p></li><li><p>NVIDIA, Form 8-K on the SB Energy residual value guaranties, Aug 17, 2026. <a href="https://www.sec.gov/Archives/edgar/data/1045810/000104581026000069/nvda-20260817.htm">https://www.sec.gov/Archives/edgar/data/1045810/000104581026000069/nvda-20260817.htm</a></p></li><li><p>NVIDIA, press release: NVIDIA Guarantees SB Energy&#8217;s PORTS-Pike Technology Campus, Aug 17, 2026. <a href="https://nvidianews.nvidia.com/news/nvidia-guarantees-sb-energy-s-ports-pike-technology-campus-in-ohio-to-exclusively-host-nvidia-ai-compute">https://nvidianews.nvidia.com/news/nvidia-guarantees-sb-energy-s-ports-pike-technology-campus-in-ohio-to-exclusively-host-nvidia-ai-compute</a></p></li><li><p>NVIDIA, press release: compute financing platforms with Apollo, BlackRock, Blackstone, Brookfield, Goldman Sachs and KKR, Aug 10, 2026. <a href="https://nvidianews.nvidia.com/news/nvidia-partners-with-apollo-blackrock-blackstone-brookfield-goldman-sachs-and-kkr-to-establish-ai-compute-infrastructure-financing-platforms-to-mobilize-over-500-billion-of-third-party-capital">https://nvidianews.nvidia.com/news/nvidia-partners-with-apollo-blackrock-blackstone-brookfield-goldman-sachs-and-kkr-to-establish-ai-compute-infrastructure-financing-platforms-to-mobilize-over-500-billion-of-third-party-capital</a></p></li><li><p>NVIDIA, press release: Vera Rubin Ramps Into Full Production, May 31, 2026. <a href="https://nvidianews.nvidia.com/news/vera-rubin-full-production-agentic-ai-factory">https://nvidianews.nvidia.com/news/vera-rubin-full-production-agentic-ai-factory</a></p></li><li><p>NVIDIA, Vera Rubin NVL72 product page and specifications table, updated Sep 11, 2026, accessed Sep 16, 2026. <a href="https://www.nvidia.com/en-us/data-center/vera-rubin-nvl72">https://www.nvidia.com/en-us/data-center/vera-rubin-nvl72</a></p></li><li><p>NVIDIA Technical Blog, Inside the NVIDIA Vera Rubin Platform, Jan 5, 2026. <a href="https://developer.nvidia.com/blog/inside-the-nvidia-rubin-platform-six-new-chips-one-ai-supercomputer/">https://developer.nvidia.com/blog/inside-the-nvidia-rubin-platform-six-new-chips-one-ai-supercomputer/</a></p></li><li><p>NVIDIA, press release: NVIDIA DSX Gives Infrastructure Builders the Playbook for AI Factories, May 31, 2026. <a href="https://nvidianews.nvidia.com/news/dsx-infrastructure-ai-factory">https://nvidianews.nvidia.com/news/dsx-infrastructure-ai-factory</a></p></li><li><p>NVIDIA Technical Blog, NVIDIA 800 VDC Architecture Will Power the Next Generation of AI Factories, updated Jan 14, 2026. <a href="https://developer.nvidia.com/blog/nvidia-800-v-hvdc-architecture-will-power-the-next-generation-of-ai-factories/">https://developer.nvidia.com/blog/nvidia-800-v-hvdc-architecture-will-power-the-next-generation-of-ai-factories/</a></p></li><li><p>Tom&#8217;s Hardware, Nvidia&#8217;s Kyber rack for Rubin Ultra reportedly delayed to 2028 (reporting a SemiAnalysis thread; includes NVIDIA&#8217;s statement), Jul 6, 2026. <a href="https://www.tomshardware.com/pc-components/gpus/nvidias-kyber-rack-for-rubin-ultra-slips-to-2028">https://www.tomshardware.com/pc-components/gpus/nvidias-kyber-rack-for-rubin-ultra-slips-to-2028</a></p></li><li><p>DCD, Nvidia&#8217;s Rubin Ultra NVL576 rack expected to be 600kW, GTC 2025 remarks. <a href="https://www.datacenterdynamics.com/en/news/nvidias-rubin-ultra-nvl576-rack-expected-to-be-600kw-coming-second-half-of-2027/">https://www.datacenterdynamics.com/en/news/nvidias-rubin-ultra-nvl576-rack-expected-to-be-600kw-coming-second-half-of-2027/</a></p></li><li><p>NVIDIA, press release: Expands AI Infrastructure Capacity in Partnership With Australia&#8217;s Data Center Ecosystem, Sep 9, 2026. <a href="https://nvidianews.nvidia.com/news/nvidia-expands-ai-infrastructure-capacity-in-partnership-with-australias-data-center-ecosystem">https://nvidianews.nvidia.com/news/nvidia-expands-ai-infrastructure-capacity-in-partnership-with-australias-data-center-ecosystem</a></p></li><li><p>Investing.com, NVIDIA at Goldman Sachs Communacopia + Technology Conference (transcript summary), Sep 10, 2026. <a href="https://www.investing.com/news/transcripts/nvidia-at-goldman-sachs-conference-huang-sees-ai-buildout-still-early-93CH-4896555">https://www.investing.com/news/transcripts/nvidia-at-goldman-sachs-conference-huang-sees-ai-buildout-still-early-93CH-4896555</a></p></li><li><p>Data Center Frontier, PORTS-Pike Takes Shape as an 8-GW AI Infrastructure Model, Aug 2026. <a href="https://www.datacenterfrontier.com/hyperscale/article/55398883/ports-pike-takes-shape-as-an-8-gw-ai-infrastructure-model">https://www.datacenterfrontier.com/hyperscale/article/55398883/ports-pike-takes-shape-as-an-8-gw-ai-infrastructure-model</a></p></li><li><p>Tom&#8217;s Hardware, Nvidia shows off Rubin Ultra with 600,000-Watt Kyber racks (GTC 2025 coverage), Mar 19, 2025. <a href="https://www.tomshardware.com/pc-components/gpus/nvidia-shows-off-rubin-ultra-with-600-000-watt-kyber-racks-and-infrastructure-coming-in-2027">https://www.tomshardware.com/pc-components/gpus/nvidia-shows-off-rubin-ultra-with-600-000-watt-kyber-racks-and-infrastructure-coming-in-2027</a></p></li><li><p>NVIDIA, Groq 3 LPX product page, updated Aug 24, 2026, accessed Sep 16, 2026. <a href="https://www.nvidia.com/en-us/data-center/lpx/">https://www.nvidia.com/en-us/data-center/lpx/</a></p></li><li><p>Spheron, NVIDIA Vera Rubin NVL72 guide (quotes the NVIDIA specification table as retrieved on Aug 16, 2026), Sep 2026. <a href="https://www.spheron.network/blog/nvidia-vera-rubin-nvl72-guide/">https://www.spheron.network/blog/nvidia-vera-rubin-nvl72-guide/</a></p></li><li><p>VideoCardz, NVIDIA Vera Rubin NVL72 Detailed (CES 2026 specifications), Jan 6, 2026. <a href="https://videocardz.com/newz/nvidia-vera-rubin-nvl72-detailed-72-gpus-36-cpus-260-tb-s-scale-up-bandwidth">https://videocardz.com/newz/nvidia-vera-rubin-nvl72-detailed-72-gpus-36-cpus-260-tb-s-scale-up-bandwidth</a></p></li><li><p>Fierce Network, Supercomputers can stay chill with hot water says Nvidia (CES 2026 keynote remarks), Jan 2026. <a href="https://www.fierce-network.com/cloud/nvidia-has-no-chill">https://www.fierce-network.com/cloud/nvidia-has-no-chill</a></p></li><li><p>SiliconReport, Nvidia Vera Rubin Explained (earlier scale-out figure and NVL144 naming history), Jul 3, 2026. <a href="https://www.siliconreport.com/nvidia-vera-rubin-everything-we-know-33727d4d">https://www.siliconreport.com/nvidia-vera-rubin-everything-we-know-33727d4d</a></p></li><li><p>Thunder Compute, Nvidia Rubin Architecture (assembly time; quotes the earlier bandwidth figures), Sep 2026. <a href="https://www.thundercompute.com/blog/nvidia-rubin-architecture">https://www.thundercompute.com/blog/nvidia-rubin-architecture</a></p></li><li><p>NVIDIA, NVLink and NVLink Switch page (describes the NVLink figure as bidirectional), accessed Sep 16, 2026. <a href="https://www.nvidia.com/en-eu/data-center/nvlink">https://www.nvidia.com/en-eu/data-center/nvlink</a></p></li><li><p>NVIDIA, Rubin platform page (Vera CPU memory bandwidth; cable-free trays, 18x faster assembly than Blackwell), accessed Sep 16, 2026. <a href="https://www.nvidia.com/en-eu/data-center/technologies/rubin/">https://www.nvidia.com/en-eu/data-center/technologies/rubin/</a></p></li></ol><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Claim dossier</h2><p>A: filed or official financial document. B: company statement or product claim, not independently verified. C: third-party reporting. M: computed here from the cited inputs, with the formula or method given. Source numbers refer to the Sources list above.</p><h3>Results and guidance</h3><ul><li><p><strong>Q2 FY27 revenue.</strong> $96,221M, up 18% q/q and 106% y/y. <em>Source 1. Tier A.</em></p></li><li><p><strong>Data Center revenue.</strong> $89,023M, up 18% q/q and 117% y/y. <em>Source 1. Tier A.</em></p></li><li><p><strong>Data Center share of revenue.</strong> 92.5% = 89,023 / 96,221; Q1 92.2%. <em>Sources 1, 4. Tier M.</em></p></li><li><p><strong>Edge Computing revenue.</strong> $7,198M, up 13% q/q and 27% y/y. <em>Source 1. Tier A.</em></p></li><li><p><strong>Revenue growth, y/y.</strong> Q1 FY26 69.2%, Q2 FY26 55.6%, Q3 FY26 62.5%, Q4 FY26 73.2%, Q1 FY27 85.2%, Q2 FY27 105.9%. <em>Sources 1, 4, 5, 6, 7, 8. Tier M.</em></p></li><li><p><strong>Q3 FY27 guidance.</strong> $108.0B plus or minus 2%; GAAP gross margin 74.0% plus or minus 50 bp. <em>Source 1. Tier A.</em></p></li><li><p><strong>Growth at the guide.</strong> 89.5% at midpoint, 93.2% at the top = guide / 57,006. <em>Sources 1, 6. Tier M.</em></p></li><li><p><strong>Revenue needed to keep accelerating.</strong> $117.3B = 57,006 x 96,221 / 46,743; 8.7% above midpoint. <em>Sources 1, 6. Tier M.</em></p></li><li><p><strong>Beats against guide midpoint.</strong> Q1 4.6% (guide $78.0B), Q2 5.7% (guide $91.0B). <em>Sources 1, 4, 5. Tier M.</em></p></li><li><p><strong>Fiscal 2028 outlook.</strong> About 70% revenue growth, supply-constrained. <em>Source 3. Tier B.</em></p></li><li><p><strong>China.</strong> Hopper shipments to China below 1% of Data Center revenue in Q2 (call); no China Data Center compute revenue in the Q3 outlook; $4.6B of H20 in Q1 FY26. <em>Sources 2, 3, 7. Tiers A, B.</em></p></li></ul><h3>Customers</h3><ul><li><p><strong>Hyperscale and ACIE, Q2 FY27.</strong> $48,710M and $40,313M. <em>Source 1. Tier A.</em></p></li><li><p><strong>Hyperscale and ACIE, Q1 FY27 as reported.</strong> $37,869M and $37,377M. <em>Source 4. Tier A.</em></p></li><li><p><strong>Hyperscale and ACIE, Q1 FY27 recast.</strong> $43,050M and $32,196M. <em>Source 1. Tier A.</em></p></li><li><p><strong>Revenue moved by the recast.</strong> $5,181M = 43,050 - 37,869; 6.9% of Q1 Data Center. <em>Sources 1, 4. Tier M.</em></p></li><li><p><strong>ACIE share of Data Center.</strong> 49.7% (Q1 reported), 42.8% (Q1 recast), 45.3% (Q2). <em>Sources 1, 4. Tier M.</em></p></li><li><p><strong>Neocloud capacity.</strong> About 3 GW at end of 2025; 8 GW expected at end of 2026. <em>Source 3. Tier B.</em></p></li><li><p><strong>Hyperscaler capital expenditure.</strong> Nearly $800B in 2026 and $1.3T in 2027 (top five); backlog above $2T. <em>Source 3. Tier B.</em></p></li><li><p><strong>AWS.</strong> Additional 2 million GPUs through Q2 FY2029. <em>Source 3. Tier B.</em></p></li></ul><h3>Vera Rubin and the rack</h3><ul><li><p><strong>Vera Rubin NVL72 specification (Sep 11 revision).</strong> 72 GPUs, 36 CPUs; 288 GB and 19.2 TB/s per GPU; 20.7 TB and 1,400 TB/s per rack; NVLink 3 TB/s and 216 TB/s; C2C 1.8 TB/s; 88 cores and 176 threads per CPU; up to 1.5 TB LPDDR5X per CPU, 54 TB per rack; scale-out 0.45 and 32.4 TB/s bidirectional; 1,296 NVIDIA and HBM4 chips; inlet 45 C. <em>Source 14. Tier B.</em></p></li><li><p><strong>Ratings per GPU.</strong> 50 PFLOPS NVFP4 inference (sparse); 35 NVFP4 training, 17.5 FP8/FP6, 4 BF16, 2 TF32 (dense); 130 TFLOPS FP32; 33 TFLOPS FP64. <em>Source 14. Tier B.</em></p></li><li><p><strong>Specification ratios.</strong> 12.5 = 3,600 / 288; 15.6% = 3 / 19.2; span 1,515 = 50,000 / 33; FP8/BF16 4.375; TF32/FP32 15.4; sparse/dense NVFP4 1.43. <em>Source 14. Tier M.</em></p></li><li><p><strong>HBM4 and memory bandwidth.</strong> HBM4 doubles HBM3e interface width; nearly 3x Blackwell memory bandwidth. <em>Source 15. Tier B.</em></p></li><li><p><strong>Vera Rubin shipments.</strong> Production shipments began in August; about 20% of Q3 Data Center revenue. <em>Source 3. Tier B.</em></p></li><li><p><strong>Vera Rubin revenue in Q3.</strong> About $20.0B = 108.0 x 0.925 x 0.20; Data Center about $99.9B. <em>Sources 1, 3. Tier M.</em></p></li><li><p><strong>Vera Rubin performance claims.</strong> Product page: up to 10x tokens per MW and one tenth cost per million tokens against GB200 NVL72 (Kimi-K2-Thinking 32K/8K), a quarter of the GPUs for a 10T MoE on 100T tokens, up to 35x per MW with LPX; call: 30x per MW and 35x lower token cost against Grace Blackwell Ultra; May: 10x agent throughput against Grace Blackwell. <em>Sources 3, 13, 14. Tier B.</em></p></li><li><p><strong>Grace Blackwell Ultra.</strong> GB300, the rack generation after GB200. <em>Source 28. Tier C.</em></p></li><li><p><strong>CPUs.</strong> Grace above $5B trailing twelve months; CPU revenue more than double in FY28. <em>Source 3. Tier B.</em></p></li><li><p><strong>Assembly and partners.</strong> Up to 18x faster assembly than Blackwell (January blog, Rubin page); over 1.5 hours to about 5 minutes (Thunder Compute); cable-free trays; more than 80 MGX partners. <em>Sources 14, 15, 29, 31. Tiers B, C.</em></p></li><li><p><strong>CES 2026.</strong> Vera Rubin described as in full production. <em>Sources 26, 27. Tier C.</em></p></li></ul><h3>The revised table</h3><ul><li><p><strong>Values before the revision.</strong> 22 TB/s and 1,580 TB/s HBM4; 260 TB/s NVLink per rack (page as retrieved Aug 16); 3.6 TB/s NVLink 6 per GPU (January blog); 28.8 TB/s scale-out per rack (July); compute and capacity as in January. <em>Sources 15, 25, 26, 28. Tiers B, C.</em></p></li><li><p><strong>Revision changes.</strong> HBM per GPU -12.7%; per rack -11.4%; NVLink per GPU -16.7%; per rack -16.9%; scale-out +12.5%. <em>Sources 14, 25, 28. Tier M.</em></p></li><li><p><strong>100 MW factory table.</strong> 40,000 GPUs with MaxLPS; 2 ZFLOPS NVFP4 inference; 1.4 ZFLOPS NVFP4 training; 700, 160, 80 EFLOPS; 12 PB HBM4 at 800 PB/s; up to 30 PB LPDDR5X; 42 PB fast memory. <em>Source 14. Tier B.</em></p></li><li><p><strong>Factory table check.</strong> 40,000 x 19.2 = 768 PB/s (listed 800); 40,000 x 22 = 880; rack product 1,382.4 against 1,400 (1.3%). <em>Source 14. Tier M.</em></p></li><li><p><strong>Critical batch B*.</strong> New: 455.7 (NVFP4, FP8 training), 208.3 (BF16, TF32); old: 397.7, 181.8; change 14.6%; sparse row refused (651 if forced). <em>Sources 14, 25. Tier M.</em></p></li></ul><h3>Memory, bandwidth and SRAM</h3><ul><li><p><strong>Bandwidth ladder.</strong> HBM4 / NVLink 6.4; / C2C 10.7; / Vera 16; / scale-out 42.7; NVLink / scale-out 6.67; BlueField-4 800 Gb/s = 0.1 TB/s. <em>Sources 13, 14, 26. Tier M.</em></p></li><li><p><strong>Vera memory bandwidth.</strong> Up to 1.2 TB/s (Rubin platform page; also reported from CES). <em>Sources 26, 31. Tier B.</em></p></li><li><p><strong>Transfer times for 100 GB.</strong> NVLink 67 ms, C2C 111 ms, scale-out 444 ms (figures halved as bidirectional, per the product page and the NVLink page); BlueField-4 1 s; HBM4 read 5.2 ms. <em>Sources 13, 14. Tier M.</em></p></li><li><p><strong>Fast memory per rack.</strong> 20.7 + 54 = 74.7 TB. <em>Source 14. Tier M.</em></p></li><li><p><strong>Groq 3 LPX.</strong> 256 LPUs; 500 MB SRAM, 150 TB/s SRAM bandwidth, 2.5 TB/s scale-up per LPU; 128 GB SRAM, 12 TB DDR5, 40 PB/s, 640 TB/s per rack. <em>Source 24. Tier B.</em></p></li><li><p><strong>LPX rack checks.</strong> 256 x 0.5 = 128 GB; 256 x 2.5 = 640 TB/s; 256 x 150 = 38.4 PB/s against 40 listed. <em>Source 24. Tier M.</em></p></li><li><p><strong>Sweep rates.</strong> Rubin 66.7/s; LPU 300,000/s; ratio 4,500; Vera 0.8/s. <em>Sources 14, 24, 26. Tier M.</em></p></li><li><p><strong>LPX against NVL72.</strong> Capacity 161.7x lower (20,700 / 128 GB); bandwidth 27.8x higher (per-chip products; 28.6x with rounded rack figures). <em>Sources 14, 24. Tier M.</em></p></li></ul><h3>Networking and Kyber</h3><ul><li><p><strong>Networking revenue.</strong> Q3 FY25 $3.1B, Q4 FY25 $3.0B, Q1 FY26 $5.0B, Q2 FY26 $7.3B, Q3 FY26 $8.2B, Q4 FY26 $11.0B, Q1 FY27 $14.8B. <em>Sources 4, 5, 6, 7. Tier A.</em></p></li><li><p><strong>Networking share and growth.</strong> 8.5% to 19.7%; 4.89x from Q4 FY25 to Q1 FY27; FY26 $31.4B, up 142%. <em>Sources 4, 5, 7. Tier M.</em></p></li><li><p><strong>Q2 FY27 networking estimate.</strong> About $17.5B = 14.8 x 1.18. <em>Sources 3, 4. Tier M.</em></p></li><li><p><strong>Co-packaged optics claims.</strong> 200 Gb/s SerDes; 5x power efficiency, 5x uptime, 1.3x faster deployment. <em>Source 13. Tier B.</em></p></li><li><p><strong>Scale-out fabrics.</strong> Quantum-X800 InfiniBand and Spectrum-X Ethernet; ConnectX-9 SuperNICs and BlueField-4 DPUs in the rack. <em>Source 14. Tier B.</em></p></li><li><p><strong>Kyber.</strong> 600 kW per rack (GTC, March 2025); reported delay to 2028; 78-layer midplane; 144 against 72 packages; NVIDIA statement. <em>Sources 18, 19, 23. Tier C.</em></p></li></ul><h3>Gigawatts, power and cooling</h3><ul><li><p><strong>PORTS-Pike site.</strong> 4.25 IT-GW initial; option 3.75 IT-GW (press release), about 3.8 GW (CFO commentary); 8 IT-GW for OpenAI; 10 GW generation; $4.2B grid; $1.5B NVIDIA investment; phases from 2028. <em>Sources 1, 11. Tier A.</em></p></li><li><p><strong>NVIDIA revenue per generation at the site.</strong> About 1.5 million GPUs; $150B to $200B. <em>Source 1. Tier B.</em></p></li><li><p><strong>Per-gigawatt arithmetic.</strong> 35.3 to 47.1 $B/GW; 40 x 4.25 = 170; 352,941 GPUs/GW; 2.83 kW/GPU; $100,000 to $133,333 per GPU. <em>Source 1. Tier M.</em></p></li><li><p><strong>Revenue opportunity per gigawatt.</strong> Hopper $18B, Grace Blackwell $25B, Vera Rubin $40B. <em>Source 3. Tier B.</em></p></li><li><p><strong>Total investment per gigawatt.</strong> About $30B five years ago, about $60B today; 40 / 60 = 66.7%. <em>Source 3. Tiers B, M.</em></p></li><li><p><strong>Watts per GPU.</strong> Factory 2.50 kW (400,000 GPUs/GW); without full MaxLPS gain 3.50 kW; Ohio 2.83 kW, 88.2% of factory density. <em>Sources 1, 14, 16. Tier M.</em></p></li><li><p><strong>800 V DC.</strong> 54 V today; up to eight shelves; up to 64U; up to 200 kg; 85% more power; 45% less copper; up to 5% efficiency; timing tied to Kyber. <em>Source 17. Tier B.</em></p></li><li><p><strong>Current and loss arithmetic.</strong> 11,111 A at 54 V; 750 A at 800 V; 14.81x; 219.5x. <em>Sources 17, 19. Tier M.</em></p></li><li><p><strong>DSX MaxLPS.</strong> 45 C cooling; up to 40% more GPUs; break-even loss 28.6% = 1 - 1/1.4. <em>Source 16. Tiers B, M.</em></p></li><li><p><strong>45 C water and chillers.</strong> Rack inlet 45 C (product page); Huang at CES: twice the power of Grace Blackwell, same 45 C water, no chillers needed. <em>Sources 14, 27. Tiers B, C.</em></p></li><li><p><strong>Water flow.</strong> 100 kW / (4.18 x 10 K) at 0.990 kg/L = 145.0 L/min; at 20 K 72.5 L/min. <em>Source: textbook constants. Tier M.</em></p></li><li><p><strong>Australia.</strong> Up to 2 GW by 2027 with eight operators. <em>Source 20. Tier B.</em></p></li><li><p><strong>AI infrastructure by 2030.</strong> $3T to $4T, repeated at Goldman Sachs conference. <em>Source 21. Tier C.</em></p></li></ul><h3>Supply, memory costs and margins</h3><ul><li><p><strong>Supply-related commitments.</strong> Q3 FY26 $50.3B, Q4 FY26 $95.2B, Q1 FY27 $119B, Q2 FY27 $279B. <em>Sources 1, 4, 5, 6. Tier A.</em></p></li><li><p><strong>Supply commitments due by FY2029.</strong> $267B of $279B (95.7%). <em>Source 1. Tier M.</em></p></li><li><p><strong>Cancelability of supply agreements.</strong> Some agreements may be cancelable, rescheduled or adjusted before firm orders. <em>Source 9. Tier A.</em></p></li><li><p><strong>Cost of revenue at the Q3 guide.</strong> $28.08B = 108.0 x (1 - 0.74); 92 / 28.08 = 3.28 quarters; two quarters 56.16. <em>Source 1. Tier M.</em></p></li><li><p><strong>Inventory.</strong> Q3 FY26 $19.8B, Q4 FY26 $21.4B, Q1 FY27 $25.8B, Q2 FY27 $31.6B. <em>Sources 1, 4, 5, 6. Tier A.</em></p></li><li><p><strong>Gross margin path.</strong> Q4 FY27 trough 71% to 72%; FY28 72% to 73%. <em>Source 3. Tier B.</em></p></li><li><p><strong>GAAP gross margin history.</strong> Q3 FY25 74.6%, Q4 FY25 73.0%, Q1 FY26 60.5%, Q2 FY26 72.4%, Q3 FY26 73.4%, Q4 FY26 75.0%, Q1 FY27 74.9%, Q2 FY27 75.0%. <em>Sources 1, 4, 5, 6, 7, 8. Tier A.</em></p></li><li><p><strong>Cost per $100 of revenue.</strong> 25.00, 26.00, 28.50, 27.50; up 4%, 14%, 10%. <em>Sources 1, 3. Tier M.</em></p></li></ul><h3>Commitments, guarantees and financing</h3><ul><li><p><strong>Commitments.</strong> 279 + 29 + 25 + 25 + 8 = $366B. <em>Source 1. Tier A.</em></p></li><li><p><strong>Additional commitments.</strong> 36 + 20 = $56B. <em>Source 1. Tier A.</em></p></li><li><p><strong>Guarantees.</strong> 3.5 + 105.0 = $108.5B. <em>Source 1. Tier A.</em></p></li><li><p><strong>Total against liabilities and revenue.</strong> $530.5B; 5.81x of $91.3B; 1.75x of $303.0B trailing revenue. <em>Sources 1, 4, 5, 6. Tier M.</em></p></li><li><p><strong>Residual value guaranties.</strong> Cap $105B; effective at lease commencement; ready-for-service expected from 2028; triggers and remedies. <em>Source 10. Tier A.</em></p></li><li><p><strong>Guarantee cap per gigawatt.</strong> $24.7B = 105 / 4.25; 52.5% to 70.0% of one generation&#8217;s revenue. <em>Source 1. Tier M.</em></p></li><li><p><strong>Lab exposure.</strong> Nearly $50B invested; about 12 GW for OpenAI; nearly 2 GW credit enhancement for another lab; labs about a quarter of next year&#8217;s business. <em>Source 3. Tier B.</em></p></li><li><p><strong>Financing platforms.</strong> More than $500B of third-party capital; MOUs subject to final agreements. <em>Source 12. Tier B.</em></p></li><li><p><strong>AI cloud agreements.</strong> Upfront sale plus revenue share above a take-or-pay floor. <em>Sources 1, 3. Tier A.</em></p></li></ul><h3>Cash and balance sheet</h3><ul><li><p><strong>Cash flow.</strong> Q2 net income $59,688M, OCF $24,077M, FCF $21,341M; OCF / net income 40.3% against 86.3% in Q1. <em>Sources 1, 4. Tiers A, M.</em></p></li><li><p><strong>Working capital in Q2.</strong> Receivables $22,346M, inventories $5,784M, prepaid and other $5,497M; equity gains $7,771M. <em>Source 2. Tier A.</em></p></li><li><p><strong>Days sales outstanding.</strong> 60 against 45; check 63,059 / 96,221 x 91 = 59.6. <em>Source 1. Tiers A, M.</em></p></li><li><p><strong>Marketable equity and non-marketable securities.</strong> 42,783 + 51,157 = $93,940M against $35,137M at Jan 25; cash and marketable debt $56,586M. <em>Source 2. Tier M.</em></p></li><li><p><strong>Equity activity, first half.</strong> Purchases $42,404M; net gains $23,707M. <em>Source 2. Tier A.</em></p></li><li><p><strong>Debt.</strong> $33,366M against $8,468M; $25.0B notes issued in Q2. <em>Sources 1, 2. Tier A.</em></p></li><li><p><strong>Groq payment.</strong> $2,944M in financing activities. <em>Source 2. Tier A.</em></p></li><li><p><strong>Capital returns.</strong> 19,732 + 6,047 = $25,779M; dividend $0.01 to $0.25 (May 18); $99.0B authorization remaining. <em>Sources 2, 4. Tier A.</em></p></li><li><p><strong>Buyback authorization.</strong> $80.0B added on May 18. <em>Source 4. Tier A.</em></p></li></ul><h3>Dates</h3><ul><li><p><strong>Dates.</strong> Q3 call November 17; GTC Berlin October 21. <em>Source 3. Tier B.</em></p></li></ul>]]></content:encoded></item><item><title><![CDATA[Unleashing OpenCL’s secrets]]></title><description><![CDATA[A technical account of the standard, its implementations, and where it actually stands in 2026.]]></description><link>https://www.thesoftwarefrontier.com/p/unleashing-opencls-secrets</link><guid isPermaLink="false">https://www.thesoftwarefrontier.com/p/unleashing-opencls-secrets</guid><dc:creator><![CDATA[Lorenzo Bradanini]]></dc:creator><pubDate>Thu, 10 Sep 2026 12:40:18 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!dRZi!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F64dee902-2387-4937-a3ea-03c91d87e6b4_2004x878.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!dRZi!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F64dee902-2387-4937-a3ea-03c91d87e6b4_2004x878.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!dRZi!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F64dee902-2387-4937-a3ea-03c91d87e6b4_2004x878.png 424w, https://substackcdn.com/image/fetch/$s_!dRZi!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F64dee902-2387-4937-a3ea-03c91d87e6b4_2004x878.png 848w, https://substackcdn.com/image/fetch/$s_!dRZi!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F64dee902-2387-4937-a3ea-03c91d87e6b4_2004x878.png 1272w, https://substackcdn.com/image/fetch/$s_!dRZi!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F64dee902-2387-4937-a3ea-03c91d87e6b4_2004x878.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!dRZi!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F64dee902-2387-4937-a3ea-03c91d87e6b4_2004x878.png" width="1456" height="638" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/64dee902-2387-4937-a3ea-03c91d87e6b4_2004x878.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:638,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;OpenCL release timeline 2008 to 2026&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="OpenCL release timeline 2008 to 2026" title="OpenCL release timeline 2008 to 2026" srcset="https://substackcdn.com/image/fetch/$s_!dRZi!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F64dee902-2387-4937-a3ea-03c91d87e6b4_2004x878.png 424w, https://substackcdn.com/image/fetch/$s_!dRZi!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F64dee902-2387-4937-a3ea-03c91d87e6b4_2004x878.png 848w, https://substackcdn.com/image/fetch/$s_!dRZi!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F64dee902-2387-4937-a3ea-03c91d87e6b4_2004x878.png 1272w, https://substackcdn.com/image/fetch/$s_!dRZi!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F64dee902-2387-4937-a3ea-03c91d87e6b4_2004x878.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Eight releases in eighteen years.</strong> The shape of the chart is the argument: a fast build-out, an over-reach that nobody could implement in full, a reset that made almost everything optional, and a 2026 release that starts requiring things again.</figcaption></figure></div><p>In May 2026, on the <strong>eve of IWOCL in Heilbronn</strong>, the Khronos OpenCL Working Group released OpenCL 3.1. It is the first version bump in almost six years, and the first since 2017 that makes the specification <em>bigger</em> rather than more permissive. </p><p>Two months later the first conformant 3.1 implementation appeared on the Khronos list: Apple M1 and M2 graphics, running Linux, driven by Mesa&#8217;s Rusticl. </p><p>Not Apple&#8217;s driver: <strong>Apple deprecated OpenCL</strong> in 2018 and never looked back. A reverse-engineered kernel driver plus a Rust userspace led by one engineer at Red Hat.</p><p>That is a reasonable summary of what OpenCL has become. It is not the thing that lost to CUDA and quietly died. It is a specification with a governance process that keeps grinding forward, implemented today mostly by people who are <strong>not the vendors that originally wrote it</strong>, running the mobile inference stacks of two of the largest silicon companies on earth, and serving as the portability substrate underneath SYCL, chipStar, and a lot of code that will never mention OpenCL in its README.</p><p>There is one constraint that explains almost everything about OpenCL, and it is worth putting up front because the rest of this article is evidence for it.</p><p>OpenCL is the only compute API that had to describe hardware it could not see. CUDA describes one company&#8217;s roadmap, and that company knows what silicon is coming. <strong>Vulkan compute describes graphics hardware.</strong> Metal describes Apple&#8217;s. </p><p>OpenCL signed a contract covering CPUs, DSPs, FPGAs, GPUs from four vendors, and processors that did not exist when the contract was written. The price of that contract is that every guarantee has to be weak enough for the weakest member, or optional.</p><p>Once you see that, the history stops looking like a series of unforced errors. <strong>OpenCL 2.0 tried to make strong guarantees universal</strong>, and half the industry declined to implement them. OpenCL 3.0 stopped pretending and made almost everything optional, which read as capitulation and was actually an accurate description of what already existed. </p><p><strong>OpenCL 3.1, in May 2026</strong>, is the first release in the standard&#8217;s history that mostly ratifies what had already shipped rather than predicting what ought to. </p><p>The performance-portability problem is the same constraint wearing different clothes: a kernel is a description of a memory hierarchy, and OpenCL cannot tell you which one you have.</p><p>I&#8217;ve tried to make this concrete rather than rhetorical. Where the article makes a claim about behaviour, I actually ran it: the kernels here were compiled and executed,<em> the accuracy bounds</em> were measured against what an implementation actually delivers, and the compile-time and dispatch-overhead arguments have numbers attached. </p><p>Those measurements come from one CPU device and are labelled as such: they <strong>demonstrate mechanisms, not hardware. </strong></p><p>I&#8217;ve also been fussy about versions and dates, because a large fraction of what is written about OpenCL online describes the world of 2015 and says so nowhere.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2><span>The origins</span></h2><p>The immediate ancestors of OpenCL are worth naming because they explain its shape. </p><p>Between 2006 and 2008 there were at least four serious attempts to make GPUs programmable for non-graphics work: <strong>Stanford&#8217;s BrookGPU</strong> and its commercial descendant at AMD, ATI&#8217;s Close to Metal, NVIDIA&#8217;s CUDA (<em>released February 2007</em>), and IBM&#8217;s toolchain for the Cell Broadband Engine. Each was tied to one vendor&#8217;s hardware. </p><p>Apple, which at the time shipped machines with GPUs from both NVIDIA and ATI and had a strategic interest in not being locked to either, wanted one API that worked on both plus the CPU.</p><p>Apple developed an initial proposal internally, with Aaftab Munshi, who had edited the<strong> OpenGL ES 1.1 and 2.0 specifications</strong>, as spec editor, and refined it with technical teams at AMD, IBM, Qualcomm, Intel, and NVIDIA before submitting it to Khronos. </p><p>The <strong>Khronos Compute Working Group</strong> was formed on 16 June 2008. It finished the technical content of OpenCL 1.0 on 18 November 2008, and the specification was approved for public release on 8 December 2008. </p><p>Five months from working group formation to a ratified cross-vendor standard is fast by any standards-body measure, and it shows in the result: OpenCL 1.0 has the feel of a design that was mostly finished before the committee got involved.</p><p><strong>Apple demonstrated a beta at WWDC in June 2008</strong>, a grid of 64 emulated Apple II machines running on the CPU to show task parallelism, and an N-body simulation on a Mac Pro&#8217;s GPU to show data parallelism, and shipped the first production implementation in Mac OS X 10.6 Snow Leopard on 28 August 2009. </p><p><strong>AMD abandoned Close to Metal and backed OpenCL</strong>. NVIDIA announced OpenCL support for its GPU Computing Toolkit the day after ratification and shipped drivers in September 2009. IBM shipped an implementation for POWER through its XL compilers in October 2009. For about eighteen months, OpenCL looked like it was going to be the way GPUs got programmed.</p><p>The 1.x line then filled in the obvious gaps:</p><p><strong>OpenCL 1.1</strong> (14 June 2010) added three-component vector types, sub-buffers, rectangular region read/write/copy for buffers, user events, <code>async_work_group_strided_copy</code>, and tighter OpenGL interop through event linking. It also allowed API calls from multiple host threads, which 1.0 had not required.</p><p><strong>OpenCL 1.2</strong> (<em>15 November 2011</em>) is the version that still matters most, because it became the mandatory baseline of OpenCL 3.0 and is therefore the safe target for portable code even today. </p><p>It added device partitioning (splitting a device into sub-devices along compute-unit or cache-hierarchy boundaries), separate compilation and linking of programs, 1D <strong>images</strong> and 1D/2D image arrays, <strong>built-in kernels</strong> and <strong>custom devices</strong>, <code>clEnqueueMigrateMemObjects</code>, <code>clEnqueueFillBuffer</code>/<code>clEnqueueFillImage</code>, DirectX 9 media surface and DirectX 11 sharing, and <code>-cl-fp32-correctly-rounded-divide-sqrt</code> for code that needs IEEE 754 semantics on single-precision division and square root rather than the looser default.</p><p>By the end of 1.2 the API had a coherent shape. What happened next is the interesting part, and we&#8217;ll come back to it.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2><span>The four models</span></h2><p>The specification organises itself around four models, and it is worth internalising them in the order the spec gives them, because most confusion about OpenCL comes from conflating two of them.</p><h3>Platform model</h3><p>A <strong>host</strong> is connected to one or more <strong>compute devices</strong>. Each device has one or more <strong>compute units</strong>, each of which has one or more <strong>processing elements</strong>. That is the entire hardware abstraction, and it is deliberately vague. </p><p>The <em>mapping to real silicon</em> is vendor-defined and it&#8217;s not always intuitive: on an NVIDIA GPU a compute unit is a streaming multiprocessor; on AMD GCN it is a compute unit containing four 16-wide SIMDs; on a CPU it is typically a hardware thread; on Adreno it is a shader processor. </p><p><code>CL_DEVICE_MAX_COMPUTE_UNITS</code> therefore does not mean the same thing across vendors, and <strong>comparing it across vendors is meaningless</strong>. The number vendors quote in marketing material, &#8220;5888 cores&#8221;, usually counts processing elements or SIMD lanes, not compute units.</p><p>Devices are typed: <code>CL_DEVICE_TYPE_CPU</code>, <code>CL_DEVICE_TYPE_GPU</code>, <code>CL_DEVICE_TYPE_ACCELERATOR</code>, <code>CL_DEVICE_TYPE_CUSTOM</code> (1.2+), plus <code>CL_DEVICE_TYPE_DEFAULT</code> and <code>CL_DEVICE_TYPE_ALL</code>. </p><p>Custom devices are the interesting outlier: they are devices that do not support the <strong>OpenCL C</strong> programming language at all and expose only <strong>built-in kernels</strong>, fixed-function or firmware-defined entry points queried through <code>CL_DEVICE_BUILT_IN_KERNELS</code>. This is how OpenCL accommodates video encoders, ISPs, and fixed-function DSP blocks.</p><p>There is a trap here worth naming, because it caught me. Through OpenCL 1.2 and the 2.x line the specification defined <code>CL_DEVICE_TYPE_ALL</code> as every device <em>except</em> custom ones, and that is still what most tutorials say.</p><p>The current unified specification does not: <code>CL_DEVICE_TYPE_ALL</code> is now simply &#8220;<em>all OpenCL devices in the platform</em>&#8221;. The custom-device constraint that survives applies to <code>CL_DEVICE_TYPE_DEFAULT</code>, which must not be a custom device unless it is the only device in the platform. </p><p>If you are carrying a decade-old mental model of that enum, it is out of date.</p><h3>Profiles</h3><p>Cutting across all of this is a distinction the article has so far skipped and most desktop developers never meet. Every platform and every device reports either <code>FULL_PROFILE</code> or <code>EMBEDDED_PROFILE</code> through <code>CL_PLATFORM_PROFILE</code> and <code>CL_DEVICE_PROFILE</code>. </p><p>The <strong>embedded profile</strong> is a formally specified relaxation of the full one, written for hardware that cannot afford the full set of guarantees: it permits round-to-zero instead of round-to-nearest as the default rounding mode for single precision, allows denormals and infinities to be handled more loosely, relaxes the<strong> accuracy requirements on several built-ins</strong>, drops 64-bit integers to optional, and lowers the minimum image dimensions and object counts an implementation must support.</p><p>This matters because &#8220;<em>OpenCL 3.0 conformant</em>&#8221; on a phone or an automotive SoC does not automatically mean the same arithmetic as OpenCL 3.0 conformant on a workstation. </p><p>It is also why Qualcomm&#8217;s statement that current Snapdragon platforms support <em>OpenCL 3.0 full profile</em> is a specific and meaningful claim rather than <strong>marketing throat-clearing</strong>. Check <code>CL_DEVICE_PROFILE</code> before you assume anything about numerics on embedded hardware.</p><p>Devices are grouped under <strong>platforms</strong>. A platform corresponds roughly to one vendor&#8217;s implementation. A single machine routinely exposes three or four: an Intel GPU platform, an NVIDIA platform, a PoCL CPU platform, a Rusticl platform. The mechanism that lets them coexist is the ICD, described in section 4.</p><h3>Execution model</h3><p>A <strong>kernel</strong> is executed over an <strong>NDRange</strong>: a 1-, 2-, or 3-dimensional index space. Each point in the index space is a <strong>work-item</strong>, and every work-item runs the same kernel body with a different global ID. </p><p>Work-items are grouped into <strong>work-groups</strong> of a size the application chooses (<em>or leaves to the implementation</em>). </p><p>Work-items in a work-group can synchronise with a <strong>barrier</strong> and share <strong>local memory</strong>; work-items in different work-groups cannot synchronise at all, and the specification gives no guarantee about the order in which work-groups execute or whether they execute concurrently.</p><p>That last point is the single most important constraint in the execution model and it is routinely violated by people porting from CPU code. There is<strong> no legal way to spin-wait</strong> in one work-group for another work-group to make progress. On many implementations it will appear to work and then deadlock on a device with fewer compute units or a different scheduler.</p><p>Since OpenCL 2.1 (and via <code>cl_khr_subgroups</code> before that) there is a third level: the <strong>sub-group</strong>, a subdivision of a work-group that maps onto the hardware&#8217;s SIMD execution width, a warp on NVIDIA, a wavefront on AMD, a subgroup on Intel and Arm and Qualcomm. </p><p>Sub-groups are guaranteed to make forward progress independently of one another within a work-group, and they support collective operations (<em>broadcast, reduce, scan, ballot, shuffle</em>) that are dramatically cheaper than the local-memory equivalents, and which <strong>became core in OpenCL 3.1.</strong></p><p>Work is submitted through a <strong>command queue</strong> attached to one context and one device. Commands are kernel executions, memory transfers, map/unmap operations, markers, and barriers. </p><p>Each returns an <strong>event</strong>, and events form a dependency graph: any enqueue can take a wait-list of events that must complete first. Queues are in-order by default; <code>CL_QUEUE_OUT_OF_ORDER_EXEC_MODE_ENABLE</code> makes them <strong>out-of-order</strong>, at which point the event graph is the only thing constraining execution order.</p><p>The <strong>context</strong> is the object that ties devices, memory objects, programs, and queues together. Memory objects belong to a context, not a device, which is how the runtime knows it is allowed to migrate a buffer between devices in the same context.</p><h3>Memory model</h3><p>Four address spaces, in decreasing scope and increasing speed:</p><p><strong>SpaceQualifierScopeTypical hardware</strong>Global<code>__global</code>All work-items, all work-groups, host-visibleDevice DRAMConstant<code>__constant</code>Read-only, all work-itemsConstant cache / DRAMLocal<code>__local</code>One work-groupScratchpad (LDS / shared memory)Private<code>__private</code>One work-itemRegisters, spilling to scratch</p><p>Since OpenCL 2.0 there is optionally a fifth: the <strong>generic address space</strong>, which lets a pointer be resolved to global, local, or private at runtime, so you can write one function that operates on all three.</p><p>Consistency is relaxed and explicit. Within a work-item, memory is consistent in program order. Across work-items in a work-group, memory is consistent only at a <code>barrier()</code>. Through work-groups, there is <strong>no consistency guarantee</strong> during a kernel&#8217;s execution at all, only at kernel boundaries and at explicit synchronisation points defined by the host API. </p><p>OpenCL 2.0 layered a formal memory model on top of this, adapted from C11, with atomic operations parameterised by memory order and <strong>memory scope</strong> (<code>memory_scope_work_item</code>, <code>_work_group</code>, <code>_device</code>, <code>_all_svm_devices</code>), which section 6 covers in detail.</p><h3>Programming model</h3><p>Two are named in the spec: data parallel (the NDRange) and task parallel (<em>a kernel enqueued with a single work-item, or independent kernels in an out-of-order queue</em>). </p><p>The task-parallel model is largely vestigial on GPUs; it mattered for the Cell processor and for CPU devices, and <code>clEnqueueTask</code> was deprecated in 2.0 in favour of just enqueuing a one-element NDRange.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2><span>The API, and what it feels like to use</span></h2><p>OpenCL is a C API with an object model built on opaque handles and manual <strong>reference counting</strong>. Every object type has <code>clRetain*</code> and <code>clRelease*</code>. </p><p>Objects are freed when their count reaches zero, but the specification is careful to say that the implementation may keep an object alive as long as it is referenced by a queued command, so releasing a buffer that a running kernel still uses is legal.</p><p>The canonical bring-up sequence is longer than people expect:</p><pre><code><code>#define CL_TARGET_OPENCL_VERSION 300
#include &lt;CL/cl.h&gt;

cl_platform_id   platform;
cl_device_id     device;
cl_int           err;

clGetPlatformIDs(1, &amp;platform, NULL);
clGetDeviceIDs(platform, CL_DEVICE_TYPE_GPU, 1, &amp;device, NULL);

cl_context context = clCreateContext(NULL, 1, &amp;device, NULL, NULL, &amp;err);

/* clCreateCommandQueue is deprecated since 1.2; this is the 2.0+ form */
cl_queue_properties qprops[] = { CL_QUEUE_PROPERTIES,
                                 CL_QUEUE_PROFILING_ENABLE, 0 };
cl_command_queue queue =
    clCreateCommandQueueWithProperties(context, device, qprops, &amp;err);

const char *src = "__kernel void saxpy(float a,                      \n"
                  "                    __global const float *x,       \n"
                  "                    __global float *y) {           \n"
                  "    size_t i = get_global_id(0);                   \n"
                  "    y[i] = fma(a, x[i], y[i]);                     \n"
                  "}                                                  \n";

cl_program program = clCreateProgramWithSource(context, 1, &amp;src, NULL, &amp;err);
err = clBuildProgram(program, 1, &amp;device, "-cl-std=CL1.2", NULL, NULL);
if (err != CL_SUCCESS) {
    size_t log_size;
    clGetProgramBuildInfo(program, device, CL_PROGRAM_BUILD_LOG,
                          0, NULL, &amp;log_size);
    char *log = malloc(log_size);
    clGetProgramBuildInfo(program, device, CL_PROGRAM_BUILD_LOG,
                          log_size, log, NULL);
    fprintf(stderr, "%s\n", log);
}

cl_kernel kernel = clCreateKernel(program, "saxpy", &amp;err);

cl_mem xbuf = clCreateBuffer(context, CL_MEM_READ_ONLY  | CL_MEM_COPY_HOST_PTR,
                             n * sizeof(float), hx, &amp;err);
cl_mem ybuf = clCreateBuffer(context, CL_MEM_READ_WRITE | CL_MEM_COPY_HOST_PTR,
                             n * sizeof(float), hy, &amp;err);

float a = 2.0f;
clSetKernelArg(kernel, 0, sizeof(float),  &amp;a);
clSetKernelArg(kernel, 1, sizeof(cl_mem), &amp;xbuf);
clSetKernelArg(kernel, 2, sizeof(cl_mem), &amp;ybuf);

size_t global = n, local = 256;
cl_event ev;
clEnqueueNDRangeKernel(queue, kernel, 1, NULL, &amp;global, &amp;local, 0, NULL, &amp;ev);
clEnqueueReadBuffer(queue, ybuf, CL_TRUE, 0, n * sizeof(float), hy, 1, &amp;ev, NULL);
</code></code></pre><p>The compiler is the <strong>core part of the runtime</strong> here: kernels are shipped as source strings and built at application startup, which is a design decision with large consequences. </p><p>And <code>clSetKernelArg</code> is stateful and untyped: the argument index and size are passed by hand, and getting either wrong is a runtime error at best and silent corruption at worst. This is the single <strong>biggest ergonomic difference from CUDA&#8217;s </strong><code>&lt;&lt;&lt;&gt;&gt;&gt;</code><strong> syntax</strong>, which type-checks arguments at compile time because the kernel and the launch site are in the same translation unit.</p><p>The C++ bindings (<code>CL/opencl.hpp</code>, formerly <code>cl2.hpp</code>, formerly <code>cl.hpp</code>) fix most of this with RAII wrappers and a variadic <code>KernelFunctor</code>. They are a <em>Khronos-maintained header-only library</em>, not part of the core spec, and they are what you should actually use from C++.</p><p>Error handling deserves a note. Every function returns or writes a <code>cl_int</code>, and there are around sixty error codes. Two behaviours cause most of the confusion:</p><ul><li><p><strong>Asynchronous errors don&#8217;t surface where you expect.</strong> A kernel that reads out of bounds does not produce an error from <code>clEnqueueNDRangeKernel</code>; it produces <code>CL_OUT_OF_RESOURCES</code> or a device reset from a later <code>clFinish</code>, or nothing at all. Attach a context error callback (the <code>pfn_notify</code> parameter of <code>clCreateContext</code>): most implementations report useful diagnostics through it and almost nobody sets it.</p></li><li><p><code>CL_INVALID_WORK_GROUP_SIZE</code><strong> has many causes.</strong> The local size must divide the global size (unless the device supports <strong>non-uniform work-groups</strong>, an OpenCL 2.0 feature that is optional in 3.0 and queried via <code>CL_DEVICE_NON_UNIFORM_WORK_GROUP_SUPPORT</code>), must not exceed <code>CL_KERNEL_WORK_GROUP_SIZE</code> for that specific kernel on that device, and must match any <code>reqd_work_group_size</code> attribute in the kernel source.</p></li></ul><p>Deprecated APIs remain callable. The headers hide 1.2-deprecated entry points unless you define <code>CL_USE_DEPRECATED_OPENCL_1_2_APIS</code>, and similarly for earlier versions. </p><p>The <code>CL_TARGET_OPENCL_VERSION</code> macro (a three-digit value: <code>120</code>, <code>200</code>, <code>300</code>) controls which version&#8217;s prototypes the unified headers expose, and setting it deliberately is good hygiene: it turns &#8220;<em>this device doesn&#8217;t support that</em>&#8221; into the <strong>worst compile error.</strong></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2><span>The ICD, and how a call reaches a driver</span></h2><p>Nothing in the OpenCL specification requires that multiple vendors&#8217; implementations coexist on one machine, but in practice they must, and the mechanism is the <strong>Installable Client Driver</strong> loader.</p><p>An application links against <code>libOpenCL.so</code> (or <code>OpenCL.dll</code>, or <code>libOpenCL.dylib</code>). That library is not a driver. It is a dispatcher, maintained by Khronos in the <code>OpenCL-ICD-Loader</code> repository, which enumerates the vendor implementations installed on the system and forwards calls to the right one. </p><p>On Linux it does this by reading <code>/etc/OpenCL/vendors/*.icd</code>, each a one-line text file containing the path of a vendor&#8217;s shared object; on Windows it reads <code>HKEY_LOCAL_MACHINE\SOFTWARE\Khronos\OpenCL\Vendors</code>. </p><p>Two environment variables control the search: <code>OCL_ICD_VENDORS</code> replaces the default directory, and <code>OCL_ICD_FILENAMES</code> takes a separator-delimited list of extra ICDs to load, which are enumerated <em>before</em> anything found by the default mechanism. </p><p><strong>Both are invaluable</strong> when you want a test suite to hit one specific implementation on a machine that has four. Layers are selected separately, through <code>OPENCL_LAYERS</code>.</p><p>Mechanically, each vendor library exports <code>clGetExtensionFunctionAddress</code> and (per <code>cl_khr_icd</code>) <code>clIcdGetPlatformIDsKHR</code>. Every <code>cl_platform_id</code> the vendor returns points to a struct whose first member is a pointer to a <strong>dispatch table</strong>. </p><p><strong>Every other OpenCL object </strong>(<em>contexts, devices, queues, buffers</em>) is likewise required to begin with a pointer to the dispatch table of the implementation that created it. So <code>clEnqueueNDRangeKernel(queue, ...)</code> in the loader is one indirect call through <code>queue-&gt;dispatch-&gt;clEnqueueNDRangeKernel</code>. </p><p>This is why passing objects from one platform into a call on another platform is <strong>undefined behaviour</strong> rather than a clean error: the loader has already dispatched before anyone checks.</p><p>Version 2.0.0 of <code>cl_khr_icd</code> added <code>clIcdGetFunctionAddressForPlatformKHR</code> and related entry points to handle extension functions cleanly, which the older scheme did badly.</p><p>The loader also supports <strong>layers</strong>, a mechanism borrowed from Vulkan. A layer is a shared object that sits between the application and the ICD and can intercept, log, modify, or synthesise OpenCL calls. This is the foundation of several genuinely useful tools:</p><ul><li><p><strong>OpenCL Intercept Layer</strong> (<em>Intel, open source, version 3.0.4 as of 2026</em>): call tracing, timing, kernel dumping, device-side printf capture, injection of modified kernel source without rebuilding the application, and USM validity checking. Runs on Windows, Linux, macOS, Android, and FreeBSD.</p></li><li><p><strong>CLVizulayer</strong> (<em>StreamHPC, presented at IWOCL 2026</em>): emits the directed acyclic graph of device submissions in Graphviz DOT format. Unlike a timeline trace, this shows the <em>constraints</em> rather than the observed order, which is how you find the case where two commands ran sequentially because the implementation felt like it rather than because you asked for it. It has been used to inspect llama.cpp, Leela Chess Zero, GROMACS and LAMMPS.</p></li></ul><p>Layers are underused. If you are debugging a portability problem across four implementations, the ability to insert instrumentation without touching either the application or the driver is worth a great deal.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2><span>The execution model in depth</span></h2><h3>NDRange geometry</h3><p><code>clEnqueueNDRangeKernel</code> takes <code>work_dim</code> (1&#8211;3), a <code>global_work_offset</code> (added in 1.1; usually <code>NULL</code>), <code>global_work_size</code>, and <code>local_work_size</code>. </p><p>Inside the kernel:</p><pre><code><code>size_t gid   = get_global_id(0);      // global index, includes offset
size_t lid   = get_local_id(0);       // index within the work-group
size_t grp   = get_group_id(0);       // work-group index
size_t gsz   = get_global_size(0);
size_t lsz   = get_local_size(0);     // actual, may differ from enqueued
size_t ngrp  = get_num_groups(0);
size_t off   = get_global_offset(0);  // 1.1+
uint   dim   = get_work_dim();
</code></code></pre><p>Passing <code>NULL</code> for <code>local_work_size</code> lets the implementation choose. How much this matters is easy to measure, and the answer has two halves.</p><p>A <strong>6.8&#215; spread</strong> between the worst and best explicit choice. This is not a knob you can leave unconsidered; it is frequently the largest single factor in a kernel&#8217;s performance, ahead of most things people spend their time on.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!vGGU!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d398f70-f81d-4409-a341-41c538b345e4_2036x1230.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!vGGU!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d398f70-f81d-4409-a341-41c538b345e4_2036x1230.png 424w, https://substackcdn.com/image/fetch/$s_!vGGU!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d398f70-f81d-4409-a341-41c538b345e4_2036x1230.png 848w, https://substackcdn.com/image/fetch/$s_!vGGU!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d398f70-f81d-4409-a341-41c538b345e4_2036x1230.png 1272w, https://substackcdn.com/image/fetch/$s_!vGGU!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d398f70-f81d-4409-a341-41c538b345e4_2036x1230.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!vGGU!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d398f70-f81d-4409-a341-41c538b345e4_2036x1230.png" width="1456" height="880" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0d398f70-f81d-4409-a341-41c538b345e4_2036x1230.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:880,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Local work-group size on a memory-bound kernel&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Local work-group size on a memory-bound kernel" title="Local work-group size on a memory-bound kernel" srcset="https://substackcdn.com/image/fetch/$s_!vGGU!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d398f70-f81d-4409-a341-41c538b345e4_2036x1230.png 424w, https://substackcdn.com/image/fetch/$s_!vGGU!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d398f70-f81d-4409-a341-41c538b345e4_2036x1230.png 848w, https://substackcdn.com/image/fetch/$s_!vGGU!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d398f70-f81d-4409-a341-41c538b345e4_2036x1230.png 1272w, https://substackcdn.com/image/fetch/$s_!vGGU!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d398f70-f81d-4409-a341-41c538b345e4_2036x1230.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Local work-group size on a memory-bound kernel.</strong> The parameter is worth more than most optimisations people spend time on, and on this device the runtime&#8217;s automatic choice beat every value chosen by hand.</figcaption></figure></div><p>And yet the implementation&#8217;s own choice beat every value I picked by hand. <strong>I&#8217;d written flatly</strong> that passing <code>NULL</code> is wrong for code you care about, and on this device that advice is (sorry) simply false: the runtime knows the vectorisation width its own compiler chose and I do not. </p><p>The honest version is much narrower: the parameter matters enormously, <code>NULL</code><strong> is a real candidate</strong> rather than a lazy default, and the only way to know is to measure both on the hardware you ship to. </p><p>Where you do choose by hand, the size is bounded by three things you should query:</p><pre><code><code>size_t max_wg, pref_mult;
cl_ulong local_used, private_used;
clGetKernelWorkGroupInfo(kernel, device, CL_KERNEL_WORK_GROUP_SIZE,
                         sizeof(max_wg), &amp;max_wg, NULL);
clGetKernelWorkGroupInfo(kernel, device,
                         CL_KERNEL_PREFERRED_WORK_GROUP_SIZE_MULTIPLE,
                         sizeof(pref_mult), &amp;pref_mult, NULL);
clGetKernelWorkGroupInfo(kernel, device, CL_KERNEL_LOCAL_MEM_SIZE,
                         sizeof(local_used), &amp;local_used, NULL);
clGetKernelWorkGroupInfo(kernel, device, CL_KERNEL_PRIVATE_MEM_SIZE,
                         sizeof(private_used), &amp;private_used, NULL);
</code></code></pre><p><code>CL_KERNEL_WORK_GROUP_SIZE</code> is per-kernel, not per-device: it accounts for the register and local-memory pressure of <em>this</em> kernel and is therefore often much smaller than <code>CL_DEVICE_MAX_WORK_GROUP_SIZE</code>. </p><p><code>CL_KERNEL_PREFERRED_WORK_GROUP_SIZE_MULTIPLE</code> is the closest thing OpenCL 1.x has to a warp-size query: in practice it is 32 on NVIDIA, 64 on GCN and typically 32 on RDNA, 8, 16 or 32 on Intel depending on the SIMD width the compiler chose, and commonly 64 or 128 on Adreno.</p><p><strong>Those are the values you will usually see</strong>, not values the specification promises; query it rather than hard-coding it, and a local size that is not a multiple of it wastes lanes.</p><p><code>CL_KERNEL_PRIVATE_MEM_SIZE</code> is the register-spill indicator and the most useful single number for diagnosing a slow kernel. If it is non-zero and large, you are spilling to scratch memory and no amount of tuning the local size will help.</p><p>OpenCL 3.1 finally standardises what everyone had been approximating: a <strong>suggested local work-group size query</strong>, promoted to core from <code>cl_khr_suggested_local_work_size</code>. The runtime tells you what it thinks the right work-group size is for a given kernel, global size, and device. It is a hint, not an oracle, but it removes a class of per-device tuning tables from application code.</p><h3>Sub-groups</h3><pre><code><code>uint sg_size  = get_sub_group_size();
uint sg_id    = get_sub_group_id();
uint sg_local = get_sub_group_local_id();
uint n_sg     = get_num_sub_groups();

float total   = sub_group_reduce_add(x);
float bcast   = sub_group_broadcast(x, 0);
int   any     = sub_group_any(pred);
float shuf    = sub_group_shuffle(x, src_lane);       // cl_khr_subgroup_shuffle
float rot     = sub_group_rotate(x, delta);           // cl_khr_subgroup_rotate
uint4 mask    = sub_group_ballot(pred);               // cl_khr_subgroup_ballot
</code></code></pre><p>Sub-group operations avoid the round trip through local memory and the barrier that a work-group reduction requires. <strong>On a modern GPU</strong> a <code>sub_group_reduce_add</code> over 32 lanes is a handful of shuffle-and-add instructions; the local-memory equivalent is a store, a barrier, log&#8322;(n) rounds of load-add-store with a barrier each, and a load. </p><p>For the reduction-heavy inner loops in attention and normalisation kernels this is the difference between competitive and not.</p><p>Before 3.1 this was a portability minefield: <code>cl_khr_subgroups</code> was optional, the extended type support was a separate extension (<code>cl_khr_subgroup_extended_types</code>), shuffles were another (<code>cl_khr_subgroup_shuffle</code>, <code>cl_khr_subgroup_shuffle_relative</code>), ballots another, and Intel had its own pre-standard <code>cl_intel_subgroups</code>. </p><p>The llama.cpp OpenCL backend requires subgroup support and therefore requires OpenCL 2.x or a 3.0 implementation that opted in. OpenCL 3.1 makes sub-groups core, including shuffles, rotations, and an expanded set of supported data types, which is the change most likely to simplify real kernel code.</p><p>You can pin the <strong>sub-group</strong> size with <code>__attribute__((intel_reqd_sub_group_size(N)))</code> on Intel, or query <code>CL_KERNEL_MAX_SUB_GROUP_SIZE_FOR_NDRANGE</code> and <code>CL_KERNEL_SUB_GROUP_COUNT_FOR_NDRANGE</code> via <code>clGetKernelSubGroupInfo</code> (2.1+).</p><h3>Queues, events, and the flush trap</h3><p>An event has five states: <code>CL_QUEUED</code>, <code>CL_SUBMITTED</code>, <code>CL_RUNNING</code>, <code>CL_COMPLETE</code>, or a negative value indicating abnormal termination. With <code>CL_QUEUE_PROFILING_ENABLE</code> set, <code>clGetEventProfilingInfo</code> returns nanosecond timestamps for <code>CL_PROFILING_COMMAND_QUEUED</code>, <code>_SUBMIT</code>, <code>_START</code>, <code>_END</code>, and (2.0+) <code>_COMPLETE</code>. <code>END - START</code> is device execution time; <code>START - SUBMIT</code> is queue latency; <code>SUBMIT - QUEUED</code> is host-side runtime overhead. </p><p>All three are worth looking at separately, a kernel that is fast but submitted late is a different problem from a kernel that is slow. The three intervals are easy to see. Timing a trivial kernel over 65,536 work-items with the queue warm:</p><p><strong>Kernel work   QUEUED&#8594;SUBMIT    SUBMIT&#8594;START     START &#8594; END</strong></p><p>1 iteration                      0.07 &#181;s                   1.75 &#181;s                                 63.9 &#181;s     </p><p>50 iterations                  0.13 &#181;s                   2.66 &#181;s                                3.6 ms</p><p>1000 iterations              0.37 &#181;s                  13.2 &#181;s                                 89.0 ms</p><p><strong>Host-side runtime cost </strong>is sub-microsecond, queue latency is single-digit microseconds, and everything else is the device. That ordering is what you want, and it is why the interesting overhead is not inside a single enqueue but in how many of them you make and how you make them.</p><p>Which is measurable too. Dispatching an empty kernel repeatedly, comparing a round trip per dispatch against submitting a batch and finishing once:</p><p><strong>Dispatches         Round-trip                               Batched                Ratio  </strong></p><p>200                     7.28 &#181;s each                               2.00 &#181;s each         3.6&#215;</p><p>1000                   7.23 &#181;s each                               1.91 &#181;s each           3.8&#215;</p><p>5000                  7.22 &#181;s each                                1.85 &#181;s each           <strong>3.9&#215;</strong></p><p></p><p><em>PoCL 3.0, CPU device.</em></p><p><strong>Nearly 4&#215; on the same work</strong>, from nothing but submission discipline, and stable across three orders of magnitude of batch size. Note that this is measured against a CPU device with no PCIe bus and no kernel-mode driver round trip; on a discrete GPU the round-trip figure is worse, not better. </p><p>Around 2 &#181;s of irreducible per-dispatch cost is also the number that makes <strong>command buffers</strong> interesting: a transformer decode step issuing three hundred kernels per token is paying roughly 600 &#181;s of pure submission before any arithmetic happens, every token.</p><p>The <strong>classic OpenCL bug </strong>is forgetting <code>clFlush</code>. Enqueuing a command does not guarantee it is submitted to the device. If you enqueue work and then block on an event with <code>clWaitForEvents</code> without flushing the queue, some implementations will hang, because the command that would signal the event was never sent. The rules:</p><ul><li><p><code>clFlush</code><em> guarantees all previously enqueued commands are submitted to the device. It does not wait.</em></p></li><li><p><code>clFinish</code><em> blocks until all previously enqueued commands have completed. It implies a flush.</em></p></li><li><p><em>Blocking enqueue calls on a queue (</em><code>clEnqueueReadBuffer</code><em> with </em><code>blocking_read = CL_TRUE</code><em>) imply a flush of that queue. </em><code>clWaitForEvents</code><em> does not: the specification says the behaviour is undefined if you wait on events from commands that have not been flushed. This bites hardest across queues, since a blocking call on queue A flushes nothing in queue B. Flush explicitly.</em></p></li></ul><p>OpenCL 3.1 fixes a <strong>genuinely subtle related hazard</strong>. Previously, polling an event&#8217;s status with <code>clGetEventInfo</code> and observing <code>CL_COMPLETE</code> did <em>not</em> by itself establish the memory ordering needed to safely read the results; you were expected to call a waiting function. </p><p>In practice a great deal of code polled and then read, and it mostly worked. In 3.1, observing that an event has reached <code>CL_COMPLETE</code> is itself a synchronisation point. </p><p>That is a spec change that legalises what people were already doing, which is the right call, but it means code written against 3.1 semantics can be subtly broken on a 3.0 driver.</p><h3>Device-side enqueue, and why it did not take</h3><p>OpenCL 2.0 introduced <strong>device-side enqueue</strong>: a kernel could enqueue further kernels onto a device-side queue without host involvement, using blocks (the Clang/Apple <code>^{}</code> extension) as the payload. </p><p>It was the answer to <strong>CUDA Dynamic Parallelism</strong>, and it was intended for irregular workloads: adaptive mesh refinement, tree traversal, anything where the amount of work is data-dependent.</p><p>It never got broad adoption, unfortunately due to a lot of reasons. The implementation burden was high (<em>it requires a device-side scheduler</em>), the syntax was unfamiliar, and the performance on the implementations that did support it was frequently worse than doing multiple host-side dispatches. </p><p>In OpenCL 3.0 it became optional, queried through <code>CL_DEVICE_DEVICE_ENQUEUE_CAPABILITIES</code>, and many implementations report zero. </p><p>It&#8217;s the clearest example of a <strong>2.0 feature that was standardised</strong> before it was proven, which is exactly the mistake the working group says its current extension-first process is designed to avoid.</p><h3>Command buffers</h3><p>The replacement for a different problem, per-submission host overhead, is <code>cl_khr_command_buffer</code>, released provisionally in November 2021 as part of OpenCL 3.0.10 and developed largely by Codeplay with Qualcomm, Arm, Intel, Tampere University, NVIDIA and Google.</p><p>The idea is the one CUDA Graphs, <strong>Vulkan command buffers</strong>, and Level Zero command lists all landed on independently: record a sequence of commands once, finalise it, then dispatch the whole thing repeatedly with a single API call. </p><p>For inference serving, where the same graph of thirty or three hundred kernels is executed per token, the per-enqueue host cost dominates at small batch sizes.</p><pre><code><code>cl_command_buffer_khr cb =
    clCreateCommandBufferKHR(1, &amp;queue, NULL, &amp;err);

clCommandNDRangeKernelKHR(cb, NULL, NULL, k0, 1, NULL, &amp;g0, &amp;l0,
                          0, NULL, NULL, NULL);
clCommandNDRangeKernelKHR(cb, NULL, NULL, k1, 1, NULL, &amp;g1, &amp;l1,
                          0, NULL, NULL, NULL);
clFinalizeCommandBufferKHR(cb);

for (int i = 0; i &lt; steps; i++)
    clEnqueueCommandBufferKHR(0, NULL, cb, 0, NULL, NULL);
</code></code></pre><p>Note that a <strong>command buffer</strong> is internally out-of-order regardless of the queue it targets; ordering comes from the sync-point dependencies you declare when recording.</p><p><code>cl_khr_command_buffer_mutable_dispatch</code> <em>(OpenCL 3.0.12, September 2022)</em> relaxes the immutability constraint so that kernel arguments, global size, local size and offsets of a recorded dispatch can be updated between replays via <code>clUpdateMutableCommandsKHR</code>, which is what you need for anything with a changing sequence length. </p><p><strong>A further extension</strong> allows a command buffer to span multiple queues and devices (added in 3.0.14).</p><p>As of OpenCL 3.1 command buffers are still an extension, but Khronos names them explicitly as one of the features in the pipeline for a future core release.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2><span>The memory model in depth</span></h2><h3>Buffers, images, pipes</h3><p>A <strong>buffer</strong> (<code>cl_mem</code> created by <code>clCreateBuffer</code>) is a linear byte range. An <strong>image</strong> (<code>clCreateImage</code>) is an opaque object with a channel order, channel data type, and dimensionality, accessed through <code>read_imagef</code> / <code>write_imagef</code> and friends, optionally through a <strong>sampler</strong> that performs normalised coordinates, addressing modes (clamp, repeat, mirror), and linear filtering in fixed-function hardware. </p><p><strong>Pipes</strong> (2.0, optional in 3.0 via <code>CL_DEVICE_PIPE_SUPPORT</code>) are FIFO objects for producer-consumer patterns between kernels; they were designed with FPGAs in mind and are rarely used on GPUs.</p><p>The buffer-versus-image decision is more consequential than it looks. On desktop GPUs, <strong>buffer loads go through the general L1/L2 path</strong>; image reads go through the texture path, which on many architectures has a separate cache, hardware address computation, hardware boundary handling, and format conversion for free. </p><p>On mobile GPUs, Adreno in particular, the texture path is substantially faster for the read patterns typical of<strong> convolution and GEMM</strong>. This is why Qualcomm&#8217;s TVM and MLC work for Adreno has a dedicated &#8220;<em>texture path</em>&#8221; and specialised layouts, and why the llama.cpp Adreno kernels are written the way they are. On <strong>CPU devices </strong>the distinction mostly evaporates and images are usually slower.</p><p>Image support is itself optional: <code>CL_DEVICE_IMAGE_SUPPORT</code> can be <code>CL_FALSE</code>. It commonly is on custom devices and on some early open-source stacks, the <strong>original Mesa Clover </strong>never supported images, which is precisely why it could not run darktable, and why Rusticl&#8217;s image support was the thing that made it useful.</p><h3>Image formats and samplers</h3><p>An image is not a typed buffer; it is a channel order plus a channel data type, and only some combinations are required. The channel orders in the specification are <code>CL_R</code>, <code>CL_A</code>, <code>CL_RG</code>, <code>CL_RA</code>, <code>CL_RGB</code>, <code>CL_RGBA</code>, <code>CL_BGRA</code>, <code>CL_ARGB</code>, <code>CL_INTENSITY</code>, <code>CL_LUMINANCE</code>, <code>CL_DEPTH</code> and <code>CL_sRGBA</code>, among others. </p><p>The data types run from <code>CL_SNORM_INT8</code> and <code>CL_UNORM_INT8</code> through the packed short formats (<code>CL_UNORM_SHORT_565</code>, <code>CL_UNORM_SHORT_555</code>, <code>CL_UNORM_INT_101010</code>) to <code>CL_SIGNED_INT8/16/32</code>, <code>CL_UNSIGNED_INT8/16/32</code>, <code>CL_HALF_FLOAT</code> and <code>CL_FLOAT</code>, with <code>CL_UNORM_INT10</code>, <code>INT12</code> and <code>INT14</code> added for higher-precision sensor data.</p><p>The normalised types do conversion in hardware: reading a <code>CL_UNORM_INT8</code> image with <code>read_imagef</code> returns floats in [0, 1] with no instruction spent on the divide. That is <strong>free range conversion</strong>, and a real reason to prefer images for pixel data. </p><p>Only a small subset of order and type combinations is guaranteed, though; everything else must be checked with <code>clGetSupportedImageFormats</code> <em>against the specific device</em>, memory flags and image type. Assuming a format exists because it is in the enum is one of the more common portability failures.</p><p>Samplers carry three orthogonal settings: normalised or unnormalised coordinates, an addressing mode for out-of-range coordinates (<code>CLK_ADDRESS_NONE</code>, <code>CLAMP</code>, <code>CLAMP_TO_EDGE</code>, <code>REPEAT</code>, <code>MIRRORED_REPEAT</code>), and a filter mode (<code>CLK_FILTER_NEAREST</code> or <code>CLK_FILTER_LINEAR</code>). All of it is fixed-function on GPUs. </p><p>A bilinear tap that would cost four loads and three lerps in a buffer kernel is one <code>read_imagef</code> with <code>CLK_FILTER_LINEAR</code>, and boundary clamping that would cost a branch per axis is free. </p><p>Samplers can be declared in the kernel as a <code>const sampler_t</code> constant or created on the host and passed in.</p><h3>Allocation flags and the map path</h3><pre><code><code>CL_MEM_READ_WRITE | CL_MEM_WRITE_ONLY | CL_MEM_READ_ONLY   /* access, from the kernel's view */
CL_MEM_USE_HOST_PTR                                        /* use this host allocation */
CL_MEM_ALLOC_HOST_PTR                                      /* allocate host-accessible memory */
CL_MEM_COPY_HOST_PTR                                       /* allocate device memory, copy in */
CL_MEM_HOST_WRITE_ONLY | CL_MEM_HOST_READ_ONLY | CL_MEM_HOST_NO_ACCESS  /* 1.2 */
</code></code></pre><p>The semantics people get wrong: <code>CL_MEM_USE_HOST_PTR</code> doesn&#8217;t mean &#8220;<em>the device will read your pointer directly</em>&#8221;. </p><p>It means the <strong>implementation may cache the contents</strong> in device memory and is required to keep the host pointer as the backing store; you must map/unmap to access it safely from the host. </p><p><code>CL_MEM_ALLOC_HOST_PTR</code> is the closest OpenCL comes to CUDA&#8217;s pinned memory: it asks the implementation to allocate memory the host can access efficiently, which on a <strong>discrete GPU</strong> usually means page-locked system memory suitable for DMA, and on an integrated GPU usually means memory both processors can access without a copy at all.</p><p>It is worth being precise about what the flag buys, because it is easy to assume it makes transfers faster on its own. Writing 64 MB into a buffer three ways on the same device:</p><p><strong>Path                                                                                  Time                Eff. Rate </strong><code>clEnqueueWriteBuffer</code><em> into a plain buffer  </em>                    14.9 ms             4.5 GB/s</p><p><code>clEnqueueWriteBuffer</code><em> into an </em><code>ALLOC_HOST_PTR</code><em> buffer</em> 29.1 ms             2.3 GB</p><p><em>/smap / write / unmap on the </em><code>ALLOC_HOST_PTR</code><em> buffer</em>     <strong>5.2 ms               12.9 GB/s</strong></p><p><em><strong>PoCL 3.0, CPU device, mean of five.</strong></em></p><p>Adding <code>CL_MEM_ALLOC_HOST_PTR</code> and changing nothing else made the copy <strong>twice as slow</strong>. The <strong>2.9&#215;</strong> win only appears when the access pattern changes to match the allocation: when you stop copying and start writing directly into mapped memory. </p><p>The flag is not an optimisation; it is a request for memory with different properties, and it pays only if you then use those properties. This generalises: <strong>allocation hints in OpenCL</strong> are contracts about placement, and a contract you do not exercise is overhead.</p><p>The idiomatic zero-copy pattern on integrated hardware is therefore <code>CL_MEM_ALLOC_HOST_PTR</code> plus <code>clEnqueueMapBuffer</code>:</p><pre><code><code>cl_mem b = clCreateBuffer(ctx, CL_MEM_READ_WRITE | CL_MEM_ALLOC_HOST_PTR,
                          n, NULL, &amp;err);
float *p = clEnqueueMapBuffer(queue, b, CL_TRUE, CL_MAP_WRITE,
                              0, n, 0, NULL, NULL, &amp;err);
/* fill p */
clEnqueueUnmapMemObject(queue, b, p, 0, NULL, NULL);
</code></code></pre><p>On a system where host and device share physical memory this can be a genuine no-copy path. On a discrete GPU it still copies, but from pinned memory, at <strong>full PCIe bandwidth</strong> rather than the roughly half you get from pageable memory.</p><p>OpenCL 3.1 clarifies the semantics of <code>CL_DEVICE_HOST_UNIFIED_MEMORY</code> so that it can be used reliably to distinguish integrated from discrete devices.</p><p>That sounds trivial, but it&#8217;s not: for a decade the flag was defined loosely enough that <strong>implementations disagreed about it</strong>, and every serious application ended up with its own heuristic (<em>usually string-matching the device name</em>) to decide whether to take the map path or the copy path, so having one query that means one thing removes real code from real applications.</p><p><code>clEnqueueMigrateMemObjects</code> (1.2) lets you explicitly move a memory object to a device, or to the host with <code>CL_MIGRATE_MEM_OBJECT_HOST</code>, ahead of the kernel that will use it. In multi-device contexts this is how you avoid a fault-driven migration in the middle of a dispatch.</p><h3>The 2.0 memory model</h3><p>OpenCL 1.x had no formal memory model, just &#8220;<strong>relaxed consistency</strong>&#8220; plus barriers and fences, described in prose. OpenCL 2.0 replaced this with a model derived from C11&#8217;s, which was the right decision and made OpenCL one of the first GPU APIs with a mathematically specified memory model.</p><p>Atomics are typed (<code>atomic_int</code>, <code>atomic_float</code>, <code>atomic_uintptr_t</code>, and <code>atomic_long</code>/<code>atomic_double</code> if supported) and take a memory order and a <strong>memory scope</strong>:</p><pre><code><code>atomic_fetch_add_explicit(&amp;counter, 1,
                          memory_order_relaxed,
                          memory_scope_work_group);

atomic_store_explicit(&amp;flag, 1,
                      memory_order_release,
                      memory_scope_device);
</code></code></pre><p>Orders are <code>relaxed</code>, <code>acquire</code>, <code>release</code>, <code>acq_rel</code>, <code>seq_cst</code>. Scopes are <code>work_item</code>, <code>sub_group</code>, <code>work_group</code>, <code>device</code>, and <code>all_svm_devices</code>. The scope is the part with no C11 analogue and it is what makes the model usable on GPUs: an atomic scoped to a work-group can be implemented in the <strong>scratchpad</strong> with no cache-coherence traffic, while an atomic scoped to <code>all_svm_devices</code> may require flushing through to system memory.</p><p>The historical wrinkle is the <strong>inclusive scope rule</strong>. Until 2026, for two atomic operations to synchronise with each other, their scopes had to match, or more precisely, both had to include each other&#8217;s work-items. </p><p>This essentially meant that a release at <code>memory_scope_work_group</code> did not necessarily synchronise with an acquire at <code>memory_scope_device</code>, even though the device scope is strictly larger. </p><p>It is a rule that <strong>made implementations easier</strong> and reasoning harder, and it produced code that specified everything at device scope out of caution, giving up the performance that scoped atomics exist to provide.</p><p>OpenCL 3.1 relaxes it: scopes no longer have to match exactly, and a finer-grained scope can satisfy a coarser-grained synchronisation requirement. </p><p>Combined with sub-groups becoming core, this makes the fine-grained synchronisation patterns used in modern GPU kernels expressible without the previous defensive over-specification.</p><h3>Shared Virtual Memory, and the USM successor</h3><p>OpenCL 2.0 added SVM in three tiers, queried through <code>CL_DEVICE_SVM_CAPABILITIES</code>:</p><blockquote><p><em>Coarse-grained buffer: CL_DEVICE_SVM_COARSE_GRAIN_BUFFER.  Pointers are valid on both sides; consistency at map/unmap and kernel boundaries.</em></p><p><em>Fine-grained buffer: CL_DEVICE_SVM_FINE_GRAIN_BUFFER. Concurrent host/device access to the same allocation, no map needed.</em></p><p><em>Fine-grained system: CL_DEVICE_SVM_FINE_GRAIN_SYSTEM.           Any malloc pointer is usable on the device.</em></p><p><em>Atomics: CL_DEVICE_SVM_ATOMICS.                                                      Cross-side atomics on SVM allocations.</em></p></blockquote><p><strong>Coarse-grained buffer SVM is widely supported</strong>. Fine-grained system SVM is rare and requires real hardware support for demand paging and coherent access to host page tables.</p><p>The problem with SVM is not the concept but the API. <code>clSVMAlloc</code> returns a pointer, but you still set it as a kernel argument through <code>clSetKernelArgSVMPointer</code>, you still have to declare indirect accesses with <code>clSetKernelExecInfo</code> and <code>CL_KERNEL_EXEC_INFO_SVM_PTRS</code>, and the capability matrix is coarse. </p><p>Intel&#8217;s answer, <code>cl_intel_unified_shared_memory</code> (added to the registry alongside OpenCL 3.0.10), is a cleaner design borrowed from what became SYCL 2020&#8217;s USM: three allocation kinds (host, device, and shared) with explicit control over placement and migration, no map/unmap, allocations associated with both a device and a context, and a richer capability query. </p><p>It is the model that <strong>oneAPI, SYCL</strong>, and by extension a lot of HPC code actually use.</p><p>The working group has been standardising this as <code>cl_khr_unified_svm</code>, and Khronos lists <strong>Unified Shared Memory</strong> as one of the extensions in flight for a future core version. </p><p>For anyone writing new OpenCL against Intel or Level Zero-adjacent stacks, USM is already the pragmatic choice; for portable code, SVM coarse-grained remains the lowest common denominator, and plain buffers remain the <em>actual</em> lowest common denominator.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2><span>OpenCL C</span></h2><p><strong>OpenCL C is C99 with removals and additions.</strong> Removed: function pointers, recursion, variable-length arrays, bit fields, most of the standard library, <code>goto</code> into blocks, and (before 2.0) program-scope variables in the global address space. </p><p>Added: address space qualifiers, vector types, a large built-in function library, work-item query functions, and a set of type qualifiers and attributes.</p><p>Vector types exist in widths 2, 3, 4, 8, and 16 for all scalar types: <code>float4</code>, <code>int8</code>, <code>uchar16</code>, <code>double2</code>. They support arithmetic elementwise, and component access through several syntaxes:</p><pre><code><code>float4 v = (float4)(1.0f, 2.0f, 3.0f, 4.0f);
float  a = v.x;          /* also .y .z .w; and .r .g .b .a since OpenCL 3.0 */
float2 b = v.xy;
float4 c = v.wzyx;       /* arbitrary swizzle, including repeats */
float  d = v.s3;         /* hex-index form, required for width 8 and 16 */
float8 e = (float8)(v, v);
float2 lo = v.lo, hi = v.hi;   /* halves */
float2 ev = v.even, od = v.odd;
</code></code></pre><p>Three-component vectors have size 16 bytes, not 12: <code>sizeof(float3) == sizeof(float4)</code>. This surprises people writing structs that cross the host/device boundary.</p><p>Whether vector types help performance is architecture-dependent, and the answer has inverted over time. On <strong>AMD&#8217;s pre-GCN VLIW architectures</strong> (<em>TeraScale</em>) and on CPU devices, vector code was essential because the compiler could not always find the parallelism itself. </p><p>On GCN, RDNA, NVIDIA, and modern Intel GPUs, the hardware is scalar-per-lane and the vector types are mostly a way to express wide loads and stores. That is still worth something: a <code>float4</code> load is one 128-bit memory instruction instead of four 32-bit ones, which matters for memory-bound kernels. </p><p>On Adreno and Mali, vector width still maps to real SIMD capability and choosing it correctly matters more.</p><h3>Precision, and the parts nobody reads</h3><p>The specification contains a table of maximum error in <strong>ULP</strong> for every math built-in, and it is the most under-appreciated part of the document. Some entries:</p><p><strong>Function                                       Max error </strong></p><p><code>x + y</code>, <code>x * y</code>, <code>fma           </code>correctly rounded</p><p><code>1.0/x</code>, <code>x / y               </code>&#8804; 2.5 ULP; correctly rounded with <code>-cl-fp32-correctly-rounded-divide-sqrt</code></p><p><code>sqrt                      </code>&#8804; 3 ULP; correctly rounded with the same flag</p><p><code>rsqrt</code>, <code>cbrt</code>, <code>log1p          </code>&#8804; 2 ULP</p><p><code>exp</code>, <code>exp2</code>, <code>exp10</code>, </p><p><code>log</code>, <code>log2</code>, <code>log10            </code>&#8804; 3 ULP</p><p><code>sin</code>, <code>cos</code>, <code>sinpi</code>, </p><p><code>cospi</code>, <code>hypot               </code>&#8804; 4 ULP</p><p><code>tan</code>, <code>tanh</code>, <code>atan</code>, <code>atanh       </code>&#8804; 5 ULP</p><p><code>pow</code>, <code>pown</code>, <code>powr</code>, <code>rootn</code>, </p><p><code>erf</code>, <code>erfc</code>, <code>tgamma           </code>&#8804; 16 ULP</p><p><code>mad                       </code>unbounded: any value is conforming</p><p><code>native_* variants</code>                 implementation-defined, no bound</p><p><code>half_* variants</code>                     &#8804; 8192 ULP</p><p><strong>Two entries deserve attention</strong>. <code>mad</code> has no accuracy requirement at all: the specification permits any result, because it exists to let the implementation pick whatever multiply-add the hardware has, fused or not. </p><p>If you want a <strong>fused multiply-add</strong> with defined semantics, write <code>fma</code>. And in double precision, division, reciprocal and square root are all <em>correctly rounded</em>, the loose bounds above are a single-precision phenomenon.</p><p>What the table does not tell you is what you will actually get, and the gap between the two is larger than most people assume. Measuring the observed error of a <strong>conformant implementation</strong> against a double-precision reference over 65,536 points per function.</p><p>Two results are worth sitting with. <code>native_sqrt</code>, <code>native_exp</code> and <code>native_log</code> returned <strong>bit-identical values</strong> to their accurate counterparts: zero of 65,536 samples differed in any bit, so this implementation simply aliases them. </p><p><code>native_sin</code> is the exception, differing on 21% of samples. <code>half_exp</code>, permitted to be wrong by 8192 ULP, was accurate to under one. So on this implementation, every fast-math variant is free accuracy, and code written to tolerate 8192 ULP is running on results good to 1.</p><p><strong>That is not a reassuring finding</strong>. It means you cannot learn anything about <code>native_</code> behaviour by testing it, because the next implementation is equally entitled to return something wildly different and still be conformant. </p><p>This is the <strong>article&#8217;s opening constraint in miniature</strong>: a standard spanning unknown hardware can specify floors and nothing else, and a floor tells you almost nothing about the room. Test on the implementation you ship against, or use the accurate functions.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!FUmf!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff890265a-1ec0-4ea6-9595-9417525b95c2_2406x1455.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!FUmf!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff890265a-1ec0-4ea6-9595-9417525b95c2_2406x1455.png 424w, https://substackcdn.com/image/fetch/$s_!FUmf!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff890265a-1ec0-4ea6-9595-9417525b95c2_2406x1455.png 848w, https://substackcdn.com/image/fetch/$s_!FUmf!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff890265a-1ec0-4ea6-9595-9417525b95c2_2406x1455.png 1272w, https://substackcdn.com/image/fetch/$s_!FUmf!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff890265a-1ec0-4ea6-9595-9417525b95c2_2406x1455.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!FUmf!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff890265a-1ec0-4ea6-9595-9417525b95c2_2406x1455.png" width="1456" height="880" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f890265a-1ec0-4ea6-9595-9417525b95c2_2406x1455.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:880,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;What the specification permits, against what one implementation delivers&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="What the specification permits, against what one implementation delivers" title="What the specification permits, against what one implementation delivers" srcset="https://substackcdn.com/image/fetch/$s_!FUmf!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff890265a-1ec0-4ea6-9595-9417525b95c2_2406x1455.png 424w, https://substackcdn.com/image/fetch/$s_!FUmf!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff890265a-1ec0-4ea6-9595-9417525b95c2_2406x1455.png 848w, https://substackcdn.com/image/fetch/$s_!FUmf!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff890265a-1ec0-4ea6-9595-9417525b95c2_2406x1455.png 1272w, https://substackcdn.com/image/fetch/$s_!FUmf!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff890265a-1ec0-4ea6-9595-9417525b95c2_2406x1455.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>What the specification permits, against what one implementation delivers.</strong> The grey bars are conformance floors; the solid bars are measured. Every function lands well inside its bound, and the unbounded native_ variants were bit-identical to the accurate ones, which is exactly why you cannot generalise from it.</figcaption></figure></div><p>The <code>native_</code> family (<code>native_sin</code>, <code>native_exp2</code>, <code>native_recip</code>, <code>native_rsqrt</code>, <code>native_divide</code>) maps directly to whatever hardware instruction exists and gives no accuracy guarantee at all. </p><p><strong>On most GPUs these are single-instruction</strong> and roughly an order of magnitude faster than the accurate versions. <code>-cl-fast-relaxed-math</code> implies <code>-cl-finite-math-only</code>, <code>-cl-unsafe-math-optimizations</code>, and permits the compiler to substitute <code>native_</code> variants globally, which is why enabling it can change results by several ULP and can turn a NaN check into dead code.</p><p>Denormal handling is implementation-defined for single precision unless the device reports <code>CL_FP_DENORM</code> in <code>CL_DEVICE_SINGLE_FP_CONFIG</code>. <code>-cl-denorms-are-zero</code> explicitly permits flush-to-zero. </p><p><strong>Most GPUs flush single-precision denormals</strong> by default and handle double-precision denormals correctly, which is the opposite of what people assume.</p><p>Double precision requires <code>cl_khr_fp64</code> (or <code>__opencl_c_fp64</code> in OpenCL C 3.0), half precision requires <code>cl_khr_fp16</code>. Neither is guaranteed. <code>half</code> as a <em>storage</em> type, via <code>vload_half</code> / <code>vstore_half</code>, which convert to and from <code>float</code>, is available without <code>cl_khr_fp16</code>; only arithmetic on <code>half</code> requires the extension. </p><p>This is a useful distinction for anyone storing FP16 weights and computing in FP32.</p><h3>Attributes and hints</h3><pre><code><code>__attribute__((reqd_work_group_size(16, 16, 1)))
__attribute__((work_group_size_hint(64, 1, 1)))
__attribute__((vec_type_hint(float4)))
__attribute__((intel_reqd_sub_group_size(16)))    /* vendor */
</code></code></pre><p><code>reqd_work_group_size</code> is a contract: enqueue with any other local size and you get <code>CL_INVALID_WORK_GROUP_SIZE</code>. </p><p>It is worth using, because it lets the compiler size local arrays statically, unroll fully, and allocate registers knowing the <strong>occupancy</strong>, which frequently produces measurably better code than the hint version.</p><h3>OpenCL C 3.0 and feature macros</h3><p>OpenCL 3.0 made almost everything from the 2.x line optional. In the language, that <strong>optionality</strong> is expressed through predefined <strong>feature-test macros</strong> named <code>__opencl_c_&lt;feature&gt;</code>, which the compiler defines with value 1 when the feature is present:</p><pre><code><code>__opencl_c_3d_image_writes
__opencl_c_atomic_order_acq_rel
__opencl_c_atomic_order_seq_cst
__opencl_c_atomic_scope_device
__opencl_c_atomic_scope_all_devices
__opencl_c_device_enqueue
__opencl_c_fp64
__opencl_c_generic_address_space
__opencl_c_images
__opencl_c_int64
__opencl_c_pipes
__opencl_c_program_scope_global_variables
__opencl_c_read_write_images
__opencl_c_subgroups
__opencl_c_work_group_collective_functions
</code></code></pre><p>So portable OpenCL C 3.0 looks like this:</p><pre><code><code>#if defined(__opencl_c_subgroups)
    float total = sub_group_reduce_add(partial);
#elif defined(__opencl_c_work_group_collective_functions)
    float total = work_group_reduce_add(partial);
#else
    /* hand-rolled local memory tree reduction */
#endif
</code></code></pre><p>On the host side there is a matching set of device queries: <code>CL_DEVICE_ATOMIC_MEMORY_CAPABILITIES</code>, <code>CL_DEVICE_ATOMIC_FENCE_CAPABILITIES</code>, <code>CL_DEVICE_DEVICE_ENQUEUE_CAPABILITIES</code>, <code>CL_DEVICE_PIPE_SUPPORT</code>, <code>CL_DEVICE_GENERIC_ADDRESS_SPACE_SUPPORT</code>, <code>CL_DEVICE_WORK_GROUP_COLLECTIVE_FUNCTIONS_SUPPORT</code>, <code>CL_DEVICE_NON_UNIFORM_WORK_GROUP_SUPPORT</code>, and <code>CL_DEVICE_OPENCL_C_ALL_VERSIONS</code> / <code>CL_DEVICE_OPENCL_C_FEATURES</code>, so an application can decide which kernel variant to build before it builds anything. <code>clang</code> exposes the same switches for offline compilation through <code>-cl-ext</code>, e.g. <code>-cl-std=CL3.0 -cl-ext=+cl_khr_fp64,+__opencl_c_fp64</code>.</p><h3>C++, twice</h3><p>OpenCL 2.2 (May 2017) introduced <strong>OpenCL C++</strong>, a static subset of C++14 as a kernel language, defined alongside <strong>SPIR-V</strong> ingestion. It is a carefully written specification. It appears never to have shipped in a production driver, there is no extension defined to detect support for it, and it was deprecated in OpenCL 3.0.</p><p>The replacement is <strong>C++ for OpenCL</strong>, which is a different thing: not a Khronos-ratified specification but a community language, documented in the OpenCL-Docs repository and implemented in upstream Clang since release 9. </p><p>It is C++17 layered over OpenCL C, backward compatible with OpenCL C source, and it works with any implementation that ingests <strong>SPIR</strong>-V, which, as of OpenCL 3.1, is all of them. Version 1.0 was published in December 2020 (compatible with OpenCL 2.0); the 2021 revision (December 2021) is compatible with OpenCL 3.0.</p><p>What it doesn&#8217;t support: virtual functions, <code>dynamic_cast</code>, non-placement <code>new</code>/<code>delete</code>, exceptions, pointers to member functions, references to functions, and the C++ standard library. </p><p><em>What it does support is templates, classes, operator overloading, lambdas, and </em><code>auto</code>, which is enough for the thing people actually want, namely writing one templated kernel body instead of four macro-expanded copies. </p><p><strong>Address spaces are extended across C++ constructs</strong>: functional casts, templates, class members, references, lambdas, and operators.</p><pre><code><code>template&lt;typename T&gt;
class complex_t {
    T re, im;
public:
    complex_t(T r, T i) : re{r}, im{i} {}
    complex_t operator*(const complex_t &amp;o) const {
        return { re*o.re - im*o.im, re*o.im + im*o.re };
    }
    T real() const { return re; }
    T imag() const { return im; }
};

__kernel void mul_sp(__global float *in, __global float *out) {
    auto i = get_global_id(0);
    auto r = complex_t{in[4*i], in[4*i+1]} * complex_t{in[4*i+2], in[4*i+3]};
    out[2*i] = r.real(); out[2*i+1] = r.imag();
}
</code></code></pre><p>Online compilation of <strong>C++ for OpenCL</strong> requires the <code>cl_ext_cxx_for_opencl</code> extension and <code>-cl-std=CLC++</code>; Arm announced support in December 2020. In practice most people compile it offline with Clang to SPIR-V, which is the flow the language was designed for.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2><span>Compilation: source, IR, and machine code</span></h2><h3>The four ways to make a program object</h3><pre><code><code>clCreateProgramWithSource(ctx, count, strings, lengths, &amp;err);
clCreateProgramWithIL(ctx, il, length, &amp;err);              /* 2.1+, SPIR-V */
clCreateProgramWithBinary(ctx, n, devs, sizes, bins, st, &amp;err);
clCreateProgramWithBuiltInKernels(ctx, n, devs, names, &amp;err);  /* 1.2+ */
</code></code></pre><p>Source is the original path and still the most common. It means the OpenCL driver contains a full C compiler, that compiler runs at application startup, and compile time is user-visible. A large kernel library (<em>llama.cpp&#8217;s OpenCL backend, a serious image processing pipeline</em>) can take seconds to build.</p><p>It is worth knowing the size of that cost rather than assuming it. Building a synthetic kernel library on PoCL&#8217;s CPU device, then rebuilding the same library from the binaries the first build produced:</p><p><strong>Kernels in the librarySource size</strong><code>clCreateProgramWithSource</code><strong> + build</strong><code>clCreateProgramWithBinary</code><strong> + buildRatio</strong>83.5 KB158.0 ms2.2 ms72&#215;3214.1 KB190.4 ms5.2 ms37&#215;6428.2 KB276.4 ms13.8 ms20&#215;</p><p>The absolute numbers are one implementation&#8217;s, and PoCL is running a full Clang and LLVM pipeline where a vendor driver would be more streamlined. </p><p>The shape is the point: compiling from source is <strong>20&#215; to 72&#215;</strong> more expensive than loading a binary here, and the gap narrows only slowly as the library grows because a large part of the source cost is fixed compiler startup. </p><p>For an application with a hundred kernels this is the difference between an instant launch and a visible pause, which is exactly why every serious OpenCL application ends up building a cache, and exactly what the OpenCL 3.1 SPIR-V mandate is for.</p><p>Binaries via <code>clCreateProgramWithBinary</code> are the obvious fix and come with a hard constraint: <strong>program binaries are not portable</strong>. They are opaque, implementation-defined blobs, valid only for the device, driver version, and build options that produced them. <code>CL_PROGRAM_BINARY_TYPE</code> tells you whether a binary is an executable, a compiled object, or a library. </p><p>The correct use is a persistent cache keyed on device name, driver version, platform version, kernel source hash, and build options, exactly what llama.cpp does with <code>GGML_OPENCL_KERNEL_CACHE_DIR</code>, defaulting to <code>%LOCALAPPDATA%\llama.cpp\cl-cache</code> on Windows and <code>~/Library/Caches/llama.cpp/cl-cache</code> on macOS, and invalidating on any of those inputs changing.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!3V6F!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F38ac4406-dcf0-4dea-9dd4-813178eb17bd_2211x1253.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!3V6F!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F38ac4406-dcf0-4dea-9dd4-813178eb17bd_2211x1253.png 424w, https://substackcdn.com/image/fetch/$s_!3V6F!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F38ac4406-dcf0-4dea-9dd4-813178eb17bd_2211x1253.png 848w, https://substackcdn.com/image/fetch/$s_!3V6F!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F38ac4406-dcf0-4dea-9dd4-813178eb17bd_2211x1253.png 1272w, https://substackcdn.com/image/fetch/$s_!3V6F!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F38ac4406-dcf0-4dea-9dd4-813178eb17bd_2211x1253.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!3V6F!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F38ac4406-dcf0-4dea-9dd4-813178eb17bd_2211x1253.png" width="1456" height="825" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/38ac4406-dcf0-4dea-9dd4-813178eb17bd_2211x1253.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:825,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Two costs that are not the kernel&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Two costs that are not the kernel" title="Two costs that are not the kernel" srcset="https://substackcdn.com/image/fetch/$s_!3V6F!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F38ac4406-dcf0-4dea-9dd4-813178eb17bd_2211x1253.png 424w, https://substackcdn.com/image/fetch/$s_!3V6F!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F38ac4406-dcf0-4dea-9dd4-813178eb17bd_2211x1253.png 848w, https://substackcdn.com/image/fetch/$s_!3V6F!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F38ac4406-dcf0-4dea-9dd4-813178eb17bd_2211x1253.png 1272w, https://substackcdn.com/image/fetch/$s_!3V6F!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F38ac4406-dcf0-4dea-9dd4-813178eb17bd_2211x1253.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Two costs that are not the kernel.</strong> Build cost is the empirical case for shipping SPIR-V; per-dispatch cost is the case for command buffers. Both measured on a CPU device, where the dispatch figure is a floor rather than a typical value.</figcaption></figure></div><p>Separate compilation and linking arrived in 1.2:</p><pre><code><code>clCompileProgram(prog, n, devs, opts, n_hdrs, hdrs, hdr_names, cb, data);
cl_program lib = clLinkProgram(ctx, n, devs, "-create-library",
                               n_in, inputs, cb, data, &amp;err);
</code></code></pre><p>This lets you build a device-side library once and link it into several programs, and it lets you pass headers by name rather than concatenating strings.</p><p>Build options worth knowing:</p><pre><code><code>-cl-std=CL1.2 | CL2.0 | CL3.0 | CLC++ | CLC++2021
-D name=value                  -I dir
-cl-single-precision-constant  -cl-denorms-are-zero
-cl-opt-disable                -cl-mad-enable
-cl-no-signed-zeros            -cl-unsafe-math-optimizations
-cl-finite-math-only           -cl-fast-relaxed-math
-cl-fp32-correctly-rounded-divide-sqrt
-cl-uniform-work-group-size    -cl-kernel-arg-info
-w  -Werror
</code></code></pre><p><code>-cl-kernel-arg-info</code> is the one people forget: without it, <code>clGetKernelArgInfo</code> cannot tell you argument names, types, or address space qualifiers, which breaks any tooling that wants to bind arguments by name.</p><h3>SPIR, SPIR-V, and the 2026 mandate</h3><p>The first attempt at a portable IR was <strong>SPIR</strong> (2012, versions 1.2 and 2.0), which was LLVM IR with an OpenCL-specific metadata layer. It inherited LLVM&#8217;s problem: LLVM IR is not a stable format and its semantics are defined by the LLVM version that produced it. Consuming SPIR meant, in practice, having a compatible LLVM inside the driver.</p><p><strong>SPIR-V</strong>, released with OpenCL 2.1 in November 2015 and adopted by Vulkan 1.0 three months later, replaced it with a purpose-built, versioned, SSA-form binary IR that is not tied to any compiler. It is shared with Vulkan, and a SPIR-V module declares which <strong>execution environment</strong> it targets, the OpenCL SPIR-V Environment Specification defines what an OpenCL implementation must accept. </p><p>The OpenCL flavour uses the <code>OpenCL</code> memory model and the <code>Kernel</code> execution model, and calls into the <em>OpenCL Extended Instruction Set for SPIR-V</em> for math built-ins.</p><p>There are now several ways to produce it:</p><ul><li><p><strong>Clang&#8217;s SPIR-V target</strong>. <code>clang -target spirv64 -c kernel.cl -o kernel.spv</code>, using the SPIR-V backend that has been maturing in upstream LLVM. This is the path that no longer requires a translator.</p></li><li><p><strong>SPIRV-LLVM-Translator</strong>, the older, still widely used tool that converts LLVM IR to SPIR-V; the standard route for <code>clang -cl-std=CL3.0 -emit-llvm</code> output and for oneAPI&#8217;s toolchain.</p></li><li><p><strong>clspv</strong>. Google&#8217;s compiler from a subset of OpenCL C to <em>Vulkan</em> compute shaders, which is a different SPIR-V dialect. Paired with <strong>clvk</strong>, a runtime that implements the OpenCL API on top of Vulkan.</p></li></ul><p>OpenCL 3.1&#8217;s headline change is that <strong>SPIR-V ingestion is mandatory</strong>. Every conformant OpenCL 3.1 implementation must accept SPIR-V kernels through <code>clCreateProgramWithIL</code>, and must additionally support the SPIR-V query extension so applications can enumerate which SPIR-V capabilities, extensions, and versions a device handles.</p><p>This is more consequential than it sounds, and it is worth being precise about why. Before 3.1, <code>clCreateProgramWithIL</code> was core in 2.1 but optional in 3.0, so a 3.0 implementation could legally accept only source. Any tool that wanted to target <strong>OpenCL as a backend</strong> therefore had to either ship an OpenCL C source generator or accept that it would not run everywhere. </p><p>That is a real tax on SYCL implementations, on chipStar (<em>which compiles CUDA and HIP to SPIR-V</em>), on Julia&#8217;s and Rust&#8217;s GPU backends, and on every domain-specific compiler. Making ingestion mandatory turns OpenCL from &#8220;<em>an API you can target if the driver cooperates</em>&#8221; into a guaranteed compilation target. </p><p>Neil Trevett, who chairs the working group, called it the most consequential change in 3.1, and on the evidence that is not marketing.</p><p>The secondary benefits are the ones that matter operationally: kernels can ship pre-compiled and <strong>pre-optimised rather than as source</strong>, which removes startup compile cost, allows ahead-of-time specialisation, and means you are not shipping your kernel source to customers.</p><h3>What the vendor stacks actually do</h3><p><strong>ImplementationFront endIRBack end</strong>NVIDIAClang-derivedNVVM (LLVM IR)PTX &#8594; SASS via ptxasIntel (NEO / compute-runtime)ClangSPIR-VIGC &#8594; Gen ISAAMD (ROCm CLR)ClangLLVM IRAMDGPU back end &#8594; GCN/RDNA ISAArm MaliClang-basedvendor IRMali ISAQualcomm AdrenoLLVM-basedvendor IRAdreno ISAPoCLClangLLVM IR / SPIR-VLLVM target back ends, plus CUDA / Level Zero / remote driversRusticlvia SPIRV-Tools / MesaSPIR-V &#8594; NIRGallium driver back endsclvk / clspvclspv (Clang-based)Vulkan-flavour SPIR-Vwhatever the Vulkan driver does</p><p>Apple&#8217;s implementation, NVIDIA&#8217;s, AMD&#8217;s, RapidMind&#8217;s and Gallium&#8217;s have all been LLVM-based from early on. </p><p>The practical implication is that &#8220;<em>OpenCL C</em>&#8221; in the field means &#8220;<em>whatever Clang&#8217;s OpenCL front end accepts, plus vendor quirks</em>&#8221; more than it means the paper specification, which is mostly good, because <strong>Clang&#8217;s OpenCL support is well maintained</strong> and tracks the spec closely, and occasionally bad, because a kernel that builds on four Clang-derived stacks may still fail on a non-LLVM one.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2><span>Performance engineering</span></h2><h3>The mental model</h3><p>The abstraction says work-items are independent. The hardware executes them in lockstep groups. Everything about OpenCL performance follows from taking the second statement seriously while writing code that satisfies the first.</p><p>Concretely: <strong>a work-group is scheduled onto one compute unit</strong> and stays there. It is subdivided into sub-groups, which are the actual scheduling and execution unit. Divergent control flow within a sub-group is executed by predicating both sides. </p><p>Memory requests from a sub-group are <strong>coalesced</strong> by the hardware into as few transactions as possible; consecutive lanes reading consecutive addresses is one transaction, consecutive lanes reading strided addresses is many.</p><p>The occupancy story is the <strong>same as CUDA&#8217;s without the vocabulary</strong>: a compute unit has a fixed register file and a fixed scratchpad, and the number of work-groups it can host concurrently is bounded by whichever runs out first. </p><p>More concurrent work-groups means more latency hiding. <code>CL_KERNEL_LOCAL_MEM_SIZE</code> and <code>CL_KERNEL_PRIVATE_MEM_SIZE</code> are how you see the pressure; <code>CL_DEVICE_LOCAL_MEM_SIZE</code> and <code>CL_DEVICE_MAX_WORK_GROUP_SIZE</code> are the budget.</p><h3>A worked GEMM</h3><p>The single-work-item-per-output-element version is the reference, and it is memory bound at an <strong>arithmetic intensity</strong> of about 1 FLOP per byte:</p><pre><code><code>__kernel void sgemm_naive(const int M, const int N, const int K,
                          __global const float *A,
                          __global const float *B,
                          __global float *C)
{
    const int col = get_global_id(0);
    const int row = get_global_id(1);
    float acc = 0.0f;
    for (int k = 0; k &lt; K; k++)
        acc += A[row*K + k] * B[k*N + col];
    C[row*N + col] = acc;
}
</code></code></pre><p>Tiling through local memory raises intensity by the tile width. Each work-group cooperatively stages a <code>TS &#215; TS</code> tile of A and of B, barriers, then every work-item does <code>TS</code> multiply-accumulates out of the scratchpad:</p><pre><code><code>#define TS 16

__attribute__((reqd_work_group_size(TS, TS, 1)))
__kernel void sgemm_tiled(const int M, const int N, const int K,
                          __global const float *A,
                          __global const float *B,
                          __global float *C)
{
    const int lx = get_local_id(0), ly = get_local_id(1);
    const int col = get_group_id(0)*TS + lx;
    const int row = get_group_id(1)*TS + ly;

    __local float Asub[TS][TS];
    __local float Bsub[TS][TS];

    float acc = 0.0f;
    for (int t = 0; t &lt; K/TS; t++) {
        Asub[ly][lx] = A[row*K + (t*TS + lx)];
        Bsub[ly][lx] = B[(t*TS + ly)*N + col];
        barrier(CLK_LOCAL_MEM_FENCE);

        #pragma unroll
        for (int k = 0; k &lt; TS; k++)
            acc = fma(Asub[ly][k], Bsub[k][lx], acc);

        barrier(CLK_LOCAL_MEM_FENCE);
    }
    C[row*N + col] = acc;
}
</code></code></pre><p><code>TS</code> is 16 here rather than the textbook 32 for a reason: a 32&#215;32 work-group is 1024 work-items, which exceeds <code>CL_DEVICE_MAX_WORK_GROUP_SIZE</code> on a great deal of <strong>mobile and embedded hardware</strong> and sits at the ceiling on most desktop GPUs. Both kernels here also assume <code>M</code>, <code>N</code> and <code>K</code> are multiples of <code>TS</code>; real code needs edge handling or padded allocations.</p><p>Both barriers are required. The second one, after the inner loop and before the next tile overwrites the scratchpad, is the one people omit, and it produces a <strong>race that is invisible </strong>on any implementation where a work-group happens to be one sub-group wide.</p><p>The tiled version is usually still short of peak, because each work-item does one FMA per two scratchpad reads. The next step is <strong>register tiling</strong>: give each work-item a <code>WPT &#215; WPT</code> block of outputs so operands loaded into registers are reused across several accumulations.</p><pre><code><code>#define TS  32     /* tile size            */
#define WPT 4      /* outputs per work-item per dimension */
#define RTS (TS/WPT)

__attribute__((reqd_work_group_size(RTS, RTS, 1)))
__kernel void sgemm_regtiled(const int M, const int N, const int K,
                             __global const float *A,
                             __global const float *B,
                             __global float *C)
{
    const int lx = get_local_id(0), ly = get_local_id(1);
    const int gx = get_group_id(0)*TS, gy = get_group_id(1)*TS;

    __local float Asub[TS][TS];
    __local float Bsub[TS][TS];

    float acc[WPT][WPT];
    #pragma unroll
    for (int a = 0; a &lt; WPT; a++)
        #pragma unroll
        for (int b = 0; b &lt; WPT; b++) acc[a][b] = 0.0f;

    for (int t = 0; t &lt; K/TS; t++) {
        #pragma unroll
        for (int a = 0; a &lt; WPT; a++)
            #pragma unroll
            for (int b = 0; b &lt; WPT; b++) {
                const int r = ly*WPT + a, c = lx*WPT + b;
                Asub[r][c] = A[(gy + r)*K + t*TS + c];
                Bsub[r][c] = B[(t*TS + r)*N + gx + c];
            }
        barrier(CLK_LOCAL_MEM_FENCE);

        #pragma unroll
        for (int k = 0; k &lt; TS; k++) {
            float av[WPT], bv[WPT];
            #pragma unroll
            for (int a = 0; a &lt; WPT; a++) av[a] = Asub[ly*WPT + a][k];
            #pragma unroll
            for (int b = 0; b &lt; WPT; b++) bv[b] = Bsub[k][lx*WPT + b];
            #pragma unroll
            for (int a = 0; a &lt; WPT; a++)
                #pragma unroll
                for (int b = 0; b &lt; WPT; b++)
                    acc[a][b] = fma(av[a], bv[b], acc[a][b]);
        }
        barrier(CLK_LOCAL_MEM_FENCE);
    }

    #pragma unroll
    for (int a = 0; a &lt; WPT; a++)
        #pragma unroll
        for (int b = 0; b &lt; WPT; b++)
            C[(gy + ly*WPT + a)*N + gx + lx*WPT + b] = acc[a][b];
}
</code></code></pre><p>Now <code>WPT</code> loads of A and <code>WPT</code> of B feed <code>WPT&#178;</code> FMAs. With <code>WPT = 4</code> that is 16 FMAs per 8 scratchpad reads.</p><p>All three kernels above were compiled and run before publication, on PoCL 3.0 targeting a Xeon CPU device, and checked against a NumPy reference at 256&#215;256: all three match to <strong>zero relative error</strong>. That is a correctness check, not a performance claim, and the CPU result below is the reason to keep those two things apart.</p><p>There is no single correct <code>TS</code> and <code>WPT</code>. The optimum depends on scratchpad size, register file size, sub-group width, cache line size, and whether the compiler decides to spill. This is exactly what <strong>CLBlast</strong> does, a tuned BLAS in OpenCL that ships per-architecture parameter sets and can autotune on unseen hardware, and what <strong>CLTune</strong> and <strong>KTT</strong> exist to automate. </p><p>The honest position is that hand-written OpenCL GEMM is a tuning problem, not a coding problem, and that any claim about &#8220;OpenCL performance&#8221; that does not say which parameters were used is not a measurement.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Lc5J!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8d3bf495-aeaa-4f44-98ae-e76fb3ba402d_2071x1238.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Lc5J!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8d3bf495-aeaa-4f44-98ae-e76fb3ba402d_2071x1238.png 424w, https://substackcdn.com/image/fetch/$s_!Lc5J!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8d3bf495-aeaa-4f44-98ae-e76fb3ba402d_2071x1238.png 848w, https://substackcdn.com/image/fetch/$s_!Lc5J!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8d3bf495-aeaa-4f44-98ae-e76fb3ba402d_2071x1238.png 1272w, https://substackcdn.com/image/fetch/$s_!Lc5J!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8d3bf495-aeaa-4f44-98ae-e76fb3ba402d_2071x1238.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Lc5J!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8d3bf495-aeaa-4f44-98ae-e76fb3ba402d_2071x1238.png" width="1456" height="870" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/8d3bf495-aeaa-4f44-98ae-e76fb3ba402d_2071x1238.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:870,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Operand reuse in the three GEMM variants&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Operand reuse in the three GEMM variants" title="Operand reuse in the three GEMM variants" srcset="https://substackcdn.com/image/fetch/$s_!Lc5J!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8d3bf495-aeaa-4f44-98ae-e76fb3ba402d_2071x1238.png 424w, https://substackcdn.com/image/fetch/$s_!Lc5J!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8d3bf495-aeaa-4f44-98ae-e76fb3ba402d_2071x1238.png 848w, https://substackcdn.com/image/fetch/$s_!Lc5J!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8d3bf495-aeaa-4f44-98ae-e76fb3ba402d_2071x1238.png 1272w, https://substackcdn.com/image/fetch/$s_!Lc5J!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8d3bf495-aeaa-4f44-98ae-e76fb3ba402d_2071x1238.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Operand reuse in the three GEMM variants.</strong> Tiling alone does not raise arithmetic intensity; it relocates the traffic from global memory to the scratchpad. Reuse only improves once each work-item holds a block of outputs in registers.</figcaption></figure></div><p>Scratchpad <strong>bank conflicts</strong> are the other thing to watch. Local memory is divided into banks (typically 32 words wide); simultaneous accesses by different lanes to different addresses in the same bank serialise. </p><p>Declaring <code>__local float Asub[TS][TS+1]</code>, padding the row stride by one, breaks the conflicting stride pattern at a cost of a few hundred bytes.</p><h3>Everything that isn&#8217;t the kernel</h3><p>For inference workloads the kernel is frequently not the bottleneck. Intel&#8217;s IWOCL 2026 talk on optimising AI workloads on their GPUs spent most of its time on three things that are not kernel code: queue model design and how to batch versus submit immediately, <strong>Ultra Low Latency Submission</strong> to cut dispatch overhead for latency-sensitive inference, and driver-level memory pooling and resource recycling to avoid the allocation churn that dominates when a model allocates and frees repeatedly. </p><p>They reported allocation overhead reductions of orders of magnitude from pooling alone.</p><p>That matches what anyone who has profiled a token-by-token decode loop finds: at batch size 1, per-dispatch host cost and allocator behaviour can be a larger share of wall time than the arithmetic. It is why command buffers exist, and why they are the extension most worth watching.</p><h3>Asynchronous copies</h3><p>For embedded and DSP targets where the scratchpad is filled by an explicit DMA engine rather than by ordinary loads, OpenCL C has:</p><pre><code><code>event_t e = async_work_group_copy(dst_local, src_global, n, 0);
event_t f = async_work_group_strided_copy(dst_local, src_global, n, stride, e);
wait_group_events(1, &amp;f);
prefetch(ptr, n);
</code></code></pre><p><code>cl_khr_extended_async_copies</code> and <code>cl_khr_async_copy_fence</code>, added in OpenCL 3.0.10, extend this with 2D and 3D copy patterns and finer-grained fencing. </p><p>These were introduced specifically for the class of embedded processors that motivated much of the 3.0 work, and on a GPU they are usually implemented as ordinary loads.</p><h3>Performance portability, honestly</h3><p>The two studies everyone quotes are worth reading rather than citing, because the numbers travel without their conditions.</p><p><strong>Karimi, Dickson and Hamze at D-Wave</strong> measured a quantum Monte Carlo kernel in both languages and reported the OpenCL kernel between about 13% and 63% slower, with end-to-end times 16% to 67% slower. Those figures are exact, and they are also from 2010, on a GeForce GTX-260, with the CUDA and OpenCL toolkits both at version 2.3. </p><p>More importantly, the paper is explicit that its OpenCL kernel is a near-identical port of the CUDA one, changed only where OpenCL forced a change: <code>__shared__</code> to <code>__local</code>, <code>threadIdx</code> to <code>get_local_id()</code>, <code>__syncthreads()</code> to <code>barrier()</code>, and one array-indexing rewrite. </p><p>It measures what a direct translation costs, which is a real and useful thing to know, and it doesn&#8217;t measure what tuned OpenCL costs. The spread also narrows as problems get larger, which the authors attribute to the kernel&#8217;s share of total runtime rising.</p><p>The <strong>2011 Delft comparison</strong> is usually summarised as CUDA leading a straightforward OpenCL translation by at most 30% on NVIDIA hardware, with the gap attributed to programming-model differences and to NVIDIA having invested more in its CUDA compiler than its OpenCL one.</p><p> I have not read that paper in full and am repeating the published summary, which is exactly the sort of secondhand number this section is warning about.</p><p>The modern version of this result is more encouraging and comes from chipStar, which compiles unmodified CUDA and HIP to OpenCL and SPIR-V. </p><p>Its <strong>IWOCL 2026 keynote </strong>reported performance competitive with vendor-native toolchains across Intel discrete and integrated GPUs, AMD and NVIDIA GPUs through Rusticl, Arm Mali-G52, RISC-V systems with PowerVR graphics, and x86 and Arm CPUs, with overhead negligible on some platforms and a reasonable trade on others. </p><p>It was validated on real codes, including a quantum chemistry package with <strong>more than 20,000 lines of GPU kernels</strong> and the libCEED finite-element library.</p><p>The cheapest demonstration of the gap is the one sitting in this article. Running the three GEMM kernels above on a CPU device (PoCL&#8217;s pthread driver on a Xeon, where <code>__local</code> memory is ordinary RAM and there is no scratchpad to tile into), the ranking inverts completely.</p><p>Every optimisation that would win on a GPU loses here, consistently: the tiled kernel by 1.45&#215; to 1.78&#215;, the register-tiled one by 1.14&#215; to 1.35&#215;, at every size tested. </p><p>Staging a tile into <code>__local</code> is a copy the CPU did not need, and the barriers are pure overhead when a work-group is executed by one thread. </p><p><strong>This is one device and it proves nothing about any GPU</strong>. What it does show is that the tuning is not incidental to the kernel: it <em>is</em> the kernel, and it is aimed at a memory hierarchy the target may not have.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!TbYy!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc4e18a50-2437-4bc9-9d7a-c030fd01e890_2048x1244.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!TbYy!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc4e18a50-2437-4bc9-9d7a-c030fd01e890_2048x1244.png 424w, https://substackcdn.com/image/fetch/$s_!TbYy!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc4e18a50-2437-4bc9-9d7a-c030fd01e890_2048x1244.png 848w, https://substackcdn.com/image/fetch/$s_!TbYy!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc4e18a50-2437-4bc9-9d7a-c030fd01e890_2048x1244.png 1272w, https://substackcdn.com/image/fetch/$s_!TbYy!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc4e18a50-2437-4bc9-9d7a-c030fd01e890_2048x1244.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!TbYy!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc4e18a50-2437-4bc9-9d7a-c030fd01e890_2048x1244.png" width="1456" height="884" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c4e18a50-2437-4bc9-9d7a-c030fd01e890_2048x1244.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:884,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;The same three kernels on a CPU device&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="The same three kernels on a CPU device" title="The same three kernels on a CPU device" srcset="https://substackcdn.com/image/fetch/$s_!TbYy!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc4e18a50-2437-4bc9-9d7a-c030fd01e890_2048x1244.png 424w, https://substackcdn.com/image/fetch/$s_!TbYy!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc4e18a50-2437-4bc9-9d7a-c030fd01e890_2048x1244.png 848w, https://substackcdn.com/image/fetch/$s_!TbYy!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc4e18a50-2437-4bc9-9d7a-c030fd01e890_2048x1244.png 1272w, https://substackcdn.com/image/fetch/$s_!TbYy!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc4e18a50-2437-4bc9-9d7a-c030fd01e890_2048x1244.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>The same three kernels on a CPU device.</strong> Measured for this article on PoCL 3.0. The ordering that holds on a GPU inverts: staging a tile into local memory is a copy a CPU never needed, and the barriers are pure overhead.</figcaption></figure></div><p>So: OpenCL gives you <strong>functional portability</strong> for free and <strong>performance portability</strong> only if you do the work. A kernel tuned for a 64-wide wavefront and a 64 KB scratchpad will run correctly and badly on a device with a 16-wide sub-group and 32 KB, or on a CPU with neither.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share&quot;,&quot;text&quot;:&quot;Share The Software Frontier&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share The Software Frontier</span></a></p><div><hr></div><h2><span>The 2.x detour and the 3.0 reset</span></h2><p>OpenCL 2.0 (18 November 2013) was an ambitious release. It added <strong>shared virtual memory</strong>, device-side enqueue, the C11-derived memory model with scoped atomics, the <strong>generic address space</strong>, <strong>pipes</strong>, program-scope global variables, and read-write images. </p><p>Every one of these was a reasonable answer to a real limitation. Together they were too much to ask of every implementer at once.</p><p>What followed is the part that shaped the next decade. AMD and Intel implemented 2.0. NVIDIA did not: it stayed on OpenCL 1.2 as its <strong>conformance</strong> level from 2015 until 2021, offering a partial 2.0 evaluation driver from February 2017 that was never conformant. Mobile vendors implemented selectively. </p><p>So a developer writing OpenCL faced a choice between targeting 1.2 and reaching everyone, or targeting 2.0 and excluding the largest installed base of discrete GPUs. </p><p>Almost everyone chose 1.2, which meant SVM, device-side enqueue, and the new memory model went largely unused, which meant there was little pressure on anyone to implement them.</p><p>OpenCL 2.1 (<em>16 November 2015</em>) added SPIR-V ingestion, sub-groups, <code>clCloneKernel</code>, low-latency device timer queries, and the OpenCL C++ kernel language. OpenCL 2.2 (16 May 2017) brought OpenCL C++ into core along with SPIR-V 1.2 and pipe storage. Both landed on an ecosystem that had not adopted 2.0, and neither changed that.</p><p>At SIGGRAPH 2017 Khronos announced that OpenCL would converge with Vulkan. This was widely reported as OpenCL being merged into Vulkan and discontinued, which was not quite what was said but was a reasonable reading of the slides. </p><p>The actual outcome was different and, in retrospect, better: the convergence happened at the <em>IR</em> level, through SPIR-V, and produced clspv and clvk, a compiler and runtime that let OpenCL C kernels execute on Vulkan drivers. Adobe used exactly this to ship <strong>Premiere Rush on Android.</strong> Meanwhile OpenCL kept its own roadmap under the working title &#8220;OpenCL Next&#8221;.</p><p>That became OpenCL 3.0, provisional on 27 April 2020 and final on 30 September 2020. Its central move was to invert the compatibility model: <strong>OpenCL 1.2 becomes the mandatory baseline, and every 2.x feature becomes optional and queryable.</strong></p><p>The immediate effect was that vendors who had been stuck could ship a &#8220;<em>current</em>&#8221; version. NVIDIA became OpenCL 3.0 conformant on Windows and Linux with the R465 driver in April 2021, covering Maxwell and later, and <strong>exposing a specific set of optional pieces</strong>: RGBA vector component naming, the <code>pragma unroll</code> hint, <code>opencl_3d_image_writes</code>, the <code>clCreate*WithProperties</code> entry points, <code>clSetContextDestructorCallback</code>, <code>clCloneKernel</code>, and <code>clEnqueueSVMMigrateMem</code>. </p><p><strong>Intel&#8217;s NEO compute-runtime</strong> had shipped OpenCL 3.0 for Tiger Lake from version 20.41 in October 2020, including complete 2.0 and 2.1 optional functionality and parts of 2.2.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!ICzn!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F14c65cfc-1409-4033-aab0-693037356d38_2278x1421.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!ICzn!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F14c65cfc-1409-4033-aab0-693037356d38_2278x1421.png 424w, https://substackcdn.com/image/fetch/$s_!ICzn!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F14c65cfc-1409-4033-aab0-693037356d38_2278x1421.png 848w, https://substackcdn.com/image/fetch/$s_!ICzn!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F14c65cfc-1409-4033-aab0-693037356d38_2278x1421.png 1272w, https://substackcdn.com/image/fetch/$s_!ICzn!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F14c65cfc-1409-4033-aab0-693037356d38_2278x1421.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!ICzn!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F14c65cfc-1409-4033-aab0-693037356d38_2278x1421.png" width="1456" height="908" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/14c65cfc-1409-4033-aab0-693037356d38_2278x1421.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:908,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;What each version requires of a conformant implementation&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="What each version requires of a conformant implementation" title="What each version requires of a conformant implementation" srcset="https://substackcdn.com/image/fetch/$s_!ICzn!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F14c65cfc-1409-4033-aab0-693037356d38_2278x1421.png 424w, https://substackcdn.com/image/fetch/$s_!ICzn!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F14c65cfc-1409-4033-aab0-693037356d38_2278x1421.png 848w, https://substackcdn.com/image/fetch/$s_!ICzn!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F14c65cfc-1409-4033-aab0-693037356d38_2278x1421.png 1272w, https://substackcdn.com/image/fetch/$s_!ICzn!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F14c65cfc-1409-4033-aab0-693037356d38_2278x1421.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>What each version requires of a conformant implementation.</strong> OpenCL 3.0 did not remove the 2.x features. It moved them out of the mandatory set and gave applications a query for each one. OpenCL 3.1 pulls a different group back in.</figcaption></figure></div><p>The so called criticism of 3.0, that &#8220;<em>everything is optional</em>&#8221; means the standard guarantees nothing, is half right. It is true that a conformant OpenCL 3.0 device may expose exactly the OpenCL 1.2 feature set and nothing more. It is also true that <strong>this describes what the ecosystem already was</strong>; 3.0 made the situation legible instead of pretending otherwise. </p><p>The query surface it introduced is <strong>genuinely usable</strong>, and the alternative, a specification that most vendors ignore, is worse.</p><p>There are real backward-compatibility traps, though, and they are not always obvious. Program-scope global variables are an example: they were introduced in OpenCL C 2.0, and under 3.0 they are optional (<code>__opencl_c_program_scope_global_variables</code>). </p><p>Code that built fine against a vendor&#8217;s OpenCL 2.0 compiler can silently produce wrong results, not a build error, on the same vendor&#8217;s 3.0 driver if the driver drops the feature and the compiler&#8217;s handling of the declaration changes. </p><p>Users of PyOpenCL&#8217;s <code>ElementwiseKernel</code> <strong>hit precisely this on NVIDIA&#8217;s R465 driver.</strong> &#8220;<em>OpenCL 1.2 applications run unchanged on OpenCL 3.0</em>&#8221; is true. &#8220;<em>OpenCL 2.x applications run unchanged</em>&#8221; is true only if the driver still supports every 2.x feature they use, and you must check.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/p/unleashing-opencls-secrets/comments&quot;,&quot;text&quot;:&quot;Leave a comment&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/p/unleashing-opencls-secrets/comments"><span>Leave a comment</span></a></p><div><hr></div><h2><span>What OpenCL 3.1 changed</span></h2><p>Released 4&#8211;5 May 2026; specification revision 3.1.1 dated 22 May 2026. The <strong>design philosophy is stated explicitly</strong> by the working group and is worth quoting in substance: features are proven in the field as extensions first, watched across multiple implementations, refined on developer feedback, and only then promoted to core. Everything mandated in 3.1 was already shipping somewhere.</p><p><strong>Mandatory SPIR-V ingestion</strong>, plus mandatory support for the SPIR-V query extension. Covered in section 8; this is the change with the largest downstream effect.</p><p><strong>Sub-groups in core</strong>, including shuffles, rotations, and an expanded set of supported data types. Applications no longer need extension guards or fallback paths for the collective operations that tuned reductions, scans, and matrix kernels are built from.</p><p><strong>Integer dot products in core</strong>, including saturating and accumulating variants, together with extended bit operations. Both map to dedicated instructions on a wide range of modern silicon, the <code>dp4a</code>-class instructions, and both are the arithmetic primitives underneath INT8 inference. Promoted from <code>cl_khr_integer_dot_product</code> and <code>cl_khr_extended_bit_ops</code>.</p><p><strong>A suggested local work-group size query in core</strong>, from <code>cl_khr_suggested_local_work_size</code>.</p><p><strong>A standard device UUID query in core</strong>, matching Vulkan&#8217;s <code>VkPhysicalDeviceIDProperties::deviceUUID</code>. This lets an application correlate the same physical device across OpenCL and Vulkan, which is required for external memory sharing and for any sane device selection policy on a multi-GPU box. Promoted from <code>cl_khr_device_uuid</code>.</p><p><code>printf</code><strong> gains </strong><code>z</code><strong> and </strong><code>t</code><strong> length modifiers</strong>, for <code>size_t</code> and <code>ptrdiff_t</code>. Device-side printf could not previously format pointer-sized values without casts or format-string tricks, a small thing that shows up constantly in debugging.</p><p><code>CL_DEVICE_HOST_UNIFIED_MEMORY</code><strong> semantics clarified</strong> so it reliably distinguishes integrated from discrete GPUs.</p><p><strong>Local memory kernel arguments may be set to zero</strong>, meaning <em>&#8220;no local memory needed&#8221;</em>. Kernels that opportunistically use the scratchpad no longer need a separate code path for the configuration where they don&#8217;t.</p><p><strong>Observing </strong><code>CL_COMPLETE</code><strong> is now a synchronisation point.</strong></p><p><strong>The inclusive scopes rule is relaxed</strong> in the memory model, so a finer-grained scope can satisfy a coarser-grained requirement.</p><p>Implementations were in flight at release from Arm, Imagination, Intel and Qualcomm, plus Rusticl, PoCL and clvk. The first conformant listing arrived on 14 July 2026: <strong>Rusticl on Apple M1/M2 hardware</strong> under Asahi Linux, with radeonsi and Zink submissions in progress. </p><p>Mesa 26.2, released 5 August 2026, ships OpenCL 3.1 support for Rusticl on Asahi, Iris, radeonsi, llvmpipe and Zink, along with <code>cl_khr_subgroup_rotate</code> on radeonsi and Iris and other subgroup extensions.</p><p>The forward roadmap named by Khronos: command buffers, <strong>unified shared memory</strong>, <strong>cooperative matrix</strong> operations, low-precision AI data types including int4 and fp8, improvements to external memory sharing, and image tiling controls. </p><p>Beyond extensions, the group says it is exploring OpenCL&#8217;s role as a substrate for higher-level programming models, in safety-critical markets, and on NPUs and RISC-V accelerators.</p><div class="directMessage button" data-attrs="{&quot;userId&quot;:201922174,&quot;userName&quot;:&quot;Lorenzo Bradanini&quot;,&quot;canDm&quot;:null,&quot;dmUpgradeOptions&quot;:null,&quot;isEditorNode&quot;:true}" data-component-name="DirectMessageToDOM"></div><div><hr></div><h2><span>Extensions</span></h2><p>Extensions are named <code>cl_khr_*</code> (ratified, cross-vendor), <code>cl_ext_*</code> (multi-vendor but not ratified), or <code>cl_&lt;vendor&gt;_*</code>. A device advertises them in <code>CL_DEVICE_EXTENSIONS</code> as a space-separated string, or since 3.0 in the structured <code>CL_DEVICE_EXTENSIONS_WITH_VERSION</code>. Language-level extensions are enabled in the kernel with <code>#pragma OPENCL EXTENSION cl_khr_fp16 : enable</code>.</p><p>Provisional extensions carry version numbers below 1.0 and are subject to change; <code>cl_khr_command_buffer</code> spent years at 0.9.x. Some are shipped behind experimental headers. </p><p>That is a<strong> feature of the process</strong>: it is how features get tested before they are mandated. But it means &#8220;supports <code>cl_khr_command_buffer</code>&#8220; is not a stable claim without a version.</p><p>The families worth knowing:</p><p><strong>Numeric and language.</strong> <code>cl_khr_fp16</code>, <code>cl_khr_fp64</code>, <code>cl_khr_int64_base_atomics</code>, <code>cl_khr_int64_extended_atomics</code>, <code>cl_khr_3d_image_writes</code>, <code>cl_khr_integer_dot_product</code>, <code>cl_khr_extended_bit_ops</code>, <code>cl_khr_expect_assume</code> (compiler hints), <code>cl_khr_kernel_clock</code> (added provisionally in OpenCL 3.0.16, April 2024, for in-kernel <strong>profiling</strong>; developed by Arm, Imagination, Intel and Qualcomm).</p><p><strong>Sub-groups.</strong> <code>cl_khr_subgroups</code> plus <code>_extended_types</code>, <code>_non_uniform_vote</code>, <code>_ballot</code>, <code>_non_uniform_arithmetic</code>, <code>_shuffle</code>, <code>_shuffle_relative</code>, <code>_clustered_reduce</code>, <code>_rotate</code>. Mostly subsumed by 3.1 core.</p><p><strong>Execution.</strong> <code>cl_khr_command_buffer</code>, <code>cl_khr_command_buffer_mutable_dispatch</code>, <code>cl_khr_command_buffer_multi_device</code>, <code>cl_khr_suggested_local_work_size</code>, <code>cl_khr_priority_hints</code>, <code>cl_khr_throttle_hints</code>.</p><p><strong>Memory.</strong> <code>cl_khr_unified_svm</code>, <code>cl_intel_unified_shared_memory</code>, <code>cl_khr_extended_async_copies</code>, <code>cl_khr_async_copy_fence</code>, <code>cl_khr_device_uuid</code>.</p><p><strong>Interop.</strong> <code>cl_khr_gl_sharing</code> and <code>cl_khr_gl_event</code> for OpenGL; <code>cl_khr_egl_image</code> and <code>cl_khr_egl_event</code> for EGL; <code>cl_khr_d3d10_sharing</code>, <code>cl_khr_d3d11_sharing</code>, <code>cl_khr_dx9_media_sharing</code> on Windows; and the modern replacements: <code>cl_khr_semaphore</code>, <code>cl_khr_external_semaphore</code> with its <code>opaque_fd</code> and <code>sync_fd</code> variants, and <code>cl_khr_external_memory</code> with <code>dma_buf</code>, <code>opaque_fd</code> and <code>win32</code> variants. </p><p>All eight were finalised together in OpenCL 3.0.16 and are the correct way to share memory and synchronisation primitives with Vulkan. NVIDIA collaborated on these specifically for Vulkan interop and ships samples using them.</p><p><strong>Machine learning.</strong> OpenCLML on Adreno, exposed through Qualcomm&#8217;s OpenCL ML SDK. It does not appear in the Khronos registry (the <code>cl_qcom_*</code> extensions registered there cover host pointers, ION and Android native buffers, and performance hints), so treat it as SDK-delivered rather than as a registry extension. </p><p>This is a vendor extension that provides accelerated neural network operations, and Qualcomm reported at IWOCL 2026 that adding new accelerated ops in OpenCLML extension version 5 doubled prefill performance for their generative AI models in TVM&#8217;s Relax pipeline.</p><p><strong>Cooperative matrix</strong>, the newest and the one that matters most for inference. On 29 April 2026 the working group published a draft <code>cl_khr_cooperative_matrix</code>, developed with Arm, Intel and Qualcomm. It lets an OpenCL implementation accept SPIR-V modules using <code>SPV_KHR_cooperative_matrix</code>, the same extension Vulkan standardised, providing cooperative load, store and multiply-add at sub-group scope, with supported matrix shapes, component types and saturation behaviours queried through a new <code>clGetDeviceCooperativeMatrixInfoKHR</code>. </p><p>A companion extension to expose the same capability directly in OpenCL C is in progress, with an <strong>RFC </strong>on <strong>LLVM Discourse proposing the Clang front-end changes</strong>: a cooperative matrix type attribute, built-in load/store/MAD functions, and lowering to SPIR-V-friendly LLVM IR through target extension types.</p><p>The absence of this has been a concrete, measurable limitation. Qualcomm&#8217;s IWOCL 2026 paper on llama.cpp says so plainly: unlike APIs with native cooperative matrix or tensor core abstractions, <strong>OpenCL has no standardised interface for them,</strong> so their team had to engineer portable GEMM implementations for dense and mixture-of-experts workloads by hand, using adaptive tiling, subgroup-aware parallelisation and device-specific kernel variants, to reach high utilisation without dedicated matrix instructions. </p><p>Every generation of hardware that adds matrix units widens the gap that this extension is meant to close.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2><span>The implementation landscape in 2026</span></h2><p><strong>NVIDIA.</strong> OpenCL 3.0 conformant since R465 (April 2021), Maxwell and later, x86/x86-64 Linux and Windows only. Functional and maintained; not where NVIDIA&#8217;s effort goes. External memory and semaphore extensions are supported for Vulkan interop.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!JNk3!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff98184e4-55ce-444a-918d-a177eb899615_2368x1279.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!JNk3!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff98184e4-55ce-444a-918d-a177eb899615_2368x1279.png 424w, https://substackcdn.com/image/fetch/$s_!JNk3!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff98184e4-55ce-444a-918d-a177eb899615_2368x1279.png 848w, https://substackcdn.com/image/fetch/$s_!JNk3!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff98184e4-55ce-444a-918d-a177eb899615_2368x1279.png 1272w, https://substackcdn.com/image/fetch/$s_!JNk3!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff98184e4-55ce-444a-918d-a177eb899615_2368x1279.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!JNk3!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff98184e4-55ce-444a-918d-a177eb899615_2368x1279.png" width="1456" height="786" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f98184e4-55ce-444a-918d-a177eb899615_2368x1279.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:786,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Advertised conformance level over time&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Advertised conformance level over time" title="Advertised conformance level over time" srcset="https://substackcdn.com/image/fetch/$s_!JNk3!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff98184e4-55ce-444a-918d-a177eb899615_2368x1279.png 424w, https://substackcdn.com/image/fetch/$s_!JNk3!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff98184e4-55ce-444a-918d-a177eb899615_2368x1279.png 848w, https://substackcdn.com/image/fetch/$s_!JNk3!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff98184e4-55ce-444a-918d-a177eb899615_2368x1279.png 1272w, https://substackcdn.com/image/fetch/$s_!JNk3!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff98184e4-55ce-444a-918d-a177eb899615_2368x1279.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Advertised conformance level over time.</strong> The flat stretch across the middle of the chart is the same seven years the release timeline shows as a gap.</figcaption></figure></div><p><strong>AMD.</strong> OpenCL ships as part of CLR (Compute Language Runtimes), the repository that also contains the HIP runtime, sharing the ROCclr device layer. Development moved there from ROCm-OpenCL-Runtime at ROCm 5.6, and kernel compilation goes through the same Clang/LLVM AMDGPU path as HIP. AMD&#8217;s own effort is overwhelmingly directed at HIP, and OpenCL rides along on the shared runtime rather than being driven forward on its own; on Linux, Rusticl has at times outperformed the ROCm OpenCL stack on the same hardware.</p><p><strong>Intel.</strong> The most complete implementation. NEO / <code>intel/compute-runtime</code> on Linux and Windows, OpenCL 3.0 with the full 2.0 and 2.1 optional feature set plus parts of 2.2, USM, and an unusually large set of vendor extensions. Intel also maintains the OpenCL Intercept Layer and contributes heavily to the working group: Ben Ashbaugh of Intel gave the OpenCL state-of-the-union at IWOCL 2026.</p><p><strong>Arm.</strong> Mali GPUs, OpenCL 3.0 conformant from Mali-G78, G310, G510, G610, G710 and G78AE onward, with G615 and G715-Immortalis listed in October 2022. Arm implemented <code>cl_ext_cxx_for_opencl</code> and is a co-author of the cooperative matrix work.</p><p><strong>Qualcomm.</strong> Adreno, OpenCL 3.0 full profile on current Snapdragon platforms, plus an OpenCL SDK, the OpenCL ML SDK, the Snapdragon Profiler, and a published Adreno OpenCL best-practices guide. Qualcomm is, on the evidence of the last two years, the most active silicon vendor in OpenCL: the llama.cpp backend, the TVM and MLC work, OpenCLML, and co-authorship of command buffers and cooperative matrix.</p><p><strong>Imagination</strong>, <strong>VeriSilicon</strong> (Vivante GPU IP, OpenCL 3.0 and 1.2 full profile on the VIP9000 series for automotive and edge AI), <strong>Texas Instruments</strong> (DSP platforms), <strong>Samsung</strong>, and <strong>Cadence</strong> all appear on the Khronos conformant list.</p><p><strong>PoCL</strong>. Portable Computing Language, MIT licensed, fifteen years old as of IWOCL 2026, developed largely at Tampere University. Conformant for CPU and Level Zero GPU targets. </p><p>PoCL 7.2-RC1 (August 2026) achieved OpenCL 3.0 conformance for CPU devices on both x86-64 (submitted on a Ryzen 9 9900X) and RISC-V (on the Star64 board and Milk-V Jupiter), and added <code>cl_khr_extended_bit_ops</code>, <code>cl_khr_device_uuid</code>, <code>cl_khr_suggested_local_work_size</code>, <code>cl_khr_integer_dot_product</code> and <code>cl_khr_kernel_clock</code>, with LLVM 22 support for CUDA and Level Zero back ends and LLVM 22/23 for the CPU back end. </p><p>PoCL also has a remote back end that transparently offloads OpenCL work to other machines over the network, with real memory management and distributed command scheduling rather than naive call forwarding.</p><p><strong>Rusticl</strong>. Mesa&#8217;s OpenCL implementation on top of Gallium, written in Rust, led by Karol Herbst at Red Hat. Merged into mainline Mesa in September 2022 and shipped in Mesa 22.3, conformant with OpenCL 3.0 in November 2022 on 12th-generation Intel graphics via the Iris driver, and the first conformant OpenCL 3.1 implementation in July 2026. </p><p>It replaces Clover, which never supported images and was effectively abandoned. It requires <code>RUSTICL_ENABLE</code> to advertise devices for most drivers, since enabling by default is a per-driver opt-in.</p><p><strong>Layered implementations</strong>, the strategic development of the last five years. <code>clvk</code> (K&#233;vin Petit) implements the OpenCL API on Vulkan, using <code>clspv</code> to compile kernels; <code>OpenCLOn12</code> layers OpenCL over Direct3D 12 through Mesa Gallium, which is how Windows on Arm got OpenCL; <code>Ancle</code> and Rusticl-over-Zink cover further combinations. </p><p>The consequence is that &#8220;does this platform have OpenCL?&#8221; increasingly means &#8220;does this platform have Vulkan or D3D12?&#8221;, and both are close to universal. chipStar&#8217;s keynote makes the point concretely: through clvk and Rusticl, CUDA code compiled to OpenCL can reach macOS GPUs by way of MoltenVK and Metal, and mobile GPUs on Android and iOS.</p><p><strong>Conformance</strong> itself is worth explaining. The Khronos <strong>Conformance Test Suite</strong> has been open source on GitHub since 2017. Passing it and submitting the results to the Adopters Program is what entitles a product to be listed as conformant and to use the trademark; you do not need to be a Khronos member to become an adopter. </p><p>This is a meaningfully stronger regime than &#8220;<em>we support OpenCL</em>&#8221;: a listing on the conformant products page is a dated, versioned claim about a specific product on specific hardware.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2><span>Where OpenCL actually runs today</span></h2><p><strong>Mobile and edge LLM inference.</strong> llama.cpp has had an OpenCL backend since late 2024, contributed by Qualcomm and built for Adreno first. It requires sub-group support. It is tuned for <code>Q4_0</code>, with <code>--pure</code> <strong>quantisation</strong> giving the best results, and supports <code>Q6_K</code> and others. </p><p>For Snapdragon X2 SoCs there is a prebuilt binary kernel library covering <code>MUL_MAT_ID</code> with <code>Q4_0</code>, <code>Q4_1</code>, <code>Q4_K</code> and <code>MXFP4</code>, distributed through Qualcomm&#8217;s software centre, plus an on-disk compiled-program cache. </p><p>At IWOCL 2026 the Qualcomm team reported a more than fourfold prefill speedup on GPT-OSS-20B with mixture-of-experts on Snapdragon X2 Elite, from roughly 120 tokens/s to over 500, through kernel restructuring, memory-access tuning and expert load balancing. </p><p>Their roadmap names FlashAttention-style kernels, more quantisation schemes, expanded INT8 paths, and auto-tuning across vendors. That paper won the IWOCL 2026 outstanding short paper award.</p><p><strong>Deep learning compilers.</strong> TVM and MLC have long-standing Adreno OpenCL support, and Qualcomm has upstreamed the Adreno enhancements (<em>texture paths, specialised layouts, memory management, OpenCLML integration</em>) into TVM&#8217;s Relax pipeline after the community deprecated Relay. Notably, they are now adding a Vulkan backend in parallel, reusing more than 90% of the target-independent optimisations, specifically to get access to Vulkan&#8217;s cooperative matmul. That is a clear signal about what the missing cooperative matrix extension costs OpenCL.</p><p><strong>CUDA portability.</strong> chipStar compiles unmodified CUDA and HIP into fat binaries built on OpenCL and SPIR-V. Unlike source-to-source translators, it preserves the programming model and produces binaries that run without recompilation across Intel, AMD, NVIDIA, Mali, PowerVR-on-RISC-V, and CPUs.</p><p><strong>SYCL.</strong> SYCL began as a <strong>single-source</strong> C++ layer over OpenCL. SYCL 2020 generalised to multiple backends (Level Zero, CUDA, HIP) but OpenCL remains a first-class one, and it is the backend that gives SYCL reach onto hardware without a vendor SYCL implementation. Intel&#8217;s &#8220;SYCL Everywhere&#8221; work at IWOCL 2026 described exactly this: LLVM evolving for SYCL compilation, SPIR-V as the IR, OpenCL as the portability layer for accelerator offload, and PoCL and Mesa quietly extending device coverage. There is a nice concrete instance of the dependency in the other direction: the FunGT graphics engine uses SYCL&#8217;s OpenCL backend specifically to reach <code>cl_khr_gl_sharing</code>, because SYCL has no native OpenGL interop.</p><p><strong>FPGAs.</strong> Intel&#8217;s FPGA SDK for OpenCL and, historically, Xilinx SDAccel (Khronos-conformant since January 2015, later folded into Vitis) used OpenCL as an HLS front end. Pipes and pipe storage exist in the specification largely because of this constituency. The centre of gravity here has moved toward vendor HLS flows and oneAPI, but OpenCL-derived tooling persists in shipping products.</p><p><strong>Safety-critical and the road not taken.</strong> Khronos says it is exploring OpenCL&#8217;s role in safety-critical markets, and there is precedent in the neighbourhood: Vulkan SC exists, SYCL SC is in development, and IWOCL 2026 carried a paper on functional-safety-oriented GPU development that migrates CUDA to SYCL and then applies static analysis to strip constructs unsafe under IEC 61508 and ISO 26262. There is no OpenCL SC. Whether one appears is a reasonable proxy for how seriously the automotive and industrial constituency takes the standard.</p><p>WebCL is the road that was not taken. Khronos formed the working group in March 2011 and published a 1.0 specification in March 2014, aiming to expose OpenCL to JavaScript. No browser shipped it; the security surface of arbitrary compute kernels in a web page proved unattractive, and Khronos now lists it among its inactive standards. The niche it aimed at is being filled by WebGPU instead.</p><p><strong>The long tail.</strong> darktable, GIMP, Blender&#8217;s historical Cycles OpenCL backend, LuxMark, Leela Chess Zero, LAMMPS, GROMACS, ffmpeg filters, Arm Compute Library, ncnn, MNN, and the TensorFlow Lite GPU delegate. This is not glamorous work and it is a great deal of deployed code.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2><span>Why CUDA really won</span></h2><p>It is worth being accurate about this rather than reaching for &#8220;NVIDIA had better marketing.&#8221;</p><p><strong>Single-source versus split-source.</strong> In CUDA, host and device code live in the same translation unit, compiled by one compiler that type-checks kernel launches. In OpenCL, kernels are strings compiled by a different compiler at runtime, with arguments bound by index and size. The gap is not aesthetic; it is a difference in how many errors the type system catches. Templates work across the boundary in CUDA and did not in OpenCL until C++ for OpenCL, which arrived in 2020 and is not ratified.</p><p><strong>Libraries.</strong> cuBLAS, cuDNN, cuFFT, cuSPARSE, Thrust, CUB, NCCL: vendor-maintained, aggressively tuned, and free. The OpenCL equivalents were clBLAS (deprecated), CLBlast (excellent, and maintained by a far smaller group), clFFT, VexCL, ArrayFire, Boost.Compute. Building a GPU application on CUDA meant assembling tuned components; on OpenCL it frequently meant writing them.</p><p><strong>Asymmetric investment, for structural reasons.</strong> NVIDIA&#8217;s compiler team optimises one language for one family of architectures. An OpenCL implementation is a compiler plus a runtime that must be correct across a specification designed to accommodate CPUs, DSPs, FPGAs and fixed-function accelerators. Vendors were also, understandably, not motivated to make the vendor-neutral API as fast as their proprietary one. A compiler team&#8217;s attention is a budget, and it goes where the strategic return is.</p><p><strong>The 2.x fragmentation.</strong> Covered above. Between 2013 and 2020 the answer to &#8220;which OpenCL version can I target?&#8221; was 1.2, which meant OpenCL was frozen at a 2011 feature set during precisely the years when GPU compute exploded.</p><p><strong>Tooling.</strong> Nsight Compute and Nsight Systems have no OpenCL equivalent. There are good OpenCL tools, the Intercept Layer, <code>clinfo</code>, <code>clpeak</code>, Oclgrind for memory error detection, Snapdragon Profiler, Intel VTune, now CLVizulayer, but they are assembled from several projects rather than shipped as one supported product.</p><p><strong>Apple&#8217;s exit.</strong> Apple invented OpenCL, held the trademark, and deprecated it in macOS 10.14 Mojave in 2018 in favour of Metal, along with OpenGL, telling developers to move computational work to Metal Performance Shaders. Losing your originator is a bad look regardless of installed base.</p><p><strong>Blender&#8217;s retirement </strong>of the OpenCL path in Cycles is the compact version of the whole story. The stated reasons were a limited kernel implementation, driver bugs, and a stalled standard, and the combination made maintenance untenable. Not one of those is a flaw in the specification. All three are consequences of the specification being implemented by people who were investing elsewhere.</p><p>What is also true, and less often said: a striking amount of what OpenCL specified early has since become the industry consensus. SPIR-V is now the IR for Vulkan, OpenCL, and SYCL, and is a target for Slang, clspv, and a growing set of DSLs. Sub-groups are warps by another name and every API now exposes them. </p><p>Command buffers are CUDA Graphs and Level Zero command lists. Unified shared memory is CUDA&#8217;s unified memory with a better capability model. Cooperative matrix is the same abstraction in Vulkan and OpenCL, standardised jointly. The ideas held up; what OpenCL never got was the distribution to go with them.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2><span>Writing OpenCL in 2026</span></h2><p>Practical positions, offered as opinion rather than doctrine:</p><p><strong>Target OpenCL 1.2 plus queried features, or 3.0 with capability checks.</strong> Do not target 2.x. Query at startup, build the kernel variant the device can run, and keep the fallback paths honest by testing them: PoCL and Rusticl on a laptop will exercise a very different feature set from a vendor GPU driver.</p><p><strong>Ship SPIR-V, once 3.1 drivers are common.</strong> Until then, ship source with a persistent <strong>binary cache</strong> keyed on device name, driver version, source hash and build options. The startup compile cost is real and users notice it.</p><p>Concretely, the startup path that makes all of this work is short enough to paste:</p><pre><code><code>/* Select a device, discover what it can actually do, and build the
   matching kernel variant. This is the whole portability story. */
cl_device_id dev = pick_device();

char prof[64], ver[128];
clGetDeviceInfo(dev, CL_DEVICE_PROFILE, sizeof(prof), prof, NULL);
clGetDeviceInfo(dev, CL_DEVICE_VERSION, sizeof(ver),  ver,  NULL);
int embedded = (strcmp(prof, "EMBEDDED_PROFILE") == 0);
int major = ver[7] - '0';                 /* "OpenCL M.m ..."         */
int minor = ver[9] - '0';

/* OpenCL 3.0+ optionality queries; on 1.2 they simply fail, which is
   the answer.  Treat a failed query as "unsupported", never as fatal. */
cl_device_atomic_capabilities atomics = 0;
cl_bool generic_as = CL_FALSE, images = CL_FALSE;
cl_device_device_enqueue_capabilities dq = 0;
clGetDeviceInfo(dev, CL_DEVICE_ATOMIC_MEMORY_CAPABILITIES,
                sizeof(atomics), &amp;atomics, NULL);
clGetDeviceInfo(dev, CL_DEVICE_GENERIC_ADDRESS_SPACE_SUPPORT,
                sizeof(generic_as), &amp;generic_as, NULL);
clGetDeviceInfo(dev, CL_DEVICE_DEVICE_ENQUEUE_CAPABILITIES,
                sizeof(dq), &amp;dq, NULL);
clGetDeviceInfo(dev, CL_DEVICE_IMAGE_SUPPORT, sizeof(images), &amp;images, NULL);

size_t next; clGetDeviceInfo(dev, CL_DEVICE_EXTENSIONS, 0, NULL, &amp;next);
char *ext = malloc(next);
clGetDeviceInfo(dev, CL_DEVICE_EXTENSIONS, next, ext, NULL);
int subgroups = strstr(ext, "cl_khr_subgroups") != NULL || (major &gt; 3)
             || (major == 3 &amp;&amp; minor &gt;= 1);
int fp16      = strstr(ext, "cl_khr_fp16")      != NULL;
int fp64      = strstr(ext, "cl_khr_fp64")      != NULL;

/* Build options carry the decisions into the kernel. */
char opts[512];
snprintf(opts, sizeof opts,
         "-cl-std=%s -DSUBGROUPS=%d -DFP16=%d -DFP64=%d -DIMAGES=%d "
         "-DEMBEDDED=%d -cl-kernel-arg-info",
         (major &gt;= 3) ? "CL3.0" : "CL1.2",
         subgroups, fp16, fp64, images == CL_TRUE, embedded);

cl_program p = build_from_cache_or_source(ctx, dev, opts);   /* see &#167;8 */
</code></code></pre><p>and the kernel side picks its own path from the same flags, preferring the mechanism the device actually has:</p><pre><code><code>inline float block_sum(float v, __local float *scratch) {
#if SUBGROUPS &amp;&amp; (defined(__opencl_c_subgroups) || defined(cl_khr_subgroups))
#  if defined(cl_khr_subgroups) &amp;&amp; !defined(__opencl_c_subgroups)
#    pragma OPENCL EXTENSION cl_khr_subgroups : enable   /* needed on 2.x */
#  endif
    return sub_group_reduce_add(v);
#elif defined(__opencl_c_work_group_collective_functions)
    return work_group_reduce_add(v);
#else
    size_t l = get_local_id(0), n = get_local_size(0);
    scratch[l] = v;
    barrier(CLK_LOCAL_MEM_FENCE);
    for (size_t s = n &gt;&gt; 1; s &gt; 0; s &gt;&gt;= 1) {
        if (l &lt; s) scratch[l] += scratch[l + s];
        barrier(CLK_LOCAL_MEM_FENCE);
    }
    return scratch[0];
#endif
}
</code></code></pre><p>Three variants of one reduction is not elegant, and it is the actual cost of the portability the standard sells. The fallback path is the one to test hardest, because it is the one that runs on the hardware you did not anticipate.</p><p><strong>Use </strong><code>-cl-kernel-arg-info</code><strong>, set a context error callback, and check </strong><code>CL_KERNEL_PRIVATE_MEM_SIZE</code><strong>.</strong> These three cost nothing and catch a disproportionate share of problems.</p><p><strong>Treat the local work size as a measured parameter, not a guess or a default.</strong> A <strong>6.8&#215; spread</strong> across plausible values is normal. Benchmark <code>NULL</code> against a handful of multiples of <code>CL_KERNEL_PREFERRED_WORK_GROUP_SIZE_MULTIPLE</code>, bounded by <code>CL_KERNEL_WORK_GROUP_SIZE</code>, or use the 3.1 suggested-size query. Sometimes the runtime wins; assume neither.</p><p><strong>Assume nothing about work-group execution order.</strong> No inter-work-group spin-waiting, ever.</p><p><strong>Treat tuning parameters as data, not code.</strong> Tile sizes, work-per-thread, vector width and unroll factors should be build-time defines selected from a per-architecture table, with <strong>autotuning</strong> as the fallback. This is what CLBlast does and it is the only approach that survives contact with five vendors.</p><p><strong>Reach for the ecosystem before writing kernels.</strong> CLBlast for BLAS, VkFFT or clFFT for transforms, Arm Compute Library on Mali, OpenCLML on Adreno. And if the actual goal is portable C++ rather than portable kernels, SYCL over an OpenCL backend is very likely the better answer: OpenCL is a good compilation target and a mediocre application-level programming model, and the working group&#8217;s own framing now treats it that way.</p><p><strong>Debug with the layer mechanism.</strong> The Intercept Layer will tell you what your application is really doing to the driver, without a rebuild.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2><span>Where this goes</span></h2><p>The OpenCL of 2026 has settled into a role nobody planned for it in 2008. It is not the API most application developers write. It is the layer that other things stand on: SYCL implementations, chipStar, TVM, MLC, llama.cpp&#8217;s mobile path, the Mesa and PoCL stacks that give OpenCL, and therefore everything above it, reach onto hardware whose vendors have shipped nothing. </p><p>The working group&#8217;s own language has shifted to match, talking about OpenCL as a substrate for higher-level programming models rather than as the model itself.</p><p>The most urgent open question is cooperative matrix. Hardware matrix units are where the FLOPs are, Vulkan standardised access to them years ago, and until the OpenCL extension is final and shipping, every inference backend on OpenCL is hand-rolling GEMM against silicon it cannot fully address. </p><p><strong>Qualcomm adding a Vulkan path to TVM alongside the OpenCL one</strong>, explicitly to get cooperative matmul, is the warning shot. The low-precision data types are the same problem one layer down: int4 and fp8 are on the roadmap, and the workloads that need cooperative matrix need those too. </p><p>Underneath both sits the question of whether the 3.1 SPIR-V mandate converts into broad driver support quickly enough to matter to the compiler authors it was written for.</p><p>The encouraging signal is that the working group has stopped repeating the 2.0 mistake. <strong>Everything mandated in 3.1 was already deployed</strong>. The extension pipeline is public, the drafts are on GitHub, the Clang RFCs are on Discourse, and the conformance suite is open source. </p><p>Whether that process moves fast enough against a competitor with a decade&#8217;s head start and no committee is a genuinely open question, and I do not think anyone should be confident either way.</p><p>What I&#8217;d say with more confidence is that the obituaries were premature by about a decade, and that if you are building anything that has to run on hardware you do not control <em>(phones, embedded silicon, whatever a customer happens to own)</em>, OpenCL is currently the only thing in the category that is both open and actually there.</p><p>Which returns to where this started. A standard that has to describe hardware it cannot see can only ever specify floors. The measurements in this article are all instances of that: <strong>an accuracy bound four orders of magnitude looser than what an implementation delivers</strong>, a work-group size the specification will not choose for you and that costs 6.8&#215; to get wrong, an allocation flag that means nothing until you change how you use it. </p><p>Every one of those is the standard declining to promise something it cannot promise for every device.</p><p>That is the failure mode and it is also the whole value. CUDA can tell you what your hardware does because NVIDIA built it. OpenCL cannot, and in exchange it is <strong>the reason a kernel written in 2011 runs today on a phone, an FPGA, a RISC-V board, and an Apple GPU</strong> under a driver Apple did not write and does not support. </p><p>The interesting question was never whether that trade was worth it in general. </p><p>It is whether it is worth it for the thing you are building, and the honest answer is that for most people writing CUDA in 2026 it is not, and for the people shipping software onto hardware they will never see it is the only trade on offer.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2><span>Appendix A: version timeline</span></h2><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!YonR!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcff0f86d-589e-4611-91d9-9bcf64d66cd9_2902x2839.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!YonR!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcff0f86d-589e-4611-91d9-9bcf64d66cd9_2902x2839.png 424w, https://substackcdn.com/image/fetch/$s_!YonR!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcff0f86d-589e-4611-91d9-9bcf64d66cd9_2902x2839.png 848w, https://substackcdn.com/image/fetch/$s_!YonR!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcff0f86d-589e-4611-91d9-9bcf64d66cd9_2902x2839.png 1272w, https://substackcdn.com/image/fetch/$s_!YonR!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcff0f86d-589e-4611-91d9-9bcf64d66cd9_2902x2839.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!YonR!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcff0f86d-589e-4611-91d9-9bcf64d66cd9_2902x2839.png" width="1456" height="1424" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/cff0f86d-589e-4611-91d9-9bcf64d66cd9_2902x2839.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1424,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:254625,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/209597506?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcff0f86d-589e-4611-91d9-9bcf64d66cd9_2902x2839.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!YonR!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcff0f86d-589e-4611-91d9-9bcf64d66cd9_2902x2839.png 424w, https://substackcdn.com/image/fetch/$s_!YonR!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcff0f86d-589e-4611-91d9-9bcf64d66cd9_2902x2839.png 848w, https://substackcdn.com/image/fetch/$s_!YonR!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcff0f86d-589e-4611-91d9-9bcf64d66cd9_2902x2839.png 1272w, https://substackcdn.com/image/fetch/$s_!YonR!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcff0f86d-589e-4611-91d9-9bcf64d66cd9_2902x2839.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2><span>Appendix B: queries worth calling at startup</span></h2><pre><code><code>/* Identity and version */
CL_PLATFORM_VERSION            CL_DEVICE_VERSION
CL_DEVICE_NAME                 CL_DEVICE_VENDOR
CL_DRIVER_VERSION              CL_DEVICE_OPENCL_C_ALL_VERSIONS   /* 3.0 */
CL_DEVICE_UUID_KHR                                               /* core in 3.1 */

/* Shape of the machine */
CL_DEVICE_MAX_COMPUTE_UNITS    CL_DEVICE_MAX_WORK_GROUP_SIZE
CL_DEVICE_MAX_WORK_ITEM_SIZES  CL_DEVICE_LOCAL_MEM_SIZE
CL_DEVICE_LOCAL_MEM_TYPE       CL_DEVICE_GLOBAL_MEM_CACHELINE_SIZE
CL_DEVICE_MAX_MEM_ALLOC_SIZE   CL_DEVICE_GLOBAL_MEM_SIZE
CL_DEVICE_HOST_UNIFIED_MEMORY  /* meaningful as of 3.1 */

/* Capabilities */
CL_DEVICE_EXTENSIONS_WITH_VERSION
CL_DEVICE_SVM_CAPABILITIES
CL_DEVICE_ATOMIC_MEMORY_CAPABILITIES
CL_DEVICE_ATOMIC_FENCE_CAPABILITIES
CL_DEVICE_DEVICE_ENQUEUE_CAPABILITIES
CL_DEVICE_PIPE_SUPPORT
CL_DEVICE_GENERIC_ADDRESS_SPACE_SUPPORT
CL_DEVICE_WORK_GROUP_COLLECTIVE_FUNCTIONS_SUPPORT
CL_DEVICE_NON_UNIFORM_WORK_GROUP_SUPPORT
CL_DEVICE_IMAGE_SUPPORT
CL_DEVICE_SINGLE_FP_CONFIG     CL_DEVICE_DOUBLE_FP_CONFIG

/* Per kernel, after build */
CL_KERNEL_WORK_GROUP_SIZE
CL_KERNEL_PREFERRED_WORK_GROUP_SIZE_MULTIPLE
CL_KERNEL_LOCAL_MEM_SIZE
CL_KERNEL_PRIVATE_MEM_SIZE
CL_KERNEL_COMPILE_WORK_GROUP_SIZE
</code></code></pre><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2><span>Appendix C: the same concept in seven APIs</span></h2><p>Every modern compute API converged on the same execution model with different nouns. This is the translation table, and it is most of what porting between them consists of.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!NjnZ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F309f1994-f4cc-49d5-9abb-980ef4aad959_2849x1846.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!NjnZ!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F309f1994-f4cc-49d5-9abb-980ef4aad959_2849x1846.png 424w, https://substackcdn.com/image/fetch/$s_!NjnZ!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F309f1994-f4cc-49d5-9abb-980ef4aad959_2849x1846.png 848w, https://substackcdn.com/image/fetch/$s_!NjnZ!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F309f1994-f4cc-49d5-9abb-980ef4aad959_2849x1846.png 1272w, https://substackcdn.com/image/fetch/$s_!NjnZ!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F309f1994-f4cc-49d5-9abb-980ef4aad959_2849x1846.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!NjnZ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F309f1994-f4cc-49d5-9abb-980ef4aad959_2849x1846.png" width="1456" height="943" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/309f1994-f4cc-49d5-9abb-980ef4aad959_2849x1846.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:943,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:179547,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/209597506?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F309f1994-f4cc-49d5-9abb-980ef4aad959_2849x1846.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!NjnZ!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F309f1994-f4cc-49d5-9abb-980ef4aad959_2849x1846.png 424w, https://substackcdn.com/image/fetch/$s_!NjnZ!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F309f1994-f4cc-49d5-9abb-980ef4aad959_2849x1846.png 848w, https://substackcdn.com/image/fetch/$s_!NjnZ!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F309f1994-f4cc-49d5-9abb-980ef4aad959_2849x1846.png 1272w, https://substackcdn.com/image/fetch/$s_!NjnZ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F309f1994-f4cc-49d5-9abb-980ef4aad959_2849x1846.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><strong>Two structural observations</strong> fall out of the table. The rows are almost perfectly aligned, which is why source-to-source translation between these APIs works as well as it does and why chipStar can compile CUDA onto OpenCL at all. And the one row where OpenCL is behind, the last one, is the row where the FLOPs are.</p><p>The <strong>source model</strong> row is the difference that decided the market. Single-source means host and device code are compiled together by one compiler, so a kernel launch is type-checked and templates cross the boundary. </p><p>Split-source means kernels are strings or binaries bound by index at runtime. The pattern is not simply proprietary-versus-open, and Metal is the counterexample that shows it: Apple controls its compiler completely and still chose <strong>split-source</strong>, because MSL is its own language in its own files. </p><p>What actually predicts the choice is whether one compiler is expected to consume both halves. CUDA, HIP and SYCL say yes and get type-checked launches and templates across the boundary. </p><p>OpenCL, Vulkan, Metal and Level Zero say no, and pay for it at every call site where an argument is bound by index instead of by name.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2><span>Appendix D: sources</span></h2><p>Khronos OpenCL Registry and the unified API, C, and SPIR-V Environment specifications (rev. 3.1.1, 22 May 2026); the Khronos blog posts announcing OpenCL 3.1 (4 May 2026) and the cooperative matrix extensions (29 April 2026); </p><p>the OpenCL-Docs, OpenCL-Headers, OpenCL-ICD-Loader and OpenCL-CTS repositories; the IWOCL 2026 program and papers, particularly the Qualcomm llama.cpp paper, the chipStar keynote, Intel&#8217;s AI-workload talk, and the CLVizulayer presentation; </p><p>NVIDIA&#8217;s OpenCL 3.0 conformance announcement (April 2021); llama.cpp&#8217;s OpenCL backend documentation; PoCL and Mesa release notes; Clang&#8217;s OpenCL support documentation; and Phoronix&#8217;s reporting on the 3.0.x and 3.1 releases and on conformance submissions.</p>]]></content:encoded></item><item><title><![CDATA[When Batching Stops Working]]></title><description><![CDATA[Why does batching stop working? We solve the compute&#8211;memory crossover and find four GPU generations become memory-bound at just 84&#8211;162 tokens of context.]]></description><link>https://www.thesoftwarefrontier.com/p/when-batching-stops-working</link><guid isPermaLink="false">https://www.thesoftwarefrontier.com/p/when-batching-stops-working</guid><dc:creator><![CDATA[Lorenzo Bradanini]]></dc:creator><pubDate>Fri, 28 Aug 2026 11:31:19 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!ZJxb!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9932372e-1c61-4393-b336-6489d1c4fb11_1536x843.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!ZJxb!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9932372e-1c61-4393-b336-6489d1c4fb11_1536x843.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!ZJxb!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9932372e-1c61-4393-b336-6489d1c4fb11_1536x843.png 424w, https://substackcdn.com/image/fetch/$s_!ZJxb!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9932372e-1c61-4393-b336-6489d1c4fb11_1536x843.png 848w, https://substackcdn.com/image/fetch/$s_!ZJxb!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9932372e-1c61-4393-b336-6489d1c4fb11_1536x843.png 1272w, https://substackcdn.com/image/fetch/$s_!ZJxb!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9932372e-1c61-4393-b336-6489d1c4fb11_1536x843.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!ZJxb!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9932372e-1c61-4393-b336-6489d1c4fb11_1536x843.png" width="1456" height="799" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/9932372e-1c61-4393-b336-6489d1c4fb11_1536x843.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:799,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2567223,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/208654603?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9932372e-1c61-4393-b336-6489d1c4fb11_1536x843.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!ZJxb!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9932372e-1c61-4393-b336-6489d1c4fb11_1536x843.png 424w, https://substackcdn.com/image/fetch/$s_!ZJxb!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9932372e-1c61-4393-b336-6489d1c4fb11_1536x843.png 848w, https://substackcdn.com/image/fetch/$s_!ZJxb!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9932372e-1c61-4393-b336-6489d1c4fb11_1536x843.png 1272w, https://substackcdn.com/image/fetch/$s_!ZJxb!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9932372e-1c61-4393-b336-6489d1c4fb11_1536x843.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>Introduction</h2><p>As many of you know, even if you&#8217;re not in the financial side of technical companies, <strong>NVIDIA published two numbers </strong>twelve hours apart this week. On Monday its engineers were at Hot Chips and they described a rack of 256 chips that holds 128 gibibytes of SRAM. </p><p>On Wednesday its CFO reported that the company has committed <strong>279 billion dollars</strong>, primarily to buying DRAM: those are not separate stories. They are the same thing said twice, once by the engineering organisation and once by finance, and what connects them is a quantity that the standard formula for it gets wrong.</p><p>Everyone who deals with<strong> inference hardware</strong> knows the critical batch: the point where a decode step stops being memory bound and starts being compute bound, B* equals peak arithmetic times bytes per element over twice the memory bandwidth. <em>On a Rubin package that is 398, </em>whereas on the <em>LPX rack it is 3.94.</em> Those are the numbers I expected to build this piece on.</p><p>They describe a machine serving no context at all. The formula is derived from a decode step that reads only weights, and once you put the KV cache back into it the expression acquires a pole. </p><p>After a certain context length the <strong>denominator goes negative </strong>and no batch size makes the machine compute bound, because each additional sequence adds more memory traffic than amortisable arithmetic.</p><p>For Gemma 4 31B, the model NVIDIA benchmarked, a Rubin package reaches that point at <strong>84 tokens</strong>. The LPX rack holds out to roughly <strong>68,000</strong>. That precise gap, and not the 101 times in the spec sheets, is what NVIDIA spent 20 billion dollars to buy.</p><p>What follows is the derivation, the money behind it, and the places where I got it wrong first.</p><h3>The short version</h3><p>I want to introduce this section in this research post, so that you are able to take the core concepts and principles in the short form. I do really think that it will help, let me know if that&#8217;s the correct approach. </p><ol><li><p>The <strong>critical batch formula</strong> in common use drops the KV term. Restored, it has a pole at the saturation context: 84 tokens on Rubin, 68,000 on the LPX rack, a ratio of 811 rather than 101.</p></li><li><p>Above saturation, raising batch buys capacity utilisation, not arithmetic utilisation. Throughput ceilings at memory bandwidth over cache bytes per sequence. For this model on an NVL72 that is about <strong>143,000 tokens per second</strong> and no scheduler setting beats it.</p></li><li><p>At 100K of context the <strong>NVL72 makes 13 times the rack tokens.</strong> The LPX makes each user&#8217;s tokens arrive 45 times faster. That is the entire trade, measured at one context length.</p></li><li><p>The LPX is a latency product, not a bandwidth product. It runs 279 times below its own bandwidth roofline, and streaming plus arithmetic account for <strong>0.43 percent </strong>of a token&#8217;s time budget.</p></li><li><p>SRAM is <strong>1.4 to 6.4 times dearer per byte </strong>than HBM4 and 642 to 2,924 times cheaper per byte per second. Every design decision in the rack follows from that one ratio.</p></li><li><p>NVIDIA&#8217;s 279 billion dollar commitment, read against its own revenue guide, implies <strong>28.2 dollars per gigabyte of HBM4</strong>, which lands within 11 percent of the price Korean and Taiwanese analysts publish. Two methods sharing no inputs.</p></li></ol><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>The ledger</h2><p>Let&#8217;s start with the disclosure, because it is unambiguous and it is filed directly with the SEC. NVIDIA reported <strong>revenue of 96,221 million dollars </strong>for the quarter ended 26 July 2026, of which 89,023 million was the Data Center section. </p><p>Gross margin was 75.0 percent on both a GAAP and a non-GAAP basis. Cost of revenue was precisely at 24,079 million.</p><p><em>Three things</em> in the same document deserve more attention than this short description.</p><p>The first one is the commitment table. In the CFO commentary, <strong>Colette Kress </strong>wrote that supply commitments rose from 119 billion dollars a quarter ago to 279 billion, and attributed the increase to the procurement of memory. </p><p>The table below breaks it out by fiscal year: 92 billion for the remainder of FY2027, 87 for FY2028, 88 for FY2029, then 6, then 5, then 1.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!THuv!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F843e3ced-8156-4c04-8318-a592b40212a1_1520x919.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!THuv!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F843e3ced-8156-4c04-8318-a592b40212a1_1520x919.png 424w, https://substackcdn.com/image/fetch/$s_!THuv!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F843e3ced-8156-4c04-8318-a592b40212a1_1520x919.png 848w, https://substackcdn.com/image/fetch/$s_!THuv!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F843e3ced-8156-4c04-8318-a592b40212a1_1520x919.png 1272w, https://substackcdn.com/image/fetch/$s_!THuv!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F843e3ced-8156-4c04-8318-a592b40212a1_1520x919.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!THuv!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F843e3ced-8156-4c04-8318-a592b40212a1_1520x919.png" width="1456" height="880" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/843e3ced-8156-4c04-8318-a592b40212a1_1520x919.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:880,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Figure 7&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 7" title="Figure 7" srcset="https://substackcdn.com/image/fetch/$s_!THuv!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F843e3ced-8156-4c04-8318-a592b40212a1_1520x919.png 424w, https://substackcdn.com/image/fetch/$s_!THuv!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F843e3ced-8156-4c04-8318-a592b40212a1_1520x919.png 848w, https://substackcdn.com/image/fetch/$s_!THuv!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F843e3ced-8156-4c04-8318-a592b40212a1_1520x919.png 1272w, https://substackcdn.com/image/fetch/$s_!THuv!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F843e3ced-8156-4c04-8318-a592b40212a1_1520x919.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Figure 1.</strong> Supply and capacity commitments by fiscal year as disclosed in the Q2 FY2027 CFO commentary. The total rose from 119 billion dollars one quarter earlier, an increase the company attributes to the procurement of memory.</figcaption></figure></div><p><strong>Ninety six percent </strong>of that money, 267 billion of the 279, lands inside three fiscal years. This isn&#8217;t a long-dated strategic reserve: it is a three year buy with a cliff at the end of it, which is what a purchase schedule looks like when you are securing a component whose supply is allocated rather than traded.</p><p>The second one is the guide, and it runs way further than the headline. NVIDIA guided <strong>Q3 revenue to 108 billion dollars </strong>and gross margin to 74.0 percent, a hundred basis points below the quarter it just reported. On the call Kress said that margin falls to 71 or 72 percent by the January quarter before recovering to 72 or 73 next year. </p><p>That is a three hundred to four hundred basis point trough, not a one hundred point wobble, and it is scheduled. On 108 billion, a hundred basis points is 1,080 million dollars of additional cost of revenue in a single quarter. </p><p>If memory is roughly 30 to 40 percent of accelerator manufacturing cost, which is the range the trade estimates converge on and which is consistent with the roughly <strong>3,250 dollars of HBM3e</strong> against 850 dollars of logic die in a B200, then the entire memory bill in that quarter is somewhere between 8.4 and 11.2 billion dollars. The guide-down is about eleven percent of it.</p><p>The third is the <strong>working capital</strong>. Days sales outstanding went from 45 to 60, attributed to extended payment terms on multi-quarter agreements. Inventory went from 25.8 to 31.6 billion. </p><p>The company issued 25 billion dollars of senior unsecured notes in a quarter in which it generated 21.3 billion of free cash flow and returned 26.0 billion to shareholders. </p><p>Separately it disclosed guarantees with a maximum gross exposure of 108.5 billion, of which 105 billion is credit support for 4.25 gigawatts at <strong>SB Energy&#8217;s PORTS-Pike campus</strong> in Ohio, leased to OpenAI for twenty years.</p><p>I&#8217;m not going to make this post about the financing structure, which is clearly outside what I can measure and out of this niche. We just want the <strong>memory number</strong>, and this is checkable against something else NVIDIA said.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Solving for the price of a gigabyte</h2><p>The commitment schedule and the revenue guide are <strong>two independent statements </strong>about the same physical object count. Force them to agree and the effective memory price falls out.</p><p>Take the revenue path first of all. </p><blockquote><p><strong>Q1 came in at 81.6 billion</strong>, Q2 at 96.2, Q3 is guided to 108.0. If Q4 grows at the same sequential rate as the Q3 guide, FY2027 lands near 407 billion. Management gave preliminary FY2028 guidance of roughly 70 percent growth, supply constrained, which puts FY2028 near 692 billion. Decay that growth by half for FY2029 and the <strong>three year total is about 2.03 trillion dollars.</strong> The 279 billion commitment is 13.7 percent of it.</p></blockquote><p>Now let&#8217;s try to convert revenue to units. </p><blockquote><p><strong>Data Center was 92.5 percent of revenue</strong> this quarter. Not all of that is GPU: there are CPUs, switches, NICs and complete systems in the number, and the CFO commentary breaks Data Center out by customer type rather than by product, so the GPU share is not disclosed. Take a band of 55 to 70 percent of Data Center revenue as GPU, and an average selling price of 40,000 to 50,000 dollars per package. The three year revenue path then implies somewhere<strong> between 20.7 and 32.9 million GPUs.</strong></p></blockquote><p>Each Rubin package carries 288 gigabytes of HBM4. If 60 to 90 percent of the 279 billion is memory, dividing through gives an <strong>effective price of 17.7 to 42.1 dollars per gigabyte</strong>, with a simple average<strong> midpoint of 28.2.</strong></p><p>What&#8217;s worth noting is that<strong> two different HBM4 prices</strong> are in wide circulation and they are not comparable, which is sort of a trap we walked into in the first edit of this piece. </p><p>The factory gate figure is a component price: it&#8217;s <strong>circa 550 dollars for a 36 gigabyte twelve-high stack</strong>, or 15.28 dollars per gigabyte. The contract figure is what NVIDIA is projected to actually pay under long term agreements. </p><p><em>Seoul Economic Daily </em>reported in July that HBM4 was moving from about 2 dollars per gigabit to 4 or 5, which is 32 to 40 dollars per gigabyte. Various analyst projections we&#8217;ve seen during the research process reported this month put it at <strong>31 to 32 dollars per gigabyte for NVIDIA</strong> specifically and 35 to 36 for other buyers.</p><p>As stated before, our reconciled figure is <strong>28.2 dollars per gigabyte</strong> with a band of 17.7 to 42.1. That lands within 11 percent of the 31 to 32 the Korean and Taiwanese analysts are publishing, from a completely independent direction: a commitment table and a revenue guide, with no memory market data used anywhere in the derivation. </p><p>The <strong>strongest result</strong> in this analysis is that two completely independent methods converge to within a tenth of each other. That agreement provides the strongest evidence that the underlying financial calculations are in some ways, real.</p><p>Let&#8217;s call it<strong> 30 dollars a gigabyte,</strong> then. At 288 gigabytes per package, the memory in a <em>single Rubin GPU costs on the order of 8,600 dollars</em>, and the memory in a Vera Rubin NVL72 rack costs on the order of 620,000. </p><p>On the factory gate basis it would be 4,400 and 317,000, which is the number to use if you want a component cost and the wrong number to use if you want NVIDIA&#8217;s bill. Either way, <strong>memory is the largest single line</strong>, and it is the line whose price is set by three suppliers in a market where HBM has gone from 8 percent of DRAM wafer output in 2024 to 23 percent in 2026.</p><p>That is the position NVIDIA is buying its way through. The interesting question is what it is building to avoid it.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>The machine</h2><p>NVIDIA licensed <strong>Groq&#8217;s LPU architecture</strong> and hired most of the company in a 20 billion dollar transaction struck at the end of December 2025. The first silicon, the LP30, was shown at GTC in March and went into full production this month. </p><p><strong>Igor Arsovski,</strong> who came over from Groq, presented the rack at Hot Chips precisely on Monday. Nebius is the named first customer. The cash flow statement carries a 2,944 million dollar line item labelled simply &#8220;<em>Groq, Inc.</em>&#8221; under financing activities.</p><p>The published rack specification is <strong>256 LP30 accelerators</strong>, 128 GB of SRAM, 40 PB/s of aggregate SRAM bandwidth, 315 PFLOPS of FP8, and 96 chip-to-chip links per die running at 112 Gbps each. The die is Samsung 4nm and the trade press puts 512 MB of SRAM and 150 TB/s on it.</p><p>Those figures do not quite close, and the way they fail to do it so is merely informative. 256 dies at 512 MB each is <strong>131.07 gigabytes in decimal units</strong>, not 128, which is exactly 128 gibibytes. So the rack figure is binary and the die figure is decimal, and the only reading under which both are correct is 512 mebibytes per die. </p><p>Similarly, <strong>40 PB/s across 256 dies is 156.25 TB/s each</strong>, not 150; and 72 Rubin GPUs at 22 TB/s is 1.584 PB/s, not the 1.6 that gets quoted. I use the exact values throughout and note where they differ from the round ones.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>One ratio we want you to think</h2><p>An LP30 contains 512 MiB of SRAM. That lets us estimate its cost within a reasonable range. A high-density six-transistor SRAM cell measures <strong>about 0.021 &#181;m&#178; on TSMC N5</strong>, a figure that remains broadly unchanged at N3E. The LP30, however, is built on Samsung 4 nm, where 4LPP uses a 54 nm contacted gate pitch versus 51 nm for N5, putting the equivalent cell closer to 0.026 &#181;m&#178;.</p><p>At 512 MiB, the<strong> LP30 contains 4,294,967,296 bits</strong>. At 0.021&#8211;0.026 &#181;m&#178; per bit, that corresponds to roughly 90&#8211;112 mm&#178; of raw SRAM cell area. Once decoders, sense amplifiers, column multiplexers, and other array overhead are included, assuming 60&#8211;75% array efficiency, the SRAM array itself comes to approximately 120&#8211;186 mm&#178;.</p><p>Wafer price cuts the other way. <strong>TSMC N5 and N4 wafers run at roughly $18,500</strong>, while Samsung typically prices 15&#8211;30 percent below TSMC at comparable nodes, putting a Samsung wafer at approximately $12,950&#8211;$15,725. Spread across 70,686 usable square millimetres, that gives roughly $22&#8211;$49 of silicon cost per die, or $44&#8211;$97 per gigabyte of SRAM, with a central case around $64.</p><p>The two corrections partly cancel. Using a TSMC SRAM cell size together with a TSMC wafer price for a Samsung-built part understates the die area by roughly a quarter while overstating the wafer cost by a similar amount. </p><p>Those errors happen to push the final estimate in opposite directions, leaving the <em>original point estimate of $66 per gigabyte surprisingly close to the corrected central case of $64.</em> The agreement is therefore accidental rather than methodological. The defensible result is a range, not a single number.</p><p>Against <strong>HBM4 at roughly $15.28 per gigabyte</strong> as a component and $31.50 under contract, SRAM comes out somewhere <strong>between 1.4 and 6.4 times more expensive per byte</strong>. And even that comparison is generous to SRAM: the calculation covers essentially bare array area, with no allowance for yield loss, peripheral circuitry, testing, packaging, or margin.</p><p>On these assumptions, the conclusion is difficult to avoid: SRAM is extraordinarily expensive memory.</p><p>Now divide by bandwidth instead of by capacity.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!KNcB!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fce7bd61d-ddea-4038-af02-9364bf30f473_1520x919.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!KNcB!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fce7bd61d-ddea-4038-af02-9364bf30f473_1520x919.png 424w, https://substackcdn.com/image/fetch/$s_!KNcB!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fce7bd61d-ddea-4038-af02-9364bf30f473_1520x919.png 848w, https://substackcdn.com/image/fetch/$s_!KNcB!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fce7bd61d-ddea-4038-af02-9364bf30f473_1520x919.png 1272w, https://substackcdn.com/image/fetch/$s_!KNcB!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fce7bd61d-ddea-4038-af02-9364bf30f473_1520x919.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!KNcB!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fce7bd61d-ddea-4038-af02-9364bf30f473_1520x919.png" width="1456" height="880" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ce7bd61d-ddea-4038-af02-9364bf30f473_1520x919.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:880,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Figure 4&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 4" title="Figure 4" srcset="https://substackcdn.com/image/fetch/$s_!KNcB!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fce7bd61d-ddea-4038-af02-9364bf30f473_1520x919.png 424w, https://substackcdn.com/image/fetch/$s_!KNcB!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fce7bd61d-ddea-4038-af02-9364bf30f473_1520x919.png 848w, https://substackcdn.com/image/fetch/$s_!KNcB!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fce7bd61d-ddea-4038-af02-9364bf30f473_1520x919.png 1272w, https://substackcdn.com/image/fetch/$s_!KNcB!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fce7bd61d-ddea-4038-af02-9364bf30f473_1520x919.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Figure 2.</strong> Left: dollars per gigabyte, SRAM array area against HBM4 on both price bases. Whiskers on the SRAM bar span the sensitivity across cell size, array efficiency and wafer price. HBM4 twelve-high contract price. Right: dollars per terabyte per second of delivered bandwidth, same two technologies, log scale. The SRAM figure is bare array area with no yield, packaging, test or margin, so it understates capacity cost and overstates the bandwidth advantage.</figcaption></figure></div><p>The <strong>512 mebibyte array delivers 156 TB/s</strong>, so the silicon under a terabyte per second of SRAM bandwidth costs 14 to 31 cents, centrally 21. The 288 gigabytes of HBM4 on a <strong>Rubin package delivers 22 TB/s</strong>, so the memory under a terabyte per second of HBM bandwidth costs about 200 dollars as a component and 412 under contract.</p><p>Take the corner <strong>least favourable to my argument</strong>, the dearest SRAM against the cheapest HBM, and the ratio is 642. Take the central cases and it is 974 against the component price and 2,009 against the contract price. Take the corner most favourable and it is 2,924. </p><p>There is no assumption inside any of these ranges under which the answer is less than two and a half orders of magnitude, which is the only precision the conclusion needs.</p><p>Stated the other way: the <strong>SRAM in an entire LPX rack is 5,600 to 12,500 dollars</strong> of silicon area, centrally about 8,200. The HBM4 in a single Vera Rubin NVL72 rack is 317,000 dollars as a component and about 653,000 at the price NVIDIA is projected to pay. The LPX rack has twenty five times the aggregate bandwidth.</p><p>Every design decision in the LPX falls out of that one ratio. Bandwidth per byte is 3,810 times higher on the SRAM machine, which means <strong>capacity per unit of bandwidth is 3,810 times lower.</strong> A Vera Rubin NVL72 holds 20.7 terabytes. The LPX holds 128 gibibytes, which is 151 times less. </p><p>You cannot put a large model in it, you cannot put a long KV cache in it, and the only workload it can run alone is one whose entire working set is under about 137 gigabytes.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>The knee</h2><p>In a piece earlier this month I derived the batch size at which a decode step stops being memory bound. A decode step reads W bytes of weights and performs <strong>2NB floating point operations</strong>, where N is the parameter count, W equals N times b bytes per element, and B is the batch. Setting compute time equal to memory time gives</p><p><code>B* = P &#183; b / (2 &#183; Bmem)</code></p><p>and because peak arithmetic P scales as 1/b on tensor cores, the product P&#183;b is a generational constant. <strong>B* is invariant to precision.</strong> Running the same model in FP4 instead of FP8 doubles your token rate and moves the knee not at all.</p><p>Both sides of that ratio have to be on the same basis, which is the part that is easy to get wrong and which we both got wrong in the earlier drafts of this piece. A sparse or compressed peak divided by a real bandwidth inflates B* by the compression ratio. Every figure below is dense.</p><p>The <strong>invariance held well for two generations.</strong> H100 sits at 295, H200 at 206, HGX B200 at 292, GB200 and MI355X at 312. All five re-derive here from datasheet FP8 dense throughput and published HBM bandwidth, so the plateau is a result rather than a citation.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!dkwU!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F00bc9876-6dbd-4183-ac5d-2c5c753eb4f1_1520x919.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!dkwU!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F00bc9876-6dbd-4183-ac5d-2c5c753eb4f1_1520x919.png 424w, https://substackcdn.com/image/fetch/$s_!dkwU!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F00bc9876-6dbd-4183-ac5d-2c5c753eb4f1_1520x919.png 848w, https://substackcdn.com/image/fetch/$s_!dkwU!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F00bc9876-6dbd-4183-ac5d-2c5c753eb4f1_1520x919.png 1272w, https://substackcdn.com/image/fetch/$s_!dkwU!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F00bc9876-6dbd-4183-ac5d-2c5c753eb4f1_1520x919.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!dkwU!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F00bc9876-6dbd-4183-ac5d-2c5c753eb4f1_1520x919.png" width="1456" height="880" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/00bc9876-6dbd-4183-ac5d-2c5c753eb4f1_1520x919.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:880,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Figure 2&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 2" title="Figure 2" srcset="https://substackcdn.com/image/fetch/$s_!dkwU!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F00bc9876-6dbd-4183-ac5d-2c5c753eb4f1_1520x919.png 424w, https://substackcdn.com/image/fetch/$s_!dkwU!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F00bc9876-6dbd-4183-ac5d-2c5c753eb4f1_1520x919.png 848w, https://substackcdn.com/image/fetch/$s_!dkwU!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F00bc9876-6dbd-4183-ac5d-2c5c753eb4f1_1520x919.png 1272w, https://substackcdn.com/image/fetch/$s_!dkwU!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F00bc9876-6dbd-4183-ac5d-2c5c753eb4f1_1520x919.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Figure 3.</strong> Weights-only critical batch by accelerator, log scale, which is the right basis for comparing machines and not for describing serving; see figure 1. All seven from published dense arithmetic and memory bandwidth. against 22 TB/s; LP30 from 315 PFLOPS FP8 against 40 PB/s across the rack. All seven derived from dense FP8 throughput and published memory bandwidth. The dotted line marks what Rubin gives if the marketed sparse NVFP4 peak is used instead, which is the ten to seven compression ratio too high.</figcaption></figure></div><p>Rubin is where the basis matters. NVIDIA markets the R100 at 50 PFLOPS of NVFP4 inference, and that figure is not dense. SemiAnalysis reports that <strong>NVIDIA brands 50 PFLOPS </strong>as the inference number while 35 PFLOPS of NVFP4 is the dense one, and that the five times over Blackwell claim compares compressed FP4 against dense FP4. Their Rubin CPX analysis puts the same part at 50 sparse and 33.3 dense on a three to two ratio.</p><p>NVIDIA&#8217;s own rack table settles it. The <strong>NVL72 is published at 3,600 PFLOPS of NVFP4 inference</strong>, 2,520 of dense NVFP4, and 1,260 of dense FP8. </p><p>Divide by 72 and the per package figures are 50 marketed, 35 dense NVFP4, and 17.5 dense FP8. The ratio between the marketed and the dense number is 3,600 over 2,520, which is ten to seven.</p><p>Both dense figures give the same answer, as they must, since P scales as 1 over b: <strong>35 PFLOPS at half a byte </strong>and 17.5 at one byte are the same product. Against 22 TB/s, B* is 398. The SemiAnalysis route through the three to two tensor core ratio gives 378. Had I used the marketed 50, I would have got 568, which is exactly the ten to seven too high.</p><p>So the <strong>knee moved 1.36 times over HGX B200</strong>, and not 1.94. That is a smaller claim than the one I started with and it is the one the arithmetic supports. To run a Rubin package at the batch where its arithmetic is fully occupied you need 398 concurrent sequences resident on it, against 292 on Blackwell.</p><p>One refinement falls out of the same correction. The invariance is not across all precisions on this part. SemiAnalysis reports that the third generation <strong>Transformer Engine doubles tensor core width for FP4 and FP8 only</strong>, leaving BF16 and TF32 to scale about 1.6 times. P&#183;b is therefore no longer one constant: B* is 398 at FP8 and at NVFP4, and near 164 at BF16. On Blackwell it was 292 at all three. Anyone still serving in BF16 is on a machine with a very different knee from the one on the slide.</p><p>An LPX rack is rated at 315 PFLOPS of FP8 across 40 PB/s of SRAM. B* equals 3.94. Per chip the answer is identical, because the ratio is scale invariant: 1.23 PFLOPS against 156 TB/s gives the same 3.94.</p><p>The ratio between the two machines is 101.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>The benchmark model is not incidental</h2><p>Gemma 4 31B was released by <strong>Google DeepMind on 2 April 2026</strong> under Apache 2.0 license. Its configuration is public: 60 layers, of which 50 are sliding window attention with a window of 1,024 and 10 are full global attention, 32 query heads, 16 KV heads, a head dimension of 256 in the sliding layers and 512 in the global ones.</p><p>The Gemma 4 technical report describes two memory optimisations on the global layers. <strong>Keys are re-used as values</strong>, and position is encoded with p-RoPE at p equal to 0.25. It then states that this reduces the global KV cache by 37.5 percent, without saying how.</p><p>That constant is derivable, and deriving it tells you what is actually stored. If you keep K and V separately you store 2d bytes per head per token. If values equal keys you would store d, a 50 percent cut. </p><p>But K carries rotary position and V must not, so the shared tensor can only be the unrotated part. With <strong>p-RoPE</strong> only a fraction p of the head dimension is rotated, so what you keep is the full unrotated tensor plus a rotated copy of the p fraction:</p><p><code>stored / baseline = d(1 + p) / 2d = (1 + 0.25) / 2 = 0.625</code></p><p>which is a <em>37.5 percent reduction,</em> exactly. The residual against the published figure is zero. So the global layers store 1.25 tensors where a conventional model stores 2.</p><p>One step here is entirely from me and Lorenzo Tettamanti and not the config&#8217;s. The published file gives <strong>16 KV heads</strong>, a head dimension of 256 and a global head dimension of 512, and does not state a separate global head count. </p><p>I take it to be 8, on the reasoning that 8 heads at 512 is the same 4,096 element width as 16 at 256, which is the pattern the 26B variant follows and the only one under which the report&#8217;s 37.5 percent lands on a round byte figure. The <strong>whole cache number below </strong>rests on that inference, so it is marked as an inference in the dossier rather than buried here.</p><p>I flag one disagreement. Several community analyses of this model use <strong>8,192 bytes per token per global layer</strong>, which corresponds to a clean 50 percent reduction and omits the rotated copy. The report&#8217;s own 37.5 percent implies 10,240. I use the report.</p><p>Run the cache arithmetic at 100K of context and the sliding layers contribute 0.84 gigabytes, capped by their 1,024 token window and therefore constant in context length, while the global layers contribute 10.24. The <strong>total is 11.1 gigabytes.</strong> A conventional model with the same head count and 60 full attention layers would need 98.3. The hybrid saves 8.87 times at 100K and 9.31 times at the model&#8217;s full 262,144.</p><p>Weights plus cache at 100K in FP8 is 41.8 gigabytes, which is 30.4 percent of the LPX rack&#8217;s 137.4 gigabytes. In <strong>BF16 it is 52.7 percent. </strong>The benchmark fits with room to spare, and it fits because Gemma 4 is close to the best case a 128 gibibyte SRAM machine could ask for: a dense model small enough to shard 256 ways, with an attention design that caps its own cache growth.</p><p>Work out what would not fit. With 100K of its own cache resident, the largest dense model an LPX rack holds is <strong>126 billion parameters at FP8 or 253 billion at FP4. </strong></p><p>The two trillion parameter figure NVIDIA quotes is 1,000 gigabytes at FP4, or 7.3 racks of SRAM, and the figure in their blog that carries it is labelled a projection of a <strong>scaled up GPT-OSS,</strong> not a measurement. That is honest labelling on their part and I want to repeat it rather than let it blur.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>The term the formula drops</h2><p>Here is where we have to correct our own framing, because the formula above is incomplete and the incompleteness is not small.</p><blockquote><p><em>B* comes from a decode step that reads W bytes of weights and performs 2NB flops. That accounting is right only if the KV cache is negligible. It is not. </em></p></blockquote><p>At batch B and context S a decode step reads W plus B times KV(S) bytes, because every resident sequence&#8217;s cache is read once per token, while the flops that batching amortises are<strong> still only 2NB</strong>. Attention&#8217;s own flops scale with context exactly as its bytes do, so attention has fixed arithmetic intensity and never becomes compute bound at any batch.</p><p>Restore the term and <strong>set compute time</strong> equal to memory time:</p><p><code>2NB / P = (W + B&#183;KV) / Bmem<br>B (2N&#183;Bmem/P &#8722; KV) = W<br>B* = b / (2&#183;Bmem/P &#8722; KV/N)</code></p><p>This reduces to the weights-only form when KV goes to zero. It also has a pole. When <strong>KV per sequence</strong> reaches 2&#183;B<sub>mem</sub>/P per parameter, the denominator vanishes and B* goes to infinity: past that context, no batch size makes the machine compute bound, because each additional sequence adds more memory traffic than it adds amortisable arithmetic.</p><p>Call that<strong> the saturation context.</strong> It is the number I should have led with.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!wwcm!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F54f6c52a-4a3f-4a72-8e7a-1167b09de760_1520x919.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!wwcm!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F54f6c52a-4a3f-4a72-8e7a-1167b09de760_1520x919.png 424w, https://substackcdn.com/image/fetch/$s_!wwcm!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F54f6c52a-4a3f-4a72-8e7a-1167b09de760_1520x919.png 848w, https://substackcdn.com/image/fetch/$s_!wwcm!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F54f6c52a-4a3f-4a72-8e7a-1167b09de760_1520x919.png 1272w, https://substackcdn.com/image/fetch/$s_!wwcm!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F54f6c52a-4a3f-4a72-8e7a-1167b09de760_1520x919.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!wwcm!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F54f6c52a-4a3f-4a72-8e7a-1167b09de760_1520x919.png" width="1456" height="880" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/54f6c52a-4a3f-4a72-8e7a-1167b09de760_1520x919.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:880,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Figure 1&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 1" title="Figure 1" srcset="https://substackcdn.com/image/fetch/$s_!wwcm!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F54f6c52a-4a3f-4a72-8e7a-1167b09de760_1520x919.png 424w, https://substackcdn.com/image/fetch/$s_!wwcm!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F54f6c52a-4a3f-4a72-8e7a-1167b09de760_1520x919.png 848w, https://substackcdn.com/image/fetch/$s_!wwcm!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F54f6c52a-4a3f-4a72-8e7a-1167b09de760_1520x919.png 1272w, https://substackcdn.com/image/fetch/$s_!wwcm!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F54f6c52a-4a3f-4a72-8e7a-1167b09de760_1520x919.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Figure 4.</strong> Critical batch with the KV term restored, as a function of context, for Gemma 4 31B with its cache in BF16. Dotted lines are the weights-only knees the standard formula gives. Dashed lines are the saturation contexts where the denominator vanishes and no batch makes the machine compute bound. Both curves use dense peak arithmetic.</figcaption></figure></div><p>For a Rubin package the headroom is 2.51 millibytes of KV per parameter, which on a 30.7 billion parameter model is 77 megabytes. Gemma 4&#8217;s cache passes<strong> 77 megabytes at 84 tokens of context.</strong> Eighty four. Past that, a Rubin GPU serving this model is memory bound at every batch size there is.</p><p>For the LPX rack the headroom is 254 millibytes per parameter, or 7.8 gigabytes of cache, which Gemma 4 reaches at about 68,000 tokens. The <strong>ratio between the two saturation contexts is 811</strong>, not the 101 that the weights-only knees give, because the sliding window caps 50 of the 60 layers and makes the relationship nonlinear.</p><p>One input to that carries more weight than I would like. The published config gives 16 KV heads and a <strong>512 dimension global head</strong>, but does not state a separate global head count, and I read it as 8 on the constant width argument above. </p><p>Take the literal 16 instead and the cache doubles to 21.3 gigabytes at 100K, Rubin&#8217;s saturation moves only from 84 tokens to 75, and the LPX&#8217;s halves from 68,000 to 34,000. The <strong>ratio becomes 451 rather than 811</strong>. Every qualitative claim survives that swing and the LPX figure is good to within a factor of two, which is the honest precision on it.</p><p>Run it across every machine in the table and the picture is worse than a two way comparison suggests.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!0FeK!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc59063d7-9c34-48aa-893e-489778f64e54_903x540.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!0FeK!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc59063d7-9c34-48aa-893e-489778f64e54_903x540.png 424w, https://substackcdn.com/image/fetch/$s_!0FeK!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc59063d7-9c34-48aa-893e-489778f64e54_903x540.png 848w, https://substackcdn.com/image/fetch/$s_!0FeK!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc59063d7-9c34-48aa-893e-489778f64e54_903x540.png 1272w, https://substackcdn.com/image/fetch/$s_!0FeK!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc59063d7-9c34-48aa-893e-489778f64e54_903x540.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!0FeK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc59063d7-9c34-48aa-893e-489778f64e54_903x540.png" width="903" height="540" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c59063d7-9c34-48aa-893e-489778f64e54_903x540.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:540,&quot;width&quot;:903,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:26881,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/208654603?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc59063d7-9c34-48aa-893e-489778f64e54_903x540.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!0FeK!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc59063d7-9c34-48aa-893e-489778f64e54_903x540.png 424w, https://substackcdn.com/image/fetch/$s_!0FeK!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc59063d7-9c34-48aa-893e-489778f64e54_903x540.png 848w, https://substackcdn.com/image/fetch/$s_!0FeK!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc59063d7-9c34-48aa-893e-489778f64e54_903x540.png 1272w, https://substackcdn.com/image/fetch/$s_!0FeK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc59063d7-9c34-48aa-893e-489778f64e54_903x540.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!DxI0!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F202361c9-edee-4e96-8fef-017e4112c741_1520x919.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!DxI0!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F202361c9-edee-4e96-8fef-017e4112c741_1520x919.png 424w, https://substackcdn.com/image/fetch/$s_!DxI0!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F202361c9-edee-4e96-8fef-017e4112c741_1520x919.png 848w, https://substackcdn.com/image/fetch/$s_!DxI0!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F202361c9-edee-4e96-8fef-017e4112c741_1520x919.png 1272w, https://substackcdn.com/image/fetch/$s_!DxI0!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F202361c9-edee-4e96-8fef-017e4112c741_1520x919.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!DxI0!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F202361c9-edee-4e96-8fef-017e4112c741_1520x919.png" width="1456" height="880" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/202361c9-edee-4e96-8fef-017e4112c741_1520x919.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:880,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Figure 9&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 9" title="Figure 9" srcset="https://substackcdn.com/image/fetch/$s_!DxI0!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F202361c9-edee-4e96-8fef-017e4112c741_1520x919.png 424w, https://substackcdn.com/image/fetch/$s_!DxI0!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F202361c9-edee-4e96-8fef-017e4112c741_1520x919.png 848w, https://substackcdn.com/image/fetch/$s_!DxI0!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F202361c9-edee-4e96-8fef-017e4112c741_1520x919.png 1272w, https://substackcdn.com/image/fetch/$s_!DxI0!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F202361c9-edee-4e96-8fef-017e4112c741_1520x919.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Figure 5.</strong> Saturation context by machine, log scale, computed for Gemma 4 31B with its cache in BF16. The band marks the range every GPU in the table falls into. Past its own line a machine is memory bound at every batch size.</figcaption></figure></div><p>Every GPU lands inside a hundred token band, across four years and three process nodes. The trend within it runs the wrong way. <strong>H200 is the best of them at 162 tokens</strong>, because it added bandwidth without adding arithmetic. Rubin is the worst at 84. </p><p>Four generations of progress moved the saturation context down 26 percent, which is the same statement as saying that compute has outrun bandwidth, made in units that matter for serving rather than in FLOPS.</p><p>So the honest version of this article&#8217;s central claim is not that the LPX has a lower knee. It is that <strong>at the context lengths the machine is sold for, the LPX still has a knee and no GPU does</strong>. Not this one, and not any of the five before it.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>What that does to the tax</h2><p>The <strong>interactivity tax</strong> I was about to compute, hardware carried per token at batch b against the same machine at its knee, doesn&#8217;t survive this. Below saturation, throughput is linear in batch. </p><p>Above it, throughput saturates at B<sub>mem</sub> divided by KV no matter how much batch you add. At 100K of context both machines are above saturation, so the binding constraint stops being the knee and becomes capacity.</p><p>Work it through at 100K. Gemma 4&#8217;s cache is 11.08 gigabytes per sequence. A Vera Rubin NVL72 holds 20.7 terabytes, so after weights it fits about 1,869 concurrent sequences and tops out <strong>near 143,000 tokens per second</strong> for the rack. An LPX rack holds 128 gibibytes, so after weights it fits 9.6 sequences.</p><p>Now compare each machine at its own best point. The Rubin rack at 1,869 sequences produces 143,000 tokens per second and gives each user 76. The LPX rack, measured, produces <strong>11,000 tokens per second </strong>and gives each user 3,431. The GPU rack makes 13 times the tokens. The LPU rack makes each user&#8217;s tokens arrive 45 times faster.</p><p>That is the trade, at one context length, with no marketing basis anywhere in it. And it is a <strong>far smaller hardware-amortisation story</strong> than the one I started with: going from the LPX&#8217;s operating batch of 3.2 up to the Rubin rack&#8217;s capacity limit improves tokens per rack-second by 1.86 times, not by 124. The weights-only figure overstates the tax at 100K context by a factor of 67.</p><p>Which leaves the LPX&#8217;s case resting almost entirely on latency rather than on hardware amortisation. That happens to be what the rest of the arithmetic says too.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share The Software Frontier&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share The Software Frontier</span></a></p><div><hr></div><h2>Where the token actually goes</h2><p>Here is where the arithmetic stops flattering the machine.</p><p>Gemma 4 31B has 30.7 billion parameters. <strong>At FP8 that is 30.7 gigabytes of weights</strong>, and at 100K of context its KV cache is 11.1 gigabytes, which I derive below. A decode step at batch one reads 41.8 gigabytes. </p><p>At 40 PB/s that takes 1.04 microseconds, which is a roofline of 957,422 tokens per second.</p><p>The measured figure is 3,431. The gap is 279 times.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Zav5!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8aa4ae77-daf1-4844-bcb4-9c7d94c82071_1520x919.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Zav5!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8aa4ae77-daf1-4844-bcb4-9c7d94c82071_1520x919.png 424w, https://substackcdn.com/image/fetch/$s_!Zav5!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8aa4ae77-daf1-4844-bcb4-9c7d94c82071_1520x919.png 848w, https://substackcdn.com/image/fetch/$s_!Zav5!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8aa4ae77-daf1-4844-bcb4-9c7d94c82071_1520x919.png 1272w, https://substackcdn.com/image/fetch/$s_!Zav5!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8aa4ae77-daf1-4844-bcb4-9c7d94c82071_1520x919.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Zav5!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8aa4ae77-daf1-4844-bcb4-9c7d94c82071_1520x919.png" width="1456" height="880" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/8aa4ae77-daf1-4844-bcb4-9c7d94c82071_1520x919.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:880,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Figure 3&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 3" title="Figure 3" srcset="https://substackcdn.com/image/fetch/$s_!Zav5!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8aa4ae77-daf1-4844-bcb4-9c7d94c82071_1520x919.png 424w, https://substackcdn.com/image/fetch/$s_!Zav5!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8aa4ae77-daf1-4844-bcb4-9c7d94c82071_1520x919.png 848w, https://substackcdn.com/image/fetch/$s_!Zav5!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8aa4ae77-daf1-4844-bcb4-9c7d94c82071_1520x919.png 1272w, https://substackcdn.com/image/fetch/$s_!Zav5!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8aa4ae77-daf1-4844-bcb4-9c7d94c82071_1520x919.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Figure 6.</strong> The 291.5 microsecond token budget at 3,431 tokens per second. Streaming 41.8 GB of weights and cache at 40 PB/s is 1.04 microseconds; 61.4 GFLOP at 315 PFLOPS is 0.195. The remainder is fixed cost that cannot be attributed further without the chip mapping, which is not published. The measurement is end to end over a network, so this bounds the machine&#8217;s internal latency from above.</figcaption></figure></div><p>The <strong>token budget is 291.5 microseconds. </strong>Streaming the weights and the cache accounts for 1.04 of it. All of the arithmetic, 61.4 gigaflops at 315 PFLOPS, accounts for 0.195. </p><p>Together, every operation that does useful work occupies <strong>0.43 percent of the budget. </strong>The other 99.57 percent is synchronisation, scheduling, and the fixed cost of moving small tensors between chips.</p><p>This is not a criticism of the design, at all. It is what NVIDIA&#8217;s own engineers say the machine is for. Their blog post on the benchmark spends its technical section on a single equation, the <strong>time to move data between two chips as A plus N over B</strong>, and argues that at high interactivity the fixed startup cost A dominates because N is tiny. My arithmetic says they are right.</p><p>How far it can be decomposed depends on the mapping, and here I have to withdraw something. An earlier version of this piece assumed pure tensor parallelism across all 256 chips with two all-reduces per layer, and derived a per-collective cost from it. </p><p>Groq&#8217;s own documentation says the machine runs <strong>pipeline parallelism layered on top of tensor parallelism</strong>: a layer is split across a group of chips and layer groups are pipelined. Without the mapping, the collective count is unknown and the per-collective figure was not supportable.</p><p>What survives does not need the mapping. Layers are a sequential dependency chain however each one is placed, so the per-layer budget is 4.86 microseconds. Streaming that layer&#8217;s share of weights and cache is 17 nanoseconds of it. <strong>Its arithmetic is 3 nanoseconds.</strong> Moving the activation between chips once, 8.2 kilobytes at the chip&#8217;s 1.34 terabytes per second of link bandwidth, is 6 nanoseconds. Add all three and you have accounted for 0.5 percent of the layer.</p><p>The pipeline structure makes this worse rather than better at the operating point being advertised. Pipelining buys throughput by overlapping stages across concurrent work; it buys a single stream nothing, because one token must still traverse every stage in order. At a concurrency of three, a 256 chip pipeline is running close to empty and paying its full depth on every token.</p><p>One caveat I cannot resolve without hardware. The <strong>3,431 figure is an end-to-end median</strong> measured by Artificial Analysis over a network, so some unknown fraction of the 291.5 microseconds is client side and never touches the rack. </p><p>My decomposition therefore bounds the machine&#8217;s internal latency from above rather than measuring it. The direction of the error makes the machine look worse than it is, and the conclusion, that this is a latency product and not a bandwidth product, survives any plausible correction.</p><p>The consequence for anyone sizing hardware is direct. At batch one this rack achieves a <strong>model FLOP utilisation of 0.067 percent.</strong> At its measured concurrency of 3.21 it achieves 0.21 percent. You are buying 315 PFLOPS and 40 PB/s in order to use two thousandths of them, and the entire value proposition is that the two thousandths arrive quickly.</p><p>There is one more gap worth naming. At <strong>100K of context the rack has room for 9.6 sequences</strong>, and if it ran all of them at the advertised 3,431 tokens per second it would produce 33,000 tokens per second rather than the 11,000 NVIDIA quotes. </p><p>It is leaving three times its own capacity ceiling unused. That is what a pipeline looks like when filling it would cost the latency the product is sold on.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/p/when-batching-stops-working/comments&quot;,&quot;text&quot;:&quot;Leave a comment&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/p/when-batching-stops-working/comments"><span>Leave a comment</span></a></p><div><hr></div><h2>The context tax</h2><p><em>Artificial Analysis </em>ran the same suite <strong>at 10K and at 100K on 20 August.</strong> The LPX went from 3,382 to 3,431 tokens per second, which is faster at ten times the context. The fastest public endpoint went from 1,402 to 870.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!MUVp!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F42ea13ca-de48-4304-8615-6c46d44f7a7a_1520x919.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!MUVp!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F42ea13ca-de48-4304-8615-6c46d44f7a7a_1520x919.png 424w, https://substackcdn.com/image/fetch/$s_!MUVp!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F42ea13ca-de48-4304-8615-6c46d44f7a7a_1520x919.png 848w, https://substackcdn.com/image/fetch/$s_!MUVp!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F42ea13ca-de48-4304-8615-6c46d44f7a7a_1520x919.png 1272w, https://substackcdn.com/image/fetch/$s_!MUVp!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F42ea13ca-de48-4304-8615-6c46d44f7a7a_1520x919.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!MUVp!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F42ea13ca-de48-4304-8615-6c46d44f7a7a_1520x919.png" width="1456" height="880" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/42ea13ca-de48-4304-8615-6c46d44f7a7a_1520x919.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:880,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Figure 8&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 8" title="Figure 8" srcset="https://substackcdn.com/image/fetch/$s_!MUVp!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F42ea13ca-de48-4304-8615-6c46d44f7a7a_1520x919.png 424w, https://substackcdn.com/image/fetch/$s_!MUVp!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F42ea13ca-de48-4304-8615-6c46d44f7a7a_1520x919.png 848w, https://substackcdn.com/image/fetch/$s_!MUVp!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F42ea13ca-de48-4304-8615-6c46d44f7a7a_1520x919.png 1272w, https://substackcdn.com/image/fetch/$s_!MUVp!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F42ea13ca-de48-4304-8615-6c46d44f7a7a_1520x919.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Figure 7.</strong> Median output tokens per second at two context lengths, measured by Artificial Analysis on 20 August 2026. The dashed line is what the decode byte count alone predicts for the public endpoint: a 22.1 percent slowdown. The observed slowdown is 37.9 percent.</figcaption></figure></div><p>My byte model says a decode step at 100K reads 1.283 times what it reads at 10K, because the <strong>sliding layers are window-capped</strong> and only the ten global layers grow. </p><p>A purely bandwidth bound machine would therefore slow by 22.1 percent. The <strong>LPX moved by plus 1.4 percent</strong>, which is inside measurement noise and confirms directly that bandwidth is not what limits it. The public endpoint slowed by 37.9 percent, which is 1.72 times what the byte count alone predicts.</p><p>That excess is the part I find most useful for practitioners. Whatever is costing the GPU endpoint an extra 16 points of throughput at long context is not the KV bytes. It is paging, chunked prefill interference, attention kernel behaviour at low batch, or a scheduler shrinking the effective batch as sequences lengthen. </p><p>Those are <strong>software problems with software fixes</strong>, and they are roughly as large as the hardware difference the whole LPX programme exists to address.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Who the four times is actually against</h2><p>NVIDIA&#8217;s headline is that the LPX produced 3,431 tokens per second against 870 for the next fastest public endpoint, a factor of 3.94. It does not name the endpoint. </p><p>ServeTheHome&#8217;s live coverage from the session says the comparison looks like a Cerebras CS-3, and the provider data supports that: <strong>Cerebras </strong>is the <strong>fastest benchmarked provider for this model at 1,191.7 tokens per second</strong>, and it is the only one whose figure is in the right range to degrade to 870 at 100K.</p><p>If that identification is right, then NVIDIA&#8217;s four times is against another SRAM machine. Against the fastest GPU-served endpoint for the same weights, which is <strong>Modular&#8217;s NVFP4 deployment at 233 tokens per second</strong>, the LPX is 14.7 times faster. Cerebras is itself 5.1 times faster than that endpoint.</p><p>We both think the correct reading is not that NVIDIA beat the field. It&#8217;s that there are two regimes, that the <strong>SRAM machines occupy one of them and every GPU occupies the other</strong>, and that NVIDIA has just spent 20 billion dollars to be present in both. </p><p>The interesting comparison in the deck is not the bar chart. It is the pareto curve NVIDIA showed at Hot Chips, where its own presenter conceded that total throughput efficiency falls as more work moves to the LPUs, and that GPUs remain the efficient choice whenever latency is negotiable.</p><div class="community-chat" data-attrs="{&quot;url&quot;:&quot;https://open.substack.com/pub/softwarefrontier/chat?utm_source=chat_embed&quot;,&quot;subdomain&quot;:&quot;softwarefrontier&quot;,&quot;pub&quot;:{&quot;id&quot;:3575776,&quot;name&quot;:&quot;The Software Frontier&quot;,&quot;author_name&quot;:&quot;Lorenzo Bradanini&quot;,&quot;author_photo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!ACM6!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F18342bf1-31cb-404a-b9e1-998a38d299bf_1200x1200.jpeg&quot;}}" data-component-name="CommunityChatRenderPlaceholder"></div><div><hr></div><h2>If you are serving models rather than buying them</h2><p>A few things in the arithmetic above are actionable without any new hardware.</p><ul><li><p><strong>The largest is the context tax</strong>. The fastest public GPU endpoint loses 37.9 percent of its throughput going from 10K to 100K of context, and the decode byte count only accounts for 22.1 of those points. The remaining 16 points are software: paging, chunked prefill interference, attention kernel behaviour at low batch, or a scheduler quietly shrinking the effective batch as sequences lengthen. That gap is roughly as large as the hardware difference this entire programme exists to address, and it is available to anyone willing to profile a long-context decode path this week.</p></li><li><p><strong>The second is that above the saturation context, batching stops buying throughput.</strong> Every serving guide tells you to raise batch until you are compute bound. On a Rubin package running a 31 billion parameter model with Gemma 4&#8217;s attention geometry, that point is at 84 tokens of context, which means it does not exist for any real workload. What batch still buys above saturation is capacity utilisation, not arithmetic utilisation, and the ceiling is memory bandwidth divided by cache bytes per sequence. For this model on an NVL72 that ceiling is about 143,000 tokens per second and no scheduler setting will beat it.</p></li></ul><p>The corollary is that <strong>quantising the KV cache is worth more than quantising the weights</strong> once you are above saturation, because the ceiling is set by cache bytes. Halving KV precision raises the ceiling by two. Halving weight precision moves it not at all.</p><ul><li><p><strong>The third is where to look for cache savings. </strong>Gemma 4&#8217;s hybrid attention makes its cache 8.87 times smaller at 100K than a conventional design with the same head count, and almost all of that comes from ten global layers being the only ones that grow with context. If you are sizing a KV budget, the layer type distribution matters more than the head count, and a model with a 5:1 local to global ratio is a fundamentally different capacity problem from one without.</p></li></ul><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>What the market pays for speed</h2><p><strong>Gemma 4 31B </strong>is a very good instrument for a price question because the <strong>weights are identical everywhere</strong>. Thirteen or so providers serve the same Apache 2.0 checkpoint. Any price difference between them is a price on the machine, not on the model.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!xp1R!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd4d99ab9-583d-494f-b46d-be4a27647f25_1520x919.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!xp1R!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd4d99ab9-583d-494f-b46d-be4a27647f25_1520x919.png 424w, https://substackcdn.com/image/fetch/$s_!xp1R!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd4d99ab9-583d-494f-b46d-be4a27647f25_1520x919.png 848w, https://substackcdn.com/image/fetch/$s_!xp1R!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd4d99ab9-583d-494f-b46d-be4a27647f25_1520x919.png 1272w, https://substackcdn.com/image/fetch/$s_!xp1R!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd4d99ab9-583d-494f-b46d-be4a27647f25_1520x919.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!xp1R!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd4d99ab9-583d-494f-b46d-be4a27647f25_1520x919.png" width="1456" height="880" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d4d99ab9-583d-494f-b46d-be4a27647f25_1520x919.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:880,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Figure 5&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 5" title="Figure 5" srcset="https://substackcdn.com/image/fetch/$s_!xp1R!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd4d99ab9-583d-494f-b46d-be4a27647f25_1520x919.png 424w, https://substackcdn.com/image/fetch/$s_!xp1R!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd4d99ab9-583d-494f-b46d-be4a27647f25_1520x919.png 848w, https://substackcdn.com/image/fetch/$s_!xp1R!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd4d99ab9-583d-494f-b46d-be4a27647f25_1520x919.png 1272w, https://substackcdn.com/image/fetch/$s_!xp1R!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd4d99ab9-583d-494f-b46d-be4a27647f25_1520x919.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Figure 8.</strong> Output price against measured output speed for one set of weights, Gemma 4 31B under Apache 2.0. The band spans the published output prices (0.34 to 0.38); the three ticks mark published speeds. Only DeepInfra publishes both, so the band is a price range and a speed range for the same regime, not four matched pairs. The LPX point is extrapolated from a two-point fit and is not a published price.</figcaption></figure></div><p>The GPU-served providers span 31.5 to 233 tokens per second, a factor of 7.4. Their<strong> output prices are 0.34 at CoreWeave, 0.38 at DeepInfra, and 0.34 to 0.35 through OpenRouter</strong>: a band of 12 percent. Fitting an elasticity to that gives 0.056, which for practical purposes is zero. Inside the GPU regime, seven times the speed is free.</p><p>The only price step in the market is at the regime boundary. Cerebras charges 0.99 per million input tokens and 1.49 per million output, against 0.08 and 0.38 at the cheapest tracked provider. </p><p>That card is confirmed rather than assumed: Artificial Analysis publishes a blended rate of 1.04 for this model on Cerebras, and 0.99 and 1.49 reproduce it exactly under their seven to two to one weighting with no cache discount. That is <strong>4.1 times the commodity output price for 5.1 times the fastest GPU speed</strong>, and 12.4 times on input.</p><p>Two things follow that we didnt expect before running the numbers.</p><blockquote><ol><li><p><em>First, the elasticity computed across the full range, 0.376, is an artifact of fitting to the two extremes. It is an upper bound on what speed is worth, not a description of the market, because the intermediate points do not lie on it. Anyone modelling a fast inference business off a smooth speed premium is modelling a curve that does not exist. What exists is a commodity floor and one occupied premium tier.</em></p></li><li><p><em>Second, input tokens carry the larger premium. Cerebras marks up input 12.4 times and output 3.9 times. That is the opposite of the intuition that a fast decoder should charge for decode, and it makes sense the moment you remember that a machine with no DRAM has to hold the prefill working set somewhere expensive.</em></p></li></ol></blockquote><p>Extrapolating the <strong>two-point curve to the LPX&#8217;s 3,431 tokens per second</strong>, which is 2.88 times Cerebras, gives an implied price near 2.22 dollars per million output tokens and 2.06 on input. </p><p>NVIDIA has not published a price card, so this is the market&#8217;s shape and not the company&#8217;s intention. I use it as a scenario, and I flag that it rests on the same two-point fit I just described as an upper bound.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Where the revenue actually sits</h2><p>Now put that against the workload NVIDIA chose to sell the machine with. From their blog: an agentic coding turn in which the <strong>model reads hundreds of files</strong>, exceeds 100K of context, and generates 5,000 tokens of reasoning and output.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!HVro!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd993a93d-fa40-4b46-b4ed-c96fc35e73d9_1520x919.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!HVro!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd993a93d-fa40-4b46-b4ed-c96fc35e73d9_1520x919.png 424w, https://substackcdn.com/image/fetch/$s_!HVro!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd993a93d-fa40-4b46-b4ed-c96fc35e73d9_1520x919.png 848w, https://substackcdn.com/image/fetch/$s_!HVro!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd993a93d-fa40-4b46-b4ed-c96fc35e73d9_1520x919.png 1272w, https://substackcdn.com/image/fetch/$s_!HVro!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd993a93d-fa40-4b46-b4ed-c96fc35e73d9_1520x919.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!HVro!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd993a93d-fa40-4b46-b4ed-c96fc35e73d9_1520x919.png" width="1456" height="880" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d993a93d-fa40-4b46-b4ed-c96fc35e73d9_1520x919.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:880,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Figure 6&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 6" title="Figure 6" srcset="https://substackcdn.com/image/fetch/$s_!HVro!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd993a93d-fa40-4b46-b4ed-c96fc35e73d9_1520x919.png 424w, https://substackcdn.com/image/fetch/$s_!HVro!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd993a93d-fa40-4b46-b4ed-c96fc35e73d9_1520x919.png 848w, https://substackcdn.com/image/fetch/$s_!HVro!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd993a93d-fa40-4b46-b4ed-c96fc35e73d9_1520x919.png 1272w, https://substackcdn.com/image/fetch/$s_!HVro!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd993a93d-fa40-4b46-b4ed-c96fc35e73d9_1520x919.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Figure 9.</strong> Share of request revenue by token type for NVIDIA&#8217;s own worked example: an agentic coding turn with more than 100K of input context generating 5,000 tokens. The rightmost bar bills a 226-turn session with cached input at ten percent of list.</figcaption></figure></div><p>Price that turn on the commodity card and 19.2 percent of the revenue is output tokens. Price it on the Cerebras card and 7.0 percent is. Price it on the extrapolated LPX card and 5.1 percent is.</p><p>This matters because of how the rack is wired. In <strong>all three co-execution configurations NVIDIA describes</strong>, prefill happens on the Vera Rubin GPUs. In prefill-decode disaggregation the GPU rack hands over the KV cache once per turn. </p><p>In attention-FFN disaggregation the GPU computes attention and holds the cache in DRAM while the LPUs run the feed forward layers. In external-drafter speculative decoding the <strong>LPUs run a draft model and the GPUs verify.</strong> In every case the input tokens, which are 81 to 95 percent of the revenue in NVIDIA&#8217;s own example, are billed against work the GPU rack does.</p><p>The LPU accelerates the small end of the bill and all of the latency. For a product whose promise is user experience that is exactly right. <strong>For a claim of ten times more revenue per watt it&#8217;s a great problem</strong>, and NVIDIA&#8217;s own footnote on that claim says as much: the figure is projected from an estimated cost-per-million-tokens tiered pricing model. It is a statement about a price card that does not exist yet.</p><p>I wanted to offer prefix caching as the counterweight here, and the arithmetic will not let me. Agentic sessions reuse their context, and NVIDIA&#8217;s own figure one shows a session growing across 226 turns. </p><p>Bill the first turn&#8217;s input at full rate and later turns at a tenth, and the <strong>output share rises from 7 percent to 42</strong>, which would be a different business and the business the LPX is actually for.</p><p>But the provider in question does not sell that. Artificial Analysis publishes a blended rate of 1.04 dollars per million for this model on Cerebras, and 0.99 in with 1.49 out reproduces it under their seven to two to one weighting only if cache hits are billed at the same 0.99 as fresh input. </p><p>Solve for the cache price and it comes back at exactly 0.99. There is no discount. So<strong> the 226 turn session bills at a 7.0 percent output share</strong>, the same as a single turn, and the counterweight I was reaching for is hypothetical on every endpoint I can price.</p><p>If a fast tier ever does offer cache pricing, the 42 percent figure is what it would look like, and that is the number to watch. It is also a business in which the cache has to live somewhere, and the somewhere is HBM on the GPU rack, which is the component under the 279 billion dollar commitment.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>What the rack has to cost</h2><p>11,000 tokens per second is 347 billion tokens per rack-year at full utilisation. At 70 percent utilisation over a four year life, at the extrapolated 2.22 dollars per million output tokens, the <strong>rack grosses 2.16 million dollars.</strong></p><p>That is the ceiling on what it can cost, before power, before the paired Rubin capacity that does its prefill, and before anyone earns a margin. At the Cerebras card of 1.49 it is 1.45 million. At commodity output pricing of 0.38 it is 369,000 dollars, which would not pay for the rack&#8217;s power.</p><p><strong>NVIDIA does not publish an LPX rack power figure</strong>, which means the 35x throughput per megawatt and 10x revenue per watt claims cannot be checked by anyone outside the company. I would treat both as unverified until a number appears. Everything else in the deck was specified to three significant figures.</p><p>So the arithmetic closes only under a conjunction: the premium tier has to hold at nearly three times Cerebras&#8217;s speed, the workload has to be cached enough that output is a real share of the bill, and the <strong>rack has to land under roughly two million dollars.</strong> None of those is absurd. All three have to be true at once.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>What we really think this is about</h2><p>Both of our reading is that the LPX is not primarily a product. It&#8217;s a hedge on a component price.</p><p><strong>NVIDIA has committed 279 billion dollars to memory</strong>, front loaded into three years, with a gross margin guide already bending under it. Memory is the input it cannot substitute, cannot second source beyond three vendors, and cannot make itself. </p><p>In that position, owning a machine whose bandwidth comes from a<strong> 4nm logic wafer rather than a stacked DRAM package </strong>is worth something independent of whether the machine sells well, because it is the only lever that converts a purchased input into a manufactured one.</p><p>The <strong>Rubin CPX supports this reading.</strong> It was announced in September 2025 as a GDDR7-based accelerator for the context phase, and it has quietly left the roadmap, displaced by the LPX. GDDR7 is still DRAM you have to buy. SRAM is area you already pay TSMC or Samsung for.</p><p>The thing I keep returning to is that the entire architecture is legible from a single number.</p><p><strong>3,810&#215; more bandwidth per byte</strong>, combined with bandwidth that is <strong>642&#8211;2,924&#215; cheaper per unit</strong> and memory that is <strong>1.4&#8211;6.4&#215; more expensive per byte</strong>, pushes the saturation context out by a factor of <strong>811&#215;</strong>. From there, the rest of the architecture follows almost inevitably.</p><p>It favors a small model. It puts attention back on the GPU because the cache cannot fit. It makes a hybrid-attention model, whose cache is roughly nine times smaller than a conventional one, the natural benchmark. It requires tensor parallelism across 256 dies. And it ultimately demands a compiler capable of scheduling <strong>320-byte transfers to the clock cycle</strong>, because at batch three, the fixed cost of moving the data is effectively the cost of generating the token itself.</p><p>That is what makes the machine interesting. These are not a collection of independent design choices. They are consequences of the same economic constraint propagating through the stack.</p><p><strong>It is a coherent machine because it is solving a procurement problem.</strong></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Where this analysis might be wrong</h2><p>The <strong>concurrency of 3.21 </strong>comes from dividing a rack throughput figure by a per-user figure that may have been measured at a different operating point. </p><p>If the <strong>11,000 tokens per second is a peak throughput number</strong> at relaxed latency rather than the aggregate at 3,431 per user, then the true concurrency at the advertised speed is lower, not higher, and the interactivity tax argument gets stronger while the revenue arithmetic gets worse. I have used the reading least favourable to my own conclusion.</p><p>The SRAM cost model uses a 5nm class bit cell on a Samsung 4nm part. If <strong>Samsung&#8217;s density is materially worse</strong>, the per-gigabyte figure rises and the per-terabyte-per-second figure rises with it. It would take roughly a factor of 900 to change the sign of the bandwidth comparison, so the conclusion is not sensitive, but the 66 dollars per gigabyte is a model and not a measurement.</p><p>The identification of <strong>Cerebras as the unnamed baseline is inference from a third party&#8217;s live blog plus provider data.</strong> If the 870 figure is some other endpoint, my point about the four times being an SRAM-to-SRAM comparison weakens.</p><p>The 28 dollar per gigabyte reconciled memory price rests on a GPU average selling price band I chose. Narrow the band and the answer moves. <strong>I reported the range </strong>rather than the point for that reason, and the fact that it lands inside the published analyst band is a check on the method, not a proof of the point estimate.</p><p>And the price extrapolation is the weakest link in the piece. I have one premium data point. One point does not make a tier.</p><p>The session revenue model originally billed repeat input at a tenth of list, assuming a prefix cache discount, and reported a 42 percent output share on that basis. My own price check rules it out: the published blended rate only reconciles if cache hits cost the same as fresh input. The 42 percent is now labelled hypothetical and the measured case is 7.0.</p><p>The latency decomposition originally attributed the <strong>unexplained 99.5 percent to a specific count of tensor parallel collectives.</strong> Groq documents pipeline parallelism on top of tensor parallelism, which makes the collective count a function of a mapping I do not have, so that attribution is gone and only the topology independent per-layer budget remains.</p><p>The larger of the two things I got wrong is in the section above: I built the argument on a critical batch formula that drops the KV term, which is fine for comparing machines and wrong for describing serving. </p><p>Restoring it moved the <strong>interactivity tax at 100K context from 124 times to 1.86 times</strong> and replaced the whole framing with the saturation context. I have left the derivation of the error in the text rather than quietly fixing the numbers, because the weights-only formula is in wide circulation and the pole is the part nobody mentions.</p><p>The other one, worth naming rather than paraphrasing, is that the first version of this piece computed Rubin&#8217;s critical batch from the <strong>marketed 50 PFLOPS NVFP4 figure</strong> and every other machine&#8217;s from a dense one. </p><p>That produced 568 instead of 398, a break of 1.94 times instead of 1.36, and an interactivity tax of 177 instead of 124. It is the <strong>same denominator mixing I have criticised in other people&#8217;s analysis</strong>, and the harness did not catch it because the harness only checked that the arithmetic followed from the constants, never that a constant meant what its name said. Every arithmetic figure now carries a declared basis and the code refuses to divide across a mismatch.</p><p>The <strong>LP30 rack&#8217;s 315 PFLOPS carries no stated basis either.</strong> I read it as dense because the LP30 is a deterministic machine with no published structured sparsity mode. If it turns out to be a sparse figure, B* halves to 1.97 and every argument here gets stronger, so the reading I chose is again the unfavourable one.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Five dated predictions</h2><ol><li><p><em>NVIDIA publishes an LPX rack power figure before the end of FY2027, and the 35x per megawatt claim is restated against a specific model and batch when it does.</em></p></li><li><p><em>The first public LPX price card, whenever it appears, prices input tokens at a larger multiple of commodity than it prices output tokens, following the Cerebras shape rather than the intuitive one.</em></p></li><li><p><em>Supply and capacity commitments fall below 200 billion dollars within four quarters, because the three year front loading in this table is a one time securing of a position rather than a new run rate.</em></p></li><li><p><em>By the end of 2027 at least one frontier open weights model ships with an attention design chosen explicitly so that its cache fits in an on-package or on-die SRAM budget, and the release notes say so.</em></p></li><li><p><em>The next NVIDIA datacenter generation after Rubin lands with a dense B* above 450. Rubin moved the knee 1.36 times in one step while HBM4 bandwidth grew 2.75 times, and the compression engine is the tell: when a vendor starts marketing a compressed peak, the dense ratio is the number under pressure.</em></p></li></ol><div><hr></div><h2>Corrections</h2><p>This is the final version of this massive research project. What changed so far during the journey:</p><ol><li><p><strong>The headline number.</strong> Rubin&#8217;s critical batch was computed from the marketed 50 PFLOPS NVFP4 figure, which is compressed rather than dense, while every other machine used a dense figure. B* was 568 and is 398. The generational break was stated as 1.94 times and is 1.36. The interactivity tax was 177 and is 124. The ratio between the two machines was 144 and is 101.</p></li><li><p><strong>The fifth prediction has been replaced.</strong> It said no generation through 2028 would bring B* under 400. The corrected figure is 398, so it was false on arrival.</p></li><li><p><strong>A secondary source contradicted a primary one.</strong> An earlier draft said Data Center Networking grew 138 percent. The CFO commentary has no networking line and 138 percent belongs to the AI Clouds, Industrial and Enterprise segment.</p></li><li><p><strong>An inference was presented as a fact.</strong> The count of 8 global KV heads in Gemma 4 31B is not in the published config. It is now disclosed where it is used and carried at tier C.</p></li><li><p><strong>A mislabelled quantity.</strong> The 4,096 in the collective payload calculation is the KV width, not the hidden size, which is not published for this checkpoint.</p></li><li><p><strong>A figure caption overclaimed.</strong> Figure 5 said the band spanned providers with both a price and a speed published. Only one provider publishes both.</p></li><li><p><strong>Five historical B* values were imported rather than derived.</strong> All five now re-derive inside the harness from datasheet dense FP8 and published bandwidth.</p></li><li><p><strong>The HBM4 price was on the wrong basis.</strong> Version 2 used 15.28 dollars per gigabyte throughout and called it a contract price. It is a factory gate component price. The contract price projected for 2026 is 31 to 32 dollars per gigabyte to NVIDIA and 35 to 36 to other buyers. Both bases are now shown. The bandwidth advantage of SRAM widens from 914x to a range of 914x to 1,885x, and the capacity penalty narrows from 4.3x to a range of 2.1x to 4.3x.</p></li><li><p><strong>The triangulation is now checked against the right number.</strong> The reconciled 28.2 dollars per gigabyte was compared to the factory gate figure and reported as 1.85 times it. Compared to the published analyst contract price it is within 11 percent, from two methods sharing no inputs, which is a far stronger result than the one originally claimed.</p></li><li><p><strong>Two unverifiable attributions were removed.</strong> A job title given for the Hot Chips presenter that no source in hand supports, and an implied claim that the Rubin package carries eight HBM4 stacks, which 288 gigabytes does not uniquely determine.</p></li><li><p><strong>The dense figure is now confirmed from NVIDIA&#8217;s own rack table</strong> rather than from a single outlet: 3,600 PFLOPS NVFP4 inference against 2,520 dense NVFP4 and 1,260 dense FP8, which is ten to seven and 17.5 PFLOPS per package.</p></li><li><p><strong>The session revenue model assumed a cache discount the provider does not offer.</strong> It billed repeat input at a tenth of list and reported a 42 percent output share. Solving the published blended rate for the cache hit price returns exactly the input price, so there is no discount. The measured case is 7.0 percent and the 42 is now labelled hypothetical. The error ran in the direction that flattered the argument being made.</p></li><li><p><strong>The one tier C input carrying a headline now ships with its sensitivity.</strong> The global KV head count is inferred, not published. Under the literal reading the LPX saturation context halves to 34,000 and the ratio falls from 811 to 451. Both are in the text.</p></li><li><p><strong>The latency decomposition assumed a topology that is not documented.</strong> It attributed the unexplained part of the token budget to a specific count of tensor parallel collectives across all 256 chips. Groq documents pipeline parallelism layered on top of tensor parallelism, so the collective count depends on a mapping that is not published. The per-collective figure is withdrawn. What remains is the per-layer budget of 4.86 microseconds, of which streaming, arithmetic and one chip-to-chip handoff together account for 0.5 percent, and that holds under any mapping.</p></li><li><p><strong>The rack break-even was given as a point.</strong> It now runs as three scenarios from 369,000 dollars to 2.16 million, because the price extrapolation behind the single figure rests on a two point fit the article itself describes as an upper bound.</p></li><li><p><strong>The central framing rested on a formula that drops the KV term.</strong> The standard critical batch B* = P&#183;b/(2&#183;B<sub>mem</sub>) assumes a decode step reads only weights. Restoring the cache term gives B* = b / (2&#183;B<sub>mem</sub>/P &#8722; KV/N), which has a pole. Past the saturation context no batch makes the machine compute bound. For Gemma 4 31B that is 84 tokens on a Rubin package and 68,000 on an LPX rack. The interactivity tax at 100K context falls from 124 times to 1.86, an overstatement of 67 times, and the article is now built on the saturation context instead. The comparison between the machines survives; the framing did not.</p></li><li><p><strong>The SRAM cost model used TSMC figures for a Samsung part.</strong> A TSMC N5 cell on a TSMC-priced wafer understated the cell area by about a quarter and overstated the wafer price by about the same, so the point estimate survived by accident rather than by method. It is now a band across cell size, array efficiency and wafer price: 44 to 97 dollars per gigabyte, centrally 64. The bandwidth advantage becomes a range of 642 to 2,924 times rather than a single figure.</p></li><li><p><strong>The margin guide runs further than reported.</strong> Version 3 stopped at the 74.0 percent Q3 guide. On the call the CFO said margin falls to 71 or 72 percent by the January quarter before recovering to 72 or 73. The trough is three to four hundred basis points, not one hundred.</p></li><li><p><strong>The SRAM bandwidth divisor contradicted the article&#8217;s own argument.</strong> The cost model divided by the trade press figure of 150 TB/s per die while the text argues the derived 156.25 is the correct one. Now 156.25 throughout.</p></li><li><p><strong>The harness gained a basis guard.</strong> Every arithmetic constant declares dense or sparse, and the code refuses to compute B* or a ratio across a mismatch. This is the check whose absence caused the first error.</p></li></ol><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><p></p>
      <p>
          <a href="https://www.thesoftwarefrontier.com/p/when-batching-stops-working">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[Distilling in depth ROCm: How it Actually Works]]></title><description><![CDATA[AMD ships 2,905,048 tuned GEMM decisions in a public git repository. Zero of CDNA 4&#8217;s 75 matrix instructions exist on CDNA 5.]]></description><link>https://www.thesoftwarefrontier.com/p/distilling-in-depth-rocm-how-it-actually</link><guid isPermaLink="false">https://www.thesoftwarefrontier.com/p/distilling-in-depth-rocm-how-it-actually</guid><dc:creator><![CDATA[Lorenzo Bradanini]]></dc:creator><pubDate>Tue, 25 Aug 2026 14:30:10 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!NY69!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4feb9b78-5127-49e8-824b-5f7fac0fa6cc_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!NY69!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4feb9b78-5127-49e8-824b-5f7fac0fa6cc_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!NY69!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4feb9b78-5127-49e8-824b-5f7fac0fa6cc_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!NY69!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4feb9b78-5127-49e8-824b-5f7fac0fa6cc_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!NY69!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4feb9b78-5127-49e8-824b-5f7fac0fa6cc_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!NY69!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4feb9b78-5127-49e8-824b-5f7fac0fa6cc_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!NY69!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4feb9b78-5127-49e8-824b-5f7fac0fa6cc_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/4feb9b78-5127-49e8-824b-5f7fac0fa6cc_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:3473148,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/212666933?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4feb9b78-5127-49e8-824b-5f7fac0fa6cc_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!NY69!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4feb9b78-5127-49e8-824b-5f7fac0fa6cc_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!NY69!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4feb9b78-5127-49e8-824b-5f7fac0fa6cc_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!NY69!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4feb9b78-5127-49e8-824b-5f7fac0fa6cc_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!NY69!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4feb9b78-5127-49e8-824b-5f7fac0fa6cc_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>Introduction</h2><p>I&#8217;ve never had a<strong> MI300X</strong> in front of me. I&#8217;m perfectly aware that&#8217;s the wrong way to start a piece about ROCm, but it&#8217;s also the exact same reason this one needs to exists. </p><p>Every time I&#8217;ve tried to <strong>read seriously about AMD&#8217;s</strong> software stack I have hit the same wall: the writing is either marketing, or it is a benchmark table, or it is a bring-up worklog by someone who rented <em>eight cards for a week</em> and is understandably more interested in getting a model to run than in explaining what the machine is. </p><p>What I wanted was the thing I&#8217;d want for any other system I do not own: a description precise enough that I could predict what it would do.</p><p>You can get closer to that than you would think without hardware, because of a <strong>property of ROCm</strong> nobody right now seems to exploit. The stack is open down to the instruction encodings, the compiler that targets it is upstream LLVM, and the compiler is an x86 program. </p><p>Everything here was produced on a single-core container with no GPU attached: <strong>clang 20.1.2 and 22.1.0</strong> from the official builds, and about three hundred megabytes of YAML cloned out of a public AMD repository. </p><p>I just compiled real code objects for <strong>five AMD datacenter targets</strong>, decoded their kernel descriptors byte by byte, and counted the entire shipped tuning surface of <strong>AMD&#8217;s GEMM library.</strong></p><p>One result surprised me a lot. From gfx90a to gfx942, every matrix instruction AMD had carried forward. In the case of gfx942 to gfx950, every single one carried forward again. </p><p>Then, from gfx950 to gfx1250, the <strong>compute die inside the MI455X</strong> that AMD launched into the <em>Helios rack</em> in July, the number that carries forward is zero.</p><p>I repeat: it&#8217;s not reduced to a numbr near zero. It&#8217;s just zero of seventy five, and the wavefront changes width at the same time. Most of this piece is about the <strong>context needed</strong> to say why that matters and what it does not mean, because the obvious reading is wrong in an interesting way.</p><p>The second thing is smaller but I think it&#8217;s more durable. <strong>AMD&#8217;s openness </strong>has been taken as a virtuous example for a decade, as though the interesting question were whether a stack is open. The interesting question is what openness lets you count. </p><p>Here it lets you count the kernel layer exactly: 16,177 generated assembly kernels and 2,905,048 measured mappings from problem shape to kernel, in a git repository anyone can clone. </p><p>There is no counterpart on the other side. NVIDIA&#8217;s equivalent is not secret so much as uncountable, because there is no file.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>What ROCm is, in the order the electrons see it</h2><p>In my personal findings, I kept seeing that the word ROCm is used for at least <strong>three different things</strong>: a kernel driver, a runtime and compiler toolchain, and a large collection of libraries that happen to be distributed together. </p><p>Talking about &#8220;<em>ROCm performance</em>&#8221; without saying which of the three you mean is how most arguments about <strong>AMD software go wrong</strong>. Here is the stack, bottom to top, with the piece of the problem each layer owns.</p><ul><li><p>At the bottom we see <code>amdgpu</code>, the kernel driver, which has been in the <em>mainline Linux tree</em> since 2015 and is the same driver that runs a gaming Radeon. Inside it lives the <strong>KFD</strong>, the compute component, which owns the queues, the doorbells, the page tables and the notion of a process having a GPU address space. </p></li><li><p>Above it in user space sits libhsakmt, historically called the <strong>Thunk</strong>, a thin ioctl wrapper, and above that ROCr, the HSA runtime, which implements a specification written by a foundation AMD co-founded in 2012 for a world that never quite arrived, and which left behind an unusually well documented dispatch model.</p></li><li><p>On top of ROCr sits <strong>HIP</strong>, which is two things wearing one name.<em> HIP, the language,</em> is a near copy of CUDA C++ with the identifiers renamed. <em>HIP the runtime</em> is a library exposing an API that is a near copy of the CUDA driver and runtime APIs with the identifiers renamed. The compiler is not a fork of anything proprietary; it is clang, with a device-side target of <code>amdgcn-amd-amdhsa</code>, plus a bitcode library of device functions called <strong>ROCm Device Libs</strong> where the math lives.</p></li><li><p>Above HIP sit the libraries, and this is where most of the engineering hours are: rocBLAS and hipBLASLt for GEMM, MIOpen for convolutions and some fused primitives, <strong>Composable Kernel </strong>as a template layer for writing fused operators, RCCL for collectives, and since 2025 AITER, which is AMD&#8217;s answer to the observation that a serving engine does not want a BLAS, it wants attention and mixture-of-experts kernels. </p></li><li><p>Above that sit the frameworks, and above those the serving engines, <strong>vLLM </strong>and <strong>SGLang</strong>, which is where a customer&#8217;s dollar actually meets the machine.</p></li></ul><p>Ten layers, counted that way. For any performance claim about AMD, the useful question is which layer it is a claim about, because they differ wildly in maturity and in rate of change. </p><p>The driver is boring and solid, and the runtime is the same. The compiler is upstream LLVM and is very good. The library layer is where the variance lives, and the variance is enormous.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>The dispatch packet is a struct, and you can read it</h2><p>Start with the thing that is most unusual about ROCm, because it sets up everything else.</p><p>When a <strong>CUDA program launches a kernel</strong>, nobody truly knows what happens between <code>&lt;&lt;&lt;&gt;&gt;&gt;</code> and the SM starting work. There is a channel, there is a doorbell, there is a scheduler; the formats are entirely IP and the only supported way to produce them is to call <em>NVIDIA&#8217;s runtime. </em></p><p>On AMD, kernel dispatch is a <strong>64 byte structure</strong> whose every field is in a published specification, written by user space into a ring buffer that user space allocated, and signalled by a store to a doorbell page that the driver mapped into the process.</p><p>The structure is the <strong>AQL kernel dispatch packet.</strong> It carries a header with the packet type and two memory fence scopes, the workgroup dimensions and grid dimensions as three 16 bit and three 32 bit fields, the sizes of the private and group segments, a pointer to the kernel object, a pointer to the kernarg buffer, and a completion signal handle. </p><p>A queue is a ring of these packets plus a read index and a write index. Submitting work is: write the packet, bump the write index, store to the doorbell.</p><p><strong>Two fields in that structure</strong> carry most of the interesting behaviour. The barrier bit, when set, says that this packet may not begin until every preceding packet in the same queue has completed, which is what makes an HSA queue behave like a CUDA stream. </p><p>The two fence scope fields say what memory ordering the hardware must establish at packet start and packet end, with values for no fence, agent scope, and system scope. Those two fields are the reason a stream on AMD has an ordering cost that is visible and adjustable rather than implicit.</p><p>The practical consequence is a line buried in AMD&#8217;s tuning documentation recommending <code>GPU_MAX_HW_QUEUES=2</code>, with the note that hardware efficiency is maximised at four or fewer HIP streams. </p><p>A HIP stream is not free the way an abstraction is free: it <strong>maps to a hardware queue</strong>, which are a finite resource arbitrated by the command processor, and oversubscribing them makes the scheduler do work that surfaces as launch latency. </p><p>The same document says plainly that ROCm serialises kernel launches across GPUs from one process, which is why <strong>RCCL wants one process per GPU</strong>. Elsewhere these would be folklore. Here they are consequences of a dispatch model you can read.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Anatomy of an AMD code object</h2><p>Now let&#8217;s compile something. This is a <strong>single MFMA loop</strong>, written in C with clang&#8217;s AMDGPU builtins, with no ROCm installation anywhere on the machine:</p><pre><code><code>typedef float    f32x16 __attribute__((ext_vector_type(16)));
typedef _Float16 f16x4  __attribute__((ext_vector_type(4)));
#define G __attribute__((address_space(1)))

__attribute__((amdgpu_kernel))
__attribute__((amdgpu_flat_work_group_size(256,256)))
void gemm_tile(G f32x16 *out, G f16x4 *a, G f16x4 *b, int n) {
  f32x16 acc = (f32x16)(0.0f);
  for (int k = 0; k &lt; n; ++k)
    acc = __builtin_amdgcn_mfma_f32_32x32x8f16(a[k], b[k], acc, 0, 0, 0);
  out[0] = acc;
}</code></code></pre><pre><code><code>clang --target=amdgcn-amd-amdhsa -mcpu=gfx942 -O3 -nogpulib \
      -fuse-ld=lld -o ko_gfx942.hsaco ko.c</code></code></pre><p>That produces <strong>exactly 4,848 bytes.</strong> What it creates is an ELF64 shared object, and not a container format, not a fat binary with a vendor magic number at the front; it&#8217;s a shared object, with <code>OS/ABI: AMDGPU_HSA</code>, <code>Machine: EM_AMDGPU</code>, fifteen sections and thirteen symbols, which <code>llvm-readobj</code> and <code>readelf</code> parse without being told anything special.</p><p>The target identity lives in the <strong>ELF header flags</strong>, and this is the first place the AMD and NVIDIA philosophies visibly diverge:</p><pre><code><code>Flags [ (0x54C)
  EF_AMDGPU_FEATURE_SRAMECC_ANY_V4 (0x400)
  EF_AMDGPU_FEATURE_XNACK_ANY_V4   (0x100)
  EF_AMDGPU_MACH_AMDGCN_GFX942     (0x04C)
]</code></code></pre><p>Three orthogonal things are encoded there. The machine, gfx942, is the ISA.<strong> XNACK</strong> is whether the code was built to tolerate page faults and retry memory operations, which matters for unified memory. </p><p><strong>SRAMECC</strong> is whether the code was built for a part with ECC on the on-chip SRAM, which costs registers. Each has three states: on, off, and any, where any means the object is compatible with either setting of the machine it lands on.</p><p>Now, let&#8217;s try with our means to compare with what I found when I looked at NVIDIA&#8217;s side of this in my earlier work. There, <strong>architecture-specific features</strong> are folded into the target name itself, which is how you get sm_90a and sm_100a, and a target with the <code>a</code> suffix is not forward compatible in the way plain PTX is. </p><p>AMD&#8217;s scheme is basically the same problem solved by making the feature axis explicit and <strong>orthogonal to the ISA version</strong>, with a documented neutral value. </p><p>Generally speaking, we&#8217;re looking at a better design, and it is worth saying that clearly because it&#8217;s one of the places where the open stack is not merely as good as the closed one, but it&#8217;s an <strong>order of magnitude better.</strong></p><p>Inside the object, each kernel has two symbols: the code, and a 64 byte kernel descriptor in <code>.rodata</code> named <code>&lt;kernel&gt;.kd</code>. The descriptor is what the command processor reads to set up a wave.</p><p> Here is the <strong>one clang emitted for gfx942</strong>, byte for byte, and decoded:</p><pre><code><code>00 00 00 00 00 00 00 00  1c 00 00 00 00 00 00 00
40 11 00 00 00 00 00 00  00 00 00 00 00 00 00 00
00 00 00 00 00 00 00 00  00 00 00 00 82 00 af 00
84 00 00 00 08 00 00 00  00 00 00 00 00 00 00 00

group_segment_fixed_size   = 0
private_segment_fixed_size = 0
kernarg_size               = 28
kernel_code_entry_offset   = 4416
compute_pgm_rsrc1          = 0x00af0082
compute_pgm_rsrc2          = 0x00000084
compute_pgm_rsrc3          = 0x00000000
kernel_code_properties     = 0x0008
kernarg_preload            = 0x0000</code></code></pre><p>The <strong>rsrc words are hardware register images.</strong> <code>rsrc1</code> bits 5:0 hold the granulated VGPR count, which is 2 here, meaning an allocation of 24 registers for a kernel that uses 20, because the encoding granule on this part is eight. </p><blockquote><p><em>Bits 9:6 hold the granulated SGPR count. Bits 17:16 and 19:18 hold the denormal mode for 32 bit and for 16 and 64 bit floats independently, both set to 3 here, meaning flush nothing. Bit 23 is IEEE mode.</em></p></blockquote><p><code>rsrc2</code> bits 7, 8 and 9 enable the workgroup id in SGPRs per dimension, and bits 12:11 say how many workitem id dimensions arrive in<strong> VGPR0</strong> through VGPR2. <code>rsrc3</code> bits 5:0 are <code>ACCUM_OFFSET</code>, the register number where the accumulator half of the file begins, and bit 16 is <code>TG_SPLIT</code>. </p><p>Here the field is 0, so AGPRs start at register 4, and the metadata agrees: 20 registers total of which 16 are AGPRs. <strong>Four architectural registers</strong> for addresses, sixteen for the accumulator, and a descriptor field that says where the boundary is.</p><p>That descriptor is the complete contract between compiler and hardware scheduler for how a wave is configured. It is documented, stable, byte-addressable, and I read it with a<strong> general purpose ELF tool.</strong> The corresponding structure exists on the other side of the market, but you meet it, if at all, through a disassembler the vendor ships. </p><p>For anyone building tooling, a profiler or a binary rewriter or a security scanner, that difference is the whole project. On AMD those are ordinary programs.</p><p>Three targets, one source. The<strong> code objects </strong>are all 4,848 bytes with 1,472 bytes of <code>.text</code>, and the text differs between gfx90a and gfx942 in 113 of those bytes, and between gfx942 and gfx950 in 33. </p><p>That is what an incremental ISA revision looks like from the outside: the same program, the same schedule, a small number of re-encoded opcodes.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>The hazard contract, which is the whole argument in miniature</h2><p>Here is the thing I keep coming back to, because it is the most clear statement of what is actually different between the two stacks, and almost nobody frames it this way.</p><p>On <strong>modern NVIDIA hardware</strong>, the dependency interlock for fixed-latency instructions is not in the silicon. It is in a control field that the assembler writes into each instruction, and getting it wrong produces wrong answers rather than slow ones. </p><p>That makes the assembler <strong>correctness-critical</strong>, and it makes the assembler&#8217;s closure a consequence of an architectural choice about where to spend area and energy, not a policy anyone can simply reverse.</p><p>AMD made the opposite choice, and it is visible in every kernel the compiler emits. Waiting is an instruction. <code>s_waitcnt</code> takes counters, <code>vmcnt</code> for vector memory, <code>lgkmcnt</code> for <strong>LDS and scalar and message traffic</strong>, <code>expcnt</code> for exports, and the operand says how many outstanding operations of that class you are willing to leave in flight. </p><p>If you write <code>s_waitcnt vmcnt(0)</code> you have waited for all of them; <code>vmcnt(1)</code> means you will proceed with one still outstanding. In my <strong>MFMA loop</strong> the compiler emits exactly two distinct forms:</p><pre><code><code>s_waitcnt lgkmcnt(0)
s_nop 4</code></code></pre><p>The <code>s_nop</code> is the second half of the contract. </p><p><strong>Matrix instructions on CDNA</strong> have structural hazards with specific, published cycle counts: after an MFMA writes an accumulator, a certain number of cycles must pass before particular kinds of reader can touch it, and if the compiler cannot fill those slots with useful work it must fill them with nothing, explicitly, by emitting a no-op with a repeat count. </p><p><strong>LLVM has a whole pass for this</strong>, the hazard recogniser, and its rules are in the source tree.</p><p>So both vendors moved hazard management out of hardware and into software. That is the shared fact, and it is a much more interesting one than &#8220;<em>AMD is open</em>&#8221;. </p><p>The difference is where in software. NVIDIA put it in a field inside the instruction word, written by a<strong> program you cannot read</strong>, which means the hazard model is enforced by a compiler. </p><p>AMD put it in separate architectural instructions, in a <strong>published ISA</strong>, which means the hazard model is enforced by the program text itself and anybody can write it.</p><p>Follow that difference to its economic conclusion and you get something non-obvious. Because <strong>AMD&#8217;s hazard model is expressible</strong>, hand-written assembly is a viable production strategy on AMD in a way it is not on NVIDIA. And AMD uses it. </p><p>The <strong>GEMM library&#8217;s kernels</strong> are generated assembly. AITER&#8217;s fastest attention paths are hand-written assembly. This is not a stopgap; it is the design. The openness did not remove the labour of getting near the metal. </p><p>It moved the labour from a compiler team to a kernel team, and made it possible to <strong>do that work outside AMD</strong>, which is a real and underrated property. What it did not do is make the work smaller.</p><p>That is the shape of the whole argument, and the rest of this piece is an attempt to put numbers on it.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Waves, SIMDs, and why 64 is not 32 twice</h2><p>One more piece of groundwork, because it is the most common source of wrong intuitions when a <strong>CUDA programmer reads AMD code</strong>.</p><p>A <strong>CDNA compute unit</strong> has four SIMD units, each sixteen lanes wide, and a wavefront is 64 work items. So one wavefront instruction occupies its SIMD for four cycles, which is the<strong> classic GCN cadence</strong>, and the natural unit of divergence is 64 rather than 32. Every lane mask is a 64 bit value in a pair of scalar registers.</p><p>The consequences propagate a long way up. A <strong>CUDA warp</strong> shuffle moves data among 32 lanes; the AMD equivalents, <code>ds_bpermute_b32</code> and the DPP and permlane families, work across 64 and do not map one to one. </p><p>A reduction written for warp size 32 does not become correct by changing a constant. A kernel assuming <code>__activemask()</code> returns 32 bits is assuming a fact about hardware, not about the language.</p><p>Alongside the vector units sits a fully separate scalar unit with its own register file and cache, and <strong>AMD code is full of scalar instructions</strong> doing address arithmetic, loop control and uniform values that on NVIDIA occupy vector registers. </p><p>That is a <strong>slight advantage</strong> and the reason the register pressure story differs: uniform work has somewhere else to live.</p><p>On gfx1250 the wave becomes 32 wide, the compute unit becomes a workgroup processor in the RDNA style, and the scalar unit&#8217;s role changes. We will get there.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Four ways to touch memory, and why the ugly one wins</h2><p>A CDNA kernel can address memory through four instruction families, and which one the compiler picks is among the larger performance levers in the stack. </p><p>It&#8217;s also where <strong>reading AMD assembly</strong> diverges most from reading NVIDIA assembly, and where a CUDA programmer&#8217;s instincts mislead.</p><p>The <code>flat_</code> family takes a 64 bit address in a pair of vector registers and works out from the address whether it means global memory, LDS or scratch. Convenient and expensive: two registers per lane, plus an aperture check. </p><p>The <code>global_</code> family takes the same 64 bit address but promises it is global, dropping the check and keeping the two registers. The <code>ds_</code> family addresses <strong>LDS with a 32 bit offset</strong> and is the only way to touch the scratchpad.</p><p>Then the <code>buffer_</code> family, which is the interesting one. A buffer instruction takes a resource descriptor of four scalar registers holding a base address, a byte count, a stride and some format fields, plus a 32 bit offset per lane. </p><p>The descriptor lives in scalar registers, shared by all 64 lanes at a cost of zero vector registers, and the byte count means that the<strong> hardware does the bounds check</strong>: an access past the end returns zero for loads and discards stores, with no branch.</p><p>So a <code>buffer_load_dwordx4</code> needs one vector register for the offset where a <code>global_load_dwordx4</code> needs two for the address. On a kernel with a dozen live pointers into a tile that is a dozen registers per lane, in an architecture whose currency is registers. </p><p>The <strong>free bounds</strong> check is worth more still, deleting the masking and predication a tiled GEMM epilogue needs for its ragged edges.</p><p>Hence AMD&#8217;s tuning advice preferring <code>buffer_load</code>, Triton&#8217;s AMD backend having a pass that converts <strong>pointer arithmetic</strong> into buffer operations when it can prove the offsets fit in 32 bits, and hand-written AMD kernels being full of them. </p><p>It is an idea that is simply better than the alternative and that nobody outside the AMD world discusses, because the alternative is what everyone learned first.</p><p>The instruction that AMD&#8217;s own guide singles out as the one that moved a <strong>reference GEMM by sixty percen</strong>t is in this family: <code>buffer_load_to_lds</code>, which reads global memory and writes LDS without the data ever entering a register. </p><p>Thats the same idea as <strong>NVIDIA&#8217;s asynchronous copy</strong>, just from a different angle, and its value is measured in registers returned to the accumulator rather than in latency hidden.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>The wait counters, in detail</h2><p>I said earlier that waiting is an instruction. It is worth being precise about what the counters count, because the <strong>semantics are unusual</strong> and they explain a class of AMD performance bug that looks like nothing on a profile.</p><p>Three counters are visible to a wave. <code>vmcnt</code> counts outstanding vector memory operations. <code>lgkmcnt</code> counts a grab bag: <strong>LDS, GDS, scalar memory reads and messages</strong>. <code>expcnt</code> counts exports and matters mostly for graphics.</p><p>The crucial property is that <code>vmcnt</code> returns in order and <code>lgkmcnt</code> does not. Vector memory completes in issue order, which is what makes <code>s_waitcnt vmcnt(3)</code> meaningful: three loads may still be in flight and you know exactly which three. </p><p>Scalar memory inside <code>lgkmcnt</code> can complete out of order, which is why you almost always see <code>lgkmcnt(0)</code> and <strong>almost never a nonzero value</strong> where scalar loads are involved. The compiler is not being lazy; it cannot express the thing you want.</p><p>That shapes how a software pipeline is written here. Prefetch depth is fine grained on the vector path and all or nothing on the scalar one, so <strong>AMD kernels push everything they can</strong> into vector memory even when the data is uniform, and reserve the scalar path for values consumed once at the top of a loop. </p><p>It also explains a specific pathology: <strong>mix LDS traffic and scalar loads inside a loop</strong> and every <code>s_waitcnt lgkmcnt(0)</code> for an LDS dependency drains the scalar loads too, needed or not. Hoisting them out fixes it, and the reason it works is a counter aliasing decision made a decade ago.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/p/distilling-in-depth-rocm-how-it-actually?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/p/distilling-in-depth-rocm-how-it-actually?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><div><hr></div><h2>Occupancy, computed rather than guessed</h2><p>Occupancy on AMD is arithmetic you can do on paper, and AMD publishes the formula, which is not something you can say about every vendor.</p><p>Each <strong>SIMD has 512 VGPRs available</strong> to it in wave64 terms, allocated to waves in blocks of sixteen. That is a different quantity from the granule of eight in the kernel descriptor: eight is how the count is encoded, sixteen is how the hardware hands registers out. </p><p>So a kernel using 170 registers is rounded to 176, and the <strong>number of waves that fit per SIMD</strong> is the floor of 512 over 176, which is two, because three times 176 is 528 and that does not fit. If you can push the kernel to 168 registers you get three waves, a fifty percent increase in latency hiding, for a change of two registers. </p><p>This is the entire reason <code>waves_per_eu</code> exists as a Triton parameter: it is a hint to the register allocator to try harder to land under a threshold that the programmer can compute and the allocator does not know about.</p><p>The workgroup level adds the scratchpad. <strong>Occupancy limited by LDS</strong> is the floor of the LDS size over the kernel&#8217;s allocation, 65,536 bytes on CDNA 3 and 163,840 on CDNA 4. </p><p>And the two limits combine through the number of waves per workgroup:</p><pre><code><code>occ_vgpr = floor(512 / roundup(vgprs, 16))     # waves per SIMD
occ_lds  = floor(LDS_total / lds_per_group)    # groups per CU
occ      = min(floor(occ_vgpr * 4 / nW), occ_lds) * nW / 4</code></code></pre><p>where <code>nW</code> is waves per workgroup and the factor of four is the SIMDs per CU.</p><p>Put the <strong>CDNA 4 numbers</strong> into that and the significance of the LDS change becomes obvious. On CDNA 3, a kernel using 32 KB of LDS gets two workgroups per CU from the scratchpad side, and that is usually the binding constraint for a tiled GEMM. On CDNA 4 the same kernel gets five. </p><p>The register side did not change, 512 either way, which is exactly why AMD&#8217;s own guidance says that<strong> compute-bound GEMM </strong>should use the same tile size on both parts despite the LDS growth: the tile is limited by registers, not by scratchpad. </p><p>What the <strong>extra LDS buys</strong> is not a bigger tile, it is a deeper pipeline, one more stage of prefetch in flight.</p><p>There is a second thing the LDS change buys that is easy to miss. On CDNA 3 the LDS has <strong>two SIMD pairs</strong> each with a 128 byte per clock bus, but the two pairs cannot both access LDS in the same cycle, so the real service rate is 128 bytes per clock. </p><p>CDNA 4 doubles it to 256. And bank conflicts eat into that badly: AMD&#8217;s Gluon guidance says an unpadded, unswizzled shared layout produces <strong>two-way to four-way conflicts</strong> and drops the effective rate to somewhere between 64 and 128 bytes per clock. </p><p>So the padding decision, which appears in the shipped kernel names as <code>LBSPPA</code> and <code>LBSPPB</code> with sixteen and eleven distinct values respectively, is worth <strong>a factor of two to four</strong> on the scratchpad path. That is why it is a tuned parameter rather than a default.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share&quot;,&quot;text&quot;:&quot;Share The Software Frontier&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share The Software Frontier</span></a></p><div><hr></div><h2>Coherence, and the part of the memory model nobody reads</h2><p>The AQL packet&#8217;s two fence scope fields point at something that HIP mostly hides and that matters enormously the moment you have more than one agent.</p><p>ROCm allocations come in two flavours. A coarse-grained allocation is coherent at the <strong>boundaries of a kernel dispatch</strong>: the acquire fence at packet start and the release fence at packet end make it visible, and inside the kernel the GPU may cache it however it likes. </p><p>A fine-grained allocation is coherent at a finer granularity, which on the GPU side means it bypasses or writes through certain caches, and which costs bandwidth. <strong>Host-visible memory</strong> that the CPU is going to poll while a kernel is running has to be fine grained. Model weights should never be.</p><p>Underneath is a memory type field, MTYPE, attached to page table entries, which determines that page&#8217;s caching behaviour from the GPU&#8217;s side. The <strong>runtime picks it from how you allocated.</strong> Getting it wrong does not produce an error. It produces a kernel that is mysteriously bandwidth starved, or a host poll loop that never sees an update.</p><p>XNACK is the other half. Enabled, the GPU can take a page fault, the driver services it, and the <em>memory operation retries</em>, which is what makes unified memory and oversubscription work. </p><p>Disabled, memory must be resident before the kernel runs and the hardware is slightly faster for not carrying the retry machinery. The setting is per boot and per device, which is why a code object built with <strong>XNACK-any</strong> exists at all: it is the build that loads either way.</p><p>The practical version is short. Weights and KV cache want coarse grained device memory; any host visible ring buffer used for scheduling wants to be fine grained and small. </p><p>If a profile shows a fraction of HBM bandwidth that makes no sense, check the allocation before you look at the kernel.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>The collective layer</h2><p>Eight MI300X modules in a node are connected by <strong>seven Infinity Fabric links each</strong>, fully connected, with aggregate peer to peer ring bandwidth of 896 GB/s on MI300X and 1,075 GB/s on MI350X. </p><p>Fully connected means a direct link between every pair, which is a different topology from a switched fabric and yields a different piece of advice.</p><p>The advice, from AMD, is to use either one GPU or all eight and to avoid collectives across two or four. The reason is direct: with all eight participating, <strong>every link in the topology carries traffic</strong>. With four, three quarters of the links are idle and you get a fraction of the potential bandwidth. </p><p>On a switched fabric this is not true in the same way, because the switch does not care which subset you use. So a tensor parallel degree of four, which is a perfectly ordinary choice on an eight way NVIDIA node, is a <strong>worse choice on an eight way AMD node</strong> than the raw bandwidth numbers suggest, and the difference is topological rather than a software deficiency.</p><p><strong>RCCL is NCCL&#8217;s counterpart </strong>and shares its heritage, algorithms and much of its API. The operational notes that matter come from the platform rather than the library: disable NUMA auto-balancing, disable PCIe access control services for multi-node, one process per GPU, and consider raising the channel count for end-to-end workloads. </p><p>That last one, <code>NCCL_MIN_NCHANNELS=112</code>, is striking if you are used to NVIDIA defaults, and it follows from having many direct links rather than a few fat ones.</p><p>Helios moves this to the rack, with <strong>Pensando handling front-end</strong>, scale-up and scale-out, 72 accelerators in the scale-up domain and 31 TB of pooled HBM4. </p><p>That is the<em> bet NVIDIA made with NVL72 two years earlier</em>, and it is the right bet: once a mixture-of-experts model&#8217;s all-to-all exceeds what a node absorbs, the rack becomes the unit of engineering. </p><p>Unfortunately, <strong>I haven&#8217;t independent data</strong> on the AMD version and would treat any number on it as provisional until somebody who sells neither rack measures both.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>The register file is the architecture</h2><p>Last year I spent a pretty long time on Blackwell&#8217;s tensor memory and concluded that <strong>TMEM is not an optimisation</strong>. The largest matrix instruction in the tcgen05 family produces an FP32 accumulator needing 256 registers per thread, and the architectural ceiling on NVIDIA is 255. </p><p>The instruction cannot exist without somewhere else to put its output. TMEM is that somewhere else: a <strong>fifth address space </strong>with its own allocator, its own load and store instructions, its own failure modes and its own compiler-injected guardrail traps.</p><p>What I did not ask at the time, because <strong>I wasn&#8217;t looking at AMD</strong>, is what the other design does about the same problem. So I directly asked the compiler. This kernel holds N independent 32x32x8 accumulators live across a loop, sixteen registers each, and I raised N until it broke.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!t0Rx!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F577e9e9d-1a18-47e0-9057-445de9cb82eb_1600x843.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!t0Rx!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F577e9e9d-1a18-47e0-9057-445de9cb82eb_1600x843.png 424w, https://substackcdn.com/image/fetch/$s_!t0Rx!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F577e9e9d-1a18-47e0-9057-445de9cb82eb_1600x843.png 848w, https://substackcdn.com/image/fetch/$s_!t0Rx!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F577e9e9d-1a18-47e0-9057-445de9cb82eb_1600x843.png 1272w, https://substackcdn.com/image/fetch/$s_!t0Rx!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F577e9e9d-1a18-47e0-9057-445de9cb82eb_1600x843.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!t0Rx!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F577e9e9d-1a18-47e0-9057-445de9cb82eb_1600x843.png" width="1456" height="767" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/577e9e9d-1a18-47e0-9057-445de9cb82eb_1600x843.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:767,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;How much accumulator one CDNA 3 lane can hold before it spills. The whole 512-entry file is allocated at twenty four accumulators and nothing spills. Twenty five is where scratch traffic starts, and thirty is clean again, which is the allocator&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="How much accumulator one CDNA 3 lane can hold before it spills. The whole 512-entry file is allocated at twenty four accumulators and nothing spills. Twenty five is where scratch traffic starts, and thirty is clean again, which is the allocator" title="How much accumulator one CDNA 3 lane can hold before it spills. The whole 512-entry file is allocated at twenty four accumulators and nothing spills. Twenty five is where scratch traffic starts, and thirty is clean again, which is the allocator" srcset="https://substackcdn.com/image/fetch/$s_!t0Rx!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F577e9e9d-1a18-47e0-9057-445de9cb82eb_1600x843.png 424w, https://substackcdn.com/image/fetch/$s_!t0Rx!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F577e9e9d-1a18-47e0-9057-445de9cb82eb_1600x843.png 848w, https://substackcdn.com/image/fetch/$s_!t0Rx!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F577e9e9d-1a18-47e0-9057-445de9cb82eb_1600x843.png 1272w, https://substackcdn.com/image/fetch/$s_!t0Rx!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F577e9e9d-1a18-47e0-9057-445de9cb82eb_1600x843.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong><span>Figure 1.</span></strong> The whole 512-entry file is allocated at twenty four accumulators and nothing spills. Twenty five is where scratch traffic starts, and thirty is clean again, which is the allocator finding a different schedule rather than the file getting bigger. </figcaption></figure></div><p>The shape of that curve is the answer. Up to eight accumulators, 128 registers of live state, everything sits in <strong>ordinary vector registers.</strong> At sixteen the accumulator alone needs 256 registers and the allocator crosses a line: 288 registers total of which 32 are AGPRs, with <code>ACCUM_OFFSET</code> set to 256 in the kernel descriptor. </p><p>At twenty four, <strong>384 registers of live accumulator</strong>, it has allocated the entire 512 entry file with 256 of those entries as AGPRs and still spills nothing. Twenty five is where scratch traffic starts.</p><p>The tail is not monotone, and it is worth saying so because I got it wrong the first time I looked<strong>. Twenty five through twenty nine spill;</strong> thirty does not, allocating 500 registers and reaching 480 of live accumulator; thirty one spills 512 bytes. </p><p>The allocator is finding a different schedule at that one point rather than the file getting bigger. The <em>defensible ceiling is twenty four accumulators.</em></p><p>So the answer is that on CDNA the same problem does not arise. A CDNA lane has 512 architectural registers, not 255. The file is <strong>one physical structure</strong> that the kernel descriptor partitions at a byte offset into a general half and an accumulator half, and matrix instructions can name the accumulator half directly. </p><blockquote><p><em>There is no new address space, no allocator, no separate instruction family for moving results in and out, no class of guardrail trap. The thing NVIDIA had to invent an address space for, AMD had already solved in 2020 by making the register file twice as deep and teaching the descriptor where to cut it.</em></p></blockquote><p>That is not a rhetorical point, and it generalises into a number worth having. If you divide the on-chip state a unit owns by the<strong> matrix throughput</strong> that unit can sustain per clock, you get bytes of working set per unit of arithmetic, which is the ratio that decides whether the accumulator fits.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!uzqG!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd76761cb-9e59-4ba6-8af7-4665c37275f1_1600x885.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!uzqG!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd76761cb-9e59-4ba6-8af7-4665c37275f1_1600x885.png 424w, https://substackcdn.com/image/fetch/$s_!uzqG!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd76761cb-9e59-4ba6-8af7-4665c37275f1_1600x885.png 848w, https://substackcdn.com/image/fetch/$s_!uzqG!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd76761cb-9e59-4ba6-8af7-4665c37275f1_1600x885.png 1272w, https://substackcdn.com/image/fetch/$s_!uzqG!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd76761cb-9e59-4ba6-8af7-4665c37275f1_1600x885.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!uzqG!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd76761cb-9e59-4ba6-8af7-4665c37275f1_1600x885.png" width="1456" height="805" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d76761cb-9e59-4ba6-8af7-4665c37275f1_1600x885.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:805,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Bytes of on-chip state per FP8 FLOP per clock per compute unit. The ratio that decides whether an accumulator can live in registers. NVIDIA has been spending it down for two generations; AMD has not.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Bytes of on-chip state per FP8 FLOP per clock per compute unit. The ratio that decides whether an accumulator can live in registers. NVIDIA has been spending it down for two generations; AMD has not." title="Bytes of on-chip state per FP8 FLOP per clock per compute unit. The ratio that decides whether an accumulator can live in registers. NVIDIA has been spending it down for two generations; AMD has not." srcset="https://substackcdn.com/image/fetch/$s_!uzqG!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd76761cb-9e59-4ba6-8af7-4665c37275f1_1600x885.png 424w, https://substackcdn.com/image/fetch/$s_!uzqG!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd76761cb-9e59-4ba6-8af7-4665c37275f1_1600x885.png 848w, https://substackcdn.com/image/fetch/$s_!uzqG!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd76761cb-9e59-4ba6-8af7-4665c37275f1_1600x885.png 1272w, https://substackcdn.com/image/fetch/$s_!uzqG!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd76761cb-9e59-4ba6-8af7-4665c37275f1_1600x885.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong><span>Figure 2.</span></strong> The ratio that decides whether an accumulator can live in registers. NVIDIA has been spending it down for two generations; AMD has not. </figcaption></figure></div><p>An MI300X compute unit carries roughly 144 bytes of register file and scratchpad for every FP8 FLOP per clock it can issue. An H100 SM carries 58. A B200 SM carries 32. </p><p>NVIDIA has cut that ratio by roughly a factor of four and a half across two generations, because <strong>matrix throughput per SM doubled twice </strong>while the register file did not move at all and shared memory did not move at all. </p><p>AMD&#8217;s CDNA 4 cut its ratio too, from 144 to 85, by doubling matrix throughput per CU while holding the register file at 512 and raising the LDS from 64 KB to 160 KB. It is still two and a half times richer than Blackwell.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!ajYp!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbe924113-8f01-4e38-ad18-1c179ed2edf5_1600x640.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!ajYp!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbe924113-8f01-4e38-ad18-1c179ed2edf5_1600x640.png 424w, https://substackcdn.com/image/fetch/$s_!ajYp!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbe924113-8f01-4e38-ad18-1c179ed2edf5_1600x640.png 848w, https://substackcdn.com/image/fetch/$s_!ajYp!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbe924113-8f01-4e38-ad18-1c179ed2edf5_1600x640.png 1272w, https://substackcdn.com/image/fetch/$s_!ajYp!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbe924113-8f01-4e38-ad18-1c179ed2edf5_1600x640.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!ajYp!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbe924113-8f01-4e38-ad18-1c179ed2edf5_1600x640.png" width="1456" height="582" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/be924113-8f01-4e38-ad18-1c179ed2edf5_1600x640.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:582,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Vector register file per accelerator, megabytes. An MI300X carries four and a half times the register capacity of an H100. That is the area AMD spent instead of inventing a new address space.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Vector register file per accelerator, megabytes. An MI300X carries four and a half times the register capacity of an H100. That is the area AMD spent instead of inventing a new address space." title="Vector register file per accelerator, megabytes. An MI300X carries four and a half times the register capacity of an H100. That is the area AMD spent instead of inventing a new address space." srcset="https://substackcdn.com/image/fetch/$s_!ajYp!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbe924113-8f01-4e38-ad18-1c179ed2edf5_1600x640.png 424w, https://substackcdn.com/image/fetch/$s_!ajYp!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbe924113-8f01-4e38-ad18-1c179ed2edf5_1600x640.png 848w, https://substackcdn.com/image/fetch/$s_!ajYp!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbe924113-8f01-4e38-ad18-1c179ed2edf5_1600x640.png 1272w, https://substackcdn.com/image/fetch/$s_!ajYp!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbe924113-8f01-4e38-ad18-1c179ed2edf5_1600x640.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong><span>Figure 3.</span></strong> An MI300X carries four and a half times the register capacity of an H100. That is the area AMD spent instead of inventing a new address space.</figcaption></figure></div><p>Read the two histories side by side and the divergence is legible. NVIDIA has been spending down <strong>on-chip state per FLOP </strong>as fast as it can and buying it back with mechanisms: asynchronous copy, then the tensor memory accelerator, then a dedicated accumulator memory. </p><p>Each is a way of not paying for registers. AMD kept paying, which is why its kernels can be written in a flatter style and <strong>why an MI300X carries 152 megabytes of vector register file</strong> against an H100&#8217;s 33.</p><p>Neither choice is obviously right. NVIDIA&#8217;s buys throughput per millimetre and pays in programming model complexity, which it absorbs by shipping libraries that hide the complexity, a business it is good at. </p><p><em>AMD&#8217;s buys a simpler programming model and pays in area</em>, which is a reasonable trade if your pitch is that the customer can write the kernel themselves.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Matrix cores, concretely</h2><p>A representative MFMA, <code>v_mfma_f32_32x32x8_f16</code>, computes a <em>32 by 32 by 8 product with FP16 inputs and an FP32 accumulator</em>, executed collectively by one wavefront, with operands and result distributed across lanes in a fixed layout the programmer has to know. </p><p>The A operand is four halves per lane and the accumulator sixteen floats per lane; 64 lanes times sixteen floats is the 1,024 values of the 32 by 32 tile.</p><p>The family is large and grew fast. In <strong>LLVM 22, gfx90a exposes 31 MFMA builtins, gfx942 39, gfx950 47,</strong> with the sparse SMFMAC family going 6, 14, 28 across the same three and conversions going 7, 15, 66. </p><p>CDNA 4&#8217;s contribution is almost entirely data types: block-scaled FP8, FP6 and FP4 in the <strong>OCP microscaling format</strong> with a shared 8 bit exponent per 32 elements, plus the scaled instructions to consume them, of which <code>v_mfma_scale_f32_16x16x128_f8f6f4</code> is the one you meet in a real MXFP4 GEMM.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!K91K!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F981131a1-f5e5-4ffc-8804-f304da4bc156_1600x885.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!K91K!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F981131a1-f5e5-4ffc-8804-f304da4bc156_1600x885.png 424w, https://substackcdn.com/image/fetch/$s_!K91K!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F981131a1-f5e5-4ffc-8804-f304da4bc156_1600x885.png 848w, https://substackcdn.com/image/fetch/$s_!K91K!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F981131a1-f5e5-4ffc-8804-f304da4bc156_1600x885.png 1272w, https://substackcdn.com/image/fetch/$s_!K91K!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F981131a1-f5e5-4ffc-8804-f304da4bc156_1600x885.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!K91K!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F981131a1-f5e5-4ffc-8804-f304da4bc156_1600x885.png" width="1456" height="805" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/981131a1-f5e5-4ffc-8804-f304da4bc156_1600x885.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:805,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;The matrix instruction surface, by family and target. MFMA and WMMA are disjoint. CDNA 5 is the first datacenter part on the WMMA side of the line.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="The matrix instruction surface, by family and target. MFMA and WMMA are disjoint. CDNA 5 is the first datacenter part on the WMMA side of the line." title="The matrix instruction surface, by family and target. MFMA and WMMA are disjoint. CDNA 5 is the first datacenter part on the WMMA side of the line." srcset="https://substackcdn.com/image/fetch/$s_!K91K!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F981131a1-f5e5-4ffc-8804-f304da4bc156_1600x885.png 424w, https://substackcdn.com/image/fetch/$s_!K91K!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F981131a1-f5e5-4ffc-8804-f304da4bc156_1600x885.png 848w, https://substackcdn.com/image/fetch/$s_!K91K!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F981131a1-f5e5-4ffc-8804-f304da4bc156_1600x885.png 1272w, https://substackcdn.com/image/fetch/$s_!K91K!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F981131a1-f5e5-4ffc-8804-f304da4bc156_1600x885.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong><span>Figure 4.</span></strong> MFMA and WMMA are disjoint. CDNA 5 is the first datacenter part on the WMMA side of the line. </figcaption></figure></div><p>Two practical facts about using these that are worth more than a page of theory.</p><p>The first is that the smaller tile usually wins. <strong>AMD&#8217;s own guidance</strong> says the 16x16 form outperforms the 32x32 form on MI300X for GEMM, including at large sizes, and gives the reason as power efficiency rather than issue rate. </p><p>This is the sort of thing that is <strong>invisible from a spec sheet</strong> and decisive in practice, and it is why Triton exposes <code>matrix_instr_nonkdim</code> as a tuning knob at all.</p><p>The second is that on CDNA 4 you should match <code>BLOCK_K</code> to the <strong>instruction&#8217;s K dimension </strong>and aim for one or two matrix instructions per K step. FP16 wants <code>v_mfma_f32_16x16x32</code> with <code>BLOCK_K</code> 64. </p><p>FP8 wants <code>v_mfma_f32_16x16x128</code> with <code>BLOCK_K</code> 128. MXFP4 wants the scaled instruction with <code>BLOCK_K</code> 256. Get that wrong and you pay pipelining overhead on every step of the loop.</p><p>A third fact says more than either about where performance comes from. AMD&#8217;s tuning guide reports that in its own <strong>reference Gluon GEMM</strong> for gfx950, switching the operand path from staging through registers to <code>buffer_load_to_lds</code>, a direct L1 to LDS asynchronous copy, moved the kernel from 697 to 1113 TFLOPS. </p><p>Sixty percent from one instruction selection decision, and the stated mechanism is that it saves <strong>roughly 100 VGPRs per wave</strong> and deletes a register movement phase from the loop. </p><p>The same document reports that remapping workgroup ids so consecutive tiles land on the same XCD cut L2 misses from circa five million to 3.1 million and added another 67 TFLOPS.</p><p>Hold that against the register file. The<strong> async copy is worth sixty percent</strong> because it returns a hundred registers to the accumulator, and registers are what this architecture spends its area on. The architecture and the kernel technique are one fact seen from two sides.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Eight dies pretending to be one GPU</h2><p>An MI300X is not a chip. It is eight compute dies on four I/O dies with eight HBM stacks, and an <strong>MI350X is eight compute dies</strong> on two I/O dies with a faster link between them. This has consequences that leak through every abstraction above it.</p><p>Each XCD has <strong>40 physical compute units</strong> of which 38 are enabled on MI300X, its own 4 MB L2, and its own hardware scheduler. The 256 MB Infinity Cache sits on the I/O dies, in front of memory, shared. </p><p>So the cache hierarchy is not a tree with a single root: it is eight private L2s and one shared last level, and two workgroups that share data are cheap if they land on the <strong>same XCD</strong> and expensive if they do not. There is no hardware mechanism that makes this decision for you. </p><p>The mechanism is that workgroups are handed to XCDs round robin, and therefore if you want<strong> two tiles co-resident</strong> you remap your program id arithmetic so that they are congruent modulo eight. </p><p>AMD&#8217;s guidance to use workgroup mapping values that are multiples of the XCD count is exactly this, and you can see it in the shipped kernel names: <code>WGMXCC8</code> appears in the <strong>GEMM library&#8217;s solution names</strong> because the number eight is baked into the tuning.</p><p>Two more things about this die that would be folklore elsewhere. Clock speed varies between XCDs on the same package by three to ten percent, <strong>XCD0 typically fastest and XCD7 slowest on MI300X</strong>, so an efficiency number computed against nominal clock is systematically optimistic and AMD tells you to compute against the slowest XCD. </p><p>And a GEMM whose leading dimension is a multiple of 512 bytes hits channel hotspotting, with the recommended fix being to pad: <code>lda = ldb = K + 128</code> when<strong> K is a multiple of 256.</strong> Every memory system has stride pathologies. It is unusual to be told about them.</p><p>Partitioning is the other side of the chiplet story. An MI300X can be presented to software as one device with<strong> eight XCDs and 192 GB</strong>, or two, or four, or eight devices with one XCD and 24 GB each, in modes named SPX, DPX, QPX and CPX, crossed with memory interleaving modes NPS1 through NPS4. </p><p>AMD recommends QPX with NPS4 on MI300X and DPX with NPS2 on MI350X. <em>For a serving fleet this is a real knob:</em> a model that fits in 24 GB gets eight independent devices per module with no cross-device traffic at all, which is a different machine from the one on the spec sheet.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>The numerics fork, which cost more than it should have</h2><p>Now the part of the story that is a genuine unforced error, and the best available evidence that a numerics contract is a moat in its own right.</p><p>FP8 has two dialects. The OCP standard E4M3 has an exponent bias of 7, supports both signed zeros, and has NaN encodings. The variant AMD implemented on CDNA 3, called <strong>E4M3FNUZ</strong>, has an exponent bias of 8, a single zero, and one NaN. The bit layout is the same. </p><p>The value the bits mean is not: read an FNUZ byte as if it were OCP and you are off by a factor of two.</p><p>CDNA 4 switched to the <strong>OCP variant</strong>. So the <em>MI300X and MI325X speak one FP8 and the MI350X and MI355X speak another</em>, and a checkpoint quantised on NVIDIA hardware speaks the second. </p><p>In practice this means every FP8 path in every framework has to be conditioned on the <strong>architecture at runtime</strong>, and the condition is a function call that asks the driver what chip it is talking to.</p><p>It went about as well as you would expect. In June 2026 a vLLM issue documented that the sparse attention wrappers for a MiniMax model classified only <code>float8_e4m3fn</code> and <code>float8_e5m2</code> as FP8, omitting <code>float8_e4m3fnuz</code>, so <strong>on gfx942 the KV cache bytes were reinterpreted in the wrong dialect</strong> before the attention kernels consumed them, with an accuracy loss on a 1,319 sample GSM8K run that the fix recovered. </p><p>In the same period an AITER issue reported that building for two architectures at once, <code>GPU_ARCHS=gfx942;gfx950</code>, made the dtype helper return the <strong>OCP type on an MI300X</strong>, because it resolved the architecture from the build configuration rather than the device. </p><p><em>Fergus Finn</em>, bringing up DeepSeek V4-Flash on a single MI300X, wrote that <strong>many of vLLM&#8217;s FP8 paths know E4M3 from E5M2</strong> but not FNUZ from OCP, and observed that MI300X is the only major accelerator where the distinction matters in practice.</p><p>That last &#8220;clause&#8221; is about the whole cost. A format that only one vendor&#8217;s one generation uses gets <strong>tested by exactly the people who have that hardware</strong>, which is a small fraction of the people who write the code. </p><p>Being different is expensive in proportion to how few of you there are, and it is expensive in the layer where it is hardest to notice, because the failure mode is not a crash, it is a slightly worse answer.</p><p>Two more numerics changes in CDNA 4 deserve a line. TF32 moved from hardware to software emulation via BF16, which sounds like a regression and is not, because <strong>BF16 matrix throughput on CDNA 4 is 4,096 FLOPs per clock per CU</strong> against CDNA 3&#8217;s TF32 rate of 1,024, so the emulated path is faster than the hardware path it replaced. </p><p>And FP64 matrix throughput halved, from 256 to 128 FLOPs per clock per CU, which is a deliberate reallocation of area away from HPC toward AI and is the reason AMD now ships a separate HPC part.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Counting the kernel layer</h2><p>Here is the measurement I most wanted to make, and the one that is only possible because the stack is open.</p><p><strong>hipBLASLt&#8217;s GEMM kernels are generated by TensileLite</strong>, an assembly generator, and the decision about which generated kernel to use for a given problem shape lives in YAML files in the repository. I cloned the logic directory and counted it.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!t5Pr!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa6f86b77-98da-49aa-9986-0ed88af06cb2_1600x885.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!t5Pr!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa6f86b77-98da-49aa-9986-0ed88af06cb2_1600x885.png 424w, https://substackcdn.com/image/fetch/$s_!t5Pr!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa6f86b77-98da-49aa-9986-0ed88af06cb2_1600x885.png 848w, https://substackcdn.com/image/fetch/$s_!t5Pr!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa6f86b77-98da-49aa-9986-0ed88af06cb2_1600x885.png 1272w, https://substackcdn.com/image/fetch/$s_!t5Pr!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa6f86b77-98da-49aa-9986-0ed88af06cb2_1600x885.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!t5Pr!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa6f86b77-98da-49aa-9986-0ed88af06cb2_1600x885.png" width="1456" height="805" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a6f86b77-98da-49aa-9986-0ed88af06cb2_1600x885.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:805,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;The shipped GEMM tuning surface in hipBLASLt, counted from the repository. Two point nine million measured decisions, in a public git repository. The equivalent number for cuBLAS is not merely secret, it is uncountable.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="The shipped GEMM tuning surface in hipBLASLt, counted from the repository. Two point nine million measured decisions, in a public git repository. The equivalent number for cuBLAS is not merely secret, it is uncountable." title="The shipped GEMM tuning surface in hipBLASLt, counted from the repository. Two point nine million measured decisions, in a public git repository. The equivalent number for cuBLAS is not merely secret, it is uncountable." srcset="https://substackcdn.com/image/fetch/$s_!t5Pr!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa6f86b77-98da-49aa-9986-0ed88af06cb2_1600x885.png 424w, https://substackcdn.com/image/fetch/$s_!t5Pr!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa6f86b77-98da-49aa-9986-0ed88af06cb2_1600x885.png 848w, https://substackcdn.com/image/fetch/$s_!t5Pr!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa6f86b77-98da-49aa-9986-0ed88af06cb2_1600x885.png 1272w, https://substackcdn.com/image/fetch/$s_!t5Pr!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa6f86b77-98da-49aa-9986-0ed88af06cb2_1600x885.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong><span>Figure 5.</span></strong> Two point nine million measured decisions, in a public git repository. The equivalent number for cuBLAS is not merely secret, it is uncountable. </figcaption></figure></div><p>145 logic files. 26,248 solutions. 16,177 distinct assembly kernel names. 2,905,048 entries mapping a problem shape to a solution index. 310 megabytes of YAML, of which 297 MB is the tuning tables.</p><p><em>Two details in that census are worth more than the headline</em>. The first is that the <strong>MI200 tuning is split into directories</strong> named <code>104CU</code> and <code>110CU</code>. The same architecture, tuned separately by how many compute units were left enabled after harvest. </p><p>Which bin of the die you bought changes which kernel is fastest for your matrix, and the library ships both tables. The second is that the older architecture has vastly more tuned shapes than the newer one: MI200 carries <strong>2.9 million exact entries across two CU counts</strong>, while the gfx942 tree in this snapshot carries 1,746, because MI300&#8217;s dispatch leans much harder on heuristics and grid-based interpolation than on exhaustive tables.</p><p>Then I did something with the kernel names, because Tensile&#8217;s names are self-documenting. A real one, unedited:</p><pre><code><code>Cijk_Ailk_Bjlk_BBS_BH_Bias_HAS_SAV_UserArgs_MT256x224x32_MI16x16x1_SN
_GRVWA8_GRVWB4_GSU2_LBSPPA2048_LBSPPB1792_LPA0_LPB32_MIWT4_14_NTC0
_NTD0_NLCA1_NLCB7_SU8_SUM0_SUS256_SVW4_VWA4_VWB2_WSGRA0_WSGRB2
_WG64_4_1_WGM304_WGMXCC8_WGMXCCG0</code></code></pre><p><code>MT256x224x32</code> is the macro tile. <code>MI16x16x1</code> is the matrix instruction. <code>GRVWA8</code> is the global read vector width for A. <code>GSU2</code> is global split-U. <code>WGM304</code> is a workgroup mapping value that happens to be the compute unit count of the part. </p><p>Parse all 1,439 distinct gfx942 names and you recover the search space directly:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Yz02!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F203c3cc2-22e2-47ec-8429-65c4a4925890_1600x2245.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Yz02!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F203c3cc2-22e2-47ec-8429-65c4a4925890_1600x2245.png 424w, https://substackcdn.com/image/fetch/$s_!Yz02!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F203c3cc2-22e2-47ec-8429-65c4a4925890_1600x2245.png 848w, https://substackcdn.com/image/fetch/$s_!Yz02!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F203c3cc2-22e2-47ec-8429-65c4a4925890_1600x2245.png 1272w, https://substackcdn.com/image/fetch/$s_!Yz02!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F203c3cc2-22e2-47ec-8429-65c4a4925890_1600x2245.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Yz02!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F203c3cc2-22e2-47ec-8429-65c4a4925890_1600x2245.png" width="1456" height="2043" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/203c3cc2-22e2-47ec-8429-65c4a4925890_1600x2245.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:2043,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;What the GEMM tuner searches, one row per parameter. Recovered from the names of the 1,439 gfx942 kernels hipBLASLt ships, which encode their own parameters. Fifty six knobs, 120 bits of space, and a shipped set that covers 2 to the &quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="What the GEMM tuner searches, one row per parameter. Recovered from the names of the 1,439 gfx942 kernels hipBLASLt ships, which encode their own parameters. Fifty six knobs, 120 bits of space, and a shipped set that covers 2 to the " title="What the GEMM tuner searches, one row per parameter. Recovered from the names of the 1,439 gfx942 kernels hipBLASLt ships, which encode their own parameters. Fifty six knobs, 120 bits of space, and a shipped set that covers 2 to the " srcset="https://substackcdn.com/image/fetch/$s_!Yz02!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F203c3cc2-22e2-47ec-8429-65c4a4925890_1600x2245.png 424w, https://substackcdn.com/image/fetch/$s_!Yz02!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F203c3cc2-22e2-47ec-8429-65c4a4925890_1600x2245.png 848w, https://substackcdn.com/image/fetch/$s_!Yz02!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F203c3cc2-22e2-47ec-8429-65c4a4925890_1600x2245.png 1272w, https://substackcdn.com/image/fetch/$s_!Yz02!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F203c3cc2-22e2-47ec-8429-65c4a4925890_1600x2245.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong><span>Figure 6.</span></strong> Recovered from the names of the 1,439 gfx942 kernels hipBLASLt ships, which encode their own parameters. Fifty six knobs, 120 bits of space, and a shipped set that covers 2 to the ten and a half of it. </figcaption></figure></div><p>Fifty six parameters take more than one value across the shipped set. Because cardinalities multiply, the<strong> natural unit is bits:</strong> a parameter with sixteen observed values contributes four. They sum to 120 bits, which is 1.3 times ten to the thirty sixth. The 1,439 kernels actually shipped are ten and a half bits of that.</p><p>Where the bits sit is as interesting as how many there are. <strong>Scheduling and non-temporal hints carry 39 of them across 24 parameters</strong>, more than tile geometry and K splitting combined. </p><p>The macro tile alone is eight bits, and the three XCD mapping parameters are twelve, which is a lot of search space spent on the fact that the chip is eight dies.</p><p>That is what the kernel layer is. Not a secret, not a compiler trick, not a set of instructions nobody knows about. It is a search over a space with about <strong>thirty six log-decades of volume</strong>, resolved by measurement on real silicon, and then frozen as a lookup table. The library is the record of the search.</p><p>Which is why I have been arguing for a while that the correct model of this moat is a labour market rather than a technology, and this measurement is the strongest version of that argument I have been able to construct. Anybody can read the ISA. Anybody can write the assembly. </p><p>What is truly expensive is running the<strong> 2.9 million benchmarks on hardware you have to own</strong>, and doing it again for every new part, and again for every harvest bin of every new part.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Four compilers, one backend</h2><p>The compiler story is simpler than people expect, because there is one backend and it is upstream.</p><p>Every path into a <strong>CDNA GPU</strong> ends at the LLVM AMDGPU target. HIP is clang. OpenMP offload is clang. Triton&#8217;s AMD backend emits LLVM IR into the same code generator. Composable Kernel is C++ templates compiled by clang. Even the assembly generators emit text the same assembler consumes. </p><p>There is no split where a public virtual ISA is compiled by a private program into the real one: the <code>.s</code> file clang produces is the instruction stream, and <code>llvm-mc</code> assembles it back.</p><p>So the surface area of AMD&#8217;s compiler is measurable, and it is growing fast. The <strong>AMDGPU builtin list went from 443 entries in LLVM 20.1.2 to 788 in 22.1.0</strong>, up 78 percent with nothing removed. Available per target: 244 on gfx90a, 274 on gfx942, 362 on gfx950, 477 on gfx1250.</p><p>Triton is the most important of the four paths, because it is the one the frameworks generate into. The <strong>AMD backend is real</strong> and it is used in production: TorchInductor generates Triton, vLLM ships Triton attention kernels for ROCm, and AMD&#8217;s own guidance for tuning them is specific in a way that tells you what the compiler is not doing for you. </p><p>Set <code>num_stages</code> to 2 for a single GEMM and 1 for two fused GEMMs, because the pipeliner&#8217;s cost model does not know the difference. </p><p>Use <code>waves_per_eu</code> to push the register allocator down to the next occupancy step, because the allocator optimises for spills rather than occupancy. Use <code>matrix_instr_nonkdim</code> to pick the matrix instruction, because the heuristic picks the larger one and the smaller one is usually faster.</p><p>Each knob is a place where the compiler has a policy and the policy is wrong often enough to warrant a flag. That is the state of the art everywhere, <strong>not an AMD failing,</strong> but the flags read as a map of the gap.</p><p>Gluon is the interesting recent addition. It ships alongside Triton, <strong>compiles through the same IR</strong>, and exposes what Triton hides: explicit layouts, explicit asynchronous copies and barriers, explicit LDS placement. </p><p>AMD&#8217;s documentation is unusually frank about when to reach for it, namely when the profiler shows you bottlenecked on layout conversions, matrix instruction selection or <strong>pipelining depth</strong>, which is to say on the three decisions Triton makes for you. The 697 to 1113 TFLOPS figures come from that tutorial. </p><p>A sixty percent gap between the natural expression and the tuned one, in a kernel the vendor wrote to demonstrate tuning, is an honest measure of how much the language is doing.</p><p><strong>Composable Kernel is the older answer:</strong> a C++ template library where a kernel is assembled from tile descriptors and instances are enumerated at build time. </p><p>TorchInductor can use it for GEMM if you add CK to the autotune backends, and it is one of AITER&#8217;s. Its weakness is the one every heavy template library has, and the fact that <strong>AITER ships Opus</strong>, a single-header alternative advertised as up to 61 times faster to build, is a fairly direct comment on it.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>What actually runs when you serve a model</h2><p><strong>AITER </strong>is the piece that changed the trajectory, and &#8220;<em>AMD&#8217;s cuDNN</em>&#8221; undersells it.</p><p>AITER is a dispatcher with five backends: Composable Kernel, hand-written assembly, Triton, a <strong>DSL called FlyDSL</strong>, and hipBLAS. For a given operator and shape it picks one, driven by CSV tables of tuned configurations merged at runtime with model-specific tables shipped alongside. </p><p>In vLLM it is one environment variable, <code>VLLM_ROCM_USE_AITER=1</code>, and it replaces GEMM, <strong>RMSNorm</strong>, mixture-of-experts and attention kernels underneath the engine without the model code changing.</p><p>Attention shows the design working. For multi-head latent attention, vLLM on ROCm offers a Triton backend and two AITER backends that differ only in the prefill path: both use the same <strong>hand-written assembly decode kernel</strong>, <code>mla_decode_fwd</code>, and vLLM&#8217;s own write-up attributes most of the 1.2 to 1.6 times speedup to that one kernel, because decode is memory bound and time per output token is decode-heavy.</p><p>What that bought is visible from outside AMD. SemiAnalysis&#8217;s InferenceX v2, published in February 2026, <strong>measured MI300X SGLang throughput roughly doubling between December 2025 and January 2026</strong>. Same silicon, same model, one month, two times the tokens. That measurement settles the question of whether the hardware was the constraint.</p><p>It also says something uncomfortable about every AMD benchmark published before it. A number measured on ROCm in November 2025 was not measuring the machine; it was <strong>measuring the kernel library at a moment in a period of rapid change</strong>. Hold that when reading anybody&#8217;s AMD versus NVIDIA comparison, including the ones I cite approvingly.</p><p>The version story compounds this. ROCm currently ships in two parallel streams, with <em>7.0 through 7.8 reserved for production and 7.9</em> and later designated as a technology preview with a different build system, so the highest version number is not the production one. </p><p>As of this writing production is 7.2.x and preview is 7.14. Meanwhile <strong>hipBLASLt&#8217;s default branch on GitHub </strong>is named <code>develop_deprecated</code> and its head commit is from June 2025. None of these are disasters, but collectively they are the texture of a stack that is being rebuilt while in flight, and they are a real cost to anybody trying to pin a reproducible configuration.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>The arithmetic, which is where AMD&#8217;s case is strongest</h2><p>Strip away the software argument for a moment and ask what the hardware is for.</p><p>The batch at which a dense GEMM stops being memory bound and starts <strong>being compute bound is a property of the part</strong>, not of the model. Call it B*, and it is peak throughput times bytes per element divided by twice the memory bandwidth. </p><p>The reason it is worth computing is that it is invariant to precision: peak throughput scales as the inverse of element width, so <em>the product is a constant for a given generation</em>, and quantising a model moves the ceiling without moving the corner.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!PaY-!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2486271d-d769-42f4-b56a-8de47f7c784c_1600x767.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!PaY-!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2486271d-d769-42f4-b56a-8de47f7c784c_1600x767.png 424w, https://substackcdn.com/image/fetch/$s_!PaY-!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2486271d-d769-42f4-b56a-8de47f7c784c_1600x767.png 848w, https://substackcdn.com/image/fetch/$s_!PaY-!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2486271d-d769-42f4-b56a-8de47f7c784c_1600x767.png 1272w, https://substackcdn.com/image/fetch/$s_!PaY-!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2486271d-d769-42f4-b56a-8de47f7c784c_1600x767.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!PaY-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2486271d-d769-42f4-b56a-8de47f7c784c_1600x767.png" width="1456" height="698" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2486271d-d769-42f4-b56a-8de47f7c784c_1600x767.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:698,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Critical batch B* = P.b / (2.BW), the batch where a dense GEMM stops being memory bound. B* is invariant to precision because peak throughput scales as the inverse of element width. Quantising moves the ceiling, not the corner.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Critical batch B* = P.b / (2.BW), the batch where a dense GEMM stops being memory bound. B* is invariant to precision because peak throughput scales as the inverse of element width. Quantising moves the ceiling, not the corner." title="Critical batch B* = P.b / (2.BW), the batch where a dense GEMM stops being memory bound. B* is invariant to precision because peak throughput scales as the inverse of element width. Quantising moves the ceiling, not the corner." srcset="https://substackcdn.com/image/fetch/$s_!PaY-!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2486271d-d769-42f4-b56a-8de47f7c784c_1600x767.png 424w, https://substackcdn.com/image/fetch/$s_!PaY-!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2486271d-d769-42f4-b56a-8de47f7c784c_1600x767.png 848w, https://substackcdn.com/image/fetch/$s_!PaY-!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2486271d-d769-42f4-b56a-8de47f7c784c_1600x767.png 1272w, https://substackcdn.com/image/fetch/$s_!PaY-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2486271d-d769-42f4-b56a-8de47f7c784c_1600x767.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong><span>Figure 7.</span></strong> B* is invariant to precision because peak throughput scales as the inverse of element width. Quantising moves the ceiling, not the corner. </figcaption></figure></div><p>MI300X sits at 247. H100 sits at 295. MI350X at 288, MI355X at 313, B200 at 281. Every one is identical across FP16, FP8 and FP4 to within one percent, a <strong>third independent confirmation</strong> of the invariance on silicon I had not previously tested it on.</p><p>So MI300X reaches the compute-bound regime at a batch about sixteen percent smaller than H100 does, a real advantage for interactive serving, and by the current generation <strong>the two vendors have converged to within ten percent.</strong> Whatever quantisation is buying, it is not an escape from the memory wall.</p><p>Then run the same arithmetic on what is shipping now. AMD&#8217;s product page gives<strong> MI455X 20 PFLOPS of FP8, 40 of FP4 and up to 23.3 TB/s of HBM4</strong>. That puts B* at 429, up thirty seven percent in one generation. NVIDIA&#8217;s published Rubin figures move in the same direction. </p><p>The corner that barely moved for two generations is now moving quickly, and it is moving the wrong way for the argument AMD has been making, because the <strong>memory-bound regime</strong> where a less mature kernel layer costs you little is the regime that is shrinking.</p><p>I would not over-read a single generation. But if I had to name the thing most likely to <strong>invalidate the AMD inference case</strong> over the next two years, it would not be software. It would be this number.</p><p>The place the arithmetic is not close is capacity.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!dqN2!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F615b9cb6-f4f5-482b-8352-9ac72943485a_1600x577.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!dqN2!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F615b9cb6-f4f5-482b-8352-9ac72943485a_1600x577.png 424w, https://substackcdn.com/image/fetch/$s_!dqN2!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F615b9cb6-f4f5-482b-8352-9ac72943485a_1600x577.png 848w, https://substackcdn.com/image/fetch/$s_!dqN2!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F615b9cb6-f4f5-482b-8352-9ac72943485a_1600x577.png 1272w, https://substackcdn.com/image/fetch/$s_!dqN2!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F615b9cb6-f4f5-482b-8352-9ac72943485a_1600x577.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!dqN2!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F615b9cb6-f4f5-482b-8352-9ac72943485a_1600x577.png" width="1456" height="525" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/615b9cb6-f4f5-482b-8352-9ac72943485a_1600x577.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:525,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Tokens of FP8 KV cache resident beside a 70B FP8 model, GQA-8. Capacity is the part of the AMD case that needs no software at all.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Tokens of FP8 KV cache resident beside a 70B FP8 model, GQA-8. Capacity is the part of the AMD case that needs no software at all." title="Tokens of FP8 KV cache resident beside a 70B FP8 model, GQA-8. Capacity is the part of the AMD case that needs no software at all." srcset="https://substackcdn.com/image/fetch/$s_!dqN2!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F615b9cb6-f4f5-482b-8352-9ac72943485a_1600x577.png 424w, https://substackcdn.com/image/fetch/$s_!dqN2!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F615b9cb6-f4f5-482b-8352-9ac72943485a_1600x577.png 848w, https://substackcdn.com/image/fetch/$s_!dqN2!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F615b9cb6-f4f5-482b-8352-9ac72943485a_1600x577.png 1272w, https://substackcdn.com/image/fetch/$s_!dqN2!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F615b9cb6-f4f5-482b-8352-9ac72943485a_1600x577.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong><span>Figure 8.</span></strong> Capacity is the part of the AMD case that needs no software at all.</figcaption></figure></div><p>Take a 70 billion parameter dense model with grouped query attention at eight key-value heads, weights and KV cache both at one byte per element, and give the runtime ninety percent of nameplate memory. </p><ul><li><p>H100 has two gigabytes left for KV after the weights, which is about twelve thousand tokens, which is one and a half concurrent requests at eight thousand tokens of context. </p></li><li><p>MI300X has 103 gigabytes left, which is 627,000 tokens, which is 76 concurrent requests. </p></li><li><p>Whereas, an MI355X has 189 gigabytes left, 1.15 million tokens, 141 requests.</p></li></ul><p>We see there&#8217;s a <strong>fifty times difference</strong> between an H100 and an MI300X in the quantity that determines how many users one accelerator can serve at once, and it required no software from anybody. </p><p>H200 closes most of it and B200 closes the rest, which is why the AMD capacity advantage was a 2024 and 2025 story more than a 2026 one, but it is also why AMD got the foothold it got: for a period, the only way to serve certain models without sharding was on AMD.</p><p>The current generation&#8217;s claim is narrower and more credible for being narrower. <strong>AMD says a Helios rack delivers up to 30 percent more tokens per dollar</strong> than the leading competitive solution, based on its own labs, with the MI455X carrying 432 GB of HBM4 against Rubin&#8217;s 288 and the rack pooling 31 TB. </p><p>Thats a capacity argument again, dressed as an economics argument, and it will be right or wrong depending on whether the kernels exist.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>What a port actually costs</h2><p>The official story about moving code from CUDA to ROCm is HIP plus hipify: run a translator, get a source tree that compiles for both, done. A figure of <strong>roughly 92 percent coverage of CUDA device APIs </strong>circulates in reseller and channel material around the MI355X launch. </p><p>I couldn&#8217;t trace it to a primary AMD statement, so treat it as folklore with a plausible magnitude rather than as a specification.</p><p>Whatever the true figure, it is the least interesting number in the discussion, because the remainder is not randomly distributed. It concentrates in exactly the places where performance lives.</p><p>Consider what does not port. <strong>Inline PTX, because there is no PTX.</strong> Warp level primitives with hardcoded masks, because the mask is 64 bits wide. </p><p>Anything using the tensor memory accelerator, <code>wgmma</code> or <code>tcgen05</code>, because those mechanisms do not exist and the matrix instructions have different shapes and register layouts. <strong>CUTLASS</strong>, where a great deal of the industry&#8217;s kernel expertise is encoded, is structurally NVIDIA-specific from the CuTe layout algebra down. </p><p>Anything assuming <strong>32 lanes per warp </strong>for a reduction is subtly wrong rather than broken, which is worse. And numerics do not port, as the FNUZ story demonstrated at length.</p><p>The honest description of a port, then: the model runs almost immediately and <em>the kernels that make it fast do not exist</em>. That matches every bring-up worklog I have read. </p><p>A model comes up in a day, then weeks go into finding which operator fell back to a slow path, which is <strong>what you would predict from a stack whose framework layer is portable</strong> and whose kernel layer is a per-architecture lookup table.</p><p>One naming artefact tells the whole story. PyTorch on ROCm still calls the device <code>cuda</code>: <code>torch.cuda.is_available()</code> returns true, <code>tensor.cuda()</code> works. </p><p>The build hipifies PyTorch&#8217;s sources and keeps the Python-facing names, because changing them would break every model script in existence. The <strong>API surface is compatible</strong> and the thing underneath is not the same machine.</p><div class="directMessage button" data-attrs="{&quot;userId&quot;:201922174,&quot;userName&quot;:&quot;Lorenzo Bradanini&quot;,&quot;canDm&quot;:null,&quot;dmUpgradeOptions&quot;:null,&quot;isEditorNode&quot;:true}" data-component-name="DirectMessageToDOM"></div><div><hr></div><h2>Profiling, and the churn</h2><p>The profiling story is the part of ROCm that has changed most in the last two years and it is worth knowing where it landed.</p><p>The current tools are rocprofv3 for counters and traces, <strong>ROCm Compute Profiler</strong> (<em>formerly Omniperf</em>) for guided kernel analysis, and <strong>ROCm Systems Profiler </strong>(<em>previously known as Omnitrace</em>) for whole-application timelines. </p><p>The generation before, <strong>rocprof, rocprofv2, ROCProfiler </strong>and <strong>ROCTracer</strong>, is deprecated with end of support announced for the second quarter of 2026, and as of the 7.2.1 notes PyTorch on ROCm still depended on ROCTracer, with a known issue tracking the migration. </p><p>That is a fair sample of the <strong>ROCm experience:</strong> the new tools are good, the migration is real, the documentation is honest about what has not moved, and something you depend on is probably still on the old path.</p><p>The counter model is conventional. Per-shader-engine and per-cache counters, collected in multiple passes with the application re-run each time, so anything nondeterministic is measured across different executions. </p><p>For a matrix kernel the ones that matter are <strong>MFMA issue counts</strong>, L2 and Infinity Cache hit rates, and memory controller requests, and the derived metrics sit close enough to Nsight Compute&#8217;s sections that the mental model transfers.</p><p>One capability has no clean counterpart: advanced thread trace, which captures per-instruction issue timing inside a wave. Combined with being able to read and rewrite the instruction stream using standard tools, it <strong>closes the loop of look at the schedule</strong>, change the schedule, measure the schedule, which is hard to close on the other side. </p><p>Whether anyone outside AMD runs that loop at scale is a separate question, and I suspect the answer is a few dozen people.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>What the numbers say about the machine, not the marketing</h2><p>One last piece of arithmetic before the fork, because it reframes the generational comparison in a way the press releases do not.</p><p>Peak throughput figures conflate three things: how many units there are, how fast they run, and <strong>how much each unit does per clock. </strong>Divide it out and you get the architectural quantity.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Ax7Z!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd4859b8d-2e79-49f2-97ec-f44d70470695_1600x885.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Ax7Z!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd4859b8d-2e79-49f2-97ec-f44d70470695_1600x885.png 424w, https://substackcdn.com/image/fetch/$s_!Ax7Z!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd4859b8d-2e79-49f2-97ec-f44d70470695_1600x885.png 848w, https://substackcdn.com/image/fetch/$s_!Ax7Z!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd4859b8d-2e79-49f2-97ec-f44d70470695_1600x885.png 1272w, https://substackcdn.com/image/fetch/$s_!Ax7Z!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd4859b8d-2e79-49f2-97ec-f44d70470695_1600x885.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Ax7Z!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd4859b8d-2e79-49f2-97ec-f44d70470695_1600x885.png" width="1456" height="805" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d4859b8d-2e79-49f2-97ec-f44d70470695_1600x885.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:805,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Matrix throughput per clock per compute unit, which removes clock and unit count. The architectural quantity behind the PFLOPS headline. CDNA 4 caught Hopper per unit per clock; Blackwell had already doubled again.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Matrix throughput per clock per compute unit, which removes clock and unit count. The architectural quantity behind the PFLOPS headline. CDNA 4 caught Hopper per unit per clock; Blackwell had already doubled again." title="Matrix throughput per clock per compute unit, which removes clock and unit count. The architectural quantity behind the PFLOPS headline. CDNA 4 caught Hopper per unit per clock; Blackwell had already doubled again." srcset="https://substackcdn.com/image/fetch/$s_!Ax7Z!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd4859b8d-2e79-49f2-97ec-f44d70470695_1600x885.png 424w, https://substackcdn.com/image/fetch/$s_!Ax7Z!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd4859b8d-2e79-49f2-97ec-f44d70470695_1600x885.png 848w, https://substackcdn.com/image/fetch/$s_!Ax7Z!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd4859b8d-2e79-49f2-97ec-f44d70470695_1600x885.png 1272w, https://substackcdn.com/image/fetch/$s_!Ax7Z!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd4859b8d-2e79-49f2-97ec-f44d70470695_1600x885.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong><span>Figure 9.</span></strong> The architectural quantity behind the PFLOPS headline. CDNA 4 caught Hopper per unit per clock; Blackwell had already doubled again.</figcaption></figure></div><p>An MI300X compute unit does 2,048 FP16 and 4,096 FP8 FLOPs per clock. An H100 SM does about 4,271 and 8,543. So per unit per clock, a CDNA 3 compute unit is half an H100 SM, and <strong>AMD reaches parity on the total by fielding 304 units</strong> against 132 and by accepting more area and more power. </p><p>CDNA 4 doubles the per-unit rate to about 8,138 FP8, which lands on top of Hopper, and adds FP4 at 16,439. Blackwell doubles again to 15,473 FP8 and 30,947 FP4.</p><p>That is the real generational story and it is more interesting than the PFLOPS. <strong>AMD spent CDNA 4 catching the previous NVIDIA generation </strong>on matrix density per unit while cutting unit count from 304 to 256, and paid for it with process, going from N5 to N3P. NVIDIA spent the same interval doubling again. </p><p>On <strong>matrix density per compute unit</strong> per clock, the gap between the two current parts is close to a factor of two, and the gap on the whole part is much smaller because of unit counts and clocks.</p><p>Which is why the AMD case has always been strongest where matrix density is not the binding constraint, which is decode. Decode is memory bound below B<em>, </em>which is <strong>313 on MI355X and 281 on B200</strong>, and the memory bandwidth is the same 8 TB/s on both. </p><p>In that regime a compute unit that does half the matrix work per clock is not costing you anything, and 288 GB against 192 is costing the other side quite a lot.</p><p>The<strong> AMD inference argument</strong>, stripped of everything else, is that a large fraction of served tokens are generated in a regime where the thing AMD is worse at does not bind and the thing AMD is better at does. </p><p>That argument was correct in 2024, it survived the H200, and the MI455X version of it, 432 GB against 288, is the same argument again. </p><p>Whether it holds depends on kernels that, as of this writing, are being written for an instruction set that did not exist eighteen months ago.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/p/distilling-in-depth-rocm-how-it-actually?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/p/distilling-in-depth-rocm-how-it-actually?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><div><hr></div><h2>Why the case is inference and not training</h2><p>There is a <strong>persistent asymmetry in AMD&#8217;s results</strong> that I have seen stated as a puzzle and that I think has a clean explanation. AMD is competitive on inference, sometimes better than competitive, but is not even close on training.</p><p>The usual explanation is that <strong>training software is harder</strong>, which is true and unhelpful. The specific version is more useful, and it falls out of the structure described so far.</p><p>Start with what a kernel layer costs to build. A serving engine&#8217;s hot path is a small set of shapes: attention in its prefill, extend and decode forms, the projections, the<strong> mixture-of-experts grouped GEMM</strong>, the normalisations, the sampling tail. Perhaps twenty operators, each with a handful of shape families fixed once you pick the model. </p><p>That is a finite, enumerable target, and it is why AITER can exist as a dispatcher with per-model CSV tables. You can <strong>hand-write assembly for a decode kernel</strong> because there is one decode kernel and it stays hot for the life of the model.</p><p>Training is not that. Backward passes roughly triple the operator count and introduce transposed shapes that hit different tuning entries. Optimiser state is bandwidth-bound elementwise work over parameter-sized tensors. </p><p>Gradient collectives at every step put the interconnect on the critical path rather than at the margin. It runs for weeks, so numerical drift invisible in a benchmark becomes a divergence at step 40,000, and it <strong>runs across thousands of accelerators</strong> where the binding constraint is not any kernel but the probability that all of them stay up. </p><p>None of that is helped by a deep tuning table, because there is no hot shape you can pay somebody to tune once.</p><p>Then the collective asymmetry. On a fully connected eight-way node, all-reduce at full width is efficient, which is what AMD&#8217;s use-eight-or-one guidance reflects. Above the node you are on the scale-out fabric, and until Helios there was <strong>no rack-scale scale-up domain at all.</strong> A tensor-parallel group that fits in a node is fine. </p><p>One that does not, or an expert-parallel all-to-all spanning racks, is the problem NVIDIA spent NVL72 on two years earlier.</p><p>Now apply the same reasoning to inference and the picture inverts. Decode is memory bound below B*, which we<strong> computed at 247 on MI300X and 313 on MI355X.</strong> Below that batch, matrix throughput per compute unit is not the constraint, which neutralises AMD&#8217;s largest architectural deficit. </p><p>Capacity is the constraint, and AMD has more of it, by 50 percent against B200 on the current part and by 50 percent again on the next one. <strong>Prefill is compute bound and AMD is worse there</strong>, but prefill is amortised across the output tokens of a request, and for the long-output workloads that dominate agentic and reasoning traffic, that amortisation is generous.</p><p>So the asymmetry is not primarily a software story. It is that inference has a small hot kernel set, tolerates a per-model tuning table, lives in a regime where <strong>AMD&#8217;s weakness does not bind,</strong> and puts a premium on the quantity AMD sells the most of.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>What the money looks like</h2><p>Let me be careful here, because I am about to do arithmetic on numbers I did not measure, and the conclusions are only as good as the inputs.</p><p>The best third-party source is <strong>SemiAnalysis&#8217;s InferenceX</strong>, which publishes a continuously updated benchmark rather than a single number, and the trajectory of that benchmark is more instructive than any point on it.</p><blockquote><p><em>On 20 May 2026 InferenceX measured MI355X on SGLang FP8 as up to about 40 percent cheaper per million tokens than B200 on GLM-5 at 8K input and 1K output, with the peak gap at 18 tokens per second per user, 22 cents per million against 30. B200 took the lead back above roughly 90 tokens per second per user. </em></p></blockquote><p>Two months later the same benchmark&#8217;s overview, running its July TCO model, listed MI355X at 35.5 cents per million on the 8K/1K FP4 path against B200 at 30.4, which is 17 percent the other way. </p><p>On the<strong> long-context multi-turn agentic scenario</strong> the gap in the published table is close to an order of magnitude in NVIDIA&#8217;s favour, and I would want to read that methodology carefully before leaning on the magnitude.</p><p>Nothing about the hardware changed in those two months. What changed is that B200&#8217;s NVFP4 path shipped for that model, and that <strong>AMD still has no disaggregation</strong> or wide expert parallel recipe for it while NVIDIA&#8217;s rack-scale version demonstrated roughly three times the throughput per GPU from wide expert parallelism alone.</p><p>That cuts against the argument I would otherwise have been tempted to make. A <strong>cost-per-token comparison</strong> between these two vendors is not a fact about the machines. It is a snapshot of which recipe landed most recently, and it has flipped sign twice inside one quarter.</p><p><strong>AMD&#8217;s Helios claim</strong>, up to 30 percent more tokens per dollar than the leading competitive solution,<em> is a very different object again</em>: a projection about a rack that began shipping at the end of the third quarter, against a competitor rack, from AMD&#8217;s own labs. Weight it accordingly.</p><p>The mechanism survives the specific numbers going stale, so state it in the abstract. <strong>Cost per token for a memory-bound decode</strong> is the accelerator hour rate divided by tokens per hour, and tokens per hour is roughly bandwidth over bytes touched per token, times the concurrency you can hold. </p><p>AMD&#8217;s argument is that the rate is lower because AMD sells at a discount and the <strong>concurrency is higher</strong> because AMD ships more memory. Neither depends on kernels being good, only on their not being bad enough to cost you the bandwidth.</p><p>That last clause is the whole game, and it is measurable. If a kernel achieves 80 percent of <strong>peak HBM bandwidth on decode</strong>, and the competitor&#8217;s achieves 90, then the entire kernel gap is 11 percent, and it is bounded above by the ratio of achieved bandwidths regardless of how much cleverness is in the other stack. </p><p>This is a much tighter bound than the equivalent for a compute-bound kernel, where the gap between a naive and a tuned implementation can be five times or more. </p><p>It is the analytical reason a memory-bound regime is forgiving of a <strong>less mature kernel layer</strong>, and it is the reason AITER could double throughput in a month and then not double it again.</p><p>The corollary is uncomfortable for the AMD case in a different direction. Once a workload moves into the <em>compute-bound regime</em>, above B*, the forgiving bound disappears and the kernel gap opens back up to whatever the tuning tables say it is. </p><p>Batch sizes are going up, prefill-heavy agentic traffic is going up, and speculative decoding exists specifically to manufacture arithmetic intensity. Every one of those trends moves served traffic toward the regime where AMD&#8217;s kernel maturity matters more, not less.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share&quot;,&quot;text&quot;:&quot;Share The Software Frontier&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share The Software Frontier</span></a></p><div><hr></div><h2>The three numbers I would want if I were buying</h2><p>If somebody put a purchase decision in front of me, I would want three measurements and would not care much about anything else.</p><ol><li><p>The <strong>first </strong>is <strong>achieved HBM bandwidth on decode for the specific model</strong>, as a fraction of nameplate, on both candidate machines, measured on the same day with the same serving engine version. That single ratio bounds the kernel gap in the regime where most tokens are produced, and it is cheap to measure.</p></li><li><p>The <strong>second </strong>is the <strong>fraction of end-to-end wall clock spent in operators that fall back</strong> to an untuned path. On AMD this is directly observable: run with the tuning tables and without, and the difference is the size of the table&#8217;s contribution. If the answer is large, you are buying a dependency on the vendor&#8217;s tuning campaign continuing.</p></li><li><p>The <strong>third </strong>is the <strong>date on every number in the deck, from either vendor.</strong> A third party measured a two times throughput change in a month and a sign flip in cost per token inside a quarter. Any figure older than that is not evidence about the machine you would receive.</p></li></ol><p>And if the machine in question is an MI455X, I would add a fourth, which is what fraction of the operator set has a gfx1250 kernel at all, as opposed to a Triton fallback. That number is knowable, it changes weekly, and as of this writing I have no way to measure it from outside.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>The fork</h2><p>Which brings us to the measurement I opened with.</p><p>I ran the entire clang builtin table against the target feature sets of <strong>five AMD datacenter targets</strong>, which is exactly the computation clang performs when it decides whether to accept a builtin, and then verified a sample of the results by compiling.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!94It!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb31f1445-2f1f-44af-89b5-a047ad959cb9_1600x885.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!94It!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb31f1445-2f1f-44af-89b5-a047ad959cb9_1600x885.png 424w, https://substackcdn.com/image/fetch/$s_!94It!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb31f1445-2f1f-44af-89b5-a047ad959cb9_1600x885.png 848w, https://substackcdn.com/image/fetch/$s_!94It!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb31f1445-2f1f-44af-89b5-a047ad959cb9_1600x885.png 1272w, https://substackcdn.com/image/fetch/$s_!94It!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb31f1445-2f1f-44af-89b5-a047ad959cb9_1600x885.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!94It!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb31f1445-2f1f-44af-89b5-a047ad959cb9_1600x885.png" width="1456" height="805" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b31f1445-2f1f-44af-89b5-a047ad959cb9_1600x885.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:805,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Matrix instructions carried across AMD datacenter generations. Every CDNA transition until now was additive. The move to gfx1250 carries nothing.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Matrix instructions carried across AMD datacenter generations. Every CDNA transition until now was additive. The move to gfx1250 carries nothing." title="Matrix instructions carried across AMD datacenter generations. Every CDNA transition until now was additive. The move to gfx1250 carries nothing." srcset="https://substackcdn.com/image/fetch/$s_!94It!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb31f1445-2f1f-44af-89b5-a047ad959cb9_1600x885.png 424w, https://substackcdn.com/image/fetch/$s_!94It!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb31f1445-2f1f-44af-89b5-a047ad959cb9_1600x885.png 848w, https://substackcdn.com/image/fetch/$s_!94It!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb31f1445-2f1f-44af-89b5-a047ad959cb9_1600x885.png 1272w, https://substackcdn.com/image/fetch/$s_!94It!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb31f1445-2f1f-44af-89b5-a047ad959cb9_1600x885.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong><span>Figure 10.</span></strong> Every CDNA transition until now was additive. The move to gfx1250 carries nothing. </figcaption></figure></div><p>From gfx90a to gfx942, 37 of 37 matrix builtins carry over, and 16 are added. From gfx942 to gfx950, 53 of 53 carry over, and 22 are added. From gfx950 to gfx1250, 0 of 75 carry over, and 73 appear that did not exist before. </p><p>The empirical check is straightforward:</p><pre><code><code>gfx90a   ACCEPTS v_mfma_f32_32x32x8f16
gfx942   ACCEPTS v_mfma_f32_32x32x8f16
gfx950   ACCEPTS v_mfma_f32_32x32x8f16
gfx1250  REJECTS: '__builtin_amdgcn_mfma_f32_32x32x8f16' needs target feature mai-insts
gfx1251  REJECTS: '__builtin_amdgcn_mfma_f32_32x32x8f16' needs target feature mai-insts</code></code></pre><p>The feature diff says the rest.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!4tpm!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9ddcee90-5854-4786-a7e3-cdebc5c7d5c4_1600x1125.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!4tpm!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9ddcee90-5854-4786-a7e3-cdebc5c7d5c4_1600x1125.png 424w, https://substackcdn.com/image/fetch/$s_!4tpm!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9ddcee90-5854-4786-a7e3-cdebc5c7d5c4_1600x1125.png 848w, https://substackcdn.com/image/fetch/$s_!4tpm!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9ddcee90-5854-4786-a7e3-cdebc5c7d5c4_1600x1125.png 1272w, https://substackcdn.com/image/fetch/$s_!4tpm!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9ddcee90-5854-4786-a7e3-cdebc5c7d5c4_1600x1125.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!4tpm!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9ddcee90-5854-4786-a7e3-cdebc5c7d5c4_1600x1125.png" width="1456" height="1024" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/9ddcee90-5854-4786-a7e3-cdebc5c7d5c4_1600x1125.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1024,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Target features lost and gained from gfx950 to gfx1250. mai-insts is the MFMA family. wavefrontsize64 is the wave. Both are on the left.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Target features lost and gained from gfx950 to gfx1250. mai-insts is the MFMA family. wavefrontsize64 is the wave. Both are on the left." title="Target features lost and gained from gfx950 to gfx1250. mai-insts is the MFMA family. wavefrontsize64 is the wave. Both are on the left." srcset="https://substackcdn.com/image/fetch/$s_!4tpm!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9ddcee90-5854-4786-a7e3-cdebc5c7d5c4_1600x1125.png 424w, https://substackcdn.com/image/fetch/$s_!4tpm!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9ddcee90-5854-4786-a7e3-cdebc5c7d5c4_1600x1125.png 848w, https://substackcdn.com/image/fetch/$s_!4tpm!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9ddcee90-5854-4786-a7e3-cdebc5c7d5c4_1600x1125.png 1272w, https://substackcdn.com/image/fetch/$s_!4tpm!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9ddcee90-5854-4786-a7e3-cdebc5c7d5c4_1600x1125.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong><span>Figure 11.</span></strong> mai-insts is the MFMA family. wavefrontsize64 is the wave. Both are on the left.</figcaption></figure></div><p>Twenty four features present on gfx950 are absent on gfx1250. Among them: <code>mai-insts</code>, which is the MFMA family; <code>wavefrontsize64</code>, replaced by <code>wavefrontsize32</code>; the entire <code>dot</code> instruction lineage; every one of CDNA 4&#8217;s block-scale conversion instructions; and <code>s_memtime</code> and <code>s_memrealtime</code>, so even reading a clock changes. </p><p><strong>Twenty two features are new</strong>, and they read like a list of things a CUDA programmer would recognise: <code>clusters</code>, a scheduling group above the workgroup of the kind Hopper introduced; <code>mcast-load-insts</code>, multicast loads; <code>vmem-pref-insts</code>, explicit prefetch; <code>tensor-cvt-lut-insts</code>; <code>transpose-load-f4f6-insts</code>; and <code>fp8e5m3-insts</code>.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/p/distilling-in-depth-rocm-how-it-actually?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/p/distilling-in-depth-rocm-how-it-actually?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><div><hr></div><h2>What replaces it, and what does not change</h2><p>Counting what disappeared is the easy half. The more useful question is what CDNA 5 puts in its place, and the compiler answers that too.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!F4oY!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4a0dccde-f55b-45b7-a3ff-25ca67fe9ecd_1600x885.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!F4oY!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4a0dccde-f55b-45b7-a3ff-25ca67fe9ecd_1600x885.png 424w, https://substackcdn.com/image/fetch/$s_!F4oY!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4a0dccde-f55b-45b7-a3ff-25ca67fe9ecd_1600x885.png 848w, https://substackcdn.com/image/fetch/$s_!F4oY!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4a0dccde-f55b-45b7-a3ff-25ca67fe9ecd_1600x885.png 1272w, https://substackcdn.com/image/fetch/$s_!F4oY!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4a0dccde-f55b-45b7-a3ff-25ca67fe9ecd_1600x885.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!F4oY!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4a0dccde-f55b-45b7-a3ff-25ca67fe9ecd_1600x885.png" width="1456" height="805" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/4a0dccde-f55b-45b7-a3ff-25ca67fe9ecd_1600x885.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:805,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Matrix instruction shapes, grouped by output tile. The 32x32 and 4x4 output tiles, which carried from GCN through all four CDNA generations, do not exist on CDNA 5. What survives is 16x16, at four times the K depth.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Matrix instruction shapes, grouped by output tile. The 32x32 and 4x4 output tiles, which carried from GCN through all four CDNA generations, do not exist on CDNA 5. What survives is 16x16, at four times the K depth." title="Matrix instruction shapes, grouped by output tile. The 32x32 and 4x4 output tiles, which carried from GCN through all four CDNA generations, do not exist on CDNA 5. What survives is 16x16, at four times the K depth." srcset="https://substackcdn.com/image/fetch/$s_!F4oY!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4a0dccde-f55b-45b7-a3ff-25ca67fe9ecd_1600x885.png 424w, https://substackcdn.com/image/fetch/$s_!F4oY!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4a0dccde-f55b-45b7-a3ff-25ca67fe9ecd_1600x885.png 848w, https://substackcdn.com/image/fetch/$s_!F4oY!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4a0dccde-f55b-45b7-a3ff-25ca67fe9ecd_1600x885.png 1272w, https://substackcdn.com/image/fetch/$s_!F4oY!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4a0dccde-f55b-45b7-a3ff-25ca67fe9ecd_1600x885.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong><span>Figure 12.</span></strong> The 32x32 and 4x4 output tiles, which carried from GCN through all four CDNA generations, do not exist on CDNA 5. What survives is 16x16, at four times the K depth. </figcaption></figure></div><p>Group the matrix builtins by the shape of the tile they produce. Every <em>AMD datacenter </em>target from gfx90a to gfx950 offers three output tiles: 4x4, 16x16 and 32x32. On gfx1250 the 4x4 and 32x32 families are gone entirely. </p><p>What survives is 16x16, plus one new rectangular 32x16, and the K dimension deepens: the <strong>largest K on a 16x16 tile goes from 64 on gfx942 to 128 on gfx950</strong> and stays at 128 on gfx1250, now on both surviving tiles.</p><p>That is a narrower and deeper instruction set. Eighteen distinct shapes on gfx950 become six on gfx1250. </p><p>Fewer tile choices means less for a tuner to search over, which is a small mercy given that the tuner has to start from nothing, and a <strong>deeper K means more operand reuse per instruction</strong>, which is the direction every matrix unit has been moving on both sides of the market.</p><p>Then the question I actually wanted answered. This whole piece has argued that AMD&#8217;s distinguishing choice is paying area for registers, and that this is why it never needed a tensor memory. </p><blockquote><p><em>Does that survive an ISA fork? </em></p></blockquote><p>So I ran the accumulator sweep again on gfx1250, with a WMMA kernel instead of an MFMA one, and pushed until the allocator gave up.</p><p>It gives up at 1,024 registers per lane. A gfx1250 lane has twice the architectural registers of a gfx942 lane, and half as many lanes per wave.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!EILJ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F483a7831-b897-4274-a0f8-48a30bb9f851_1600x640.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!EILJ!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F483a7831-b897-4274-a0f8-48a30bb9f851_1600x640.png 424w, https://substackcdn.com/image/fetch/$s_!EILJ!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F483a7831-b897-4274-a0f8-48a30bb9f851_1600x640.png 848w, https://substackcdn.com/image/fetch/$s_!EILJ!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F483a7831-b897-4274-a0f8-48a30bb9f851_1600x640.png 1272w, https://substackcdn.com/image/fetch/$s_!EILJ!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F483a7831-b897-4274-a0f8-48a30bb9f851_1600x640.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!EILJ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F483a7831-b897-4274-a0f8-48a30bb9f851_1600x640.png" width="1456" height="582" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/483a7831-b897-4274-a0f8-48a30bb9f851_1600x640.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:582,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Vector register state addressable by one wave or warp, bytes. AMD halved the wave and doubled the per-lane file. The product did not move. The ISA forked completely and the resource budget did not move at all.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Vector register state addressable by one wave or warp, bytes. AMD halved the wave and doubled the per-lane file. The product did not move. The ISA forked completely and the resource budget did not move at all." title="Vector register state addressable by one wave or warp, bytes. AMD halved the wave and doubled the per-lane file. The product did not move. The ISA forked completely and the resource budget did not move at all." srcset="https://substackcdn.com/image/fetch/$s_!EILJ!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F483a7831-b897-4274-a0f8-48a30bb9f851_1600x640.png 424w, https://substackcdn.com/image/fetch/$s_!EILJ!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F483a7831-b897-4274-a0f8-48a30bb9f851_1600x640.png 848w, https://substackcdn.com/image/fetch/$s_!EILJ!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F483a7831-b897-4274-a0f8-48a30bb9f851_1600x640.png 1272w, https://substackcdn.com/image/fetch/$s_!EILJ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F483a7831-b897-4274-a0f8-48a30bb9f851_1600x640.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong><span>Figure 13.</span></strong> AMD halved the wave and doubled the per-lane file. The product did not move. The ISA forked completely and the resource budget did not move at all. </figcaption></figure></div><p>512 registers times 64 lanes is 32,768 words. 1,024 times 32 is 32,768 words. A wave on CDNA 3 and a wave on CDNA 5 address exactly the same 128 kilobytes of vector register state. </p><p><strong>AMD threw away the entire matrix instruction set</strong>, changed the wave width, moved to a workgroup processor, adopted clusters and multicast loads, and did not move the register budget by one byte.</p><p>I did not expect that, and I think it is the most informative single fact in the whole census. <strong>The fork is an encoding fork</strong>. The architectural philosophy underneath it, spend transistors on register file and let the kernel keep its working set in registers, is intact. </p><p>Whatever else changes, an AMD wave will still be able to <strong>hold an accumulator that an NVIDIA warp cannot</strong>, by roughly four times, and AMD will still not need a separate address space to put it in.</p><p>One last thing the compiler gives away. The hazard contract survives as an idea and not one instruction of it survives as text:</p><pre><code><code>gfx942   s_waitcnt lgkmcnt(0)   s_nop 4
gfx950   s_waitcnt lgkmcnt(0)   s_nop 5
gfx1250  s_wait_kmcnt 0x0   s_wait_xcnt 0x0   s_delay_alu instid0(VALU_DEP_1)</code></code></pre><p>The <em>unified GCN counters split into separate ones per traffic class</em>, a counter that did not exist before appears, and the structural no-op is replaced by an explicit ALU dependency delay. </p><p>Same principle, every mnemonic renamed, the counters re-partitioned. Anyone who has memorised the <strong>waiting rules for CDNA </strong>has to memorise them again.</p><p>I want to be careful about credit here, because the qualitative version of this observation is not mine. <strong>Chips and Cheese read the gfx1250 patches in LLVM in July</strong> and described the workgroup processor structure and the matrix unit shapes. </p><p>SemiAnalysis, in its Advancing AI coverage three weeks ago, stated plainly that MI455X uses a completely different ISA from MI355X, that every kernel must be independently rewritten and tuned, and characterised gfx1250&#8217;s ISA as close to Hopper&#8217;s. <strong>Both were there before me. </strong></p><p>What I have added is the number, the method that produces it in about four seconds on a laptop, and the <strong>family-level decomposition </strong>that shows the two matrix instruction sets are strictly disjoint rather than overlapping.</p><p>And one corroboration worth flagging. SemiAnalysis reported that gfx1250&#8217;s matrix engine speaks <strong>NVFP4 natively</strong>, citing a scale-format enum that includes an e5m3 member and a gfx1250-compiled NVFP4 GEMM already shipping inside AITER. </p><p>My feature census independently shows <code>fp8e5m3-insts</code> present on gfx1250 and absent from <strong>every CDNA target before it</strong>. Two different artefacts, same conclusion.</p><p>Now the interpretation, and this is where I think the obvious reading is wrong.</p><p>The obvious reading is that AMD has thrown away its software investment, and against the census that reading has force:<strong> 16,177 tuned assembly kernels and 2.9 million shape mappings </strong>are worth nothing on a target that cannot encode the instructions they are built from. AITER&#8217;s fastest paths are hand-written wave64 assembly. Every one is a rewrite.</p><p>But look at what the corpus is. Not knowledge, a lookup table. The knowledge sits in the generator, in the parameter space, in knowing which of those 57 dimensions matter and how they interact, and in the harness that runs the search. </p><p>Tensile is a program that emits kernels, and retargeting a generator is work but not the same work as rediscovering what to generate. What has to run again is the campaign, <strong>on hardware AMD must own</strong>, for every part and every harvest bin. That is expensive in machine time and calendar time rather than in insight.</p><p>Which is precisely what AMD&#8217;s July announcements are about, and it took me a while to see them as a coherent response rather than as agent marketing. </p><p><strong>ROCm.ai bundles a CLI,</strong> a set of AMD-authored skills for coding agents, and Hyperloom, an open-source agentic system whose stated job is automating end-to-end inference workload optimisation.</p><p> Read against the census, that is not a developer-relations play. It is an attempt to industrialise the search, because the search is the cost, and because a company that has just invalidated its own tuning corpus and has <strong>fewer internal cluster hours</strong> than its competitor needs the search to get cheaper by more than the corpus got smaller.</p><p>Whether it works is empirical, and the evidence so far cuts both ways.</p><p>On the encouraging side, <strong>AMD&#8217;s GEAK v4 write-up</strong> reports serving throughput gains on real workloads rather than on kernel microbenchmarks: 60 percent on Qwen3.5-27B-FP8, 96 percent on Qwen3-14B-FP8, 42 percent on Minimax-M3-MXFP8. </p><p>Those are <em>AMD&#8217;s own numbers</em> on AMD&#8217;s own hardware, so discount accordingly, but they are end-to-end serving throughput on named models, which is the hard version of the claim and not the easy one.</p><p>Against that, <strong>KernelBench-Verified </strong>showed that when the baseline is TF32 and the tests are hidden, the best model produces a geometric mean of 0.88 times rather than the 1.43 the earlier literature reported, which suggests much of the reported gain was the distance between PyTorch&#8217;s defaults and PyTorch&#8217;s available settings.</p><p>My reconciliation last year was that the value sits in the harness rather than the model: the profiler, the knowledge base, the end-to-end gate. <strong>GEAK v4&#8217;s own description</strong> of itself is now unusually direct evidence for that. </p><p>It puts control flow, budget loops, fan-out, verification and stop conditions in deterministic JavaScript, and invokes the model only for structured technical judgment. </p><blockquote><p><em>That is a harness with a model bolted into it at the points where judgment is needed, which is exactly the shape the KernelBench correction implies you need.</em></p></blockquote><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Where I might be wrong</h2><p>The matrix builtin count is a count of compiler builtins, not of encodable instructions. <strong>Clang exposes builtins for the instructions LLVM has intrinsics for</strong>, and it is possible that gfx1250 encodes something MFMA-shaped that has no builtin and that AMD reaches through inline assembly.</p><p> I looked for a fallback and did not find one, and the feature bit <code>mai-insts</code> being absent is a stronger signal than a missing builtin, but absence of a builtin is not proof of absence of an encoding.</p><p>I have not run a single instruction on <strong>AMD hardware</strong>. Everything here is compiler behaviour, published specification, repository content and arithmetic. </p><p>Compiler behaviour is a good proxy for what is encodable and a poor proxy for what is fast. Where I quote performance numbers they are somebody else&#8217;s, and I have tried to say whose.</p><p>The <em>Tensile census is a snapshot of one branch of one repository</em>, and that branch is called <code>develop_deprecated</code> with a head commit from June 2025. </p><p>The gfx942 numbers in particular are almost certainly not the current state, and the gfx950 tuning did not exist in the tree I cloned. </p><p>I believe the shape of the result, that the tuning surface is millions of measured decisions over a <strong>space of about ten to the twenty second</strong>, is robust to the snapshot. The specific integers are not.</p><p>The 120 bit parameter space is an upper bound computed by summing the <strong>log of observed cardinalities</strong>, and it assumes an independence that certainly does not hold: many combinations are invalid, and Tensile&#8217;s protocol exists precisely to avoid enumerating them. </p><p>Read it as the volume the search has to reason about, not the number of legal kernels. It also treats the macro tile as <strong>one parameter with 249 observed values </strong>rather than three separate dimensions, which makes it conservative in the other direction.</p><p>The<em> bytes-of-state-per-FLOP ratio </em>is a derived quantity with a choice baked in, namely that scratchpad and register file are commensurable. They are not, exactly. </p><p>LDS is shared across a workgroup and registers are private to a lane, and <strong>228 KB of Hopper shared memory</strong> is not interchangeable with 228 KB of register file. </p><p>The ratio is a way of seeing a trend across generations, and I would not defend a comparison between two parts that <em>differed by ten percent on it. </em>Finally, the reading that AMD&#8217;s corpus is a lookup table and the knowledge is in the generator is an argument, not a measurement. </p><p>Somebody with a Tensile retargeting on their hands could tell me it took eighteen months, and I would have no basis to argue.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Five predictions, dated</h2><p>I score these publicly, so here they are with a resolution date and a criterion.</p><p>By the end of 2027, hipBLASLt or its successor will ship a gfx1250 logic tree whose count of tuned problem shapes is at least a quarter of the gfx942 tree&#8217;s at the same point in that part&#8217;s life. If the retarget is as cheap as I argue, this is easy. If it is not, this is where it shows.</p><ul><li><p>By the end of 2027, AMD will not ship a compatibility layer that lets an MFMA-based kernel run unmodified on gfx1250. Emulation is possible and would be slow enough to be pointless, and I do not think they will bother.</p></li><li><p>By mid 2027, a published Hyperloom or GEAK result will report an end-to-end inference gain above 15 percent on gfx1250 with no human kernel author in the loop, on a model somebody else chose. The same claim on gfx950 is already met on AMD&#8217;s own numbers; the question is whether it transfers to a target with no tuning corpus behind it, which is the load-bearing question for the whole strategy.</p></li><li><p>By the end of 2027, at least one serious third-party project will publish a kernel for gfx1250 that beats AMD&#8217;s own library on a shape that matters, and will do it in Gluon or Triton rather than assembly. The open ISA has never produced a competitive third-party GEMM. If it is going to, the moment when the vendor&#8217;s own corpus is empty is the moment.</p></li><li><p>By the end of 2028, NVIDIA will not have moved hazard management out of the instruction control field and into architectural instructions. The choice is an area and energy decision from 2012 and there is no sign of a reversal.</p></li><li><p>And the one I got wrong last time. In July 2025 I predicted at least one more CUDA-compatibility project would lose its funding. Qualcomm announced an all-stock acquisition of Modular on 24 June 2026, valued near 3.9 billion dollars, and closed it on 29 July. That is not losing your funding.</p></li></ul><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/p/distilling-in-depth-rocm-how-it-actually?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/p/distilling-in-depth-rocm-how-it-actually?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><div><hr></div><h2>What the whole thing adds up to</h2><p>I came into this expecting to write that AMD&#8217;s openness is undervalued, and I am leaving with a more specific and less comfortable version of that.</p><p>The<strong> openness is real</strong> and deeper than the marketing suggests. The dispatch packet is a documented struct. The code object is an ELF file standard tools read. The target feature axes are orthogonal with a neutral value, which is better engineering than the alternative. </p><p>The <strong>hazard model is expressible in the instruction stream</strong>, so hand-written assembly is a viable production strategy and third parties can pursue it. The compiler is upstream. That I could do all of the work in this piece on a machine with no GPU in it is a direct demonstration of the value.</p><p><strong>Openness makes that cost visible rather than smaller</strong>. That is still the more useful of the two situations, because a measurable cost is one somebody can attack with automation.</p><p>Then AMD did the one thing that makes the cost bigger, which is change ISA families between generations, at the moment it is shipping into gigawatt-scale commitments. Zero of seventy five. I do not think that was a mistake, exactly. </p><p><strong>Converging the datacenter and consumer instruction sets</strong> is the right long-run move for a company that cannot afford two of everything, and gaining clusters, multicast loads, prefetch and a native NVFP4 path is worth something real. </p><p>But it means the next eighteen months of AMD&#8217;s software story is a rerun of the tuning campaign, with agents in the loop instead of engineers, on fewer machines than the competitor has.</p><p><strong>It is a real bet</strong>, it is falsifiable on a timescale of months, and we will know.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Confidence dossier</h2><p>Every load-bearing claim in this piece, properly graded manually.</p><p><strong><span>M </span></strong>gfx950 to gfx1250 carries over 0 of 75 matrix builtins; gfx90a to gfx942 and gfx942 to gfx950 carry 100 percentm1_isa_census.py, clang 22.1.0</p><p><strong><span>M </span></strong>clang rejects __builtin_amdgcn_mfma_f32_32x32x8f16 on gfx1250 and gfx1251 with &#8216;needs target feature mai-insts&#8217;m1_isa_census.py.</p><p><strong><span>M </span></strong>gfx1250 loses 24 gfx950 features including mai-insts, wavefrontsize64, all dot-insts, all CDNA4 scale conversions, s_memtime, s_memrealtimem1_isa_census.py</p><p><strong><span>M </span></strong>gfx1250 gains 22 features including clusters, mcast-load-insts, vmem-pref-insts, tensor-cvt-lut-insts, transpose-load-f4f6-insts, fp8e5m3-insts, wavefrontsize32m1_isa_census.py</p><p><strong><span>M </span></strong>AMDGPU builtins in clang: 443 at LLVM 20.1.2, 788 at LLVM 22.1.0, 345 added, 0 removedm1_isa_census.py. </p><p><strong><span>M </span></strong>Builtins available per target: 244 gfx90a, 274 gfx942, 362 gfx950, 477 gfx1250m1_isa_census.py. </p><p><strong><span>M </span></strong>MFMA builtins per target: 31 / 39 / 47 / 0. SMFMAC: 6 / 14 / 28 / 0. WMMA: 0 / 0 / 0 / 73m1_isa_census.py</p><p><strong><span>M </span></strong>Target feature counts: gfx90a 27, gfx942 34, gfx950 48, gfx1250 46, gfx1251 46m1_isa_census.py. </p><p><strong><span>M </span></strong>gfx942 has xf32-insts; gfx950 does not. gfx950 adds 15 features over gfx942clang 20.1.2 feature diff.  </p><p><strong><span>M </span></strong>A gfx942 lane holds 24 independent 32x32x8 MFMA accumulators (384 registers) with zero spill and the whole 512-entry file allocated; 25 is where scratch traffic startsm5_cdna5.py, full sweep.</p><p><strong><span>M </span></strong>The gfx942 spill tail is not monotone: 25 to 29 spill, 30 does not (500 registers, 480 of accumulator), 31 and 32 spill againm5_cdna5.py, every N from 1 to 34.</p><p><strong><span>M </span></strong>A gfx1250 lane holds 1,024 architectural registers in wave32; the allocator caps vgpr_count at 1024 and first spills at 128 WMMA accumulatorsm5_cdna5.py plus a manual probe to N=256. </p><p><strong><span>M </span></strong>512 registers x 64 lanes and 1,024 x 32 lanes are both 32,768 words, so a CDNA 3 wave and a CDNA 5 wave address the same 128 KB of vector register statearithmetic on the two measured ceilings. </p><p><strong><span>M </span></strong>Output tiles per target: gfx90a, gfx942 and gfx950 all offer 4x4, 16x16 and 32x32; gfx1250 offers 16x16 and a new 32x16 onlym5_cdna5.py. </p><p><strong><span>M </span></strong>Distinct matrix shapes fall from 18 on gfx950 to 6 on gfx1250; largest K on a 16x16 tile is 64 on gfx942 and 128 on gfx950 and gfx1250m5_cdna5.py. </p><p><strong><span>M </span></strong>The hazard instructions are all renamed on gfx1250: s_waitcnt lgkmcnt and s_nop become s_wait_kmcnt, s_wait_xcnt and s_delay_alum5_cdna5.py. </p><p><strong><span>M </span></strong>Back-computed FLOPs per clock per unit land within 0.7 percent of a power of two on every AMD part and 4 to 6 percent off on H100 and B200, so at least one published NVIDIA input is not the number usedm4_model.py. </p><p><strong><span>M </span></strong>At 16 accumulators the allocator reports 288 registers of which 32 AGPR, and sets ACCUM_OFFSET to 256m2_hazards_registers.py</p><p><strong><span>M </span></strong>gfx90a, gfx942 and gfx950 emit .amdhsa_accum_offset and .agpr_count; gfx1250 emits neither and defaults to wave32m2_hazards_registers.py. </p><p><strong><span>M </span></strong>A gfx942 code object for a single MFMA loop is 4,848 bytes: ELF64, OS/ABI AMDGPU_HSA, 15 sections, 13 symbolsllvm-readobj on ko_gfx942.hsaco.</p><p><strong><span>M </span></strong>e_flags 0x54C decodes to EF_AMDGPU_MACH_AMDGCN_GFX942 plus XNACK_ANY_V4 plus SRAMECC_ANY_V4llvm-readobj --file-headers.</p><p><strong><span>M </span></strong>The kernel descriptor is 64 bytes in .rodata; decoded rsrc1 0x00af0082, rsrc2 0x00000084, rsrc3 0x00000000, properties 0x0008kernel descriptor decode.</p><p><strong><span>M </span></strong>Same source, 1,472 bytes of .text on all three CDNA targets; 113 bytes differ gfx90a to gfx942, 33 bytes gfx942 to gfx950llvm-objcopy byte diff. </p><p><strong><span>M </span></strong>hipBLASLt logic tree: 145 files, 26,248 solutions, 16,177 distinct kernel names, 2,905,048 shape mappings, 310 MBm3 census, develop_deprecated at 3a609b0.</p><p><strong><span>M </span></strong>MI200 tuning is split by compute unit count into 104CU and 110CU directories with separate tablesrepository layout.</p><p><strong><span>M </span></strong>The 1,439 distinct gfx942 kernel names encode 56 tuning parameters with more than one observed value, summing to 119.9 bits, which is 1.28e36m6_tuning_space.py. </p><p><strong><span>M </span></strong>The shipped set is 1,439 kernels, 10.5 bits, which is 1.1e-33 of the spacem6_tuning_space.py.</p><p><strong><span>M </span></strong>Bits by group: scheduling and hints 38.7 over 24 parameters, split along K 18.6 over 9, tile geometry 17.8 over 5, vector widths 17.6 over 9, LDS layout 15.4 over 6, XCD mapping 11.9 over 3m6_tuning_space.py.</p><p><strong><span>M </span></strong>Macro tile MT takes 249 distinct values (7.96 bits); WGM 32; GSU 30; WGMXCCG 20; LBSPPA 16; MIWT 14m6_tuning_space.py. </p><p><strong><span>M </span></strong>B* is identical across FP16, FP8 and FP4 to within 1 percent on every part measuredm4_model.py. </p><p><strong><span>M </span></strong>B*: MI300X 246.6, MI325X 217.8, MI350X 287.5, MI355X 312.5, H100 295.4, H200 206.1, B200 281.2m4_model.py.</p><p><strong><span>M </span></strong>Bytes of on-chip state per FP8 FLOP per clock per unit: MI300X 144.0, MI355X 84.6, H100 58.0, B200 32.0m4_model.py. </p><p><strong><span>M </span></strong>Vector register file per chip: MI300X 152 MB, MI355X 128 MB, H100 33 MB, B200 37 MBm4_model.py. </p><p><strong><span>M </span></strong>70B FP8 with GQA-8 at 90 percent capacity: H100 12,207 KV tokens, MI300X 627,441, MI355X 1,154,785m4_model.py. </p><p><strong><span>M </span></strong>Matrix throughput per clock per unit: MI300X 4,096 FP8, MI355X 8,138, H100 8,543, B200 15,473m4_model.py. </p><p><strong><span>M </span></strong>hipBLASLt has 180 remote branches; its default branch is named develop_deprecated with head from 2025-06-20git ls-remote. </p><p><strong><span>A </span></strong>MI455X: 432 GB HBM4, up to 23.3 TB/s, 40 PFLOPS FP4, 20 PFLOPS FP8, 8 XCDs, CDNA 5; the 19.6 TB/s figure that circulated before launch is supersededAMD MI400 series product page. </p><p><strong><span>M </span></strong>B* for MI455X is 429 at both FP8 and FP4, up 37 percent on MI355X, so the memory-bound regime is shrinkingm4_model.py with AMD published figures.</p><p><strong><span>M </span></strong>MI455X holds 1,945,801 tokens of FP8 KV beside a 70B FP8 model at 90 percent of capacitym4_model.py. </p><p><strong><span>A </span></strong>Helios delivers up to 1.4 exaFLOPS FP8 and 2.9 exaFLOPS FP4 with 31 TB of HBM4AMD MI400 series product page.</p><p><strong><span>A </span></strong>MI355X: 256 CU, 288 GB HBM3E, 8.0 TB/s, 160 KB LDS per CU, 2.5 PF FP16 dense, 10 PF FP4, 1,400 WAMD product page and ROCm workload optimization guide.</p><p><strong><span>A </span></strong>MI300X: 304 CU, 192 GB HBM3, 5.3 TB/s, 64 KB LDS per CU, 4 IODs, 8 XCDs, 256 MB Infinity CacheROCm workload optimization guide.</p><p><strong><span>A </span></strong>CDNA 3 uses FP8 FNUZ variants; CDNA 4 uses OCP variants. TF32 moves to software emulation via BF16 on CDNA 4ROCm workload optimization guide. </p><p><strong><span>A </span></strong>FP64 matrix halves on CDNA 4, 128 versus 256 FLOPs per clock per CUROCm workload optimization guide. </p><p><strong><span>A </span></strong>E4M3FN and E4M3FNUZ share a bit layout, differ in exponent bias by one, so a misread byte is off by a factor of twoAMD Matrix Core blog; Fergus Finn. </p><p><strong><span>A </span></strong>buffer_load_to_lds saves about 100 VGPR per wave and moved a reference gfx950 GEMM from 697 to 1113 TFLOPSROCm workload optimization guide, Gluon section. </p><p><strong><span>A </span></strong>XCD-aware workgroup remapping cut L2 misses from about 5M to 3.1M and added about 67 TFLOPSROCm workload optimization guide. </p><p><strong><span>A </span></strong>XCD clocks vary 3 to 10 percent on one package; XCD0 typically fastest, XCD7 slowest on MI300XROCm workload optimization guide. </p><p><strong><span>A </span></strong>A GEMM stride that is a multiple of 512 bytes causes channel hotspotting on MI300; pad to K+128 when K%256==0ROCm workload optimization guide. </p><p><strong><span>A </span></strong>MI16x16 outperforms MI32x32 on MI300X for GEMM, attributed to power efficiencyROCm workload optimization guide. </p><p><strong><span>A </span></strong>ROCm serialises kernel launches across GPUs from one process; RCCL wants one process per GPU; GPU_MAX_HW_QUEUES=2 recommendedROCm workload optimization guide. </p><p><strong><span>A </span></strong>ROCm 7.0 to 7.8 is the production stream, 7.9 and later is technology preview; production is 7.2.x, preview reached 7.14ROCm release version pages. </p><p><strong><span>A </span></strong>Helios: 72 MI455X, 18 EPYC Venice, up to 2.9 EF FP4, 31 TB pooled HBM4, MI455X at 432 GBAMD Advancing AI 2026 materials. </p><p><strong><span>A </span></strong>AMD claims Helios delivers up to 30 percent more tokens per dollar than the leading competitive solutionAMD press release, 23 July 2026. </p><p><strong><span>A </span></strong>ROCm.ai comprises ROCm CLI, AMD Skills for coding agents, and Hyperloom, an open-source agentic optimisation systemAMD newsroom, 23 July 2026. </p><p><strong><span>B </span></strong>vLLM issue 45562 and PR 45720: FNUZ KV bytes read as FN on gfx942 degraded GSM8K accuracy; the dtype fix recovered itvLLM GitHub, June 2026. </p><p><strong><span>B </span></strong>AITER issue 3807: building for gfx942 and gfx950 together returned the OCP dtype on an MI300XROCm/aiter GitHub, June 2026. </p><p><strong><span>B </span></strong>Both AITER MLA backends share the assembly decode kernel mla_decode_fwd; most of the 1.2 to 1.6x gain is attributed to itvLLM blog, 27 February 2026. </p><p><strong><span>B </span></strong>MI300X SGLang throughput roughly doubled between December 2025 and January 2026, attributed to AITERSemiAnalysis InferenceX v2, via HyperAccel analysis. </p><p><strong><span>B </span></strong>MI455X is gfx1250 and MI430X is gfx1251; gfx1250 is WGP-based, wave32, and listed as an APU in LLVMChips and Cheese, July 2026. </p><p><strong><span>B </span></strong>gfx1250&#8217;s matrix engine supports NVFP4 natively; a gfx1250 NVFP4 GEMM code object already ships in AITERSemiAnalysis, Advancing AI 2026 coverage. </p><p><strong><span>B </span></strong>SemiAnalysis states every MI355X kernel must be independently rewritten and tuned for MI455XSemiAnalysis, Advancing AI 2026 coverage. </p><p><strong><span>B </span></strong>GEAK v4 reports serving throughput gains of 60 percent on Qwen3.5-27B-FP8, 96 percent on Qwen3-14B-FP8 and 42.2 percent on Minimax-M3-MXFP8, self-measuredAMD GEAK v4 technical article, 23 July 2026. </p><p><strong><span>A </span></strong>GEAK v4 puts control flow, budget loops, fan-out, verification and stop conditions in deterministic JavaScript and invokes the model only for structured technical judgmentAMD GEAK v4 technical article. </p><p><strong><span>A </span></strong>Hyperloom orchestrates five components: TraceLens for bottleneck identification, GEAK and Arbor for optimisation in parallel, over Magpie and IntelliKit for profilingROCm Hyperloom documentation and repository. </p><p><strong><span>B </span></strong>KernelBench-Verified: best model reaches 0.88x geomean against a TF32 baseline with hidden testsarXiv:2607.16241. </p><p><strong><span>C </span></strong>The CDNA corpus is a lookup table and the reusable asset is the generator plus the search harness, so retargeting costs machine time rather than insightauthor&#8217;s argument. </p><p><strong><span>C </span></strong>ROCm.ai and Hyperloom are best read as an attempt to industrialise the tuning search in response to the gfx1250 retargetauthor&#8217;s inference from timing and content. </p><p><strong><span>C </span></strong>The absence of an MFMA fallback encoding on gfx1250 is inferred from the absent mai-insts feature bit, not provedauthor&#8217;s inference.</p><p><strong><span>B </span></strong>InferenceX, 20 May 2026: MI355X SGLang FP8 up to about 40 percent cheaper per million tokens than B200 on GLM-5 8K/1K, peak gap at 18 tok/s/user, 22 cents against 30; B200 ahead above about 90 tok/s/userInferenceX blog post, measured 2026-05-20. </p><p><strong><span>B </span></strong>InferenceX overview, July 2026 TCO model: MI355X 35.5 cents per million on 8K/1K FP4 against B200 at 30.4, i.e. 17 percent more expensive; the long-context agentic scenario lists a far larger gapInferenceX overview page. </p><p><strong><span>B </span></strong>The reversal is attributed to B200&#8217;s NVFP4 path shipping and to AMD lacking a disaggregation or wide expert parallel recipe for that modelInferenceX blog post, same source. </p><p><strong><span>A </span></strong>Qualcomm acquired Modular in an all-stock deal valued near 3.9 billion dollars, announced 24 June 2026, closed 29 July 2026Qualcomm and Modular announcements, SEC filing. </p><p><strong><span>C </span></strong>A memory-bound decode bounds the kernel gap above by the ratio of achieved HBM bandwidths, which is why an immature kernel layer costs less in that regimeauthor&#8217;s argument. </p><p><strong><span>C </span></strong>The inference and training asymmetry follows from hot-kernel-set size, the memory-bound regime and capacity, not primarily from software maturityauthor&#8217;s argument. </p><p><strong><span>C </span></strong>A figure of about 92 percent CUDA device API coverage for HIP circulates in channel material and could not be traced to a primary AMD sourceauthor&#8217;s search. </p><p><strong><span>A </span></strong>The kernel descriptor for the sample kernel sets ACCUM_OFFSET to 0, so AGPRs begin at register 4, and the metadata agrees: 20 registers total of which 16 are AGPRskernel descriptor decode, cross-checked against .agpr_count. </p><p><strong><span>A </span></strong>vmcnt returns in order and lgkmcnt does not, which is why nonzero lgkmcnt waits are rare when scalar loads are involvedLLVM AMDGPUUsage and generated code. </p><p><strong><span>A </span></strong>Unpadded shared layouts cause 2-way to 4-way LDS bank conflicts, cutting the effective rate from 256 B/cycle to 64-128ROCm workload optimization guide, Gluon section. </p><p><strong><span>A </span></strong>Eight MI300X modules are fully connected by seven Infinity Fabric links each; AMD advises using one GPU or all eight for collectivesROCm workload optimization guide. </p><p><strong><span>A </span></strong>rocprof, rocprofv2, ROCProfiler and ROCTracer are deprecated with end of support announced for 2026 Q2; PyTorch still depended on ROCTracer at 7.2.1ROCm 7.2.1 release notes. </p><p><strong><span>D </span></strong>AMD&#8217;s long-run direction is a single unified datacenter and consumer ISA, of which gfx1250 is the first datacenter memberspeculation consistent with the feature set. </p><h3>Tier Definition </h3><p><strong><span>M. </span></strong>Measured here. Produced by a script in this piece on a machine with no GPU, reproducible from the appendix.</p><p><strong><span>A. </span></strong>Primary and verifiable. Vendor documentation, specification, or repository content read directly.</p><p><strong><span>B. </span></strong>Credible secondary. Reported by an identified party with a method, not independently reproduced here.</p><p><strong><span>C. </span></strong>Inference. Follows from the evidence but is an argument rather than a measurement.</p><p><strong><span>D. </span></strong>Speculation. Stated as such.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><p></p>
      <p>
          <a href="https://www.thesoftwarefrontier.com/p/distilling-in-depth-rocm-how-it-actually">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[Exploring how Triton actually compiles]]></title><description><![CDATA[Same source, same block sizes, same warp count. An integer the programmer never writes decides whether the loop is pipelined at all, and a second one decides whether you are on the tensor core.]]></description><link>https://www.thesoftwarefrontier.com/p/exploring-how-triton-actually-compiles</link><guid isPermaLink="false">https://www.thesoftwarefrontier.com/p/exploring-how-triton-actually-compiles</guid><dc:creator><![CDATA[Lorenzo Bradanini]]></dc:creator><pubDate>Sat, 22 Aug 2026 12:53:23 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!slye!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F16cd618f-3249-4877-962b-2bf0eaa535a7_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!slye!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F16cd618f-3249-4877-962b-2bf0eaa535a7_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!slye!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F16cd618f-3249-4877-962b-2bf0eaa535a7_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!slye!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F16cd618f-3249-4877-962b-2bf0eaa535a7_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!slye!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F16cd618f-3249-4877-962b-2bf0eaa535a7_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!slye!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F16cd618f-3249-4877-962b-2bf0eaa535a7_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!slye!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F16cd618f-3249-4877-962b-2bf0eaa535a7_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/16cd618f-3249-4877-962b-2bf0eaa535a7_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2454267,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/208654595?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F16cd618f-3249-4877-962b-2bf0eaa535a7_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!slye!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F16cd618f-3249-4877-962b-2bf0eaa535a7_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!slye!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F16cd618f-3249-4877-962b-2bf0eaa535a7_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!slye!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F16cd618f-3249-4877-962b-2bf0eaa535a7_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!slye!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F16cd618f-3249-4877-962b-2bf0eaa535a7_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>Why was my kernel slower than an identical kernel?</h2><p>The honest version of how this piece started is that I was deeply annoyed. I had a <strong>Triton kernel </strong>that was slower than a kernel that someone else had written, and the two files differed by nothing I could point at. </p><p>Same block sizes. Same <code>num_stages</code>. Same <code>num_warps</code>. I read both of them several times, in the way you do when you are convinced you are<strong> missing something obvious</strong> and you are not yet willing to consider that the thing you are missing might not be in the file.</p><p>I didn&#8217;t have a GPU that weekend. So I did the thing I have been doing for a while now with <strong>NVIDIA&#8217;s toolchain</strong>, which essentially consiosts to install the compiler and take it apart on the host, and it turned out that the answer was not in either file.</p><p>The setup works better than it has any right to. Triton&#8217;s entire compilation pipeline runs on the CPU. <strong>Source becomes TTIR</strong>, then TTIR becomes TTGIR,  which becomes LLVM IR; this piece of code gets tanslated into PTX, and PTX becomes a cubin through <code>ptxas</code>, which is an x86 binary shipped inside the wheel. </p><p>Nothing in that chain touches a device. You need a GPU to <em>run</em> a Triton kernel. You do not need one to find out what the compiler decided, and what the compiler decided is most of the story.</p><pre><code><span># 197 MB from PyPI. No CUDA install, no driver, no device.</span>
pip install triton==3.7.1 --no-deps --target ./t

<span># the full pipeline, on one CPU core</span>
ck = triton.compile(src, target=GPUTarget(<span>&#8220;cuda&#8221;</span>, 90, 32),
                    options={<span>&#8220;num_warps&#8221;</span>: 8, <span>&#8220;num_stages&#8221;</span>: 3})
list(ck.asm.keys())
<span># [&#8217;source&#8217;, &#8216;ttir&#8217;, &#8216;ttgir&#8217;, &#8216;llir&#8217;, &#8216;ptx&#8217;, &#8216;cubin&#8217;]</span>

<span># and, from the same wheel, the same source, a different vendor</span>
ck = triton.compile(src, target=GPUTarget(<span>&#8220;hip&#8221;</span>, <span>&#8220;gfx942&#8221;</span>, 64), ...)
<span># [&#8217;source&#8217;, &#8216;ttir&#8217;, &#8216;ttgir&#8217;, &#8216;llir&#8217;, &#8216;amdgcn&#8217;, &#8216;hsaco&#8217;]</span></code></pre><p>What follows is a census of the decisions Triton makes for you, measured from Triton, on the theory that a decision you can&#8217;t see is more interesting than one you can. </p><p>It&#8217;s not an introduction, neither a benchmark. As said before, we have no hardware here and we make no claims about wall clock speed. Everything is what the compiler emits.</p><p>Three questions organise it, and we found answers to all three that we did not expect:</p><blockquote><ul><li><p><strong>What decides whether your loop is pipelined?</strong> Not <code>num_stages</code>. An alignment attribute the JIT infers from your pointers at launch, with a threshold at four bytes and a ramp to sixteen.</p></li><li><p><strong>What decides which matrix instruction you get?</strong> On Hopper, <code>num_warps</code>. Below four warps Triton silently drops off the warpgroup path and back onto an Ampere-era instruction.</p></li><li><p><strong>Where is the compiler&#8217;s real work?</strong> Not in the arithmetic. Eighty eight percent of the passes Triton contributes, over and above what it inherits from MLIR, are about where data lives and when it moves.</p></li></ul></blockquote><p><strong>On method.</strong> We wrote the verification harness before the prose, a rule we adopted after the last piece. It is 117 assertions against a live Triton install and it exits non-zero if any of them fail. </p><p>Building it caught a number we had already written down wrong, for the second time in three articles. Every figure below is tagged in the dossier, and you can reproduce some of the outputs too. </p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Thirty thousand lines of Python on 180 megabytes of C++</h2><p>We keep seeing that Triton is described by everyone, almost universally, as a <strong>Python DSL</strong>. That is both true and wrong. It is correct for the part you write, but it&#8217;s a poor description of the whole artifact. </p><p>The Python frontend in 3.7.1 is 115 files and 29,984 lines. It sits on <code>libtriton.so</code>, which is 461,559,216 bytes as shipped and 180,128,520 stripped of debug information.</p><p>What is interesting to compare actually, is with the closed compiler underneath, because the usual framing has Triton as a thin open layer over an<strong> opaque assembler</strong>. </p><p>Stripped Triton is 4.35 times the size of the Blackwell <code>ptxas</code> it feeds. This is a huge number.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!qILz!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb35f00c-708f-441a-8b00-e77b797120f6_1820x842.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!qILz!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb35f00c-708f-441a-8b00-e77b797120f6_1820x842.png 424w, https://substackcdn.com/image/fetch/$s_!qILz!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb35f00c-708f-441a-8b00-e77b797120f6_1820x842.png 848w, https://substackcdn.com/image/fetch/$s_!qILz!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb35f00c-708f-441a-8b00-e77b797120f6_1820x842.png 1272w, https://substackcdn.com/image/fetch/$s_!qILz!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb35f00c-708f-441a-8b00-e77b797120f6_1820x842.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!qILz!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb35f00c-708f-441a-8b00-e77b797120f6_1820x842.png" width="1456" height="674" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/db35f00c-708f-441a-8b00-e77b797120f6_1820x842.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:674,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;wheel anatomy&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="wheel anatomy" title="wheel anatomy" srcset="https://substackcdn.com/image/fetch/$s_!qILz!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb35f00c-708f-441a-8b00-e77b797120f6_1820x842.png 424w, https://substackcdn.com/image/fetch/$s_!qILz!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb35f00c-708f-441a-8b00-e77b797120f6_1820x842.png 848w, https://substackcdn.com/image/fetch/$s_!qILz!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb35f00c-708f-441a-8b00-e77b797120f6_1820x842.png 1272w, https://substackcdn.com/image/fetch/$s_!qILz!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb35f00c-708f-441a-8b00-e77b797120f6_1820x842.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Figure 1.</strong> The Python frontend everyone describes Triton by is the smallest component in the wheel. Sizes are on a logarithmic axis; <code>libtriton.so</code> is 461,559,216 bytes before stripping.</figcaption></figure></div><p>The ratio of 80/1 between the two backend directories will get quoted by everyone without its caveat, but in this case the caveat we are about to explore is not a measure of engineering effort.</p><p>As you may know, AMD&#8217;s code generation happens <strong>inside LLVM,</strong> which is already linked into <code>libtriton.so</code>; on the other side, NVIDIA&#8217;s requires shipping proprietary binaries that cannot be linked. </p><p>Thedirectories are said to be described as asymmetric, just because the licensing is asymmetric. </p><p>The<strong> narrow version of the claim</strong> survives and is still worth something: a portable compiler carries 219 MB of one vendor&#8217;s closed tooling inside its own wheel, and carries none of the other vendor&#8217;s because there is nothing closed to carry.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Which half of this stack is slow?</h2><p>I assumed multiple times, before measuring, that <code>ptxas</code> was the expensive part, and for a good reason. It is the closed one, it does register allocation and<strong> instruction scheduling</strong>, and it is the thing people complain about when Triton compile times come up. </p><p>Spoiler alert: it is not the expensive part. Compiling the 128 by 128 by 64 matmul for <code>sm_90a</code>, five cold runs with the cache cleared between them, taking medians:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!TLWC!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe594bf2b-839e-48a7-b080-476ff4d16fc5_1820x530.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!TLWC!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe594bf2b-839e-48a7-b080-476ff4d16fc5_1820x530.png 424w, https://substackcdn.com/image/fetch/$s_!TLWC!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe594bf2b-839e-48a7-b080-476ff4d16fc5_1820x530.png 848w, https://substackcdn.com/image/fetch/$s_!TLWC!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe594bf2b-839e-48a7-b080-476ff4d16fc5_1820x530.png 1272w, https://substackcdn.com/image/fetch/$s_!TLWC!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe594bf2b-839e-48a7-b080-476ff4d16fc5_1820x530.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!TLWC!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe594bf2b-839e-48a7-b080-476ff4d16fc5_1820x530.png" width="1456" height="424" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e594bf2b-839e-48a7-b080-476ff4d16fc5_1820x530.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:424,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;compile time split&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="compile time split" title="compile time split" srcset="https://substackcdn.com/image/fetch/$s_!TLWC!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe594bf2b-839e-48a7-b080-476ff4d16fc5_1820x530.png 424w, https://substackcdn.com/image/fetch/$s_!TLWC!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe594bf2b-839e-48a7-b080-476ff4d16fc5_1820x530.png 848w, https://substackcdn.com/image/fetch/$s_!TLWC!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe594bf2b-839e-48a7-b080-476ff4d16fc5_1820x530.png 1272w, https://substackcdn.com/image/fetch/$s_!TLWC!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe594bf2b-839e-48a7-b080-476ff4d16fc5_1820x530.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Figure 2.</strong> Triton&#8217;s own MLIR pipeline accounts for four fifths of compile time. This is a lower bound on the Triton share: the outer timing includes a Python AST walk while the inner one is a bare <code>ptxas -O3</code> invocation.</figcaption></figure></div><p>This matters for <strong>two concrete reasons. </strong></p><ol><li><p>The first is more practical: if you are waiting on Triton compile times, tuning <code>ptxas</code> flags is not where the win is. </p></li><li><p>The second is more interesting. The open part of this stack is where the complexity is, by a factor of four in time and a factor of four in binary size.</p></li></ol><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>How much of &#8220;portable&#8221; survives contact with the backend?</h2><p>Take just a single one matmul. Compile it for Ampere, Hopper, Blackwell datacenter, Blackwell consumer, and<strong> three AMD targets.</strong> The TTIR is byte for byte identical across all seven, at 165 lines. </p><p>Everything the language expresses is architecture neutral. Every architectural decision happens below it. What comes out the other end is not neutral in any sense.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!WJdx!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda74bf6b-cd37-4c63-9ab9-ecbe8a3feb0e_1820x936.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!WJdx!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda74bf6b-cd37-4c63-9ab9-ecbe8a3feb0e_1820x936.png 424w, https://substackcdn.com/image/fetch/$s_!WJdx!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda74bf6b-cd37-4c63-9ab9-ecbe8a3feb0e_1820x936.png 848w, https://substackcdn.com/image/fetch/$s_!WJdx!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda74bf6b-cd37-4c63-9ab9-ecbe8a3feb0e_1820x936.png 1272w, https://substackcdn.com/image/fetch/$s_!WJdx!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda74bf6b-cd37-4c63-9ab9-ecbe8a3feb0e_1820x936.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!WJdx!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda74bf6b-cd37-4c63-9ab9-ecbe8a3feb0e_1820x936.png" width="1456" height="749" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/da74bf6b-cd37-4c63-9ab9-ecbe8a3feb0e_1820x936.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:749,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;seven targets&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="seven targets" title="seven targets" srcset="https://substackcdn.com/image/fetch/$s_!WJdx!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda74bf6b-cd37-4c63-9ab9-ecbe8a3feb0e_1820x936.png 424w, https://substackcdn.com/image/fetch/$s_!WJdx!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda74bf6b-cd37-4c63-9ab9-ecbe8a3feb0e_1820x936.png 848w, https://substackcdn.com/image/fetch/$s_!WJdx!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda74bf6b-cd37-4c63-9ab9-ecbe8a3feb0e_1820x936.png 1272w, https://substackcdn.com/image/fetch/$s_!WJdx!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda74bf6b-cd37-4c63-9ab9-ecbe8a3feb0e_1820x936.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Figure 3.</strong> Register counts, where they apply: 229, 128, 152 and 221 across the four NVIDIA targets. <code>sm_100a</code> additionally allocates 128 columns of tensor memory, an address space the other six targets do not have.</figcaption></figure></div><p>We get four observations, in ascending order of how much they should bother you.</p><ol><li><p><strong>The accumulator moves.</strong> Registers on Ampere, fed from shared memory by a warpgroup instruction on Hopper, in tensor memory on Blackwell datacenter, a fifth address space that did not exist two years ago. The programmer wrote <code>tl.dot(a, b, acc)</code> in every case.</p></li><li><p><strong>Blackwell consumer does not use wgmma.</strong> It falls back to <code>mma.sync</code>, the same instruction Ampere uses, and allocates no tensor memory. We reached the same conclusion from the opposite direction in Issue 09 by disassembling cubins. Here it falls out of a Triton compile: <code>sm_120a</code> emits zero <code>wgmma</code> and zero <code>tcgen05</code>. The warpgroup instruction has no successor on that part. Only <code>mma.sync</code> spans Hopper and both Blackwells.</p></li><li><p><strong>The AMD path is a real path, and it is target-specialised too.</strong> Same 165-line TTIR, three CDNA generations, three different TTGIRs. gfx950 issues 24 <code>v_mfma</code> where gfx942 issues 48, because the instruction is wider. And gfx942 gets a layout that no other target in the entire compiler uses:</p><pre><code><span>// layout kinds present in TTGIR, per target</span>
gfx90a  blocked, dot_op, shared_memory, slice, swizzled_shared, amd_mfma
gfx942  blocked, dot_op, shared_memory, slice, swizzled_shared, amd_mfma,
        linear, <span>amd_rotating_shared</span>
gfx950  blocked, dot_op, shared_memory, slice, swizzled_shared, amd_mfma,
        linear</code></pre><p>The rotating shared layout cuts LDS writes from 54 to 26 on the same kernel. It exists for one hardware generation. This is the concrete form of what portability costs: the IR is shared, the language is shared, and the layouts are per-silicon.</p></li><li><p><strong>Register pressure is a function of pipelining, not of architecture.</strong> The figures above are from a build that pipelines. The same kernel compiled without the pipeline uses 226, 227, 244 and 234 registers on the four NVIDIA targets. On Hopper that is 227 against 128. Staging operands through shared memory buys back nearly half the register file. So pipelining is not simply a cost paid in shared memory, and the tradeoff is two-sided.</p></li></ol><p>Which brings us to the question I could not answer about my own kernel: what makes the compiler pipeline the loop in the first place?</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/p/exploring-how-triton-actually-compiles?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/p/exploring-how-triton-actually-compiles?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><div><hr></div><h2>An integer you never write decides your vector width</h2><p>The short answer is <code>tt.divisibility</code>, an attribute attached to the kernel&#8217;s pointer and integer arguments. You don&#8217;t write it. Triton&#8217;s JIT computes it at launch by looking at the runtime value of each argument and checking what it divides by. </p><p>If your tensor happens to be <strong>sixteen-byte aligned</strong>, your loads vectorize and your loop pipelines. If it does not, the same Python file compiles to a kernel that is correct, slower, and gives no indication that anything was declined.</p><p>In the first draft of this piece we wrote that the threshold was sixteen. <strong>That was wrong</strong>, or rather it was one point on a curve we had not measured. So we swept it. The structure is much better than &#8220;<em>sixteen or nothing</em>&#8221;, and it is exactly what the hardware would predict if you knew where to look.</p><ul><li><p><strong>The cliff is at four bytes.</strong> That is the minimum granularity of <code>cp.async</code>, which moves 4, 8 or 16 bytes per thread and nothing else. Below four, the compiler cannot express the copy as an asynchronous one, so it cannot pipeline, so <code>num_stages</code> becomes inert at every value. The threshold is not a heuristic. It is an instruction encoding.</p></li><li><p><strong>Between four and sixteen it is a ramp, not a switch.</strong> Vector width tracks alignment one for one: 4 bytes gives <code>[1,2]</code>, 8 gives <code>[1,4]</code>, 16 gives <code>[1,8]</code>. The count of <code>cp.async</code> instructions falls correspondingly, 70 to 38 to 22, because the same bytes move in fewer, wider transactions.</p></li><li><p><strong>Above sixteen it saturates.</strong> A divisibility of 32 buys nothing, because 16 bytes is the widest vector load the ISA has. The compiler stops asking for more alignment than it can use.</p></li><li><p><strong>Shared memory is identical for every pipelined width.</strong> 81,920 bytes at divisibility 4, 8, 16 and 32. Vector width changes how the bytes travel, not how many buffers exist. Those are two independent decisions the same attribute happens to gate.</p></li></ul><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!WzFw!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd806529c-a451-487f-aa5f-fe97b2a9e138_1820x863.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!WzFw!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd806529c-a451-487f-aa5f-fe97b2a9e138_1820x863.png 424w, https://substackcdn.com/image/fetch/$s_!WzFw!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd806529c-a451-487f-aa5f-fe97b2a9e138_1820x863.png 848w, https://substackcdn.com/image/fetch/$s_!WzFw!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd806529c-a451-487f-aa5f-fe97b2a9e138_1820x863.png 1272w, https://substackcdn.com/image/fetch/$s_!WzFw!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd806529c-a451-487f-aa5f-fe97b2a9e138_1820x863.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!WzFw!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd806529c-a451-487f-aa5f-fe97b2a9e138_1820x863.png" width="1456" height="690" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d806529c-a451-487f-aa5f-fe97b2a9e138_1820x863.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:690,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;divisibility cliff&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="divisibility cliff" title="divisibility cliff" srcset="https://substackcdn.com/image/fetch/$s_!WzFw!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd806529c-a451-487f-aa5f-fe97b2a9e138_1820x863.png 424w, https://substackcdn.com/image/fetch/$s_!WzFw!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd806529c-a451-487f-aa5f-fe97b2a9e138_1820x863.png 848w, https://substackcdn.com/image/fetch/$s_!WzFw!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd806529c-a451-487f-aa5f-fe97b2a9e138_1820x863.png 1272w, https://substackcdn.com/image/fetch/$s_!WzFw!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd806529c-a451-487f-aa5f-fe97b2a9e138_1820x863.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Figure 4.</strong> Vector width scales linearly with proven alignment from 4 bytes to 16 and then stops, because 16 bytes is the widest vector load in the ISA. Below 4 bytes the pipeliner cannot run at all, since <code>cp.async</code> has no 2-byte form. Shared memory is 81,920 bytes at every pipelined width: vector width and buffer depth are independent decisions the same attribute happens to gate.</figcaption></figure></div><p>Without the attribute, <code>num_stages</code> is inert. It is accepted, recorded in the kernel metadata, and ignored. There is no warning.</p><p>So <em>what does this mean for someone writing kernels? </em>It means the shape of an allocation upstream of your kernel is a performance parameter of your kernel, and the kernel author has no lever, because the lever is not in the kernel. </p><p>A <strong>tensor produced by a slice</strong>, a view at an odd offset, or a buffer carved from a pool at an unaligned boundary takes the slow path silently. That is what had happened to me. The other person&#8217;s kernel was not better. Their tensors were.</p><p><em>Is this a bug?</em> We do not think so, and it is worth being clear about why. Specialize on what you can prove, fall back to what is always safe, is the correct design for a <strong>JIT that has to be correct</strong> on every input. The complaint is narrower: the fallback is unobservable from the language. </p><p>There is no warning, no metadata field the user reads, no <code>assert_aligned</code> to make the assumption explicit and fail loudly when it breaks. The compiler knows it declined. It just does not say so.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>What does a stage actually cost?</h2><p>The folklore answer is that stages multiply your operand footprint, so shared memory goes as <code>num_stages &#215; (A + B)</code>. That is not what happens inthe real world. We swept five block shapes with deliberately asymmetric operands to decompose it.</p><p>Every measured value matches one closed form, across all fifteen combinations:</p><pre><code>shared = max( num_stages &#215; A_tile + B_tile , C_tile )</code></pre><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!HDL4!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa97673e7-05d7-4b2f-bb19-d7817d0c8cf6_1820x894.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!HDL4!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa97673e7-05d7-4b2f-bb19-d7817d0c8cf6_1820x894.png 424w, https://substackcdn.com/image/fetch/$s_!HDL4!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa97673e7-05d7-4b2f-bb19-d7817d0c8cf6_1820x894.png 848w, https://substackcdn.com/image/fetch/$s_!HDL4!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa97673e7-05d7-4b2f-bb19-d7817d0c8cf6_1820x894.png 1272w, https://substackcdn.com/image/fetch/$s_!HDL4!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa97673e7-05d7-4b2f-bb19-d7817d0c8cf6_1820x894.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!HDL4!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa97673e7-05d7-4b2f-bb19-d7817d0c8cf6_1820x894.png" width="1456" height="715" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a97673e7-05d7-4b2f-bb19-d7817d0c8cf6_1820x894.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:715,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;buffering rule&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="buffering rule" title="buffering rule" srcset="https://substackcdn.com/image/fetch/$s_!HDL4!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa97673e7-05d7-4b2f-bb19-d7817d0c8cf6_1820x894.png 424w, https://substackcdn.com/image/fetch/$s_!HDL4!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa97673e7-05d7-4b2f-bb19-d7817d0c8cf6_1820x894.png 848w, https://substackcdn.com/image/fetch/$s_!HDL4!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa97673e7-05d7-4b2f-bb19-d7817d0c8cf6_1820x894.png 1272w, https://substackcdn.com/image/fetch/$s_!HDL4!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa97673e7-05d7-4b2f-bb19-d7817d0c8cf6_1820x894.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Figure 5.</strong> Fifteen shape and stage combinations, every one matching the closed form exactly. At <code>BLOCK_K = 32</code> the epilogue scratch dominates the loop allocation, so two stage counts produce byte-identical kernels.</figcaption></figure></div><p>The marginal cost of a stage is the A tile alone. B is not multi-buffered. <strong>The TTGIR says so directly:</strong></p><pre><code><span>// 128x256x64, num_stages = 4, sm_90a</span>
%a = ttg.local_alloc : () -&gt; !ttg.memdesc&lt;<span>4x</span>128x64xf16, #shared, #smem, mutable&gt;
%b = ttg.local_alloc %b_reg : (...) -&gt; !ttg.memdesc&lt;64x256xf16, #shared, #smem&gt;</code></pre><p><strong>A is allocated once</strong>, outside the loop, four deep, mutable. B is allocated inside the loop from a register tensor, one deep. So A travels global to shared asynchronously and B travels global to register to shared synchronously, in a Hopper kernel where the matrix instruction reads both operands from shared memory.</p><p>We do not have a confident explanation. Both loads carry the <strong>same divisibility attributes</strong>, both index expressions have the same shape, and both operands feed the same <code>warp_group_dot</code>. </p><p>It is stable across shapes and stage counts, which argues against an accident of one configuration, but we are reporting it as measured rather than as a rule of the pipeliner. If you know why, we would like to hear it, and we will publish the correction.</p><p>The <code>max</code> term matters more than it looks. At <code>BLOCK_K = 32</code> the epilogue scratch for storing C dominates the loop allocation entirely, so <code>num_stages</code> of 2 and 3 produce byte identical kernels: 32,768 either way. </p><p><strong>Anyone autotuning that shape</strong> is paying compile time to distinguish two configurations that are the same kernel.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>What happens below four warps?</h2><p>This one we found by accident, sweeping the other option people sweep blindly, and it is the second answer that is not in the file.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!H-XI!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F691da6f0-f0e6-4e17-bd17-f9525a1ba9c1_1820x904.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!H-XI!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F691da6f0-f0e6-4e17-bd17-f9525a1ba9c1_1820x904.png 424w, https://substackcdn.com/image/fetch/$s_!H-XI!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F691da6f0-f0e6-4e17-bd17-f9525a1ba9c1_1820x904.png 848w, https://substackcdn.com/image/fetch/$s_!H-XI!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F691da6f0-f0e6-4e17-bd17-f9525a1ba9c1_1820x904.png 1272w, https://substackcdn.com/image/fetch/$s_!H-XI!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F691da6f0-f0e6-4e17-bd17-f9525a1ba9c1_1820x904.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!H-XI!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F691da6f0-f0e6-4e17-bd17-f9525a1ba9c1_1820x904.png" width="1456" height="723" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/691da6f0-f0e6-4e17-bd17-f9525a1ba9c1_1820x904.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:723,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;num warps instruction&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="num warps instruction" title="num warps instruction" srcset="https://substackcdn.com/image/fetch/$s_!H-XI!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F691da6f0-f0e6-4e17-bd17-f9525a1ba9c1_1820x904.png 424w, https://substackcdn.com/image/fetch/$s_!H-XI!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F691da6f0-f0e6-4e17-bd17-f9525a1ba9c1_1820x904.png 848w, https://substackcdn.com/image/fetch/$s_!H-XI!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F691da6f0-f0e6-4e17-bd17-f9525a1ba9c1_1820x904.png 1272w, https://substackcdn.com/image/fetch/$s_!H-XI!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F691da6f0-f0e6-4e17-bd17-f9525a1ba9c1_1820x904.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Figure 6.</strong> <code>warpsPerCTA</code> and <code>instrShape</code> also move: <code>[1,1]</code> and <code>[16,8]</code> at one warp, <code>[8,1]</code> and <code>[16,128,16]</code> at eight, <code>[8,2]</code> and <code>[16,64,16]</code> at sixteen. Three tensor core configurations across five values of one integer.</figcaption></figure></div><p>At <code>num_warps</code> of 1 or 2, Triton does not emit <code>wgmma</code>. It emits MMA version 2, which is <code>mma.sync</code> with an <code>instrShape</code> of <code>[16, 8]</code>, the Ampere-era instruction. The <strong>reason is structural and unavoidable</strong>: <code>wgmma</code> is a warpgroup instruction and a warpgroup is four warps. </p><p>Below four warps there is <strong>no warpgroup</strong>, so there is no warpgroup instruction, so you are on the previous generation of tensor core path on brand new silicon.</p><p>On Hopper, <code>num_warps</code> is not a parallelism knob: it&#8217;s an instruction selection knob with a hard threshold at four, and nothing in the language says so.</p><p>The register column is the other half of it. Total register file consumption, which is registers per thread times threads per block, is 2,048 at one and two warps and 32,640 at four. A <strong>sixteen-fold jump</strong> for a doubling of the warp count, because crossing the threshold changes which instruction runs and therefore how much state has to be live. </p><p>And at exactly four warps the compiler lands on <strong>255 registers per thread</strong>, which is the architectural ceiling, with zero spills. It is sitting on the edge. Eight warps halves it to 128.</p><p>Then at sixteen warps the <code>instrShape</code> narrows from <code>[16, 128, 16]</code> to <code>[16, 64, 16]</code> and <code>warpsPerCTA</code> becomes <code>[8, 2]</code>: the compiler splits the N dimension across two warp columns and each warpgroup gets a narrower instruction. <strong>Three different tensor core</strong> configurations across five values of one integer.</p><p>So if you autotune <code>num_warps</code> over <code>[1, 2, 4, 8]</code>, which is a common default, half your search space is not testing parallelism. It is testing a different instruction. <em>Does that matter for the result?</em> Not necessarily, since the autotuner measures wall clock and does not care why one config is faster. </p><p>It<strong> matters for the interpretation</strong>, and it matters for anyone reasoning about the space rather than searching it exhaustively, which is everyone with a compile budget.</p><div><hr></div><h2>If it is not arithmetic, what is it?</h2><p><strong>Triton 3.7.1 registers 82 distinct passes </strong>across its core and its two backends. We classified every one by hand, into four buckets, and shipped the classification in the repository so it can be argued with line by line. </p><p>The headline number depends entirely on the classification, so it should be possible to check the classification.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!83Yr!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4285fd51-3f61-49b5-bebf-53219c499565_1820x858.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!83Yr!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4285fd51-3f61-49b5-bebf-53219c499565_1820x858.png 424w, https://substackcdn.com/image/fetch/$s_!83Yr!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4285fd51-3f61-49b5-bebf-53219c499565_1820x858.png 848w, https://substackcdn.com/image/fetch/$s_!83Yr!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4285fd51-3f61-49b5-bebf-53219c499565_1820x858.png 1272w, https://substackcdn.com/image/fetch/$s_!83Yr!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4285fd51-3f61-49b5-bebf-53219c499565_1820x858.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!83Yr!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4285fd51-3f61-49b5-bebf-53219c499565_1820x858.png" width="1456" height="686" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/4285fd51-3f61-49b5-bebf-53219c499565_1820x858.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:686,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;pass census&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="pass census" title="pass census" srcset="https://substackcdn.com/image/fetch/$s_!83Yr!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4285fd51-3f61-49b5-bebf-53219c499565_1820x858.png 424w, https://substackcdn.com/image/fetch/$s_!83Yr!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4285fd51-3f61-49b5-bebf-53219c499565_1820x858.png 848w, https://substackcdn.com/image/fetch/$s_!83Yr!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4285fd51-3f61-49b5-bebf-53219c499565_1820x858.png 1272w, https://substackcdn.com/image/fetch/$s_!83Yr!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4285fd51-3f61-49b5-bebf-53219c499565_1820x858.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Figure 7.</strong> Every pass reachable from the NVIDIA or AMD backend, classified by hand. The full name lists ship in <code>m2_pipeline.py</code> so the classification can be disputed line by line, which matters because the headline number depends on it.</figcaption></figure></div><p>Strip out the borrowed machinery, which tells you nothing about Triton because every MLIR project has it, and 60 passes remain. Of those, 53 are about <strong>data placement </strong>and movement. Five are about arithmetic. Two are checks.</p><p>Triton is a layout compiler with an arithmetic language attached, not an arithmetic compiler with a layout system attached.</p><p>This is consistent with the project&#8217;s own bug data. The linear layouts paper from the Triton team reports that that <strong>12 percent of issues</strong> filed against the Triton repository are layout related, and that the pre-linear-layout system suffered a <strong>quadratic blow-up</strong> in the number of layout-to-layout conversions that had to be implemented by hand.</p><p>The architecture specific branching goes further than lowering. The TTGIR pipeline in the NVIDIA backend is a<strong> three-way conditional on compute capability</strong>, and the branches are not the same length. The Ampere and Hopper branch makes 9 pass registration calls. </p><p>The <strong>Blackwell branch</strong> makes 14, adding accumulator initialization, tensor memory hoisting twice, promotion of the left operand into tensor memory, automatic warp specialization, partition warp optimization, and tensor memory token removal. </p><p>A &#8220;<em>portable</em>&#8221; compiler with a per-architecture pipeline is portable in the sense that it will produce working code everywhere, not in the sense that it does the same thing everywhere.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Eleven bits describe a two thousand element tile</h2><p>The reason the placement machinery is tractable at all is a change that landed in Triton over the <strong>last two years</strong> and that we think is the most interesting idea in the project. </p><p>A layout is a map from hardware coordinates, which is to say the triple of register index, lane index and warp index, to positions in a logical tensor.</p><p>The old way to represent such a map was a family of hand-written attributes, one per pattern, plus a conversion routine for each ordered pair of patterns. That is <strong>quadratic in the number of patterns</strong> and it is where the bugs lived.</p><p>The new representation treats the hardware coordinate as a vector of bits and the layout as a linear map over the field with two elements. Because it is linear, it is a matrix. Because it is a matrix, <strong>composition is multiplication</strong> and a conversion between two layouts is one composed with the inverse of the other. One function instead of a table.</p><p>The bindings are exposed to Python, so this is not a description, it is something you can execute:</p><pre><code><span>from</span> triton._C.libtriton.linear_layout <span>import</span> LinearLayout

<span># the #blocked layout the compiler chose for the A operand in section 04:</span>
<span># 8 elements per thread on the fast axis, lanes split 4 by 8, 8 warps</span>
reg  = LinearLayout.identity_1d(8, <span>&#8220;register&#8221;</span>, <span>&#8220;dim1&#8221;</span>)
lane = LinearLayout.identity_1d(8, <span>&#8220;lane&#8221;</span>,     <span>&#8220;dim1&#8221;</span>)
slow = LinearLayout.identity_1d(4, <span>&#8220;lane&#8221;</span>,     <span>&#8220;dim0&#8221;</span>)
warp = LinearLayout.identity_1d(8, <span>&#8220;warp&#8221;</span>,     <span>&#8220;dim0&#8221;</span>)
ll = (reg * lane) * (slow * warp)

ll.is_surjective(), ll.is_injective()   <span># True, True</span>
sum(len(b) <span>for</span> _, b <span>in</span> ll.bases)     <span># 11</span></code></pre><p>Eleven basis vectors, each one telling you where a single bit of the hardware index lands in the tensor:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!oedl!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F643b4ada-7800-4f8c-a46f-ccb411970160_1820x925.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!oedl!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F643b4ada-7800-4f8c-a46f-ccb411970160_1820x925.png 424w, https://substackcdn.com/image/fetch/$s_!oedl!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F643b4ada-7800-4f8c-a46f-ccb411970160_1820x925.png 848w, https://substackcdn.com/image/fetch/$s_!oedl!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F643b4ada-7800-4f8c-a46f-ccb411970160_1820x925.png 1272w, https://substackcdn.com/image/fetch/$s_!oedl!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F643b4ada-7800-4f8c-a46f-ccb411970160_1820x925.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!oedl!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F643b4ada-7800-4f8c-a46f-ccb411970160_1820x925.png" width="1456" height="740" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/643b4ada-7800-4f8c-a46f-ccb411970160_1820x925.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:740,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;layout bases&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="layout bases" title="layout bases" srcset="https://substackcdn.com/image/fetch/$s_!oedl!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F643b4ada-7800-4f8c-a46f-ccb411970160_1820x925.png 424w, https://substackcdn.com/image/fetch/$s_!oedl!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F643b4ada-7800-4f8c-a46f-ccb411970160_1820x925.png 848w, https://substackcdn.com/image/fetch/$s_!oedl!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F643b4ada-7800-4f8c-a46f-ccb411970160_1820x925.png 1272w, https://substackcdn.com/image/fetch/$s_!oedl!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F643b4ada-7800-4f8c-a46f-ccb411970160_1820x925.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Figure 8.</strong> The blocked layout Triton selects for the A operand once the alignment hint is present. A layout over 2^k hardware slots needs exactly k basis vectors, so the representation is logarithmic in the size of the thing it describes where a table would be linear.</figcaption></figure></div><p>The scaling is the point.<strong> A layout over 2<sup>k</sup> hardware </strong>slots needs exactly k basis vectors, which we checked from 16 elements to 65,536: 4 vectors and 16 vectors respectively. The representation is logarithmic in the size of the thing it describes, and a table would be linear. A <em>65,536 element</em> tile is sixteen integers.</p><p>Conversion falls out of the same algebra. Take an identity assignment of 32 lanes and a <strong>bit-reversed assignment </strong>of the same 32 lanes, which is the shape of a swizzle, and ask for the map between them:</p><pre><code>ident = LinearLayout.identity_1d(32, <span>&#8220;lane&#8221;</span>, <span>&#8220;dim0&#8221;</span>)
rev   = LinearLayout.from_bases([(<span>&#8220;lane&#8221;</span>, [[16],[8],[4],[2],[1]])], [<span>&#8220;dim0&#8221;</span>], [32])
conv  = ident.invert_and_compose(rev)
<span># bases: lane[0]-&gt;16, lane[1]-&gt;8, lane[2]-&gt;4, lane[3]-&gt;2, lane[4]-&gt;1</span></code></pre><p>No case analysis, no table entry, no new pass. This is why the placement machinery can be <strong>53 pasasses</strong> instead of 53 passes plus a combinatorial explosion of conversion special cases, and it is the piece of Triton that we would expect to outlive the language it currently serves.</p><p>The limitation is stated openly by the authors and shows up immediately in practice: <strong>the algebra is over powers of two.</strong> Non power of two shapes have to be padded and masked, and operations like slicing and flipping are affine rather than linear, <code>y = Ax + b</code> rather than <code>y = Ax</code>, so they sit outside the framework as originally formulated.</p><p>It is worth asking how general this is, because if it is general it outlives Triton. The signs are that it is. </p><p>Work published in January 2026 gives a categorical account of CuTe layouts and observes that <strong>Triton&#8217;s F<sub>2</sub> layouts</strong> compose naturally with swizzles, which CuTe layouts generally cannot express, while being less expressive in the other direction because of the power-of-two constraint and the inability to scale by a non-power-of-two. </p><p>A separate 2026 proposal, Axe, argues for a <strong>single unified layout abstraction </strong>across ML compilers and cites the linear layouts work as prior art. Two independent groups converging on &#8220;layout is the object worth formalising&#8221; is the same conclusion the pass census reaches by counting.</p><div><hr></div><h2>Why does one wheel need two assemblers?</h2><p>The wheel ships two copies of <code>ptxas</code>, at CUDA 12.8.93 and CUDA 13.1.80, and picks between them with one line:</p><pre><code><span>def</span> get_ptxas(arch: int) -&gt; knobs.NvidiaTool:
    <span>return</span> knobs.nvidia.ptxas_blackwell <span>if</span> arch &gt;= 100 <span>else</span> knobs.nvidia.ptxas</code></pre><p>The reason is that the <strong>two target</strong> sets are disjoint at both ends. CUDA 12.8 still accepts ten targets that 13.1 dropped, including all of Maxwell, Pascal and Volta plus <code>sm_101</code> and <code>sm_101a</code>.</p><p>CUDA 13.1 adds twelve that 12.8 has never heard of, including <code>sm_88</code>, <code>sm_103</code>, <code>sm_110</code>, <code>sm_121</code> and the entire family-compatible class. Eleven targets are common to both. </p><p>Supporting the range Triton claims to support requires carrying two closed assemblers, and there is <strong>no version of CUDA that covers it.</strong></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!8zEi!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61aecfa7-8af3-43ff-8e83-1d0cd13eaf70_1820x764.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!8zEi!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61aecfa7-8af3-43ff-8e83-1d0cd13eaf70_1820x764.png 424w, https://substackcdn.com/image/fetch/$s_!8zEi!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61aecfa7-8af3-43ff-8e83-1d0cd13eaf70_1820x764.png 848w, https://substackcdn.com/image/fetch/$s_!8zEi!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61aecfa7-8af3-43ff-8e83-1d0cd13eaf70_1820x764.png 1272w, https://substackcdn.com/image/fetch/$s_!8zEi!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61aecfa7-8af3-43ff-8e83-1d0cd13eaf70_1820x764.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!8zEi!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61aecfa7-8af3-43ff-8e83-1d0cd13eaf70_1820x764.png" width="1456" height="611" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/61aecfa7-8af3-43ff-8e83-1d0cd13eaf70_1820x764.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:611,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;assembler targets&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="assembler targets" title="assembler targets" srcset="https://substackcdn.com/image/fetch/$s_!8zEi!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61aecfa7-8af3-43ff-8e83-1d0cd13eaf70_1820x764.png 424w, https://substackcdn.com/image/fetch/$s_!8zEi!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61aecfa7-8af3-43ff-8e83-1d0cd13eaf70_1820x764.png 848w, https://substackcdn.com/image/fetch/$s_!8zEi!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61aecfa7-8af3-43ff-8e83-1d0cd13eaf70_1820x764.png 1272w, https://substackcdn.com/image/fetch/$s_!8zEi!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61aecfa7-8af3-43ff-8e83-1d0cd13eaf70_1820x764.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Figure 9.</strong> The five family-compatible targets CUDA 13 introduced trade some architecture-specific capability for forward compatibility within a family. Triton ships the assembler that accepts them and asks for none of them.</figcaption></figure></div><p>That last row is the one worth dwelling on. Here is how Triton picks the target string it hands to <code>ptxas</code>:</p><pre><code><span>def</span> sm_arch_from_capability(capability: int):
    <span># TODO: Handle non-&#8221;a&#8221; sms</span>
    suffix = <span>&#8220;a&#8221;</span> <span>if</span> capability &gt;= 90 <span>else</span> <span>&#8220;&#8221;</span>
    <span>return</span> f<span>&#8220;sm_{capability}{suffix}&#8221;</span></code></pre><p>Unconditionally. Every kernel Triton compiles for Hopper or later goes to an <code>a</code> suffixed target. The <code>a</code> targets are the ones that expose architecture-specific instructions and, in exchange, give up <strong>PTX forward compatibility</strong>: a cubin built for <code>sm_90a</code> will not load on a later architecture, and the embedded PTX will not JIT forward either. </p><p>We argued in the compiler moat piece that PTX forward compatibility does not cover the instructions that matter. Here is the same claim from the other side, in <strong>four lines of Python</strong>, with a live TODO on top of it.</p><p>CUDA 13 shipped the middle option. The <code>f</code> targets keep forward compatibility within an <strong>architecture family </strong>while still admitting most of the family&#8217;s instructions. Triton ships the assembler that accepts them and asks for none of them. </p><p>Whether that is a deliberate choice about capability or simply work nobody has done, the effect on users is the same: your Triton kernels are locked to the <strong>exact architecture</strong> they were built for, and the lock is a hardcoded string with a comment saying somebody should look at it.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>What is the compiler deciding for you, exactly?</h2><p>Inside the same wheel, sharing the same compiler and the same IR, is a second language. <strong>Gluon is a lower-level DSL</strong> that hands the programmer the decisions Triton makes automatically. </p><p>Its own tutorial says the quiet part plainly: the Triton compiler generates efficient code across a<strong> wide range of kernels</strong> but can be beaten by hand-tuned low-level code, and when that happens there is little the user can do.</p><p>That gives us a way to size the automation without arguing about it. Whatever Gluon exposes and Triton does not is, by construction, a decision the Triton compiler is making on your behalf. So we counted.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!kzjr!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89d1f233-51ad-418f-8ef9-5d929c3d16d6_1820x1045.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!kzjr!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89d1f233-51ad-418f-8ef9-5d929c3d16d6_1820x1045.png 424w, https://substackcdn.com/image/fetch/$s_!kzjr!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89d1f233-51ad-418f-8ef9-5d929c3d16d6_1820x1045.png 848w, https://substackcdn.com/image/fetch/$s_!kzjr!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89d1f233-51ad-418f-8ef9-5d929c3d16d6_1820x1045.png 1272w, https://substackcdn.com/image/fetch/$s_!kzjr!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89d1f233-51ad-418f-8ef9-5d929c3d16d6_1820x1045.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!kzjr!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89d1f233-51ad-418f-8ef9-5d929c3d16d6_1820x1045.png" width="1456" height="836" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/89d1f233-51ad-418f-8ef9-5d929c3d16d6_1820x1045.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:836,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;gluon delta&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="gluon delta" title="gluon delta" srcset="https://substackcdn.com/image/fetch/$s_!kzjr!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89d1f233-51ad-418f-8ef9-5d929c3d16d6_1820x1045.png 424w, https://substackcdn.com/image/fetch/$s_!kzjr!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89d1f233-51ad-418f-8ef9-5d929c3d16d6_1820x1045.png 848w, https://substackcdn.com/image/fetch/$s_!kzjr!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89d1f233-51ad-418f-8ef9-5d929c3d16d6_1820x1045.png 1272w, https://substackcdn.com/image/fetch/$s_!kzjr!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89d1f233-51ad-418f-8ef9-5d929c3d16d6_1820x1045.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Figure 10.</strong> Gluon ships inside the same wheel and compiles through the same passes to the same PTX. What it removes is the inference, which makes this list a measured inventory of the automation rather than an argument about it.</figcaption></figure></div><p><code>triton.language</code> exposes 122 public symbols. <code>gluon.language</code> exposes 147, sharing 76 with Triton and adding 71 of its own. Strip the Python plumbing and roughly forty are substantive. </p><p>Those forty are a precise inventory of the automation: choose a layout, place a tile in <strong>shared or tensor memory</strong>, pick a matrix instruction, insert a fence, decide whether to specialize warps, and check for bank conflicts.</p><p>A language that grows a second, lower-level language inside itself has told you where its ceiling is. The interesting part is that <strong>Gluon reuses the entire compiler</strong>, so the ceiling is in the inference, not in the representation.</p><p>That last distinction is what separates Gluon from the usual story of an abstraction failing.<strong> Triton&#8217;s IR </strong>can express everything Gluon can express, and Gluon compiles through the same passes to the same PTX. </p><p>What Gluon removes is the inference step: the guessing about which layout, how many buffers, which instruction. What is not always good enough is the search over it, which is a much more tractable problem and a much better place for a moat to be than in a representation.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Is the ceiling in the representation or in the search?</h2><p>Triton&#8217;s ceiling is in the inference, not the representation, on the grounds that <em>Gluon reuses the entire compiler</em> and only removes the guessing. That is an argument from architecture. </p><blockquote><p><em>Is there independent evidence?</em></p></blockquote><p>There is, and it is unusually direct. In December 2025 a group from Stanford, Toronto and NVIDIA published <strong>Twill</strong>, which formulates software pipelining and warp specialization as a single joint optimization problem and solves it with an off-the-shelf constraint solver rather than with heuristics. </p><p>Their framing of the problem is worth reading against the measurements above: they describe the current state as a mix of <strong>brittle compilation heuristics and fallible human intuition</strong>, with little insight into the space of solutions, and they point at the year that elapsed between Hopper shipping and FlashAttention-3 arriving with a hand-designed schedule for it.</p><p>Two critical findings in that paper bear directly on this one.</p><ol><li><p><strong>The first concerns what Triton&#8217;s pipeliner is. </strong>The authors report that Triton&#8217;s Blackwell backend heuristically applies the FlashAttention-3 pipelining strategy. Not derives, applies. A schedule that a human designed for one kernel on one architecture is baked in as the backend&#8217;s default plan, which is a reasonable engineering decision and also exactly the thing sections 05 through 07 keep running into: the compiler is not searching, it is pattern matching against a small number of known-good plans, and whether you land on one is gated by conditions the language does not surface.</p></li><li><p><strong>The second is harsher, and we quote its shape carefully because it is someone else&#8217;s measurement and not ours.</strong> When the Twill authors tried to have Triton compile the schedules their solver found, they report that Triton made incorrect decisions in memory allocation, layout conversion and synchronization placement, and that the result was either a compile failure or poorly performing code. Their workaround was to hand-translate their pipelined IR into CUDA C++. On Blackwell attention their solver completed in 19 seconds and found a strategy that beat Triton substantially and was competitive with cuDNN and FlashAttention-4, and the strategy it discovered was the same one the FlashAttention-4 authors had arrived at by hand.</p></li></ol><p>A constraint solver rediscovered a hand-tuned schedule in nineteen seconds. The obstacle to using it was not the schedule. It was getting the compiler to lower it.</p><p>This is corroboration and <strong>it is also a caveat on our own framing</strong>. We said the ceiling is in the inference rather than the representation. Twill&#8217;s experience says the lowering has gaps too: a schedule that is expressible in principle was not compilable in practice. Those are different problems with different fixes. </p><p>The <strong>optimistic reading</strong>, which we lean toward but hold loosely, is that lowering gaps are the kind of thing that gets closed by ordinary engineering, whereas a <em>representation that cannot express the schedule at all would be a structural problem.</em></p><p>Gluon existing, and shipping precisely the primitives that Twill needed to place by hand, is weak evidence for the optimistic reading.</p><p>There is a second literature pointing the same way from further out. He and Yoneki&#8217;s<strong> CuAsmRL, at CGO 2025</strong>, intercepts the cubin Triton produces, disassembles it, and has a reinforcement learning agent mutate the SASS schedule that <code>ptxas -O3</code> already optimized. </p><p>They report up to 26 percent improvement and 9 percent on average, a geometric mean of 1.09 times, transparently, on <strong>kernels Triton had already compiled</strong> as well as it knows how.</p><p>Two caveats we would want if someone quoted this at us. Their evaluation used Triton 2.1.0 and <code>ptxas</code> 12.2 on an A100, which is several compiler generations behind everything measured in this piece, so the number should not be read as a current gap. </p><p>And the <strong>optimization happens below PTX</strong>, which is precisely where Triton has no visibility at all: it is a measurement of what the whole stack leaves on the table, not of what Triton specifically gets wrong. What survives both caveats is the direction. There was single-digit percent lying underneath an <code>-O3</code> schedule, found by search.</p><p>And the automated-kernel-generation literature, <strong>TritonBench</strong> and AutoTriton among others, consistently finds generating good Triton harder than generating good CUDA. </p><p>That is a strange result for a <strong>higher-level language</strong> until you notice what is actually being generated: not a program, but a set of hints to a heuristic, several of which are the invisible switches of sections 05 and 07.</p><div><hr></div><h2>Six minutes of CPU before a single kernel runs</h2><p>Because the compiler is a host program, the cost of an autotune space can be priced without a GPU. </p><p>We took a <strong>conventional matmul space,</strong> three block M by three block N by three block K by three stage counts by two warp counts, which is 162 configurations, sampled 24 of them and compiled each.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!CY5l!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82ab1769-3bb0-4574-8532-c9a10038812f_1820x754.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!CY5l!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82ab1769-3bb0-4574-8532-c9a10038812f_1820x754.png 424w, https://substackcdn.com/image/fetch/$s_!CY5l!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82ab1769-3bb0-4574-8532-c9a10038812f_1820x754.png 848w, https://substackcdn.com/image/fetch/$s_!CY5l!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82ab1769-3bb0-4574-8532-c9a10038812f_1820x754.png 1272w, https://substackcdn.com/image/fetch/$s_!CY5l!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82ab1769-3bb0-4574-8532-c9a10038812f_1820x754.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!CY5l!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82ab1769-3bb0-4574-8532-c9a10038812f_1820x754.png" width="1456" height="603" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/82ab1769-3bb0-4574-8532-c9a10038812f_1820x754.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:603,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;autotune cost&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="autotune cost" title="autotune cost" srcset="https://substackcdn.com/image/fetch/$s_!CY5l!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82ab1769-3bb0-4574-8532-c9a10038812f_1820x754.png 424w, https://substackcdn.com/image/fetch/$s_!CY5l!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82ab1769-3bb0-4574-8532-c9a10038812f_1820x754.png 848w, https://substackcdn.com/image/fetch/$s_!CY5l!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82ab1769-3bb0-4574-8532-c9a10038812f_1820x754.png 1272w, https://substackcdn.com/image/fetch/$s_!CY5l!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82ab1769-3bb0-4574-8532-c9a10038812f_1820x754.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Figure 11.</strong> Shared memory across the sampled space ranges from 20,480 to 344,064 bytes. Two of the 24 exceed the 232,448 byte limit an H100 gives a single block: they compile successfully, cost their full compile time, and fail at launch.</figcaption></figure></div><p>Two things follow. The first is that autotuning has a <strong>substantial fixed cost </strong>that is paid in host CPU and is invisible in any GPU-side measurement. Six minutes for one operator, on one shape, on one architecture, before measuring anything. </p><p>A serving stack with a <strong>dozen tuned operators</strong> and a per-architecture rebuild is spending real time on this, and it is time that does not show up in a kernel benchmark.</p><p>The second is that a meaningful fraction of the space is not launchable. Two of our 24 samples committed <strong>more than 232,448 bytes of shared memory</strong>, which is above what an H100 will give a single block. Those configurations compile successfully, cost their full compile time, and fail at launch. </p><p>Shared memory is known to the compiler at the end of compilation, so this is not information that has to be discovered on hardware. It could be a pre-filter. It is not.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Is 80 percent the right way to read a range?</h2><p>The most useful recent number on Triton&#8217;s absolute performance comes from Yadav, Zhao and Kumar at Wisconsin-Milwaukee and Illinois Tech, who ran GEMM, <strong>fused multi-head attention and end-to-end LLM inference </strong>on an H100 NVL, a B200 and an RTX PRO 6000 Blackwell Server Edition, in BF16 and FP16. </p><p>Triton sustains 62 to 101 percent of cuBLAS across all three with no architecture-specific tuning. NVIDIA&#8217;s own CuTile reaches 52 to 79 percent of <strong>cuBLAS on GEMM in 22 lines of Python against 123 for WMMA</strong>, and on B200 its attention kernel hits 1,007 TFLOP/s, beating FlashAttention-2 by 2.5 times in 60 lines. </p><p>On the RTX PRO 6000 the same<strong> CuTile attention kernel gets 53 percent of FlashAttention-2.</strong> An earlier result from the Triton-distributed work puts Triton GEMM at roughly 95 percent of cuBLAS and CUTLASS on H800.</p><p>One caveat, which the authors raise themselves and which we would have raised anyway: their H100 ran PyTorch 2.7.1 with CUDA 12.6 while <strong>both Blackwell machines</strong> ran PyTorch 2.8.0 with CUDA 12.8. They flag it as a possible confound in cross-GPU comparison. </p><p>It does not affect the within-GPU comparisons between Triton, CuTile and cuBLAS, which is what we are using the paper for, but anyone quoting the 62 to 101 band as a clean cross-architecture result should know it is not one.</p><p>The temptation is to average that band and call Triton an eighty percent solution. We think the width of the band is the finding, not its centre. A compiler that ranges from 62 to 101 percent depending on shape and architecture is not delivering eighty percent of peak. It is delivering<strong> peak on some inputs and losing a third on others</strong>, and the variance is not something the programmer can see from the source.</p><p>Everything in the preceding sections is a mechanism for that variance, and this is the part where the <strong>compile-time census </strong>earns its keep. The divisibility hint moves the vector width by a factor of eight and gates pipelining altogether. <code>num_warps</code> below four takes you off the warpgroup instruction entirely. </p><p>The buffering rule means some stage counts are byte identical kernels while others triple the footprint. The <code>max</code> term means the epilogue can dominate a <strong>small-K shape</strong>. The architecture branch means Blackwell runs five more passes than Hopper, and one AMD generation gets a layout no other target has.</p><p>None of these are visible at the call site. A band from 62 to 101 percent is what you would expect from a compiler whose output is <strong>controlled by four or five invisible switches,</strong> some of which are set by your allocator rather than by you.</p><div><hr></div><h2>Where does the seam sit after Tile IR?</h2><p>In January 2026 NVIDIA published a Triton backend that emits <strong>CUDA Tile IR instead of PTX</strong>, as an incubator repository under the <code>triton-lang</code> organisation, enabled with <code>ENABLE_TILE=1</code>, requiring CUDA 13.1 and Blackwell. Helion has a backend for it too.</p><p>Read against the rest of this piece, that is not a portability feature. It is a proposal to move the boundary. Today the <strong>boundary between Triton and NVIDIA sits at PTX</strong>, and everything in sections 04 through 07, the layouts, the buffering, the pipelining, the placement passes, happens above it, inside Triton. </p><p>Tile IR sits above PTX and its type system already encodes tile semantics. A Triton that lowers to Tile IR hands a large part of its placement work to a closed lowering compiler, which is the exact machinery this piece has spent thirty pages measuring.</p><p>We wrote in July that Tile IR was the new PTX: publish the interface, keep the lowering, one level higher than before, as a response to <strong>Triton&#8217;s position inside PyTorch. </strong>A first-party Triton backend for it, six months later, is consistent with that reading. </p><p>NVIDIA&#8217;s own documentation for the backend notes that Tile IR in CUDA 13.1 does not support <code>num_warps</code>, replacing it with an <code>occupancy</code> attribute, and that tensor-of-pointer patterns, which is the ordinary way people write Triton, perform poorly and should be rewritten to use the TMA descriptor APIs. </p><p>Both are the interface asserting itself over the language.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Where we might be wrong</h2><p><strong>The B operand asymmetry might be our kernel, not the pipeliner.</strong> We observe that only A is multi-buffered, across five shapes and three stage counts. We do not know why, and we did not test enough distinct kernels to separate a property of the pipeliner from a property of our index expressions. If the modulo we use for bounds wrapping on the N axis defeats the analysis for B specifically, the closed form in section 06 is a fact about our kernel.</p><p><strong>Everything here is compile-time observation presented next to performance claims.</strong> We measured what the compiler emits, not what a GPU does with it. It is possible, though we think unlikely given that the unpipelined kernel has no asynchronous copies at all, that the gap is smaller in wall clock than in shared memory. We have no hardware and we are not pretending otherwise. This is the single largest weakness of the piece and the obvious thing to fix: the same harness with a device attached would settle it.</p><p><strong>The pass classification is a judgement call and the headline moves with it.</strong> We put <code>add_fuse_nested_loops</code> and <code>add_triton_licm</code> in the borrowed bucket even though both exist in Triton partly to enable pipelining, which biases 88.3 percent downward. Moving those plus two or three similar calls pushes it past 90. Someone arguing the other way could reclassify <code>add_accelerate_matmul</code> and <code>add_optimize_dot_operands</code> as arithmetic and pull it to 85. The full list is in the repository so the argument can be had with the names visible.</p><p><strong>The compile-time split is a lower bound on the Triton share, not a precise ratio.</strong> The outer measurement includes Python AST walking and object construction; the inner one is a bare <code>ptxas</code> process. A fairer accounting would instrument the MLIR pass manager. We would expect that to move the number somewhat and not to change the direction.</p><p><strong>The 80 to 1 backend size ratio is close to meaningless</strong> and we include it only because it is measured and will be quoted anyway. It measures licensing, not effort.</p><p><strong>The Twill results are someone else&#8217;s measurements on hardware we do not have.</strong> We are relying on them for a claim central to section 12. Two of the seven authors work at NVIDIA, which cuts both ways: they have unusually good access, and they have an interest in the conclusion that heuristic compilers leave performance on the table.</p><p><strong>We are one version deep.</strong> All of this is Triton 3.7.1. The pipeline has been rewritten more than once in two years, the AMD backend changed substantially in 3.7, and the numbers in sections 06, 08 and 11 should be assumed to drift. Rerun the harness rather than trusting the figures.</p><p><strong>The four-byte cliff is an inference from one instruction&#8217;s encoding.</strong> We observe the threshold and we observe that <code>cp.async</code> has 4, 8 and 16 byte forms. The causal claim connecting them is ours, not something we read in the compiler. It is a very short inference and we could still be wrong about the mechanism while being right about the number.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Five dated predictions</h2><ol><li><p><strong>By the end of 2027, Triton will request non-</strong><code>a</code><strong> targets for at least one architecture family.</strong> The TODO is four lines long, the shipped assembler already accepts the <code>f</code> targets, and the pressure comes from anyone shipping precompiled kernels. 60 percent.</p></li><li><p><strong>Alignment will become assertable in the Triton language before the inference is removed.</strong> Some annotation or type-level marker letting an author declare alignment rather than having it guessed at launch, plus a diagnostic when the pipeliner declines. 50 percent by end of 2027, and we would rather be wrong in the direction of it happening sooner.</p></li><li><p><strong>The Tile IR backend will not be merged into mainline Triton before 2028.</strong> It is an incubator repository, it is Blackwell only, and merging it puts a closed lowering path inside the project whose institutional value is being the open one. 70 percent.</p></li><li><p><strong>A solver-based scheduler will ship inside a mainstream kernel compiler by the end of 2027.</strong> Twill demonstrated 19-second solve times for a problem currently handled by baked-in heuristics, and the gap it exposed is the kind that gets closed. 55 percent.</p></li><li><p><strong>Gluon&#8217;s public surface will grow faster than Triton&#8217;s over the next eighteen months.</strong> Measured in public symbols. If the ceiling is in the inference rather than the representation, the escape hatch is where the work goes. 65 percent.</p></li></ol><p>The July predictions from the compiler moat piece are on the record and two resolved early, one in our favour and one against, which we scored in the GPU software gap piece. These get the same treatment.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/p/exploring-how-triton-actually-compiles?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/p/exploring-how-triton-actually-compiles?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><div><hr></div><h2>Confidence dossier</h2><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!5_WQ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F12b19001-7a1d-46b4-a030-eb77e76d5a52_2160x3566.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!5_WQ!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F12b19001-7a1d-46b4-a030-eb77e76d5a52_2160x3566.png 424w, https://substackcdn.com/image/fetch/$s_!5_WQ!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F12b19001-7a1d-46b4-a030-eb77e76d5a52_2160x3566.png 848w, https://substackcdn.com/image/fetch/$s_!5_WQ!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F12b19001-7a1d-46b4-a030-eb77e76d5a52_2160x3566.png 1272w, https://substackcdn.com/image/fetch/$s_!5_WQ!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F12b19001-7a1d-46b4-a030-eb77e76d5a52_2160x3566.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!5_WQ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F12b19001-7a1d-46b4-a030-eb77e76d5a52_2160x3566.png" width="1456" height="2404" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/12b19001-7a1d-46b4-a030-eb77e76d5a52_2160x3566.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:2404,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:933347,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/208654595?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F12b19001-7a1d-46b4-a030-eb77e76d5a52_2160x3566.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!5_WQ!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F12b19001-7a1d-46b4-a030-eb77e76d5a52_2160x3566.png 424w, https://substackcdn.com/image/fetch/$s_!5_WQ!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F12b19001-7a1d-46b4-a030-eb77e76d5a52_2160x3566.png 848w, https://substackcdn.com/image/fetch/$s_!5_WQ!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F12b19001-7a1d-46b4-a030-eb77e76d5a52_2160x3566.png 1272w, https://substackcdn.com/image/fetch/$s_!5_WQ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F12b19001-7a1d-46b4-a030-eb77e76d5a52_2160x3566.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Software Frontier is a reader-supported publication. To continue reading and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h2></h2>
      <p>
          <a href="https://www.thesoftwarefrontier.com/p/exploring-how-triton-actually-compiles">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[DeepSeek V4-Flash: The Cost of Deciding What to Read]]></title><description><![CDATA[284 billion parameters rebuilt from the published constants, a million-token cache in 3.37 GiB, and the arithmetic showing that 4/5 of the attention budget is spent choosing what to attend to.]]></description><link>https://www.thesoftwarefrontier.com/p/deepseek-v4-flash-the-cost-of-deciding</link><guid isPermaLink="false">https://www.thesoftwarefrontier.com/p/deepseek-v4-flash-the-cost-of-deciding</guid><dc:creator><![CDATA[Lorenzo Bradanini]]></dc:creator><pubDate>Mon, 10 Aug 2026 06:45:30 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!i_aj!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F58602cdd-1158-491a-827b-650143b365c9_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!i_aj!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F58602cdd-1158-491a-827b-650143b365c9_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!i_aj!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F58602cdd-1158-491a-827b-650143b365c9_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!i_aj!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F58602cdd-1158-491a-827b-650143b365c9_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!i_aj!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F58602cdd-1158-491a-827b-650143b365c9_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!i_aj!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F58602cdd-1158-491a-827b-650143b365c9_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!i_aj!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F58602cdd-1158-491a-827b-650143b365c9_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/58602cdd-1158-491a-827b-650143b365c9_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2550191,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/209988515?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F58602cdd-1158-491a-827b-650143b365c9_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!i_aj!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F58602cdd-1158-491a-827b-650143b365c9_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!i_aj!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F58602cdd-1158-491a-827b-650143b365c9_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!i_aj!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F58602cdd-1158-491a-827b-650143b365c9_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!i_aj!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F58602cdd-1158-491a-827b-650143b365c9_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>I&#8217;ll show you seven numbers this piece derives</h2><div class="callout-block" data-callout="true"><p><strong>284.202 B</strong> ; reconstructed backbone against a published 284 B</p><p><strong>39.3 %</strong> ; of activated parameters are attention, which is 1.74 percent of the weights</p><p><strong>3.37 GiB</strong> ; of KV for a 1,048,576-token sequence, 1.96 percent of a GQA-8 baseline</p><p><strong>79 %</strong> ; of attention FLOPs at 1M are the selector, not the attention</p><p><strong>4096</strong> ; FLOPs per byte Flash needs from its interconnect, 1.5x Pro&#8217;s demand</p><p><strong>5,504</strong> ; tokens of recompute restore a million-token prefix, a 191x saving</p><p><strong>$2.92</strong> ; per GPU-hour is what 1M context pays; owners clear it, renters do not</p></div><p>I didn&#8217;t set out to write about <strong>DeepSeek V4-Flash,</strong> but i just wanted to check a precise number. The technical report gives the model&#8217;s constants in a paragraph on page 25 and its headline size, 284B total and 13B activated, on page 4, and I wanted to know whether the two agreed before I trusted anything else in the document. </p><p>They agree to seven hundredths of one percent, but only once you notice that the headline quietly leaves out the multi-token prediction module and that<strong> Heavily Compressed Attention</strong> carries half the key-value projections that Compressed Sparse Attention does. </p><p>Doing the raw math anyway shows something that the report never states: attention is 1.7 percent of this model&#8217;s weights and 39 percent of the weights it touches per token.</p><p>That is the kind of fact that changes what you build. It means the sparsity everyone is now discussing on reddit, lives <strong>entirely in the expert bank</strong>, that the attention stack is as dense as it has ever been, and that there is a hard floor under how cheap a model in this family can get. It also doesnt appear in any of the two dozen write-ups of this model I read before starting.</p><p>So this is more a <strong>reconstruction </strong>rather than a summary. Every architectural number below was just rebuilt from the published constants and checked against something DeepSeek or the <strong>vLLM team</strong> published independently. </p><p>Where the reconstruction disagrees with what is currently written about V4-Flash on the open web, I say so. Where it disagrees with the report, I say that too.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>What shipped, and what the number on the card means</h2><p>As of now, there are<strong> two DeepSeek-V4-Flash releases</strong> and they are the same weights. The preview landed on 24 April 2026 with the 58 page technical report and DeepSeek-V4-Pro. </p><p>The official release, tagged 0731, landed on 31 July 2026 as a public beta of the API. DeepSeek&#8217;s changelog states that <strong><span>DeepSeek-V4-Flash-0731</span></strong> has the same model structure and the same size as the preview and that only the post-training was rerun. I&#8217;ve not yet seen a new architecture, or a new parameter count; even the price hasn&#8217;t changed.</p><p>That is convenient for a piece like this, because every architectural fact in the April report still describes the model you can call today. It also means the <strong>agent scores DeepSeek published with the 0731 release</strong>, Terminal Bench 2.1 at 82.7 and DeepSWE at 54.4, are alignment results rather than efficiency results. </p><p>They were measured with DeepSeek&#8217;s own harness in what the changelog calls minimal mode at the max reasoning tier, temperature 1.0, top_p 0.95. The harness has not shipped. </p><p>Agent scores move by ten points on harness changes. Just treat them as a <strong>statement of intent.</strong></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!10Gx!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcc9d9f9d-0547-44ef-b260-8398c5da5a95_1600x832.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!10Gx!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcc9d9f9d-0547-44ef-b260-8398c5da5a95_1600x832.png 424w, https://substackcdn.com/image/fetch/$s_!10Gx!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcc9d9f9d-0547-44ef-b260-8398c5da5a95_1600x832.png 848w, https://substackcdn.com/image/fetch/$s_!10Gx!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcc9d9f9d-0547-44ef-b260-8398c5da5a95_1600x832.png 1272w, https://substackcdn.com/image/fetch/$s_!10Gx!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcc9d9f9d-0547-44ef-b260-8398c5da5a95_1600x832.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!10Gx!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcc9d9f9d-0547-44ef-b260-8398c5da5a95_1600x832.png" width="1456" height="757" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/cc9d9f9d-0547-44ef-b260-8398c5da5a95_1600x832.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:757,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;DeepSeek-V4-Flash, published configuration. Report section 4.2.1, stated directly.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="DeepSeek-V4-Flash, published configuration. Report section 4.2.1, stated directly." title="DeepSeek-V4-Flash, published configuration. Report section 4.2.1, stated directly." srcset="https://substackcdn.com/image/fetch/$s_!10Gx!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcc9d9f9d-0547-44ef-b260-8398c5da5a95_1600x832.png 424w, https://substackcdn.com/image/fetch/$s_!10Gx!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcc9d9f9d-0547-44ef-b260-8398c5da5a95_1600x832.png 848w, https://substackcdn.com/image/fetch/$s_!10Gx!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcc9d9f9d-0547-44ef-b260-8398c5da5a95_1600x832.png 1272w, https://substackcdn.com/image/fetch/$s_!10Gx!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcc9d9f9d-0547-44ef-b260-8398c5da5a95_1600x832.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>DeepSeek-V4-Flash, published configuration. Report section 4.2.1, stated directly.</em></p><p>I want to focus on one correction, before we go further and deeper. A widely cited <strong>third-party guide</strong> states that DeepSeek publishes the layer arrangement for Pro but not for Flash, and advises readers to treat Flash&#8217;s layer count as unconfirmed. </p><p>That is<strong> totally wrong. </strong>The report gives Flash 43 layers and hidden dimension 4096 in the same paragraph that gives it 284B parameters. Pro gets 61 layers and 7168. </p><p>The arrangements differ in one way that matters: Pro&#8217;s first two layers are HCA, Flash&#8217;s first two are pure sliding window attention with no compression at all.</p><div><hr></div><h2>Rebuilding the model from its constants</h2><p>Every expert is a <strong>SwiGLU block</strong>, so three matrices of 4096 by 2048, which is precisely 25.17M parameters. With 256 routed experts plus one shared expert in all 43 blocks, the expert bank alone is 278.108B. </p><p>That is 97.8 percent of the model and it takes one line of arithmetic. Everything interesting is (in my opinion) in the <strong>remaining 2.2 percent.</strong></p><p>The attention shapes are not what you would guess from a normal transformer, and <strong>CSA</strong> and <strong>HCA </strong>are not the same size. CSA computes two independent key-value streams, so four projection matrices from equations 9 and 10 of the report. </p><p>HCA computes one, so two matrices, from equations 20 and 21. Getting this wrong is the difference between a reconstruction that lands and one that does not.</p><pre><code>CSA layer
  W^aKV, W^bKV, W^aZ, W^bZ    4 x (4096 x 512)   =  8.389 M   two overlapped KV streams
  B^a, B^b                     2 x (4 x 512)      =  0.004 M   learnable positional bias
  W^DQ                         4096 x 1024        =  4.194 M   query down-projection
  W^UQ                         1024 x (512 x 64)  = 33.554 M   query up-projection
  W^IUQ                        1024 x (128 x 64)  =  8.389 M   indexer query up-projection
  W^w                          4096 x 64          =  0.262 M   per-head indexer gate
  grouped output, 8 groups     8 x (4096 x 1024)  = 33.554 M
  final output projection      8192 x 4096        = 33.554 M
  attention sink logits        64                 =  0.000 M
                                                    ---------
                                                    121.901 M

HCA layer   one KV stream, no indexer, bias over m&#8217; = 128         109.117 M
SWA layer   uncompressed KV, no compression weights, no indexer   106.954 M</code></pre><p>Manifold-constrained hyper-connections add 393K per residual junction. <strong>Router gates add 45M</strong> across the model. Embeddings, using the DeepSeek-V3 tokenizer with a handful of added context-construction tokens, add 530M on each side. </p><p>The report says the vocabulary remains 128K, which I read as the 129,280 entries V3 used.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!sWAa!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc88d9bcd-3f58-4e01-957b-c952ded8e67b_1600x619.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!sWAa!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc88d9bcd-3f58-4e01-957b-c952ded8e67b_1600x619.png 424w, https://substackcdn.com/image/fetch/$s_!sWAa!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc88d9bcd-3f58-4e01-957b-c952ded8e67b_1600x619.png 848w, https://substackcdn.com/image/fetch/$s_!sWAa!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc88d9bcd-3f58-4e01-957b-c952ded8e67b_1600x619.png 1272w, https://substackcdn.com/image/fetch/$s_!sWAa!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc88d9bcd-3f58-4e01-957b-c952ded8e67b_1600x619.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!sWAa!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc88d9bcd-3f58-4e01-957b-c952ded8e67b_1600x619.png" width="1456" height="563" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c88d9bcd-3f58-4e01-957b-c952ded8e67b_1600x619.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:563,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Where 284 billion parameters actually sit. Reconstructed from the published constants; the residual against the published figure is 0.07 percent.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Where 284 billion parameters actually sit. Reconstructed from the published constants; the residual against the published figure is 0.07 percent." title="Where 284 billion parameters actually sit. Reconstructed from the published constants; the residual against the published figure is 0.07 percent." srcset="https://substackcdn.com/image/fetch/$s_!sWAa!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc88d9bcd-3f58-4e01-957b-c952ded8e67b_1600x619.png 424w, https://substackcdn.com/image/fetch/$s_!sWAa!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc88d9bcd-3f58-4e01-957b-c952ded8e67b_1600x619.png 848w, https://substackcdn.com/image/fetch/$s_!sWAa!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc88d9bcd-3f58-4e01-957b-c952ded8e67b_1600x619.png 1272w, https://substackcdn.com/image/fetch/$s_!sWAa!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc88d9bcd-3f58-4e01-957b-c952ded8e67b_1600x619.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>Where 284 billion parameters actually sit. Reconstructed from the published constants; the residual against the published figure is 0.07 percent.</em></p><p>284.202B against a published 284B,<em> a 0.07 percent residual</em> on a reconstruction with no free parameters. </p><p>The <strong>MTP head is real</strong>, it is 6.6B, vLLM will use it for speculative decoding, and it is not in the number on the model card. DeepSeek did the same with V3, whose Hugging Face repository showed 685B against a stated 671B.</p><p>There is a second confirmation of the reconstruction hiding in an unlikely place. In the <strong>determinism section</strong>, discussing why they cannot use split-k for <em>one particular GEMM</em>, the report mentions in passing that &#8220;<em>mHC involves a matrix multiplication with an output dimension of only 24.</em>&#8221; </p><p>With n_hc = 4, the dynamic parameterisation generates A in R^4, B in R^{4x4} and C in R^4 from one flattened input. Four plus sixteen plus four is twenty-four. </p><p>The three mappings are produced by a single GEMM, and a throwaway sentence in a section about floating-point associativity confirms the shape.</p><p>The <strong>activated count</strong> follows. Seven experts fire per token, six routed plus the shared one, which is 7.575B. Everything else in the forward pass is dense: all 4.956B of attention, all of mHC, all of the router gates. </p><p>That is <strong>12.610B</strong> excluding embeddings and 13.140B counting the output head, against a published 13B.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!zTVu!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe778cefd-9d87-4c3f-9dff-dad878a474c0_1600x806.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!zTVu!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe778cefd-9d87-4c3f-9dff-dad878a474c0_1600x806.png 424w, https://substackcdn.com/image/fetch/$s_!zTVu!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe778cefd-9d87-4c3f-9dff-dad878a474c0_1600x806.png 848w, https://substackcdn.com/image/fetch/$s_!zTVu!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe778cefd-9d87-4c3f-9dff-dad878a474c0_1600x806.png 1272w, https://substackcdn.com/image/fetch/$s_!zTVu!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe778cefd-9d87-4c3f-9dff-dad878a474c0_1600x806.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!zTVu!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe778cefd-9d87-4c3f-9dff-dad878a474c0_1600x806.png" width="1456" height="733" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e778cefd-9d87-4c3f-9dff-dad878a474c0_1600x806.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:733,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Attention weights are a rounding error in the checkpoint and a plurality of the arithmetic at decode. Reconstructed from the constants in section 4.2.1.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Attention weights are a rounding error in the checkpoint and a plurality of the arithmetic at decode. Reconstructed from the constants in section 4.2.1." title="Attention weights are a rounding error in the checkpoint and a plurality of the arithmetic at decode. Reconstructed from the constants in section 4.2.1." srcset="https://substackcdn.com/image/fetch/$s_!zTVu!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe778cefd-9d87-4c3f-9dff-dad878a474c0_1600x806.png 424w, https://substackcdn.com/image/fetch/$s_!zTVu!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe778cefd-9d87-4c3f-9dff-dad878a474c0_1600x806.png 848w, https://substackcdn.com/image/fetch/$s_!zTVu!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe778cefd-9d87-4c3f-9dff-dad878a474c0_1600x806.png 1272w, https://substackcdn.com/image/fetch/$s_!zTVu!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe778cefd-9d87-4c3f-9dff-dad878a474c0_1600x806.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>Attention weights are a rounding error in the checkpoint and a plurality of the arithmetic at decode. Reconstructed from the constants in section 4.2.1.</em></p><p>Attention is 1.74 percent of the weight bank and 39.3 percent of what gets touched per token. That&#8217;s most load-bearing fact about serving this model and it<strong> appears nowhere</strong> in the report, because the report presents attention as the thing being optimised away rather than as a fixed cost that survives the optimisation.</p><blockquote><p><em>The sparsity everyone is discussing lives entirely in the expert bank. The attention stack is as dense as it has ever been.</em></p></blockquote><p>The consequence is a floor. <em>Push MoE sparsity as far as you like</em>, drop from six routed experts to four, go <strong>from 256 experts to 512 </strong>at the same activation count, and the activated parameter count will not fall below roughly 5.0B, because the attention stack is dense and it is 4.956B of it.</p><p>Every decode step of every request reads all of it. At 12.6B activated you are already 40 percent of the way to that floor. A hypothetical<strong> V4-Flash-Nano </strong>with two routed experts would be a 10.1B-activated model, not a 4B one.</p><p>It also reframes what the hybrid attention is for. It is not there to make attention cheap in absolute terms. It is there to stop attention&#8217;s <em>state</em> from growing without bound, which is a <strong>memory problem </strong>rather than a FLOPs problem, and the FLOPs bill arrives anyway. </p><p>We will come back later to how large it gets.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Compressed Sparse Attention</h2><p>CSA is <strong>five separate ideas stacked</strong>, and it is often described as one. Taken apart, each of them is simple, and two of them are the sort of thing you only invent after you have been beaten by the alternative.</p><h3>Key and value are the same vector</h3><p>V4 stores one cached entry per compressed position, not two. Here, the key and the value are literally the <strong>same tensor</strong>, used in both roles by a multi-query attention. That is an immediate halving of the cache before any compression happens at all.</p><p>It should not work. Attention output is normally translation invariant: rotate the query by <strong>R(t)</strong> and the key by <strong>R(s)</strong> and the score depends on <em>R(t-s), which is relative.</em> Values carry no rotation, so the output carries none. </p><p>Share the key and the value, and the value inherits the key&#8217;s rotation, the output acquires an <strong>absolute position</strong> through R(s), and shifting the whole sequence changes the answer.</p><p>The fix is in section 2.3.3 and it is one line: apply RoPE with position minus i to the last 64 dimensions of each head&#8217;s output. Because R is orthogonal and <strong>R(t)<sup>-1</sup>R(s) = R(s-t)</strong>, rotating the output backwards restores relative positioning, and the contribution of each cached entry ends up depending on its distance from the query again. </p><p>The vLLM team&#8217;s write-up derives the same result from the other direction and calls it <strong>inverse RoPE</strong>. Their implementation fuses it into the FP8 quantisation ahead of the output projection, worth two to three times over doing the two separately.</p><p><strong>Two times the cache</strong>, bought with one elementwise kernel. This is the cheapest trick in the model and it is the one that has attracted the least attention.</p><h3>The compressor has two streams and they overlap</h3><p>CSA does not average four tokens into one. It computes two independent projections of the <strong>hidden state, C<sup>a</sup> and C<sup>b</sup></strong>, each with its own learned compression weights <strong>Z<sup>a</sup> and Z<sup>b</sup></strong>, then takes a softmax across the concatenation of 2m weights and forms a weighted sum over both streams. </p><p>Compressed entry i draws C<sup>a</sup> from positions [mi, m(i+1)] and C<sup>b</sup> from positions [m(i-1), mi].</p><p>With m = 4 that makes every compressed entry a data-dependent weighted sum of eight consecutive tokens taken at a stride of four. <strong>vLLM names this attention type <span>c4a</span> </strong>and documents it as a weighted sum of 8 uncompressed tokens with a stride of 4, which is exactly equations 11 and 12.</p><p>The overlap is the point. A hard boundary every four tokens cuts arbitrary spans in arbitrary places and the model is blind across the seam. Overlapping means every token appears in <strong>two compressed entries </strong>under different weights, while the sequence still shrinks by exactly four because the stride is four. The receptive field of an entry is eight; the compression ratio is four.</p><p>The compression weights are per dimension. The softmax runs over <strong>2m elements independently</strong> for each of the 512 channels, so each channel picks its own mixture of the eight tokens. </p><p>This is a learned per-channel pooling rather than a summarisation step, and describing it as &#8220;remembering the paragraph&#8217;s key point&#8221; undersells it by a wide margin.</p><h3>The Lightning Indexer</h3><p>After compression a one-million-token context still holds 262,144 entries in every CSA layer. Attending densely to a quarter of a million entries is not obviously better than attending densely to a million, so <strong>CSA runs DeepSeek Sparse Attention</strong> over the compressed stream: a cheap scoring pass picks the top 512, and the real attention runs only over those.</p><p>The scorer builds 64 low-rank query heads of dimension 128 from the same compressed query latent the main attention uses, computes a rectified dot product against a separately compressed indexer key for each block, and sums across heads with a learned per-head gate:</p><pre><code>c^Q_t   = h_t &#183; W^DQ                                   shared with the main attention queries
q^I_t   = c^Q_t &#183; W^IUQ                                64 indexer heads of dimension 128
w^I_t   = h_t &#183; W^w                                    one gate per indexer head, from the hidden state
I(t,s)  = sum_h  w^I(t,h) &#183; ReLU( q^I(t,h) &#183; K^IComp_s )</code></pre><p>Two choices there are worth stopping on. <strong>ReLU rather than softmax leaves the score unnormalised</strong>, so a head can contribute nothing to a block rather than merely down-weighting it, and heads can veto. And the gate w is produced from the hidden state by a 4096 by 64 matrix, so the model decides per token which of its 64 scoring heads to trust. </p><p>That is a router in everything but name, sitting in front of the attention, and it is trained end to end with it.</p><p><strong>Flash&#8217;s top-k is 512</strong>. Pro&#8217;s is 1024. V3.2&#8217;s was 2048. The report is direct about the reason: a smaller top-k greatly improves efficiency on short and medium texts, which is where the traffic is.</p><h3>Grouped output projection</h3><p>Sixty-four heads at head dimension 512 produce <strong>32,768 values per token. </strong>A conventional output projection would be 32,768 by 4096, which is 134M parameters per layer, more than everything else in the attention block put together. </p><p><strong>V4 splits the heads into g = 8 groups of eight</strong>, projects each group&#8217;s 4096 values to 1024, concatenates the eight results into 8192, and projects that to 4096. Total 67.1M, half the naive cost, and the bottleneck at 8192 acts as a constraint on how freely heads can mix.</p><p>Pro uses <strong>g = 16 at the same d_g = 1024</strong>, so a 16,384-wide concatenation from 128 heads. The report also applies RMSNorm per head on the queries and on the single head of the compressed KV entries just before the core attention, which is the same numerical hygiene MLA needed, for the same reason, and which turns out to matter for the optimiser.</p><h3>Attention sink, and heads that abstain</h3><p>Both CSA and HCA carry a set of learnable sink logits, one per head. For head h, <strong>Exp(z&#8217;<sub>h</sub>) is added to the denominator </strong>of the softmax and to nothing else:</p><pre><code>s(h,i,j) = Exp(z(h,i,j)) / ( sum_k Exp(z(h,i,k)) + Exp(z&#8217;_h) )</code></pre><p>The report&#8217;s description of what this buys is unusually blunt. It allows each query head to make its total attention score not equal to one, &#8220;<em>and even to be near 0.</em>&#8221; A head with a large sink logit contributes almost nothing to the output no matter what is in the context. </p><p>Given that <strong>this model asks 64 heads per layer </strong>to attend to a top-512 selection out of a quarter of a million compressed blocks, giving heads a principled way to decline is not decoration. </p><p>It is what stops a head from being forced to spend its mass on whichever blocks the indexer happened to hand it.</p><h3>The sliding window is not an optimisation</h3><p>Every compressed layer also keeps <em>128 uncompressed tokens in a window</em>, concatenated with the selected compressed entries before the softmax. This is a correctness requirement and the reason is causality.</p><p>A <strong>compressed entry i in <span>c128a</span></strong> summarises positions 128i through 128(i+1)-1. A query at position t may only use information derived from positions at or before t, so it cannot use entry i unless 128(i+1)-1 is at or before t. </p><p>A query sitting anywhere inside the current block therefore has no compressed entry it is permitted to read. Without the window, tokens 1 through 127 of every block would attend to no local context whatsoever. </p><p>The window covers the distance between the query and the most recent legal compression boundary, and 128 is exactly m&#8217; for that reason.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Heavily Compressed Attention</h2><p>HCA is the same compressor at m&#8217; = 128, with two differences that both simplify it. There is<strong> one key-value stream</strong> instead of two, so no overlap and no seam handling. And there is no indexer, so no top-k. It attends densely over everything it holds.</p><p>The arithmetic explains the second choice. A <strong>one-million-token context under 128x compression</strong> yields 8,192 entries, and eight thousand keys is an ordinary attention problem, shorter than most models&#8217; native context. Sparsity would buy nothing; vLLM implements it as a sparse-attention call with top-k set to 8192, a selection that selects everything, purely so one kernel serves both paths.</p><p>Dropping the overlap is defensible because HCA is not responsible for local detail. The sliding window handles that and the interleaved CSA layers handle the middle range. <strong>HCA&#8217;s job is coarse global memory</strong>, and boundary precision on a 128-token block does not matter much when the question is what the document was about.</p><p>What the pair produces is a genuinely two-rate memory. In Flash&#8217;s stack, two sliding-window layers are followed by 41 alternating layers, giving<strong> 21 CSA and 20 HCA</strong>. </p><p>Every token&#8217;s representation passes through 21 layers that can retrieve 512 four-token spans from anywhere in the history and 20 layers that see all 8,192 coarse summaries at once. Neither works alone. </p><p>Sparse selection over fine blocks has recall problems on diffuse queries; dense attention over coarse blocks cannot resolve a specific line of code.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>The router, and the three ways V4 changed it</h2><p>Every write-up of this model spends its length on the attention and treats the mixture of experts as inherited furniture. It is not. <strong>Section 2.1 makes four changes to routing</strong>, and one of them removes a whole class of layer that every DeepSeek model before this one had.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!kL1g!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83871c94-9dd8-471d-91b4-d1b1b9cae9af_1600x450.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!kL1g!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83871c94-9dd8-471d-91b4-d1b1b9cae9af_1600x450.png 424w, https://substackcdn.com/image/fetch/$s_!kL1g!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83871c94-9dd8-471d-91b4-d1b1b9cae9af_1600x450.png 848w, https://substackcdn.com/image/fetch/$s_!kL1g!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83871c94-9dd8-471d-91b4-d1b1b9cae9af_1600x450.png 1272w, https://substackcdn.com/image/fetch/$s_!kL1g!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83871c94-9dd8-471d-91b4-d1b1b9cae9af_1600x450.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!kL1g!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83871c94-9dd8-471d-91b4-d1b1b9cae9af_1600x450.png" width="1456" height="409" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/83871c94-9dd8-471d-91b4-d1b1b9cae9af_1600x450.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:409,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Forty-three layers, two independent schedules. Reconstructed from report sections 2.1 and 4.2.1; the CSA and HCA assignment within the interleave is inferred, and the split is either 21 to 20 or the reverse.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Forty-three layers, two independent schedules. Reconstructed from report sections 2.1 and 4.2.1; the CSA and HCA assignment within the interleave is inferred, and the split is either 21 to 20 or the reverse." title="Forty-three layers, two independent schedules. Reconstructed from report sections 2.1 and 4.2.1; the CSA and HCA assignment within the interleave is inferred, and the split is either 21 to 20 or the reverse." srcset="https://substackcdn.com/image/fetch/$s_!kL1g!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83871c94-9dd8-471d-91b4-d1b1b9cae9af_1600x450.png 424w, https://substackcdn.com/image/fetch/$s_!kL1g!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83871c94-9dd8-471d-91b4-d1b1b9cae9af_1600x450.png 848w, https://substackcdn.com/image/fetch/$s_!kL1g!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83871c94-9dd8-471d-91b4-d1b1b9cae9af_1600x450.png 1272w, https://substackcdn.com/image/fetch/$s_!kL1g!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83871c94-9dd8-471d-91b4-d1b1b9cae9af_1600x450.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>Forty-three layers, two independent schedules. Reconstructed from report sections 2.1 and 4.2.1; the CSA and HCA assignment within the interleave is inferred, and the split is either 21 to 20 or the reverse.</em></p><p><strong>There are no dense FFN layers.</strong> V3 kept ordinary dense feed-forward blocks in its first three Transformer layers, on the standard argument that early layers do generic work and routing them wastes capacity. </p><p>V4 replaces those with MoE layers that use <em>hash routing</em>: the target experts for a token are fixed by a hash of its token ID, following Roller&#8217;s Hash Layers work from 2021. The Hugging Face implementation makes the mechanism concrete. </p><p><strong>Routing type</strong> is set per layer through <span>mlp_layer_types</span>, and a hash layer resolves its experts through a frozen <span>tid2eid</span> lookup shipped inside the checkpoint.</p><p>The details that make this more than a curiosity is that only the <em>selection</em> is static. The learned gate still produces the per-expert scores that weight the chosen experts. So a <strong>hash layer </strong>is not an un-routed layer; it is a layer where the router has been told which experts to consider and gets to decide how much to trust each one. </p><p>That converts the hardest part of early-layer routing, an <strong>unstable argmax over 256 options</strong> while the model knows nothing, into a fixed assignment with a learnable mixture on top.</p><p>It also <strong>does something useful</strong> for the infrastructure. A hash of the token ID is known before the forward pass reaches the layer, which means dispatch for the first three layers can be planned as soon as the tokens are known rather than after the previous block finishes. </p><h4>three layers where the overlap is free</h4><ul><li><p><strong>The affinity function changed.</strong> <em>V3 scored expert affinity with a sigmoid. V4 uses the square root of a softplus. Both are positive and monotone, so the ranking behaviour is similar, but the tails are not. A sigmoid saturates at one, so once an expert is clearly the best its score stops responding and the gradient through it vanishes. Softplus does not saturate, and the square root damps its growth to sublinear without ever flattening. The practical effect is that a strongly preferred expert keeps receiving gradient signal instead of going quiet, which is exactly the failure mode you would expect to precede the routing-driven loss spikes described two sections later.</em></p></li><li><p><strong>The routing target cap is gone.</strong> <em>V3 constrained how many nodes a token&#8217;s experts could be spread across, through the <span>n_group</span> and <span>topk_group</span> parameters, because unconstrained routing means a token&#8217;s six experts can live on six different machines and the all-to-all cost is set by the worst case. V4 drops the constraint entirely and says the parallelism strategy was redesigned to pay for it. Which is the same trade appearing again in a different costume. The wave-partitioned mega-kernel in section 3.1 exists so that communication hides under computation. Once it does, capping communication to protect throughput stops being necessary, and the model gets its routing freedom back. Every efficiency result in this report buys an architectural freedom somewhere else, and the report never quite says so.</em></p></li><li><p><strong>Load balancing stays auxiliary-loss-free, with one addition.</strong> <em>V4 keeps V3&#8217;s scheme, where a per-expert bias is added to the score for the purpose of top-k selection and excluded from the gating weight, so balance is enforced without a gradient term competing with the language modelling objective. The Hugging Face implementation keeps it as an <span>e_score_correction_bias</span> buffer that shifts the argmax without carrying gradients. On top of that V4 adds a mild sequence-wise balance loss whose only job is to stop a single sequence collapsing onto a handful of experts, with a weight of 0.0001 and a bias update speed of 0.001. Global balance from a bias, local balance from a loss, and the loss is small enough to be a guardrail rather than an objective.</em></p></li></ul><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Manifold-Constrained Hyper-Connections</h2><p>Hyper-connections widen the residual stream from one vector to n_hc vectors and let the model learn the mixing, which decouples residual width from hidden size. </p><p><strong>The update is X<sub>l+1</sub> = B<sub>l</sub>X<sub>l</sub> + C<sub>l</sub>F<sub>l</sub>(A<sub>l</sub>X<sub>l</sub>),</strong> with A projecting the widened stream down to the layer input, C projecting the layer output back up, and B mixing the residual lanes among themselves.</p><p>DeepSeek&#8217;s stated problem with plain hyper-connections is that stacking them is numerically unstable. B is applied at every junction, so across 86 junctions the model computes a product of <strong>86 learned matrices. </strong>If the spectral norm of B exceeds 1 by any margin, the product diverges. If it falls below 1, the signal dies.</p><p>The fix is to constrain B to the <strong>Birkhoff polytope</strong>, the set of doubly stochastic matrices: nonnegative, every row and column summing to one. Two properties make this the right set rather than a convenient one. </p><p>A <strong>doubly stochastic matrix</strong> has spectral norm exactly 1, so the residual mapping is non-expansive by construction and neither the forward nor the backward pass can blow up through it. And the set is closed under multiplication, so a product of 86 of them is still doubly stochastic. The stability is structura-l, not empirical.</p><p>Projection onto the polytope uses <strong>Sinkhorn-Knopp:</strong> exponentiate the raw matrix for positivity, then alternate row and column normalisation, twenty times. A and C get a sigmoid, with C scaled by two so it can express amplification up to a factor of two while staying nonnegative, which rules out lanes cancelling one another.</p><p>One implementation detail worth recording because the paper skips it: the widened stream has to collapse before the output. A final hyper-head folds the <strong>four residual lanes</strong> back into a single sequence just ahead of the model norm, so the language modelling head sees an ordinary hidden state and nothing downstream needs to know the residual was ever four vectors wide.</p><p><strong>At n_hc = 4 this costs 393K parameters per junction and 34M across the model</strong>, about a hundredth of one percent of the weights. The cost is elsewhere. The residual stream is four times wider in activation memory, twenty Sinkhorn iterations sit on the critical path of every junction, and pipeline communication between stages grows. </p><p>DeepSeek&#8217;s answer is obviously <strong>fused kernels</strong>, a recomputation strategy that checkpoints most inter-layer hidden states and all normalised layer inputs while leaving compute-intensive operations alone, and an adjustment to the DualPipe 1F1B overlap so parts of mHC run concurrently with the pipeline. </p><p>The number they report for the whole apparatus is 6.7 percent of the overlapped 1F1B stage.</p><p><strong>SGLang found the other end</strong> of the same problem at serving time. In low-latency decode the batch is small, the <strong>pre-GEMM</strong> that feeds the Sinkhorn normalisation has almost no parallelism, and it becomes the bottleneck. Their answer was to split the K dimension of that GEMM across CTAs. </p><p>Which is where the determinism section&#8217;s remark about an output dimension of 24 comes from: the<strong> GEMM is small enough that split-k is compulsory</strong> and split-k is non-deterministic, so they emit each split separately and reduce in a following kernel.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Muon, and the thing it did not break</h2><p>V4 trains with Muon for most parameters and AdamW for the embedding, the prediction head, the static biases and gating factors of mHC, and all <strong>RMSNorm weights</strong>. Momentum 0.95, weight decay 0.1, update RMS rescaled to 0.18 so the AdamW learning rate schedule could be reused unchanged.</p><p>The orthogonalisation is a hybrid Newton-Schulz, ten iterations in two stages. The first eight use coefficients (<em>3.4445, -4.7750, 2.0315</em>), which converge fast and overshoot; the final two use (<em>2, -1.5, 0.5</em>), which are gentler and settle the singular values precisely at one. </p><p>Splitting the schedule this way is the practical answer to a known tension in <strong>Newton-Schulz</strong>: aggressive coefficients reach the neighbourhood quickly and oscillate there, conservative ones land cleanly but slowly.</p><p>The detail worth flagging is a negative result. Muon has a documented pathology where orthogonalised updates keep singular values near uniform, query and key norms drift up together, and pre-softmax logits reach values low precision cannot hold. Moonshot&#8217;s answer in the Kimi work was<strong> QK-Clip</strong>.</p><p>DeepSeek states plainly that they do not use it, because the attention architecture<strong> already applies RMSNorm</strong> to the queries and the KV entries, which prevents the logits from exploding in the first place.</p><p>So the RMSNorm in section 2.3.3, which reads like routine hygiene, is load-bearing for the optimiser choice. That is a co-design decision presented as a footnote.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>How the sparsity was actually taught</h2><p>You can&#8217;t train a top-k selector from scratch. If the indexer is random at initialisation, the <strong>attention sees 512 random blocks</strong>, the gradient signal for choosing better blocks is buried, and the model learns to ignore the compressed path entirely.</p><p>The schedule in section 4.2.2 solves this in stages, and it is worth reading as a recipe rather than a list. Flash starts at <strong>sequence length 4K and extends through 16K and 64K to 1M. </strong>Attention is dense for the first 1T tokens. Sparse attention is introduced at the 64K stage, and before it is switched on there is a short stage that warms up the lightning indexer alone. Then sparse attention runs for the rest of training, which is most of 32T tokens.</p><p>The ordering is the interesting part. Dense first so the model learns what to attend to; then an <strong>indexer warmup </strong>so the selector learns to imitate the dense attention&#8217;s choices; then sparsity, at a sequence length long enough that selection matters and short enough that dense supervision was still affordable to produce. </p><p><strong>Pro gets a longer dense stage than Flash</strong>, which is what you would expect if the dense phase is the expensive part and the larger model needs more of it.</p><p>Two other schedule details are worth having. Batch size ramps to 75.5M tokens and stays there. Learning rate warms over 2000 steps to 2.7e-4, holds, then decays to 2.7e-5 on a cosine near the end. </p><p><strong>MTP loss weight is 0.3</strong> for most of training and drops to 0.1 when the learning rate starts decaying, which reads as a decision to stop letting the speculative head pull on the backbone once the model is being finished.</p><h3>The instability, and the two things that fixed it</h3><p>The report is unusually candid here. Training was unstable, rollbacks did not prevent recurrence, and the spikes were <strong>consistently traced to outliers in the MoE layers</strong>, with the routing mechanism itself appearing to make the outliers worse. </p><p>Two techniques fixed it and <strong>DeepSeek </strong>says openly that they do not have a theory for why.</p><div class="callout-block" data-callout="true"><p><strong>Anticipatory routing</strong>. It decouples the routing decision from the backbone update. At step t the model computes features with current parameters but routes with parameters from step t minus delta, and to avoid loading weights twice they fetch step t&#8217;s data early and cache the routing indices during the earlier step&#8217;s forward pass. That costs about 20 percent of wall time, so they do not run it continuously: an automatic detector triggers a short rollback and switches the mode on when a spike occurs, then reverts after a period. Amortised, the overhead is close to nothing. The mechanism is worth thinking about. If routing and features update together, a token that starts going to a bad expert gets a gradient that makes both the expert worse and the routing decision more confident, which is a positive feedback loop with no damping. Freezing the routing for a few steps breaks the loop by making the router a fixed target that the experts have to fit rather than a moving one that co-adapts.</p></div><div class="callout-block" data-callout="true"><p><strong>SwiGLU clamping</strong> is the blunt half. The linear component is clamped to [-10, 10] and the gate component is capped at 10, throughout the training of both models. Clamping an activation is what you reach for after watching a run diverge, and DeepSeek reports it eliminates outliers without compromising performance.</p></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>What the cache costs, checked four ways</h2><p>The report gives ratios against DeepSeek-V3.2 and<strong> vLLM gives absolute figures for Pro</strong>. Nobody publishes the absolute figure for Flash. It is derivable, and the derivation can be validated before it is used.</p><p>Build a byte model from the published dimensions. Section 2.3.3 fixes the rotary dimension at exactly 64, and section 2.3.4 specifies bf16 for the <strong>RoPE dimensions</strong> and FP8 for the rest, so a shared key-value entry of head dimension 512 costs 64 times 2 plus 448, which is 576 bytes. </p><p>The <strong>indexer cache runs in FP4</strong> under quantization-aware training, so a 128-dimension indexer key costs 64 bytes.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!yAdk!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F343d20eb-08c4-4d61-9833-c0699949033e_1600x760.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!yAdk!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F343d20eb-08c4-4d61-9833-c0699949033e_1600x760.jpeg 424w, https://substackcdn.com/image/fetch/$s_!yAdk!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F343d20eb-08c4-4d61-9833-c0699949033e_1600x760.jpeg 848w, https://substackcdn.com/image/fetch/$s_!yAdk!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F343d20eb-08c4-4d61-9833-c0699949033e_1600x760.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!yAdk!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F343d20eb-08c4-4d61-9833-c0699949033e_1600x760.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!yAdk!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F343d20eb-08c4-4d61-9833-c0699949033e_1600x760.jpeg" width="1456" height="692" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/343d20eb-08c4-4d61-9833-c0699949033e_1600x760.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:692,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:74905,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/209988515?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F343d20eb-08c4-4d61-9833-c0699949033e_1600x760.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!yAdk!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F343d20eb-08c4-4d61-9833-c0699949033e_1600x760.jpeg 424w, https://substackcdn.com/image/fetch/$s_!yAdk!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F343d20eb-08c4-4d61-9833-c0699949033e_1600x760.jpeg 848w, https://substackcdn.com/image/fetch/$s_!yAdk!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F343d20eb-08c4-4d61-9833-c0699949033e_1600x760.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!yAdk!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F343d20eb-08c4-4d61-9833-c0699949033e_1600x760.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>Two against figures vLLM published independently, two against ratios stated in the report itself.</em></p><p>V3.2 caches an MLA latent of 512 dimensions plus 64 RoPE and a 128-dimension indexer key. In bf16 that is 1,408 bytes per token per layer, which over 61 layers at 1,048,576 tokens is 83.88 GiB<strong>. vLLM publishes 83.9. V4 at Pro&#8217;s 30 CSA</strong> and 31 HCA layers gives 9.62 GiB in bf16. vLLM publishes 9.62. Applying the same model at production precision, V3.2 comes to 49.12 GB and Flash to 3.62 GB, a ratio of 13.57 to one, against the 13.7x the report prints on <strong>Figure 1. </strong></p><p>And section 3.6.2 says the uncompressed sliding-window state would be roughly eight times the volume of the compressed state; the byte model says 7.2.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!tIaC!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6bd4b030-c6d5-4bbd-89ea-d81f308c65a9_1600x880.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!tIaC!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6bd4b030-c6d5-4bbd-89ea-d81f308c65a9_1600x880.png 424w, https://substackcdn.com/image/fetch/$s_!tIaC!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6bd4b030-c6d5-4bbd-89ea-d81f308c65a9_1600x880.png 848w, https://substackcdn.com/image/fetch/$s_!tIaC!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6bd4b030-c6d5-4bbd-89ea-d81f308c65a9_1600x880.png 1272w, https://substackcdn.com/image/fetch/$s_!tIaC!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6bd4b030-c6d5-4bbd-89ea-d81f308c65a9_1600x880.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!tIaC!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6bd4b030-c6d5-4bbd-89ea-d81f308c65a9_1600x880.png" width="1456" height="801" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6bd4b030-c6d5-4bbd-89ea-d81f308c65a9_1600x880.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:801,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;One 1,048,576-token sequence. The model reproduces vLLM's published V3.2 and V4-Pro figures before being applied to Flash.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="One 1,048,576-token sequence. The model reproduces vLLM's published V3.2 and V4-Pro figures before being applied to Flash." title="One 1,048,576-token sequence. The model reproduces vLLM's published V3.2 and V4-Pro figures before being applied to Flash." srcset="https://substackcdn.com/image/fetch/$s_!tIaC!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6bd4b030-c6d5-4bbd-89ea-d81f308c65a9_1600x880.png 424w, https://substackcdn.com/image/fetch/$s_!tIaC!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6bd4b030-c6d5-4bbd-89ea-d81f308c65a9_1600x880.png 848w, https://substackcdn.com/image/fetch/$s_!tIaC!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6bd4b030-c6d5-4bbd-89ea-d81f308c65a9_1600x880.png 1272w, https://substackcdn.com/image/fetch/$s_!tIaC!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6bd4b030-c6d5-4bbd-89ea-d81f308c65a9_1600x880.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>One 1,048,576-token sequence. The model reproduces vLLM&#8217;s published V3.2 and V4-Pro figures before being applied to Flash.</em></p><p>So: <strong>3.372 GiB per one-million-token sequence</strong>, or 3,453 bytes for each token of context across the whole 43-layer stack. Against the baseline the report chooses, bf16 grouped-query attention with 8 KV heads at head dimension 128, which is 176,128 bytes per token and 172 GiB for the same sequence, Flash comes in at 1.96 percent. </p><p>The report claims <strong>approximately 2 percent</strong>. Derived independently, it holds.</p><p>Three multipliers produce that and only one of them is the headline. Fifty-one times from sequence-axis compression and the interleave, two times from sharing key and value, and slightly under two times from mixed FP8 and FP4 storage. </p><p><strong>Top-k selection, the mechanism everyone names when they describe this model, saves no memory at all</strong>. It saves bandwidth at read time. The memory win is compression, sharing and precision, in that order.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>The indexer is O(n), and that is the ceiling</h2><p>Here is what the efficiency section does not say. Top-k selection bounds the cost of attending. It does not bound the cost of selecting. To pick the <strong>best 512 of 262,144 compressed entries</strong>, the indexer scores all of them. That scan is linear in context length and it never becomes sparse.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!oHYg!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff9f3b59e-3f9b-4c48-bc29-9819d9b45f64_1600x922.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!oHYg!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff9f3b59e-3f9b-4c48-bc29-9819d9b45f64_1600x922.png 424w, https://substackcdn.com/image/fetch/$s_!oHYg!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff9f3b59e-3f9b-4c48-bc29-9819d9b45f64_1600x922.png 848w, https://substackcdn.com/image/fetch/$s_!oHYg!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff9f3b59e-3f9b-4c48-bc29-9819d9b45f64_1600x922.png 1272w, https://substackcdn.com/image/fetch/$s_!oHYg!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff9f3b59e-3f9b-4c48-bc29-9819d9b45f64_1600x922.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!oHYg!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff9f3b59e-3f9b-4c48-bc29-9819d9b45f64_1600x922.png" width="1456" height="839" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f9f3b59e-3f9b-4c48-bc29-9819d9b45f64_1600x922.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:839,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;The matmul side is flat by construction. The attention side is linear in context because of the indexer scan and the HCA dense pass.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="The matmul side is flat by construction. The attention side is linear in context because of the indexer scan and the HCA dense pass." title="The matmul side is flat by construction. The attention side is linear in context because of the indexer scan and the HCA dense pass." srcset="https://substackcdn.com/image/fetch/$s_!oHYg!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff9f3b59e-3f9b-4c48-bc29-9819d9b45f64_1600x922.png 424w, https://substackcdn.com/image/fetch/$s_!oHYg!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff9f3b59e-3f9b-4c48-bc29-9819d9b45f64_1600x922.png 848w, https://substackcdn.com/image/fetch/$s_!oHYg!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff9f3b59e-3f9b-4c48-bc29-9819d9b45f64_1600x922.png 1272w, https://substackcdn.com/image/fetch/$s_!oHYg!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff9f3b59e-3f9b-4c48-bc29-9819d9b45f64_1600x922.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>The matmul side is flat by construction. The attention side is linear in context because of the indexer scan and the HCA dense pass.</em></p><p>At 4K of context, attention is 9 percent of per-token arithmetic and the model behaves like a cheap 13B. At 32K it is 18 percent. At 128K it is 39 percent. <strong>Around 213K tokens</strong> the attention overtakes the entire 284B expert bank, and at 1M it is 82 percent of the work.</p><p><strong> The model is linear in context, not sub-linear.</strong> Compression changed the constant by roughly fifty. It did not change the exponent.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!rOt3!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F86ee4412-8b4d-4163-9535-4c522b82f00f_1600x867.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!rOt3!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F86ee4412-8b4d-4163-9535-4c522b82f00f_1600x867.png 424w, https://substackcdn.com/image/fetch/$s_!rOt3!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F86ee4412-8b4d-4163-9535-4c522b82f00f_1600x867.png 848w, https://substackcdn.com/image/fetch/$s_!rOt3!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F86ee4412-8b4d-4163-9535-4c522b82f00f_1600x867.png 1272w, https://substackcdn.com/image/fetch/$s_!rOt3!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F86ee4412-8b4d-4163-9535-4c522b82f00f_1600x867.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!rOt3!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F86ee4412-8b4d-4163-9535-4c522b82f00f_1600x867.png" width="1456" height="789" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/86ee4412-8b4d-4163-9535-4c522b82f00f_1600x867.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:789,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;At 32K of context, half the attention arithmetic is spent deciding what to attend to. At 1M, four fifths of it is.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="At 32K of context, half the attention arithmetic is spent deciding what to attend to. At 1M, four fifths of it is." title="At 32K of context, half the attention arithmetic is spent deciding what to attend to. At 1M, four fifths of it is." srcset="https://substackcdn.com/image/fetch/$s_!rOt3!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F86ee4412-8b4d-4163-9535-4c522b82f00f_1600x867.png 424w, https://substackcdn.com/image/fetch/$s_!rOt3!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F86ee4412-8b4d-4163-9535-4c522b82f00f_1600x867.png 848w, https://substackcdn.com/image/fetch/$s_!rOt3!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F86ee4412-8b4d-4163-9535-4c522b82f00f_1600x867.png 1272w, https://substackcdn.com/image/fetch/$s_!rOt3!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F86ee4412-8b4d-4163-9535-4c522b82f00f_1600x867.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>At 32K of context, half the attention arithmetic is spent deciding what to attend to. At 1M, four fifths of it is.</em></p><p><strong>Half the attention FLOPs at 32K are the indexer. Four fifths at 1M.</strong> The thing that makes attention sparse is the dominant cost of the attention.</p><blockquote><p><em>Four fifths of the cost of attention in this model is the cost of deciding what to attend to.</em><strong><span>Derived from the published dimensions</span></strong></p></blockquote><p>DeepSeek clearly knew. The indexer runs in FP4 while the main attention runs in FP8, which is the more aggressive precision going to the larger consumer. <strong>Flash&#8217;s top-k came down to 512 while Pro&#8217;s is 1024</strong>, because reducing k shrinks the attention pass and does nothing at all to the scan, and Flash needs the attention pass small relative to its smaller matmul side. </p><p>And the <strong>compression rate m = 4 </strong>is not primarily a memory optimisation: the scan runs over n/m keys, so compressing by four is a four times discount on selection, with the cache saving arriving as a side effect.</p><p>If someone finds a sub-linear selector, this architecture gets a second life. There is already a paper trying, from a group that fine-tuned V4-Flash with a <strong>Neural Memory Indexer </strong>that predicts and prefetches only the query-critical KV chunks, reporting comparable benchmark scores at 13.5 percent of the GPU memory. </p><p>Its authors are explicit that the work was constrained by resources and cut short, with the indexer trained on frozen keys and no end-to-end optimisation against the backbone. As a result it is <strong>a direction rather than a result.</strong> It is the right direction.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>FP4, down to the index scores</h2><p>The so-called<strong> Quantization-aware training</strong> is applied during post-training to two things: the MoE expert weights, and the query-key path of the CSA indexer, where activations are cached, loaded and multiplied entirely in FP4.</p><p>The expert-weight scheme has a property worth stating because it explains why the whole thing was affordable. Master weights are held in FP32, quantised to MXFP4, then dequantised back to FP8 for the actual computation, and <em><strong>the FP4 to FP8 dequantisation is lossless</strong></em>. </p><p><strong>FP8 in E4M3</strong> has two more exponent bits than FP4 in E2M1, so as long as the ratio between the largest and smallest scale factors of the FP4 sub-blocks, which are 1 by 32 tiles, inside a given FP8 quantisation block, which is 128 by 128, stays under a threshold, the finer scale information is absorbed entirely by the wider dynamic range. </p><p>DeepSeek verified their weights satisfy the condition. The consequence is that the <strong>entire QAT pipeline</strong> reuses the existing FP8 training framework without modification, with a straight-through estimator carrying gradients back to the FP32 masters, and no need to requantise transposed weights.</p><p>Then there is one number in that section that deserves its own paragraph. They also quantise the index scores themselves, the output of the lightning indexer, from FP32 to BF16. </p><p>That gives a <strong>two times speedup on the top-k selector while preserving a 99.7 percent recall rate of KV entries.</strong> Given that the selector is the dominant term in attention cost at long context, halving it for three tenths of a percent of recall is the highest-leverage line in the report.</p><p>During rollouts and any inference-only forward pass, including teachers and reference models,<em> real FP4 weights are used rather than simulated quantisation</em>, so sampling behaviour during RL is identical to deployment behaviour. </p><p>That is a correctness argument dressed as an efficiency one, and it matters: a policy trained against a simulated-quantisation rollout is optimising a model that will never be served.</p><div><hr></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><h2>The mega-kernel, and the balance point DeepSeek wants hardware to hit</h2><p>Expert parallelism needs an all-to-all dispatch and an all-to-all combine per MoE layer, and the <strong>conventional implementation</strong> runs communication and computation as separate serial kernels, which leaves both the interconnect and the SMs idle half the time. </p><p>DeepSeek&#8217;s answer fuses them into one pipelined kernel and then partitions the experts into waves. As soon as every expert in a wave has its tokens, that wave computes, while the next wave&#8217;s tokens are still in flight and the previous wave&#8217;s results are being sent back. In steady state all three proceed at once.</p><p>The <strong>reported gains are 1.50 to 1.73 times against strong non-fused baselines</strong> for general inference, and up to 1.96 times for latency-sensitive work like RL rollouts and high-speed agent serving, where batches are small and long-tailed and the pipeline has the most idle time to recover. </p><p>The report&#8217;s own figure puts the theoretical ceiling of the wave scheme at 1.92 times against 1.42 for Comet, which overlaps dispatch with the first linear and the second linear with combine but not at wave granularity, and it evaluates both in the V4-Flash configuration specifically. The implementation is open, as <strong>MegaMoE inside DeepGEMM</strong>, and it was validated on both NVIDIA GPUs and Huawei Ascend NPUs.</p><p>Then comes the paragraph I think is the most economically consequential in the entire report, and it is addressed to hardware vendors rather than to users.</p><p>Communication hides under computation when<strong> C/B is at most V<sub>comp</sub>/V<sub>comm</sub>,</strong> where C is peak compute and B is interconnect bandwidth. For a DeepSeekMoE layer each token-expert pair costs 6hd<sub>ff</sub> FLOPs across the gate, up and down projections, and 3h bytes of traffic, being h bytes of FP8 dispatch and 2h bytes of BF16 combine. </p><p>The h cancels. The condition collapses to:</p><pre><code>C / B  &lt;=  2 * d_ff</code></pre><p>The report evaluates this for Pro, whose <strong>expert intermediate dimension is 3072</strong>, and gets 6144 FLOPs per byte, then observes that once bandwidth clears that threshold it stops being the bottleneck and further silicon spent on it brings diminishing returns. Their recommendation to hardware designers is to target the balance point rather than scale bandwidth unconditionally.</p><blockquote><p><em>The report contains the equation that says the cheaper model is the harder one to host. It does not evaluate it for the cheaper model.</em><strong><span>Report section 3.1</span></strong></p></blockquote><p>Flash&#8217;s expert intermediate dimension is 2048. Run the same derivation and Flash&#8217;s balance point is 4096 FLOPs per byte, which is two thirds of Pro&#8217;s.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!5DGw!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc3fc06f6-13ba-4789-95b3-0f2ff5cec30c_1600x748.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!5DGw!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc3fc06f6-13ba-4789-95b3-0f2ff5cec30c_1600x748.png 424w, https://substackcdn.com/image/fetch/$s_!5DGw!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc3fc06f6-13ba-4789-95b3-0f2ff5cec30c_1600x748.png 848w, https://substackcdn.com/image/fetch/$s_!5DGw!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc3fc06f6-13ba-4789-95b3-0f2ff5cec30c_1600x748.png 1272w, https://substackcdn.com/image/fetch/$s_!5DGw!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc3fc06f6-13ba-4789-95b3-0f2ff5cec30c_1600x748.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!5DGw!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc3fc06f6-13ba-4789-95b3-0f2ff5cec30c_1600x748.png" width="1456" height="681" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c3fc06f6-13ba-4789-95b3-0f2ff5cec30c_1600x748.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:681,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Derived from the report's own condition at 4.5 PFLOP/s of dense FP8 per GPU. Smaller experts do less arithmetic per byte moved.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Derived from the report's own condition at 4.5 PFLOP/s of dense FP8 per GPU. Smaller experts do less arithmetic per byte moved." title="Derived from the report's own condition at 4.5 PFLOP/s of dense FP8 per GPU. Smaller experts do less arithmetic per byte moved." srcset="https://substackcdn.com/image/fetch/$s_!5DGw!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc3fc06f6-13ba-4789-95b3-0f2ff5cec30c_1600x748.png 424w, https://substackcdn.com/image/fetch/$s_!5DGw!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc3fc06f6-13ba-4789-95b3-0f2ff5cec30c_1600x748.png 848w, https://substackcdn.com/image/fetch/$s_!5DGw!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc3fc06f6-13ba-4789-95b3-0f2ff5cec30c_1600x748.png 1272w, https://substackcdn.com/image/fetch/$s_!5DGw!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc3fc06f6-13ba-4789-95b3-0f2ff5cec30c_1600x748.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>Derived from the report&#8217;s own condition at 4.5 PFLOP/s of dense FP8 per GPU. Smaller experts do less arithmetic per byte moved.</em></p><p><strong>The smaller model is the harder one to interconnect.</strong> Flash needs 1.5 times more bandwidth per unit of compute than Pro does, because its experts do less arithmetic for the same number of bytes dispatched and combined. </p><p>At a B200&#8217;s dense FP8 throughput, <strong>Flash wants about 1.10 TB/s per GPU</strong> to hide its all-to-all and Pro wants about 0.73. NVLink 5 covers both comfortably. A PCIe Gen5 box does not cover either, and misses Flash by a factor of seventeen.</p><p>Anyone sizing a deployment on the assumption that the cheaper model is the easier one to host has the relationship backwards, and the report contains the equation that says so.</p><p>Three other proposals in that section are worth recording because they are a roadmap. DeepSeek asks for more power headroom, on the grounds that <strong>extreme fusion drives compute</strong>, memory and network to high load simultaneously and power throttling becomes the limiter, which is a real and underdiscussed consequence of fusing everything. </p><p>They use pull-based communication, where each GPU reads from remote GPUs, because fine-grained push carries too much notification latency, and they ask for <strong>lower-latency cross-GPU signalling</strong> so push becomes viable. </p><p>And they <strong>propose replacing SwiGLU</strong> with a cheap elementwise activation with no exponential and no division, because that lightens post-GEMM work and, under a fixed parameter budget, removing the gate projection lets d<sub>ff</sub> grow, which pushes the balance point up and relaxes the bandwidth requirement further.</p><p>That last one is a description of the next model.<br></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>TileLang, and a theorem prover in the compiler</h2><p>Section 3.2 is the part of this report that a compiler person should read twice. DeepSeek&#8217;s architecture, written naively, <strong>decomposes into hundreds of fine-grained Torch ATen operators</strong>, and they replaced most of them with fused kernels written in TileLang, a tile-level DSL, rather than in CUDA.</p><p>They did two things to TileLang along the way are more interesting than the choice itself.</p><div class="callout-block" data-callout="true"><p><strong>Host codegen.</strong> As accelerators get faster, CPU-side orchestration becomes the ceiling for small kernels, and the usual source is host-side logic such as runtime contract checks written in Python for flexibility. DeepSeek co-generates the device kernel and a lightweight host launcher at the IR level, embedding data types, rank and shape constraints and stride and layout assumptions parsed from the frontend, then lowers the launcher to host source on TVM-FFI, whose compact calling convention and zero-copy tensor interop keep the overhead small. Validation and argument marshalling happen in generated C rather than in Python. Their measurement: <em>CPU-side validation drops from tens or hundreds of microseconds per invocation to under one</em>.</p><p>That is a two-order-of-magnitude reduction in a cost that most people do not measure at all, and it is the sort of thing that only shows up when your model has enough small kernels for launch overhead to dominate. Which this one does, by construction.</p></div><div class="callout-block" data-callout="true"><p><strong>Z3 in the algebraic system.</strong> TileLang kernels are full of complex tensor index arithmetic, and passes like layout inference, memory hazard detection and bound analysis all need to prove properties of integer expressions before they are allowed to fire. Weak integer reasoning means conservative passes means slower kernels. DeepSeek integrated the Z3 SMT solver into TileLang&#8217;s algebraic system, translating integer expressions into quantifier-free non-linear integer arithmetic, which handles the ordinary linear index algebra through ILP and the harder cases, such as vectorising over variable tensor shapes, through genuine non-linear reasoning. They report a few seconds of added compilation time and improvements across vectorisation, barrier insertion and simplification.</p></div><p><strong>A production LLM shipped with an SMT solver inside its kernel compiler. </strong>This is the direction I have argued the field goes: the bottleneck in kernel performance is not the language, it is how much the compiler can prove, and buying proving power off the shelf is cheaper than hand-writing the kernel.</p><p>The numerics policy in the same section is equally deliberate. <strong>Fast-math is disabled at the compiler level</strong> by default, precision-affecting approximations are opt-in frontend operators, and IEEE-compliant intrinsics with explicit rounding modes are available when strict semantics are required. </p><p>They also align TileLang&#8217;s algebraic simplification and lowering rules with <em>NVCC </em>so that kernels can be validated<strong> bit-for-bit against hand-written CUDA baselines</strong>, with layout annotations available to pin down lowering decisions and hold accumulation order constant. Accuracy by default, speed by opt-in, which is the opposite of the usual arrangement.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share&quot;,&quot;text&quot;:&quot;Share The Software Frontier&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share The Software Frontier</span></a></p><div><hr></div><h2>Bitwise batch invariance, and why they wanted it</h2><p>Batch invariance means a given token&#8217;s output is bitwise identical regardless of where it sits in a batch. Almost nobody ships this, because it costs performance, and <strong>DeepSeek&#8217;s reasoning for paying</strong> is that they wanted bitwise alignment across pre-training, post-training and inference, which makes loss spikes diagnosable and post-training behaviour consistent with what gets served.</p><p>Getting it required giving up two standard optimisations and then engineering the loss back out.</p><blockquote><p><strong>Attention.</strong> Split-KV, which spreads one sequence&#8217;s attention across many SMs to balance load, is not batch invariant. Abandoning it causes wave quantisation, where the final partially filled wave of thread blocks leaves most of the GPU idle. DeepSeek&#8217;s answer is a dual-kernel decode: a first kernel computes an entire sequence&#8217;s attention inside a single SM, giving throughput on fully occupied waves, and a second kernel spreads one sequence across multiple SMs to shorten the trailing partial wave. The two are engineered to have the <em>same accumulation order</em> so their outputs are bit-identical, and the second uses distributed shared memory within thread-block clusters to exchange partial results across SMs at speed. The reported overhead of batch-invariant decoding after this is negligible.</p></blockquote><blockquote><p><strong>Matrix multiplication.</strong> cuBLAS cannot be made batch invariant, so it is replaced end to end by DeepGEMM. Split-k, which is how you get performance at very small batch, is also not batch invariant, so it is dropped in most scenarios and the resulting loss is recovered by other means.</p></blockquote><p>Determinism is a separate problem from batch invariance and it comes from accumulation order, usually via atomic addition in the backward pass. </p><p>Sparse attention backward normally uses atomicAdd to accumulate KV gradients, so they <strong>allocate a separate accumulation buffer per SM</strong> and do a global deterministic summation afterwards. MoE backward is non-deterministic because SMs from different ranks negotiate write positions into the same receiving buffer, so they pre-process token order within each rank and isolate buffers across ranks. </p><p>And the <strong>mHC GEMM with its output dimension of 24 is small enough that split-k is unavoidable</strong>, so each split is emitted separately and reduced deterministically in a following kernel.</p><p>There is a payoff for this that shows up two sections later, in the rollout service. Because generation can be preempted at any time on their cluster, they keep a token-granular write-ahead log per request and resume from it. </p><p>The report explains why <strong>they cannot simply regenerate interrupted requests </strong>from scratch, and the argument is a good one: shorter responses are more likely to survive an interruption, so regenerating the survivors biases the training distribution towards short outputs. </p><p>They note that a batch-invariant deterministic stack could fix this instead by regenerating with a consistent sampler seed, but that this still costs a <strong>full re-decode</strong>, so the log wins. Batch invariance is what makes that alternative even expressible.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Two cache hierarchies, and lcm(4, 128)</h2><p>The <strong>hybrid attention breaks the assumption PagedAttention is built on</strong>, which is that every layer&#8217;s KV state has the same shape and the same eviction policy. </p><p>In V4 the compression ratios differ per layer, the indexer carries its own embedding size, the sliding-window layers have their own hit and eviction rules, and there is a <strong>rolling residual of tokens</strong> not yet numerous enough to compress that has to live somewhere.</p><p>DeepSeek splits it in two. A classical paged KV cache holds the compressed CSA and HCA entries. A separate <em>state cache</em> holds the sliding-window entries and the uncompressed tail, on the argument that both are a function only of the current position, which makes them a<strong> state-space model rather than a growing history</strong>, so a fixed-size pool can be pre-allocated and assigned per sequence.</p><p>The block geometry falls out of the two compression rates. A block has to cover a whole number of compressed entries in every layer, so it must span a multiple of the least common multiple of m and m&#8217;. </p><p><strong>For Flash that is lcm(4, 128) = 128 original tokens</strong>, giving 32 CSA entries and exactly one HCA entry per block. vLLM independently chose 256 native positions, which is two of DeepSeek&#8217;s minimum blocks, giving 64 c4a entries and 2 c128a entries.</p><p><strong>vLLM&#8217;s implementation notes</strong> are the best serving document published on this model, and their three decisions are worth having next to DeepSeek&#8217;s. </p><p>One logical block size in native token positions for every compressed layer, so slot mapping, scheduler accounting and prefix-hit detection use one unit instead of branching on the compression ratio. The compressor&#8217;s rolling residual registered as<strong> sliding-window KV with sliding_window set to the compression stride</strong>, rather than as a side buffer, so prefix caching lands on block boundaries and disaggregated prefill ships it through the existing SWA transfer path instead of a second one. </p><p>And a page-size argument: <strong>page size is block_size times compress_ratio times entry_size, all three are controllable</strong>, and chosen carefully the five cache kinds collapse into three buckets, each backed by one pool, sized once at load, with no runtime repartitioning and no cross-kind fragmentation.</p><p>On the kernel side vLLM fuses three groups: compressor with RMSNorm, RoPE and cache insertion, all elementwise, for 1.4 to 3 times; inverse RoPE with <strong>FP8 quantisation ahead of the output projection, for 2 to 3 times</strong>; and a horizontal fusion of query normalisation, KV RoPE and sliding-window key insertion using static warp-ID dispatch, each warp working independently on a query head or a key head with no cross-warp communication, for 10 to 20 times over the naive version. </p><p>The <strong>indexer then runs on its own CUDA stream </strong>alongside KV compression and window insertion, worth 5 to 6 percent end to end at low batch.</p><p>SGLang went at the FP4 weights instead, pairing MXFP8 activations with MXFP4 expert weights through <strong>FlashInfer&#8217;s TRTLLM-Gen fused MoE backend, and splitting K across CTAs</strong> in the mHC pre-GEMM for exactly the small-batch parallelism problem described earlier.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/p/deepseek-v4-flash-the-cost-of-deciding?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/p/deepseek-v4-flash-the-cost-of-deciding?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><div><hr></div><h2>What a cache hit actually costs</h2><p>DeepSeek stores all compressed CSA and HCA entries to disk. </p><p>When a request hits a stored prefix, <strong>those entries are read back rather than recomputed</strong>, up to the last complete compression block; the tail of an incomplete block still has to be recomputed because uncompressed entries are not stored.</p><p>The sliding-window state is the problem. It is uncompressed and it exists in every layer, so storing it for every token would be roughly eight times the volume of everything else. </p><p>The report gives <strong>precisely three strategies</strong> with different trade-offs, and the third is the one that makes the economics work.</p><ol><li><p><strong>Full SWA caching</strong> stores everything, so a hit reads the last n_win tokens of the prefix and recomputes nothing. Zero redundancy, but only a sliver of what was written is ever read, which is a write-heavy unbalanced access pattern that SSDs handle badly.</p></li><li><p><strong>Periodic checkpointing</strong> saves the window state every p tokens, loads the nearest checkpoint on a hit and recomputes the tail, with p tuning the storage against compute trade.</p></li><li><p><strong>Zero SWA caching</strong> stores none of it. Here is the argument, and it is the neatest piece of reasoning in the report. Each token&#8217;s sliding-window entry in a given layer depends only on the window entries of the previous layer, which span n_win tokens. So the dependency cone going back through L layers is exactly n_win times L tokens wide. Recompute that many tokens and the entire window state is restored.</p></li></ol><blockquote><p><em>A million-token prefix comes back for the price of five and a half thousand.</em><strong><span>Report section 3.6.2, zero SWA caching</span></strong></p></blockquote><p>For Flash, n_win times L is 128 times 43, which is 5,504 tokens. Restoring the full sliding-window state of a one-million-token prefix costs a 5,504-token recompute, about 140 TFLOP, against the 26.6 PFLOP a cold prefill of that prefix would cost. <strong>That is a factor of 191, and it is why a cache hit can be sold for a fiftieth of a cache miss.</strong></p><p>The whole thing only works because the compressed state is small. Storing 172 GiB per session on disk is possible, but reading it back at request time competes with the prefill it was meant to replace. At 3.62 GB it does not. </p><p>Compressing the sequence axis is <strong>what turns disk from a bad idea into the cheapest tier in the hierarchy</strong>, and the $0.0028 line on the rate card is the direct commercial expression of that.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/p/deepseek-v4-flash-the-cost-of-deciding?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/p/deepseek-v4-flash-the-cost-of-deciding?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><div><hr></div><h2>Post-training, and the parts of it that are cost mechanisms</h2><p>The pipeline broadly follows <em>DeepSeek-V3.2 with one substitution the report calls critical</em>: the mixed reinforcement learning stage is replaced entirely by on-policy distillation. </p><p>Domain specialists are trained separately, each through supervised fine-tuning then<strong> GRPO against domain-specific rewards</strong>, and then more than ten of them are merged into one student by having the student sample its own trajectories and minimise reverse KL against the relevant teacher.</p><p>They use full-vocabulary logit distillation rather than the usual token-level KL estimate, on the grounds that the cheap estimator has high gradient variance and destabilises training. Making that affordable took two tricks. </p><p>Teacher weights live in centralised distributed storage and are loaded on demand with <strong>ZeRO-like sharding. </strong></p><p>And rather than materialising logits for a vocabulary above 100k across more than ten teachers, they cache only the last-layer teacher hidden states and <strong>reconstruct logits on the fly through the prediction head at training time</strong>, ordering training samples by teacher index so exactly one teacher head is resident on device at a time. The KL itself is a TileLang kernel.</p><p>Three things from the post-training section have direct consequences for what you pay.</p><h3>The reasoning ladder is a context ladder</h3><p>Three modes, trained as separate RL configurations with distinct length penalties and context windows, then unified. Non-think evaluates at 8K of context,<strong> Think High at 128K, Think Max at 384K.</strong> Non-think emits an empty reasoning block and goes straight to the summary. </p><p>Think Max additionally prepends a system instruction, printed verbatim in the report, which tells the model that shortcuts are not permitted and that it must document every intermediate step, considered alternative and rejected hypothesis.</p><p>That instruction is <strong>why Artificial Analysis</strong> measures this model generating 210M output tokens across its index against a 100M median. The verbosity is not a training accident. </p><p><strong>It is an instruction, in the system prompt, that DeepSeek wrote and that you are billed for at $0.28 per million. </strong>Anyone running Flash at max effort and complaining about token burn is paying for a behaviour they asked for by name.</p><h3>The tool schema is XML on purpose</h3><p>V4 introduces a tool-call format built on a dedicated <span>|DSML|</span> token with XML-shaped invocations rather than JSON. </p><p>String parameters go through as-is with an explicit <span>string=&#8221;true&#8221;</span> flag; <strong>everything else is JSON-encoded with the flag false. </strong>The stated reason is that XML mitigates escaping failures and reduces tool-call errors.</p><p>This is a small decision with a large downstream effect. Every escaping failure in a JSON tool call is a wasted turn, and a wasted turn in an agent loop costs a full round of prefill plus decode. Reducing tool-call error rate is a cost reduction that never appears on a rate card.</p><h3>Interleaved thinking, and a warning inside it</h3><p>V3.2 kept reasoning traces across tool-result rounds but discarded them when a new user message arrived. V4 keeps everything, across user message boundaries, for <strong>tool-calling conversations</strong>, so a long-horizon agent maintains one cumulative chain of thought instead of reconstructing its state each turn. </p><p>General conversation keeps the old discarding behaviour, on the reasonable grounds that persistent traces buy little there and cost context.</p><p>The <strong>warning is in the same paragraph </strong>and it is easy to miss. Agent frameworks that simulate tool interactions through user messages, and the report names Terminus, will not trigger the tool-calling context path and therefore will not get the persistence. </p><p>DeepSeek&#8217;s own recommendation for those frameworks is to use non-think models. If you are <strong>benchmarking Flash inside a harness </strong>that fakes tools as user turns, you are measuring the wrong path.</p><h3>Quick Instruction, which is an inference-economics feature wearing a post-training costume</h3><p>In a chat product, a handful of auxiliary decisions run before the real response: <strong>whether to trigger a web search</strong>, what the query should be, how authoritative a source needs to be, what domain the request belongs to, whether a pasted URL should be fetched. </p><p>The standard answer is a separate small model, which means a second prefill of the same prompt because it cannot share the big model&#8217;s KV cache.</p><p>DeepSeek trained special tokens for each of those tasks and appends them to the input sequence directly. The<strong> auxiliary task runs on the already-computed KV cache. </strong>There is no second prefill, several of the tasks run in parallel, the user-perceived time to first token drops, and there is no small model to maintain.</p><p>The published tokens are <strong><span>|action|</span>, <span>|title|</span>, <span>|query|</span>, <span>|authority|</span>, <span>|domain|</span>, <span>|extracted_url|</span> and <span>|read_url|</span></strong>. Read the list and it is obvious this is DeepSeek&#8217;s own chat product spilling into the model card, which is exactly what makes it interesting. </p><p>It is a vertical integration of the router into the weights, and it deletes a whole class of serving infrastructure that everyone else runs.</p><h3>The sandbox, briefly</h3><p>Agentic RL needs somewhere to execute, and DeepSeek built a platform they call DSec:<strong> three Rust components</strong> on top of their 3FS distributed filesystem, running hundreds of thousands of concurrent sandbox instances per cluster. </p><ul><li><p>One Python SDK abstracts four execution substrates behind one API, switchable by a parameter. </p></li><li><p>Function Call dispatches stateless invocations to a pre-warmed pool with no cold start. </p></li><li><p>Container is Docker-compatible with EROFS on-demand image loading. MicroVM is Firecracker for security-sensitive high-density work. </p></li><li><p>FullVM is QEMU for arbitrary guest operating systems. </p></li><li><p>Base images sit on 3FS as read-only layers shared across instances, writes go to a local copy-on-write layer, and snapshots chain, which gets them millisecond-scale resumption.</p></li></ul><p>Each sandbox keeps a globally ordered trajectory log of every command and result, which serves three purposes: fast-forwarding after a preemption by <strong>replaying cached results rather than re-executing non-idempotent commands</strong>, provenance for every state change, and deterministic replay of any historical session.</p><p>None of this is in the model. All of it is why the model has agent scores.</p><div class="community-chat" data-attrs="{&quot;url&quot;:&quot;https://open.substack.com/pub/softwarefrontier/chat?utm_source=chat_embed&quot;,&quot;subdomain&quot;:&quot;softwarefrontier&quot;,&quot;pub&quot;:{&quot;id&quot;:3575776,&quot;name&quot;:&quot;The Software Frontier&quot;,&quot;author_name&quot;:&quot;Lorenzo Bradanini&quot;,&quot;author_photo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!ACM6!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F18342bf1-31cb-404a-b9e1-998a38d299bf_1200x1200.jpeg&quot;}}" data-component-name="CommunityChatRenderPlaceholder"></div><div><hr></div><h2>What a node holds</h2><p>Now the economics, and they start with capacity rather than speed.</p><p><strong>Four B200 at 180 GB usable is 720 GB of HBM</strong>. The weights ship natively mixed: FP4 for routed experts, FP8 for attention, norms and router. That is 139 GB plus 6 GB, call it 145 GB resident. </p><p>At 85 percent HBM utilisation, leaving room for activations, the four-times-wider mHC residual stream and the compressor states, the KV pool is roughly 467 GB.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!yz23!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F375bb791-df0e-4820-9b0a-07b8547e8d77_1600x882.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!yz23!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F375bb791-df0e-4820-9b0a-07b8547e8d77_1600x882.png 424w, https://substackcdn.com/image/fetch/$s_!yz23!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F375bb791-df0e-4820-9b0a-07b8547e8d77_1600x882.png 848w, https://substackcdn.com/image/fetch/$s_!yz23!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F375bb791-df0e-4820-9b0a-07b8547e8d77_1600x882.png 1272w, https://substackcdn.com/image/fetch/$s_!yz23!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F375bb791-df0e-4820-9b0a-07b8547e8d77_1600x882.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!yz23!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F375bb791-df0e-4820-9b0a-07b8547e8d77_1600x882.png" width="1456" height="803" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/375bb791-df0e-4820-9b0a-07b8547e8d77_1600x882.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:803,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;The same node, the same money, the same power draw. The difference is which axis was compressed.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="The same node, the same money, the same power draw. The difference is which axis was compressed." title="The same node, the same money, the same power draw. The difference is which axis was compressed." srcset="https://substackcdn.com/image/fetch/$s_!yz23!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F375bb791-df0e-4820-9b0a-07b8547e8d77_1600x882.png 424w, https://substackcdn.com/image/fetch/$s_!yz23!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F375bb791-df0e-4820-9b0a-07b8547e8d77_1600x882.png 848w, https://substackcdn.com/image/fetch/$s_!yz23!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F375bb791-df0e-4820-9b0a-07b8547e8d77_1600x882.png 1272w, https://substackcdn.com/image/fetch/$s_!yz23!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F375bb791-df0e-4820-9b0a-07b8547e8d77_1600x882.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>The same node, the same money, the same power draw. The difference is which axis was compressed.</em></p><p><strong>129 concurrent one-million-token sessions on one four-GPU node</strong>. A V3.2-style stack on the same hardware holds five. At 128K, the working length of most real agent traffic, it is 1,031.</p><p>This is the actual product. Not the million-token window as a marketing number, but the ability to keep a thousand long-lived agent sessions warm on one node without evicting anyone. </p><p>Eviction is what makes long-context serving expensive, because every eviction is a <strong>re-prefill a</strong>nd a re-prefill of 128K tokens costs more than the entire conversation that followed it. Compressing the sequence axis turns that from a scheduling problem into a non-problem.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>The roofline, and what $0.28 buys</h2><p>Decode on a large MoE is bandwidth-bound. At any batch deep enough to hit every expert, the node streams the full expert bank once per step: <strong>278.1B parameters at FP4 is 139 GB</strong>, plus the dense stack replicated across four data-parallel ranks, giving 164 GB of weight traffic per step against 32 TB/s of aggregate bandwidth. A 5.12 ms floor, or 195 steps per second.</p><p>Artificial Analysis measures 122.7 output tokens per second per user on DeepSeek&#8217;s own API, which is 8.15 ms per step. <strong>The model achieves 63 percent of peak HBM bandwidth.</strong> That is a good number for something with this much elementwise work between matmuls, and it is a direct vindication of the fusion work in both vLLM and DeepSeek&#8217;s own kernels. </p><p>Holding that efficiency and adding KV read traffic gives node throughput, and node throughput times $0.28 per million gives revenue.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!EU1M!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2848795e-f386-486d-b1ba-eadeb90eb0c5_1600x862.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!EU1M!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2848795e-f386-486d-b1ba-eadeb90eb0c5_1600x862.png 424w, https://substackcdn.com/image/fetch/$s_!EU1M!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2848795e-f386-486d-b1ba-eadeb90eb0c5_1600x862.png 848w, https://substackcdn.com/image/fetch/$s_!EU1M!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2848795e-f386-486d-b1ba-eadeb90eb0c5_1600x862.png 1272w, https://substackcdn.com/image/fetch/$s_!EU1M!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2848795e-f386-486d-b1ba-eadeb90eb0c5_1600x862.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!EU1M!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2848795e-f386-486d-b1ba-eadeb90eb0c5_1600x862.png" width="1456" height="784" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2848795e-f386-486d-b1ba-eadeb90eb0c5_1600x862.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:784,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Break-even GPU rate from output revenue alone. The gold band is the rentable B200 market, roughly $3.35 reserved to $6.35 median as of August 2026.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Break-even GPU rate from output revenue alone. The gold band is the rentable B200 market, roughly $3.35 reserved to $6.35 median as of August 2026." title="Break-even GPU rate from output revenue alone. The gold band is the rentable B200 market, roughly $3.35 reserved to $6.35 median as of August 2026." srcset="https://substackcdn.com/image/fetch/$s_!EU1M!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2848795e-f386-486d-b1ba-eadeb90eb0c5_1600x862.png 424w, https://substackcdn.com/image/fetch/$s_!EU1M!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2848795e-f386-486d-b1ba-eadeb90eb0c5_1600x862.png 848w, https://substackcdn.com/image/fetch/$s_!EU1M!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2848795e-f386-486d-b1ba-eadeb90eb0c5_1600x862.png 1272w, https://substackcdn.com/image/fetch/$s_!EU1M!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2848795e-f386-486d-b1ba-eadeb90eb0c5_1600x862.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>Break-even GPU rate from output revenue alone. The gold band is the rentable B200 market, roughly $3.35 reserved to $6.35 median as of August 2026.</em></p><p>At 32K of context you need a sustained decode batch of about 256 before <strong>$0.28 per million covers a mid-market B200 at $5.50 per GPU-hour.</strong> At 512 you clear 57 percent gross margin, at 1024 you clear 77. </p><p>At 1M of context it never clears: the KV pool caps the batch at 129 sequences and the break-even rate there is $2.92 per GPU-hour, below the cheapest reserved B200 anyone publishes.</p><p>That is the analysis I would&#8217;ve published if I had stopped there, and it is wrong in two ways that point in opposite directions. It prices only output tokens, which understates the revenue. And <strong>it prices GPUs at rental rates</strong>, which overstates the cost for the only company that matters here.</p><h3>Nobody buys only output tokens</h3><p>Real traffic has a shape. A coding agent sends tens of thousands of input tokens for every thousand it gets back, most of the input is a prefix it sent before, and DeepSeek charges for all three streams at three different prices. </p><p>A <strong>node has one time budget </strong>and has to split it between prefilling input and decoding output, so the sustainable output rate falls as the input ratio rises while total revenue climbs.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!q8Ak!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd8d8f28e-d3d1-4f3e-89a1-58f899f5d4e0_1600x894.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!q8Ak!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd8d8f28e-d3d1-4f3e-89a1-58f899f5d4e0_1600x894.png 424w, https://substackcdn.com/image/fetch/$s_!q8Ak!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd8d8f28e-d3d1-4f3e-89a1-58f899f5d4e0_1600x894.png 848w, https://substackcdn.com/image/fetch/$s_!q8Ak!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd8d8f28e-d3d1-4f3e-89a1-58f899f5d4e0_1600x894.png 1272w, https://substackcdn.com/image/fetch/$s_!q8Ak!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd8d8f28e-d3d1-4f3e-89a1-58f899f5d4e0_1600x894.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!q8Ak!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd8d8f28e-d3d1-4f3e-89a1-58f899f5d4e0_1600x894.png" width="1456" height="814" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d8d8f28e-d3d1-4f3e-89a1-58f899f5d4e0_1600x894.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:814,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;32K context, batch 1024, node time split between prefill and decode. Cache hits consume neither prefill nor decode, only a disk read.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="32K context, batch 1024, node time split between prefill and decode. Cache hits consume neither prefill nor decode, only a disk read." title="32K context, batch 1024, node time split between prefill and decode. Cache hits consume neither prefill nor decode, only a disk read." srcset="https://substackcdn.com/image/fetch/$s_!q8Ak!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd8d8f28e-d3d1-4f3e-89a1-58f899f5d4e0_1600x894.png 424w, https://substackcdn.com/image/fetch/$s_!q8Ak!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd8d8f28e-d3d1-4f3e-89a1-58f899f5d4e0_1600x894.png 848w, https://substackcdn.com/image/fetch/$s_!q8Ak!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd8d8f28e-d3d1-4f3e-89a1-58f899f5d4e0_1600x894.png 1272w, https://substackcdn.com/image/fetch/$s_!q8Ak!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd8d8f28e-d3d1-4f3e-89a1-58f899f5d4e0_1600x894.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>32K context, batch 1024, node time split between prefill and decode. Cache hits consume neither prefill nor decode, only a disk read.</em></p><p>Output-only pricing understates this card by roughly a factor of two. At a 20-to-1 input ratio with no caching, break-even is $55 per GPU-hour rather than $28. </p><p><strong>At 200-to-1 with a 90 percent hit rate, which is what a long-running agent on a stable repository actually looks like, it is $64</strong>. Every one of those numbers is an order of magnitude above the rental market.</p><p>Cache hits are the mechanism. They consume no prefill and no decode, only a disk read of <strong>compressed entries</strong> plus a 5,504-token recompute, so raising the hit rate raises sustainable throughput without raising cost. </p><p>The $0.0028 price is low because the marginal cost is low, and it is worth having because it makes the traffic denser.</p><h3>DeepSeek is not renting these GPUs</h3><p>A rental rate contains a lessor&#8217;s margin, a scarcity premium that has been visibly volatile all year, and the lessor&#8217;s own financing cost. DeepSeek owns its fleet, and an owner&#8217;s cost is amortisation plus power.</p><p>Take a B200 at roughly $38,000 of capex, three years of life, 80 percent utilisation: $1.81 per GPU-hour. Add a kilowatt at PUE 1.3 and eight cents a kilowatt-hour: ten cents. </p><p>Call it $1.91 per GPU-hour all in, against a rental band of $3.35 to $6.35. <strong>An owner&#8217;s floor sits somewhere between 1.8 and 3.3 times below the price of renting the same silicon</strong><mark data-color="rgb(244, 229, 189)" style="background-color: rgb(244, 229, 189); color: rgb(0, 0, 0);">.</mark></p><p>Note what power is and is not in that sum. A four-GPU node draws about 5.2 kW including overhead, which costs 42 cents an hour, which is under two percent of what the same node costs to rent. Power is not a meaningful line in the financial model. </p><p>It is a meaningful line in the physical one, and the report says so: the same paragraph that asks hardware vendors for a balance point also asks for more power headroom, because <strong>extreme kernel fusion drives compute</strong>, memory and network to high load simultaneously and power throttling becomes the limiter. Fusing everything is how you become thermally bound rather than bandwidth bound.</p><h3>Which changes the verdict on the million-token window</h3><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!fBIg!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc68cd709-ac92-4feb-8935-aab8909feb74_1600x838.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!fBIg!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc68cd709-ac92-4feb-8935-aab8909feb74_1600x838.png 424w, https://substackcdn.com/image/fetch/$s_!fBIg!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc68cd709-ac92-4feb-8935-aab8909feb74_1600x838.png 848w, https://substackcdn.com/image/fetch/$s_!fBIg!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc68cd709-ac92-4feb-8935-aab8909feb74_1600x838.png 1272w, https://substackcdn.com/image/fetch/$s_!fBIg!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc68cd709-ac92-4feb-8935-aab8909feb74_1600x838.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!fBIg!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc68cd709-ac92-4feb-8935-aab8909feb74_1600x838.png" width="1456" height="763" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c68cd709-ac92-4feb-8935-aab8909feb74_1600x838.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:763,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;The break-even line is what $0.28 per million pays at 1M context with the batch capped at 129 by the KV pool.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="The break-even line is what $0.28 per million pays at 1M context with the batch capped at 129 by the KV pool." title="The break-even line is what $0.28 per million pays at 1M context with the batch capped at 129 by the KV pool." srcset="https://substackcdn.com/image/fetch/$s_!fBIg!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc68cd709-ac92-4feb-8935-aab8909feb74_1600x838.png 424w, https://substackcdn.com/image/fetch/$s_!fBIg!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc68cd709-ac92-4feb-8935-aab8909feb74_1600x838.png 848w, https://substackcdn.com/image/fetch/$s_!fBIg!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc68cd709-ac92-4feb-8935-aab8909feb74_1600x838.png 1272w, https://substackcdn.com/image/fetch/$s_!fBIg!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc68cd709-ac92-4feb-8935-aab8909feb74_1600x838.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>The break-even line is what $0.28 per million pays at 1M context with the batch capped at 129 by the KV pool.</em></p><p>Serving a million-token context at $0.28 per million output tokens needs a cost basis under $2.92 per GPU-hour. Nobody renting Blackwell has one. DeepSeek does, with room.</p><blockquote><p><em>The million-token window is not a loss leader. It is a moat, and a precisely shaped one. </em><strong><span>Owner economics against rental economics</span></strong></p></blockquote><p>So the million-token window is not a loss leader. It is a moat, and a precisely shaped one. The twenty providers currently reselling the preview weights on <strong>OpenRouter at 37 percent below DeepSeek&#8217;s</strong> own card can do that at short context, where deep batches make the arithmetic work on rented hardware. </p><p>They <strong>structurally cannot do it at a million tokens</strong>, at that price, on rented Blackwell, because the KV pool caps the batch and the capped batch does not clear the rent. Which is presumably why none of them advertise it.</p><p>An open-weights MIT model whose most expensive capability is only economic for an owner-operator is a strange and rather elegant object. DeepSeek gave away the architecture, the <strong>mega-kernel and the kernel library,</strong> and kept the one thing that does not fit in a repository: a fleet, a request volume large enough to make the on-disk cache pay, and a cost basis a third of what anyone else can rent.</p><h3>What the whole card is betting on</h3><p>Put the three prices next to the three cost structures and the strategy is legible. Prefill is compute-bound and enormously profitable: at<strong> 35 percent MFU a node prefills roughly 485,000 tokens per second</strong>, which at $0.14 per million is $244 an hour against a node that costs $22 to rent and $8 to own. Decode is bandwidth-bound and needs deep batches. </p><p>Cache hits cost almost nothing and are priced almost at nothing, which makes them a throughput multiplier rather than a revenue line.</p><p>The card rewards exactly one traffic shape: high input-to-output ratio, high prefix reuse, deep concurrency, moderate context. </p><p>That is coding-agent and tool-calling traffic, which is what the 0731 post-training pass targeted, what the<em> native Responses API and Codex adaptation are for, and what Quick Instruction and interleaved thinking were built to make cheaper.</em> The architecture, the post-training, the serving stack and the price list are one artefact pointed at one workload.</p><p>Point a different workload at it, <strong>long single-turn document analysis with no prefix reuse</strong> and shallow concurrency, and the margin thins toward nothing even for DeepSeek. The card does not distinguish, yet. I do not expect that to last.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Reading the benchmarks in full</h2><p>The report&#8217;s summary of its own results is accurate and selective, and the difference is worth spending a section on.</p><p>Start with the base models, which are <strong>the cleanest comparison in the document </strong>because all three ran in one internal harness with identical settings. </p><p>V4-Flash-Base carries 13B activated against V3.2-Base&#8217;s 37B, and 284B total against 671B. It should lose. It mostly does not.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!GNaW!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10f7e37b-1aa1-4f20-893d-5c23df20b66a_1600x864.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!GNaW!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10f7e37b-1aa1-4f20-893d-5c23df20b66a_1600x864.png 424w, https://substackcdn.com/image/fetch/$s_!GNaW!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10f7e37b-1aa1-4f20-893d-5c23df20b66a_1600x864.png 848w, https://substackcdn.com/image/fetch/$s_!GNaW!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10f7e37b-1aa1-4f20-893d-5c23df20b66a_1600x864.png 1272w, https://substackcdn.com/image/fetch/$s_!GNaW!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10f7e37b-1aa1-4f20-893d-5c23df20b66a_1600x864.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!GNaW!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10f7e37b-1aa1-4f20-893d-5c23df20b66a_1600x864.png" width="1456" height="786" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/10f7e37b-1aa1-4f20-893d-5c23df20b66a_1600x864.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:786,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Report Table 1. Same harness, same settings. Flash carries 35 percent of V3.2's activated parameters and 42 percent of its total.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Report Table 1. Same harness, same settings. Flash carries 35 percent of V3.2's activated parameters and 42 percent of its total." title="Report Table 1. Same harness, same settings. Flash carries 35 percent of V3.2's activated parameters and 42 percent of its total." srcset="https://substackcdn.com/image/fetch/$s_!GNaW!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10f7e37b-1aa1-4f20-893d-5c23df20b66a_1600x864.png 424w, https://substackcdn.com/image/fetch/$s_!GNaW!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10f7e37b-1aa1-4f20-893d-5c23df20b66a_1600x864.png 848w, https://substackcdn.com/image/fetch/$s_!GNaW!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10f7e37b-1aa1-4f20-893d-5c23df20b66a_1600x864.png 1272w, https://substackcdn.com/image/fetch/$s_!GNaW!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10f7e37b-1aa1-4f20-893d-5c23df20b66a_1600x864.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>Report Table 1. Same harness, same settings. Flash carries 35 percent of V3.2&#8217;s activated parameters and 42 percent of its total.</em></p><blockquote><p>Plus 6.8 on FACTS Parametric, plus 6.7 on HumanEval, plus 4.5 on LongBench-V2, plus 3.5 on MultiLoKo, plus 2.8 on MMLU-Pro. </p></blockquote><p>Getting more world knowledge out of fewer total parameters is the surprising one, because knowledge retention is supposed to scale with parameter count and Flash has 42 percent of them.</p><p>And then two regressions that nobody discussing this model has mentioned. <strong>BigCodeBench falls 7.1 points, from 63.9 to 56.8</strong>. MATH falls 3.1, from 60.5 to 57.4. Those are not noise, they are the two largest deltas in the table after FACTS, and one of them is a coding benchmark on a model being marketed for coding agents. </p><p><strong>Post-training clearly recovers a great deal of it</strong>, since the instructed model&#8217;s code-agent scores are strong. But the base model is worse at BigCodeBench than its predecessor and the report does not discuss why.</p><p>My guess, and it is a guess: BigCodeBench is a three-shot benchmark on library-heavy code, which is a knowledge task about API surfaces rather than a reasoning task, and it is the <strong>sort of long-tail recall that a smaller expert bank should hurt. </strong></p><p>The world-knowledge gains going the other way argue against that, which is why it stays a guess.</p><p>Then the long-context claim. The report&#8217;s own summary says V4-Pro-Max delivers strong results with a one-million-token window, <em>&#8220;surpassing even Gemini-3.1-Pro on academic benchmarks.&#8221; </em>That is totally true.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!q6i3!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a4cc0b7-6824-44c9-a006-558630822962_1600x718.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!q6i3!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a4cc0b7-6824-44c9-a006-558630822962_1600x718.png 424w, https://substackcdn.com/image/fetch/$s_!q6i3!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a4cc0b7-6824-44c9-a006-558630822962_1600x718.png 848w, https://substackcdn.com/image/fetch/$s_!q6i3!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a4cc0b7-6824-44c9-a006-558630822962_1600x718.png 1272w, https://substackcdn.com/image/fetch/$s_!q6i3!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a4cc0b7-6824-44c9-a006-558630822962_1600x718.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!q6i3!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a4cc0b7-6824-44c9-a006-558630822962_1600x718.png" width="1456" height="653" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/9a4cc0b7-6824-44c9-a006-558630822962_1600x718.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:653,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Report Table 6, DeepSeek's own numbers, standardised configuration across models.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Report Table 6, DeepSeek's own numbers, standardised configuration across models." title="Report Table 6, DeepSeek's own numbers, standardised configuration across models." srcset="https://substackcdn.com/image/fetch/$s_!q6i3!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a4cc0b7-6824-44c9-a006-558630822962_1600x718.png 424w, https://substackcdn.com/image/fetch/$s_!q6i3!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a4cc0b7-6824-44c9-a006-558630822962_1600x718.png 848w, https://substackcdn.com/image/fetch/$s_!q6i3!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a4cc0b7-6824-44c9-a006-558630822962_1600x718.png 1272w, https://substackcdn.com/image/fetch/$s_!q6i3!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a4cc0b7-6824-44c9-a006-558630822962_1600x718.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>Report Table 6, DeepSeek&#8217;s own numbers, standardised configuration across models.</em></p><p>It is also the smaller half of the story. On LongMRCR at 1M, Claude Opus 4.6 scores 92.9, V4-Pro-Max 83.5, Gemini-3.1-Pro 76.3. On CorpusQA at 1M, Opus 71.7, V4-Pro-Max 62.0, Gemini 53.8. DeepSeek beats Gemini on both, comfortably, and <strong>loses to Opus by 9.4 and 9.7 points, in their own table, measured with their own harness</strong>. </p><p>Choosing Gemini as the comparison in the summary is a defensible framing decision and it is a framing decision.</p><p>The report is honest in other places where it did not have to be. It <strong>reports Terminal-Bench 2.0 at 67.9</strong> on the original dataset while noting environment issues raised by another lab and disclosing that on the Verified subset V4-Pro scores about 72.0, which is higher. </p><p>It leaves cells blank for K2.6 and GLM-5.1 rather than filling them, saying those APIs were too busy to answer. It states that <strong>its reasoning performance trails the frontier by roughly three to six months.</strong> </p><p>That last sentence is in the report&#8217;s own summary of its own results and it is a more useful number than most third-party analysis of the same question.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/p/deepseek-v4-flash-the-cost-of-deciding?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/p/deepseek-v4-flash-the-cost-of-deciding?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><div><hr></div><h2>What this does to everyone else&#8217;s floor</h2><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!9O26!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F63f0fd58-5b62-4abd-9e19-078a2c22b67d_1600x864.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!9O26!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F63f0fd58-5b62-4abd-9e19-078a2c22b67d_1600x864.png 424w, https://substackcdn.com/image/fetch/$s_!9O26!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F63f0fd58-5b62-4abd-9e19-078a2c22b67d_1600x864.png 848w, https://substackcdn.com/image/fetch/$s_!9O26!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F63f0fd58-5b62-4abd-9e19-078a2c22b67d_1600x864.png 1272w, https://substackcdn.com/image/fetch/$s_!9O26!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F63f0fd58-5b62-4abd-9e19-078a2c22b67d_1600x864.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!9O26!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F63f0fd58-5b62-4abd-9e19-078a2c22b67d_1600x864.png" width="1456" height="786" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/63f0fd58-5b62-4abd-9e19-078a2c22b67d_1600x864.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:786,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;List rates, August 2026. The gap to the cheapest Western flagship is more than two orders of magnitude.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="List rates, August 2026. The gap to the cheapest Western flagship is more than two orders of magnitude." title="List rates, August 2026. The gap to the cheapest Western flagship is more than two orders of magnitude." srcset="https://substackcdn.com/image/fetch/$s_!9O26!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F63f0fd58-5b62-4abd-9e19-078a2c22b67d_1600x864.png 424w, https://substackcdn.com/image/fetch/$s_!9O26!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F63f0fd58-5b62-4abd-9e19-078a2c22b67d_1600x864.png 848w, https://substackcdn.com/image/fetch/$s_!9O26!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F63f0fd58-5b62-4abd-9e19-078a2c22b67d_1600x864.png 1272w, https://substackcdn.com/image/fetch/$s_!9O26!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F63f0fd58-5b62-4abd-9e19-078a2c22b67d_1600x864.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>List rates, August 2026. The gap to the cheapest Western flagship is more than two orders of magnitude.</em></p><p>Reuters, reporting <strong>Artificial Analysis figures,</strong> put V4-Flash at roughly 3 cents to complete the Intelligence Index battery, against 86 cents for Kimi K3, $1.86 for GPT-5.6 Sol and $3.15 for Claude Fable 5. </p><p>On the index itself Flash scores 50, tying Gemini 3.6 Flash, one point behind <strong>GLM-5.2 and Muse Spark 1.1</strong>, seven behind Kimi K3, and nine or more behind Opus 5, Fable 5 and GPT-5.6.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!egWI!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6e2f797-4468-4f8c-bf30-47125bab71f8_1600x836.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!egWI!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6e2f797-4468-4f8c-bf30-47125bab71f8_1600x836.png 424w, https://substackcdn.com/image/fetch/$s_!egWI!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6e2f797-4468-4f8c-bf30-47125bab71f8_1600x836.png 848w, https://substackcdn.com/image/fetch/$s_!egWI!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6e2f797-4468-4f8c-bf30-47125bab71f8_1600x836.png 1272w, https://substackcdn.com/image/fetch/$s_!egWI!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6e2f797-4468-4f8c-bf30-47125bab71f8_1600x836.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!egWI!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6e2f797-4468-4f8c-bf30-47125bab71f8_1600x836.png" width="1456" height="761" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b6e2f797-4468-4f8c-bf30-47125bab71f8_1600x836.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:761,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Artificial Analysis, August 2026. The vertical axis is nine points wide. The horizontal axis is two orders of magnitude.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Artificial Analysis, August 2026. The vertical axis is nine points wide. The horizontal axis is two orders of magnitude." title="Artificial Analysis, August 2026. The vertical axis is nine points wide. The horizontal axis is two orders of magnitude." srcset="https://substackcdn.com/image/fetch/$s_!egWI!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6e2f797-4468-4f8c-bf30-47125bab71f8_1600x836.png 424w, https://substackcdn.com/image/fetch/$s_!egWI!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6e2f797-4468-4f8c-bf30-47125bab71f8_1600x836.png 848w, https://substackcdn.com/image/fetch/$s_!egWI!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6e2f797-4468-4f8c-bf30-47125bab71f8_1600x836.png 1272w, https://substackcdn.com/image/fetch/$s_!egWI!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6e2f797-4468-4f8c-bf30-47125bab71f8_1600x836.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>Artificial Analysis, August 2026. The vertical axis is nine points wide. The horizontal axis is two orders of magnitude.</em></p><p>The standard rebuttal to a cheap model is that it burns more tokens, so cost per task closes the gap that cost per token opens. It does not work here and the numbers say so precisely. </p><p><strong>Flash generated 210M output tokens on the index against a 100M median</strong>, so it is 2.1 times more verbose than the typical model on identical work, for the documented reason that its Think Max system prompt instructs it to be. It is still 105 times cheaper per task than Fable 5, whose output token costs 179 times more. </p><p>Doubling token count against a 179x price advantage leaves 89x. The measured 105x is in that neighbourhood, and <strong>the verbosity tax is nowhere near large enough to matter</strong>.</p><p>What follows for the market is narrower than the headline suggests.</p><p>The nine-point index gap is not a rounding error. It is the <em>difference between a model that finishes a hard task and one that plausibly fails it, </em>and on hard agentic work a failed run costs more than the token price of a successful one. </p><p>At the top of the market price is not the binding constraint, and a 105x discount on a wrong answer is not a discount.</p><blockquote><p><em>A 105x discount on a wrong answer is not a discount.</em><strong><span>On where the price pressure actually lands</span></strong></p></blockquote><p>The pressure is on the middle. Every workload running on a flagship because nobody bothered to route it, every classification and extraction and summarisation and <strong>first-draft call</strong>, is now paying somewhere between 60 and 105 times more than it needs to. </p><p>Routing is the mechanism that transfers that value and routing is engineering work most teams have not done. The models that should be nervous are not Opus 5 and Fable 5. </p><p>They are Haiku, Sonnet, the Flash and Terra and Luna tiers, <strong>everything between $1 and $5 per million output tokens</strong> where the capability gap to V4-Flash is small or negative and the price gap is 20x or more.</p><p><strong>Two second-order effects </strong>are worth flagging. OpenRouter currently lists twenty providers serving the preview weights at $0.088 in and $0.176 out, 37 percent below DeepSeek&#8217;s own card. </p><p>That is an MIT-licensed model being resold below the price its author charges. Given the<em> break-even analysis above</em>, those hosts are running deeper batches, or cheaper capacity, or buying share. Only the first is durable.</p><p>And the weights are MIT, the mega-kernel is open inside DeepGEMM, the kernel library is open, and the fine-grained expert-parallel scheme was validated on Huawei Ascend as well as NVIDIA. <strong>The architecture is not the moat. The traffic shape is</strong>, and so is the on-disk cache that only pays off at DeepSeek&#8217;s request volume.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Where we might be wrong</h2><p>The CSA and HCA layer split for Flash is inferred. The report says the first two layers are pure sliding window and the rest interleave, which for 41 remaining layers gives either 21 CSA and 20 HCA or the reverse. </p><p><strong>We assumed the interleave opens with CSA.</strong> If it opens with HCA the cache figure moves from 3.372 to 3.216 GiB, about 5 percent, the parameter reconstruction moves by 13M, and nothing in the argument changes.</p><p>The B200 configuration is nameplate: 180 GB usable per GPU, 8 TB/s, four GPUs, 85 percent HBM utilisation, and a 63 percent achieved bandwidth fraction calibrated against a single third-party throughput measurement. </p><p>A production deployment with prefill and decode disaggregated across separate node pools has <strong>different and probably better economics than the single-pool model we used</strong>, and DeepSeek describes that disaggregation themselves.</p><p>The owner-cost figure of $1.91 per GPU-hour is built from a $38,000 capex assumption, a three-year life and 80 percent utilisation, none of which DeepSeek publishes. Capex at $30,000 gives $1.53 and at $45,000 gives $2.24, so the <strong>conclusion that an owner clears the 1M break-even survives the range</strong>, but only just at the top of it. The utilisation assumption is the fragile one: at 50 percent utilisation the figure is $2.99 and the verdict flips.</p><p>The blended-revenue model splits node time between prefill and decode on a single pool. Real serving disaggregates them across separate node pools with different shapes, which changes the split and generally improves it. </p><p><strong>It also assumes cache hits consume no node time beyond the disk read</strong>, which understates their cost at high hit rates where storage bandwidth starts to bind.</p><p>The FLOPs reconstruction lands at 12.2 percent of V3.2 at 1M against the report&#8217;s 10 percent. The report measures in equivalent FP8 FLOPs while <strong>V4&#8217;s experts run FP4</strong>, which has identical peak throughput on current silicon but which the report says could be a third cheaper on future hardware. </p><p>That accounting difference is the likely source. My V3.2 indexer model may also be too generous.</p><p>The prefill MFU of 35 percent is an assumption, not a measurement, and both prefill revenue and the blended break-even scale linearly with it. At <strong>20 percent the prefill figure drops from $244 an hour to $140</strong> and every blended break-even in the traffic-mix chart falls by roughly a third.</p><p>Our reading of <strong>BigCodeBench&#8217;s regression as an API-recall effect</strong> is speculation and I have flagged it as such in the text. I would drop it entirely if the world-knowledge results did not cut the other way.</p><p>And the largest one. Every agent benchmark in the 0731 release is vendor-reported, measured with a <strong>harness DeepSeek has announced but not shipped</strong>, at maximum reasoning effort, with sampling parameters DeepSeek chose. Two of the nine suites are DeepSeek&#8217;s own internal sets. I have reproduced none of them and neither has anyone else.</p><h2>Seven predictions, dated</h2><ol><li><p><strong>By June 2027</strong>, at least one major Western lab ships a production model that compresses the KV cache along the sequence axis rather than the head axis. The mechanism is cheap, the memory win is fifty times, and it has now been demonstrated at 32T tokens of pre-training rather than in an ablation.</p></li><li><p><strong>By 31 December 2026</strong>, someone publishes a sub-linear top-k selector for compressed attention, most likely hierarchical or learned-hash, and demonstrates it on V4&#8217;s open weights. The indexer being four fifths of attention cost at long context is too visible a target, and the FlashMemory work has already aimed at it from the prefetch side.</p></li><li><p><strong>By March 2027</strong>, DeepSeek introduces context-length-tiered pricing, a peak-hour surcharge that actually activates, or both. A flat rate across a thirty times span of serving cost is not stable, and a surcharge has already been announced without being switched on.</p></li><li><p><strong>By September 2027</strong>, at least one of Anthropic, OpenAI or Google cuts a mid-tier model&#8217;s output price by more than 50 percent without a corresponding capability release. The squeeze is on the middle of the ladder, not the top.</p></li><li><p><strong>By the end of December 2027</strong>, DeepSeek-V5 or its equivalent replaces SwiGLU with a gate-free elementwise activation. The report asks for this explicitly in its hardware proposals, gives the reason (removing the gate projection lets the intermediate dimension grow under a fixed budget, which raises the interconnect balance point), and labs that publish that kind of request are usually describing work already underway.</p></li><li><p><strong>By late June 2027</strong>, an SMT or ILP solver appears in the compilation pipeline of at least one other major inference stack. Z3 inside TileLang&#8217;s algebraic system is the first production instance I know of, the payoff is a few seconds of compile time for stronger vectorisation and bound analysis, and the idea travels.</p></li><li><p><strong>By 31 December 2027</strong>, no model with a one-million-token context window is profitable at that context length on <em>rented</em> NVIDIA hardware at published rates, while remaining profitable for owner-operators. The KV pool caps concurrency, concurrency is what pays for decode, and the gap between owning and renting is wider than the margin. This is the prediction I expect to age worst, and it fails if HBM capacity per package jumps faster than I think.</p></li></ol><div class="directMessage button" data-attrs="{&quot;userId&quot;:201922174,&quot;userName&quot;:&quot;Lorenzo Bradanini&quot;,&quot;canDm&quot;:null,&quot;dmUpgradeOptions&quot;:null,&quot;isEditorNode&quot;:true}" data-component-name="DirectMessageToDOM"></div><div><hr></div><h2>Confidence dossier</h2><p><em><strong>Tier A &#183; Stated in a primary source &#183; 13 claims</strong></em></p><ul><li><p><em><strong>43 layers, d=4096, m=4, m&#8217;=128, top-k 512, 256+1 experts, 6 activated</strong><br><span>Report section 4.2.1, stated directly</span></em></p></li><li><p><em><strong>Partial RoPE on exactly the last 64 dimensions; inverse RoPE at position -i on the output</strong><br><span>Report section 2.3.3, stated directly</span></em></p></li><li><p><em><strong>HCA has one KV stream and no indexer; CSA has two overlapping streams</strong><br><span>Report equations 9 to 12 versus 20 to 23</span></em></p></li><li><p><em><strong>Attention sink logits per head, added to the softmax denominator</strong><br><span>Report equation 27</span></em></p></li><li><p><em><strong>Hybrid Newton-Schulz, 8 steps at (3.4445, -4.7750, 2.0315) then 2 at (2, -1.5, 0.5)</strong><br><span>Report section 2.4</span></em></p></li><li><p><em><strong>No QK-Clip, because RMSNorm on queries and KV entries already bounds the logits</strong><br><span>Report section 2.4, stated as a deliberate omission</span></em></p></li><li><p><em><strong>Dense attention for the first 1T tokens; sparsity introduced at 64K with an indexer warmup</strong><br><span>Report section 4.2.2</span></em></p></li><li><p><em><strong>Index scores quantised FP32 to BF16: 2x on top-k, 99.7 percent KV recall preserved</strong><br><span>Report section 3.4</span></em></p></li><li><p><em><strong>Interconnect condition C/B &lt;= 2 d_ff; 6144 FLOPs per byte for Pro</strong><br><span>Report section 3.1, derived and evaluated there</span></em></p></li><li><p><em><strong>Three on-disk SWA strategies; zero-caching needs n_win x L tokens of recompute</strong><br><span>Report section 3.6.2</span></em></p></li><li><p><em><strong>Quick Instruction tokens reuse the existing KV cache to avoid a second prefill</strong><br><span>Report section 5.1.1, Table 5</span></em></p></li><li><p><em><strong>0731 is the same structure and size as the preview; post-training only</strong><br><span>DeepSeek API changelog, 31 July 2026</span></em></p></li><li><p><em><strong>Rate card $0.14 / $0.0028 / $0.28 per million</strong><br><span>Artificial Analysis and DeepSeek changelog, mutually consistent</span></em></p></li></ul><p><em><strong>Tier B &#183; Derived here and cross-checked &#183; 8 claims</strong></em></p><ul><li><p><em><strong>Backbone reconstructs to 284.20B; MTP module excluded from the headline</strong><br><span>Derived here from published constants, 0.07 percent residual</span></em></p></li><li><p><em><strong>Attention is 1.74 percent of weights and 39.3 percent of activated</strong><br><span>Derived here; follows from the reconstruction</span></em></p></li><li><p><em><strong>mHC GEMM output dimension 24 = n_hc + n_hc squared + n_hc</strong><br><span>Derived here; matches the figure quoted in report section 3.3</span></em></p></li><li><p><em><strong>3.372 GiB of KV per 1M sequence at production precision</strong><br><span>Derived here; four independent cross-checks, worst 10 percent</span></em></p></li><li><p><em><strong>Indexer is 79 percent of attention FLOPs at 1M, 50 percent at 32K</strong><br><span>Derived here from published dimensions</span></em></p></li><li><p><em><strong>Attention overtakes the expert bank at roughly 213K tokens</strong><br><span>Derived here; sensitive to the CSA/HCA split assumption</span></em></p></li><li><p><em><strong>Flash&#8217;s balance point is 4096 FLOPs per byte, 1.5x more demanding than Pro</strong><br><span>Derived here by applying the report&#8217;s own condition to Flash&#8217;s d_ff</span></em></p></li><li><p><em><strong>5,504-token recompute restores the full SWA state, a 191x saving on prefill</strong><br><span>Derived here from n_win, L and the activated parameter count</span></em></p></li></ul><p><em><strong>Tier C &#183; Model output, assumptions named &#183; 7 claims</strong></em></p><ul><li><p><em><strong>129 concurrent 1M sessions on one 4xB200 node</strong><br><span>Model; depends on 180 GB usable and 85 percent utilisation</span></em></p></li><li><p><em><strong>63 percent of peak HBM bandwidth achieved at low batch</strong><br><span>Inferred from one third-party throughput measurement</span></em></p></li><li><p><em><strong>1M context clears at owner cost and not at any rental rate</strong><br><span>Model output; hinges on the $38k / 3yr / 80 percent capex assumption</span></em></p></li><li><p><em><strong>Blended break-even is 32 to 64 dollars per GPU-hour on agentic traffic mixes</strong><br><span>Model output; single-pool prefill and decode, 35 percent prefill MFU</span></em></p></li><li><p><em><strong>Owner cost near $1.91 per GPU-hour against a $3.35 to $6.35 rental band</strong><br><span>Model output; capex and utilisation assumed, power from published TDP</span></em></p></li><li><p><em><strong>Break-even needs batch 256 at 32K against a $5.50 GPU-hour</strong><br><span>Model output; single-pool assumption, no disaggregation</span></em></p></li><li><p><em><strong>Prefill grosses $244 an hour against a $22 node</strong><br><span>Model output; 35 percent MFU is assumed, not measured</span></em></p></li></ul><p><em><strong>Tier D &#183; Unreproduced or speculative &#183; 3 claims</strong></em></p><ul><li><p><em><strong>BigCodeBench regression is a long-tail API recall effect</strong><br><span>Speculation, contradicted by the world-knowledge results</span></em></p></li><li><p><em><strong>Agent benchmark figures for the 0731 release</strong><br><span>Vendor reported, unshipped harness, two internal suites</span></em></p></li><li><p><em><strong>OpenRouter hosts undercutting DeepSeek by 37 percent are unprofitable</strong><br><span>Speculative; their batch depth and capacity costs are unknown</span></em></p></li></ul><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/p/deepseek-v4-flash-the-cost-of-deciding?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/p/deepseek-v4-flash-the-cost-of-deciding?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><div><hr></div><h2>Reproducing this</h2><p>Everything numerical above comes from three short scripts with no dependencies beyond the standard library. None of it needs a GPU, because none of it is a measurement of a running model. </p><p>It is arithmetic over published constants, validated against figures published independently by vLLM and against ratios stated in the report itself.</p><p>The parameter reconstruction. Note the asymmetry between CSA and HCA in the projection count, which is the thing that is easy to get wrong:</p><pre><code>L, d = 43, 4096
n_h, c, d_c = 64, 512, 1024          # query heads, head dim, query compression dim
n_hI, c_I   = 64, 128                # indexer heads, indexer head dim
g, d_g      = 8, 1024                # output projection groups
m, mp       = 4, 128                 # CSA and HCA compression rates
n_exp, n_sh, n_act, d_ff = 256, 1, 6, 2048
n_hc, V     = 4, 129280
n_swa, n_csa, n_hca = 2, 21, 20      # 2 SWA + 41 interleaved

per_expert = 3 * d * d_ff                        # SwiGLU: gate, up, down
moe_tot, moe_act = (n_exp+n_sh)*per_expert, (n_act+n_sh)*per_expert

def attn(kind):
    if   kind == &#8216;csa&#8217;: p = 4*d*c + 2*m*c        # two KV streams + positional biases
    elif kind == &#8216;hca&#8217;: p = 2*d*c + mp*c         # one KV stream  + positional bias
    else:               p = 1*d*c                # pure SWA, uncompressed
    p += d*d_c + d_c*(c*n_h)                     # query down then up
    if kind == &#8216;csa&#8217;:
        p += d_c*(c_I*n_hI) + d*n_hI             # indexer queries and per-head gate
    p += g*((n_h//g)*c)*d_g + (g*d_g)*d          # grouped output projection
    return p + n_h                               # attention sink logits

mhc = 2*(n_hc*d)*n_hc + (n_hc*d)*(n_hc**2)       # W_pre, W_post, W_res
att = n_swa*attn(&#8217;swa&#8217;) + n_csa*attn(&#8217;csa&#8217;) + n_hca*attn(&#8217;hca&#8217;)
backbone = L*moe_tot + att + 2*L*mhc + L*d*n_exp + 2*V*d
active   = L*moe_act + att + 2*L*mhc + L*d*n_exp

print(backbone/1e9, active/1e9)                  # 284.202  12.610
print(n_hc + n_hc**2 + n_hc)                     # 24, matching the mHC GEMM in section 3.3</code></pre><p>The KV byte model, with all four calibration checks it has to pass before being used:</p><pre><code>GiB, N = 1024**3, 1_048_576
c, c_I = 512, 128

def layer(kind, ent, idx, m=4, mp=128, n_win=128):
    if kind == &#8216;c4a&#8217;:   return (N//m)  * (ent + idx)
    if kind == &#8216;c128a&#8217;: return (N//mp) * ent
    if kind == &#8216;swa&#8217;:   return n_win * ent

# check 1: V3.2 bf16, MLA 512 latent + 64 rope, indexer 128, 61 layers
assert abs(61*(576*2 + 128*2)*N/GiB - 83.9) &lt; 0.1              # vLLM publishes 83.9

# check 2: V4-Pro bf16, 30 c4a + 31 c128a, key and value shared
pro = 30*layer(&#8217;c4a&#8217;, c*2, c_I*2) + 31*layer(&#8217;c128a&#8217;, c*2, 0)
assert abs(pro/GiB - 9.62) &lt; 0.01                              # vLLM publishes 9.62

# production precision: 64 rope dims bf16 + 448 fp8 = 576 B; fp4 indexer = 64 B
ent, idx = 64*2 + (c-64)*1, c_I//2
flash = 21*layer(&#8217;c4a&#8217;, ent, idx) + 20*layer(&#8217;c128a&#8217;, ent, 0) + 43*layer(&#8217;swa&#8217;, ent, 0)
v32   = 61*(64*2 + 512 + 128)*N

# check 3: the report&#8217;s own Figure 1 ratio
print(v32/flash)                       # 13.57  vs the report&#8217;s &#8220;13.7x smaller&#8221;

# check 4: the report&#8217;s claim that uncompressed SWA state is ~8x the compressed state
print(43*N*ent / (flash - 43*layer(&#8217;swa&#8217;, ent, 0)))            # 7.2  vs &#8220;approximately 8&#8221;

print(flash/GiB, flash/N)              # 3.372 GiB, 3453 bytes per context token
print(100*flash / (43*(2*8*128*2)*N))  # 1.96 percent of bf16 GQA-8, report says ~2</code></pre><p>The decode roofline and break-even, for substituting your own hardware and rates:</p><pre><code>NG, BW = 4, 8.0e12                                   # four B200, 8 TB/s each
w_step  = 278.1e9*0.5 + NG*6.2e9                     # FP4 experts + FP8 dense per rank
t_floor = w_step / (NG*BW)                           # 5.12 ms
eff     = t_floor / (1/122.7)                        # 0.63, from the measured 122.7 tok/s

def kv_bytes(ctx, ent=576, idx=64):
    return 21*((512+128)*ent + (ctx//4)*idx) + 20*(max(1, ctx//128)*ent + 128*ent)

def breakeven(batch, ctx, price=0.28):
    t = ((w_step + batch*kv_bytes(ctx)) / (NG*BW)) / eff
    return (batch/t) * 3600/1e6 * price / NG         # dollars per GPU-hour

print(breakeven(512, 32768))      # 14.76  clears the market comfortably
print(breakeven(256, 32768))      #  7.64  clears a mid-market B200
print(breakeven(129, 1048576))    #  2.93  clears nothing you can rent

# the report&#8217;s interconnect condition, applied to Flash
for name, d_ff in ((&#8221;Flash&#8221;, 2048), (&#8221;Pro&#8221;, 3072)):
    print(name, 2*d_ff, &#8220;FLOP/Byte -&gt;&#8221;, 4.5e15/(2*d_ff)/1e12, &#8220;TB/s at 4.5 PFLOP/s FP8&#8221;)</code></pre><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Sources</h2><ul><li><p><em>DeepSeek-AI. DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence. arXiv:2606.19348, 26 April 2026. Sections 2.2 to 2.4 for architecture, 3.1 for expert parallelism and the interconnect balance point, 3.2 for TileLang, 3.3 for batch invariance and determinism, 3.4 for FP4 QAT, 3.5 for the training framework, 3.6 for KV cache management and on-disk storage, 4.2 for the constants and the training schedule, 5.1 for post-training, 5.2 for RL infrastructure and DSec, 5.3 for evaluation.</em></p></li><li><p><em>vLLM Team. DeepSeek V4 in vLLM: Efficient Long-context Attention, 24 April 2026. Appendix contains the KV arithmetic used here for calibration, the derivation of why inverse RoPE is needed when key and value are shared, and the exact top-k values for c4a and c128a.</em></p></li><li><p><em>LMSYS. DeepSeek-V4 on Day 0: From Fast Inference to Verified RL with SGLang and Miles, 25 April 2026, for the MXFP8 by MXFP4 fused MoE path and the split-K mHC pre-GEMM kernel.</em></p></li><li><p><em>DeepSeek API changelog, 31 July 2026, for the 0731 release scope, the agent benchmark table and the harness settings.</em></p></li><li><p><em>Artificial Analysis, model page for DeepSeek V4 Flash 0731, accessed 5 August 2026, for Intelligence Index v4.1, output speed, time to first token, token volume and the rate card.</em></p></li><li><p><em>Reuters, via Quartz and Business Standard, 3 August 2026, for the cost-per-task comparison across V4-Flash, Kimi K3, GPT-5.6 Sol and Claude Fable 5.</em></p></li><li><p><em>Hugging Face model cards for <span>deepseek-ai/DeepSeek-V4-Flash</span> and <span>deepseek-ai/DeepSeek-V4-Pro</span>, for the mixed FP4 and FP8 weight format and the reference inference implementation.</em></p></li><li><p><em>DeepGEMM pull request 304, for the open-sourced MegaMoE fused expert-parallel kernel.</em></p></li><li><p><em>FlashMemory-DeepSeek-V4: Lightning Index Ultra-Long Context via Lookahead Sparse Attention, arXiv:2606.09079, for the Neural Memory Indexer follow-up and its own account of its limitations.</em></p></li><li><p><em>getdeploying.com B200 index (4 August 2026), gpuprice.fyi B200 index (31 July 2026) and published neocloud rate cards, for the $3.35 to $6.35 band.</em></p></li><li><p><em>Anthropic, OpenAI, Moonshot and Z.ai published rate cards as of 4 August 2026, for the price ladder.</em></p></li></ul>]]></content:encoded></item><item><title><![CDATA[Invite your friends to read The Software Frontier]]></title><description><![CDATA[A warm thank you]]></description><link>https://www.thesoftwarefrontier.com/p/invite-your-friends-to-read-the-software</link><guid isPermaLink="false">https://www.thesoftwarefrontier.com/p/invite-your-friends-to-read-the-software</guid><dc:creator><![CDATA[Lorenzo Bradanini]]></dc:creator><pubDate>Tue, 04 Aug 2026 06:26:25 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!SAY7!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F54550d86-2756-4131-8818-956604f6749d_608x608.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2>A warm thank you</h2><p>Thanks so much for reading <a href="https://www.thesoftwarefrontier.com/">The Software Frontier</a> ! Your huge support keeps us motivated to do this work.</p><p>If you do really enjoy <strong>The Software Frontier</strong>, we would be really happy if you invited friends to subscribe and read with us. If you refer friends, you will receive a few benefits that give you special access to The Software Frontier.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>How to participate </h2><p><strong>1. Share The Software Frontier. </strong>When you use the referral link below, or the &#8220;Share&#8221; button on any post, you'll get credit for any new subscribers. Simply send the link in a text, email, or share it on social media with friends.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/leaderboard?&amp;utm_source=post&quot;,&quot;text&quot;:&quot;Refer a friend&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/leaderboard?&amp;utm_source=post"><span>Refer a friend</span></a></p><p>2.<strong> Earn benefits.</strong> When more friends use your referral link to subscribe (free or paid), you&#8217;ll receive special benefits.</p><ul><li><p>Get a 1 month comp for 3 referrals</p></li><li><p>Get a 3 month comp for 5 referrals</p></li><li><p>Get a 6 month comp for 25 referrals</p></li></ul><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/leaderboard?&amp;utm_source=post&quot;,&quot;text&quot;:&quot;Visit the leaderboard&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/leaderboard?&amp;utm_source=post"><span>Visit the leaderboard</span></a></p><p>To learn more, check out <a href="https://support.substack.com/hc/en-us/articles/16142857300372">Substack&#8217;s FAQ</a>.</p><p>Thank you for helping get the word out about The Software Frontier! Without you, none of this could have been made possible. </p><p>Lorenzo Bradanini and Lorenzo Tettamanti. </p>]]></content:encoded></item><item><title><![CDATA[How Blackwell’s Tensor Memory Actually Works ]]></title><description><![CDATA[Blackwell's largest matrix instruction needs 256 registers per thread. The ceiling is 255.]]></description><link>https://www.thesoftwarefrontier.com/p/how-blackwells-tensor-memory-actually</link><guid isPermaLink="false">https://www.thesoftwarefrontier.com/p/how-blackwells-tensor-memory-actually</guid><dc:creator><![CDATA[Lorenzo Bradanini]]></dc:creator><pubDate>Mon, 03 Aug 2026 15:07:50 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!zfIO!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4a8ed67-c0a5-4192-8a9c-b5a9b3edd44c_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!zfIO!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4a8ed67-c0a5-4192-8a9c-b5a9b3edd44c_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!zfIO!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4a8ed67-c0a5-4192-8a9c-b5a9b3edd44c_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!zfIO!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4a8ed67-c0a5-4192-8a9c-b5a9b3edd44c_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!zfIO!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4a8ed67-c0a5-4192-8a9c-b5a9b3edd44c_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!zfIO!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4a8ed67-c0a5-4192-8a9c-b5a9b3edd44c_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!zfIO!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4a8ed67-c0a5-4192-8a9c-b5a9b3edd44c_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e4a8ed67-c0a5-4192-8a9c-b5a9b3edd44c_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2322458,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/208027653?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4a8ed67-c0a5-4192-8a9c-b5a9b3edd44c_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!zfIO!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4a8ed67-c0a5-4192-8a9c-b5a9b3edd44c_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!zfIO!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4a8ed67-c0a5-4192-8a9c-b5a9b3edd44c_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!zfIO!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4a8ed67-c0a5-4192-8a9c-b5a9b3edd44c_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!zfIO!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4a8ed67-c0a5-4192-8a9c-b5a9b3edd44c_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>For nine years the register file per SM didn&#8217;t grow of a byte, while tensor throughput per clock doubled almost four times. Blackwell&#8217;s answer was to take out the accumulator out of the register file, and to put it in an address space with its own allocator, its own barrier, and no coherence with anything. This is what that memory is, what the compiler actually emits for it, and the invariance that explains why it had to exist.</em></p><h2>CUDA Mastery 2026</h2><p>This article is a narrow part of one architecture. The <strong>actual guide</strong> is the wide one: 34 chapters and thousands words on CUDA 13.x, from memory model up through <strong>Hopper</strong> and <strong>Blackwell tensor cores</strong>, written the same way as everything here. Compile it, disassemble it, show the command, then explain what happened.</p><p>It also has a corrections page at the end. An early edition listed H100 sparse tensor rates as dense ones, and that single mislabel propagated into three roofline figures before a reader caught it. </p><p>That figures are fixed and the mistake is documented rather than quietly removed, which is <strong>the standard </strong>we would want from anyone selling us a technical book.</p><p><a href="https://lorenzobrada.gumroad.com/l/cuda_mastery">Get CUDA Mastery guide </a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share The Software Frontier&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share The Software Frontier</span></a></p><div><hr></div><h2>Introduction</h2><p>The very first thing I did with Blackwell was to compile something for it on a machine that has <strong>no GPU</strong> in it at all. That is not a counter sense. </p><p><code>ptxas</code>, <code>nvdisasm</code> and <code>cuobjdump</code> are ordinary x86 programs, they ship as Python wheels on PyPI, and they will <strong>lower PTX </strong>for compute capability 10.0 on a laptop with a 40 mb download and no driver. You can&#8217;t run the result, but what you can do is to read every instruction, which for my own research, turned out to be the more useful half.</p><p>I wanted to know what a <code>tcgen05.alloc</code> costs, first was a small quest, but it got way bigger. It&#8217;s important because that&#8217;s the instruction that reserves <strong>Tensor Memory on Blackwell</strong>, it&#8217;s in every CUTLASS kernel for the architecture, and every explanation out there said the same <em>three things</em>: it takes a column count, it must be issued by one warp, and the result comes back through shared memory. </p><p>Nobody explained in full what the machine does. So, I wrote a <strong>small PTX kernel</strong> around it, assembled it for <code>sm_100a</code>, and disassembled the result.</p><p>What I saw was not an allocation in the sense a systems programmer means: it&#8217;s a <strong>uniform-datapath atomic</strong> against a per-SM pool, wrapped by the compiler in a spin loop with a <code>NANOSLEEP</code> backoff, guarded by three distinct trap handlers that <code>ptxas</code> injects on your behalf, with names like <code>__cuda_sm10x_tcgen05_guardrail_trap_unallocated_columns_being_dealloced</code>. </p><p>The tensor cores now have a memory allocator, with contention, with a retry path, and with a compiler-inserted runtime safety checker for use-after-free.</p><p>That&#8217;s the shape of the thing this article is about. Blackwell <strong>added an address space</strong> with its own instruction family, not just a cache with its own bus. On top of it we have its own allocator, a barrier, its own access-permission model, and zero coherence with anything else on the chip.</p><p>Almost everything written about Blackwell treats Tensor Memory as an implementation detail of the new MMA, but I found out it&#8217;s the opposite.</p><p>The <strong>MMA changed</strong> because the memory changed, which consequentially changed due to an arithmetic problem that had been building since 2017, and the consequences run outward through occupancy, epilogue design, quantization format choice, kernel portability, and eventually the cost per million tokens of anything you serve on hardware.</p><p>Neither of us owns a B200, unfortunately. Everything here that is measured was measured with a toolchain and a disassembler and is <strong>reproducible from the appendix</strong>. Derived things are in the open with the arithmetic on the page. All data taken from somebody else&#8217;s hardware is attributed and tiered in a dossier at the end. </p><p>Where the published literature and our disassembly disagree,(<em>spoiler alert: in two places they do</em>) both are shown.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>The number that did not move</h2><p>We write this very often: every <strong>streaming multiprocessor</strong> NVIDIA has launched since Volta has exactly 65,536 registers of 32 bits. That is 262,144 bytes of register file per SM, and NVIDIA&#8217;s own tuning guides give the same 64 K figure for Ampere and, in the Blackwell tuning guide, for compute capability 10.0. </p><p>The same documents states the<strong> 255 register per thread ceiling</strong>, which we&#8217;ll see in use. what we have is: five architectures, three process nodes, four HBM generations, and the largest piece of storage in the SM has not changed capacity by one byte in nine years.</p><p>Over the same period <strong>tensor throughput per SM</strong> per clock went up eight times, and the doubling is exact rather than approximate; NVIDIA&#8217;s own numbers make this easily checkable without trusting anybody&#8217;s marketing. </p><p>A Volta tensor core does <strong>64 fused multiply-adds</strong> per clock, eight of them per SM, so 1,024 FP16 FLOPs per clock per SM. Ampere doubled it to 2,048, then Hopper doubled it again to 4,096. </p><p>Multiply those out and the datasheets fall out to three digits: 108 A100 SMs times 2,048 times 1.410 GHz is 312 teraflops, which is exactly the published A100 figure. 132 H100 SXM SMs times 4,096 times 1.830 GHz is 989.7 teraflops against a published 989.4.</p><p>Run the same identity backwards on Blackwell. The B200 is two dies of 80 SMs with 74 enabled on each, so 148, and 2.25 petaflops dense FP16. That requires 8,192 FLOPs per clock per SM at 1.86 GHz. </p><p>Another doubling, and not only ours: the authors of the <strong>JAX scaling book </strong>run the same division and land on 2,048 FLOPs per tensor core per cycle across four tensor cores per SM, which is the same figure. So the ratio that actually matters, register file bytes per FP16 FLOP per clock, has fallen from 256 on Volta to 32 on Blackwell.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!vClk!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf0eac48-a9af-4255-afb6-0b6d447f9cd8_1800x800.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!vClk!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf0eac48-a9af-4255-afb6-0b6d447f9cd8_1800x800.png 424w, https://substackcdn.com/image/fetch/$s_!vClk!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf0eac48-a9af-4255-afb6-0b6d447f9cd8_1800x800.png 848w, https://substackcdn.com/image/fetch/$s_!vClk!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf0eac48-a9af-4255-afb6-0b6d447f9cd8_1800x800.png 1272w, https://substackcdn.com/image/fetch/$s_!vClk!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf0eac48-a9af-4255-afb6-0b6d447f9cd8_1800x800.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!vClk!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf0eac48-a9af-4255-afb6-0b6d447f9cd8_1800x800.png" width="1456" height="647" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/cf0eac48-a9af-4255-afb6-0b6d447f9cd8_1800x800.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:647,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Chart showing register file per SM flat at 262144 bytes from Volta to Blackwell while FP16 tensor FLOPs per clock per SM double each generation from 1024 to 8192&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Chart showing register file per SM flat at 262144 bytes from Volta to Blackwell while FP16 tensor FLOPs per clock per SM double each generation from 1024 to 8192" title="Chart showing register file per SM flat at 262144 bytes from Volta to Blackwell while FP16 tensor FLOPs per clock per SM double each generation from 1024 to 8192" srcset="https://substackcdn.com/image/fetch/$s_!vClk!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf0eac48-a9af-4255-afb6-0b6d447f9cd8_1800x800.png 424w, https://substackcdn.com/image/fetch/$s_!vClk!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf0eac48-a9af-4255-afb6-0b6d447f9cd8_1800x800.png 848w, https://substackcdn.com/image/fetch/$s_!vClk!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf0eac48-a9af-4255-afb6-0b6d447f9cd8_1800x800.png 1272w, https://substackcdn.com/image/fetch/$s_!vClk!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf0eac48-a9af-4255-afb6-0b6d447f9cd8_1800x800.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em><strong><sup>Figure 1.</sup></strong><sup> Register file capacity per SM against FP16 tensor throughput per SM per clock. The bars are flat by architectural fact. The line doubles every generation. Register file bytes per FLOP per clock: 256, 128, 64, 32.</sup></em></p><p>You can absorb a gap like that for one generation by being clever, and NVIDIA did, twice. Ampere added <strong>asynchronous copy</strong> so global to shared traffic stopped passing through registers. </p><p>Hopper added the Tensor Memory Accelerator so the address arithmetic for a tiled copy stopped consuming a warp&#8217;s registers, and added <code>setmaxnreg</code> so a<strong> producer warpgroup</strong> could donate its register budget to a consumer at runtime. Both are the same move: find something sitting in registers for no good reason and evict it.</p><p>By Hopper the evictable things were gone. What remained was the one thing that genuinely belongs to the math: the accumulator.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>The tile that cannot exist</h2><p>Here is the constraint that decided Blackwell&#8217;s design, and it is a single line of arithmetic against a limit you can make the assembler confirm.</p><p>On Hopper, <code>wgmma.mma_async</code> accumulates into registers owned by the 128 threads of a warpgroup. The largest shape is <code>m64n256k16</code>. That accumulator is 16,384 FP32 values, which over 128 threads is 128 registers per thread. </p><p>As stated in the subtitle, the architectural ceiling is 255, and <code>ptxas</code> will tell you so directly:</p><pre><code><code>$ ptxas -arch=sm_100a -maxrregcount=255 probe.ptx -o /dev/null
$ ptxas -arch=sm_100a -maxrregcount=256 probe.ptx -o /dev/null
ptxas warning : Too big maxrregcount value specified 256, will be ignored</code></code></pre><p>So Hopper&#8217;s largest MMA already spends half of every thread&#8217;s addressable register space on the output tile, before any addressing, predication, loop state or <strong>epilogue arithmetic.</strong></p><p>Now&#8230; let&#8217;s double it, which is what <code>tcgen05.mma</code> does. The largest single-CTA UMMA atom is <code>m128n256k16</code>, twice the area of the largest WGMMA atom. Its accumulator is 32,768 FP32 values. Over a warpgroup that is <strong>256 registers per thread</strong>, against a ceiling of 255.</p><p>Blackwell&#8217;s headline matrix instruction produces a result that is, by one register, unrepresentable in the programming model of every NVIDIA GPU that came before it.</p><p>This is a representability problem and not a turning tradeoff. There is no register allocation, no spill policy, <strong>no compiler heroics</strong> that make a 128 by 256 FP32 tile live in the fragments of a warpgroup, because the fragment model tops out below the tile. </p><p>If you want that instruction to exist, its output must go somewhere that is not the register file. Everything else about Tensor Memory follows from that sentence.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!ygwX!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6477b278-49a2-430f-8aae-b9bffd458da2_1800x760.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!ygwX!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6477b278-49a2-430f-8aae-b9bffd458da2_1800x760.png 424w, https://substackcdn.com/image/fetch/$s_!ygwX!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6477b278-49a2-430f-8aae-b9bffd458da2_1800x760.png 848w, https://substackcdn.com/image/fetch/$s_!ygwX!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6477b278-49a2-430f-8aae-b9bffd458da2_1800x760.png 1272w, https://substackcdn.com/image/fetch/$s_!ygwX!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6477b278-49a2-430f-8aae-b9bffd458da2_1800x760.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!ygwX!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6477b278-49a2-430f-8aae-b9bffd458da2_1800x760.png" width="1456" height="615" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6477b278-49a2-430f-8aae-b9bffd458da2_1800x760.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:615,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Bar chart of registers per thread required to hold FP32 accumulator tiles, with the 255 register ceiling crossed by the two largest Blackwell tiles&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Bar chart of registers per thread required to hold FP32 accumulator tiles, with the 255 register ceiling crossed by the two largest Blackwell tiles" title="Bar chart of registers per thread required to hold FP32 accumulator tiles, with the 255 register ceiling crossed by the two largest Blackwell tiles" srcset="https://substackcdn.com/image/fetch/$s_!ygwX!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6477b278-49a2-430f-8aae-b9bffd458da2_1800x760.png 424w, https://substackcdn.com/image/fetch/$s_!ygwX!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6477b278-49a2-430f-8aae-b9bffd458da2_1800x760.png 848w, https://substackcdn.com/image/fetch/$s_!ygwX!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6477b278-49a2-430f-8aae-b9bffd458da2_1800x760.png 1272w, https://substackcdn.com/image/fetch/$s_!ygwX!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6477b278-49a2-430f-8aae-b9bffd458da2_1800x760.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em><strong><sup>Figure 2.</sup></strong><sup> The forcing function. Hopper&#8217;s largest MMA already spent half the addressable register space per thread on the output tile. Blackwell&#8217;s largest single-CTA MMA needs one register more than the architecture allows.</sup></em></p><p>There is a second argument in the same direction, less absolute but more expensive in practice. A register is <strong>thread-private</strong>, so an MMA that accumulates into registers is an operation the owning threads must be present for. </p><p>On Hopper this is a scheduling tax: a warpgroup issues <code>wgmma</code>, and although the <em>instruction is asynchronous</em>, that warpgroup cannot go and do something else with those registers, because the registers <em>are</em> the accumulator. </p><p><strong>Warp specialization </strong>on Hopper is largely a set of arrangements to ensure the warps holding accumulators are not the warps doing anything interesting. Move the accumulator out and the whole class of tricks becomes unnecessary.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/p/how-blackwells-tensor-memory-actually?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/p/how-blackwells-tensor-memory-actually?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><div><hr></div><h2>What Tensor Memory actually is</h2><p>Tensor Memory is 256 kilobytes per SM, organized as <strong>128 lanes by 512 columns</strong> of 32 bit cells. A TMEM address is a 32 bit word whose high 16 bits are the lane index and low 16 bits are the column index. It is not a linear address space with a base and an offset, cause it&#8217;s a coordinate. </p><p>One consequence worth internalising before reading any <strong>CUTLASS TMEM layout</strong>: a stride of 65,536 in a TMEM tensor is not a large jump in memory, it is a step of exactly one lane.</p><p><em>Four properties</em> matter more than the capacity, and none are properties of any other memory on the chip.</p><ul><li><p><strong>It is allocated, not addressed.</strong> You call <code>tcgen05.alloc</code> with a column count, which must be a power of two and at least 32, and the hardware returns a base address which it writes into shared memory for you. Allocation is by column, and a column is all 128 lanes: there is no way to reserve part of one. You free it with <code>tcgen05.dealloc</code>, from the same warp that allocated it, or the columns stay claimed.</p></li><li><p><strong>Nothing computes on it.</strong> The only instructions that touch TMEM are the <code>tcgen05</code> family. No ALU op, no <code>ld.shared</code>, no <code>ldmatrix</code>, no <code>cp.async</code>, no atomic, no texture path. Every pre-processing step happens before data enters and every post-processing step after it leaves. TMEM is the one memory on a Blackwell SM that a general purpose instruction cannot see.</p></li><li><p><strong>Access is partitioned by warp, in hardware.</strong> When threads read or write TMEM explicitly, warp 0 of a warpgroup reaches only lanes 0 to 31, warp 1 only lanes 32 to 63, and so on. This is the access model, not a guideline. One warp physically cannot read a full 128 lane accumulator tile, so draining one requires a whole warpgroup by construction, and CUTLASS&#8217;s <code>make_tmem_copy</code> is hardcoded to four warps for exactly that reason.</p></li><li><p><strong>It is a per-SM pool with a hard ceiling.</strong> 512 columns, shared by every CTA resident on that SM, and the largest UMMA accumulator occupies exactly 256 of them.</p></li></ul><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!yMB9!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4faa0c3f-8b88-495a-9c08-417abcab5cf8_1800x920.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!yMB9!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4faa0c3f-8b88-495a-9c08-417abcab5cf8_1800x920.png 424w, https://substackcdn.com/image/fetch/$s_!yMB9!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4faa0c3f-8b88-495a-9c08-417abcab5cf8_1800x920.png 848w, https://substackcdn.com/image/fetch/$s_!yMB9!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4faa0c3f-8b88-495a-9c08-417abcab5cf8_1800x920.png 1272w, https://substackcdn.com/image/fetch/$s_!yMB9!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4faa0c3f-8b88-495a-9c08-417abcab5cf8_1800x920.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!yMB9!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4faa0c3f-8b88-495a-9c08-417abcab5cf8_1800x920.png" width="1456" height="744" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/4faa0c3f-8b88-495a-9c08-417abcab5cf8_1800x920.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:744,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Diagram of the Tensor Memory grid, 128 lanes by 512 columns, showing a 128 by 256 accumulator allocation, scale factor columns, and the warp lane partition&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Diagram of the Tensor Memory grid, 128 lanes by 512 columns, showing a 128 by 256 accumulator allocation, scale factor columns, and the warp lane partition" title="Diagram of the Tensor Memory grid, 128 lanes by 512 columns, showing a 128 by 256 accumulator allocation, scale factor columns, and the warp lane partition" srcset="https://substackcdn.com/image/fetch/$s_!yMB9!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4faa0c3f-8b88-495a-9c08-417abcab5cf8_1800x920.png 424w, https://substackcdn.com/image/fetch/$s_!yMB9!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4faa0c3f-8b88-495a-9c08-417abcab5cf8_1800x920.png 848w, https://substackcdn.com/image/fetch/$s_!yMB9!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4faa0c3f-8b88-495a-9c08-417abcab5cf8_1800x920.png 1272w, https://substackcdn.com/image/fetch/$s_!yMB9!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4faa0c3f-8b88-495a-9c08-417abcab5cf8_1800x920.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em><strong><sup>Figure 3</sup></strong><sup>. The ledger. Every Blackwell GEMM kernel is, underneath, a plan for carving up 512 columns. The largest accumulator takes half, block scale factors take a few more, and what is left is what you have for pipelining or for a second resident CTA.</sup></em></p><p>The instruction that uses all this is <code>tcgen05.mma</code>, which <strong>CUTLASS calls UMMA</strong>. Its operand rules invert everything before them. Operand A may be in shared memory or in Tensor Memory. Operand B has to be in shared memory. The accumulator must be in Tensor Memory. Registers appear nowhere in that sentence.</p><p>And it is issued by <strong>one thread</strong>. Not a warp, not a warpgroup: a single elected thread on behalf of the whole CTA, or of a pair of CTAs under <code>cta_group::2</code>, where two SMs sharing a texture processing cluster cooperate on one logical tile. </p><p>The consequence is visible in CUTLASS: the CuTe atom&#8217;s <code>ThrID</code>, which was <code>Layout&lt;_32&gt;</code> for <strong>warp-level MMA </strong>and <code>Layout&lt;_128&gt;</code> for Hopper&#8217;s warpgroup MMA, is now <code>Layout&lt;_1&gt;</code>, and the thread layouts have been repurposed as layouts of the CTAs collaborating on the instruction.</p><p>The abstraction the entire programming model is named after has been vacated at the top of the pipeline. There is still a thread. It does not do the math, does not own the inputs, and does not own the result.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>What the compiler actually emits</h2><p>This is the part we did ourselves, so it is the part we trust most. The toolchain is<strong> three PyPI wheels</strong> and no GPU: <code>ptxas</code> 12.9.86 from <code>nvidia-cuda-nvcc-cu12</code>, and <code>nvdisasm</code> and <code>cuobjdump</code> 13.3.73 from their standalone packages. The full command sequence is in the appendix.</p><p>The probe does the minimum honest thing: reserve 128 columns of <strong>Tensor Memory</strong>, read back the base address, issue one UMMA into it, commit through an mbarrier, drain one fragment to registers, store it, free the columns.</p><pre><code><code>// probe.ptx, assembled with ptxas 12.9.86 for sm_100a
.version 8.6
.target sm_100a
.address_size 64

.visible .entry umma_probe(.param .u64 p_out, .param .u64 p_adesc, .param .u64 p_bdesc)
{
    .reg .b32  %r&lt;16&gt;;   .reg .b64 %rd&lt;8&gt;;   .reg .pred %p&lt;2&gt;;
    .shared .align 16 .b32 tmem_slot[4];
    .shared .align 8  .b64 mbar[1];

    ld.param.u64 %rd1, [p_out];
    ld.param.u64 %rd2, [p_adesc];
    ld.param.u64 %rd3, [p_bdesc];

    mov.u32 %r1, 128;
    tcgen05.alloc.cta_group::1.sync.aligned.shared::cta.b32 [tmem_slot], %r1;
    tcgen05.relinquish_alloc_permit.cta_group::1.sync.aligned;
    ld.shared.b32 %r2, [tmem_slot];

    mov.u32 %r3, 0;
    setp.eq.u32 %p1, %r3, 0;
    tcgen05.mma.cta_group::1.kind::f16 [%r2], %rd2, %rd3, %r3, %p1;

    mbarrier.init.shared::cta.b64 [mbar], 1;
    tcgen05.commit.cta_group::1.mbarrier::arrive::one.shared::cluster.b64 [mbar];

    tcgen05.ld.sync.aligned.32x32b.x1.b32 {%r10}, [%r2];
    tcgen05.wait::ld.sync.aligned;
    st.global.u32 [%rd1], %r10;

    mov.u32 %r11, 128;
    tcgen05.dealloc.cta_group::1.sync.aligned.b32 %r2, %r11;
    ret;
}</code></code></pre><p><em>Three facts fall out before the disassembler is even involved.</em></p><ol><li><p><strong>The minimum PTX ISA version is 8.6</strong>, which is CUDA 12.8. Below that the assembler names the requirement precisely: <code>Feature 'tcgen05.alloc' requires PTX ISA .version 8.6 or later</code>. We mention this because at least one widely linked public write-up puts it at 8.4, and 8.4 does not assemble.</p></li><li><p><code>tcgen05.commit</code><strong> requires </strong><code>.shared::cluster</code>, and rejects <code>.shared::cta</code> with <code>State space incorrect for instruction 'tcgen05.commit'</code>, even in a kernel with a trivial cluster and <code>cta_group::1</code>. The tensor core completion path is a cluster level mechanism whether or not you asked for a cluster.</p></li><li><p><code>ptxas</code><strong> does not validate the column count, even when it is a compile time constant.</strong> We assembled the probe with 16, 48, 96 and 1,024 columns, all of which violate the documented power-of-two and minimum-32 rules, and all of which assembled without a warning. The rule is enforced by hardware at runtime, not by the compiler. Given that the failure mode is a trap handler, this is a class of bug that only exists on a machine you may not own.</p></li></ol><h3>The allocator</h3><p>Here is what <code>tcgen05.alloc</code> becomes. Address arithmetic trimmed, control flow still intact.</p><pre><code><code>// nvdisasm -c probe_sm100a.cubin, excerpt
        ELECT P0, URZ, PT ;
   @!P0 BRA `(.L_x_2) ;
        DEPBAR.LE SB0, 0x36 ;
        UTCATOMSWS.FIND_AND_SET.ALIGN UP0, UR4, UR4 ;      // claim an aligned run of columns
        PLOP3.LUT P0, PT, PT, PT, UP0, 0x80, 0x8 ;
        SEL R0, RZ, 0xffffffff, !P0 ;
        ISETP.NE.AND P0, PT, R0, RZ, PT ;
    @P0 BRA `(.L_x_3) ;                                    // claimed, continue
.L_x_4:
        NANOSLEEP 0x64 ;                                    // back off, then retry
        UTCATOMSWS.FIND_AND_SET.ALIGN UP0, UR4, UR4 ;
        ...
   @!P0 BRA `(.L_x_4) ;                                    // spin
.L_x_3:
        ATOMS.OR RZ, [UR4], R2 ;                           // record the claimed mask in SMEM
        STS [UR6], R0 ;                                    // publish the base address
        ...
        UVIRTCOUNT.DEALLOC.SMPOOL 0x80 ;                   // SM pool virtual counter</code></code></pre><p>The claim is a <code>UTCATOMSWS.FIND_AND_SET.ALIGN</code>, a uniform-datapath atomic that searches a per-SM pool for a free, aligned run of columns and sets them. </p><p>On failure the compiler emits a <code>NANOSLEEP</code> of 0x64 units and retries indefinitely. So a Blackwell kernel that allocates Tensor Memory has a <strong>spin loop</strong> on its critical path that no source line asked for, and the pool is genuinely contended, otherwise the retry path would not be there.</p><p>Then there are the <strong>guardrails</strong>. <code>ptxas</code> injects three named trap handlers and fifteen references to them, and it does so identically at every optimisation level from <code>-O0</code> to <code>-O3</code>:</p><pre><code><code>$__internal_0_$__cuda_sm10x_tcgen05_guardrail_trap_col_being_dealloced_not_returned_by_alloc
$__internal_1_$__cuda_sm10x_tcgen05_guardrail_trap_phase_invalid_during_alloc
$__internal_2_$__cuda_sm10x_tcgen05_guardrail_trap_unallocated_columns_being_dealloced</code></code></pre><p>Read those names as a bug taxonomy. Freeing columns you did not allocate. Freeing a column that came from a different allocation. Allocating from an invalid phase. </p><p>That is a <strong>use-after-free checker</strong>, a double-free checker and a state machine assertion, compiled into every kernel, unconditionally. NVIDIA does not do this casually, and the fact that they did it tells you what the failure modes look like in practice.</p><blockquote><p><em><strong>Why this matters beyond trivia</strong></em></p><p><em>Every mental model of a GPU kernel assumes resources are assigned at launch. Registers and shared memory are fixed by the compiler and the launch configuration, and occupancy is computed from them before a single instruction runs. Tensor Memory is the first first-class SM resource that is acquired at runtime, can block, can fail, and is arbitrated by an atomic. It moves part of the occupancy calculation out of the launch and into the kernel body, where no static tool can see it.</em></p></blockquote><p>One number for scale. The probe above, whose actual work is one matrix multiply and one 32 bit drain, compiles to <strong>152 SASS instructions at </strong><code>-O3</code> and 400 at <code>-O0</code>, using 14 registers and 24 bytes of shared memory. </p><p>Two of those 152 are the tensor core; the rest is allocation, election, barrier phase tracking, address reconstruction and guardrails.</p><h3>The instruction</h3><pre><code><code>// dense f16
UTCHMMA      gdesc[UR14], gdesc[UR16], tmem[UR9], tmem[URZ], idesc[URZ], UPT ;
// same source with .sp: the sparsity metadata takes the fourth tmem slot
UTCHMMA      gdesc[UR14], gdesc[UR16], tmem[UR7], tmem[UR4], idesc[UR5], UPT ;
// same source with .cta_group::2
UTCHMMA.2CTA gdesc[UR12], gdesc[UR14], tmem[UR8], tmem[URZ], idesc[URZ], UPT ;
// NVFP4, block16 scaling: scale factors are a fifth tmem operand
UTCOMMA.4X   gdesc[UR14], gdesc[UR16], tmem[UR8], tmem[URZ], idesc[URZ], tmem[UR8], UPT ;</code></code></pre><p>Every operand is a descriptor or a Tensor Memory coordinate, and every one lives in a <strong>uniform register</strong>, the UR file, not the per-thread register file. The <code>U</code> prefix is the same <code>U</code> as in <code>UMOV</code>, <code>ULEA</code> and <code>UIADD3</code>: the scalar, warp-uniform datapath NVIDIA added in Turing for address arithmetic. Blackwell&#8217;s matrix multiply runs entirely on it. There is no vector register in the instruction at all.</p><p>That is the<strong> cleanest evidence</strong> we have found that the CTA, not the thread, is now the unit of tensor computation. It is not an abstraction in CUTLASS or a convenience in PTX. It is visible in which register file the opcode reads.</p><p>One clarification the SASS supports and the secondary literature often gets wrong: under <code>cta_group::2</code> the instruction is not issued by both CTAs in lock step. CUTLASS&#8217;s own <strong>two-SM tutorial states</strong> that only one of the two peer CTAs executes it, and names that one the leader. A single thread, in a single CTA, drives a matrix multiply spanning two SMs.</p><p>The dense form passes <code>tmem[URZ]</code>, the zero register, in the fourth slot; the sparse form fills it with a real address. </p><p>So <strong>structured sparsity metadata</strong> also lives in Tensor Memory, alongside the accumulator and, for block scaled kinds, alongside the scale factors. Three different kinds of state, one pool of 512 columns.</p><h3>The opcode family</h3><p>We compiled the probe once per qualifier and disassembled each result. This mapping is measured.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!TFo_!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15f09d28-0351-4bd9-9668-8393d162d188_2700x2244.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!TFo_!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15f09d28-0351-4bd9-9668-8393d162d188_2700x2244.png 424w, https://substackcdn.com/image/fetch/$s_!TFo_!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15f09d28-0351-4bd9-9668-8393d162d188_2700x2244.png 848w, https://substackcdn.com/image/fetch/$s_!TFo_!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15f09d28-0351-4bd9-9668-8393d162d188_2700x2244.png 1272w, https://substackcdn.com/image/fetch/$s_!TFo_!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15f09d28-0351-4bd9-9668-8393d162d188_2700x2244.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!TFo_!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15f09d28-0351-4bd9-9668-8393d162d188_2700x2244.png" width="1456" height="1210" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/15f09d28-0351-4bd9-9668-8393d162d188_2700x2244.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1210,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:89069,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/208027653?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15f09d28-0351-4bd9-9668-8393d162d188_2700x2244.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!TFo_!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15f09d28-0351-4bd9-9668-8393d162d188_2700x2244.png 424w, https://substackcdn.com/image/fetch/$s_!TFo_!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15f09d28-0351-4bd9-9668-8393d162d188_2700x2244.png 848w, https://substackcdn.com/image/fetch/$s_!TFo_!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15f09d28-0351-4bd9-9668-8393d162d188_2700x2244.png 1272w, https://substackcdn.com/image/fetch/$s_!TFo_!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15f09d28-0351-4bd9-9668-8393d162d188_2700x2244.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>This <strong>disagrees with the published literature</strong> in one place. The Delaware microbenchmarking paper, which is otherwise the most useful public measurement of this hardware and which we lean on later, reports in its Table IV that <code>tcgen05.mma</code> lowers to <code>HMMA</code>, <code>QMMA</code>, <code>OMMA</code> and <code>IMMA</code>, the same opcode names Volta through Hopper used. </p><p>On <strong>our disassembly</strong> it does not. The distinction is not cosmetic: <code>HMMA</code> reads and writes vector registers, <code>UTCHMMA</code> touches none. The most likely explanation is that the table was written from the family names rather than from a fresh disassembly.</p><h3>Which chips can run any of this</h3><p>Compiling against every target the assembler accepts produces a matrix sharper than the marketing, and the sharpest row is the one nobody mentions.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!StRx!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F551c9e65-c5d4-4eff-96b6-763c7bbb8028_900x592.svg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!StRx!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F551c9e65-c5d4-4eff-96b6-763c7bbb8028_900x592.svg 424w, https://substackcdn.com/image/fetch/$s_!StRx!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F551c9e65-c5d4-4eff-96b6-763c7bbb8028_900x592.svg 848w, https://substackcdn.com/image/fetch/$s_!StRx!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F551c9e65-c5d4-4eff-96b6-763c7bbb8028_900x592.svg 1272w, https://substackcdn.com/image/fetch/$s_!StRx!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F551c9e65-c5d4-4eff-96b6-763c7bbb8028_900x592.svg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!StRx!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F551c9e65-c5d4-4eff-96b6-763c7bbb8028_900x592.svg" width="1456" height="958" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/551c9e65-c5d4-4eff-96b6-763c7bbb8028_900x592.svg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:958,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:7800,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/svg+xml&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/208027653?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F551c9e65-c5d4-4eff-96b6-763c7bbb8028_900x592.svg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!StRx!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F551c9e65-c5d4-4eff-96b6-763c7bbb8028_900x592.svg 424w, https://substackcdn.com/image/fetch/$s_!StRx!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F551c9e65-c5d4-4eff-96b6-763c7bbb8028_900x592.svg 848w, https://substackcdn.com/image/fetch/$s_!StRx!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F551c9e65-c5d4-4eff-96b6-763c7bbb8028_900x592.svg 1272w, https://substackcdn.com/image/fetch/$s_!StRx!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F551c9e65-c5d4-4eff-96b6-763c7bbb8028_900x592.svg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><pre><code><code>$ ptxas -arch=sm_100  probe.ptx  -o /dev/null
ptxas error : Instruction 'tcgen05.alloc' not supported on .target 'sm_100'

$ ptxas -arch=sm_100a legacy.ptx -o /dev/null
ptxas error : Instruction 'wgmma.fence' not supported on .target 'sm_100a'

$ ptxas -arch=sm_120a legacy.ptx -o /dev/null
ptxas error : Instruction 'wgmma.fence' not supported on .target 'sm_120a'</code></code></pre><p>Three things to take from that table.</p><p>First, <code>wgmma</code> is not deprecated on Blackwell. It is <strong>removed</strong>, and it is removed from every Blackwell target including the consumer one. A Hopper kernel built on warpgroup MMA does not run slower on Blackwell, it fails at assembly time on all of <code>sm_100a</code>, <code>sm_103a</code> and <code>sm_120a</code>. </p><p>This is a <strong>harder break than any</strong> NVIDIA has shipped in the tensor core era, and it is worth stating precisely because the commonly repeated version of this claim, that consumer Blackwell falls back to <code>mma.sync</code> and <code>wgmma</code>, is wrong on the second half. There is no <code>wgmma</code> to fall back to.</p><p>Second, <code>sm_100</code> without the trailing <code>a</code>, the forward compatible target whose PTX a future driver may recompile for a future chip, has neither <code>tcgen05</code> nor <code>wgmma</code>. </p><p>The <strong>entire tensor path</strong> that defines Blackwell is available only on architecture specific targets, which by NVIDIA&#8217;s documented rules are not forward compatible with anything.</p><p>Third, the only instruction family that spans Hopper, datacenter Blackwell and consumer Blackwell is <code>mma.sync</code>, the warp-level, register-resident, <code>m16n8k16</code> class of instruction that predates all of this. That is the portable subset now. It is also the one whose accumulator lives in the register file, which is the constraint we said the architecture outgrew.</p><p>The <strong>portable path</strong> and the fast path have fully separated. What runs everywhere is the instruction whose limits forced Tensor Memory into existence.</p><p>The practical consequence is a development loop, not a benchmark. You cannot write, debug or profile a <code>tcgen05</code> kernel on a workstation. Not slowly, not at reduced fidelity, not at all. The instructions do not exist on hardware you can buy without a data center behind it. </p><p>In our compiler moat piece we argued the <strong>durable advantage sits at the </strong><code>ptxas</code><strong> and SASS layer</strong> and that the counter-technology is research rather than a new language. This adds a cruder second mechanism that has nothing to do with compilers: iteration on the instructions that matter now requires an allocation of scarce hardware.</p><p>We should immediately weaken that. The barrier is money and patience rather than access. B200 time is rentable by the hour, and the author of the best <code>tcgen05</code> tutorial we have read, writing as <strong>gau-nernst</strong>, reports reaching 98 percent of cuBLAS on a 4096 cubed problem using rented capacity. </p><p>A single motivated person got there in a blog series. That is way a much weaker moat than a first reading suggests.</p><h3>The cost of getting the answer back</h3><p><code>tcgen05.ld</code> takes a shape and a repetition count, and the count determines how many 32 bit values land in each thread&#8217;s registers. </p><p>We swept the count from 1 to 128, <strong>forced every loaded value</strong> to stay live by storing all of them, and read the allocation from <code>ptxas -v</code>.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!pZEC!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72920abc-baa3-4c1e-9133-956eea449577_1800x760.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!pZEC!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72920abc-baa3-4c1e-9133-956eea449577_1800x760.png 424w, https://substackcdn.com/image/fetch/$s_!pZEC!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72920abc-baa3-4c1e-9133-956eea449577_1800x760.png 848w, https://substackcdn.com/image/fetch/$s_!pZEC!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72920abc-baa3-4c1e-9133-956eea449577_1800x760.png 1272w, https://substackcdn.com/image/fetch/$s_!pZEC!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72920abc-baa3-4c1e-9133-956eea449577_1800x760.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!pZEC!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72920abc-baa3-4c1e-9133-956eea449577_1800x760.png" width="1456" height="615" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/72920abc-baa3-4c1e-9133-956eea449577_1800x760.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:615,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Measured register usage per thread as a function of the tcgen05.ld repetition count, rising from 14 registers at x1 to 134 registers at x128&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Measured register usage per thread as a function of the tcgen05.ld repetition count, rising from 14 registers at x1 to 134 registers at x128" title="Measured register usage per thread as a function of the tcgen05.ld repetition count, rising from 14 registers at x1 to 134 registers at x128" srcset="https://substackcdn.com/image/fetch/$s_!pZEC!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72920abc-baa3-4c1e-9133-956eea449577_1800x760.png 424w, https://substackcdn.com/image/fetch/$s_!pZEC!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72920abc-baa3-4c1e-9133-956eea449577_1800x760.png 848w, https://substackcdn.com/image/fetch/$s_!pZEC!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72920abc-baa3-4c1e-9133-956eea449577_1800x760.png 1272w, https://substackcdn.com/image/fetch/$s_!pZEC!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72920abc-baa3-4c1e-9133-956eea449577_1800x760.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em><strong><sup>Figure 4</sup></strong><sup> The epilogue tax, measured. A single ldtm.x128 costs 134 registers per thread. Across a warpgroup that is 67 KiB of the 256 KiB register file, just over a quarter of it, checked out purely to hold data on its way from one on-chip memory to another.</sup></em></p><p>The curve is N plus six, and <code>ptxas</code> emits one <code>LDTM.xN</code> rather than N loads. So the <strong>register file did not stop being the bottleneck</strong>. It stopped being the bottleneck for the multiply and became the bottleneck for the epilogue.</p><p>On a Hopper kernel the accumulator is already in registers when you want to scale it, add a bias, apply an activation and cast down. On Blackwell you must<strong> first pay to bring it back</strong>, and the wider you pay the less room remains to do anything with the result. </p><p>Every fused epilogue on this architecture is a negotiation between drain width and register headroom, and no source-level construct expresses it.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Why it had to be a separate memory</h2><p><strong>previous sections </strong>explained why the accumulator left the register file. It does not explain why the destination is a new addressable memory rather than a hidden internal buffer. </p><p>For that you need <strong>the traffic number</strong>, and the traffic number turns out to be more interesting than we expected when we first ran it.</p><p>Start with the shape rule, which is the load bearing fact. In CUTLASS&#8217;s Blackwell MMA traits the K extent of a UMMA atom is not a free parameter. It is computed:</p><pre><code><code>// cute/atom/mma_traits_sm100.hpp
// Logical shape-K is always 256bits, transform to units of elements
static constexpr int K = 256 / cute::sizeof_bits&lt;ValTypeA&gt;::value;</code></code></pre><p>K is always 256 bits, 32 bytes, of operand. FP16 gives K equal to 16. FP8 gives 32. FP4 gives 64. The K dimension scales as the inverse of element width, exactly.</p><p>Now compute the accumulator traffic. <strong>One UMMA</strong> over an M by N tile with FP32 accumulation reads the tile and writes it back, which is 8MN bytes, and performs 2MNK floating point operations. So:</p><pre><code><code>accumulator bytes per FLOP  =  8MN / (2MNK)  =  4 / K  =  (bits per element) / 64

FP16   K=16   0.2500 B/FLOP   x  2.25 PFLOP/s  =  562.5 TB/s
FP8    K=32   0.1250 B/FLOP   x  4.50 PFLOP/s  =  562.5 TB/s
FP4    K=64   0.0625 B/FLOP   x  9.00 PFLOP/s  =  562.5 TB/s

per SM (148):                                       3.80 TB/s
B200 HBM3e:                                         8.00 TB/s
ratio, chip accumulator traffic to HBM:               70 x</code></code></pre><p>The tile shape cancels, the clock cancels, and so does the precision. <strong>Accumulator traffic on a B200 is 562.5 terabytes per second at every precision the tensor cores support</strong>, because every time NVIDIA doubled the math rate by halving the element width, they simultaneously doubled K, which halved the accumulator traffic per FLOP. The two effects cancel to the digit.</p><p>That is not a coincidence and it is not a rounding artifact. It is the constraint the 256 bit K rule exists to satisfy. </p><p>Whatever structure holds the accumulator has to sustain a fixed bandwidth, and the format ladder was designed so that <strong>going from FP16 to FP4 buys four times the math</strong> without asking the accumulator path for a single additional byte per second.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!1r1K!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc8dd318-a701-4316-a238-1956fb4462b8_1800x800.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!1r1K!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc8dd318-a701-4316-a238-1956fb4462b8_1800x800.png 424w, https://substackcdn.com/image/fetch/$s_!1r1K!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc8dd318-a701-4316-a238-1956fb4462b8_1800x800.png 848w, https://substackcdn.com/image/fetch/$s_!1r1K!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc8dd318-a701-4316-a238-1956fb4462b8_1800x800.png 1272w, https://substackcdn.com/image/fetch/$s_!1r1K!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc8dd318-a701-4316-a238-1956fb4462b8_1800x800.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!1r1K!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc8dd318-a701-4316-a238-1956fb4462b8_1800x800.png" width="1456" height="647" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/bc8dd318-a701-4316-a238-1956fb4462b8_1800x800.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:647,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Chart showing math throughput rising four times from FP16 to FP4 while accumulator traffic stays constant at 562.5 terabytes per second&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Chart showing math throughput rising four times from FP16 to FP4 while accumulator traffic stays constant at 562.5 terabytes per second" title="Chart showing math throughput rising four times from FP16 to FP4 while accumulator traffic stays constant at 562.5 terabytes per second" srcset="https://substackcdn.com/image/fetch/$s_!1r1K!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc8dd318-a701-4316-a238-1956fb4462b8_1800x800.png 424w, https://substackcdn.com/image/fetch/$s_!1r1K!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc8dd318-a701-4316-a238-1956fb4462b8_1800x800.png 848w, https://substackcdn.com/image/fetch/$s_!1r1K!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc8dd318-a701-4316-a238-1956fb4462b8_1800x800.png 1272w, https://substackcdn.com/image/fetch/$s_!1r1K!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc8dd318-a701-4316-a238-1956fb4462b8_1800x800.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em><strong><sup>Figure 5.</sup></strong><sup> The invariance. This is the strongest single argument for why Tensor Memory is a fixed size, fixed geometry, precision-agnostic array: the bandwidth it must sustain does not depend on what numbers you put in it.</sup></em></p><h3>The other half, which explains where each operand lives</h3><p>Run the same calculation on the inputs and the architecture stops looking like a set of choices and starts looking like a single constraint solved twice.</p><p>One <strong>UMMA</strong> reads an M by K slab of A and a K by N slab of B, so operand bytes are (M + N) times K times the element width, against the same 2MNK operations. The K cancels here too:</p><pre><code><code>operand bytes per FLOP  =  (M+N)K b / (2MNK)  =  (M+N) b / (2MN)     // b = bytes per element

for the m128n256 tile:
FP16   b=2      0.01172 B/FLOP   x  2.25 PFLOP/s  =  26.4 TB/s
FP8    b=1      0.00586 B/FLOP   x  4.50 PFLOP/s  =  26.4 TB/s
FP4    b=0.5    0.00293 B/FLOP   x  9.00 PFLOP/s  =  26.4 TB/s

accumulator traffic / operand traffic  =  562.5 / 26.4  =  21.3 x</code></code></pre><p>Invariant again, and for the same reason: halving the element width halves the operand bytes per FLOP at exactly the rate it doubles the FLOPs. So the precision ladder from FP16 down to FP4 is <strong>bandwidth neutral at both ends of the datapath</strong>. </p><p>Blackwell quadrupled its peak math without asking either the operand path or the accumulator path for one additional byte per second.</p><p>That is also the answer to a question the operand rules raise and never explain. <em>Why must B sit in shared memory while D must sit somewhere else entirely?</em> Because the accumulator moves twenty one times the traffic. </p><p>Operands are read once per instruction and are narrow by construction; the <strong>accumulator is read and written in full</strong> every single time, in FP32, no matter how few bits the inputs have. A general purpose memory can serve the first workload. Nothing general purpose serves the second.</p><p>It also explains a fact that otherwise looks like an oversight: Blackwell&#8217;s shared memory did not grow. <strong>228 kilobytes per SM</strong>, identical to Hopper, in a generation that doubled math per SM per clock. It did not need to. </p><p>And where operand pressure did rise, NVIDIA answered with sharing rather than capacity: under <code>cta_group::2</code> two SMs in a <strong>texture processing cluster </strong>consume the same operands for one logical tile, which halves the per-SM operand traffic instead of doubling the memory that carries it.</p><p>562.5 terabytes per second across the chip is seventy times the entire HBM bandwidth of a B200, and 3.8 terabytes per second per SM sustained. Neither a cache nor a register file survives that. </p><p>What survives is a small, <strong>banked array</strong> physically adjacent to the consumer, addressed in the coordinate system the datapath already uses, and free of every general purpose obligation: no coherence, no cache tags, no arbitrary indexing, no participation in the memory model, no ability to be read by an ALU. </p><p>Which is precisely the <strong>list of things TMEM cannot do.</strong> The restrictions are not a first generation compromise to be relaxed later. They are the reason the number is achievable.</p><p>It is worth putting our figure next to the one measurement that exists. The Delaware group reports roughly 16 terabytes per second of TMEM read bandwidth on a B200. That is <strong>about four times </strong>the sustained accumulator requirement we derive, which is the right shape of answer: a peak port figure with headroom for drains overlapping accumulation, not a number that contradicts ours. </p><p>We would rather show both than pretend they measure the same thing.</p><p><strong>A caution about that paper&#8217;s peaks</strong>The same paper reports achieved throughputs as percentages of theoretical peak: FP4 at 7,700 TFLOPS being 96.2 percent, FP16 at 1,929.6 being 96.5 percent. Those imply peaks of about 8,004 and 1,999 TFLOPS, where the B200 datasheet says 9,000 and 2,250 dense. </p><p>Both of their implied peaks are <strong>11.1 percent below the datasheet</strong>, consistently, which is what you get from assuming a clock about 11 percent lower than the one the datasheet figures use. Their measurements are probably fine. </p><p>Their percentages are relative to a different baseline than NVIDIA&#8217;s, and should not be read as 96 percent of the number on the box.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Occupancy stops meaning what it meant</h2><p>Occupancy is the<strong> oldest performance heuristic</strong> in CUDA: resident warps per SM, limited by registers and shared memory, more of them meaning more latency to hide. </p><p>On Blackwell tensor kernels it is close to useless, and the reason is that a resource nobody&#8217;s occupancy calculator models now binds before the ones it does.</p><p>Start with what<strong> NVIDIA documents </strong>for compute capability 10.0, in the Blackwell tuning guide. </p><ul><li><p><em>Register file: 64 K 32-bit registers per SM, unchanged. </em></p></li><li><p><em>Maximum concurrent warps per SM: 64, unchanged since Volta. </em></p></li><li><p><em>Maximum thread blocks per SM: 32. </em></p></li><li><p><em>Shared memory capacity per SM: 228 kilobytes, the same as Hopper, with 227 addressable by a single block after CUDA&#8217;s 1 kilobyte reservation. </em></p></li><li><p><em>Combined L1, texture and shared memory: 256 kilobytes, also the same as Hopper.</em></p></li></ul><p>Read that list again with the previous sections in mind. Across a generation that doubled tensor throughput per SM per clock, <strong>not one of the classical occupancy resources grew</strong>. </p><p>The only capacity Blackwell added to the SM is the 256 kilobytes of Tensor Memory, and Tensor Memory is the one resource with an allocation rule that quantizes hard.</p><p>Columns come in powers of two, minimum 32, from a pool of 512, and <strong>every column carries all 128 lanes</strong>. So the pool admits 16 concurrent allocations at the smallest legal size, 8 at 64 columns, 4 at 128, 2 at 256, and 1 if you take the whole thing. There is no middle. And 16, the best case, is already half the hardware ceiling of 32 blocks per SM. </p><p>The moment your tile needs more than the minimum allocation, which any tile worth issuing a UMMA for does, <strong>Tensor Memory</strong> is the binding constraint and nothing else is close.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!v2Yk!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcacc8382-819f-430a-8d28-4985ba5a559f_1800x800.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!v2Yk!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcacc8382-819f-430a-8d28-4985ba5a559f_1800x800.png 424w, https://substackcdn.com/image/fetch/$s_!v2Yk!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcacc8382-819f-430a-8d28-4985ba5a559f_1800x800.png 848w, https://substackcdn.com/image/fetch/$s_!v2Yk!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcacc8382-819f-430a-8d28-4985ba5a559f_1800x800.png 1272w, https://substackcdn.com/image/fetch/$s_!v2Yk!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcacc8382-819f-430a-8d28-4985ba5a559f_1800x800.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!v2Yk!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcacc8382-819f-430a-8d28-4985ba5a559f_1800x800.png" width="1456" height="647" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/cacc8382-819f-430a-8d28-4985ba5a559f_1800x800.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:647,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Chart of concurrent tensor memory allocations permitted by a 512 column pool, falling from sixteen at thirty two columns to one at five hundred and twelve, against a hardware ceiling of thirty two thread blocks per SM&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Chart of concurrent tensor memory allocations permitted by a 512 column pool, falling from sixteen at thirty two columns to one at five hundred and twelve, against a hardware ceiling of thirty two thread blocks per SM" title="Chart of concurrent tensor memory allocations permitted by a 512 column pool, falling from sixteen at thirty two columns to one at five hundred and twelve, against a hardware ceiling of thirty two thread blocks per SM" srcset="https://substackcdn.com/image/fetch/$s_!v2Yk!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcacc8382-819f-430a-8d28-4985ba5a559f_1800x800.png 424w, https://substackcdn.com/image/fetch/$s_!v2Yk!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcacc8382-819f-430a-8d28-4985ba5a559f_1800x800.png 848w, https://substackcdn.com/image/fetch/$s_!v2Yk!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcacc8382-819f-430a-8d28-4985ba5a559f_1800x800.png 1272w, https://substackcdn.com/image/fetch/$s_!v2Yk!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcacc8382-819f-430a-8d28-4985ba5a559f_1800x800.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em><strong><sup>Figure 6.</sup></strong><sup> What the pool admits. This is an upper bound from the allocation rule, not a measurement of scheduler behaviour. Whether the hardware will actually co-resident that many CTAs is a separate question, and one we cannot settle without a B200.</sup></em></p><p>That caveat is not decoration. Colfax&#8217;s post on the workstation part, contrasting it with the datacenter part, states flatly that on SM10x <code>tcgen05.mma</code> is locked to one CTA per SM. <strong>Colfax&#8217;s own tutorial kernel</strong> allocates all 512 columns for a single 128 by 256 tile and never revisits the question. </p><p>And the <strong>PTX manual </strong>describes <code>tcgen05.relinquish_alloc_permit</code> as a promise that the CTA will make no further allocations, which Colfax glosses as allowing future CTAs to queue up for the same SM. The word queue is doing a lot of work in that sentence. </p><p>It is <em>consistent with a scheduler </em>that admits a new CTA only when the pool can serve it, which would collapse figure 6 toward 1 for any realistic tile regardless of the arithmetic.</p><p>We cannot resolve that without hardware, and we would rather show the bound and name the uncertainty than assert a residency number we have not seen. What survives either reading is the <strong>shape of the problem</strong>: the resource that limits parallelism on a Blackwell SM is acquired at runtime, quantizes in powers of two, and does not appear in any static occupancy model.</p><p>Now the part that makes low occupancy survivable, which is the more interesting half.</p><p>What occupancy hid was <em>memory latency</em>: a warp stalls on a load, the scheduler runs another warp. In a <strong>Blackwell GEMM</strong> the loads are done by the Tensor Memory Accelerator, asynchronously, into shared memory, signalled by an mbarrier. </p><p>The math is done by a <strong>single elected thread</strong> issuing an asynchronous instruction that reads shared memory and writes Tensor Memory. Neither heavy operation is a thread stalling on anything. Latency is hidden by <em>pipelining within one CTA</em>, using barriers and multiple buffers, rather than by <em>switching between CTAs</em>. </p><p>The author writing as gau-nernst describes exactly this in a working kernel: multiple <code>tcgen05.mma</code> in flight, one mbarrier per stage, so different MMA stages can be waited on independently.</p><p>The <strong>Delaware measurements</strong> make the same point from the other side. Single instruction latency for <code>tcgen05.mma</code> is 11.0 to 11.4 clocks and nearly flat across tile shapes from <code>m64n64k16</code> to <code>m256n256k16</code>. Hopper&#8217;s <code>wgmma</code> scales linearly with tile width, 32 clocks at <code>m64n64k16</code> and 128 at <code>m64n256k16</code>. </p><p>A flat latency across a <strong>sixteen fold range</strong> of tile area is the signature of a spatial array, not a deeper pipeline.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!2tlO!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89cac211-672c-4619-8795-02efdb7d63b2_1800x720.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!2tlO!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89cac211-672c-4619-8795-02efdb7d63b2_1800x720.png 424w, https://substackcdn.com/image/fetch/$s_!2tlO!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89cac211-672c-4619-8795-02efdb7d63b2_1800x720.png 848w, https://substackcdn.com/image/fetch/$s_!2tlO!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89cac211-672c-4619-8795-02efdb7d63b2_1800x720.png 1272w, https://substackcdn.com/image/fetch/$s_!2tlO!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89cac211-672c-4619-8795-02efdb7d63b2_1800x720.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!2tlO!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89cac211-672c-4619-8795-02efdb7d63b2_1800x720.png" width="1456" height="582" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/89cac211-672c-4619-8795-02efdb7d63b2_1800x720.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:582,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Grouped bar chart comparing single instruction latency of Hopper wgmma, which scales with tile width, against Blackwell tcgen05 which stays flat near eleven cycles&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Grouped bar chart comparing single instruction latency of Hopper wgmma, which scales with tile width, against Blackwell tcgen05 which stays flat near eleven cycles" title="Grouped bar chart comparing single instruction latency of Hopper wgmma, which scales with tile width, against Blackwell tcgen05 which stays flat near eleven cycles" srcset="https://substackcdn.com/image/fetch/$s_!2tlO!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89cac211-672c-4619-8795-02efdb7d63b2_1800x720.png 424w, https://substackcdn.com/image/fetch/$s_!2tlO!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89cac211-672c-4619-8795-02efdb7d63b2_1800x720.png 848w, https://substackcdn.com/image/fetch/$s_!2tlO!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89cac211-672c-4619-8795-02efdb7d63b2_1800x720.png 1272w, https://substackcdn.com/image/fetch/$s_!2tlO!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89cac211-672c-4619-8795-02efdb7d63b2_1800x720.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em><strong><sup>Figure 7.</sup></strong><sup> Hopper&#8217;s MMA latency grows with the tile. Blackwell&#8217;s does not. Combined with single thread issue and a TMEM resident accumulator, this is what makes very low CTA counts survivable.</sup></em></p><p>So the tuning knobs invert. On Hopper you asked how many warpgroups fit and how to specialize them. On Blackwell you ask how many columns your accumulator needs, how many buffers fit in what remains, and whether the tile is <strong>worth a CTA pair</strong>. </p><p>The Nsight metric that used to matter, achieved occupancy, tells you almost nothing. The metric that matters has no counter.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Quantization became an allocation problem</h2><p>Here is the part we think is genuinely underappreciated, and it is the point where <strong>this architecture stops being a matter for kernel authors</strong> and starts being a matter for anyone who chooses a serving format.</p><p>Blackwell implements block scaling in hardware. A block scaled MMA computes <em>D = C + (A x SFA) times (B x SFB)</em>, where SFA and SFB are vectors of scale factors, one per group of 16 or 32 elements along K. </p><p>The scale factors are not folded in beforehand and they are not applied afterwards. They are consumed by the tensor core as it runs. And, per the PTX ISA, <strong>they are consumed from Tensor Memory</strong>.</p><p>My disassembly shows this directly. The block scaled variants take a third and fourth <code>tmem[]</code> operand:</p><pre><code><code>// kind::mxf4nvf4.block_scale.scale_vec::4X, NVFP4 with block16 scaling
UTCOMMA.4X gdesc[UR14], gdesc[UR16], tmem[UR8], tmem[URZ], idesc[URZ], tmem[UR8], UPT ;

// kind::mxf4.block_scale.scale_vec::2X, MXFP4 with block32 scaling
UTCOMMA    gdesc[UR14], gdesc[UR16], tmem[UR8], tmem[URZ], idesc[URZ], tmem[UR8], UPT ;</code></code></pre><p>Read the consequence carefully. Your choice of numerical format now consumes the same scarce, power of two, 32 column granular, 512 column per SM resource that your accumulator consumes. </p><p>Finer scaling is not just more metadata bandwidth from HBM. It is columns you cannot use for the accumulator, for double buffering, or for a <strong>second CTA</strong>.</p><p>The two formats differ exactly where it hurts. MXFP4 follows the Open Compute microscaling specification: blocks of 32, scale in E8M0, a bare power of two exponent. </p><p><strong>NVFP4</strong> is NVIDIA&#8217;s own: blocks of 16, scale in E4M3, a real floating point number with a mantissa, plus a second level FP32 tensor-wide scale. Per the PTX rules, block32 must pair with E8M0, while block16 may use E8M0 or E4M3.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!MHla!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ea35e78-0dc7-4686-ba58-ac915d1957db_1800x760.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!MHla!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ea35e78-0dc7-4686-ba58-ac915d1957db_1800x760.png 424w, https://substackcdn.com/image/fetch/$s_!MHla!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ea35e78-0dc7-4686-ba58-ac915d1957db_1800x760.png 848w, https://substackcdn.com/image/fetch/$s_!MHla!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ea35e78-0dc7-4686-ba58-ac915d1957db_1800x760.png 1272w, https://substackcdn.com/image/fetch/$s_!MHla!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ea35e78-0dc7-4686-ba58-ac915d1957db_1800x760.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!MHla!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ea35e78-0dc7-4686-ba58-ac915d1957db_1800x760.png" width="1456" height="615" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/1ea35e78-0dc7-4686-ba58-ac915d1957db_1800x760.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:615,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Chart of effective bits per parameter and scale factor overhead for MXFP8, MXFP4 and NVFP4&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Chart of effective bits per parameter and scale factor overhead for MXFP8, MXFP4 and NVFP4" title="Chart of effective bits per parameter and scale factor overhead for MXFP8, MXFP4 and NVFP4" srcset="https://substackcdn.com/image/fetch/$s_!MHla!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ea35e78-0dc7-4686-ba58-ac915d1957db_1800x760.png 424w, https://substackcdn.com/image/fetch/$s_!MHla!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ea35e78-0dc7-4686-ba58-ac915d1957db_1800x760.png 848w, https://substackcdn.com/image/fetch/$s_!MHla!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ea35e78-0dc7-4686-ba58-ac915d1957db_1800x760.png 1272w, https://substackcdn.com/image/fetch/$s_!MHla!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ea35e78-0dc7-4686-ba58-ac915d1957db_1800x760.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em><strong><sup>Figure 8</sup></strong><sup> NVFP4 halves the block size and puts a mantissa in the scale, which is why it is more accurate than MXFP4. It also doubles the scale factor footprint, and that footprint lives in the same 512 column pool as the accumulator.</sup></em></p><p>This is a co-design decision hiding inside a numerics decision. NVIDIA defined a format whose <strong>accuracy advantage</strong> over the open standard comes from finer blocks and richer scales, then built the only silicon where those scales are consumed directly out of a dedicated on-chip memory rather than being unpacked into registers first. </p><p>On hardware without that memory, the same format is a software dequantization problem with a register cost. On Blackwell it is an operand.</p><p>The Delaware group reports the accuracy side: FP8 costs about 2 percent perplexity on Mistral 7B and Mixtral 8x7B, FP4 costs 8 to 9 percent, and their FP4 throughput is 2.5 times FP16 on the dense model and 2.7 times on the mixture of experts model. </p><p>Whether <strong>8 percent perplexity</strong> is acceptable is a per-layer question and always was. My point is narrower: on this architecture the answer is also a per-SM capacity planning question, because the scale factors and the accumulator compete for the same 512 columns.</p><p>Numerics stopped being a property of the model and became a property of the memory allocator.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>What this does to serving economics</h2><p>Everything above is architecture. Here is the part that shows up on an invoice.</p><p><strong>Two rules</strong> from earlier sections collide here. The UMMA shape table offers exactly two values of M for a single CTA, 64 and 128; there is nothing smaller. </p><p>And allocation is by whole column, so the accumulator&#8217;s footprint is fixed by the tile you chose, not by the rows you filled.</p><p>During prefill neither rule bites. M is the <strong>token count of a chunk</strong>, thousands of rows, and the tile is full. During decode both bite at once. A decode step is a GEMM whose M is the number of sequences you are batching, and it is skinny by construction, so you pick the 64-row tile and leave most of it empty.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!20ki!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9031d001-8aac-4140-8df0-e8d90d0da36c_1800x800.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!20ki!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9031d001-8aac-4140-8df0-e8d90d0da36c_1800x800.png 424w, https://substackcdn.com/image/fetch/$s_!20ki!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9031d001-8aac-4140-8df0-e8d90d0da36c_1800x800.png 848w, https://substackcdn.com/image/fetch/$s_!20ki!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9031d001-8aac-4140-8df0-e8d90d0da36c_1800x800.png 1272w, https://substackcdn.com/image/fetch/$s_!20ki!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9031d001-8aac-4140-8df0-e8d90d0da36c_1800x800.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!20ki!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9031d001-8aac-4140-8df0-e8d90d0da36c_1800x800.png" width="1456" height="647" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/9031d001-8aac-4140-8df0-e8d90d0da36c_1800x800.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:647,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Chart showing what fraction of the smallest legal UMMA accumulator tile carries a real sequence during decode, rising from 1.6 percent at batch one to 100 percent at batch sixty four&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Chart showing what fraction of the smallest legal UMMA accumulator tile carries a real sequence during decode, rising from 1.6 percent at batch one to 100 percent at batch sixty four" title="Chart showing what fraction of the smallest legal UMMA accumulator tile carries a real sequence during decode, rising from 1.6 percent at batch one to 100 percent at batch sixty four" srcset="https://substackcdn.com/image/fetch/$s_!20ki!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9031d001-8aac-4140-8df0-e8d90d0da36c_1800x800.png 424w, https://substackcdn.com/image/fetch/$s_!20ki!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9031d001-8aac-4140-8df0-e8d90d0da36c_1800x800.png 848w, https://substackcdn.com/image/fetch/$s_!20ki!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9031d001-8aac-4140-8df0-e8d90d0da36c_1800x800.png 1272w, https://substackcdn.com/image/fetch/$s_!20ki!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9031d001-8aac-4140-8df0-e8d90d0da36c_1800x800.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em><strong><sup>Figure 9</sup></strong><sup> The decode penalty, stated against the most favourable tile the instruction set offers. At a batch of eight, an ordinary steady state for a latency sensitive endpoint, an eighth of the reserved accumulator carries anything. The rest is allocated, idle, and unavailable to any other block on that SM.</sup></em></p><p>Be precise about what is and is not new, because it is easy to overstate. The 64-row floor is not new: Hopper&#8217;s <code>wgmma</code> had exactly the same minimum M, and skinny GEMMs have underused tensor cores since Volta. </p><p>That is why <strong>prefill and decode</strong> <strong>get disaggregated</strong> onto separately provisioned pools in the first place, which we argued at length in a previous issue.</p><p>What is new is where the waste is recorded. Previously it was purely a <em>throughput</em> loss: you issued an MMA whose M dimension was mostly zeros and you got a fraction of peak flops, and the moment the instruction retired the machine was free again. </p><p>Now the same tile also holds a <em>capacity</em> reservation, in a 512 column pool, acquired through an atomic that other CTAs may be spinning on, held for the tile&#8217;s lifetime, and invisible to every static occupancy model. Underutilisation stopped being a transient and became an allocation.</p><p><strong>Three practical consequences</strong> follow, and we hold them with decreasing confidence.</p><ol><li><p><strong>The first is that Blackwell widens the gap between prefill and decode economics</strong> rather than narrowing it, in a generation whose marketing is entirely about inference. The FP4 tensor cores are a prefill and large-batch story, and section 05 is the reason: the precision ladder is bandwidth neutral on chip, so what FP4 buys you is math you can only spend if you have rows to fill. Decode has neither. It remains bound by HBM bandwidth for weight movement, and now carries an on-chip capacity reservation on top. This is consistent with the Delaware measurements from a different angle: as precision drops from FP16 to FP4 their measured memory bandwidth utilization <em>falls</em> from 67 percent to 48 percent, which is what it looks like when a workload stops being bandwidth bound and starts being bound by something else.</p></li><li><p><strong>The second</strong> <strong>is that this pushes harder</strong> <strong>toward every technique that manufactures M</strong>. Speculative decoding turns one sequence into k candidate tokens per step, and on Blackwell it is not only amortizing weight reads across more rows, it is filling rows of a tile that was reserved whether or not you filled them. </p><blockquote><p><em><strong>Figure 9</strong>, read the other way, is a chart of how much accumulator a speculative draft gets for free. We modelled speculation as a throughput question in a previous issue. There is a second term in that model now, and it points the same way.</em></p></blockquote></li><li><p><strong>The third, and the one we are least sure of, is that grouped GEMM for mixture of experts becomes a harder allocation problem</strong> than it was. Each expert&#8217;s tile has its own M, determined by routing, varying per step. If your kernel allocates for the worst case it wastes columns on every expert that got fewer tokens; if it allocates per expert it pays the atomic repeatedly. </p></li></ol><p>The one worklog we have found on optimizing NVFP4 grouped GEMM on Blackwell, by Mufeez Amjad, lands on <code>cta_group::2</code> with a shared TMEM allocation across the CTA pair, and describes the shared allocation as one of the main benefits rather than the wider math tile. That is a hint about where the pressure actually is.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>What Rubin says about whether this was a one-off</h2><p>A reasonable objection to everything above is that <strong>Tensor Memory might be a Blackwell specific hack</strong>, a way to ship a doubled MMA without redesigning the register file, and that the next architecture folds it back into something more general. </p><p>NVIDIA published enough about <strong>Rubin on 21 July 2026 </strong>to answer that. It is not going away, and every disclosed change points at the numbers in section 05.</p><p>Three details in NVIDIA&#8217;s own post are load bearing.</p><ol><li><p><strong>Tensor Memory gained a consumer that is not the MMA.</strong> In the long context attention path, the intermediate scores from the dense QK transpose are, in NVIDIA&#8217;s words, loaded from Tensor Memory into a structured 2 to 4 sparse compressed form, generating both the nonzero values and the metadata. TMEM is no longer only where accumulators land. It is a staging tier that a hardware compression path reads from. That is the same direction our disassembly already showed on Blackwell, where sparsity metadata and block scale factors are already TMEM operands.</p></li><li><p><strong>Rubin doubles the K dimension per tensor core instruction.</strong> NVIDIA frames this as fewer K loop iterations and less loop overhead, which is true and is the reason a reader would care. Run it through section 05 and it is also something else. Accumulator bytes per FLOP is 4 over K. Doubling K halves it. If Rubin&#8217;s NVFP4 K goes from 64 to 128, accumulator traffic per unit of math drops by half at exactly the moment the math rate goes up. That is the same invariance trick applied across a generation instead of across a precision ladder, and it is the strongest evidence we have that accumulator bandwidth is a first order design constraint at NVIDIA rather than a consequence.</p></li><li><p><strong>Softmax became the bottleneck, so they widened it.</strong> Rubin raises exponential throughput per clock per SM by 2 times for FP32 and 4 times for BF16 against Blackwell, with Blackwell Ultra at 2 times for both. You only build that if the matrix path has already pulled far enough ahead that the transcendental path is what is left.</p></li></ol><p>Alongside those, Rubin adds inline descriptor updates for the Tensor Memory Accelerator, so a <strong>mixture of experts kernel </strong>keeps one descriptor and overrides the pointer and stride fields in the instruction instead of rewriting a descriptor in memory per expert, and counted writes for <strong>device initiated NVLink transfers</strong> so the receiver tracks completion without the barrier, acknowledgement and atomic flag sequence.</p><p>Every one of those is the same move: take a coordination cost that was paid in <strong>general purpose instructions and registers</strong>, and pay it in a dedicated mechanism instead. Tensor Memory was the first large instance. Rubin is the pattern applied to descriptors, to synchronization, to sparsity metadata and to the softmax path.</p><p>One sentence for what NVIDIA is doing to the SM across these two generations: they are disassembling the<strong> general purpose core</strong> into special purpose engines connected by dedicated memories and barriers, and leaving the threads to do bookkeeping. </p><p>The execution model is being hollowed out from the inside while its surface syntax stays the same.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Where this could be wrong</h2><ul><li><p><strong>The accumulator traffic derivation assumes a full read and write per instruction.</strong> If the tensor core retains partial accumulator state internally across the K extent, or streams only the sub-tile currently in flight, then 562.5 terabytes per second is an upper bound. What makes us reasonably confident is not the absolute number but the invariance: the cancellation between element width and K is exact, holds at three precisions, and falls out of a rule visible in CUTLASS source. A different traffic model would change the constant and probably not the invariance. Note that this is a revision of an earlier draft of ours, which held K fixed at 16 across precisions and therefore produced a per-SM figure four times too large at FP4.</p></li><li><p><strong>We are reading the SASS of a probe, not of a real kernel.</strong> A kernel with one MMA gives <code>ptxas</code> no scheduling problem to solve. The allocation spin loop may be scheduled very differently, hoisted, or dominated by something else in a CUTLASS mainloop with a producer warpgroup and four pipeline stages. We are confident the instruction sequence exists and that the guardrails are unconditional. We are not confident about its cost in situ, and the 152 instruction figure is a property of our probe, not of production kernels.</p></li><li><p><strong>Figure 6 is an upper bound, not a residency measurement.</strong> The allocation arithmetic is exact, and the documented block and warp ceilings for compute capability 10.0 are exact. What we cannot verify is whether the scheduler will actually co-resident the CTAs the pool arithmetically admits. One published source states that SM10x is locked to one CTA per SM outright, and the PTX description of <code>relinquish_alloc_permit</code> hints at queueing rather than co-residency. If that reading is right, the correct version of figure 6 is a flat line at 1, the argument gets stronger rather than weaker, and our chart is still wrong.</p></li><li><p><strong>Figure 9 measures reservation, not throughput, and only for the tile we chose.</strong> The 64 row floor is documented and the arithmetic against it is exact, but a real decode kernel may prefer a different shape, may reuse one allocation across many tiles, or may batch several matrices into a single call in ways that change what fraction of the pool sits idle at any instant. The claim we will defend is narrow: below 64 sequences the accumulator is reserved for rows that do not exist. What that costs in dollars depends on a serving stack we have not profiled.</p></li><li><p><strong>The Blackwell per-SM per-clock figure of 8,192 is inferred, not published.</strong> Volta, Ampere and Hopper rates come from NVIDIA whitepapers and reproduce the published teraflops to three digits. For Blackwell we ran the identity backwards from 2.25 petaflops over 148 SMs, which requires 8,192 FLOPs per clock at 1.86 GHz. If the real SM count in the shipping part differs, or if the datasheet figure assumes a different clock, the doubling claim survives but the exact number moves.</p></li><li><p><strong>The portability finding is solid; the moat reading of it is not.</strong> Section 04 already discounts it and we would discount it further rather than less. The instruction family genuinely does not exist outside datacenter Blackwell. What follows from that commercially is much weaker than it first sounds.</p></li></ul><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share The Software Frontier&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share The Software Frontier</span></a></p><div><hr></div><h2>Four predictions, each with its failure condition</h2><p>Deliberately conservative, and each one checkable by a specific date against a specific artifact.</p><ol><li><p><strong>NVIDIA will not expose Tensor Memory to a general purpose instruction through the end of 2028.</strong> No <code>ld</code>, <code>st</code>, atomic or ALU operand outside the <code>tcgen05</code> family, in Rubin or its successor. The invariance in section 05 is the reason: general purpose addressability is incompatible with the bandwidth. <em>Wrong if</em> a PTX ISA revision adds any TMEM access outside the dedicated opcode family.</p></li><li><p><strong>By 31 December 2027, no open source compiler will generate a </strong><code>tcgen05</code><strong> GEMM within 10 percent of CUTLASS on a mainstream shape without hand written PTX in its lowering path.</strong> Reaching the instruction is already happening. Reaching it through a general lowering that models column allocation, the warp-lane access partition and the drain width tradeoff is a different problem. <em>Wrong if</em> a mainline release of Triton, Mojo or tinygrad hits the bar with a pure compiler path.</p></li><li><p><strong>Nsight Compute will ship a Tensor Memory occupancy or column pressure section before the end of 2028.</strong> The hardware is already tracking the pool, since the allocator does a find-and-set against it, and there is currently no static way to reason about a resource acquired at runtime. <em>Wrong if</em> no NVIDIA profiler release by then reports TMEM allocation state.</p></li><li><p><strong>NVFP4 will remain the default four bit format in NVIDIA&#8217;s own inference libraries through 2027, and MXFP4 will not displace it there.</strong> Not an accuracy argument: the block16 path is the one with a dedicated opcode modifier and a hardware operand path on the silicon that matters. <em>Wrong if</em> TensorRT-LLM or NVIDIA&#8217;s quantization tooling defaults to block32 for four bit weights.</p></li></ol><p>We have dropped two predictions that appeared in an earlier draft, on model architectures being shaped around column boundaries and on serving share ratios between formats, because neither had a failure condition we could actually check.</p><div class="directMessage button" data-attrs="{&quot;userId&quot;:201922174,&quot;userName&quot;:&quot;Lorenzo Bradanini&quot;,&quot;canDm&quot;:null,&quot;dmUpgradeOptions&quot;:null,&quot;isEditorNode&quot;:true}" data-component-name="DirectMessageToDOM"></div><div><hr></div><h2>Confidence dossier</h2><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!uiBL!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1d544e02-60d4-4544-88ef-dd831cd8be82_2700x2472.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!uiBL!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1d544e02-60d4-4544-88ef-dd831cd8be82_2700x2472.png 424w, https://substackcdn.com/image/fetch/$s_!uiBL!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1d544e02-60d4-4544-88ef-dd831cd8be82_2700x2472.png 848w, https://substackcdn.com/image/fetch/$s_!uiBL!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1d544e02-60d4-4544-88ef-dd831cd8be82_2700x2472.png 1272w, https://substackcdn.com/image/fetch/$s_!uiBL!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1d544e02-60d4-4544-88ef-dd831cd8be82_2700x2472.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!uiBL!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1d544e02-60d4-4544-88ef-dd831cd8be82_2700x2472.png" width="1456" height="1333" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/1d544e02-60d4-4544-88ef-dd831cd8be82_2700x2472.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1333,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:192773,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/208027653?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1d544e02-60d4-4544-88ef-dd831cd8be82_2700x2472.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!uiBL!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1d544e02-60d4-4544-88ef-dd831cd8be82_2700x2472.png 424w, https://substackcdn.com/image/fetch/$s_!uiBL!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1d544e02-60d4-4544-88ef-dd831cd8be82_2700x2472.png 848w, https://substackcdn.com/image/fetch/$s_!uiBL!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1d544e02-60d4-4544-88ef-dd831cd8be82_2700x2472.png 1272w, https://substackcdn.com/image/fetch/$s_!uiBL!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1d544e02-60d4-4544-88ef-dd831cd8be82_2700x2472.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Reproduce every tier A row on a machine with no GPU</h2><p>Everything in tier A came out of the following, in a clean container, in under five minutes, with no CUDA installation and no driver.</p><pre><code><code># 1. the toolchain, as ordinary x86 binaries from PyPI
pip download nvidia-cuda-nvcc-cu12 --no-deps -d /tmp/w            # ptxas 12.9.86
pip download nvidia-cuda-nvdisasm nvidia-cuda-cuobjdump --no-deps -d /tmp/w   # 13.3.73
mkdir -p /tmp/tk &amp;&amp; cd /tmp/tk &amp;&amp; for f in /tmp/w/*.whl; do unzip -oq "$f"; done
PTXAS=/tmp/tk/nvidia/cuda_nvcc/bin/ptxas
DIS=/tmp/tk/nvidia/cu13/bin/nvdisasm

# A1, A2, A3, A4: assemble the probe and read the machine code
$PTXAS -arch=sm_100a probe.ptx -o probe.cubin &amp;&amp; $DIS -c probe.cubin | less
$DIS -c probe.cubin | grep -oE '__cuda_sm10x_tcgen05_[a-z_]+' | sort -u
for o in 0 1 2 3; do $PTXAS -O$o -arch=sm_100a probe.ptx -o g.cubin
  echo "-O$o $($DIS -c g.cubin | grep -c guardrail) refs, \
$($DIS -c g.cubin | grep -cE '^\s+/\*[0-9a-f]{4}\*/') instructions"; done

# A1 continued: one build per .kind qualifier
for k in f16 tf32 f8f6f4 i8; do sed "s/kind::f16/kind::$k/" probe.ptx &gt; k.ptx
  $PTXAS -arch=sm_100a k.ptx -o k.cubin &amp;&amp; $DIS -c k.cubin | grep -oE 'UTC[A-Z0-9.]+'; done

# A5: three instruction families across five targets
for a in sm_90a sm_100a sm_103a sm_100 sm_120a; do
  for f in probe legacy_wgmma legacy_mmasync; do
    printf '%-9s %-16s ' $a $f; $PTXAS -arch=$a $f.ptx -o /dev/null 2&gt;&amp;1 | head -1; echo; done; done

# A6: minimum PTX ISA version
for v in 8.3 8.4 8.5 8.6 8.7 8.8; do sed "s/^.version 8.6/.version $v/" probe.ptx &gt; v.ptx
  echo -n "$v "; $PTXAS -arch=sm_100a v.ptx -o /dev/null 2&gt;&amp;1 | head -1; echo; done

# A7: the epilogue register curve. widen tcgen05.ld to .xN and store every register
for n in 1 2 4 8 16 32 64 128; do $PTXAS -arch=sm_100a -v ld_$n.ptx -o /dev/null 2&gt;&amp;1 | grep Used; done

# A8, A9: the ceilings the compiler will and will not enforce
$PTXAS -arch=sm_100a -maxrregcount=256 probe.ptx -o /dev/null
for c in 16 48 96 1024; do sed "s/mov.u32 %r1, 128;/mov.u32 %r1, $c;/" probe.ptx &gt; a.ptx
  $PTXAS -arch=sm_100a a.ptx -o /dev/null &amp;&amp; echo "$c columns: accepted"; done</code></code></pre><p><strong>Two practical notes</strong>. <code>ptxas</code> from the 12.9 wheel caps at PTX ISA 8.8, which covers <code>sm_103a</code> and nothing above it, so pull a newer <code>nvidia-cuda-nvcc</code> for later targets. And <code>nvdisasm</code> from the CUDA 13 wheels reads CUDA 12 cubins, which is convenient because the CUDA 13 nvcc wheel did not build in our container while the standalone disassembler wheels did.</p><p>The <strong>value of this workflow</strong> is not that it replaces a GPU. It is that it separates two questions that get conflated constantly in GPU writing: what the machine <em>does</em>, which needs hardware, and what the compiler <em>emits</em>, which does not. </p><p>A surprising share of public claims about the CUDA moat are <strong>claims of the second kind</strong>, and the second kind is checkable by anyone with forty megabytes of disk.</p><div><hr></div><h2>Bibliography</h2><ol><li><p>NVIDIA, <em>Parallel Thread Execution ISA</em>: tensor memory addressing, tcgen05 MMA and its kind shapes, shared memory descriptors, instruction descriptors, data path layout organization, and the tcgen05 memory consistency model. The primary source for everything structural here.</p></li><li><p>NVIDIA, <em>Blackwell Tuning Guide</em> and <em>Ampere Tuning Guide</em>. Register file size, the register per thread ceiling, warp and thread block limits per SM, and shared memory capacities for compute capabilities 8.0, 10.0 and 12.0.</p></li><li><p>NVIDIA Volta, Ampere and Hopper architecture whitepapers. Source for 1,024, 2,048 and 4,096 FP16 tensor FLOPs per clock per SM, each of which reproduces the corresponding published teraflops figure to three digits.</p></li><li><p>NVIDIA, <em>HGX B200 datasheet</em>. Dense and sparse rates per precision, with the explicit note that dense is half of sparse, plus memory capacity and bandwidth.</p></li><li><p>Ryo, <em>CUTLASS Tutorial: Writing GEMM Kernels Using Tensor Memory For NVIDIA Blackwell GPUs</em>, Colfax Research, April 2025, updated November 2025. The clearest published account of TMEM allocation, UMMA operand rules and the CuTe abstractions over both.</p></li><li><p>Colfax Research, <em>CUTLASS Tutorial: Hardware-supported Block-scaling with NVIDIA Blackwell GPUs</em>. Scale vector rules and TMEM layouts of scale factors.</p></li><li><p>Colfax Research, <em>NVFP4 Blockscaled GEMM on NVIDIA RTX Pro Blackwell GPUs (SM12x)</em>, June 2026. The explicit statement that SM12x has neither tcgen05 nor TMEM, and that SM10x is locked to one CTA per SM.</p></li><li><p>NVIDIA, <em>CUTLASS</em> source, <code>include/cute/atom/mma_traits_sm100.hpp</code> and <code>copy_traits_sm100.hpp</code>. The 256 bit K rule, the TMEM copy atoms, and the ThrID change.</p></li><li><p>Aaron Jarmusch and Sunita Chandrasekaran, University of Delaware, <em>Microbenchmarking NVIDIA&#8217;s Blackwell Architecture: An in-depth Architectural Analysis</em>, arXiv 2512.02189. The only systematic public measurement of B200 tensor core latency, TMEM behaviour and FP4 accuracy we are aware of.</p></li><li><p>NVIDIA, <em>Inside NVIDIA Rubin GPU Architecture: Powering the Era of Agentic AI</em>, 21 July 2026.</p></li><li><p>NVIDIA, <em>CUTLASS documentation, Blackwell SM100 functionality</em>. The seven tcgen05.mma instructions and the block scaled data types.</p></li><li><p>SemiAnalysis, <em>Dissecting NVIDIA Blackwell: Tensor Cores, PTX Instructions, SASS, Floorsweep, Yield</em>. The CuTe ThrID observation and the TPC scoped CTA pair framing.</p></li><li><p>gau-nernst, <em>tcgen05 for dummies</em>, December 2025. A working tutorial in plain CUDA and inline PTX.</p></li><li><p>Mufeez Amjad, <em>Optimizing NVFP4 Grouped GEMM on Blackwell</em>, worklog, March 2026. The CTA pair and shared TMEM allocation observation in section 08.</p></li><li><p>Rouhani et al., <em>Microscaling Data Formats for Deep Learning</em>, arXiv 2310.10537. The MX specification behind MXFP4 and MXFP8.</p></li><li><p>NVIDIA developer forums, thread on computing tensor core FP16 throughput per SM per clock from whitepaper figures. The identity used in section 01.</p></li><li><p>Chips and Cheese, <em>Nvidia&#8217;s B200: Keeping the CUDA Juggernaut Rolling</em>, December 2025. Die level SM counts: 74 enabled of 80 per die.</p></li><li><p>Austin et al., <em>How To Scale Your Model</em>, GPU chapter. Independent derivation of Blackwell&#8217;s per-SM per-clock tensor rate from the same datasheet figures.</p></li><li><p>Our own previous work: <em>How the NVIDIA Compiler Moat Actually Works</em>, <em>How CUDA Binaries Actually Work</em>, <em>The Split and the Seam</em> on prefill and decode disaggregation, and <em>The Draft and the Ledger</em> on speculative decoding economics.</p><p></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share The Software Frontier&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share The Software Frontier</span></a></p><p></p></li></ol>]]></content:encoded></item><item><title><![CDATA[How CUDA Binaries Actually Work]]></title><description><![CDATA[A byte-level surgical dissection of the cubin and fatbin formats, and the second encoding nobody documents.]]></description><link>https://www.thesoftwarefrontier.com/p/how-cuda-binaries-actually-work</link><guid isPermaLink="false">https://www.thesoftwarefrontier.com/p/how-cuda-binaries-actually-work</guid><dc:creator><![CDATA[Lorenzo Bradanini]]></dc:creator><pubDate>Fri, 31 Jul 2026 12:45:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!xYOt!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F68806f07-c2e9-44a7-a9db-9dd6ff086db4_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!xYOt!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F68806f07-c2e9-44a7-a9db-9dd6ff086db4_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!xYOt!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F68806f07-c2e9-44a7-a9db-9dd6ff086db4_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!xYOt!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F68806f07-c2e9-44a7-a9db-9dd6ff086db4_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!xYOt!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F68806f07-c2e9-44a7-a9db-9dd6ff086db4_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!xYOt!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F68806f07-c2e9-44a7-a9db-9dd6ff086db4_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!xYOt!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F68806f07-c2e9-44a7-a9db-9dd6ff086db4_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/68806f07-c2e9-44a7-a9db-9dd6ff086db4_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2689230,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/207127260?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F68806f07-c2e9-44a7-a9db-9dd6ff086db4_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!xYOt!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F68806f07-c2e9-44a7-a9db-9dd6ff086db4_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!xYOt!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F68806f07-c2e9-44a7-a9db-9dd6ff086db4_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!xYOt!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F68806f07-c2e9-44a7-a9db-9dd6ff086db4_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!xYOt!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F68806f07-c2e9-44a7-a9db-9dd6ff086db4_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>Intro</h2><p>I spent a good part of last year writing an <strong>x86-64 assembler</strong> by hand, in C, for no more reason than wanting to know what an object file really is. The thing that surprised me was not the instruction encoding, but was the metadata. </p><p>A <strong>modern object file</strong> is essentially a negotiation between the compiler and the loader about few things the machine code itself cannot say properly: where the arguments are, how much stack to reserve, which symbols are entry points, what has to be patched before anything runs.</p><p>So when I started reading GPU binaries, the question I kept asking more and more was not &#8220;<em>what does this instruction do.</em>&#8221; It was &#8220;<em>what is this file promising the driver</em>.&#8221;</p><p>The answer turns out to be&#8230; well, a lot, and almost none of it is written down. NVIDIA just documents the container in one sentence and the disassembler in twelve pages. And that&#8217;s basically it. </p><p>The <a href="https://docs.nvidia.com/cuda/cuda-binary-utilities/">CUDA Binary Utilities</a> manual says a cubin is &#8220;<em>an ELF-formatted file which consists of CUDA executable code sections as well as other sections containing symbols, relocators, debug info, etc.</em>&#8221; That is the specification, and the word &#8220;etc.&#8221; is where the whole platform lives.</p><p>I want you to imagine this piece as a medical dissection, if it&#8217;s possible to say so. Every single number, hex dump and section listing below, came out of a container on <strong>my own machine</strong>, and the exact commands are in the appendix so you can disagree with me precisely. </p><p>Two things made it possible: first thing is that you do not need a GPU to compile, link or disassemble CUDA binaries: <code>ptxas</code>, <code>fatbinary</code>, <code>nvdisasm</code> and <code>cuobjdump</code> are a bunch of ordinary x86 programs. The second is that the entire toolchain is on PyPI, so getting a byte-exact CUDA 13.3.73 install is actually just one <code>pip install</code>.</p><h3>Methodology</h3><p>Everything measured here is from CUDA 13.3.73 (built 9 June 2026) installed from the <code>nvidia-cuda-nvcc</code>, <code>nvidia-cuda-cuobjdump</code> and <code>nvidia-cuda-nvdisasm</code> wheels on x86-64 Linux, plus the cuBLAS 13.6.0.2 wheel for the library dissection. </p><p>I do not currently have a GPU, which means every claim here is about what the toolchain <em>emits</em> and not about what silicon does with it. There are a few parts where I&#8217;m reading structure out of bytes instead of documentation, dont&#8217;t worry: I&#8217;ll say so and I give it like a confidence tier at the end. </p><p>NVIDIA publishes no specification for the cubin section layout, the <code>.nv.info</code> attribute encoding, the fatbin entry header or the Mercury sections, and all the interpretations are mine, unless specified.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Seven processes to compile one file</h2><p>I want you to introduce the format and the pipeline now, because the shape of the output is derived from it. Keep in mind that <code>nvcc</code> is not a compiler, it&#8217;s just a driver that<strong> orchestrates other programs</strong>, and it will show you its plan.</p><pre><code>$ nvcc -arch=sm_90 -c -o k.o saxpy.cu --dryrun
gcc      -D__CUDA_ARCH_LIST__=900 -E -x c++ -D__CUDACC__ ...        # host preprocess
<strong>cudafe++</strong> --c++17 --static-host-stub --device-hidden-visibility ...  # split host/device
gcc      -D__CUDA_ARCH__=900 -E -x c++ -DCUDA_DOUBLE_MATH_FUNCTIONS # device preprocess
<strong>cicc</strong>     ... saxpy.cudafe1.gpu -o saxpy.ptx                        # C++ front end -&gt; PTX
<strong>ptxas</strong>    -arch=sm_90 -m64 saxpy.ptx -o saxpy.sm_90.cubin           # PTX -&gt; SASS
<strong>fatbinary</strong> -64 --cicc-cmdline=... --image3=...  -o saxpy.fatbin.c    # package
gcc      -c -x c++ saxpy.cudafe1.cpp -o k.o                        # host compile</code></pre><p>We have <strong>two preprocessor passes</strong> over the same source with &#8220;different macros&#8221;, then a source-to-source splitter, a front end that emits PTX, an assembler that (you know) emits machine code, a packager that turns the machine code into a C array, and a host compiler that swallows the result. </p><p>The <em>artifacts in the middle</em> are real files you can save with <code>--keep</code>, and each one is a place where the format is decided.</p><p>The only thing that matters is where the close part actually is. We alredy know that <code>cicc</code> is the <strong>NVVM-based front end</strong> and it produces PTX, which is a published virtual instruction set, with all docs that you can read. But for <code>ptxas,</code>which takes PTX and produces the thing this article is about, we don&#8217;t know nothing. </p><p>Everything that <strong>works below PTX </strong>is the vendor&#8217;s private business, and everything upstream has alredy been made public: Clang, Triton, Julia, Mojo and tinygrad all emit PTX, and NVIDIA itself ships <code>libNVVM</code>, so they can do it. </p><p>The interesting feature in the CUDA stack is just one process wide.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>The container is an ELF file that lies about being one</h2><p>Now, let&#8217;s compile a kernel to a standalone cubin and we will see that the <code>file</code> recognizes it immediately.</p><pre><code>$ nvcc -arch=sm_90 -cubin -o saxpy.sm90.cubin saxpy.cu
$ file saxpy.sm90.cubin
saxpy.sm90.cubin: ELF 64-bit LSB executable, <strong>NVIDIA CUDA architecture</strong>, version 1,
                  statically linked, not stripped</code></pre><p>It is a <strong>real ELF64</strong>, little endian, and <code>readelf</code> will walk its section table happily. Three fields in the 64-byte header carry the CUDA-specific part.</p><pre><code>e_ident:  7f 45 4c 46 02 01 01 <strong>41</strong> <strong>08</strong> 00 00 00 00 00 00 00
                              |  |
                              |  +-- EI_ABIVERSION = 8
                              +----- EI_OSABI      = 0x41
e_type    = 2        (ET_EXEC; ET_REL when compiled with -rdc=true)
e_machine = 190      (0xbe, EM_CUDA)
e_flags   = <strong>0x06005a04</strong></code></pre><p><code>EM_CUDA = 190</code> is the one part of this that is available to everybody: it sits in LLVM&#8217;s <code>BinaryFormat/ELF.h</code> alongside every other machine type, because LLVM needs to read these files. </p><p>The <strong>OS ABI byte</strong> is where things get strange. That same header defines two CUDA values, <code>ELFOSABI_CUDA = 51</code> and <code>ELFOSABI_CUDA_V2 = 41</code>, both written in decimal. Every cubin my 13.3 toolchain produces carries 65, which is <code>0x41</code>. 51 is <code>0x33</code>, so the first constant is the older value written in decimal, and the second&#8230;. looks like <code>0x41</code> transcribed as if it were decimal. </p><p>Obv.  I don&#8217;t know where is the mistake, but if you are writing a tool that <strong>sniffs cubins</strong>, sniff for the byte <code>0x41</code> and don&#8217;t trust either constant. The ABI version byte is 8.</p><p> I know that<code> e_flags</code> is where the architecture lives, and the field names are recoverable from the disassembler: <code>EF_CUDA_SM</code>, <code>EF_CUDA_PTX_SM</code>, <code>EF_CUDA_64BIT_ADDRESS</code> and <code>EF_CUDA_ACCELERATORS</code>, alongside an <code>EF_CUDA_SM10</code> through <code>EF_CUDA_SM121</code> enumeration that covers fifteen years of hardware in one list. </p><p>The layout is easy peasy to recover by<strong> sweeping every target</strong> the compiler supports.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!36uk!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F05834540-c5d9-49a8-a0a9-f60a401bb806_1939x933.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!36uk!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F05834540-c5d9-49a8-a0a9-f60a401bb806_1939x933.png 424w, https://substackcdn.com/image/fetch/$s_!36uk!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F05834540-c5d9-49a8-a0a9-f60a401bb806_1939x933.png 848w, https://substackcdn.com/image/fetch/$s_!36uk!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F05834540-c5d9-49a8-a0a9-f60a401bb806_1939x933.png 1272w, https://substackcdn.com/image/fetch/$s_!36uk!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F05834540-c5d9-49a8-a0a9-f60a401bb806_1939x933.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!36uk!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F05834540-c5d9-49a8-a0a9-f60a401bb806_1939x933.png" width="1456" height="701" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/05834540-c5d9-49a8-a0a9-f60a401bb806_1939x933.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:701,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Left: the e_flags word of nine cubins broken into four bytes, showing the compute capability in byte 1 and a byte 0 that changes at Blackwell. Right: cubin size for each target.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Left: the e_flags word of nine cubins broken into four bytes, showing the compute capability in byte 1 and a byte 0 that changes at Blackwell. Right: cubin size for each target." title="Left: the e_flags word of nine cubins broken into four bytes, showing the compute capability in byte 1 and a byte 0 that changes at Blackwell. Right: cubin size for each target." srcset="https://substackcdn.com/image/fetch/$s_!36uk!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F05834540-c5d9-49a8-a0a9-f60a401bb806_1939x933.png 424w, https://substackcdn.com/image/fetch/$s_!36uk!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F05834540-c5d9-49a8-a0a9-f60a401bb806_1939x933.png 848w, https://substackcdn.com/image/fetch/$s_!36uk!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F05834540-c5d9-49a8-a0a9-f60a401bb806_1939x933.png 1272w, https://substackcdn.com/image/fetch/$s_!36uk!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F05834540-c5d9-49a8-a0a9-f60a401bb806_1939x933.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>The same four-line kernel, one target per row. Bits 8 to 15 hold the compute capability as a plain integer. Byte 0 flips from 0x04 to 0x02 at the Blackwell boundary, and the file grows by 44% across that same boundary for identical source.</em></p><p>This is actually very good, but at least three things don&#8217;t match reality: </p><blockquote><p>First, the <strong>virtual architecture</strong> is not in <code>e_flags</code> anymore, despite the name in the field. Compile the same kernel three ways, with the <em>PTX architecture</em> set to <code>compute_75</code>, <code>compute_80</code> and <code>compute_90</code> but the real target fixed at <code>sm_90</code>, and all three cubins carry <code>e_flags = 0x06005a04</code>. </p><p>The virtual architecture actually moved into a <strong>note section,</strong> where <code>cuobjdump</code> reports it as <code>CUDA Virtual SM: sm_75</code> and so on. This used to be visible: CUDA 8-era disassembly printed a line reading <code>.headerflags @"EF_CUDA_SM20 EF_CUDA_PTX_SM(EF_CUDA_SM20)"</code>, and current <code>nvdisasm</code> prints a plain <code>.target sm_90</code> instead. If you have old tooling that reads the <strong>PTX architecture</strong> out of the flags word, you have seen zero for some years.</p></blockquote><blockquote><p>Second, the <strong>architecture-conditional</strong> suffixes do not show up here either. <code>sm_90</code> and <code>sm_90a</code> produce identical <code>e_flags</code>, and so do all three of <code>sm_100</code>, <code>sm_100a</code> and <code>sm_100f</code>. For a kernel that uses no special features the <code>sm_100</code> and <code>sm_100f</code> cubins are the same program <strong>byte for byte</strong>, and the 210 bytes that do differ between the two files are all downstream of one string: the note section records <code>-arch sm_100f</code> instead of <code>-arch sm_100</code>, which is one byte longer and shifts everything after it. </p><p>That is in line with <strong>what NVIDIA says</strong> about the feature, which is that <code>code=sm_100</code> and <code>code=sm_100f</code> are aliases producing the same cubin when no family-specific features are in play. The suffix is recorded, but the recording is somewhere else and is &#8220;triggered&#8221; only when the kernel actually uses something.</p></blockquote><blockquote><p>Third, CUDA 13&#8217;s supported target list is very short, and that surprised even me. Turing is now the floor. <code>ptxas</code> in 13.3 accepts exactly <code>sm_75, sm_80, sm_86, sm_87, sm_88, sm_89, sm_90, sm_90a</code>, then the three-digit generation: <code>sm_100, sm_103, sm_110, sm_120, sm_121</code> each with <code>a</code> and <code>f</code> variants, plus the matching <code>lto_*</code> targets. </p><p>Maxwell, Pascal and Volta were removed in CUDA 13.0, which NVIDIA&#8217;s release notes state directly: <strong>offline compilation</strong> and library support for those architectures are gone completely, and 12.x is the last line that can target them.</p></blockquote><p>You have to know at least two numbers if you ever have to explain a build matrix. <code>sm_110</code> is Jetson AGX Thor, renamed from <code>sm_101</code> when CUDA 13.0 shipped, which means build scripts that hardcoded <code>101</code> broke on a rename rather than on a chip. NVIDIA states the rename in the<strong> nvCOMPDx release notes</strong> rather than anywhere prominent. </p><p>And <code>sm_88</code>, compute capability 8.8, is said by the community compatibility references to be the Nintendo Switch 2, which I cannot verify from the toolchain itself: <code>ptxas</code> will happily build for it and tells you nothing about what it is. </p><p>If that claim is right, the CUDA binary format has a target for a games console sitting in the same list as your H100.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share The Software Frontier&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share The Software Frontier</span></a></p><div><hr></div><h2>Fifteen sections, but only one is code</h2><p>Below I will show you the full section table of a cubin containing one trivial kernel, a <strong>SAXPY with four parameters</strong>, compiled for Hopper.</p><pre><code>idx name                                   type        off   size  align
  1 .shstrtab                              STRTAB       64    295      1
  2 .strtab                                STRTAB      406    369      1
  3 .symtab                                SYMTAB      776    240      8
  4 .debug_frame                           PROGBITS   1016    104      1
  5 .note.nv.tkinfo                        NOTE       1120    164      4
  6 .note.nv.cuinfo                        NOTE       1284     32      4
  7 .nv.info                               CUDA_INFO  1316     36      4
  8 .nv.compat                             CUDA_COMPAT_INFO
                                                      1352     28      4
  9 .nv.info._Z5saxpyifPKfPf               CUDA_INFO  1380    132      4
 10 .nv.callgraph                          CUDA_CALLGRAPH
                                                      1512     32      4
 11 .rela.debug_frame                      RELA       1544     24      8
 12 <strong>.text._Z5saxpyifPKfPf</strong>                   PROGBITS   1664    512    128
 13 .nv.shared.reserved.0                   NOBITS     2176      0      1
 14 .nv.constant0._Z5saxpyifPKfPf           PROGBITS   2176    552      4</code></pre><p><em>We have 512 bytes of machine code inside a 3,968 byte file. The italicised sections are the CUDA-specific ones, with section types in the 0x70000000 range that generic ELF tools print as &#8220;unknown&#8221;.</em></p><p>The <strong>code section</strong> is said to be &#8220;<em>per kernel</em>&#8221; and named after the mangled symbol, which is why a file with ten kernels has ten <code>.text.*</code> sections and ten of most other things too. Alignment is 128 bytes, one instruction cache line&#8217;s worth of paranoia.</p><p>Then, in a random order of how much they surprised me:</p><h3>.note.nv.tkinfo records how the binary was built</h3><p>This is a standard ELF note with owner string <code>NVIDIA Corp</code>, and <code>cuobjdump</code> will print it for you.</p><pre><code>$ cuobjdump -elf saxpy.sm90.cubin | grep -A5 &#8220;CUDA Toolkit Information&#8221;
  NVIDIA Corp   140   NVIDIA CUDA Toolkit Information
    Note Version: 2
    Tool Name: <strong>ptxas</strong>
    Tool Version: Cuda compilation tools, release 13.3, V13.3.73
    Tool Branch: Build cuda_13.3.r13.3/compiler.38244171_0
    Tool Command Line Arguments: <strong>-arch sm_90 -m 64</strong></code></pre><p>Every cubin has the <strong>exact compiler build</strong> that produced it and the flags it was invoked with. That's a provenance record, and it&#8217;s sitting in every shipped binary on every machine learning platform you have ever installed and used. </p><p>If you want to know which toolkit built the kernels inside somebody&#8217;s wheel, no need to ask them. There is also a <code>-verbose-tkinfo</code> mode in the assembler that emits, in its own words, &#8220;<em>object name and command line arguments which contains all arguments having file format,</em>&#8221; meaning paths. I havent seen it on by default, and I&#8217;d look before shipping a binary built with it.</p><h3>.nv.constant0 is the launch frame, and it keeps growing</h3><p>Kernel parameters don&#8217;t come in registers, but in<strong> constant bank 0</strong>, and the cubin declares a per-kernel section for it. On Hopper we said this section is 552 bytes for a kernel whose parameters occupy 24. The user arguments start at offset <code>0x210</code>; everything below that is owned by the driver and the ABI.</p><p>You can watch the exact boundary move across generations. It is at <code>0x160</code> for Turing through Ada, <code>0x210</code> for Hopper, and <code>0x380</code> for Blackwell and later, with the section growing from 380 to 556 to 924 bytes for the same kernel. </p><p>This is the sort of constant that <strong>reverse-engineering</strong> projects hardcode and then rediscover every two years.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!sEJJ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb71d2985-8ddd-4a60-a6a0-989bc0ca09c0_1503x878.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!sEJJ!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb71d2985-8ddd-4a60-a6a0-989bc0ca09c0_1503x878.png 424w, https://substackcdn.com/image/fetch/$s_!sEJJ!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb71d2985-8ddd-4a60-a6a0-989bc0ca09c0_1503x878.png 848w, https://substackcdn.com/image/fetch/$s_!sEJJ!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb71d2985-8ddd-4a60-a6a0-989bc0ca09c0_1503x878.png 1272w, https://substackcdn.com/image/fetch/$s_!sEJJ!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb71d2985-8ddd-4a60-a6a0-989bc0ca09c0_1503x878.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!sEJJ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb71d2985-8ddd-4a60-a6a0-989bc0ca09c0_1503x878.png" width="1456" height="851" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b71d2985-8ddd-4a60-a6a0-989bc0ca09c0_1503x878.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:851,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Two bar charts: kernel parameter base offset in constant bank 0 and total .nv.constant0 size, both by architecture, showing steps at Hopper and Blackwell.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Two bar charts: kernel parameter base offset in constant bank 0 and total .nv.constant0 size, both by architecture, showing steps at Hopper and Blackwell." title="Two bar charts: kernel parameter base offset in constant bank 0 and total .nv.constant0 size, both by architecture, showing steps at Hopper and Blackwell." srcset="https://substackcdn.com/image/fetch/$s_!sEJJ!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb71d2985-8ddd-4a60-a6a0-989bc0ca09c0_1503x878.png 424w, https://substackcdn.com/image/fetch/$s_!sEJJ!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb71d2985-8ddd-4a60-a6a0-989bc0ca09c0_1503x878.png 848w, https://substackcdn.com/image/fetch/$s_!sEJJ!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb71d2985-8ddd-4a60-a6a0-989bc0ca09c0_1503x878.png 1272w, https://substackcdn.com/image/fetch/$s_!sEJJ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb71d2985-8ddd-4a60-a6a0-989bc0ca09c0_1503x878.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">The driver&#8217;s half of the launch. The implicit state a kernel launch carries has grown 2.4x since Turing while the user-visible argument list stayed the same size. Clusters, tensor memory, distributed shared memory and grid constants all have to be described somewhere, and this is where.<span>Measured. Base offset read from the EIATTR_PARAM_CBANK attribute; 24 bytes of user arguments in every case. The space below the base belongs to the driver and the ABI.</span></figcaption></figure></div><p>You can see the effect in the disassembly, which is the easiest way I know to make constant bank 0 sort of&#8230;. real. Here is the same kernel on Hopper and on Thor.</p><pre><code>sm_90                                    sm_110
LDC   R1, c[0x0][0x28]                   LDC   R1, c[0x0][0x37c]
S2R   R0, SR_TID.X                       S2R   R0, SR_TID.X
S2UR  UR4, SR_CTAID.X                    S2UR  UR4, SR_CTAID.X
LDC   R7, c[0x0][RZ]                     <strong>LDCU</strong>  UR5, c[0x0][0x380]
IMAD  R7, R7, UR4, R0                    LDC   R7, c[0x0][0x360]
ULDC  UR4, c[0x0][0x210]                 IMAD  R7, R7, UR4, R0
ISETP.GE.AND P0, PT, R7, UR4, PT         ISETP.GE.AND P0, PT, R7, UR5, PT
@P0   EXIT                               @P0   EXIT</code></pre><p><em>Left: the first parameter is loaded from 0x210. Right: from 0x380, using <strong>LDCU</strong>, an instruction that loads a constant straight into a uniform register and that does not exist in the Hopper instruction set reference. Even the stack pointer moved, from 0x28 to 0x37c.</em></p><h3>The symbol table has its own vocabulary</h3><p>Kernels are ordinary <strong>global function symbols </strong>with two CUDA-specific decorations, but if you look carefully, ther's one symbol in a cubin that has nothing to do with your code.</p><pre><code>index  value   size   info  other  shndx  name
  0x3      0      0    0x3      0    0xc  .text._Z5saxpyifPKfPf
  0x4      0    0x4   0x21      0      0  .nv.reservedSmem.offset0
  0x5      0      0   0x20   0xa0    0xd  __nv_reservedSMEM_offset_0_alias
  0x8      0  0x200   0x12   0x10    0xc  <strong>_Z5saxpyifPKfPf</strong>
  0x9      0      0    0x3      0    0xe  .nv.constant0._Z5saxpyifPKfPf</code></pre><p>The kernel symbol carries <code>st_other = 0x10</code>, which <code>nvdisasm</code> prints as <code>STO_CUDA_ENTRY STV_DEFAULT</code>: a visibility byte reused to mark launchable entry points, which is how the driver tells a kernel from a device function without parsing names. </p><p>The reserved shared memory symbols are the interesting thing here. Even a kernel that declares &#8220;no shared memory&#8221; unexpectedly gets a <strong>reservation record</strong>, and on Blackwell that reservation has a <strong>non-zero size:</strong> somehw you see 64 bytes of shared memory that belong to the platform, not to you, aliased under a name beginning with <code>__nv_</code>. Anyone computing occupancy from source is computing it from the wrong number.</p><h3>The same source is not the same program</h3><p>I decided to write a whole paragraph to the architecture sweep, because it shows how much the<strong> cubin is influenced</strong> by the <strong>specific chip</strong>, rather than by your code. </p><p>The four-line SAXPY compiles to 24 instruction slots on <code>sm_86</code> and 48 on <code>sm_87</code>, Orin, and the difference is not optimization.</p><pre><code>sm_86 opcode histogram:   8 NOP   2 S2R   2 MOV   2 LDG.E   2 IMAD.WIDE  ...
sm_87 opcode histogram:  15 <strong>BMOV.32.CLEAR</strong>  12 NOP   2 S2R   2 LDG.E  ...</code></pre><p>We see<strong> fifteen instructions</strong> clearing barrier state at kernel entry, on one embedded part, for a kernel that uses no barriers. </p><p>In other words, that piece of code inside the kernel isn&#8217;t there because your program desperately needs it. It&#8217;s here because NVIDIA is paying, through software, the price of a <strong>silicon limitation</strong>. And that proves that SASS  is dependent not only on the PTX, but also on the hardware, the drivers and the workarounds embedded into the toolchain. </p><h3>.nv.callgraph exists because device code can call things</h3><p>Fixed 8-byte entries, and for a leaf kernel it is four rows of negative sentinels:</p><pre><code><code>.nv.callgraph
&lt;0,-1&gt;   &lt;0,-2&gt;   &lt;0,-3&gt;   &lt;0,-4&gt;</code></code></pre><p>The sentinels are the interesting part: a leaf kernel does not get an empty section, but four explicit &#8220;<em>no callee of kind N</em>&#8221; markers. This suggests that <code>.nv.callgraph</code> is not simply a list of callees, but a structured description of call relationships understood by the runtime.</p><p>That metadata matters because features such as indirect calls, dynamic parallelism, and <strong>separately compiled device functions</strong> require the driver to know the worst-case call depth before launch so it can size the per-thread stack correctly. </p><p> Even a kernel that calls nothing still participates in this contract, which is why absence is encoded explicitly rather than by omitting the section entirely.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>.nv.info is where the kernel describes itself</h2><p>This is the section that let you see a loadable cubin, rather than merely readable. The format is a simple <strong>tag-length-value stream</strong>: one byte of format, one byte of attribute, then either nothing, one byte, two bytes, or a 16-bit length followed by that many bytes. </p><p>And&#8230;.thirty-six bytes of it, at file level, look like this:</p><pre><code>04 <strong>2f</strong> 08 00 | 08 00 00 00  0a 00 00 00      EIATTR_REGCOUNT
04 <strong>11</strong> 08 00 | 08 00 00 00  00 00 00 00      EIATTR_FRAME_SIZE
04 <strong>12</strong> 08 00 | 08 00 00 00  00 00 00 00      EIATTR_MIN_STACK_SIZE
 ^  ^  ^
 |  |  +-- value length (EIFMT_SVAL)
 |  +----- attribute code
 +-------- format: 1 = none, 2 = one byte, 3 = two bytes, 4 = length-prefixed</code></pre><p>The first value that you read is a symbol table index, the second is the payload. The symbol 8 in each column is the kernel, and it needs 10 registers, no stack frame, no minimum stack. You can check that decode against the tool, because <code>cuobjdump -elf</code> knows the names.</p><pre><code>$ cuobjdump -elf saxpy.sm90.cubin
.nv.info
  Attribute: <strong>EIATTR_REGCOUNT</strong>       Value: function: _Z5saxpyifPKfPf(0x8)  register count: 10
  Attribute: <strong>EIATTR_FRAME_SIZE</strong>     Value: function: _Z5saxpyifPKfPf(0x8)  frame size: 0x0
  Attribute: <strong>EIATTR_MIN_STACK_SIZE</strong> Value: function: _Z5saxpyifPKfPf(0x8)  min stack size: 0x0</code></pre><p>The per-kernel section is way richer. Let&#8217;s see a four-parameter SAXPY:</p><pre><code>.nv.info._Z5saxpyifPKfPf
  EIATTR_LANGUAGE            PTX
  EIATTR_CUDA_API_VERSION    0x85                      (133 = CUDA 13.3)
  EIATTR_KPARAM_INFO         Ordinal 0x3  Offset 0x10  Size 0x8  Space CBANK
  EIATTR_KPARAM_INFO         Ordinal 0x2  Offset 0x8   Size 0x8  Space CBANK
  EIATTR_KPARAM_INFO         Ordinal 0x1  Offset 0x4   Size 0x4  Space CBANK
  EIATTR_KPARAM_INFO         Ordinal 0x0  Offset 0x0   Size 0x4  Space CBANK
  EIATTR_SPARSE_MMA_MASK     0x0
  EIATTR_MAXREG_COUNT        0xff
  <strong>EIATTR_MERCURY_ISA_VERSION 1.1</strong>
  EIATTR_EXIT_INSTR_OFFSETS  0x70  0x120
  EIATTR_CBANK_PARAM_SIZE    0x18                      (24 bytes of parameters)
  EIATTR_PARAM_CBANK         0x9   0x180210            (symbol 9, size 0x18 at 0x210)
  EIATTR_SW_WAR              0x8
  EIATTR_NVSAL_SW_WAR        0x1</code></pre><p><strong>Kernel parameters</strong> are described entirely through size-and-offset metadata, with entries stored in something called &#8220;<em>reverse ordinal order</em>&#8221;. </p><p>This means the host runtime doesn&#8217;t have to understand the original programming language or the function signature. It can construct the kernel&#8217;s parameter buffer directly from the metadata, placing each argument at the correct location before launch. The compiled binary therefore exposes a<strong> language-independent ABI</strong> that the driver and runtime can consume.</p><p><code>EIATTR_EXIT_INSTR_OFFSETS</code> contains the byte offset of every <code>EXIT</code> instruction generated for the kernel. At first glance this is considered to be like a minor detail, but it&#8217;s very useful for tooling. </p><p>A profiler, tracer, or instrumentation framework can find every<strong> kernel return point</strong> immediately, without performing a full disassembly or reconstructing the control-flow graph. If a tool have to inject timing code, coverage probes, or custom telemetry before kernel termination, these offsets provide the exact insertion points.</p><p><code>EIATTR_SW_WAR</code> is a different kind of purpose here. It encodes software workarounds for known hardware errata as a bitmask. Modern GPUs sometimes require<strong> compiler- or driver-level mitigations</strong> to avoid wrong behavior in specific corner cases. </p><p>Rather than hard-coding these decisions somewhere else, the compiler do a record of the workarounds that are needed, in the binary metadata. The driver then inspects the flags at load time and enables the correct mitigation paths for the architecture in place.</p><p>All these attributes show you that <strong>CUDA metadata</strong> is not  just a bunch of descriptions; it forms actively some parts of the contract between the compiler, runtime, driver, profiling tools, and the GPU itself, carrying the information needed to launch kernels, instrument execution, and safely navigate hardware-specific quirks.</p><p>Now a question arises naturally: <em>how big is this vocabulary?</em> The disassembler contains the enum names, so you can just ask it, and you&#8217;ll find out.</p><pre><code>$ strings -a nvdisasm | grep -o &#8220;EIATTR_[A-Z0-9_]*&#8221; | sort -u | wc -l
<strong>113</strong></code></pre><p>Now you see it: 113 attribute kinds. A few of them are a decent map of what the hardware has learned to do since 2010. </p><ul><li><p><code>EIATTR_TCGEN05_1CTA_USED</code> and <code>EIATTR_TCGEN05_2CTA_USED</code> flag use of Blackwell&#8217;s fifth-generation tensor cores, and the fact that there are separate flags for the one-CTA and two-CTA forms tells you the driver has to care which. </p></li><li><p><code>EIATTR_CTA_PER_CLUSTER</code>, <code>EIATTR_MAX_CLUSTER_RANK</code>, <code>EIATTR_EXPLICIT_CLUSTER</code> and <code>EIATTR_BLOCKS_ARE_CLUSTERS</code> are the Hopper cluster model. </p></li><li><p><code>EIATTR_STACK_CANARY_TRAP_OFFSETS</code> means device code has stack canaries now, and indeed <code>nvcc</code> has a flag for them. </p></li><li><p><code>EIATTR_COROUTINE_RESUME_ID_OFFSETS</code> is there for something that has not been announced. </p></li><li><p>And a small family of attributes are named after bug numbers: <code>EIATTR_SW1850030_WAR</code>, <code>EIATTR_SW2393858_WAR</code>, <code>EIATTR_SW2861232_WAR</code>, <code>EIATTR_WAR5829587_NEEDED</code>. </p></li></ul><p>There is  also an internal ticket for each somewhere, and the workaround is load-bearing enough to have its own slot in the binary format.</p><blockquote><p><em>The format is not a description of a program. It is a contract between a compiler and a driver that ship on different schedules and are written by the same company.</em></p></blockquote><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Twenty-one bits per instruction that the hardware does not check</h2><p>Starting with Volta, SASS instructions are 128 bits, and a piece of every instruction word is not the instruction at all, because it&#8217;s the <strong>scheduling metadata</strong> that the assembler computes and the hardware obeys without verifying. </p><p>You will find plenty of documentation about the structure of the field in the <strong>microbenchmarking literature</strong>, rather than by the vendor: a paper by Jia and colleagues describes stall counts, a yield flag, separate read and write dependency barriers, a wait mask that is a bitmask because an instruction can wait on several barriers at once, and reuse flags. </p><p>The exact bit positions below are the ones I decoded, and I trust them because the result is coherent with the docs: stall count in bits 105 to 108, yield flag at 109, write barrier index at 110 to 112, read barrier index at 113 to 115, a six-bit wait mask at 116 to 121, and four reuse flags at 122 to 125. Bits 126 and 127 come out zero on every instruction of every kernel I looked at, on five architectures.</p><p>You can decode it yourself with twenty lines of Python, and I think everyone who works on inference should do it at least once, because the mechanism explains more about the CUDA moat than any benchmark. Here is a real <strong>Hopper GEMM prologue</strong> with the fields pulled out.</p><pre><code>off   stall  y  wr  rd    wait  reuse   instruction
0000      1  1   -   -  000000   0000   LDC R1, c[0x0][0x28]
0010      1  1   <strong>0</strong>   -  000000   0000   S2R R9, SR_CTAID.Y
0020      1  1   -   -  000000   0000   ULDC UR4, c[0x0][0x228]
0030      1  1   -   -  000000   0000   ULDC.64 UR6, c[0x0][0x208]
0040      1  1   -   -  000000   0000   MOV R0, UR4
0050      1  1   <strong>0</strong>   -  000000   0000   S2R R13, SR_TID.Y
0060      2  1   -   -  000000   0000   HFMA2.MMA R16, -RZ, RZ, 0, 0
0070      1  1   -   -  000000   0000   ISETP.GE.AND P0, PT, R0, 0x1, PT
0080      1  1   <strong>1</strong>   -  000000   0000   S2R R17, SR_CTAID.X
0090      1  1   <strong>1</strong>   -  000000   0000   S2R R7, SR_TID.X
00a0      2  0   -   -  <strong>000001</strong>   0000   LEA R9, R9, R13, 0x5
00b0      8  0   -   -  <strong>000010</strong>   0000   LEA R11, R17, R7, 0x5</code></pre><p>Two special register reads allocate barrier 0, and the<strong> LEA</strong> that consumes them waits on barrier 0. Two more allocate barrier 1, and the next LEA waits on barrier 1. The dependency is not discovered at run time, but written into the binary.</p><p>That is the whole argument, in twelve instructions. For <strong>variable-latency instructions</strong> the assembler allocates one of six barriers and makes consumers wait on a bitmask. </p><p>For <strong>fixed-latency instructions</strong> it does not use barriers at all, it just writes a number of cycles into the stall field and the scheduler holds off that long. There is no interlock checking this: if the number is wrong, you do not get a stall, you get wrong answers. </p><p>Across a <strong>full tiled GEMM kernel</strong> the distribution looks like this.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!wUyx!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa8e2d257-93b7-427a-9a55-163592501b79_1522x854.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!wUyx!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa8e2d257-93b7-427a-9a55-163592501b79_1522x854.png 424w, https://substackcdn.com/image/fetch/$s_!wUyx!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa8e2d257-93b7-427a-9a55-163592501b79_1522x854.png 848w, https://substackcdn.com/image/fetch/$s_!wUyx!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa8e2d257-93b7-427a-9a55-163592501b79_1522x854.png 1272w, https://substackcdn.com/image/fetch/$s_!wUyx!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa8e2d257-93b7-427a-9a55-163592501b79_1522x854.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!wUyx!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa8e2d257-93b7-427a-9a55-163592501b79_1522x854.png" width="1456" height="817" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a8e2d257-93b7-427a-9a55-163592501b79_1522x854.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:817,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Histogram of encoded stall counts and a bar chart of how often each control field is used across one kernel.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Histogram of encoded stall counts and a bar chart of how often each control field is used across one kernel." title="Histogram of encoded stall counts and a bar chart of how often each control field is used across one kernel." srcset="https://substackcdn.com/image/fetch/$s_!wUyx!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa8e2d257-93b7-427a-9a55-163592501b79_1522x854.png 424w, https://substackcdn.com/image/fetch/$s_!wUyx!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa8e2d257-93b7-427a-9a55-163592501b79_1522x854.png 848w, https://substackcdn.com/image/fetch/$s_!wUyx!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa8e2d257-93b7-427a-9a55-163592501b79_1522x854.png 1272w, https://substackcdn.com/image/fetch/$s_!wUyx!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa8e2d257-93b7-427a-9a55-163592501b79_1522x854.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">One 32x32 tiled float GEMM for sm_90, 168 instruction slots. Most instructions issue back to back with a one cycle stall. Write barriers appear on 26% of slots, read barriers on none, and the yield flag on 71%. The same decode works unchanged on sm_100, sm_110 and sm_120, which is a small piece of evidence that Blackwell did not change the scheduling contract. <span>Measured: one 32x32 float tiled GEMM compiled for sm_90, 168 instruction slots including padding. Counts are slots, decoded by direct parse of the .text section.</span></figcaption></figure></div><h3>The linker can rewrite the control bits</h3><p>The relocation namespace makes the point better than I can. Pull the type names out of <code>nvlink</code> and there are 119 of them, and most encode a bit position in the name.</p><pre><code>$ strings -a nvlink | grep -o &#8220;R_CUDA_[A-Z0-9_]*&#8221; | sort -u | wc -l
<strong>119</strong>

R_CUDA_ABS32_23        patch 32 bits at bit offset 23
R_CUDA_ABS32_HI_32     high half of a 64-bit address, at bit 32
R_CUDA_CONST_FIELD19_20  19-bit constant bank field at bit 20
R_CUDA_ABS55_16_34     55-bit value split across bits 16 and 34
R_CUDA_INSTRUCTION128  a whole 128-bit instruction word
R_CUDA_PCREL_IMM24_23  24-bit program counter relative branch at bit 23
<strong>R_CUDA_YIELD_CLEAR_PRED4_87</strong>   clear 4 bits at bit 87
<strong>R_CUDA_YIELD_OPCODE9_0</strong>        rewrite a 9-bit opcode field at bit 0</code></pre><p>On a CPU, relocations patch whole bytes because operands are byte-aligned. In this case, the operands are bit fields inside a <strong>128-bit word</strong>, so the relocation type has to name the field, and there is a separate type for the same width at each position it can occur. </p><p>The last two are the ones that really stopped me. A relocation that clears the yield predicate, and one that <strong>rewrites an opcode</strong>, mean the device linker&#8217;s contract includes editing scheduling and instruction selection after the assembler has finished. </p><p>The<strong> 21 control bits</strong> are not a compiler-internal detail, which is nested at the end of <code>ptxas</code>, because they are a key part of the object format, and a later stage is allowed to change them.</p><p>I have argued many times before that this field, not the language and not PTX, is the <strong>load-bearing part</strong> of the CUDA moat, and taking cubins apart has not changed my mind. It has sharpened one point though. </p><p>The reason the assembler is fully close is not that NVIDIA wants to hide optimizations and reverse-engineering, at least not the first point. What we also need to keep in mind is that the <code>ptxas</code> is<strong> correctness-critical</strong> infrastructure. </p><p>An external tool that emits SASS is not competing with a compiler on quality of code, it is taking over responsibility for a hazard-avoidance protocol with no runtime check and <strong>no error reporting</strong>, which is extremely dangerous. </p><p>NVIDIA&#8217;s own programming guide is blunt about the consequence: you read that binary compatibility is promised only for binaries created by <strong>NVIDIA tools:</strong> manual editing or generating binary code is not supported, and compatibility promises are invalidated if binaries are modified in any way. </p><p>Read that as a description of the <strong>support boundary</strong> rather than a threat, and the closure looks less like strategy and more like the only sentence a vendor can write once it has moved the interlock into software.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share The Software Frontier&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share The Software Frontier</span></a></p><div><hr></div><h2>Blackwell cubins carry a second copy of every kernel</h2><p>When I thought about writing this part, I didn&#8217;t even know what to expect. It was a complete surprise to me. Let&#8217;s do a concrete example: compile the same <strong>trivial kernel</strong> for Hopper and for Blackwell and count sections: 15 for <code>sm_90</code>, 22 for <code>sm_100</code>. The seven extra sections are not debug info.</p><pre><code>idx name                                   type              flags     size
 12 .text._Z5saxpyifPKfPf                  PROGBITS              6      512
 13 .nv.shared.reserved.0                  NOBITS                3       64
 14 .nv.constant0._Z5saxpyifPKfPf          PROGBITS             42      920
 15 <strong>.nv.capmerc.text._Z5saxpyifPKfPf</strong>       0x70000016     10000000      258
 16 <strong>.nv.merc.debug_frame</strong>                   PROGBITS       10000000      112
 17 <strong>.nv.merc.nv.info</strong>                       0x70000083     10000000       36
 18 <strong>.nv.merc.nv.info._Z5saxpyifPKfPf</strong>       0x70000083     10000040      144
 19 <strong>.nv.merc.rela.debug_frame</strong>              0x70000082     10000040       24
 20 <strong>.nv.merc.nv.shared.reserved.0</strong>          0x70000015     10000003        0
 21 <strong>.nv.merc.symtab</strong>                        0x70000085     10000000      216</code></pre><p>You are now staring at a <strong>complete parallel object</strong>: its own symbol table, its own info sections, its own relocations, its own shared memory reservation, and a &#8220;<em>capmerc</em>&#8221; text section. Section flag bit 0x10000000 marks the whole family.</p><p>The prefix is <code>merc</code>, and the name shows up in three other places. It is in the per-kernel info as <code>EIATTR_MERCURY_ISA_VERSION 1.1</code>. It is in a new section called <code>.nv.compat</code>, which exists on every target but says more on Blackwell, and it is all over the assembler binary.</p><pre><code>$ cuobjdump -elf saxpy.sm100a.cubin | sed -n &#8216;/nv.compat/,/nv.info\._/p&#8217;
.nv.compat
  <strong>EICOMPAT_ATTR_CUDA_ACCELERATOR_TARGET</strong>                 0x1
  EICOMPAT_ATTR_ISA_CLASS                                0x1
  <strong>EICOMPAT_ATTR_INST_TCGEN05_MMA</strong>                        0x5
  EICOMPAT_ATTR_MERCURY_ISA_MAJOR_MINOR_VERSION_V2       1.1
  EICOMPAT_ATTR_MERCURY_ISA_MAJOR_MINOR_VERSION_V1       1.1
  EICOMPAT_ATTR_INST_TENSORMAP_V1                        0x0
  <strong>EICOMPAT_ATTR_CAN_FASTPATH_FINALIZE</strong>                   0x9 0x0</code></pre><p>There is the missing suffix. <code>EICOMPAT_ATTR_CUDA_ACCELERATOR_TARGET</code> is 1 for <code>sm_100a</code> and 0 for both <code>sm_100</code> and <code>sm_100f</code>, and the accelerated build declares <code>EICOMPAT_ATTR_INST_TCGEN05_MMA</code>, a capability it was permitted to use even though this kernel does not use it.</p><p>To really find out what the block really tracks, you need to build a kernel that needs a family-specific instruction. One line of <strong>inline PTX</strong> will do: <code>tcgen05.fence::before_thread_sync</code>, part of Blackwell&#8217;s fifth-generation tensor core interface. </p><p>Compiling it for <strong>five targets </strong>gives the feature tiers as a straight experiment rather than as a diagram.</p><pre><code>$ for a in sm_100 sm_100f sm_100a sm_103f sm_120a; do nvcc -arch=$a -cubin -o t.cubin tc.cu; done

sm_100    <strong>Instruction &#8216;tcgen05.fence&#8217; not supported on .target &#8216;sm_100&#8217;</strong>
sm_100f   OK  (5136 bytes)
sm_100a   OK  (5136 bytes)
sm_103f   OK  (5136 bytes)
sm_120a   <strong>Instruction &#8216;tcgen05.fence&#8217; not supported on .target &#8216;sm_120a&#8217;</strong></code></pre><p>The baseline target refuses the instruction, but the<strong> family target</strong> accepts it, and so does the family target for the other member of the same family. </p><p>And the fully accelerated target for consumer Blackwell refuses it, because <code>sm_120</code> is a different family and doesn&#8217;t have the hardware, which is the concrete answer to the frequently asked question of why a 5090 is not a small B200.</p><p>Now look at what the successful builds wrote into the object.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!nssc!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fffb50e64-a1b5-490f-926d-d21b1414ccb0_1939x1111.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!nssc!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fffb50e64-a1b5-490f-926d-d21b1414ccb0_1939x1111.png 424w, https://substackcdn.com/image/fetch/$s_!nssc!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fffb50e64-a1b5-490f-926d-d21b1414ccb0_1939x1111.png 848w, https://substackcdn.com/image/fetch/$s_!nssc!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fffb50e64-a1b5-490f-926d-d21b1414ccb0_1939x1111.png 1272w, https://substackcdn.com/image/fetch/$s_!nssc!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fffb50e64-a1b5-490f-926d-d21b1414ccb0_1939x1111.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!nssc!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fffb50e64-a1b5-490f-926d-d21b1414ccb0_1939x1111.png" width="1456" height="834" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ffb50e64-a1b5-490f-926d-d21b1414ccb0_1939x1111.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:834,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:158227,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/207127260?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fffb50e64-a1b5-490f-926d-d21b1414ccb0_1939x1111.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!nssc!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fffb50e64-a1b5-490f-926d-d21b1414ccb0_1939x1111.png 424w, https://substackcdn.com/image/fetch/$s_!nssc!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fffb50e64-a1b5-490f-926d-d21b1414ccb0_1939x1111.png 848w, https://substackcdn.com/image/fetch/$s_!nssc!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fffb50e64-a1b5-490f-926d-d21b1414ccb0_1939x1111.png 1272w, https://substackcdn.com/image/fetch/$s_!nssc!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fffb50e64-a1b5-490f-926d-d21b1414ccb0_1939x1111.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>That is the piece I was missing when I first wrote this section, infact I had to rebuild it. The compat block is not a record of your compiler flags; it is a <strong>graded statement</strong> about what the machine code inside requires, computed from the code, and the flags only set the ceiling on what the assembler was allowed to reach for. </p><p>A resolver reading <code>ISA_CLASS = 1</code> knows this object is safe anywhere in the generation; reading 2, it knows to check the family. This is the <strong>minimum viable metadata</strong> for the compatibility promise NVIDIA made in CUDA 12.9, and it is sitting in a 32-byte section nobody documents.</p><p>What reads it becomes obvious once you look inside <code>ptxas</code>. The strings are unambiguous.</p><pre><code>$ strings -a ptxas | grep -iE &#8220;merc|finaliz&#8221; | sort -u
[Finalizer] fastpath optimization applied for <strong>off-target %u -&gt; %u finalization</strong>
Generate Capsule Mercury
Specify the type of target ELF binary kind. Default on sm100+ is <strong>capmerc</strong>
Self check for capsule mercury (capmerc)
Specify the opportunistic finalization level. 0=default, 1=no opportunistic
  finalization, <strong>2=intra family finalization only, or 3=intra and inter family
  finalization</strong>
Turns off the fast-path finalization optimization (allows normal refinaization)
Specify the &#8216;sm_&#8217; name of the target architecture. If not specified, default
  behavior is <strong>on-target finalization</strong>
R_MERCURY_NONE  R_MERCURY_G64  R_MERCURY_ABS64  R_MERCURY_ABS32  R_MERCURY_ABS16 ...</code></pre><p>Read together with the sections, this describes a two-stage back end. Mercury is an encoding of a kernel that is <strong>below PTX, </strong>but above final machine code. </p><p>A capsule wraps it with the metadata a finalizer needs: its own symbol table, register and barrier counts, shared memory usage, and its own relocation types. &#8220;<em>Finalization</em>&#8221; is the step that turns a capsule into <strong>executable SASS</strong> for a specific chip. On-target finalization is the ordinary case. </p><p><strong>Off-target finalization</strong> is the same capsule being finalized for a different chip than the one it was compiled for, with a fast path when the assembler can prove the existing encoding is already valid. And the levels of &#8220;<em>opportunistic finalization</em>&#8221; go from none, to within a family, to <em>across</em> families.</p><p>That is what the <code>f</code> targets are made of. NVIDIA introduced family-conditional compilation in CUDA 12.9 and actually described it in terms of feature sets: an <code>f</code> binary may use the <strong>architecture-specific features</strong> that are common to a whole family, and will run on later members of that family. </p><p>The public framing is a compatibility promise. The mechanism, as far as I can see it in the binary, is that the cubin ships a <strong>re-finalizable representation </strong>next to the SASS, and the driver can produce correct machine code for a family member the compiler never saw.</p><h4>Where I am reading, not knowing</h4><p>Unfortunately, I havent observed finalization happen. At the moment, don&#8217;t own a Blackwell GPU, and none of this machinery has the <code>ptxas</code> command line: passing <code>--binary-kind</code> gets you <code>Unknown option</code>, so the option table I am quoting belongs to an internal or driver-side entry point. </p><p>What I am confident about is the <strong>presence and shape of the artifacts</strong>. An independent reverse-engineering effort on <code>ptxas</code> 13.0 reached the same reading and adds detail I could not verify, including a 328-byte capsule descriptor and a compilation-knob snapshot inside it. Treat the purpose as tier C, so a speculative claim.</p><p>There is more of it in the section type table than any single cubin shows. <code>cuobjdump</code> has to be able to print a name for every section type it might meet, so the names are in the binary, and the Mercury family is large.</p><pre><code>$ strings -a cuobjdump | grep -oE &#8220;CUDA_[A-Z_]{3,}&#8221; | sort -u | grep -i merc
CUDA_CAPMERC
CUDA_MERCURY
CUDA_MERCURY_CONSTANT_DRIVER      CUDA_MERCURY_CONSTANT_PARAMS
CUDA_MERCURY_CONSTANT_IMGHDR      CUDA_MERCURY_CONSTANT_PIC
CUDA_MERCURY_CONSTANT_OPT         CUDA_MERCURY_CONSTANT_TOOLS
                                  CUDA_MERCURY_CONSTANT_USER
CUDA_MERCURY_RESOLVED_RELA
<strong>CUDA_MERCURY_SASS_MAP</strong></code></pre><p><strong>Seven </strong>distinct <strong>constant bank</strong> <strong>section types</strong>, split by who owns the data: driver, image header, optimizer, parameters, position independent code, tools, user. A resolved-relocation type. And a section type called <code>CUDA_MERCURY_SASS_MAP</code>, which is hard to read as anything other than a correspondence between Mercury entities and SASS. </p><p>A format that needs a<strong> map to SASS</strong> is not an annotation on SASS, because it&#8217;s a representation of the program in its own right, and the map exists so that debuggers, profilers and the finalizer can move between the two.</p><p>All of which makes the size behaviour the really awkward part, because it argues against the simplest reading of what I just described.</p><p>If the capsule were a second copy of the instruction stream, in whatever denser encoding, its size would have to <strong>follow the SASS.</strong> Longer kernel, longer capsule. That is not what happens. Across seven kernels the SASS spans 384 to 2,048 bytes, a factor of 5.3, while the capsules span only 130 to 350, a factor of 2.7. </p><p>Correlations are carried almost entirely by a single long kernel: the coefficient across all seven is 0.76, and dropping that one point takes it to 0.33. Now, if you rank the kernels by <strong>capsule size</strong> and you get an ordering with no obvious relationship to how much code they contain.</p><p>The cleanest way to see it is a controlled pair. Two of these kernels compile to exactly 384 bytes of SASS. One stores a float, the other multiplies and adds in double precision. <strong>Same code size</strong>, same instruction count, and their capsules are 130 and 194 bytes, a 49% difference. Meanwhile a kernel using <code>sqrtf</code> compiles to 896 bytes of SASS, more than twice the double-precision kernel, and carries a <em>smaller</em> capsule at 184.</p><p>I want to be careful here, because when I first wrote this section I claimed the ordering tracked the variety of operations a kernel performs, and then I checked. It does not. Correlating capsule size against the <strong>number of distinct opcodes</strong> in each kernel gives 0.03, which is nothing at all. I could not find anything that predicts capsule size, and with seven kernels I would not trust a pattern even if I had found one.</p><p>What the double-precision pair does suggest is that whatever the capsule enumerates is closer to a set of requirements than to a program, and <strong>double precision</strong> is a plausible thing to find in such a set. Double-precision throughput is one of the sharpest differences within a Blackwell family: the datacenter parts and the consumer parts are not the same machine in that respect. </p><p>A representation whose job is to let a driver produce correct code for a family member the compiler never saw would need to record exactly this sort of thing, and would not need to grow much when you add another hundred instructions of the same kind.</p><p>So the sizes push against &#8220;<em>a second copy of every instruction</em>&#8221; and toward &#8220;<em>a description of what this code needs.</em>&#8221; They cannot separate the two readings that remain. A compact whole-kernel encoding in a much denser format, and a <strong>fixup table</strong> over the existing SASS naming only the sites a re-finalizer would have to revisit, would both be small and would both fail to scale with instruction count. </p><p>The section type called <code>CUDA_MERCURY_SASS_MAP</code> is the one piece of evidence that leans toward the first, since a table of fixups would not obviously need a map. I cannot settle it from sizes alone, and I doubt anyone can without a driver to watch.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!QrF2!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe20c3f10-bfea-46a3-b840-90d90891b8a3_1503x884.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!QrF2!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe20c3f10-bfea-46a3-b840-90d90891b8a3_1503x884.png 424w, https://substackcdn.com/image/fetch/$s_!QrF2!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe20c3f10-bfea-46a3-b840-90d90891b8a3_1503x884.png 848w, https://substackcdn.com/image/fetch/$s_!QrF2!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe20c3f10-bfea-46a3-b840-90d90891b8a3_1503x884.png 1272w, https://substackcdn.com/image/fetch/$s_!QrF2!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe20c3f10-bfea-46a3-b840-90d90891b8a3_1503x884.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!QrF2!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe20c3f10-bfea-46a3-b840-90d90891b8a3_1503x884.png" width="1456" height="856" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e20c3f10-bfea-46a3-b840-90d90891b8a3_1503x884.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:856,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Scatter plot of Mercury capsule size against SASS size for seven kernels, showing a nearly flat relationship.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Scatter plot of Mercury capsule size against SASS size for seven kernels, showing a nearly flat relationship." title="Scatter plot of Mercury capsule size against SASS size for seven kernels, showing a nearly flat relationship." srcset="https://substackcdn.com/image/fetch/$s_!QrF2!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe20c3f10-bfea-46a3-b840-90d90891b8a3_1503x884.png 424w, https://substackcdn.com/image/fetch/$s_!QrF2!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe20c3f10-bfea-46a3-b840-90d90891b8a3_1503x884.png 848w, https://substackcdn.com/image/fetch/$s_!QrF2!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe20c3f10-bfea-46a3-b840-90d90891b8a3_1503x884.png 1272w, https://substackcdn.com/image/fetch/$s_!QrF2!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe20c3f10-bfea-46a3-b840-90d90891b8a3_1503x884.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Seven kernels on sm_100. If the capsule were a compact second encoding of the instruction stream, these points would sit on a line through the origin. They do not. A double-precision kernel with 384 bytes of SASS carries a 194 byte capsule; a sqrtf kernel with 896 bytes of SASS carries 184. Whatever the capsule enumerates, it is closer to a set of properties than to a program.<span>Measured on sm_100. The double-precision kernel has the smallest SASS and nearly the largest capsule; the sqrtf kernel has the largest SASS and one of the smallest.</span></figcaption></figure></div><p>What it costs is easier to state. On the ten-kernel workload, Mercury sections add 4.2 KB to a 57.8 KB cubin, about 7%, and Blackwell cubins are 2.0x the size of Turing cubins for identical source.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Flhn!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0cce32f6-7c9d-4089-8db0-8fe09f17b5fc_1487x952.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Flhn!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0cce32f6-7c9d-4089-8db0-8fe09f17b5fc_1487x952.png 424w, https://substackcdn.com/image/fetch/$s_!Flhn!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0cce32f6-7c9d-4089-8db0-8fe09f17b5fc_1487x952.png 848w, https://substackcdn.com/image/fetch/$s_!Flhn!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0cce32f6-7c9d-4089-8db0-8fe09f17b5fc_1487x952.png 1272w, https://substackcdn.com/image/fetch/$s_!Flhn!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0cce32f6-7c9d-4089-8db0-8fe09f17b5fc_1487x952.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Flhn!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0cce32f6-7c9d-4089-8db0-8fe09f17b5fc_1487x952.png" width="1456" height="932" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0cce32f6-7c9d-4089-8db0-8fe09f17b5fc_1487x952.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:932,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Stacked bar chart of cubin size composition across twelve architectures, showing text, constant bank, Mercury sections, metadata, debug and ELF scaffolding.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Stacked bar chart of cubin size composition across twelve architectures, showing text, constant bank, Mercury sections, metadata, debug and ELF scaffolding." title="Stacked bar chart of cubin size composition across twelve architectures, showing text, constant bank, Mercury sections, metadata, debug and ELF scaffolding." srcset="https://substackcdn.com/image/fetch/$s_!Flhn!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0cce32f6-7c9d-4089-8db0-8fe09f17b5fc_1487x952.png 424w, https://substackcdn.com/image/fetch/$s_!Flhn!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0cce32f6-7c9d-4089-8db0-8fe09f17b5fc_1487x952.png 848w, https://substackcdn.com/image/fetch/$s_!Flhn!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0cce32f6-7c9d-4089-8db0-8fe09f17b5fc_1487x952.png 1272w, https://substackcdn.com/image/fetch/$s_!Flhn!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0cce32f6-7c9d-4089-8db0-8fe09f17b5fc_1487x952.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Ten kernels, twelve targets, same source file. The SASS roughly doubles from Turing to Blackwell for reasons that have nothing to do with the format, the launch frame more than doubles, and the Mercury sections appear from sm_100 onwards. Nothing here is compressed: this is the cubin as ptxas writes it.<span>Measured with nvcc -cubin. Section sizes read from the ELF section headers; &#8216;ELF scaffolding&#8217; is the file remainder (string tables, symtab, headers, padding).</span></figcaption></figure></div><p>There is <strong>one more hint </strong>about direction of travel, in the disassembler rather than the assembler. <code>nvdisasm</code> in 13.3 has an option called <code>--no-vliw</code>, described as &#8220;<em>conventional mode; disassemble paired instructions in normal syntax, instead of VLIW syntax.</em>&#8221; </p><p>Nothing I compiled for any current target produced paired output. An option to <em>turn off</em> VLIW syntax implies a target where VLIW syntax is the default. </p><p>I would not build a thesis on a <strong>command-line flag,</strong> but if a future architecture issues statically paired instructions, the tooling has already been taught the word.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>The fatbin is a filesystem, and the driver is the resolver</h2><p>A cubin runs on exactly <strong>one compute capability</strong>. Shipping software means shipping several, plus PTX for the ones that do not exist yet, and the container for that is the fatbin. </p><p>It starts with a 16-byte header whose magic is <code>0xBA55ED50</code>, and it is followed by a run of entries, each with its own header and payload.</p><pre><code>$ nvcc -gencode arch=compute_90,code=sm_90 \
       -gencode arch=compute_100,code=[sm_100,compute_100] -fatbin -o saxpy.fatbin saxpy.cu

kind     arch  hdrsz  payload   compressed  uncompressed  flags
CUBIN      90     64     3968            0             0  0x11
CUBIN     100    112     5712            0             0  0x1000011
PTX       100     80      400          399           922  0x8011</code></pre><p>The entry header fields are recoverable by differential compilation: change one thing about the build, diff the bytes, name the field.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!kd9n!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb4c40fd0-9a96-48aa-a305-15247265966b_2163x1460.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!kd9n!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb4c40fd0-9a96-48aa-a305-15247265966b_2163x1460.png 424w, https://substackcdn.com/image/fetch/$s_!kd9n!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb4c40fd0-9a96-48aa-a305-15247265966b_2163x1460.png 848w, https://substackcdn.com/image/fetch/$s_!kd9n!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb4c40fd0-9a96-48aa-a305-15247265966b_2163x1460.png 1272w, https://substackcdn.com/image/fetch/$s_!kd9n!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb4c40fd0-9a96-48aa-a305-15247265966b_2163x1460.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!kd9n!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb4c40fd0-9a96-48aa-a305-15247265966b_2163x1460.png" width="1456" height="983" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b4c40fd0-9a96-48aa-a305-15247265966b_2163x1460.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:983,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:285278,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/207127260?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb4c40fd0-9a96-48aa-a305-15247265966b_2163x1460.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!kd9n!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb4c40fd0-9a96-48aa-a305-15247265966b_2163x1460.png 424w, https://substackcdn.com/image/fetch/$s_!kd9n!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb4c40fd0-9a96-48aa-a305-15247265966b_2163x1460.png 848w, https://substackcdn.com/image/fetch/$s_!kd9n!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb4c40fd0-9a96-48aa-a305-15247265966b_2163x1460.png 1272w, https://substackcdn.com/image/fetch/$s_!kd9n!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb4c40fd0-9a96-48aa-a305-15247265966b_2163x1460.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>Reconstructed by construction, not from documentation. Every row was fixed by building the same kernel with one thing changed and diffing the header bytes. Fields I did not manage to move are omitted.</em></p><h3>PTX is not the cheap option</h3><p>The standard advice is to always ship PTX for the newest virtual architecture so future GPUs have something to JIT. The advice generally is right, and the reason people give for it is usually wrong. </p><p><strong>PTX is not a compact representation</strong>. For the ten-kernel workload the PTX is 37.3 KB against a 35.9 KB Hopper cubin, and on Blackwell it is 48.0 KB against 57.8 KB. It is the same order of magnitude in both directions, and after compression the gap narrows further because PTX is text and compresses beautifully.</p><p>What PTX buys is <strong>coverage of chips</strong> that do not exist yet, at the cost of a JIT compile on first use, per process, unless the driver&#8217;s compute cache is warm. </p><p>What it does not buy is what most people think: <strong>forward compatibility </strong>does not extend to the instructions anyone actually cares about, because PTX built for an architecture-conditional target is not forward compatible at all. </p><p>If your kernel needs <code>wgmma</code> or <code>tcgen05</code>, it needs <code>sm_90a</code> or <code>sm_100a</code>, and the compatibility story ends there. The <code>f</code> targets exist precisely because that cliff was too steep, and the Mercury capsule is how they made it less steep.</p><p>That extended header is the good part. The <code>sm_100</code> entry above has a 112-byte header instead of 64, and the extra 48 bytes are a length-prefixed copy of the <code>.nv.compat</code> capability block from inside the cubin.</p><pre><code>header bytes at offset 0x40:
48 00 00 00  20 00 00 00     -- payload at 0x48, length 0x20
02 09 00 00  02 02 01 00  03 0d 01 01  03 07 01 01
02 03 00 00  04 0b 08 00  09 00 00 00  00 00 00 00
   ^^ the same TLV stream that .nv.compat holds inside the ELF</code></pre><p>The capability claim is duplicated into the index so that whatever picks an entry never has to open the ELF. That is a loader optimization, and it is also a small confirmation of what the compat block is for: it is the thing a resolver matches against a device before committing to a payload.</p><h3>Compression is off by default for the part that matters</h3><p>CUDA has a <code>--compress-mode</code> flag with values <code>none</code>, <code>speed</code>, <code>balance</code>, <code>size</code> and <code>default</code>. Running all five over the same two-cubin-plus-PTX fatbin gives a result worth knowing.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Pt5x!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6f49747e-ba08-4693-89fe-0fb22883c10c_1699x753.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Pt5x!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6f49747e-ba08-4693-89fe-0fb22883c10c_1699x753.png 424w, https://substackcdn.com/image/fetch/$s_!Pt5x!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6f49747e-ba08-4693-89fe-0fb22883c10c_1699x753.png 848w, https://substackcdn.com/image/fetch/$s_!Pt5x!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6f49747e-ba08-4693-89fe-0fb22883c10c_1699x753.png 1272w, https://substackcdn.com/image/fetch/$s_!Pt5x!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6f49747e-ba08-4693-89fe-0fb22883c10c_1699x753.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Pt5x!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6f49747e-ba08-4693-89fe-0fb22883c10c_1699x753.png" width="1456" height="645" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6f49747e-ba08-4693-89fe-0fb22883c10c_1699x753.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:645,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Horizontal bar chart of fatbin size under four compression modes, with the size mode dramatically smaller.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Horizontal bar chart of fatbin size under four compression modes, with the size mode dramatically smaller." title="Horizontal bar chart of fatbin size under four compression modes, with the size mode dramatically smaller." srcset="https://substackcdn.com/image/fetch/$s_!Pt5x!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6f49747e-ba08-4693-89fe-0fb22883c10c_1699x753.png 424w, https://substackcdn.com/image/fetch/$s_!Pt5x!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6f49747e-ba08-4693-89fe-0fb22883c10c_1699x753.png 848w, https://substackcdn.com/image/fetch/$s_!Pt5x!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6f49747e-ba08-4693-89fe-0fb22883c10c_1699x753.png 1272w, https://substackcdn.com/image/fetch/$s_!Pt5x!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6f49747e-ba08-4693-89fe-0fb22883c10c_1699x753.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">In default, speed and balance mode, only the PTX gets compressed and both cubins are stored raw. Only -compress-mode=size compresses SASS, and on this input it takes the fatbin from 10,352 bytes to 3,000.<span>Measured. In default, speed and balance mode the two cubins are stored raw and only the PTX is compressed; -compress-mode=size compresses the cubins as well.</span></figcaption></figure></div><p>The flag values behave like two different algorithms rather than one algorithm at three settings: <code>speed</code> sets flag bit 13 and gets 504 bytes out of 922, while <code>balance</code> and <code>size</code> set bit 15 and get 400. </p><p>On the larger <strong>ten-kernel workload</strong> across twelve targets, <code>size</code> mode takes 537 KB down to 75 KB, an 86% reduction, for no measurable build time cost at this scale. </p><p>If you ship CUDA binaries and have not tried it, that is the cheapest win in this article.</p><h3>How it gets into your executable</h3><p>The host side is one of the few places NVIDIA actually publishes the format, in <code>fatbinary_section.h</code> in the toolkit include directory.</p><pre><code>#define FATBINC_MAGIC   0x466243B1
#define FATBINC_VERSION 1
typedef struct {
  int magic;
  int version;
  const unsigned long long* data;
  void *filename_or_fatbins;
} __fatBinC_Wrapper_t;

#define FATBIN_CONTROL_SECTION_NAME  &#8220;.nvFatBinSegment&#8221;
#define FATBIN_DATA_SECTION_NAME     &#8220;.nv_fatbin&#8221;</code></pre><p>Compile a <code>.cu</code> file to an object and you get exactly that: a <strong>24-byte wrapper </strong>in <code>.nvFatBinSegment</code> holding the magic and a relocation pointing at the fatbin bytes in <code>.nv_fatbin</code>, plus a section called <code>__nv_module_id</code>, plus a constructor.</p><pre><code>$ readelf -x .nvFatBinSegment saxpy.o
  0x00000000 <strong>b1436246</strong> 01000000 00000000 00000000 .CbF............
$ readelf -rW saxpy.o | grep nvFatBinSegment
  R_X86_64_64  .nv_fatbin + 0

$ cat saxpy.cudafe1.stub.c        # generated by nvcc --keep
static void __nv_cudaEntityRegisterCallback(void **__T0) {
  __cudaRegisterEntry(__T0, (void(*)(int,float,const float*,float*))saxpy,
                      _Z5saxpyifPKfPf, (-1)); }
static void <strong>__sti____cudaRegisterAll</strong>(void) __attribute__((__constructor__));
static void __sti____cudaRegisterAll(void) {
  <strong>__cudaRegisterBinary</strong>(__nv_cudaEntityRegisterCallback); }

void __device_stub__Z5saxpyifPKfPf(int p0, float p1, const float *p2, float *p3){
  __cudaLaunchPrologue(4);
  __cudaSetupArgSimple(p0, 0UL);  __cudaSetupArgSimple(p1, 4UL);
  __cudaSetupArgSimple(p2, 8UL);  __cudaSetupArgSimple(p3, 16UL);
  __cudaLaunch((char *)saxpy); }</code></pre><p>So a CUDA program registers its device code from an ELF constructor, before <code>main</code>. The <strong>argument offsets in the stub</strong>, 0, 4, 8 and 16, are the same offsets the cubin declared in <code>EIATTR_KPARAM_INFO</code>. </p><p>Host and device agree on the layout because both were generated from the same front end, and the binary carries the agreement in two places.</p><p>Two variants are worth noting because they trip people up. With <code>-rdc=true</code>, the device code is relocatable, the cubin becomes <code>ET_REL</code> with <code>.rela.text.*</code> sections and an <code>EIATTR_EXTERNS</code> attribute, and it lands in a differently named section, <code>__nv_relfatbin</code>, until the device link step turns it into a normal <code>.nv_fatbin</code>. The cost shows up in the machine code, and it is not subtle:</p><pre><code>$ nvcc -arch=sm_90 -rdc=true -c a.cu b.cu &amp;&amp; nvcc -arch=sm_90 -dlink -o dl.o a.o b.o
$ cuobjdump -sass dl.o
    Function : _Z5applyPfPKffi
        /*0100*/  <strong>CALL.ABS.NOINC</strong> 0x0 ;
    Function : _Z5scaleff
        /*0000*/  FFMA R4, R4, R5, 1 ;
        /*0010*/  <strong>RET.ABS.NODEC</strong> R20 0x0 ;</code></pre><p>A one-instruction device function became a real call, with a stack frame, a return address register and a separate <code>.text</code> section, where whole-program compilation would have inlined it into a single FFMA. </p><p>That is what <code>-dlto</code> exists to undo: with LTO the fatbin entry is kind 8 and <strong>the instruction selection</strong> is deferred to link time so the inlining can happen after all translation units are visible. </p><p>In my sandbox the LTO device link fails for want of a driver-side link library, so I can report the container shape (1,944 compressed bytes expanding to 2,548) but not the linked output.</p><h3>What the driver does with all this</h3><p>The loading path is the reason the format looks the way it does, and it is worth stating plainly because the layering is easy to lose. The constructor calls <code>__cudaRegisterBinary</code> with a pointer to the wrapper. </p><p>The runtime <strong>hands the fatbin to the driver</strong>, which reads the index and selects one payload for the device it is about to use: an exact-architecture cubin if there is one, otherwise a compatible one, otherwise PTX to be compiled on the spot. The selected payload becomes a module. </p><p>The module&#8217;s <code>.nv.info</code> tells the driver how many registers each kernel needs, how much shared memory to reserve, how big a stack to allocate given the <strong>call graph</strong>, and where in constant bank 0 to write the arguments. Only then can a launch happen, and a launch is essentially a memcpy of the argument block into the frame the cubin described, followed by a grid dispatch.</p><p>Everything in the binary that is not machine code exists to let that sequence happen without the driver consulting anything else. That is why the <strong>parameter layout is in the object</strong> rather than in a header, why the exit offsets are precomputed, why the capability claim is duplicated into the fatbin index, and why the reserved shared memory has a symbol. A loader that had to infer any of it would be slower and more fragile.</p><p><strong>Two changes to this path</strong> are worth knowing. Since CUDA 12.0 there is a second module abstraction, the library and kernel API (<code>cuLibraryLoadData</code>, <code>cuKernel</code>), which decouples a loaded artifact from a specific context, and it exists largely because the old model made lazy loading awkward. </p><p>And since 12.2 on Linux and 12.3 everywhere, lazy loading is the default: modules load on first use of a symbol from them, and kernels within a module load on first <code>cuModuleGetFunction</code>. </p><p>The requirement is a runtime of 11.7 or later, statically linked into whatever built the library, which means an old dependency in your stack can quietly opt you out. <strong>NVIDIA&#8217;s own documentation</strong> puts the requirement plainly: libraries compiled against older runtimes load all modules eagerly.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>What the format costs, in bytes and seconds</h2><p>All of the above has a price, and it is paid by everyone who ships a wheel.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Lwpg!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa0373bfe-1212-465b-8b0b-289abb835e94_1487x885.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Lwpg!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa0373bfe-1212-465b-8b0b-289abb835e94_1487x885.png 424w, https://substackcdn.com/image/fetch/$s_!Lwpg!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa0373bfe-1212-465b-8b0b-289abb835e94_1487x885.png 848w, https://substackcdn.com/image/fetch/$s_!Lwpg!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa0373bfe-1212-465b-8b0b-289abb835e94_1487x885.png 1272w, https://substackcdn.com/image/fetch/$s_!Lwpg!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa0373bfe-1212-465b-8b0b-289abb835e94_1487x885.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Lwpg!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa0373bfe-1212-465b-8b0b-289abb835e94_1487x885.png" width="1456" height="867" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a0373bfe-1212-465b-8b0b-289abb835e94_1487x885.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:867,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Two panels: fatbin size versus number of targets, and nvcc wall clock versus number of targets.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Two panels: fatbin size versus number of targets, and nvcc wall clock versus number of targets." title="Two panels: fatbin size versus number of targets, and nvcc wall clock versus number of targets." srcset="https://substackcdn.com/image/fetch/$s_!Lwpg!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa0373bfe-1212-465b-8b0b-289abb835e94_1487x885.png 424w, https://substackcdn.com/image/fetch/$s_!Lwpg!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa0373bfe-1212-465b-8b0b-289abb835e94_1487x885.png 848w, https://substackcdn.com/image/fetch/$s_!Lwpg!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa0373bfe-1212-465b-8b0b-289abb835e94_1487x885.png 1272w, https://substackcdn.com/image/fetch/$s_!Lwpg!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa0373bfe-1212-465b-8b0b-289abb835e94_1487x885.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Ten kernels. Each additional real target is a full independent compilation and a full independent copy. The relationship is linear in both size and time, which is the whole problem: there is no sharing between targets, because there is nothing to share.<span>Measured: CUDA 13.3.73, 10 templated kernels (4 GEMM, 4 elementwise, 2 reductions), single-threaded nvcc, no GPU present.</span></figcaption></figure></div><p>This is why PyTorch&#8217;s release engineering discussions read the way they do. When the<strong> CUDA 12.8 build matrix</strong> had to add Blackwell, the team explicitly rejected keeping the older architectures because of binary size. The trade is not subtle: adding a target adds a copy of every kernel in the library, and machine learning libraries have a great many kernels.</p><p>You can see the endpoint of that logic by taking apart what NVIDIA itself ships. I pulled the <strong>cuBLAS 13.6.0.2 wheel</strong> from PyPI and parsed the <code>.nv_fatbin</code> sections directly.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!uff2!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F529ff6e2-6245-42e5-a852-82e3382b2846_2276x1486.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!uff2!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F529ff6e2-6245-42e5-a852-82e3382b2846_2276x1486.png 424w, https://substackcdn.com/image/fetch/$s_!uff2!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F529ff6e2-6245-42e5-a852-82e3382b2846_2276x1486.png 848w, https://substackcdn.com/image/fetch/$s_!uff2!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F529ff6e2-6245-42e5-a852-82e3382b2846_2276x1486.png 1272w, https://substackcdn.com/image/fetch/$s_!uff2!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F529ff6e2-6245-42e5-a852-82e3382b2846_2276x1486.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!uff2!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F529ff6e2-6245-42e5-a852-82e3382b2846_2276x1486.png" width="1456" height="951" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/529ff6e2-6245-42e5-a852-82e3382b2846_2276x1486.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:951,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:234858,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/207127260?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F529ff6e2-6245-42e5-a852-82e3382b2846_2276x1486.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!uff2!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F529ff6e2-6245-42e5-a852-82e3382b2846_2276x1486.png 424w, https://substackcdn.com/image/fetch/$s_!uff2!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F529ff6e2-6245-42e5-a852-82e3382b2846_2276x1486.png 848w, https://substackcdn.com/image/fetch/$s_!uff2!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F529ff6e2-6245-42e5-a852-82e3382b2846_2276x1486.png 1272w, https://substackcdn.com/image/fetch/$s_!uff2!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F529ff6e2-6245-42e5-a852-82e3382b2846_2276x1486.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Two libraries, 6,860 fatbin entries, 178 MB on disk expanding to 1.8 GB of device code. And <strong>every single entry</strong> is compressed, which tells you that NVIDIA does not ship its own libraries with the default compression mode.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!FkGO!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa6869b1f-f7ef-48eb-97aa-14c055cc5911_1530x952.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!FkGO!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa6869b1f-f7ef-48eb-97aa-14c055cc5911_1530x952.png 424w, https://substackcdn.com/image/fetch/$s_!FkGO!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa6869b1f-f7ef-48eb-97aa-14c055cc5911_1530x952.png 848w, https://substackcdn.com/image/fetch/$s_!FkGO!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa6869b1f-f7ef-48eb-97aa-14c055cc5911_1530x952.png 1272w, https://substackcdn.com/image/fetch/$s_!FkGO!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa6869b1f-f7ef-48eb-97aa-14c055cc5911_1530x952.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!FkGO!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa6869b1f-f7ef-48eb-97aa-14c055cc5911_1530x952.png" width="1456" height="906" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a6869b1f-f7ef-48eb-97aa-14c055cc5911_1530x952.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:906,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Left: device code per architecture for the two cuBLAS libraries. Right: section sizes inside libcublasLt.so.13.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Left: device code per architecture for the two cuBLAS libraries. Right: section sizes inside libcublasLt.so.13." title="Left: device code per architecture for the two cuBLAS libraries. Right: section sizes inside libcublasLt.so.13." srcset="https://substackcdn.com/image/fetch/$s_!FkGO!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa6869b1f-f7ef-48eb-97aa-14c055cc5911_1530x952.png 424w, https://substackcdn.com/image/fetch/$s_!FkGO!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa6869b1f-f7ef-48eb-97aa-14c055cc5911_1530x952.png 848w, https://substackcdn.com/image/fetch/$s_!FkGO!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa6869b1f-f7ef-48eb-97aa-14c055cc5911_1530x952.png 1272w, https://substackcdn.com/image/fetch/$s_!FkGO!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa6869b1f-f7ef-48eb-97aa-14c055cc5911_1530x952.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Hopper is still the largest single target in cuBLASLt at 37.4 MB of compressed SASS, more than consumer Blackwell and datacenter Blackwell individually. The right panel is the part I did not expect: device code is only 27% of the file.<span>Measured on the cuBLAS 13.6.0.2 wheel from PyPI. Right panel counts only file-resident sections, which sum to 491.4 MB of the 492.7 MB file; .bss and .lbss add another 99 MB of runtime memory but zero bytes on disk. cuBLASLt holds 5,599 fatbin entries (5,309 cubins, 290 PTX), 132.3 MB compressed, expanding to 1,361 MB.</span></figcaption></figure></div><p>The right-hand panel is worth sitting with. In cuBLASLt, 101 MB is host code, and 78 MB is a section called <code>.cask_resource</code>. </p><p>One detail worth noting before anyone adds those numbers up: the library also declares 99 MB of <code>.bss</code> and <code>.lbss</code>, and those are NOBITS sections, which means they cost 99 MB of memory at run time and zero bytes on disk. </p><p>Counting only the sections that actually occupy file space gets you 491.4 MB of the 492.7 MB file, and the rest is headers and padding. I cannot tell you <strong>what CASK stands for,</strong> because NVIDIA has never said, and the name only surfaces publicly in cuDNN and TensorRT diagnostics. </p><p>What I can tell you is that the strings in that section include 2,429 mangled C++ names beginning <code>N11cask</code>, so it is a namespace, and that the rest of the section is a catalog. Pull strings out of it and you get entries like this:</p><pre><code>cutlass3x_sm103_bstensorop_s256x256x96gemm_block_scaled_ue4m3xe2m1_ue4m3xe2m1_
  f32_bf16_ue4m3xe2m1_256x256x768_0_tnn_align...</code></pre><p>A CUTLASS 3.x kernel for <code>sm_103</code> doing a block-scaled GEMM with FP8 and FP4 operand types at a 256x256x96 tile. </p><p>Counting distinct kernel-shaped names across the resource and read-only sections gives <strong>25,401</strong>, with the generation prefixes still visible in the histogram: 5,329 mentioning <code>sm90</code>, 5,015 <code>cutlass3x</code>, 1,956 <code>ampere</code>, 808 <code>volta</code>, 274 <code>turing</code>.</p><p>That number is the moat, expressed as an artifact count rather than an argument. <strong>Not one clever kernel</strong>. Twenty-five thousand parameterised ones, plus 78 MB of metadata describing which to pick, plus 101 MB of host code doing the picking. A competitor can match the code generator. </p><p>Matching the catalog and the selection heuristics is a different kind of work, and it is the kind that does not benefit much from being smart.</p><p>NVIDIA is visibly feeling the weight of it. cuDNN 9.5 shipped a build configuration called <code>GRAPH_JIT_ONLY</code> that keeps the runtime kernel generation engines and drops the <strong>precompiled ones </strong>specifically to cut binary size, which is the vendor arriving at the same trade its users have been making by hand: generate at run time or ship the archive.</p><p>It also explains why lazy loading became the default. If a process eagerly loaded <strong>all 5,309 cubins</strong> in cuBLASLt it would spend its startup budget on kernels it will never call. </p><p>Lazy loading arrived as opt-in in CUDA 11.7, became the Linux default in 12.2 and the universal default in 12.3, and defers module and kernel loading until first use. </p><p>In a world where a <strong>single library</strong> holds thousands of independently loadable objects, that stops being an optimization and becomes a requirement.</p><h3>Debug information is not free either</h3><p>For the same ten kernels on <code>sm_90</code>: 35.9 KB release, 75.6 KB with <code>-lineinfo</code>, 205.9 KB with <code>-G</code>. The last one is not just extra sections. </p><p>It changes code generation, taking <code>.text</code> from 16.0 KB to 57.5 KB, because full device debug turns off most optimization. <code>-lineinfo</code> is the option you want in production builds: it doubles the cubin and leaves the code alone.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/p/how-cuda-binaries-actually-work?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/p/how-cuda-binaries-actually-work?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><div><hr></div><h2>Why any of this matters if you are not writing a compiler</h2><p>Three reasons, in increasing order of how much I care about them.</p><p>The first is operational. Binary size and load time are real costs in inference deployments, container images and cold starts, and they are <strong>almost entirely determined by choices</strong> in the format: how many targets, PTX or not, which compression mode, lazy or eager. Those are four flags. Most teams have never touched them.</p><p>The second is diagnostic. A cubin tells you its register count, its shared memory usage, its stack frame, its parameter layout, the toolkit that built it and the <strong>flags it was built with</strong>, and you can read all of that out of a shipped wheel without running anything. </p><p>If you are trying to work out why somebody else&#8217;s kernel spills, or which architectures a vendor actually optimized for versus which they merely support, the <strong>binary is a better source</strong> than the changelog. The <code>-res-usage</code> flag on <code>cuobjdump</code> is the fastest occupancy audit I know.</p><h4>Auditing somebody else&#8217;s wheel</h4><pre><code>cuobjdump -lelf libfoo.so          # which architectures, how many cubins
cuobjdump -lptx libfoo.so          # is there PTX for anything newer
cuobjdump -res-usage libfoo.so     # registers, shared memory, stack, per kernel
cuobjdump -elf libfoo.so | grep -A5 &#8220;Toolkit Information&#8221;   # which ptxas built it
readelf -SW libfoo.so | grep nv_fatbin                      # how much of the file is device code</code></pre><p>Five commands, no GPU, no source. The last one is the one that surprises people: if <code>.nv_fatbin</code> is 85% of a library, you are shipping a kernel archive with a thin API on top, and any conversation about image size has to start there.</p><p>The third is strategic, and it is why I went looking in the first place. </p><p>There is a long line of work reverse-engineering these formats: <code>decuda</code> for the earliest chips, <code>asfermi</code> for Fermi, <code>KeplerAs</code>, Scott Gray&#8217;s <code>maxas</code> for <strong>Maxwell </strong>(the toolchain behind the fast convolution kernels of the deep learning era), <code>turingas</code>, <code>CuAssembler</code> for everything through Ampere, NVBit for dynamic instrumentation, and now work lifting SASS back into typed compiler IR. </p><p>Every one of these projects <strong>had to rediscover the container </strong>before it could touch the code, and every one of them targets an architecture at least two generations behind current silicon. That lag is the moat&#8217;s actual shape. </p><p>It is <strong>not that the format is unknowable</strong>, instead it&#8217;s that knowing it is a full-time job that resets on a yearly cadence, and the number of people doing it is small enough to name.</p><p>Mercury is the part of this that I think changes the picture, and not in the direction I expected. Publishing an interface and keeping the lowering is NVIDIA&#8217;s oldest move: <strong>PTX is public</strong>, <code>ptxas</code> is not. </p><p>What the capsule does is insert a second, private layer at exactly the point where a competitor would want to attach, and it does it for a reason customers asked for: one binary that keeps working on the next chip in the family. The <em>compatibility benefit</em> is real. </p><p>The side effect is that the interesting <strong>representation of a Blackwell kernel </strong>is no longer the SASS you can disassemble, and the format that carries it has no public name, no documented encoding, and its own relocation types.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Four things I would change, and one I would not</h2><p>Having read a few hundred of these files, I have opinions about the format that have nothing to do with the moat.</p><p><strong>Publish the metadata, keep the encoding.</strong> <em>There is no competitive information in </em><code>EIATTR_KPARAM_INFO</code><em>. It is a calling convention. Documenting the </em><code>.nv.info</code><em> attribute codes and the fatbin entry header would cost NVIDIA nothing and would stop every profiler, virtualization layer, checkpointer and build tool from rediscovering the same twenty structs. The instruction encoding is a different argument and I understand the refusal. The container is not.</em></p><p><strong>Make </strong><code>size</code><strong> the default compression mode.</strong> <em>The current default compresses PTX and leaves the SASS raw, which is the wrong way round: SASS is the bulk and the part that never gets read by a human. NVIDIA already ships its own libraries with everything compressed. Every framework wheel on PyPI is paying for a default its vendor does not use.</em></p><p><strong>Give the toolkit a way to strip PTX.</strong> <em>There is </em><code>nvprune</code><em> for removing architectures, but the common case in a container build is &#8220;I know exactly which GPUs this image runs on, remove the JIT fallback.&#8221; Doing that today means unpacking and repacking fatbins with </em><code>nvFatbin</code><em>, which is a strange amount of work for what should be a link flag.</em></p><p><strong>Version the format visibly.</strong><em> Between the ABI version byte, the OS ABI byte that changed value, the toolkit note, the API version attribute and the Mercury ISA version, there are five different notions of &#8220;which format is this&#8221; in one file, and no single field that a tool can check to know whether it will understand what follows. A cubin that a future tool cannot parse should say so in its first sixteen bytes.</em></p><p>What <strong>I would not change personally</strong> is the decision to make the binary self-describing rather than pushing the information into headers or a side-channel. </p><p>It is the reason <strong>NVBit can instrument a kernel</strong> it has never seen, the reason <code>cuobjdump -res-usage</code> works on a stranger&#8217;s wheel, and the reason this article was possible on a machine with no GPU in it. Verbose, redundant, self-describing binaries are a gift to everyone downstream, including the people trying to work out what you shipped. </p><p>Most of what is wrong with the CUDA binary format is that it is undocumented. Very little of it is that it is badly designed.</p><h4><span>Where I might be wrong</span></h4><ul><li><p><strong>The capsule might be much less than I think.</strong> If it turns out to be a small fixup table rather than a re-encodable program, then &#8220;second copy of every kernel&#8221; is the wrong phrase and the right one is &#8220;compatibility annotations.&#8221; The sublinear size scaling in the capsule plot above is evidence for the smaller reading; the existence of a section type called <code>CUDA_MERCURY_SASS_MAP</code> is evidence for the larger one. I could not settle it.</p></li><li><p><strong>I have no Blackwell hardware.</strong> Nothing here was executed. A claim about what a driver does with a capsule is, from my position, a claim about what an artifact appears designed for. Everything about run time behaviour, including JIT cost and the actual effect of lazy loading, is cited rather than measured.</p></li><li><p><strong>The OS ABI byte is a mess and I may be reading the mess wrong.</strong> LLVM lists 51 and 41; cubins carry 65. My reading is that the second constant is <code>0x41</code> written as decimal, but it is equally possible NVIDIA changed the value again and LLVM has not caught up. I only checked 13.3 output.</p></li><li><p><strong>The fatbin flag bits are inferred from a handful of data points.</strong> I set them by construction using the compression modes and the LTO path; other bits in that word certainly mean things I did not exercise. Byte 0 of <code>e_flags</code>, which flips from 4 to 2 at Blackwell, I could not explain at all.</p></li><li><p><strong>My library numbers are one version of one vendor library.</strong> cuBLAS is the extreme case for kernel count. Generalizing &#8220;27% device code&#8221; to other libraries would be wrong.</p></li><li><p><strong>The feature-tier experiment used one instruction.</strong> <code>tcgen05.fence</code> behaves as described, and I have no reason to think the mechanism differs for other family-specific instructions, but one probe is one probe.</p></li></ul><h3><span>Predictions</span></h3><ol><li><p><strong>By the end of 2027</strong>, at least one open-source project will publish a working parser for the Mercury capsule format, and it will be motivated by GPU virtualization or checkpointing rather than by performance.</p></li><li><p><strong>NVIDIA will not document the capsule encoding</strong>, on the same terms it has never documented SASS encodings, through at least CUDA 15.</p></li><li><p><strong>By the end of 2027</strong>, <code>-compress-mode=size</code> or its equivalent will be the default in at least one major framework&#8217;s release build, driven by wheel size limits rather than by anyone reading the flag documentation.</p></li><li><p><strong>The next architecture after Blackwell will ship a SASS encoding with statically paired instructions</strong>, and <code>nvdisasm</code>&#8216;s VLIW syntax will be why we find out.</p></li><li><p><strong>Family-conditional targets will absorb the architecture-conditional ones in practice</strong>: by the end of 2028, the fraction of shipped cubins carrying <code>a</code> suffixes will fall as libraries move to <code>f</code>, because one binary per family is a better deal than one per chip.</p></li></ol><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Dossier</h2><p><strong>Tier A, measured directly and reproducible from the appendix. </strong></p><ul><li><p><em>ELF header field values and the </em><code>e_flags</code><em> architecture encoding, including the absence of the virtual architecture from that word. </em></p></li><li><p><em>Section inventories and sizes for all twelve targets. </em></p></li><li><p><em>The </em><code>.nv.info</code><em> and </em><code>.nv.compat</code><em> attribute contents as decoded by </em><code>cuobjdump</code><em>, including the </em><code>ISA_CLASS</code><em> transition when a family-specific instruction is used. </em></p></li><li><p><em>The rejection of </em><code>tcgen05.fence</code><em> on </em><code>sm_100</code><em> and </em><code>sm_120a</code><em> and its acceptance on </em><code>sm_100f</code><em>, </em><code>sm_100a</code><em> and </em><code>sm_103f</code><em>. </em></p></li><li><p><em>The 113 attribute names and the 119 relocation names. Parameter base offsets and constant bank sizes. </em></p></li><li><p><em>Control field bit positions and their per-kernel distribution. </em></p></li><li><p><em>Fatbin entry field offsets, compression flag bits and all size and timing measurements. </em></p></li><li><p><em>The </em><code>CALL.ABS.NOINC</code><em> cost of relocatable device code. </em></p></li><li><p><em>The cuBLAS section and entry statistics, the 25,401 kernel names and the 2,429 </em><code>cask</code><em> mangled names. </em></p></li><li><p><em>The presence, sizes and scaling of the Mercury sections.</em></p></li></ul><p><strong>Tier B, documented by NVIDIA, in shipped headers, or in third-party source.</strong><span> </span></p><ul><li><p><em>The fatbin wrapper struct and section names (</em><code>fatbinary_section.h</code><em>). </em></p></li><li><p><em>The family and architecture-conditional target semantics, including the statement that </em><code>sm_100</code><em> and </em><code>sm_100f</code><em> are aliases absent family features (Programming Guide, PTX ISA, the CUDA 12.9 blog post). </em></p></li><li><p><em>Lazy loading history and defaults. </em></p></li><li><p><em>Dropped architecture support in CUDA 13.0 and the Thor renumbering. </em></p></li><li><p><code>EM_CUDA</code><em>, </em><code>ELFOSABI_CUDA</code><em> and </em><code>ELFOSABI_CUDA_V2</code><em> in LLVM. </em></p></li><li><p><em>cuDNN&#8217;s </em><code>GRAPH_JIT_ONLY</code><em> configuration. </em></p></li><li><p><em>The structure of the control field, which is documented in the literature rather than by the vendor.</em></p></li></ul><p><strong>Tier C, inferred from artifacts.</strong></p><ul><li><p><em>Mercury is a re-finalizable intermediate representation and that the capsule exists to serve family-conditional compatibility. </em></p></li><li><p><em>That the duplicated compat block in the fatbin entry header is a loader fast path. </em></p></li><li><p><code>ptxas</code><em>&#8216;s hidden finalizer options correspond to a driver-side code path. </em></p></li><li><p><code>.cask_resource</code><em> is a kernel selection catalog. </em></p></li><li><p><em>That the </em><code>ELFOSABI_CUDA_V2</code><em> constant is a hexadecimal value written as decimal.</em></p></li></ul><p><strong>Tier D, speculation, flagged as such in the text.</strong></p><ul><li><p><code>--no-vliw</code><em> implies a future paired-issue ISA. </em></p></li><li><p><code>EIATTR_COROUTINE_RESUME_ID_OFFSETS</code><em> points at an unannounced feature. </em></p></li><li><p><code>sm_88</code><em> is the Nintendo Switch 2, which comes from community references and not from anything I could check.</em></p></li></ul><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Reproduce this</h2><p>No GPU required. On <strong>x86-64 Linux:</strong></p><pre><code>pip install nvidia-cuda-nvcc==13.3.73 nvidia-cuda-cuobjdump==13.3.73 \
            nvidia-cuda-nvdisasm==13.3.73 nvidia-cuda-crt==13.3.73 \
            nvidia-cuda-runtime nvidia-cuda-cccl
export PATH=$(python3 -c &#8220;import sysconfig,os;print(os.path.join(sysconfig.get_paths()[&#8217;purelib&#8217;],&#8217;nvidia/cu13/bin&#8217;))&#8221;):$PATH

cat &gt; saxpy.cu &lt;&lt;&#8217;EOF&#8217;
#include &lt;cuda_runtime.h&gt;
__global__ void saxpy(int n, float a, const float* x, float* y) {
    int i = blockIdx.x * blockDim.x + threadIdx.x;
    if (i &lt; n) y[i] = a * x[i] + y[i];
}
EOF

# the container and its metadata
nvcc -arch=sm_90  -cubin -o k90.cubin  saxpy.cu
nvcc -arch=sm_100 -cubin -o k100.cubin saxpy.cu
cuobjdump -elf k90.cubin | less
diff &lt;(cuobjdump -elf k90.cubin | sed -n &#8216;/^Sections:/,/^$/p&#8217;) \
     &lt;(cuobjdump -elf k100.cubin | sed -n &#8216;/^Sections:/,/^$/p&#8217;)

# the capability claim, and the accelerator bit
for a in sm_100 sm_100a sm_100f; do nvcc -arch=$a -cubin -o c.cubin saxpy.cu
  echo &#8220;== $a&#8221;; cuobjdump -elf c.cubin | grep -A1 EICOMPAT | grep -E &#8220;Attr|Value&#8221;; done

# the attribute namespace and the finalizer options
strings -a $(command -v nvdisasm) | grep -o &#8220;EIATTR_[A-Z0-9_]*&#8221; | sort -u
strings -a $(command -v ptxas)    | grep -iE &#8220;capmerc|finaliz|R_MERCURY&#8221;

# the feature tiers, as an experiment rather than a diagram
cat &gt; tc.cu &lt;&lt;&#8217;EOF&#8217;
#include &lt;cuda_runtime.h&gt;
__global__ void k(unsigned* o){
  asm volatile(&#8221;tcgen05.fence::before_thread_sync;&#8221; ::: &#8220;memory&#8221;);
  o[0]=1;
}
EOF
for a in sm_100 sm_100f sm_100a sm_103f sm_120a; do printf &#8220;%-9s &#8220; $a
  nvcc -arch=$a -cubin -o t.cubin tc.cu 2&gt;&amp;1 | head -1; done
for a in sm_100f sm_100a; do nvcc -arch=$a -cubin -o t_$a.cubin tc.cu
  echo &#8220;== $a&#8221;; cuobjdump -elf t_$a.cubin | grep -A2 EICOMPAT | grep -E &#8220;Attribute|Value&#8221;; done

# is the virtual architecture in e_flags? (no)
for v in 75 80 90; do nvcc -gencode arch=compute_$v,code=sm_90 -cubin -o v.cubin saxpy.cu
  python3 -c &#8220;import struct;print(&#8217;compute_$v -&gt; 0x%08x&#8217;%struct.unpack_from(&#8217;&lt;I&#8217;,open(&#8217;v.cubin&#8217;,&#8217;rb&#8217;).read(),48)[0])&#8221;
  cuobjdump -elf v.cubin | grep &#8220;Virtual SM&#8221;; done

# the field names nvdisasm knows
strings -a $(command -v nvdisasm) | grep -o &#8220;EF_CUDA_[A-Z0-9_]*&#8221; | sort -u

# scheduling control bits, 21 bits starting at 105 of each 128-bit word
nvdisasm -c -hex k90.cubin | head -40
python3 - &lt;&lt;&#8217;PY&#8217;
import struct
d=open(&#8217;k90.cubin&#8217;,&#8217;rb&#8217;).read()
o,=struct.unpack_from(&#8217;&lt;Q&#8217;,d,0x28); es,n,si=struct.unpack_from(&#8217;&lt;HHH&#8217;,d,0x3a)
sh=[struct.unpack_from(&#8217;&lt;IIQQQQIIQQ&#8217;,d,o+i*es) for i in range(n)]
st=sh[si][4]; nm=lambda x: d[st+x:d.index(b&#8217;\0&#8217;,st+x)].decode()
for s in sh:
    if not nm(s[0]).startswith(&#8217;.text.&#8217;): continue
    b=d[s[4]:s[4]+s[5]]
    for i in range(0,len(b),16):
        lo,hi=struct.unpack_from(&#8217;&lt;QQ&#8217;,b,i); w=(hi&lt;&lt;64)|lo
        print(f&#8221;{i:04x} stall={(w&gt;&gt;105)&amp;0xf} yield={(w&gt;&gt;109)&amp;1} &#8220;
              f&#8221;wr={(w&gt;&gt;110)&amp;7} rd={(w&gt;&gt;113)&amp;7} wait={(w&gt;&gt;116)&amp;0x3f:06b} &#8220;
              f&#8221;reuse={(w&gt;&gt;122)&amp;0xf:04b}&#8221;)
PY

# the fatbin container and the cost of coverage
nvcc -gencode arch=compute_90,code=sm_90 \
     -gencode arch=compute_100,code=[sm_100,compute_100] -fatbin -o f.fatbin saxpy.cu
for m in none speed balance size default; do
  nvcc -compress-mode=$m -gencode arch=compute_90,code=sm_90 \
       -gencode arch=compute_100,code=[sm_100,compute_100] \
       -fatbin -o f_$m.fatbin saxpy.cu; ls -l f_$m.fatbin; done

# the host side
nvcc -arch=sm_90 -c -o k.o saxpy.cu
readelf -SW k.o | grep -E &#8220;nv_fatbin|nvFatBinSegment|module_id&#8221;
readelf -x .nvFatBinSegment k.o
nvcc -arch=sm_90 --keep -c -o /dev/null saxpy.cu &amp;&amp; cat saxpy.cudafe1.stub.c</code></pre><p>The fatbin entry parser used for the library dissection is thirty lines: walk from the <code>0xBA55ED50</code> header, then for each entry read the kind at offset 0, header size at 4, padded payload size at 8, compressed size at 0x10, architecture at 0x1c, flags at 0x28, and skip <code>header_size + payload_size</code> to the next one.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2></h2>
      <p>
          <a href="https://www.thesoftwarefrontier.com/p/how-cuda-binaries-actually-work">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[How the NVIDIA Compiler Moat Actually Works]]></title><description><![CDATA[Inside the NVIDIA Compiler Moat: ptxas, SASS, and the 21 Bits Nobody Else Can Write]]></description><link>https://www.thesoftwarefrontier.com/p/how-the-nvidia-compiler-moat-actually</link><guid isPermaLink="false">https://www.thesoftwarefrontier.com/p/how-the-nvidia-compiler-moat-actually</guid><dc:creator><![CDATA[Lorenzo Bradanini]]></dc:creator><pubDate>Mon, 27 Jul 2026 11:08:36 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!D2wR!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdc74089c-cb20-4762-9ab1-4acb811e1150_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!D2wR!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdc74089c-cb20-4762-9ab1-4acb811e1150_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!D2wR!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdc74089c-cb20-4762-9ab1-4acb811e1150_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!D2wR!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdc74089c-cb20-4762-9ab1-4acb811e1150_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!D2wR!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdc74089c-cb20-4762-9ab1-4acb811e1150_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!D2wR!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdc74089c-cb20-4762-9ab1-4acb811e1150_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!D2wR!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdc74089c-cb20-4762-9ab1-4acb811e1150_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/dc74089c-cb20-4762-9ab1-4acb811e1150_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2763373,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/206014254?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdc74089c-cb20-4762-9ab1-4acb811e1150_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!D2wR!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdc74089c-cb20-4762-9ab1-4acb811e1150_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!D2wR!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdc74089c-cb20-4762-9ab1-4acb811e1150_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!D2wR!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdc74089c-cb20-4762-9ab1-4acb811e1150_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!D2wR!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdc74089c-cb20-4762-9ab1-4acb811e1150_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>Intro</h2><p>I spent, more or less, about a year writing an <strong>x86-64 assembler</strong> in C. Did Lexer, expression evaluator, instruction encoder, ELF writer, relocations, symbol resolution, eventually a macro preprocessor. The thing I keep coming back to from that year is how little information the assembler needed to know.</p><p>x86-64 encoding is somwthing baroque. <strong>ModRM</strong> and SIB are a whole mess, REX prefixes leak into everything, and the exact same mnemonic can have six encodings depending on operand width and register numbering. </p><p>But none of that is semantic. I emit the bytes and the machine figures out the rest in some ways. If I put two dependent instructions back to back, the processor detects the hazard, stalls, and my program suddently becomes slow. It is almost never wrong. The whole apparatus of scoreboarding, register renaming and<strong> out-of-order</strong> issue exists so that the assembler can behave like a dumb table lookup with a symbol table stapled on.</p><p>Then I started reading SASS, and that assumption stopped holding.</p><p>On modern NVIDIA hardware, if the compiler gets the scheduling metadata wrong, you do not get a slow kernel, you get directly a wrong answer. The <strong>dependency interlock</strong> for fixed-latency instructions is not in the hardware. </p><p>It is basically in a 21-bit field that <code>ptxas</code> writes into every instruction, and that field is undocumented, unstable across various architectures, and produced by exactly one program on earth.</p><p>Just that, and <strong>not the C++ dialect</strong>, and not <code>cudaMalloc</code>, is where I believe the compiler moat actually lives. What follows is a personal attempt to be precise about which layer is load-bearing and which layers people keep mistaking for it, because the public version of this argument is usually pitched at the wrong altitude.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Five moats wearing one name</h2><p>&#8220;<strong>CUDA</strong>&#8221; is, in various conversations, a giant bag holding at least five separate things, with wildly different levels of defensibility. We have:</p><p><strong>1. The source language.</strong> <code>__global__</code>, <code>threadIdx</code>, <code>&lt;&lt;&lt;grid, block&gt;&gt;&gt;</code>, a C++ dialect with a few extensions. This is the" &#8220;easiest&#8221; thing everyone points at and, by design, it is the least defensible layer. Clang has compiled CUDA for more than a decade. The NVPTX backend is just upstream LLVM. <code>libNVVM</code> is a documented public API and NVVM IR is a specified subset of LLVM IR. NVIDIA gave the frontend away years ago, precisely because giving away the frontend costs them nothing.</p><p><strong>2. The virtual ISA and the JIT distribution channel.</strong> PTX plus the compiler that ships inside <code>libcuda.so</code>. This is a genuine moat, one that I&#8217;ve personally studied in the past few months, but it is a <em>distribution</em> moat, not a performance one. It is the reason a fatbinary built in 2019 still runs on Blackwell.</p><p><strong>3. </strong><code>ptxas</code><strong> and SASS.</strong> This is interesting. We&#8217;re talking about a closed optimizing compiler that does register allocation, scheduling, and the encoding of hardware control state, pointing to an instruction set with no public specification and no official assembler. </p><p><strong>4. The libraries and the layout algebra.</strong> cuBLAS, cuDNN, CUTLASS, CuTe, NCCL, TensorRT-LLM. Thousands of pre-tuned kernel variants plus a template metaprogramming layer that encodes the tensor core&#8217;s operand layout requirements. It&#8217;s partly open, but the real thing is that it is entirely co-designed with silicon that is not public.</p><p><strong>5. The numerics contract and the install base.</strong> Now we&#8217;re serious. Reduction orders, TF32 defaults, FP8 and NVFP4 scaling recipes, atomics behaviour, the accumulated set of hyperparameter recipes that were tuned on top of all of it. I don&#8217;t know why yobody writes about this one, because it&#8217;s worth more than you can think.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!CrYh!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc9c0c1c8-c242-4531-b00c-dd379e1c5245_1634x912.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!CrYh!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc9c0c1c8-c242-4531-b00c-dd379e1c5245_1634x912.png 424w, https://substackcdn.com/image/fetch/$s_!CrYh!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc9c0c1c8-c242-4531-b00c-dd379e1c5245_1634x912.png 848w, https://substackcdn.com/image/fetch/$s_!CrYh!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc9c0c1c8-c242-4531-b00c-dd379e1c5245_1634x912.png 1272w, https://substackcdn.com/image/fetch/$s_!CrYh!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc9c0c1c8-c242-4531-b00c-dd379e1c5245_1634x912.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!CrYh!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc9c0c1c8-c242-4531-b00c-dd379e1c5245_1634x912.png" width="1456" height="813" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c9c0c1c8-c242-4531-b00c-dd379e1c5245_1634x912.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:813,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Figure 3. The five things people mean when they say CUDA, scored on whether the specificat&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 3. The five things people mean when they say CUDA, scored on whether the specificat" title="Figure 3. The five things people mean when they say CUDA, scored on whether the specificat" srcset="https://substackcdn.com/image/fetch/$s_!CrYh!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc9c0c1c8-c242-4531-b00c-dd379e1c5245_1634x912.png 424w, https://substackcdn.com/image/fetch/$s_!CrYh!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc9c0c1c8-c242-4531-b00c-dd379e1c5245_1634x912.png 848w, https://substackcdn.com/image/fetch/$s_!CrYh!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc9c0c1c8-c242-4531-b00c-dd379e1c5245_1634x912.png 1272w, https://substackcdn.com/image/fetch/$s_!CrYh!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc9c0c1c8-c242-4531-b00c-dd379e1c5245_1634x912.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong><span>Figure 1 </span></strong>The five things people mean when they say CUDA, scored on whether the specification is public, whether the source is available, whether anyone has independently reimplemented it, and whether it survives a new architecture unchanged.<span>Copy imageDownload PNG</span></figcaption></figure></div><p>Layer 1 is defintely gone. Layer 2 is a legal and logistical lock, not a technical one. Layers 3, 4 and 5 are &#8220;<em>the moat</em>&#8221;, and they defend against completely different attacks. </p><p>The vast majority &#8220;<em>CUDA alternative</em>&#8221; projects pick a layer, do good work, and then basically die on a different layer they did not budget for.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/p/how-the-nvidia-compiler-moat-actually?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/p/how-the-nvidia-compiler-moat-actually?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><div><hr></div><h2>The frontend left the building a long time ago</h2><p><code>nvcc</code> is not a compiler. It is a driver script, and you can watch it work:</p><pre><code><code>nvcc --dryrun -arch=sm_90 kernel.cu 2&gt;&amp;1 | grep -v '^#\$ *[A-Z_]*=' 
</code></code></pre><p>What you get back is a pipeline: <code>cudafe++</code> splits host from device, <code>cicc</code> (<em>the NVVM-based device compiler</em>) turns device code into PTX, the <code>ptxas</code> compiler turns PTX into a <code>cubin</code>, <code>fatbinary</code> staples the cubins and the PTX together, <code>nvlink</code> handles<strong> device linking</strong>, and your host compiler of choice gently handles all the rest. Six things doing 6 separate tasks. </p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Oo-2!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbcceed06-56f0-4264-a99a-71142732f51e_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Oo-2!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbcceed06-56f0-4264-a99a-71142732f51e_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!Oo-2!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbcceed06-56f0-4264-a99a-71142732f51e_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!Oo-2!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbcceed06-56f0-4264-a99a-71142732f51e_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!Oo-2!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbcceed06-56f0-4264-a99a-71142732f51e_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Oo-2!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbcceed06-56f0-4264-a99a-71142732f51e_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/bcceed06-56f0-4264-a99a-71142732f51e_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Figure 1. Five independent frontends converge on one virtual ISA and one closed backend. T&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 1. Five independent frontends converge on one virtual ISA and one closed backend. T" title="Figure 1. Five independent frontends converge on one virtual ISA and one closed backend. T" srcset="https://substackcdn.com/image/fetch/$s_!Oo-2!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbcceed06-56f0-4264-a99a-71142732f51e_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!Oo-2!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbcceed06-56f0-4264-a99a-71142732f51e_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!Oo-2!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbcceed06-56f0-4264-a99a-71142732f51e_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!Oo-2!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbcceed06-56f0-4264-a99a-71142732f51e_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong><span>Figure 2 </span></strong>Five independent frontends converge on one virtual ISA and one closed backend. The driver carries a second copy of that backend, which is why even a fully open toolchain has an NVIDIA compiler in its runtime path.<span>Copy imageDownload PNG</span></figcaption></figure></div><p>One of those, <code>cicc,</code> is an LLVM. NVIDIA has been explicit about this for the last few years: <strong>NVVM IR </strong>is just LLVM IR with a documented set of intrinsics and address space conventions, and <code>libNVVM</code> is a shipped, documented library you can use. </p><p>If you don&#8217;t want to use it, Clang&#8217;s own CUDA support will be happy to emit PTX for you, and it has been extreme quality, at least since Google put it there to build <strong>TensorFlow</strong>.</p><p>So the frontend nowadays is a commodity. That is why Triton, Mojo, tinygrad, Julia etc&#8230; can emit PTX, all without hurting NVIDIA&#8217;s position. Every one of those projects walks up to the same closed door at the end of the hallway.</p><p>I just want to be careful here, because &#8220;<em>the frontend is commoditized</em>&#8221; is often stated as if it were an indictment of NVIDIA&#8217;s openness, but it&#8217;s not. Ceding the frontend was correct and (probably) a deliberate choice. </p><p>Various frontends just increase the <strong>number of paths into PTX</strong>, and every path into PTX eventually terminates in a program NVIDIA controls.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>The 21-bit field</h2><p>Here is the part that changed how I think about this.</p><p><strong>Kepler</strong> version removed hardware dependency checking for fixed-latency instructions and moved the responsibility into the compiler. Then we had <strong>Maxwell </strong>refining it. </p><p><strong>Volta</strong> made it permanent by widening the instruction word from 64 bits to 128 bits and embedding the control payload directly in every instruction rather than in separate scheduling words interleaved every few instructions.</p><p>The reconstructed layout of that payload, for Volta and later, looks more or less like this:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!WZA9!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8107d78d-1c19-4c0c-ba5e-d0c60ab675d6_1634x836.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!WZA9!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8107d78d-1c19-4c0c-ba5e-d0c60ab675d6_1634x836.png 424w, https://substackcdn.com/image/fetch/$s_!WZA9!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8107d78d-1c19-4c0c-ba5e-d0c60ab675d6_1634x836.png 848w, https://substackcdn.com/image/fetch/$s_!WZA9!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8107d78d-1c19-4c0c-ba5e-d0c60ab675d6_1634x836.png 1272w, https://substackcdn.com/image/fetch/$s_!WZA9!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8107d78d-1c19-4c0c-ba5e-d0c60ab675d6_1634x836.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!WZA9!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8107d78d-1c19-4c0c-ba5e-d0c60ab675d6_1634x836.png" width="1456" height="745" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/8107d78d-1c19-4c0c-ba5e-d0c60ab675d6_1634x836.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:745,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Figure 2. The scheduling payload ptxas writes into every instruction. The stall count is t&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 2. The scheduling payload ptxas writes into every instruction. The stall count is t" title="Figure 2. The scheduling payload ptxas writes into every instruction. The stall count is t" srcset="https://substackcdn.com/image/fetch/$s_!WZA9!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8107d78d-1c19-4c0c-ba5e-d0c60ab675d6_1634x836.png 424w, https://substackcdn.com/image/fetch/$s_!WZA9!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8107d78d-1c19-4c0c-ba5e-d0c60ab675d6_1634x836.png 848w, https://substackcdn.com/image/fetch/$s_!WZA9!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8107d78d-1c19-4c0c-ba5e-d0c60ab675d6_1634x836.png 1272w, https://substackcdn.com/image/fetch/$s_!WZA9!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8107d78d-1c19-4c0c-ba5e-d0c60ab675d6_1634x836.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong><span>Figure 3 </span></strong>The scheduling payload ptxas writes into every instruction. The stall count is the field that turns a compiler into a component of the machine: for fixed latency instructions there is no hardware interlock behind it.<span>Copy imageDownload PNG</span></figcaption></figure></div><p>First. four bits of operand reuse cache hints. Then, a six-bit mask of <strong>scoreboard barriers </strong>this instruction waits on. Now on, two three-bit indices naming the barriers it sets for variable-latency results. A yield bit that tells the warp scheduler whether to prefer this warp or switch away. </p><p>And finally we have four bits of static stall count: the number of cycles the scheduler must wait before issuing the next instruction from this warp.</p><p>That last field is the one that matters for real. For fixed-latency instructions, the hardware does not check read-after-write hazards at all. The microarchitecture research on this is mostly unambiguous: if you set the stall counter wrong and the program produces incorrect results, because there is <strong>no scoreboard behind </strong>you to catch it. </p><p>The register status table and the wires from the issue logic to it were purposefully deleted from the design, and the area and energy they would have cost were spent on something else. The compiler is the scoreboard.</p><p>On x86, the ISA is a contract about <em>meaning</em> and the microarchitecture is free to reorder underneath it as long as the contract holds. On NVIDIA hardware since Kepler, a <strong>meaningful piece </strong>of the microarchitecture&#8217;s correctness lives in bits the compiler writes. </p><p><code>ptxas</code> is not a tool that targets the machine, because it is a component of the machine that happens to run on your workstation. Those are the obvious consequences. </p><p>You cannot open source that boundary without publishing the timing characteristics of every functional unit in <strong>every SM variant</strong> of every SKU, including the ones you have not announced. And if you publish it, you have frozen it, because now third-party code depends on it and any change basically breaks correctness, not just performance. </p><p>The closedness is not primarily a business decision. It is downstream of an architectural decision that was made for <strong>area and power reasons </strong>around 2012, and the strategic value came along for the ride.</p><p>I find that way interesting than the general story, even thoigh it&#8217;s more depressing, because it means the moat is not a policy anyone can be lobbied to reverse.</p><h3>Occupancy is a compiler policy, not a hardware property</h3><p>The other thing <code>ptxas</code> decides on your behalf, and the one you can actually measure without a disassembler, is <strong>how many registers</strong> each thread gets.</p><p>We alredy know that an H100 SM has a <strong>256 KB register file</strong>: 65,536 registers of 32 bits, allocated to warps at a granularity of eight registers per thread. The SM can hold at most 64 resident warps. Those three numbers set the whole occupancy calculation, and the only free variable is the register count <code>ptxas</code> chose.</p><p>Take a kernel where <code>ptxas -v</code> reports 168 registers per thread. Each warp consumes <code>168 x 32 = 5,376</code> registers, so the SM can hold <code>floor(65536 / 5376) = 12</code> warps, which is 12 of 64, or 18.75 percent occupancy. Now suppose you pass <code>-maxrregcount=128</code>. </p><p>The same warp now costs 4,096 registers, 16 warps fit, occupancy goes to roughly 25 percent, and <code>ptxas</code> pays for the difference by <strong>pushing the excess to local memory</strong>, which is to say to L1 and then to L2 and then, if you are unlucky and you have a giant working set, to HBM.</p><p>Neither number is right on its own. It depends on what the kernel spends its time waiting for. A kernel that <strong>waits on memory</strong> wants more warps. When one warp stalls on a load, the scheduler switches to another one, and if enough warps are resident, the waiting never shows up in the runtime. Occupancy is what keeps that kernel fed.</p><p>On the opposite side, a kernel that lives in the tensor cores usually wants the registers. It already hides its own waiting, by fetching the next tile while the current one is being multiplied. <strong>It has a pipeline</strong>, so it has nothing to gain from having other warps to switch to, and it will run happily at <em>eighteen percent occupancy.</em></p><p>That tradeoffs are the single most consequential decision in GPU performance work, and they are made by a <strong>heuristic inside a binary </strong>that I believe is impossible to read, tune, or replace. <code>-maxrregcount</code> and <code>__launch_bounds__</code> are the two knobs you get, and they are blunt.</p><p>I want to highlight this because it is the everyday version of the argument. You don&#8217;t need to care about control bits to be affected by the closed backend. If you have ever watched a <strong>two-line source change move register count</strong> from 168 to 176 and cost you a resident warp, you have already been on the wrong side of this wall.</p><p>There is no official SASS assembler. Instead, there&#8217;s a disassembler, <code>nvdisasm</code>, and <code>cuobjdump --dump-sass</code>, and the documentation for the instruction set amounts to a table of mnemonics with one-line descriptions. The opcode encodings, the<strong> control field layout</strong>, and the exact latency table are reverse engineered by a small group of people, generation by generation.</p><p>The lineage of that work is short enough to list: <code>asfermi</code> for Fermi, Scott Gray&#8217;s <code>maxas</code> for Maxwell, <code>KeplerAs</code>, <code>TuringAs</code>, <code>CuAssembler</code> for Pascal all the way through Ampere, <code>openptxas</code> under the <strong>gpuocelot umbrella</strong> for Maxwell and Pascal. Every one of these is a heroic, partial, architecture-locked effort that goes stale the moment a new chip make it to the public.</p><p>The payoff of that is real, which is the frustrating part in my opinion. Gray&#8217;s <code>maxas</code> SGEMM reached figures close to theoretical peak on GM204, ahead of the cuBLAS of that era, purely by scheduling better than <code>ptxas</code> did. </p><p>More recently there is paper work applying reinforcement learning to SASS schedules directly (<code>CuAsmRL</code>), reordering instructions inside <code>cubin</code>s that <code>ptxas</code> already produced and finding wins. If <code>ptxas</code> were optimal, that entire research direction would return zeros.</p><p>I have personally enjoyed one single detail, that I&#8217;m gonna show you below:</p><pre><code><code>/*0080*/ LDG.E.128 R8,  desc[UR6][R2.64]    ;  /* [----:B----:R-:W2:-:S01] */
/*0090*/ LDG.E.128 R12, desc[UR6][R4.64]    ;  /* [----:B----:R-:W3:-:S01] */
/*00a0*/ IMAD.WIDE R2,  R0, R7, c[0x0][0x168] ;/* [----:B----:R-:-:-:S02] */
/*00b0*/ HMMA.16816.F32 R20, R8, R12, R20   ;  /* [--23:B----:R-:-:-:S04] */
                          ^          ^
                          |          waits on write barriers 2 and 3
                          |          set by the two loads above
                          stall 4 cycles before the next issue
</code></code></pre><p><em>This is a reconstructed fragment in the shape that </em><code>cuobjdump --dump-sass</code><em> prints, not literal output from a specific compilation, but the bracketed control payload follows the documented conventions. </em></p><p><br>Please <strong>take a deep look</strong> at the two loads first. Each one sets a write barrier, 2 and 3 respectively, because a global load has variable latency and the compiler cannot know statically when the data lands. </p><p>Then the <code>IMAD</code> in between carries no barrier at all <em>(only a stall count) </em>because integer multiply-add is fixed latency and the compiler knows exactly how long it takes. The tensor core instruction waits on both barriers before it issues.</p><p>Now notice what is doing the work. The <code>S02</code> on the <code>IMAD</code> is not an optimization hint. It is the compiler asserting the <strong>pipeline depth of the integer unit</strong> on this specific SM. Change the SM, keep the number, and the machine reads a register before the previous instruction has written it. </p><p>One last thing: <strong>NVIDIA Research&#8217;s </strong>own SASS instrumentation framework, SASSI, was distributed as a <em>closed-source fork of </em><code>ptxas</code><em> binaries</em>. Their own researchers, inside the company, shipped a patched binary rather than source. That suggests you something about how the artifact is treated internally.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share The Software Frontier&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share The Software Frontier</span></a></p><div><hr></div><h2>PTX is a treaty, not an ISA</h2><p>I often read posts about people describing PTX as &#8220;<em>NVIDIA&#8217;s assembly language,</em>&#8221; which is so wrong in a way that hides the actual mechanism. <strong>PTX is a virtual ISA</strong> with unbounded virtual registers, no notion of the physical register file, no scheduling, and no control bits. </p><p>It is much closer to LLVM IR than to x86 assembly. Writing PTX by hand does not get you anywhere near the metal; it gets you a <em>slightly lower-level input</em> to the same closed optimizer.</p><p><strong>DeepSeek&#8217;s V3</strong> work is the canonical example, and it is routinely misreported. They partitioned twenty of the H800&#8217;s 132 SMs to handle inter-node communication, used warp specialization, and wrote custom PTX to reduce L2 pressure and interference with the compute kernels. </p><p>That is obv. excellent engineering under export-control constraints, but is not an escape from CUDA. It is a <em>deeper</em> commitment to CUDA: PTX pins you to NVIDIA <strong>harder than CUDA C++</strong> does, simply because CUDA C++ at least has clean-room reimplementations and PTX has one consumer.</p><p>What PTX actually buys NVIDIA is the forward compatibility story, and it is a genuinely excellent piece of platform engineering.<strong> Ship PTX inside your fatbinary</strong>, and when your binary lands on hardware it has never seen, the driver compiles it at load time. <code>libcuda.so</code> contains a compiler. Every NVIDIA GPU in the world is shipped with a copy of the last-mile toolchain baked into the driver.</p><p>The <strong>strategic consequences</strong> of that are larger than the technical ones:</p><ul><li><p><em>Any stack that emits PTX, no matter how open, has a closed NVIDIA compiler in its runtime path. Triton is open source and ships PTX to </em><code>ptxas</code><em>. tinygrad in </em><code>PTX=1</code><em> mode does the same. Open at the top, closed at the bottom.</em></p></li><li><p><em>Compatibility is a one-way ratchet. Old PTX runs on new hardware. New PTX never runs on old hardware. Every generation, the set of instructions you need to hit peak throughput moves into a </em><code>.target sm_XXa</code><em> gate that the previous generation&#8217;s driver cannot parse.</em></p></li><li><p><em>The world&#8217;s binaries contain PTX. That corpus is a compatibility obligation NVIDIA has taken on, and it is also an asset nobody else can serve.</em></p></li></ul><h3>The compatibility promise does not cover the instructions that matter</h3><p>This is the part of the PTX story I have not seen written down anywhere, but I do believe it changes the conclusion.</p><p>PTX <strong>forward compatibility</strong> applies to base architecture targets. <code>sm_90</code>, <code>sm_100</code>, and so on. But the instructions that reach peak throughput do not live there. <code>wgmma</code> requires <code>sm_90a</code>. The <code>tcgen05</code> family requires <code>sm_100a</code> or <code>sm_103a</code>. </p><p>That &#8220;<code>a</code> suffix&#8221; means architecture-specific, and code compiled for an <code>a</code> target is explicitly <em>not</em> forward compatible: it <strong>runs on that architecture</strong> and nowhere else, ever. NVIDIA later added <code>f</code> family targets to soften this within a family, which is an admission that the problem is real.</p><p>Put the two facts next to each other. The forward compatibility that everyone cites as CUDA&#8217;s great platform virtue covers the instruction set you use when you <strong>do not care about performance</strong>. The moment you write a kernel that actually saturates the tensor cores, you have opted out of it, and you are shipping architecture-locked binaries exactly like everyone else.</p><p>So the<strong> JIT compatibility story</strong> is not a technical moat at all; it is a convenience for the long tail, and the long tail is not where the money is. What it does buy NVIDIA is that the world&#8217;s software ships PTX, and PTX has one consumer.</p><p>And then there is the legal layer, which people mention and then move past, maybe too quickly. The <strong>CUDA EULA</strong> has, since 2021 online and since the 11.6 installed files, carried a restriction on reverse engineering, decompiling or disassembling the output generated using SDK elements for the purpose of translating those artifacts to target a non-NVIDIA platform. </p><p>Read it as written: it is a restriction on what you may do with <em>the output of the compiler</em>, not just the compiler. Whatever its enforceability, and I am not a lawyer, its practical effect is a chilling one on exactly the class of project that would otherwise attack layer 2.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>The libraries are an admission, not a flex</h2><p>The standard framing is that <strong>cuBLAS</strong> and <strong>cuDNN</strong> are evidence of NVIDIA&#8217;s software superiority. I want you to read them the opposite way, like I do.</p><p>If <code>nvcc</code> and <code>ptxas</code> could reliably generate peak code from readable source, cuBLAS would be a thin wrapper over a few templated kernels. Instead it is a multi-hundred-megabyte binary containing<strong> thousands of pre-compiled kernel variants</strong>, selected at runtime by heuristics, many of them tuned by hand at the SASS level by people who work at NVIDIA. </p><p>cuDNN is even worse. The size of those <code>.so</code> files is a direct measurement of the gap between what the compiler can do and what the hardware can do.</p><p>That <strong>gap is the real product. </strong>NVIDIA is not selling you a compiler that makes your code fast. It is selling you the output of a large, expensive, permanently-employed team of kernel engineers, wrapped in an API, in a form you cannot fork.</p><h3>And the gap is widening, on purpose</h3><p>The<strong> tensor core</strong> programming model has changed shape three times in six years, and each change moved way further away from anything a general compiler can target from ordinary source.</p><p>Ampere gave us <code>mma</code>, warp-level, operands in registers, with layout requirements you could hold in your head if you tried. Hopper introduced <code>wgmma</code>, warpgroup-scoped and asynchronous, plus TMA for <strong>bulk asynchronous copies</strong> driven by descriptors, plus <code>mbarrier</code> for the synchronization that async now requires. </p><p>Blackwell deprecated <code>wgmma</code> entirely and introduced the <code>tcgen05</code> family: a dedicated <strong>on-chip Tensor Memory</strong> for accumulators, a <em>single thread</em> issuing the MMA on behalf of everyone, and <code>cta_group::2</code> letting two CTAs on a TPC cooperate on one MMA with shared operands.</p><p>Read that again as a compiler person. The CuTe MMA atoms for Blackwell use a thread ID layout of one, where Hopper used 128. At peak performance, the <strong>SIMT model is a fiction.</strong> </p><p>What is actually happening is that one elected lane fires a descriptor at an accelerator, and the rest of the code is a software pipeline built out of barriers, async copies and a memory space whose allocation you manage explicitly.</p><p>This is why the ceiling moved from the compiler to the <strong>layout algebra. </strong>You will see some numbers, because the abstraction hides how physical this is. </p><blockquote><p><em>Tensor Memory on Blackwell is 256 KB per SM, organised as 128 lanes of 512 columns of 32 bits. You allocate it explicitly, in columns, in power-of-two units with a minimum of 32. </em></p><p><em>A single warp can only reach the 32 lanes matching its position in the warpgroup, which is why the copy instruction has multicast modes to duplicate a tile across all four lane quadrants. </em></p><p><em>Block-scaled narrow formats add their own geometry on top: MXFP8 carries one E8M0 scale per block of 32 elements, NVFP4 carries an E4M3 scale per block of 16 plus a second-level FP32 tensor scale, and the scale factors themselves have to be staged into tensor memory in a specific layout before the MMA will consume them.</em></p></blockquote><p>None of that is expressible as &#8220;<em>the compiler will figure it out</em>&#8221;. It is a data layout problem with <strong>hardware-imposed</strong> alignment constraints that propagate back into your choice of tile shape, which propagates back into your choice of pipeline depth, which propagates back into your register budget.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!22VA!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9626b009-bde1-49eb-8d33-56bab7959aec_1634x912.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!22VA!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9626b009-bde1-49eb-8d33-56bab7959aec_1634x912.png 424w, https://substackcdn.com/image/fetch/$s_!22VA!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9626b009-bde1-49eb-8d33-56bab7959aec_1634x912.png 848w, https://substackcdn.com/image/fetch/$s_!22VA!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9626b009-bde1-49eb-8d33-56bab7959aec_1634x912.png 1272w, https://substackcdn.com/image/fetch/$s_!22VA!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9626b009-bde1-49eb-8d33-56bab7959aec_1634x912.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!22VA!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9626b009-bde1-49eb-8d33-56bab7959aec_1634x912.png" width="1456" height="813" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/9626b009-bde1-49eb-8d33-56bab7959aec_1634x912.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:813,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Figure 4. Ampere issued an MMA from a warp, Hopper from a warpgroup, Blackwell from a sing&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 4. Ampere issued an MMA from a warp, Hopper from a warpgroup, Blackwell from a sing" title="Figure 4. Ampere issued an MMA from a warp, Hopper from a warpgroup, Blackwell from a sing" srcset="https://substackcdn.com/image/fetch/$s_!22VA!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9626b009-bde1-49eb-8d33-56bab7959aec_1634x912.png 424w, https://substackcdn.com/image/fetch/$s_!22VA!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9626b009-bde1-49eb-8d33-56bab7959aec_1634x912.png 848w, https://substackcdn.com/image/fetch/$s_!22VA!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9626b009-bde1-49eb-8d33-56bab7959aec_1634x912.png 1272w, https://substackcdn.com/image/fetch/$s_!22VA!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9626b009-bde1-49eb-8d33-56bab7959aec_1634x912.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong><span>Figure 4 </span></strong>Ampere issued an MMA from a warp, Hopper from a warpgroup, Blackwell from a single elected lane. Meanwhile the operands and the accumulator migrated out of the register file entirely.<span>Copy imageDownload PNG</span></figcaption></figure></div><p>This is why the ceiling moved from the compiler to the layout algebra. You cannot express a <strong>competitive Blackwell GEMM</strong> by writing loops and hoping. </p><p>You express it as a set of layouts, and the algebra of those layouts is <strong>CuTe</strong>, and CuTe is co-designed with the silicon by the people who designed the silicon, eighteen months before you can buy it.</p><p>The single most clarifying data point I have seen on this: FlashAttention 4, the first version written in <strong>CuTe DSL</strong> rather than C++ templates, reportedly hits about 1613 TFLOP/s on B200 for bf16 head-dim 128 causal attention at 8K sequence length, roughly 71 percent of peak, about 1.3 times faster than cuDNN 9.13, and around <em>2.7 times faster</em> than Triton on the same hardware.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!vR4q!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2987cdf1-5bb5-4e21-ade3-583932f9909b_1634x836.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!vR4q!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2987cdf1-5bb5-4e21-ade3-583932f9909b_1634x836.png 424w, https://substackcdn.com/image/fetch/$s_!vR4q!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2987cdf1-5bb5-4e21-ade3-583932f9909b_1634x836.png 848w, https://substackcdn.com/image/fetch/$s_!vR4q!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2987cdf1-5bb5-4e21-ade3-583932f9909b_1634x836.png 1272w, https://substackcdn.com/image/fetch/$s_!vR4q!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2987cdf1-5bb5-4e21-ade3-583932f9909b_1634x836.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!vR4q!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2987cdf1-5bb5-4e21-ade3-583932f9909b_1634x836.png" width="1456" height="745" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2987cdf1-5bb5-4e21-ade3-583932f9909b_1634x836.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:745,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Figure 5. One benchmark, three conclusions: the closed library is not the ceiling, the til&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 5. One benchmark, three conclusions: the closed library is not the ceiling, the til" title="Figure 5. One benchmark, three conclusions: the closed library is not the ceiling, the til" srcset="https://substackcdn.com/image/fetch/$s_!vR4q!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2987cdf1-5bb5-4e21-ade3-583932f9909b_1634x836.png 424w, https://substackcdn.com/image/fetch/$s_!vR4q!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2987cdf1-5bb5-4e21-ade3-583932f9909b_1634x836.png 848w, https://substackcdn.com/image/fetch/$s_!vR4q!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2987cdf1-5bb5-4e21-ade3-583932f9909b_1634x836.png 1272w, https://substackcdn.com/image/fetch/$s_!vR4q!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2987cdf1-5bb5-4e21-ade3-583932f9909b_1634x836.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong><span>Figure 5 </span></strong>One benchmark, three conclusions: the closed library is not the ceiling, the tile level DSL is the productive authoring layer, and the portable tile language is a long way off peak on the newest tensor cores.<span>Copy imageDownload PNG</span></figcaption></figure></div><p>Three conclusions fall out of one benchmark. NVIDIA&#8217;s own closed, hand-tuned attention library was not the ceiling. </p><p>The<strong> tile-level DSL</strong>, not the C++ template library, is now the productive authoring layer. And Triton, the great portable hope, is nowhere near peak on the newest tensor cores, because its abstractions do not yet expose what <code>tcgen05</code> and TMA require.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>What is actually in a cubin</h2><p>I want to discuss and be clear about one thing worth opening up, because it is the part that felt most familiar to me coming from ELF.</p><p>A <code>cubin</code> is an <strong>ELF64 object.</strong> S ame header layout you already know, with a CUDA-specific machine type. Inside, you get a <code>.text</code> section per kernel holding the SASS, a <code>.nv.constant0</code> section per kernel holding the kernel parameters and grid constants (<em>this is why parameters arrive via </em><code>c[0x0][...]</code><em> in the disassembly rather than in registers</em>), <code>.nv.info</code> sections carrying metadata like register counts and parameter layout, plus the usual relocation sections. </p><p>The fatbinary that <code>nvcc</code> embeds in your host object is itself a container of these, in a <code>.nv_fatbin</code> section, with a small descriptor segment telling the runtime what is inside.</p><p>You can walk all of it here:</p><pre><code><code>cuobjdump -elf a.out            # fatbin contents, section by section
nvdisasm -elf kernel.cubin      # ELF structure plus disassembly
readelf -S kernel.cubin         # it really is just ELF
</code></code></pre><p>The greatest absent, once you have looked, is the same one as before. <strong>Every part of this container is inspectable</strong>: the relocations and the metadata are readable, the instruction stream is disassemblable. What you can&#8217;t do is produce a valid one from modified SASS, because the only program that writes the <code>.text</code> section is the one you do not have.</p><p>Coming from x86, that asymmetry is the strangest part of the whole stack. I can hand-assemble an ELF object, link it, and run it. Here I can read everything and write nothing.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>What is not being told</h2><p>Most of the commentary stops at &#8220;<em>CUDA is sticky.</em>&#8221; Here are the parts I think are underweighted, roughly in order of how much they change the conclusion.</p><h3>The compiler is correctness-critical, not just performance-critical</h3><p>Again, sorry for that, but I have not seen this stated plainly anywhere outside the microarchitecture literature. Every discussion of open SASS toolchains treats the problem as &#8220;<em>we would generate slower code</em>&#8221;. </p><p>The actual problem is that a <strong>third-party SASS assembler</strong> that mis-models the latency of one functional unit on one SKU generates <em>silently incorrect</em> results, in a data-dependent way, on a subset of hardware. </p><p>That is a categorically harder engineering problem than performance parity, and it is the reason no serious commercial effort has attempted it. The risk profile is closer to writing a memory controller than to writing a backend.</p><h3>The moat is a labor market</h3><p>Strip away the branding and ask what actually<strong> cannot be replicated</strong> and you&#8217;ll find out that is not the compiler source. </p><p>It is the few hundred people worldwide who can look at a <code>cuobjdump</code> listing, reason about the control bits, and know from experience that a particular shared memory access pattern will hit a bank conflict on this generation but not the last. </p><p>A large fraction of them work at NVIDIA. Most of the rest work at four or five labs and two or three GPU startups. That reframing matters because it tells you what the actual counter-technology is. Many have attemped to address with it, but no, it&#8217;s not a new language. </p><p>It is just search: autotuners, superoptimizers, and increasingly <strong>LLM-driven kernel generation</strong>, all of which convert scarce expert labour into abundant compute. If the moat is a hiring problem, then the thing that erodes it is the thing that makes kernel expertise reproducible.</p><h3>There is a telemetry loop nobody else has</h3><p>Nsight, <strong>DCGM</strong>, the driver, customer escalations from every large training run on the planet. NVIDIA sees the performance pathologies of the entire industry&#8217;s kernels, continuously, at a scale no competitor approaches. </p><p>Then it fixes the top ones in the next library release and, where the fix has to be in hardware, in the next architecture. Every &#8220;<em>CUDA maturity</em>&#8221; argument is really an argument about the number of laps this loop has run.</p><h3>The collectives are compiled kernels, and they eat your SMs</h3><p>I listed the interconnect <strong>under layer 5 </strong>and then almost walked past it, which would have been a mistake, because NCCL is where the compiler moat and the network moat turn out to be exactly the same thing.</p><p>A collective is not a library call that hands work to a NIC, it&#8217;s a set of CUDA kernels. NCCL builds rings and trees over the <strong>discovered topology</strong>, splits each collective across some number of channels, and every channel is a resident thread block. </p><p>Which means bandwidth on the wire is bought with SMs that are then not available for compute. It also means the protocol matters: <strong>NCCL </strong>picks between a low latency mode that trades payload efficiency for skipping memory fences, a 128 byte variant that recovers most of the bandwidth but wants <strong>NVLink underneath</strong>, and a plain mode that gets full bandwidth at higher latency. </p><p>The choice is made by heuristics against measured topology, and the tuning knobs are environment variables. We see probably at least two consequences that are worth holding onto.</p><p>First, this is exactly what DeepSeek was doing when they carved<strong> twenty SMs</strong> out of 132 for communication. Their work was quite revolutionary because they were taking<strong> manual control</strong> of a tradeoff that NCCL normally makes for you, because on export-limited H800s with halved inter-GPU bandwidth, NCCL&#8217;s defaults were wrong for their shape. All of that, as stated before, working with CUDA, not around it or without it.</p><p>Second, <strong>in-network reduction</strong> changes the arithmetic entirely. When the switch itself can perform the reduction, the collective stops being a kernel that moves the same bytes several times and becomes a kernel that ships bytes once. </p><p>That capability lives in the NVSwitch and InfiniBand silicon, it is exposed through the same library, and there is no version of &#8220;<em>port your kernels</em>&#8221; that gets you it. AMD&#8217;s answer, <strong>RCCL</strong>, is a real port of the same design, and the vLLM tuning guidance for it involves pushing the channel count up sharply on <strong>MI300X</strong>, which tells you how much of the performance is in the topology-specific constants rather than the algorithm.</p><p>If you are keeping a scorecard of which moats are compiler-shaped, this one is a hybrid, and it is the least portable thing in the stack.</p><p>Port a model from <strong>cuBLAS </strong>to hipBLASLt and the arithmetic is not bit-identical for a few reasons. Split-k strategies differ, even reduction orders do not behave the same. </p><blockquote><p><em>TF32 is on by default in one place and not another. </em></p><p><em>FP8 scaling granularity, per-tensor versus per-block versus the MX and NVFP4 formats, is a semantic choice baked into the kernels.</em></p></blockquote><p>For inference this usually shows up as a small quality wobble that people can get along with. For training, it shows up as a recipe that diverges at step 40,000 for <strong>some random reasons</strong> nobody can bisect, on a piece of hardware where you have less tooling. </p><p>The cost of that risk almost never appears in the total-cost-of-ownership spreadsheets that compare dollars per FLOP, and it is a real reason large labs pay the NVIDIA premium.</p><h3>Watch which artifact they do not publish</h3><p>There is a pattern in NVIDIA&#8217;s open sourcing, and once you see it you can predict the next move. They publish the specification and the dialect, while keeping the lowering.</p><p><strong>NVVM IR</strong>: specified, with a public library. <code>cicc</code>: closed. PTX: exhaustively documented, hundreds of pages, updated every release. <code>ptxas</code>: closed. CUTLASS: open, permissively licensed, genuinely excellent. </p><p>The SASS it lowers to: closed. And now<strong> CUDA Tile IR</strong>: the MLIR dialect, the Python bindings, the bytecode format and the conformance suite are on GitHub. The compiler that turns tile bytecode into <code>cubin</code> is not.</p><p>This is a coherent, repeated strategy. Publishing the interface grows the number of producers. Keeping the lowering means every producer terminates in your binary. </p><p>It is the same move IBM made with channel architecture and the same move ARM makes with the ISA versus the implementations, and NVIDIA runs it more cleanly than either.</p><h3>AMD is the counterexample that breaks the simple story</h3><p>Here is the fact that should discipline every &#8220;<em>closed compiler is the moat</em>&#8221; argument: AMD publishes its ISA documents. The <strong>AMDGPU </strong>backend is upstream LLVM. ROCm is open source, top to bottom, including the compiler and the runtime. </p><p>You can read the machine encodings for CDNA3 and CDNA4 in a PDF on AMD&#8217;s website But (unexpectedly) even this has not collapsed the moat.</p><p>So openness is not the mechanism. If it were, AMD would have won already a few years ago. The mechanism is the <strong>co-design loop</strong>, the library labour, the numerics contract, the install base and the interconnect, and NVIDIA&#8217;s closed compiler is a symptom of the same architectural strategy that produces the rest of it. </p><p>Anyone whose plan is &#8220;<em>make an open CUDA</em>&#8221; has already been run as an experiment, by a company with $25 billion of revenue, for at least a decade.</p><h3>The moat is thinning at the top and thickening at the bottom</h3><p><strong>Two things are true at once</strong> and they point in opposite directions, which is again interesting because it sounds counterintuitive and understandable at the same time.</p><p>At the top, the default path <strong>from PyTorch to the GPU</strong> is now a portable tile language. <code>torch.compile</code> emits Triton, not CUDA C++. Liger, FlexAttention and a large fraction of production fused kernels are&#8230;. yes, again, Triton. That layer is vendor-neutral by construction, and it is where most new kernel code is being written.</p><p>At the bottom, the instructions you need to reach peak on the newest silicon are <strong>less reachable from a generic compiler </strong>than at any point since 2007. TMA descriptors, tensor memory allocation, CTA-pair MMA, block-scaled narrow formats with alignment constraints that constrain your tile shapes. </p><p>The gap between &#8220;<em>compiles and runs</em>&#8221; and &#8220;<em>hits peak</em>&#8221; is the widest it has ever been in 19 years.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!q8BO!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6bd9d02f-7d35-46fd-a5a8-f9d1c3fb3f46_1634x931.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!q8BO!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6bd9d02f-7d35-46fd-a5a8-f9d1c3fb3f46_1634x931.png 424w, https://substackcdn.com/image/fetch/$s_!q8BO!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6bd9d02f-7d35-46fd-a5a8-f9d1c3fb3f46_1634x931.png 848w, https://substackcdn.com/image/fetch/$s_!q8BO!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6bd9d02f-7d35-46fd-a5a8-f9d1c3fb3f46_1634x931.png 1272w, https://substackcdn.com/image/fetch/$s_!q8BO!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6bd9d02f-7d35-46fd-a5a8-f9d1c3fb3f46_1634x931.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!q8BO!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6bd9d02f-7d35-46fd-a5a8-f9d1c3fb3f46_1634x931.png" width="1456" height="830" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6bd9d02f-7d35-46fd-a5a8-f9d1c3fb3f46_1634x931.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:830,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Figure 9. The binding constraint has moved four times in nineteen years. It sat on the sou&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 9. The binding constraint has moved four times in nineteen years. It sat on the sou" title="Figure 9. The binding constraint has moved four times in nineteen years. It sat on the sou" srcset="https://substackcdn.com/image/fetch/$s_!q8BO!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6bd9d02f-7d35-46fd-a5a8-f9d1c3fb3f46_1634x931.png 424w, https://substackcdn.com/image/fetch/$s_!q8BO!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6bd9d02f-7d35-46fd-a5a8-f9d1c3fb3f46_1634x931.png 848w, https://substackcdn.com/image/fetch/$s_!q8BO!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6bd9d02f-7d35-46fd-a5a8-f9d1c3fb3f46_1634x931.png 1272w, https://substackcdn.com/image/fetch/$s_!q8BO!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6bd9d02f-7d35-46fd-a5a8-f9d1c3fb3f46_1634x931.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong><span>Figure 6 </span></strong>The binding constraint has moved four times in nineteen years. It sat on the source language only during the first five, and the two most recent moves went in opposite directions: down into the tensor core, and up into a new virtual ISA.<span>Copy imageDownload PNG</span></figcaption></figure></div><p>Now the interesting question is not whether the moat holds. It is more which layer the industry ends up standardizing on, and NVIDIA has just made a very large move to answer that question in its own favour.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Tile IR: the moat gets rebuilt one level up</h2><p>The most important compiler event of the last twelve months got almost no coverage outside compiler circles and a restricted group of GPU programmers.</p><p>CUDA 13.1 introduced <strong>CUDA Tile</strong>, a tile-based programming model with its own intermediate representation. Tile IR is an MLIR dialect. The dialect, the Python bindings, the bytecode serializer and a conformance test suite are open source at <code>NVIDIA/cuda-tile</code>. </p><p>The bytecode format is explicitly <strong>specified as stable</strong>, versioned, and forward and backward compatible: bytecode emitted by an older compiler is readable by a newer compiler or driver, and a driver accepts bytecode up to its supported version.</p><p>That paragraph should sound familiar, because it is the <strong>PTX contract</strong>, restated one abstraction level higher, in MLIR, with a published binary encoding. Tile IR is the new PTX.</p><p>The user-facing language is cuTile, a <strong>Python DSL</strong>, with a Julia binding already shipping at near parity on simple kernels. And critically, NVIDIA wrote a Triton backend that targets Tile IR, so Triton programs can be compiled through the new path instead of through PTX. </p><p>Meta has filed an RFC to add a cuTile backend to <strong>PyTorch Inductor</strong>, with the stated motivation that Tile IR gives bytecode portability, performance portability across GPU generations, and access to a tile-specific optimizing compiler, and that it is a prerequisite for pointing Helion at Tile IR as well.</p><p>Pro tip: read the sequence of moves rather than the individual announcements.</p><p>Triton became the default codegen target of PyTorch. That made the tile abstraction, the layer where new kernels get written, and that layer is portable across vendors. It was the most serious <strong>structural threat</strong> to the moat in fifteen years, because it made the frontend irrelevant <em>and</em> gave AMD and Intel a credible path to the same source.</p><p>NVIDIA&#8217;s response was not to fight Triton: it was to build a <strong>better tile IR,</strong> open source the dialect and the spec, write the Triton backend themselves, and offer everyone a <strong>faster path</strong> through their own infrastructure. If it works, the portable tile layer that was supposed to route around NVIDIA becomes a frontend for NVIDIA&#8217;s IR, and the optimizing compiler underneath, the one that turns tiles into <code>cubin</code>, is closed, just like <code>ptxas</code>.</p><p>It is the cleanest execution of the &#8220;<em>publish the interface, keep the lowering</em>&#8221; strategy I have seen so far, and it happened while everyone was arguing about whether ROCm had caught up.</p><p>I want to flag my uncertainty honestly: it is early. Tile IR shipped in December 2025, the <strong>Inductor RFC</strong> is in progress, and I have not seen independent third-party benchmarks of cuTile against hand-written CuTe on Blackwell at production shapes. </p><p>NVIDIA&#8217;s own claim is that cuTile is competitive with cuBLAS for large GEMMs and cuDNN for large attention. Competitive with cuDNN is a lower bar than it used to be, given FA4.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>The alternatives, honestly</h2><p>Every project below is real engineering by capable people. What I want to do is say which layer each one attacks and where it runs out of road, because most of them are described by their advocates as if they attack all five.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!vtGj!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbfdc9857-fe72-401d-8dd7-9f39f30d08d0_1634x988.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!vtGj!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbfdc9857-fe72-401d-8dd7-9f39f30d08d0_1634x988.png 424w, https://substackcdn.com/image/fetch/$s_!vtGj!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbfdc9857-fe72-401d-8dd7-9f39f30d08d0_1634x988.png 848w, https://substackcdn.com/image/fetch/$s_!vtGj!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbfdc9857-fe72-401d-8dd7-9f39f30d08d0_1634x988.png 1272w, https://substackcdn.com/image/fetch/$s_!vtGj!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbfdc9857-fe72-401d-8dd7-9f39f30d08d0_1634x988.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!vtGj!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbfdc9857-fe72-401d-8dd7-9f39f30d08d0_1634x988.png" width="1456" height="880" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/bfdc9857-fe72-401d-8dd7-9f39f30d08d0_1634x988.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:880,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Figure 7. Where each project actually pushes, and where it stops. Most of them are pitched&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 7. Where each project actually pushes, and where it stops. Most of them are pitched" title="Figure 7. Where each project actually pushes, and where it stops. Most of them are pitched" srcset="https://substackcdn.com/image/fetch/$s_!vtGj!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbfdc9857-fe72-401d-8dd7-9f39f30d08d0_1634x988.png 424w, https://substackcdn.com/image/fetch/$s_!vtGj!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbfdc9857-fe72-401d-8dd7-9f39f30d08d0_1634x988.png 848w, https://substackcdn.com/image/fetch/$s_!vtGj!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbfdc9857-fe72-401d-8dd7-9f39f30d08d0_1634x988.png 1272w, https://substackcdn.com/image/fetch/$s_!vtGj!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbfdc9857-fe72-401d-8dd7-9f39f30d08d0_1634x988.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong><span>Figure 7 </span></strong>Where each project actually pushes, and where it stops. Most of them are pitched as if they cover the whole stack.<span>Copy imageDownload PNG</span></figcaption></figure></div><h3>HIP and ROCm</h3><blockquote><p><strong>Attacks:</strong> layer 1 and layer 4. A CUDA-shaped API, a source translation tool, and a library set with matching names.</p></blockquote><p>ROCm in 2026 is a genuinely different product from ROCm in 2023, and people who formed their opinion during the Frontier era are working from stale data. </p><p>ROCm 7.2 ships tuned hipBLASLt GEMM kernels for FP8, BF16 and FP16 on MI350 and MI355X. AITER provides fused kernels that, inside vLLM, deliver large multiples over the legacy ROCm attention path. vLLM now carries seven attention backends on ROCm. </p><p>MI355X ships 288GB of HBM3E, and capacity per package is a real architectural advantage that removes tensor parallelism from model sizes where NVIDIA needs it.</p><blockquote><p><strong>Where it runs out of road:</strong> the API-clone strategy structurally guarantees you arrive second. <code>hipify</code> ignores inline PTX, which means exactly the kernels that were worth hand-optimizing are the ones that do not port. Matrix core layouts (<code>MFMA</code>) differ enough from <code>mma</code>, <code>wgmma</code> and <code>tcgen05</code> that a &#8220;<em>ported</em>&#8221; kernel is a rewrite at the layout level. And the performance story has historically been per-model tuning: a specific configuration of a specific model at a specific batch size is fast because someone at AMD made it fast, which is a different asset from a compiler that makes your unseen kernel fast.</p><p><strong>Honest read:</strong> on standard transformer inference, the gap is small enough that the decision is now economic rather than technical. On training, on novel architectures, and on anything requiring bespoke kernels, it is not close.</p></blockquote><h3>SCALE (Spectral Compute)</h3><blockquote><p><strong>Attacks:</strong> layer 1 and layer 3 simultaneously, which makes it the most technically interesting project in this space.</p></blockquote><p>SCALE is a clean-room, Clang and LLVM based drop-in replacement for <code>nvcc</code> that compiles nvcc-dialect CUDA source, including inline PTX, directly to <strong>AMD machine code</strong>. This is not translation, not even middleware: an actual second implementation of the CUDA language. </p><p>The company is small, London based, founded 2018, and funded its own compiler work off consulting. Published benchmarks claim roughly <strong>5.94 times</strong> over HIP on AMD hardware, and as of May 2026 they are claiming speedups over <code>nvcc</code> itself on B300 with CUDA 13.</p><p>The detail I keep turning over is that they joined NVIDIA Inception in June 2026. A company whose product is CUDA portability, partnering with NVIDIA. Which makes sense if you believe NVIDIA&#8217;s interest is in CUDA remaining the standard more than in CUDA remaining exclusive, but it is a strange shape in my view.</p><blockquote><p><strong>Where it runs out of road:</strong> the CUDA-X library surface. There are hundreds of libraries above the language, and a compiler that perfectly compiles the language still leaves you needing cuDNN, cuTENSOR, cuDF, NCCL and the rest. Spectral is working on delegating those to ROCm equivalents, which means SCALE inherits ROCm&#8217;s library ceiling at exactly the point where it matters most.</p></blockquote><h3>ZLUDA</h3><blockquote><p><strong>Attacks:</strong> layer 2. Binary and PTX level translation, no recompilation.</p></blockquote><p>ZLUDA is the project people cite most and understand least. It gained ROCm 7 support in December 2025 and shipped v6 in June 2026. </p><p>That release also announced that commercial funding had ended again and the project is back to being one person&#8217;s weekend work, which is why v6&#8217;s headline features are <strong>32-bit PhysX</strong> and Blender textures rather than anything relevant to inference.</p><blockquote><p><strong>Where it runs out of road:</strong> everywhere, and the reasons are structural rather than technical. Binary translation is a treadmill against a vendor who ships a new virtual ISA extension every generation and a whole new IR every few years. The EULA restriction points directly at this approach. It has now lost funding twice, from two different backers. I would treat the binary translation path as closed, not because the engineering is bad, but because the economics have been tested twice and failed twice.</p></blockquote><h3>Triton</h3><blockquote><p><strong>Attacks:</strong> layer 1 and, at least partially, layer 4.</p></blockquote><p>Triton is the most consequential thing on this list because of where it sits: it is what <code>torch.compile</code> emits, which means it is the default kernel authoring layer for most of the industry whether they know it or not. It runs on NVIDIA, AMD and Intel from the same source.</p><blockquote><p><strong>Where it runs out of road:</strong> two places. On NVIDIA it emits PTX, so it never escapes <code>ptxas</code>; it relocates the frontend, it does not cross the wall. And on the newest hardware it is a long way from peak, because the abstractions do not expose TMEM, CTA-pair MMA and the descriptor-driven async machinery that Blackwell needs. FlashAttention 4 leaving Triton for CuTe DSL is the load-bearing evidence there.</p></blockquote><p>And now NVIDIA is offering to compile it through Tile IR, which resolves the second problem in exchange for reintroducing the first at a higher altitude.</p><h3>Helion, TileLang, Pallas, and the search-based approach</h3><blockquote><p><strong>Attacks:</strong> the labour market, which as argued above is the real moat.</p></blockquote><p><strong>Helion </strong>is a Python DSL from Meta&#8217;s PyTorch compiler team that sits above Triton: you write PyTorch-shaped code with tile loops and it autotunes over hundreds of generated Triton implementations, spending roughly ten minutes of compute per kernel to do it. </p><p>The reported numbers are the interesting part. Around 1.05 times over <code>torch.compile</code> and 1.44 times over hand-written Triton on AMD on average, with individual kernels much higher, and on H100 roughly 1.2 to 1.85 times over Triton.</p><p><strong>Beating hand-written Triton </strong>on average, by searching. That is the anti-moat technology, and it is not a language feature. It is a compute substituting for expertise, and it gets better every time GPUs get cheaper, which is a nice recursion.</p><p>The same logic explains AMD&#8217;s agentic kernel work (<strong>GEAK</strong>), LLM kernel generators, and the growing body of RL-over-schedules research including the SASS-level work mentioned earlier. If your competitor&#8217;s advantage is several hundred irreplaceable engineers, the correct response is not to hire several hundred engineers. It is to make the engineers less necessary.</p><blockquote><p><strong>Where it runs out of road:</strong> search needs a correct, expressive, fast-to-evaluate space to search over, and on Blackwell the space that contains peak performance is only reachable through the vendor&#8217;s layout algebra. You can search brilliantly within Triton&#8217;s expressible set and still be 2.7 times off.</p></blockquote><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!xQrV!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2630a6f8-94c7-4c08-b0a0-6a3d8ca99a29_1634x912.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!xQrV!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2630a6f8-94c7-4c08-b0a0-6a3d8ca99a29_1634x912.png 424w, https://substackcdn.com/image/fetch/$s_!xQrV!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2630a6f8-94c7-4c08-b0a0-6a3d8ca99a29_1634x912.png 848w, https://substackcdn.com/image/fetch/$s_!xQrV!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2630a6f8-94c7-4c08-b0a0-6a3d8ca99a29_1634x912.png 1272w, https://substackcdn.com/image/fetch/$s_!xQrV!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2630a6f8-94c7-4c08-b0a0-6a3d8ca99a29_1634x912.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!xQrV!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2630a6f8-94c7-4c08-b0a0-6a3d8ca99a29_1634x912.png" width="1456" height="813" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2630a6f8-94c7-4c08-b0a0-6a3d8ca99a29_1634x912.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:813,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Figure 8. Reported multiples from vendors and projects, on different hardware, at differen&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 8. Reported multiples from vendors and projects, on different hardware, at differen" title="Figure 8. Reported multiples from vendors and projects, on different hardware, at differen" srcset="https://substackcdn.com/image/fetch/$s_!xQrV!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2630a6f8-94c7-4c08-b0a0-6a3d8ca99a29_1634x912.png 424w, https://substackcdn.com/image/fetch/$s_!xQrV!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2630a6f8-94c7-4c08-b0a0-6a3d8ca99a29_1634x912.png 848w, https://substackcdn.com/image/fetch/$s_!xQrV!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2630a6f8-94c7-4c08-b0a0-6a3d8ca99a29_1634x912.png 1272w, https://substackcdn.com/image/fetch/$s_!xQrV!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2630a6f8-94c7-4c08-b0a0-6a3d8ca99a29_1634x912.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong><span>Figure 8 </span></strong>Reported multiples from vendors and projects, on different hardware, at different shapes, against different baselines. Useful as a map of who is claiming what, not as a comparison.<span>Copy imageDownload PNG</span></figcaption></figure></div><h3>tinygrad</h3><blockquote><p><strong>Attacks:</strong> layers 1, 2 and 3, by refusing the premise.</p></blockquote><p>tinygrad generates its own kernels, can render directly to PTX, and has experimental backends that talk to AMD and NVIDIA hardware without the vendor userspace runtime at all. </p><p>Whatever you think of the framework&#8217;s ambitions, the AMD path is an existence proof that the vendor compiler and userspace stack are not physically required.</p><blockquote><p><strong>Where it runs out of road:</strong> peak tensor core utilization on the newest silicon, and the entire ecosystem surface. It is a demonstration of what is possible for a small team, not a production alternative for a lab training frontier models. Its real contribution is epistemic: it proves the stack is thinner than the marketing implies.</p></blockquote><h3>Mojo and MAX</h3><blockquote><p><strong>Attacks:</strong> layers 1 and 4, with an MLIR-native language designed for kernel authoring across vendors.</p></blockquote><p>Modular&#8217;s own writing on this is refreshingly unromantic: they note that write-once-run-everywhere in GPU land has usually meant write-once-run-slowly-everywhere, that CUTLASS does not attempt portability beyond NVIDIA and is often locked within a generation, and that Triton&#8217;s performance degrades off NVIDIA. </p><p>Their bet is that a properly designed <strong>compile-time metaprogramming language</strong> can express hardware-specific structure without forking the source.</p><blockquote><p><strong>Where it runs out of road:</strong> it is a proprietary language asking developers to adopt it on faith, competing against a Python DSL layer that PyTorch already emits by default. The technology is pretty good. The distribution problem is brutal.</p></blockquote><h3>The vertical integrators: TPU, Trainium, and friends</h3><blockquote><p><strong>Attacks:</strong> the entire column, by not participating in the compatibility game at all.</p></blockquote><p>Google&#8217;s TPU plus XLA is the only stack that has demonstrably trained frontier models at scale outside NVIDIA, sustained, for years. It did not do that by cloning CUDA. </p><p>It did it by owning the chip, the interconnect, the compiler, one framework, and one enormous first-party customer whose workloads shape the hardware roadmap. <strong>Trainium plus NKI </strong>is the same play at an earlier stage.</p><p>The lesson is uncomfortable for everyone selling portability: the proven counter to vertical integration is vertical integration. Compatibility layers have a zero-for-many record. Owning the column has a one-for-one record.</p><blockquote><p><strong>Where it runs out of road:</strong> you cannot buy this. It requires a decade, a captive workload, and a willingness to eat several generations of worse hardware.</p></blockquote><h3>The ones I am not covering properly, and why</h3><p>Three more categories exist and deserve at least an honest pointer rather than silence.</p><p><strong>Huawei&#8217;s CANN</strong> with the Ascend NPUs is the most complete non-Western vertical stack I&#8217;ve ever seen, with its own kernel language, its own graph compiler, and a captive domestic customer base that gives it the one thing compatibility layers never have, which is a workload that shapes the roadmap. </p><p>I do not have enough independent benchmark data to say anything useful about where it sits on peak utilization, and I would rather say that than guess randomly. The same caveat applies to<strong> Moore Threads</strong> and Biren, both of which ship CUDA-adjacent programming models.</p><p>On the standards side, chipStar compiles HIP and CUDA to SPIR-V for OpenCL and Level Zero targets, and Intel&#8217;s SYCLomatic migrates CUDA source to SYCL with a reported majority of code converting automatically and a manual remainder. Both are layer 1 plays with the layer 4 problem intact.</p><p>And there is a fourth category I have deliberately excluded from the comparison table because it is not a competitor at all: NVIDIA&#8217;s own Python surface. <code>cuda.core</code>, <strong>Numba</strong>, Warp, <code>nvmath</code>, and now cuTile are a coordinated effort to make sure that when kernel authoring moves fully into Python, which it is doing, it moves into NVIDIA&#8217;s Python rather than someone else&#8217;s.</p><h3>The cautionary tale: OpenCL</h3><p>Worth remembering, because the failure mode repeats. OpenCL standardized the API and the language and left the libraries, the tuning, and the <strong>last-mile compiler</strong> quality to each vendor. </p><p>Portable source, unportable performance, no cuBLAS equivalent, no tensor core story, committee cadence against NVIDIA&#8217;s annual co-design cadence. </p><p>SYCL is a better designed successor with the same structural problem: it is a standard for the layer that was never the moat.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>What replacing the kernel layer actually costs</h2><p>Let me put a rough number on the library moat, since that is the part people assume is unbuyable.</p><p>Take LLM inference specifically, not the whole CUDA-X surface. The kernels that matter for a modern decoder-only model at serving time are: <strong>dense GEMM </strong>in several shapes, grouped or batched GEMM for MoE, prefill attention, decode attention with paged KV, an MLA variant if you serve DeepSeek-shaped models, quantize and dequantize paths for FP8 and FP4, <strong>RMSNorm</strong>, RoPE, fused SwiGLU, sampling, KV cache gather and copy, speculative verification, and the collectives. </p><p>Call it 30 to 50 kernels that appear in profiles, of which eight or so carry most of the time.</p><p>Assume a fully loaded peak-quality kernel engineer at somewhere between $350,000 and $600,000, call it $450,000 to be some kind of &#8220;average&#8221;. Assume a top-tier attention or<strong> GEMM kernel</strong> for a new architecture takes three to nine engineer-months including autotuning, numerics validation and the long tail of shape coverage, and the secondary kernels take one to two.</p><ul><li><p><em>Eight critical kernels at six months each: roughly $1.8M</em></p></li><li><p><em>Forty secondary kernels at 1.5 months each: roughly $2.2M</em></p></li><li><p><em>Test infrastructure, numerics validation, CI across shapes and dtypes: call it 50 percent overhead</em></p></li></ul><p>That lands somewhere around $6M to $15M per hardware generation to build a competitive <em>inference</em> kernel layer from scratch. For a company selling accelerators, that is a rounding error. It is roughly one percent of a single generation&#8217;s mask set and tape-out cost.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!MXjK!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2a4139cb-ecc9-4d45-a7a6-7eca5dfc9eea_1634x912.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!MXjK!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2a4139cb-ecc9-4d45-a7a6-7eca5dfc9eea_1634x912.png 424w, https://substackcdn.com/image/fetch/$s_!MXjK!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2a4139cb-ecc9-4d45-a7a6-7eca5dfc9eea_1634x912.png 848w, https://substackcdn.com/image/fetch/$s_!MXjK!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2a4139cb-ecc9-4d45-a7a6-7eca5dfc9eea_1634x912.png 1272w, https://substackcdn.com/image/fetch/$s_!MXjK!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2a4139cb-ecc9-4d45-a7a6-7eca5dfc9eea_1634x912.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!MXjK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2a4139cb-ecc9-4d45-a7a6-7eca5dfc9eea_1634x912.png" width="1456" height="813" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2a4139cb-ecc9-4d45-a7a6-7eca5dfc9eea_1634x912.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:813,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Figure 6. A stated model, not a measurement. The inference kernel surface is small enough &quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 6. A stated model, not a measurement. The inference kernel surface is small enough " title="Figure 6. A stated model, not a measurement. The inference kernel surface is small enough " srcset="https://substackcdn.com/image/fetch/$s_!MXjK!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2a4139cb-ecc9-4d45-a7a6-7eca5dfc9eea_1634x912.png 424w, https://substackcdn.com/image/fetch/$s_!MXjK!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2a4139cb-ecc9-4d45-a7a6-7eca5dfc9eea_1634x912.png 848w, https://substackcdn.com/image/fetch/$s_!MXjK!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2a4139cb-ecc9-4d45-a7a6-7eca5dfc9eea_1634x912.png 1272w, https://substackcdn.com/image/fetch/$s_!MXjK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2a4139cb-ecc9-4d45-a7a6-7eca5dfc9eea_1634x912.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong><span>Figure 9 </span></strong>A stated model, not a measurement. The inference kernel surface is small enough for any accelerator vendor to buy outright, which is exactly why AMD is close on inference and not close on training.<span>Copy imageDownload PNG</span></figcaption></figure></div><p>Which explains something that otherwise looks strange: AMD is actually close on inference and not close on training. The inference kernel surface is small enough to buy. </p><p>The training surface, plus the full CUDA-X library set, plus the numerics validation across a decade of recipes, is somewhere between ten and fifty times that, and the <strong>co-design loop</strong> that makes the kernels good on day one instead of month nine is not purchasable at any price.</p><p>Treat these figures as a stated model with stated assumptions, not as a measurement. The point is the ratio, not the absolute value: the part everyone talks about is the cheap part.</p><h2>Where I might be wrong</h2><p>Four ways this argument could be badly calibrated, in descending order of how much they would cost me.</p><p><strong>The backend gap may be small.</strong> My case for layer 3 rests partly on the <code>maxas</code> era, which was Maxwell, and partly on recent reinforcement learning work over SASS schedules that finds wins by reordering <code>ptxas</code> output. If the residual there is three to five percent on typical kernels rather than the double digits the Maxwell results implied, then <code>ptxas</code> is close enough to optimal that the backend is not the moat at all, the libraries are, and I have spent a lot of words on the wrong floor. </p><p>I have not seen a rigorous modern measurement of that residual across a representative kernel set, and I would very much like to.</p><p><strong>The labour market framing may be a category error.</strong> I argue the scarce resource is people who can read control bits, and therefore that search and autotuning erode it. </p><p>But the advantage might be organisational rather than headcount: eighteen months of pre-silicon access, a hardware team down the hall, and the ability to change the chip when the kernel is awkward. If that is the real mechanism, then no amount of autotuning compute closes it, because the loop being run is design, not optimization.</p><p><strong>Tile IR may be commoditization rather than capture.</strong> I read the open dialect, the published bytecode format and the conformance suite as a way of owning the layer Triton was about to own. The more generous reading is that an MLIR dialect with a stable binary encoding and a test suite is precisely the artifact another vendor needs in order to target the same source, and NVIDIA has just handed it over. Both readings are consistent with the facts as of today. Which one is right will be visible in whether anyone outside NVIDIA ships a Tile IR backend.</p><p><strong>And the boring one.</strong> I do not have a B200. Everything in this piece at the Blackwell layer comes from documentation, from other people&#8217;s benchmarks, and from reading disassembly conventions rather than fresh disassembly. The control field layout is reverse engineered by researchers, not published. </p><p>Where I have marked things Tier B or C, that is what the marking means, and I would treat the Hopper and Blackwell specifics as directionally right rather than precisely right until someone with the hardware checks them.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Five dated calls</h2><p>Predictions with dates, so this piece can be scored rather than admired.</p><p><strong>1. No third party ships a production-quality SASS assembler for Hopper or Blackwell before 2030.</strong> Production quality meaning it passes a public correctness suite across a non-trivial kernel set on more than one SKU. Falsified by any project that does.</p><p><strong>2. NVIDIA never open sources the Tile IR to cubin lowering.</strong> The dialect, the bytecode, the spec and the conformance suite stay open. The optimizing backend does not. Falsified trivially and publicly if wrong.</p><p><strong>3. Before the end of 2027, an automated kernel generator beats a vendor library on a headline GEMM or attention shape on non-NVIDIA hardware, with reproducible numbers.</strong> The search-versus-labour thesis stands or falls on this one.</p><p><strong>4. AMD reaches within fifteen percent of NVIDIA on tokens per second per dollar for standard transformer inference on a like-for-like generation by end of 2027, and does not close the equivalent gap on training in the same window.</strong> The asymmetry is the prediction, not the number.</p><p><strong>5. At least one more CUDA compatibility project loses its funding before the end of 2027.</strong> ZLUDA has now done it twice. The economics of chasing a virtual ISA that moves every eighteen months have been tested and they do not work, and I expect the next attempt to discover this independently.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/p/how-the-nvidia-compiler-moat-actually?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/p/how-the-nvidia-compiler-moat-actually?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><div><hr></div><h2>What would actually work (for real)</h2><p>I am sorry to say it, but I do not think the moat gets breached. I think it gets routed around, in specific places, for specific workloads, and here is where I would put effort if that were my job.</p><p><strong>Own the tile layer, and make it a spec with teeth.</strong> Not a language, not a framework: a specification with a versioned binary encoding and a conformance test suite that vendors must pass. NVIDIA just told you this layer is the strategic one by shipping exactly that artifact. The mistake would be to let Tile IR become the only one.</p><p><strong>Make numerics a specification too.</strong> This is the gap nobody is filling. A conformance suite that pins reduction order tolerances, accumulation types, scaling granularity for narrow formats, and determinism guarantees, with reference outputs. Right now &#8220;<em>matches cuBLAS</em>&#8221; is folklore transmitted through issue threads. Turning it into a testable contract would remove a real and unpriced switching cost.</p><p><strong>Attack at inference, not at parity.</strong> The kernel surface is small, the workloads are known, the customers are cost-sensitive, and memory capacity per package is a lever NVIDIA does not always win. Trying to be CUDA-complete is the losing strategy; being excellent at forty kernels is the winning one.</p><p><strong>Spend compute instead of headcount.</strong> Every dollar into autotuning, superoptimization and learned scheduling attacks the actual scarce resource. Helion beating hand-written Triton on average is the proof of concept. This is the only lever in the list that gets cheaper over time.</p><p><strong>And keep the compiler argument in proportion.</strong> The binding constraint on who gets to serve AI workloads in 2026 is HBM supply, advanced packaging capacity, rack-scale networking and power. A perfect compiler on a chip you cannot buy, in a rack you cannot cool, connected by a fabric that stalls at 72 GPUs, wins nothing. The compiler moat is real and it is roughly the fourth most important moat NVIDIA has.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Check it yourself</h2><p>Nothing above requires trusting me. The relevant artifacts are all inspectable with the toolkit you already have installed.</p><p><strong>See the driver script decompose:</strong></p><pre><code><code>nvcc --dryrun -arch=sm_90 kernel.cu
</code></code></pre><p>Every stage, in order, with the real command lines. <code>cicc</code>, <code>ptxas</code>, <code>fatbinary</code>, <code>nvlink</code>.</p><p><strong>Watch the virtual ISA:</strong></p><pre><code><code>nvcc -arch=sm_90 -ptx kernel.cu -o kernel.ptx
</code></code></pre><p>Note the unbounded <code>%r</code> virtual registers and the total absence of scheduling information.</p><p><strong>Watch the last mile do the work:</strong></p><pre><code><code>ptxas -arch=sm_90 -v kernel.ptx -o kernel.cubin
</code></code></pre><p>The <code>-v</code> output tells you registers used, spill stores, spill loads, shared memory. That register count is a policy decision <code>ptxas</code> made on your behalf, and it sets your occupancy.</p><p><strong>Look at the control bits:</strong></p><pre><code><code>nvdisasm -c -g kernel.cubin
cuobjdump --dump-sass kernel.cubin
</code></code></pre><p>The bracketed fields next to each instruction are the scheduling payload. Compare the same kernel compiled with and without <code>-maxrregcount</code> and watch the stall counts and barrier usage change.</p><p><strong>Confirm the frontend is a commodity:</strong></p><pre><code><code>clang++ -x cuda --cuda-gpu-arch=sm_90 --cuda-device-only -S kernel.cu -o clang.ptx
</code></code></pre><p>Diff that against <code>nvcc</code>&#8216;s PTX. They differ. Then push both through <code>ptxas</code> and compare the SASS. That comparison is the whole argument of this piece in two commands: two independent frontends, one backend, and the backend is where the differences stop mattering.</p><p><strong>Find the boundary:</strong> try to turn a modified SASS listing back into a <code>cubin</code> using only NVIDIA tooling. There is no such command. That absence is the moat.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Confidence dossier</h2><p><strong><span>Tier A</span></strong>, directly verifiable from primary sources or first-party documentation</p><ul><li><p><code>nvcc</code> is a driver invoking <code>cicc</code>, <code>ptxas</code>, <code>fatbinary</code> and <code>nvlink</code>; observable via <code>--dryrun</code></p></li><li><p>NVVM IR is a documented LLVM IR subset with a public <code>libNVVM</code> API; Clang&#8217;s NVPTX backend is upstream</p></li><li><p>PTX is a virtual ISA; the driver JIT-compiles it, providing forward compatibility</p></li><li><p>There is no NVIDIA-provided SASS assembler; <code>nvdisasm</code> and <code>cuobjdump</code> disassemble only</p></li><li><p>Blackwell deprecated <code>wgmma</code> and introduced <code>tcgen05</code>, Tensor Memory, single-thread MMA issue and <code>cta_group::2</code></p></li><li><p>CUDA Tile IR shipped in CUDA 13.1; the MLIR dialect, bytecode format and conformance suite are open source; the bytecode is specified as stable and versioned</p></li><li><p>A Triton to Tile IR backend exists, authored by NVIDIA</p></li><li><p>The CUDA EULA restricts translating SDK-generated output artifacts to non-NVIDIA platforms</p></li><li><p><code>wgmma</code> requires the <code>sm_90a</code> architecture-specific target and <code>tcgen05</code> requires <code>sm_100a</code> or <code>sm_103a</code>; code for <code>a</code> targets is not forward compatible</p></li><li><p>A <code>cubin</code> is an ELF64 object with per-kernel <code>.text</code>, <code>.nv.constant0</code> and <code>.nv.info</code> sections</p></li><li><p>Blackwell Tensor Memory is 256 KB per SM, 128 lanes by 512 columns of 32 bits, allocated in power-of-two column counts</p></li><li><p>MXFP8 uses one E8M0 scale per 32 elements; NVFP4 uses an E4M3 scale per 16 elements plus a tensor-level FP32 scale</p></li></ul><p><strong><span>Tier B</span></strong>, well-supported by multiple independent secondary sources or peer-reviewed measurement</p><ul><li><p>The Volta-and-later control field layout: reuse, wait mask, read and write barrier indices, yield, stall count, packed into the instruction word</p></li><li><p>Fixed-latency instructions lack hardware RAW interlocks; incorrect stall counts produce incorrect results</p></li><li><p>Instruction encoding widened from 64 to 128 bits at Volta</p></li><li><p>H100 register file of 65,536 32-bit registers per SM, 64 maximum resident warps, eight-register allocation granularity, and the occupancy arithmetic that follows from them</p></li><li><p>FlashAttention 4 in CuTe DSL outperforming cuDNN 9.13 and Triton by the margins cited</p></li><li><p>ROCm 7.2 library and AITER performance improvements on MI300 through MI355X</p></li><li><p>Helion&#8217;s reported speedups over Triton and <code>torch.compile</code></p></li><li><p>NCCL&#8217;s channel-per-thread-block structure, its protocol selection, and the fact that collectives consume SMs that are then unavailable for compute</p></li><li><p>In-network reduction in NVSwitch and InfiniBand switch silicon</p></li><li><p>SCALE&#8217;s benchmark claims against HIP and <code>nvcc</code></p></li></ul><p><strong><span>Tier C</span></strong>, single-source, historical, or reported rather than independently verified</p><ul><li><p><code>maxas</code> SGEMM efficiency figures on GM204 versus contemporaneous cuBLAS</p></li><li><p>The annotated SASS fragment, which is reconstructed in the documented format rather than captured from a specific compilation</p></li><li><p>DeepSeek&#8217;s specific SM partitioning and custom PTX details, as described in the V3 paper and subsequent reporting</p></li><li><p>Spectral Compute&#8217;s NVIDIA Inception membership and its strategic reading</p></li><li><p>cuTile&#8217;s competitiveness with cuBLAS and cuDNN at large shapes, which is a first-party claim awaiting third-party benchmarks</p></li></ul><p><strong><span>Tier D</span></strong>, my model or my opinion, flagged as such</p><ul><li><p>The $6M to $15M per-generation estimate for an inference kernel layer, and the 10x to 50x multiplier for the full surface</p></li><li><p>The claim that layer 3 is the load-bearing wall rather than layer 1</p></li><li><p>The reading of Tile IR as a strategic response to Triton&#8217;s position in PyTorch</p></li><li><p>The assertion that the moat is best modelled as a labour market</p></li><li><p>The ranking of the compiler moat as roughly fourth behind supply, packaging and networking</p></li><li><p>All five dated calls, which are forecasts and should be read as such</p></li><li><p>The four self-criticisms, which are my own estimate of where this argument is weakest</p></li></ul><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Bibliography</h2><p><strong>Primary documentation</strong></p><ul><li><p>NVIDIA, <em>Parallel Thread Execution ISA</em>. https://docs.nvidia.com/cuda/parallel-thread-execution/</p></li><li><p>NVIDIA, <em>CUDA Binary Utilities</em> (<code>cuobjdump</code>, <code>nvdisasm</code>, SASS instruction set tables). https://docs.nvidia.com/cuda/cuda-binary-utilities/</p></li><li><p>NVIDIA, <em>NVVM IR Specification</em>. https://docs.nvidia.com/cuda/nvvm-ir-spec/</p></li><li><p>NVIDIA, <em>CUDA Compiler Driver NVCC</em>. https://docs.nvidia.com/cuda/cuda-compiler-driver-nvcc/</p></li><li><p>NVIDIA, <em>License Agreement for NVIDIA Software Development Kits</em>. https://docs.nvidia.com/cuda/eula/index.html</p></li><li><p>NVIDIA, <em>CUDA Tile</em>. https://developer.nvidia.com/cuda/tile</p></li><li><p>NVIDIA, <em>Tile IR: Binary Format</em>. https://docs.nvidia.com/cuda/tile-ir/latest/sections/bytecode.html</p></li><li><p>NVIDIA, <em>cuTile Python documentation</em>. https://docs.nvidia.com/cuda/cutile-python/</p></li><li><p>NVIDIA, <em>CUTLASS documentation: tcgen05 MMA Programming Guide</em>. https://docs.nvidia.com/cutlass/latest/media/docs/pythonDSL/mma_docs/tcgen05_programming.html</p></li><li><p>NVIDIA, <em>CUTLASS: Blackwell SM100 functionality</em>. https://docs.nvidia.com/cutlass/4.2.1/media/docs/cpp/blackwell_functionality.html</p></li><li><p>AMD, <em>ROCm documentation: vLLM inference optimization</em>. https://rocm.docs.amd.com/en/latest/how-to/rocm-for-ai/inference-optimization/vllm-optimization.html</p></li></ul><p><strong>Microarchitecture and reverse engineering</strong></p><ul><li><p>Jia, Maggioni, Staiger, Scarpazza, <em>Dissecting the NVIDIA Volta GPU Architecture via Microbenchmarking</em>, arXiv:1804.06826. https://arxiv.org/pdf/1804.06826</p></li><li><p>Huerta, Abaie et al., <em>Analyzing Modern NVIDIA GPU cores</em>, arXiv:2503.20481. https://arxiv.org/pdf/2503.20481</p></li><li><p><em>Dissecting and Modeling the Architecture of Modern GPU Cores</em>, MICRO 58, 2025. https://dl.acm.org/doi/10.1145/3725843.3756041</p></li><li><p>Luo et al., <em>Benchmarking and Dissecting the Nvidia Hopper GPU Architecture</em>, arXiv:2402.13499</p></li><li><p><em>CuAsmRL: Optimizing GPU SASS Schedules via Deep Reinforcement Learning</em>, arXiv:2501.08071. https://arxiv.org/html/2501.08071v1</p></li><li><p><em>GPA: A GPU Performance Advisor Based on Instruction Sampling</em>, arXiv:2009.04061</p></li><li><p>Finn, <em>What happens when you run a CUDA kernel</em>. https://fergusfinn.com/blog/what-happens-when-you-run-a-gpu-kernel/</p></li></ul><p><strong>Assembler and toolchain projects</strong></p><ul><li><p>Gray, <em>maxas</em>, Maxwell assembler and control code notes. https://github.com/NervanaSystems/maxas/wiki/Control-Codes</p></li><li><p><em>CuAssembler</em>. https://github.com/cloudcores/CuAssembler</p></li><li><p><em>openptxas</em>, gpuocelot. https://github.com/gpuocelot/openptxas</p></li><li><p>NVlabs, <em>SASSI: Flexible GPGPU instrumentation</em>, distributed as a closed-source <code>ptxas</code> fork. https://github.com/NVlabs/SASSI</p></li><li><p>NVIDIA, <em>cuda-tile</em>. https://github.com/NVIDIA/cuda-tile</p></li></ul><p><strong>Kernel authoring layers</strong></p><ul><li><p>Colfax Research, <em>CUTLASS Tutorial: Writing GEMM Kernels Using Tensor Memory for NVIDIA Blackwell GPUs</em>. https://research.colfax-intl.com/cutlass-tutorial-writing-gemm-kernels-using-tensor-memory-for-nvidia-blackwell-gpus/</p></li><li><p>Colfax Research, <em>CUTLASS Tutorial: Hardware-supported Block-scaling with NVIDIA Blackwell GPUs</em>. https://research.colfax-intl.com/cutlass-tutorial-hardware-supported-block-scaling-with-nvidia-blackwell-gpus/</p></li><li><p>SemiAnalysis, <em>Dissecting Nvidia Blackwell: Tensor Cores, PTX Instructions, SASS, Floorsweep, Yield</em>. https://newsletter.semianalysis.com/p/dissecting-nvidia-blackwell-tensor</p></li><li><p>PyTorch, <em>Helion: A High-Level DSL for Performant and Portable ML Kernels</em>. https://pytorch.org/blog/helion/</p></li><li><p>pytorch/helion repository. https://github.com/pytorch/helion</p></li><li><p>PyTorch RFC, <em>Inductor cuTile Backend</em>, issue 175311. https://github.com/pytorch/pytorch/issues/175311</p></li><li><p>NVIDIA Developer Blog, <em>Advancing GPU Programming with the CUDA Tile IR Backend for OpenAI Triton</em>. https://developer.nvidia.com/blog/?p=112268</p></li><li><p>Red Hat Emerging Technologies, <em>From hand-tuned to generated: a reproducible Triton GPU kernel benchmark across different vendors</em>. https://next.redhat.com/2026/02/12/from-hand-tuned-to-generated-a-reproducible-triton-gpu-kernel-benchmark-across-different-vendors/</p></li><li><p>Modular, <em>Structured Mojo Kernels Part 4: Portability and the Road Ahead</em>. https://www.modular.com/blog/structured-mojo-kernels-part-4-portability-and-the-road-ahead</p></li><li><p><em>Mojo: MLIR-Based Performance-Portable HPC Science Kernels on GPUs for the Python Ecosystem</em>, arXiv:2509.21039</p></li></ul><p><strong>Alternatives and portability</strong></p><ul><li><p>Spectral Compute, <em>SCALE</em>. https://scale-lang.com/ and https://github.com/spectral-compute/scale-docs</p></li><li><p>HPCwire, <em>Spectral Compute Aims to Set CUDA Free. Will It Succeed?</em>, July 2026. https://www.hpcwire.com/2026/07/09/spectral-compute-aims-to-set-cuda-free-will-it-succeed/</p></li><li><p><em>SCALE: Ahead-Of-Time Compilation of CUDA for AMD GPUs</em>, Middleware 2024 demo track. https://dl.acm.org/doi/10.1145/3704440.3704782</p></li><li><p>ZLUDA, <em>Update Q1 and Q2 2026: back to the roots</em>. https://vosen.github.io/ZLUDA/blog/zluda-update-q1q2-2026/</p></li><li><p>Phoronix, <em>ZLUDA For CUDA On Non-NVIDIA GPUs Enables AMD ROCm 7 Support</em>. https://www.phoronix.com/news/ZLUDA-ROCm-7</p></li><li><p>Tom&#8217;s Hardware, <em>Nvidia bans using translation layers for CUDA software</em>, March 2024. https://www.tomshardware.com/pc-components/gpus/nvidia-bans-using-translation-layers-for-cuda-software-to-run-on-other-chips-new-restriction-apparently-targets-zluda-and-some-chinese-gpu-makers</p></li><li><p>vLLM Blog, <em>Beyond Porting: How vLLM Orchestrates High-Performance Inference on AMD ROCm</em>. https://vllm.ai/blog/2026-02-27-rocm-attention-backend</p></li><li><p>AMD, <em>Speed is the Moat: Inference Performance on AMD GPUs</em>. https://www.amd.com/en/developer/resources/technical-articles/2026/inference-performance-on-amd-gpus.html</p></li><li><p>tinygrad repository and release notes. https://github.com/tinygrad/tinygrad</p></li><li><p>Pall et al., <em>GROMACS on AMD GPU-Based HPC Platforms: Using SYCL for Performance and Portability</em>, arXiv:2405.01420</p></li></ul><p><strong>Model and systems work referenced</strong></p><ul><li><p>DeepSeek-AI, <em>DeepSeek-V3 Technical Report</em>, arXiv:2412.19437</p></li><li><p>Tom&#8217;s Hardware, <em>DeepSeek&#8217;s AI breakthrough bypasses industry-standard CUDA for some functions</em>. https://www.tomshardware.com/tech-industry/artificial-intelligence/deepseeks-ai-breakthrough-bypasses-industry-standard-cuda-uses-assembly-like-ptx-programming-instead</p></li><li><p>Dao et al., FlashAttention series, most recently FlashAttention 4 on Blackwell in CuTe DSL</p></li><li><p>Bauer et al., <em>Singe: Leveraging Warp Specialization for High Performance on GPUs</em>, PPoPP 2014</p></li></ul><p><em>Corrections and disagreements are welcome and get published. If you have measured </em><code>ptxas</code><em> output against hand-scheduled SASS on Hopper or Blackwell, I would like to see the numbers.</em></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share The Software Frontier&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share The Software Frontier</span></a></p><p></p>]]></content:encoded></item><item><title><![CDATA[Kimi K3: 2.8 Trillion Parameters, Four Bits at a Time]]></title><description><![CDATA[Moonshot just announced the largest open-weight model ever built. The parameter count is the headline. The serving stack is the story.]]></description><link>https://www.thesoftwarefrontier.com/p/kimi-k3-28-trillion-parameters-four</link><guid isPermaLink="false">https://www.thesoftwarefrontier.com/p/kimi-k3-28-trillion-parameters-four</guid><dc:creator><![CDATA[Lorenzo Bradanini]]></dc:creator><pubDate>Wed, 22 Jul 2026 09:53:06 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!B2ac!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F73922128-9535-44e5-b887-ad4a9f5ccea0_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!B2ac!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F73922128-9535-44e5-b887-ad4a9f5ccea0_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!B2ac!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F73922128-9535-44e5-b887-ad4a9f5ccea0_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!B2ac!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F73922128-9535-44e5-b887-ad4a9f5ccea0_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!B2ac!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F73922128-9535-44e5-b887-ad4a9f5ccea0_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!B2ac!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F73922128-9535-44e5-b887-ad4a9f5ccea0_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!B2ac!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F73922128-9535-44e5-b887-ad4a9f5ccea0_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/73922128-9535-44e5-b887-ad4a9f5ccea0_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2860025,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/206014241?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F73922128-9535-44e5-b887-ad4a9f5ccea0_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!B2ac!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F73922128-9535-44e5-b887-ad4a9f5ccea0_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!B2ac!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F73922128-9535-44e5-b887-ad4a9f5ccea0_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!B2ac!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F73922128-9535-44e5-b887-ad4a9f5ccea0_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!B2ac!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F73922128-9535-44e5-b887-ad4a9f5ccea0_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>Intro</h2><p>Moonshot AI released <strong>Kimi K3 </strong>precisely on July 16. The headline number is 2.8 trillion parameters, which makes it the largest open-weight model ever announced. </p><p>We spent the launch week reading almost everything published about it, and the more we read, the less the parameter count felt like the point. The point is that a <strong>2.8T model</strong> is being served today, at Claude Sonnet prices, with a flat rate across a one million token context window. </p><p>The reasons that is possible sit exactly at the layer this newsletter cares about: attention design, expert routing, <strong>quantization format</strong>, and cache economics.</p><p>A caveat before the numbers. The technical report is not out yet. Moonshot&#8217;s launch blog states that architecture, training, and <strong>evaluation details</strong> will arrive alongside the report, and several figures below come from community analysis of the launch documentation rather than from a paper. </p><p>Where we compute something ourselves, the assumptions are stated inline and marked as ours, and there is a confidence appendix at the end. </p><p>This is a <strong>launch-week analysis</strong>, not a postmortem.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>What shipped</h2><p>The facts first, because launch weeks always tend to generate a lot of useless fog. The model went live on July 16 through kimi.com, the Kimi apps, Kimi Work, Kimi Code, and the <strong>Moonshot API</strong> at api.moonshot.ai under the <strong>model ID kimi-k3. </strong></p><p>It is also listed on <strong>OpenRouter</strong> at the same rates as the direct API. What did not ship is the checkpoint: Moonshot says the full weights land by July 27, and until the files appear on Hugging Face, K3 is a hosted model you can call but not download. </p><p>Verdent&#8217;s <strong>launch guide </strong>reports the weights are coming under a modified MIT license, which would basically match the K2 line, but the license text is unconfirmed until the release itself.</p><p>The specification, per Moonshot: 2.8 trillion total parameters, a mixture of experts <strong>activating 16 of 896 experts per token</strong>, native vision input rather than a bolted-on adapter, and a context window of 1,048,576 tokens. The model reasons before every answer. At launch, reasoning effort is locked to max, with lower effort modes promised in later updates. That detail matters more than it sounds, and we will get to it in the pricing section.</p><p>The timing is not subtle either. VentureBeat notes the release landed just a few days before the <strong>World Artificial Intelligence Conference</strong> in Shanghai, and frames it as a comeback move for a company whose position had eroded badly during DeepSeek&#8217;s rise. </p><p>Moonshot is backed by <strong>Alibaba</strong>, which also builds the <strong>Qwen line</strong>, so the Chinese open-model field is now crowded with players who share an investor and compete anyway. Moonshot&#8217;s own blog claims that for nine of the past twelve months, Kimi models have held the frontier of open-source scale. </p><p>Whatever you think of that framing, the release cadence behind it is real: K2 in July 2025 at one trillion total parameters, K2 Thinking in November with <strong>quantization-aware INT4</strong>, and now K3 at nearly three times the K2 total.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/p/kimi-k3-28-trillion-parameters-four?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/p/kimi-k3-28-trillion-parameters-four?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><div><hr></div><h2>The shape of the model</h2><p>Three important architectural pieces define K3: the <strong>sparsity of the expert layer</strong>, the hybrid attention stack, and a modified residual path. Each one is a serving decision as much as a modeling decision.</p><p>My advice is to start with sparsity, because the trend line is the clearest signal of where frontier MoE design is going. K2 activated 8 of 384 experts, roughly 32B active parameters out of 1T total, an activation ratio around 3.1 percent. K3 activates 16 of 896. </p><p>Moonshot has not published the active parameter count; the community model card analysis on Hugging Face estimates <strong>roughly 50B active equivalent</strong>, which against 2.8T total is circa 1.8 percent. Total capacity nearly tripled while active compute grew by maybe half. </p><p>That is the whole game: parameters are cheap to store and expensive to move, so you scale what sits in memory and hold nearly flat what has to cross the datapath every token. </p><p><strong>DeepSeek</strong> formalized this direction at <strong>671B total </strong>and <strong>5.5 percent active</strong>; Moonshot is now running it harder than anyone with an open checkpoint. </p><p>The lineage is worth one sentence: the sparse expert layer goes back to Shazeer&#8217;s 2017 outrageously large networks paper, was made trainable at datacenter scale by <strong>GShard </strong>and <strong>Switch</strong>, and was pushed toward fine-grained experts with a shared trunk by the DeepSeek line; K3 extends that same axis to 896, and the references at the end walk the chain to the primary sources.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!QV-V!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ec1f9f7-2bc0-4286-a8c1-9616aea80ab6_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!QV-V!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ec1f9f7-2bc0-4286-a8c1-9616aea80ab6_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!QV-V!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ec1f9f7-2bc0-4286-a8c1-9616aea80ab6_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!QV-V!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ec1f9f7-2bc0-4286-a8c1-9616aea80ab6_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!QV-V!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ec1f9f7-2bc0-4286-a8c1-9616aea80ab6_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!QV-V!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ec1f9f7-2bc0-4286-a8c1-9616aea80ab6_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6ec1f9f7-2bc0-4286-a8c1-9616aea80ab6_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;The sparsity frontier: open flagship MoE models&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="The sparsity frontier: open flagship MoE models" title="The sparsity frontier: open flagship MoE models" srcset="https://substackcdn.com/image/fetch/$s_!QV-V!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ec1f9f7-2bc0-4286-a8c1-9616aea80ab6_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!QV-V!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ec1f9f7-2bc0-4286-a8c1-9616aea80ab6_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!QV-V!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ec1f9f7-2bc0-4286-a8c1-9616aea80ab6_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!QV-V!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ec1f9f7-2bc0-4286-a8c1-9616aea80ab6_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 1. The sparsity frontier: open flagship MoE models.</figcaption></figure></div><p>The expert layer runs on what Moonshot calls <strong>Stable LatentMoE</strong>. According to the launch documentation, routing happens in a latent space rather than directly on token representations, load is managed by a mechanism called <strong>Quantile Balancing,</strong> and overflow tokens are soft-dropped rather than hard-bounced. </p><p>All three claims obviously await the technical report, but the intent is legible: at 896 experts, classical <em>routing instability</em> and load skew are the failure modes that kill both training and serving, and everything named here is aimed at them. </p><p>If you have ever watched a MoE deployment where two hot experts saturate their devices while the other <strong>few hundred idle</strong>, you know why a lab would lead with the word Stable.</p><p>The attention stack is a hybrid, and it has a paper trail. K3 is built on <strong>Kimi Delta Attention</strong>, KDA, which Moonshot introduced with the Kimi Linear report in late 2025. </p><p>KDA is essentially a gated refinement of DeltaNet: a delta-rule state update with fine-grained per-channel decay instead of a scalar forget gate, implemented with chunked parallel kernels so the recurrence trains at practical speed. </p><p>In the Kimi Linear configuration these layers were interleaved with full attention at a <em>three to one ratio</em>, the global layers used multi-head latent attention, and the paper reported on the order of a <strong>75 percent KV cache reduction</strong> and up to a reported 6x decode throughput gain at million-token lengths against a full-attention baseline. </p><p>K3&#8217;s global layers use Gated MLA, a gated variant of the same latent attention. Moonshot has not confirmed K3&#8217;s layer ratio, but the lineage tells you the design logic: KDA layers carry a fixed-size recurrent state that costs the same at position one million as at position one hundred, and only the global layers accumulate <strong>position-indexed KV</strong>, compressed into latents at that. </p><p>The million-token window is not priced flat out of generosity. It is priced flat because most of the stack does not pay for length.</p><p>For readers who want the update rule, KDA in the Kimi Linear formulation maintains a matrix-valued state S, rewritten each step as</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;plaintext&quot;,&quot;nodeId&quot;:&quot;0007b882-8309-414c-9f06-2ae450d85604&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-plaintext">S_t = (I - beta_t k_t k_t^T) Diag(alpha_t) S_{t-1} + beta_t k_t v_t^T</code></pre></div><p>read right to left: <strong>decay the previous state channel</strong> by channel through Diag(alpha_t), erase the stale association along the incoming key direction, then write the new key-value pair at strength beta_t. </p><p>The per-channel diagonal gate is the refinement over <strong>Gated DeltaNet&#8217;s scalar decay</strong>, restoring some of the selectivity full attention gets for free, and the chunked kernel that makes the recurrence trainable at scale, a specialized diagonal-plus-low-rank formulation, is the same computational shape the prefill discussion below returns to.</p><p>The third piece is <strong>Attention Residuals</strong>, AttnRes, described as a drop-in replacement for standard residual connections. Instead of every layer adding onto one uniformly accumulated stream, each layer can selectively retrieve representations from arbitrary earlier layers. </p><p>The claimed motivation is depth: in <strong>very deep MoE stacks</strong>, different experts fire at different depths, and a selective residual path keeps early information reachable without forcing it through every intermediate transformation. </p><p>It is the kind of change that sounds small and touches everything, and it is the single item on this list we most want to see ablated in the report.</p><p>Around the edges: a custom activation Moonshot calls SiTU, a sigmoid tanh unit replacing the usual <strong>SwiGLU family</strong>, and a per-head variant of the Muon optimizer, scheduling learning rates at the level of individual attention heads. </p><p>Muon is a K2-era inheritance; per-head scheduling is new. Moonshot&#8217;s aggregate claim is that the architectural and data changes together give K3 roughly <em>2.5 times the scaling efficiency of K2</em>, meaning more capability per unit of training compute. That number is <strong>vendor-measured</strong>, marked as unverified, and should be treated accordingly.</p><p>The vision path deserves more than a scope note, because it is half the modality story and it touches everything above. What is known: <strong>vision is native </strong>rather than adapter-based per the community model card reading, continuing the <strong>multimodal line</strong> Moonshot opened with K2.5, which means image tokens enter the same transformer stack and the same 896-expert router as text, and the multimodal scores ship on day one rather than in a point release. </p><p>The API mechanics are documented even where the economics are not: images go in as <strong>base64 payloads </strong>or platform file references, public image URLs are not accepted, and OpenRouter lists image input at the standard rates. </p><p>The agent framing leans on vision explicitly, a model pitched for iterating against screenshots, logs, and runtime feedback rather than captioning, and Moonshot&#8217;s own multimodal numbers, <strong>81.6 on MMMU-Pro </strong>and 94.3 on MathVision, sit in the launch tables.</p><p>What is not known at all is the <strong>exchange rate</strong>. The rate card is a single line, and whether an image bills at a multiplier over the <em>$3 text rate</em> or simply lands as however many tokens the tokenizer emits is unconfirmed, a gap the pricing trackers have flagged since day one. </p><p>The serving consequences do not wait for the answer. Whatever the patch geometry turns out to be, images are prefill load: a screenshot-heavy agent loop is a long-prompt workload wearing a different costume, and it collides with <strong>prefix caching </strong>in a way text does not, because a refreshed screenshot mid-transcript invalidates every cached token downstream of it. </p><p>The append-only discipline from the harness section applies doubly to media: attach new frames at the tail, never edit them in place.</p><p>And there is a broad research question waiting in the checkpoint. Early fusion means text and vision tokens share the router and the expert pool. </p><p>Whether the 896 experts partition by modality, some effectively becoming vision specialists, or blend, is exactly the kind of question the July 27 weights let anyone answer with an afternoon of <strong>routing statistics,</strong> and the answer feeds straight back into pruning, because a modality-partitioned pool compresses very differently from a blended one.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Four bits from the start</h2><p>The decision that makes 2.8T servable is not in the attention stack. It&#8217;s in the number format. K3 was trained with quantization-aware training from the <strong>supervised fine-tuning stage</strong> onward: MXFP4 for weights, MXFP8 for activations. </p><p>This is not a post-training quantization pass applied to a bf16 checkpoint. The model learned to live inside <strong>four-bit weights</strong>, compensating for quantization error during training instead of absorbing it afterward.</p><p>Terms, briefly, because the acronym is doing a lot of work. MXFP4 is the four-bit member of the <strong>OCP Microscaling family</strong>: weights are grouped into blocks of 32, each element is an E2M1 float, one sign bit, two exponent bits, one mantissa bit, and every block shares a single 8-bit power-of-two scale. </p><p>The element grid has eight magnitudes, 0, 0.5, 1, 1.5, 2, 3, 4, 6, spaced logarithmically rather than uniformly. Activations ride one tier up as <strong>MXFP8</strong>, eight-bit elements under the same block-scale scheme, where the precision matters for numerical stability. </p><p>The migration from K2 Thinking, which shipped INT4 quantization-aware training in November 2025, to floating-point four bits here is a change of grid, and grids have consequences you can measure. So we measured, on a synthetic but realistic weight distribution, a <strong>Gaussian with and without </strong>a sprinkle<strong> </strong>of<strong> 30-sigma outliers</strong>, block size 32, both formats implemented per spec:</p><pre><code><code># MXFP4 (E2M1 elements, shared E8M0 block scale) vs INT4 groupwise, block size 32.
# Question: what does each grid do to a realistic weight distribution?
import numpy as np
rng = np.random.default_rng(7)

E2M1 = np.array([0.0, 0.5, 1.0, 1.5, 2.0, 3.0, 4.0, 6.0])   # magnitudes

def quant_mxfp4(w, block=32):
    out = np.empty_like(w)
    for i in range(0, w.size, block):
        b = w[i:i+block]
        amax = np.abs(b).max()
        if amax == 0:
            out[i:i+block] = 0; continue
        shared_exp = np.floor(np.log2(amax)) - 2          # emax of E2M1 is 2
        scale = 2.0 ** shared_exp                          # E8M0: power of two only
        x = np.clip(np.abs(b) / scale, 0, 6.0)
        idx = np.abs(x[:, None] - E2M1[None, :]).argmin(1) # round to nearest level
        out[i:i+block] = np.sign(b) * E2M1[idx] * scale
    return out

def quant_int4(w, block=32):
    out = np.empty_like(w)
    for i in range(0, w.size, block):
        b = w[i:i+block]
        amax = np.abs(b).max()
        if amax == 0:
            out[i:i+block] = 0; continue
        scale = amax / 7.0                                 # fp16 scale, uniform grid
        out[i:i+block] = np.clip(np.round(b / scale), -7, 7) * scale
    return out

def report(name, w):
    for label, q in (("mxfp4", quant_mxfp4(w)), ("int4-g32", quant_int4(w))):
        rel = np.linalg.norm(w - q) / np.linalg.norm(w)
        small = np.abs(w) &lt; np.quantile(np.abs(w), 0.99)   # error on the 99% mass
        rel_small = np.linalg.norm(w[small] - q[small]) / np.linalg.norm(w[small])
        dead = (q[small] == 0).mean() * 100                # small weights crushed to zero
        print(f"{name:18s} {label:9s} relRMSE {rel:.4f}  relRMSE(99% mass) {rel_small:.4f}  zeroed {dead:5.2f}%")

w_gauss = rng.normal(0, 0.02, 1 &lt;&lt; 20)
report("gaussian", w_gauss)

w_out = w_gauss.copy()                                     # 0.2% outliers at 30x sigma
hit = rng.choice(w_out.size, w_out.size // 500, replace=False)
w_out[hit] = rng.choice([-1, 1], hit.size) * 0.6
report("gaussian+outliers", w_out)
</code></code></pre><p>Executed output:</p><pre><code><code>gaussian           mxfp4     relRMSE 0.1141  relRMSE(99% mass) 0.1102  zeroed  8.19%
gaussian           int4-g32  relRMSE 0.0970  relRMSE(99% mass) 0.1012  zeroed 13.37%
gaussian+outliers  mxfp4     relRMSE 0.1919  relRMSE(99% mass) 0.2353  zeroed 13.01%
gaussian+outliers  int4-g32  relRMSE 0.1499  relRMSE(99% mass) 0.2585  zeroed 18.40%
</code></code></pre><p>Read the table honestly and neither format dominates. </p><p>INT4 with a <strong>floating-point scale per block</strong> really wins global reconstruction error in both regimes, because its fifteen uniform levels and exact scale beat eight logarithmic levels under a power-of-two scale on Gaussian mass. </p><p>What the FP4 grid buys is essentially at the bottom of the distribution: it zeroes meaningfully fewer small weights, 8.2 versus 13.4 percent on clean data, 13.0 versus 18.4 with outliers, and it carries the bulk of the mass with lower error when an <strong>outlier inflates a block&#8217;s scale</strong>, because the log-spaced levels keep resolution near zero exactly where a uniform grid goes coarse. </p><p>Small weights are where accumulated damage shows up at scale, so this is not a cosmetic difference, but it is also not the headline reason for the format. </p><p>The headline reasons are that the exponent-only scale is free to apply in hardware, that <strong>Blackwell tensor cores</strong> execute block-scaled MXFP4 natively at full rate, with the Hugging Face community analysis reporting the same for<em> AMD&#8217;s MI400 class</em>, and that quantization-aware training exists precisely to claw back whatever the grid loses. </p><p>A lab that wanted a research artifact would release <strong>bf16</strong> and let the community fight over quants. A lab that wants its model actually served releases the four-bit weights it trained, in the format the current accelerator generation runs at full rate.</p><p>The arithmetic consequence: 2.8 trillion parameters at four bits is <strong>1.4 TB of element payload</strong>, and closer to 1.5 TB once the per-block scales ride along, since a 32-element block carries 136 bits, 4.25 effective bits per weight. </p><p>The same model in FP16 would be about 5.6 TB. That factor of nearly four compounds through the memory system: fewer devices to hold the model, and a<strong> quarter of the bandwidth per token </strong>spent reading weights. Be precise about what the four does and does not buy on the wire, though. </p><p>Expert dispatch and combine move activations, and activations are MXFP8, so all-to-all traffic halves relative to a bf16 model rather than quartering. The<strong> full 4x applies to weight movemen</strong>t: initial loading, host offload, expert migration and rebalancing. </p><p>Weights and activations shrink by different factors, and conflating them is how serving estimates go wrong by 2x.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Qfcd!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F25ef7b1b-c40c-4cb5-98f5-5822efaae7fe_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Qfcd!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F25ef7b1b-c40c-4cb5-98f5-5822efaae7fe_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!Qfcd!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F25ef7b1b-c40c-4cb5-98f5-5822efaae7fe_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!Qfcd!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F25ef7b1b-c40c-4cb5-98f5-5822efaae7fe_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!Qfcd!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F25ef7b1b-c40c-4cb5-98f5-5822efaae7fe_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Qfcd!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F25ef7b1b-c40c-4cb5-98f5-5822efaae7fe_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/25ef7b1b-c40c-4cb5-98f5-5822efaae7fe_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Weight storage by precision&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Weight storage by precision" title="Weight storage by precision" srcset="https://substackcdn.com/image/fetch/$s_!Qfcd!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F25ef7b1b-c40c-4cb5-98f5-5822efaae7fe_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!Qfcd!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F25ef7b1b-c40c-4cb5-98f5-5822efaae7fe_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!Qfcd!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F25ef7b1b-c40c-4cb5-98f5-5822efaae7fe_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!Qfcd!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F25ef7b1b-c40c-4cb5-98f5-5822efaae7fe_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 2. Weight storage by precision.</figcaption></figure></div><p>There is a subtler consequence for the open release. When the weights drop on July 27, they will be born four-bit. </p><p>There will be no ambiguity about which community quant is faithful, because the <strong>FP4 checkpoint is the reference</strong>, not a lossy derivative. The flip side is that the fine-tuning story gets strange: LoRA and QLoRA on top of an already-quantized MoE base is thinly explored territory, a point the Hugging Face overview raises as an open research question. </p><p>The first <strong>serious K3 fine-tunes</strong> will be experiments in method, not just in data. One loose thread already dangles here. One launch tracker describes OpenRouter&#8217;s route as Moonshot&#8217;s hosted INT4 endpoint. </p><p>That may be nothing more than four-bit shorthand for the MXFP4 checkpoint, or it may mean the serving fleet runs a different quantization than the training format, and the difference matters to anyone who plans to <strong>benchmark API behavior</strong> against the weights they download. File it with the questions the report owes.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>What it takes to serve it</h2><p>Napkin arithmetic first, ours and labeled as such, all of it downstream of the ~50B active estimate. </p><p>A single decode token&#8217;s forward pass touches 16 experts per MoE layer plus the shared trunk, on the order of 25 GB of weight reads at four bits. Against the <strong>roughly 8 TB/s of HBM bandwidth</strong> on a current Blackwell-class part, that is about three milliseconds of pure weight traffic, an upper bound near<strong> 320 tokens per second </strong>even if one device could hold the whole model, which it cannot. </p><p>The ratio is the point of the exercise: at FP16 the same pass would read 100 GB per token and the model would be bandwidth-strangled on any realistic hardware. Four-bit weights are <em>not an optimization here</em>. They are the enabling condition.</p><p>That 25 GB figure is also a batch-size-one number, and batch size one is the worst case this architecture has. Here is the mechanism, because it is the actual <strong>thesis of extreme sparsity</strong>. Dense-layer weights are read once per forward pass and amortize cleanly across every token in the batch.</p><p> Expert weights amortize only when tokens share experts, and with 16 of 896 routing, sharing takes scale. </p><p>Rather than assert the curve, we simulated it, sixteen trials per point, with two popularity models: uniform routing, which is what Quantile Balancing is trying to buy, and a moderately skewed distribution with a <strong>coefficient of variation around 0.7</strong>, which is what real routers produce when you let them:</p><pre><code><code># How expert weight traffic per token falls with batch size, and what routing
# skew does to it. E=896, k=16, ~1.5 GB per expert at MXFP4 (assumption, tier E).
import numpy as np
rng = np.random.default_rng(3)

E, K, GB_PER_EXPERT, TRIALS = 896, 16, 1.5, 16
batches = [1, 8, 64, 256, 1024, 4096]

def run(pop):                       # pop: expert popularity distribution, sums to 1
    uniq_gb, imb = [], []
    for B in batches:
        u, m = [], []
        for _ in range(TRIALS):
            draws = np.concatenate([rng.choice(E, K, replace=False, p=pop)
                                    for _ in range(B)])
            counts = np.bincount(draws, minlength=E)
            u.append((counts &gt; 0).sum())
            m.append(counts.max() / (B * K / E))          # max load over mean load
        uniq_gb.append(np.mean(u) * GB_PER_EXPERT / B)
        imb.append(np.mean(m))
    return uniq_gb, imb

uniform = np.full(E, 1 / E)
skewed = rng.dirichlet(np.full(E, 2.0))                    # cv ~0.7, mild hotness
g_u, i_u = run(uniform)
g_s, i_s = run(skewed)

print(f"{'batch':&gt;6} {'uniform GB/tok':&gt;15} {'imbalance':&gt;10} {'skewed GB/tok':&gt;14} {'imbalance':&gt;10}")
for j, B in enumerate(batches):
    print(f"{B:&gt;6} {g_u[j]:&gt;15.2f} {i_u[j]:&gt;10.1f} {g_s[j]:&gt;14.2f} {i_s[j]:&gt;10.1f}")
print(f"floor at large batch: {E * GB_PER_EXPERT:.0f} GB / B")
print("EP sizing at MXFP4: EP64 -&gt;", E // 64, "experts/GPU,", E // 64 * GB_PER_EXPERT,
      "GB | EP128 -&gt;", E // 128, "experts/GPU,", E // 128 * GB_PER_EXPERT, "GB")
</code></code></pre><p>Executed output:</p><pre><code><code> batch  uniform GB/tok  imbalance  skewed GB/tok  imbalance
     1           24.00       56.0          24.00       56.0
     8           22.72       14.9          21.82       18.8
    64           14.42        4.9          12.65        7.4
   256            5.20        2.7           4.80        5.5
  1024            1.31        1.8           1.30        4.6
  4096            0.33        1.4           0.33        4.6
floor at large batch: 1344 GB / B
EP sizing at MXFP4: EP64 -&gt; 14 experts/GPU, 21.0 GB | EP128 -&gt; 7 experts/GPU, 10.5 GB
</code></code></pre><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!RkV6!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F77019cbc-c703-4f89-8cff-24420681c437_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!RkV6!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F77019cbc-c703-4f89-8cff-24420681c437_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!RkV6!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F77019cbc-c703-4f89-8cff-24420681c437_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!RkV6!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F77019cbc-c703-4f89-8cff-24420681c437_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!RkV6!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F77019cbc-c703-4f89-8cff-24420681c437_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!RkV6!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F77019cbc-c703-4f89-8cff-24420681c437_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/77019cbc-c703-4f89-8cff-24420681c437_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Why extreme sparsity demands batch&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Why extreme sparsity demands batch" title="Why extreme sparsity demands batch" srcset="https://substackcdn.com/image/fetch/$s_!RkV6!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F77019cbc-c703-4f89-8cff-24420681c437_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!RkV6!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F77019cbc-c703-4f89-8cff-24420681c437_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!RkV6!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F77019cbc-c703-4f89-8cff-24420681c437_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!RkV6!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F77019cbc-c703-4f89-8cff-24420681c437_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 3. Why extreme sparsity demands batch.</figcaption></figure></div><p>Three things in that table are worth internalizing. First, the amortization cliff: per-token expert traffic barely moves from batch 1 to batch 8, because <strong>eight tokens drawing 128 expert slots</strong> land on about 120 distinct experts, which is to say almost no sharing at all. </p><p>It falls to 5 GB per token at batch 256, where the batch has touched nearly the whole 1,344 GB expert pool once, and approaches the 1,344 divided by B floor beyond a thousand concurrent tokens. </p><p>Second, the <strong>imbalance column,</strong> max expert load over mean, is the one your step time actually obeys, because the slowest device sets the pace of every synchronous layer.</p><p>Under uniform routing it decays toward 1.4 as the law of large numbers kicks in. Under skew it plateaus at about 4.6 and never recovers: the <strong>hot expert stays hot</strong>, no matter how big the batch gets. Third, notice the trade hiding between the columns: skew slightly reduces traffic, popular experts get reused, while wrecking balance. </p><p>Quantile Balancing, read through this table, is Moonshot choosing the uniform column, paying the <strong>worst-case traffic </strong>to buy the converging imbalance, because a 4.6x hot spot at expert parallelism width is a 4.6x tax on every token. </p><p>The structural conclusion stands either way: a 1.8 percent activation ratio is only cheap at large batch under wide expert parallelism, where every device holds a<strong> slice of the pool</strong>, reads its residents once per step, and serves whichever tokens the all-to-all delivers. </p><p>At EP64 that slice is 14 experts and 21 GB per GPU; at EP128, 7 experts and 10.5 GB, numbers that fit comfortably beside a trunk shard and a KV allocation. </p><p><strong>Extreme sparsity</strong> does not just permit big-batch serving. It demands it, which is why models shaped like this favor operators with deep pools of concurrent traffic and punish anyone trying to run them hot for three users.</p><p>The all-to-all itself can be sized on the back of the same napkin. Per MoE layer, each token&#8217;s hidden state ships to its 16 expert devices and 16 partial outputs ship back: roughly 2 x k x d_model bytes at FP8. </p><p><strong>K2&#8217;s hidden width was 7,168</strong> and K3&#8217;s is unpublished, so treat this as parameterized: 2 x 16 x 7,168 is about 229 KB per token per MoE layer, double K2&#8217;s 8-way routing at the same width. </p><p>Across an assumed 60 MoE layers that is roughly 14 MB of fabric traffic per generated token, and at <strong>15,000 tokens per second</strong> of aggregate throughput the expert-parallel group is moving on the order of 200 GB/s. Inside an NVLink domain at 1.8 TB/s per GPU that is background noise. </p><p>Stretched across racks on 400G links at 50 GB/s each, it is the budget. We have written before about the all-to-all tax and the rack-scale networking built to pay it; K3 is precisely the class of model that hardware exists for.</p><p>If the 2 x k x d term feels abstract, here is the entire MoE data path, route, dispatch, <strong>grouped GEMM</strong>, combine, in one page of numpy, checked against a dense reference:</p><pre><code><code># The whole MoE data path in one page: route, dispatch, grouped GEMM, combine.
# Toy sizes so it prints; the shapes are the ones a serving engine juggles.
import numpy as np
rng = np.random.default_rng(0)

T, D, E, K = 8, 16, 32, 4                  # tokens, hidden, experts, top-k
x = rng.normal(size=(T, D)).astype(np.float32)
W = rng.normal(size=(E, D, D)).astype(np.float32) * 0.1    # one matrix per expert
router = rng.normal(size=(D, E)).astype(np.float32)

logits = x @ router                                        # [T, E]
topk = np.argsort(-logits, axis=1)[:, :K]                  # [T, K] expert ids
gates = np.exp(logits[np.arange(T)[:, None], topk])
gates /= gates.sum(1, keepdims=True)                       # softmax over the k winners

# dispatch: replicate each token k times, then sort rows by destination expert
flat_expert = topk.ravel()                                 # [T*K]
flat_token  = np.repeat(np.arange(T), K)                   # [T*K]
order = np.argsort(flat_expert, kind="stable")             # the permutation
xp = x[flat_token[order]]                                  # [T*K, D] permuted buffer

# grouped GEMM: one segment per expert, ragged sizes
counts = np.bincount(flat_expert, minlength=E)
offsets = np.concatenate([[0], np.cumsum(counts)])
yp = np.empty_like(xp)
for e in range(E):                                         # in CUDA: one grouped kernel
    s, t = offsets[e], offsets[e + 1]
    if t &gt; s:
        yp[s:t] = xp[s:t] @ W[e]

# combine: unpermute and gate-weighted sum back to [T, D]
y = np.zeros_like(x)
np.add.at(y, flat_token[order], yp * gates.ravel()[order][:, None])

# reference: dense loop over tokens and their experts
y_ref = np.zeros_like(x)
for t in range(T):
    for j in range(K):
        y_ref[t] += gates[t, j] * (x[t] @ W[topk[t, j]])
print("matches dense loop:", np.allclose(y, y_ref, atol=1e-5))
print("permuted buffer rows:", xp.shape[0], f"({K}x the token count)")
bytes_per_tok = 2 * K * D * xp.itemsize
print(f"wire bytes per token, this layer: 2 x k x d x {xp.itemsize} = {bytes_per_tok}")
print("tokens per expert:", counts.tolist())
</code></code></pre><p>Executed output:</p><pre><code><code>matches dense loop: True
permuted buffer rows: 32 (4x the token count)
wire bytes per token, this layer: 2 x k x d x 4 = 512
tokens per expert: [2, 0, 1, 1, 0, 2, 0, 0, 2, 0, 1, 1, 1, 2, 0, 1, 2, 1, 1, 1, 0, 2, 0, 0, 2, 0, 0, 2, 2, 3, 1, 1]
</code></code></pre><p>Every serving-engine complication is already visible in the toy. The permuted buffer holds k times the token count, so top-16 doubles the activation memory and <strong>permutation traffic of a top-8 model</strong> before a single expert FLOP is spent. </p><p>The per-expert segments are ragged, 0 to 3 tokens here, thousands in production, which is why the <strong>expert matmul </strong>is a grouped GEMM with runtime-sized groups rather than a clean batched one. And the scatter-add in the combine is the reduction the fabric has to carry back. </p><p>Scale the token count by six orders of magnitude and the expert count by 28, and the sort, the offsets, and the segment loop become the router kernel, the dispatch layout, and the<strong> grouped GEMM</strong> that the serving stacks will spend the next quarter optimizing.</p><p>Memory next. Weights at ~1.49 TB fit inside a single 8x B200 node, 1,536 GB, with about <strong>50 GB to spare</strong>, and that sentence is a trap: the spare must hold every KV block, activation buffer, routing table, and graph on the box, so a one-node K3 is a demo, not a deployment. </p><p>Moonshot&#8217;s own launch guidance recommends a supernode of at least 64 accelerators, which community sizing reads as eight nodes of eight 80 GB GPUs, around<strong> 5 TB aggregate</strong>, a floor with headroom for cache and parallelism. </p><p>On a GB200 NVL72 the weights occupy roughly a tenth of the rack&#8217;s 13.4 TB of HBM. The DeepSeek 671B generation already made rack-scale expert parallelism a normal conversation; K3 raises the <strong>resident-byte requirement</strong> by roughly 4x&#8230;. and makes the full NVL72 look like the natural unit for a single replica rather than an extravagance.</p><p>The long-context economics deserve their own bytes, because the flat 1M pricing rests on them. In the DeepSeek V3 configuration that MLA descends from, each token stores a compressed latent per global layer: a <strong>512-wide KV latent</strong> plus a 64-wide decoupled positional key, 576 elements, 576 bytes at FP8. </p><p>If K3 carries something like the Kimi Linear ratio, call it 15 to 20 global layers out of the stack, a full million-token sequence holds roughly <strong>9 to 12 GB of KV</strong>, ours again and resting on the assumed ratio. A comparable dense model, 60 layers of grouped-query attention with 8 KV heads of dimension 128 at FP16, holds about 246 KB per token, 246 GB at the same window, per sequence. </p><p>That is a factor of about 25, and it is the difference between a million-token session being a scheduling problem and being a physical impossibility. The KDA layers contribute a fixed-size state regardless of length, which is a rounding error next to either number. </p><p>This is the <strong>architecture showing</strong> through the rate card: Moonshot can price the window flat because most of the model does not pay for it.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!M_21!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb59ad582-3fe1-44b7-8cb6-7ca45f4ab3f0_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!M_21!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb59ad582-3fe1-44b7-8cb6-7ca45f4ab3f0_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!M_21!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb59ad582-3fe1-44b7-8cb6-7ca45f4ab3f0_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!M_21!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb59ad582-3fe1-44b7-8cb6-7ca45f4ab3f0_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!M_21!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb59ad582-3fe1-44b7-8cb6-7ca45f4ab3f0_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!M_21!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb59ad582-3fe1-44b7-8cb6-7ca45f4ab3f0_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b59ad582-3fe1-44b7-8cb6-7ca45f4ab3f0_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;KV cache at the full window&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="KV cache at the full window" title="KV cache at the full window" srcset="https://substackcdn.com/image/fetch/$s_!M_21!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb59ad582-3fe1-44b7-8cb6-7ca45f4ab3f0_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!M_21!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb59ad582-3fe1-44b7-8cb6-7ca45f4ab3f0_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!M_21!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb59ad582-3fe1-44b7-8cb6-7ca45f4ab3f0_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!M_21!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb59ad582-3fe1-44b7-8cb6-7ca45f4ab3f0_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 4. KV cache at the full window.</figcaption></figure></div><p>Prefill is the other half, and it is why the serving system is disaggregated. Filling the full window on a cache miss costs $3.15 of input at the sticker rate, $0.31 on a hit, and the compute behind the miss is brutal. </p><p>Linear-layer FLOPs alone run <strong>about 2 x 50B x 1M tokens</strong>, on the order of 10^17, several seconds on a full H100 node at perfect efficiency and the better part of half a minute at realistic utilization, before the attention term, which grows quadratically on the global layers and dominates at this length even confined to a quarter of the stack. </p><p>The <strong>KDA layers</strong> prefill differently again: a chunked scan, parallel within each chunk with recurrent state carried across chunk boundaries, a kernel shape with nothing in common with FlashAttention. </p><p><em>Prefill at 1M</em> is a compute event; decode is a bandwidth event; the two want different hardware shapes and different batching. Moonshot serves K3 on Mooncake, its <strong>KVCache-centric</strong> disaggregated architecture, which splits prefill and decode across separate node pools and treats the cache as a first-class pooled resource between them. </p><p>Mooncake is not opaque vendor infrastructure: the design won the <strong>best paper award at FAST 25</strong>, the code is open-sourced, and its transfer engine has been integrated into the mainstream serving stacks, including vLLM and SGLang. </p><p>The company reports cache hit rates above 90 percent on coding workloads. That figure is vendor-reported, but it is plausible for agent traffic, where the <strong>same system prompt</strong>, repository context, and conversation prefix recur on every loop iteration, and the published architecture is engineered to produce exactly that number. </p><p>The hit rate is what makes the pricing model work, which brings us to the bill by way of the software that has to exist first.</p><blockquote><p><em>Extreme sparsity does not just permit big-batch serving. It demands it.</em></p></blockquote><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>What has to change in the serving stacks</h2><p>Self-hosting K3 is not a config file away, and it is worth being specific about where the engineering actually lands, because the July 27 story will be written in pull requests.</p><p>Start at the router. Every MoE layer produces logits of shape tokens by 896, takes a top-16, and must land tokens in per-expert segments fast enough <strong>not to shadow the GEMMs</strong>. </p><p>The grouped GEMM behind it now has up to 896 runtime-sized groups per layer per rank, and the tokens-per-expert variance the toy above showed becomes <strong>tile quantization waste</strong> at CUDA scale: an expert with 37 tokens still occupies whole tensor-core tiles. </p><p>The mature answers are persistent grouped kernels that pull ragged work from a queue instead of launching per group, and the CUTLASS grouped GEMM machinery most stacks already wrap. </p><p>What changes for K3 is scale, 896 groups against the 256 the current kernels were tuned around, and precision, because the groups now <strong>multiply block-scaled FP4 weights </strong>against FP8 activations.</p><p>Precision is where the hardware generations split. Blackwell&#8217;s fifth-generation tensor cores execute block-scaled MXFP4 natively, scales applied in the MMA itself, so B200-class serving runs the checkpoint as stored. </p><p>Hopper has no FP4 MMA path at all: an H100 deployment dequantizes weights to FP8 or BF16 in registers on the way into the matmul, the <strong>Marlin </strong>and <strong>Machete family </strong>of mixed-input kernels, which exist and are fast for dense weight-only GEMMs but are immature for 896-group grouped MoE at launch. </p><p>That dequant tax is not cosmetic. It is compute spent on format conversion inside the innermost loop, and it lands directly on the self-hosting ledger below: the 37 tokens per second per GPU that beats the sticker price assumes the<strong> grouped W4A8 path</strong> works well on Hopper, and today that is an assumption.</p><p>The communication layer has a shelf to raid but not a finished product. DeepSeek open-sourced DeepEP, NVSHMEM-based dispatch and combine kernels with FP8 payloads, <strong>IBGDA for the internode path</strong>, and a low-latency mode for decode, and it is the obvious starting point. </p><p>It was also built and tuned around the 256-expert, top-8 shape of the V3 generation. K3 doubles the per-token payload with top-16 and multiplies the destination fan-out with 896 experts, so the kernel launch geometry, the <strong>SM budget </strong>reserved <strong>for communication</strong>, and the overlap schedule that hides all-to-all latency behind expert compute all need retuning. </p><p>The dual-microbatch overlap trick, one microbatch&#8217;s communication hidden under another&#8217;s GEMMs, matters more here, not less, because the fabric term we sized above grows with k while the compute term holds near flat.</p><p>The subtlest problem is the cache manager, and it is the one we would watch first. Serving frameworks assume a homogeneous per-layer paged KV cache; the hybrid allocators built for the <strong>Mamba-generation models </strong>already broke that assumption, and K3 stresses it further: paged 576-byte latents for the global layers next to fixed-size recurrent state for the KDA layers. </p><p>Decode is straightforward. Prefix caching is not, because the entire cache-hit economy assumes you can resume from a stored prefix, and for a recurrent layer that means storing the state at the reuse boundary, not just the tokens. </p><p>A KV pool like Mooncake&#8217;s holds paged latents naturally; checkpointed KDA states at chunk boundaries are a second object type with different lifetime and size semantics, and <strong>how Moonshot handles state-restore</strong> for the linear layers is, to us, the single most interesting undisclosed detail in the serving stack. </p><p>Whoever implements it in the open stacks will be making a real design decision, not just a port.</p><p>The good news is how much is already on the shelf. The chunked delta-rule kernels for KDA shipped in the open with the Kimi Linear release through the flash-linear-attention line. </p><p>FlashMLA covers the global layers on Hopper. DeepEP covers the fabric, pending retuning. Mooncake&#8217;s transfer engine is merged where it needs to be. </p><p>What does not exist yet, anywhere public, is the integration: one engine that schedules block-scaled grouped GEMMs, hybrid attention with two kernel families, a two-type cache, and disaggregated state transfer, for a 1.5 TB checkpoint. </p><p>That is the <strong>artifact to watch</strong> for in the weeks after July 27, and its arrival date, more than the weights themselves, decides when open K3 is real.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>The bill</h2><p>The rate card, from Moonshot&#8217;s platform pages:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Xn2c!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa7260193-a4e1-4c92-a46a-995dc07fd591_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Xn2c!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa7260193-a4e1-4c92-a46a-995dc07fd591_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!Xn2c!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa7260193-a4e1-4c92-a46a-995dc07fd591_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!Xn2c!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa7260193-a4e1-4c92-a46a-995dc07fd591_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!Xn2c!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa7260193-a4e1-4c92-a46a-995dc07fd591_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Xn2c!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa7260193-a4e1-4c92-a46a-995dc07fd591_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a7260193-a4e1-4c92-a46a-995dc07fd591_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:81860,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/206014241?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa7260193-a4e1-4c92-a46a-995dc07fd591_1634x893.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Xn2c!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa7260193-a4e1-4c92-a46a-995dc07fd591_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!Xn2c!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa7260193-a4e1-4c92-a46a-995dc07fd591_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!Xn2c!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa7260193-a4e1-4c92-a46a-995dc07fd591_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!Xn2c!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa7260193-a4e1-4c92-a46a-995dc07fd591_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>All figures per million tokens. Web search tool calls bill separately at $0.004 each, though the tool ships flagged as being updated and one guide advises against production use for now. Vision input pricing is unconfirmed, per the vision section above.</em></p><p>The first thing the sticker tells you is positioning. K2 built its reputation as the cheap frontier model. K3 abandons that identity: at $3 and $15 it matches <strong>Claude Sonnet 5 </strong>to the cent, undercuts Opus 4.8 at $5 and $25 and GPT-5.6 Sol at $5 and $30, and costs several times its own K2.6 sibling.</p><p>Within the Chinese open cohort it is the expensive option, with DeepSeek V4 Pro at roughly a sixth of its input price and GLM 5.2 around half. Moonshot is telling you it believes K3 competes on capability, and it has priced away the discount narrative to prove it.</p><p>The second thing is that the sticker is not the effective price. The cache tier is. At a <strong>90 percent hit rate</strong>, effective input cost is 0.9 times $0.30 plus 0.1 times $3.00, about $0.57 per million tokens. </p><p>Agent workloads with stable prefixes live near that floor. Caching is automatic, with <strong>no cache ID or TTL management</strong>, which removes the usual engineering excuse for missing it, and OpenRouter&#8217;s traffic view says the same thing from the demand side: average realized K3 prices run 60 to 80 percent below list once caching is counted. </p><p>Two more levers are worth knowing. K3 ships with no batch tier, the 60 percent batch discount on Moonshot&#8217;s menu covers only the K2 line, and the API defaults max_completion_tokens to 131,072, configurable up to the full window, a cap worth leaving in place until you have watched what max-effort thinking does to a bill.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!3nL3!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe5513bb3-2038-4118-b06b-ec2249018314_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!3nL3!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe5513bb3-2038-4118-b06b-ec2249018314_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!3nL3!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe5513bb3-2038-4118-b06b-ec2249018314_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!3nL3!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe5513bb3-2038-4118-b06b-ec2249018314_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!3nL3!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe5513bb3-2038-4118-b06b-ec2249018314_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!3nL3!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe5513bb3-2038-4118-b06b-ec2249018314_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e5513bb3-2038-4118-b06b-ec2249018314_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;K3 effective input price vs cache hit rate&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="K3 effective input price vs cache hit rate" title="K3 effective input price vs cache hit rate" srcset="https://substackcdn.com/image/fetch/$s_!3nL3!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe5513bb3-2038-4118-b06b-ec2249018314_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!3nL3!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe5513bb3-2038-4118-b06b-ec2249018314_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!3nL3!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe5513bb3-2038-4118-b06b-ec2249018314_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!3nL3!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe5513bb3-2038-4118-b06b-ec2249018314_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 5. K3 effective input price vs cache hit rate.</figcaption></figure></div><p>The third thing cuts the other way. K3 always reasons, at max effort, and reasoning tokens bill as output at $15. <strong>Artificial Analysis measured 130 million output tokens </strong>to run its Intelligence Index evaluation, against a median near 63 million for comparable reasoning models. </p><p>The same measurements put decode speed at 62 tokens per second, under the 72.7 median for the price tier, and split latency into the two numbers an always-thinking model needs: about 2 seconds to the first streamed token, 34.24 seconds to the first answer token, which at the measured decode rate is roughly 2,000 thinking tokens spent before the answer begins on a standard 10,000-token workload. </p><p>Running the full index cost $2,709.75, nearly all of it output. Independent testing has recorded over thirteen thousand thinking tokens on a single short task, roughly twenty cents of deliberation before the first answer token. </p><p>Artificial Analysis puts the blended cost at $2.31 per million tokens on a 7:2:1 cache, input, output mix, but your mix will vary, and a maximally verbose reasoner shifts the mix toward the expensive column. &#8220;<em>Price it on your own traces</em>&#8221; is easy to say obv, so here is the whole model in twenty lines, run against three workload archetypes at launch rates:</p><pre><code><code># Request-level cost model for kimi-k3 at launch pricing.
# $3/MTok input on miss, $0.30 on cache hit, $15/MTok output.
# Thinking is always on and bills as output.
IN_MISS, IN_HIT, OUT = 3.00, 0.30, 15.00

def request_cost(in_tokens, cached_frac, answer_tokens, thinking_tokens):
    inp = in_tokens * ((1 - cached_frac) * IN_MISS + cached_frac * IN_HIT)
    out = (answer_tokens + thinking_tokens) * OUT
    return inp / 1e6, out / 1e6

archetypes = [
    # name, input, cached share, answer, thinking, requests
    ("short chat turn",      2_000, 0.20,   300, 13_000, 1),
    ("agent loop, 40 iters", 80_000, 0.95,   500,  6_000, 40),
    ("1M-doc analysis",     900_000, 0.00, 2_000, 10_000, 1),
]
print(f"{'workload':&lt;22}{'$ input':&gt;9}{'$ output':&gt;9}{'$ total':&gt;9}{'output share':&gt;14}")
for name, i, c, a, th, n in archetypes:
    ci, co = request_cost(i, c, a, th)
    ci, co = ci * n, co * n
    print(f"{name:&lt;22}{ci:&gt;9.3f}{co:&gt;9.3f}{ci+co:&gt;9.2f}{co/(ci+co):&gt;13.0%}")
</code></code></pre><p>Executed output:</p><pre><code><code>workload                $ input $ output  $ total  output share
short chat turn           0.005    0.200     0.20          98%
agent loop, 40 iters      1.392    3.900     5.29          74%
1M-doc analysis           2.700    0.180     2.88           6%
</code></code></pre><p>The <strong>distribution of pain</strong> is the finding. On a small chat turn, deliberation is 98 percent of the bill: the answer costs half a cent and the thinking costs twenty. </p><p>A forty-iteration agent loop, the workload K3 is aimed at, still spends three of every four dollars on output even with a 95 percent cache hit rate doing everything right on the input side. Only the <strong>cold million-token document</strong> flips the ratio, and that shape pays $2.70 of prefill precisely once, after which it becomes an agent loop too. </p><p>The lever ordering falls out directly: for two of the three archetypes that matter, the promised low and high effort modes are worth more than any cache optimization you can do, and until they ship, the K2.6 rate card at $0.95 and $4 remains the right tool for every request that does not need a frontier model deliberating at full depth.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Pf9j!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F798b78b6-a19c-47fd-ba61-e468e0d66318_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Pf9j!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F798b78b6-a19c-47fd-ba61-e468e0d66318_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!Pf9j!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F798b78b6-a19c-47fd-ba61-e468e0d66318_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!Pf9j!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F798b78b6-a19c-47fd-ba61-e468e0d66318_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!Pf9j!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F798b78b6-a19c-47fd-ba61-e468e0d66318_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Pf9j!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F798b78b6-a19c-47fd-ba61-e468e0d66318_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/798b78b6-a19c-47fd-ba61-e468e0d66318_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Where the money goes by workload&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Where the money goes by workload" title="Where the money goes by workload" srcset="https://substackcdn.com/image/fetch/$s_!Pf9j!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F798b78b6-a19c-47fd-ba61-e468e0d66318_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!Pf9j!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F798b78b6-a19c-47fd-ba61-e468e0d66318_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!Pf9j!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F798b78b6-a19c-47fd-ba61-e468e0d66318_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!Pf9j!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F798b78b6-a19c-47fd-ba61-e468e0d66318_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 6. Where the money goes by workload.</figcaption></figure></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>The self-hosting ledger</h2><p>The weight release invites an obvious question: </p><blockquote><p><em>at what point does running K3 yourself beat paying Moonshot? </em></p></blockquote><p>The napkin again, ours. Take the community floor configuration, 64 H100-class GPUs, and a rental price of $2 per GPU-hour: $128 per hour for the cluster. </p><p>To generate tokens cheaper than the $15 output rate, the cluster must sustain about 2,370 tokens per second aggregate, or 37 per GPU. That is an achievable number for a<strong> well-tuned MoE deployment</strong> with real batch depth, with the caveat the kernel section earned: on Hopper it assumes the grouped W4A8 dequant path matures, because every register spent converting FP4 is throughput the break-even does not get. </p><p>To beat the $2.31 blended rate that cached API traffic actually pays, the same cluster must sustain <strong>about 15,400 tokens per second</strong>, 240 per GPU, which is a very hard number for a model this size on Hopper-generation hardware. </p><p>Independent color corroborates the memory wall from an unexpected direction: an early community stress test found a 1.5 TB unified-memory Mac Studio cluster at the edge of what a K3-class million-token stack can carry, which is the same wall the 1.49 TB storage figure predicts.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!BdUF!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0424b430-a537-408b-9641-07fff9bfd4ea_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!BdUF!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0424b430-a537-408b-9641-07fff9bfd4ea_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!BdUF!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0424b430-a537-408b-9641-07fff9bfd4ea_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!BdUF!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0424b430-a537-408b-9641-07fff9bfd4ea_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!BdUF!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0424b430-a537-408b-9641-07fff9bfd4ea_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!BdUF!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0424b430-a537-408b-9641-07fff9bfd4ea_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0424b430-a537-408b-9641-07fff9bfd4ea_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Self-hosting K3: 64x H100 at $2 per GPU-hour&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Self-hosting K3: 64x H100 at $2 per GPU-hour" title="Self-hosting K3: 64x H100 at $2 per GPU-hour" srcset="https://substackcdn.com/image/fetch/$s_!BdUF!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0424b430-a537-408b-9641-07fff9bfd4ea_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!BdUF!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0424b430-a537-408b-9641-07fff9bfd4ea_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!BdUF!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0424b430-a537-408b-9641-07fff9bfd4ea_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!BdUF!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0424b430-a537-408b-9641-07fff9bfd4ea_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 7. Self-hosting K3: 64x H100 at $2 per GPU-hour.</figcaption></figure></div><p>The comparison is loose by construction: the API&#8217;s $15 covers output alone while your cluster serves whole requests, the blended target moves with your cache profile, and <strong>$2 per H100-hour</strong> is a spot figure that swings with region and term. </p><p>The asymmetry survives all of that. Self-hosting K3 competes with Moonshot&#8217;s sticker prices and loses badly to Moonshot&#8217;s cache economics, because the cache discount is not a margin decision you can copy, it is a serving architecture, <strong>Mooncake&#8217;s pooled KV </strong>plus a workload with 90 percent prefix reuse, that a single-tenant cluster cannot replicate without the traffic to feed it. </p><p>The teams for whom self-hosting pencils out are the ones who need it for reasons the rate card does not price: data residency (<em>the hosted API processes in Singapore, per an EU procurement review</em>), fine-tuned variants, latency floors, or sovereignty. </p><p>For everyone else, the July 27 download is leverage in a pricing negotiation, not a cost reduction.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Benchmarks, with the usual caution</h2><p>Moonshot&#8217;s own evaluation places K3 at frontier level but behind the two strongest closed models, Claude Fable 5 and GPT-5.6 Sol, an admission worth respecting because labs rarely volunteer it. </p><p><strong>On GDPval-AA v2</strong>, the aggregate knowledge-work evaluation, the reported ordering is Fable 5 Max at 1815, GPT-5.6 Sol Max at 1747.8, K3 at 1687, and Opus 4.8 at 1600. On the science and multimodal side, Moonshot reports <em>93.5 on GPQA-Diamond</em>, <strong>81.6 on MMMU-Pro</strong>, and 94.3 on MathVision.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!oLLt!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F724f72f0-1247-4494-a8bf-1965da6f2f7e_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!oLLt!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F724f72f0-1247-4494-a8bf-1965da6f2f7e_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!oLLt!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F724f72f0-1247-4494-a8bf-1965da6f2f7e_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!oLLt!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F724f72f0-1247-4494-a8bf-1965da6f2f7e_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!oLLt!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F724f72f0-1247-4494-a8bf-1965da6f2f7e_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!oLLt!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F724f72f0-1247-4494-a8bf-1965da6f2f7e_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/724f72f0-1247-4494-a8bf-1965da6f2f7e_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;GDPval-AA v2 as reported&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="GDPval-AA v2 as reported" title="GDPval-AA v2 as reported" srcset="https://substackcdn.com/image/fetch/$s_!oLLt!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F724f72f0-1247-4494-a8bf-1965da6f2f7e_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!oLLt!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F724f72f0-1247-4494-a8bf-1965da6f2f7e_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!oLLt!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F724f72f0-1247-4494-a8bf-1965da6f2f7e_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!oLLt!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F724f72f0-1247-4494-a8bf-1965da6f2f7e_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 8. GDPval-AA v2 as reported.</figcaption></figure></div><p>The coding numbers are where it gets interesting. K3 posts 88.3 on Terminal-Bench 2.1, half a point behind GPT-5.6 Sol, and takes the top score outright on Program Bench at 77.8 and on SWE Marathon at 42.0. SWE Marathon is the<strong> long-horizon one</strong>, and a model leading there while merely placing elsewhere is consistent with the design: a million-token window that holds an entire repository plus the full history of a long session is worth more on sustained work than on sprints. </p><p>FrontierSWE lands at 81.2, against a reported 86.6 for Fable 5 on the same benchmark, and DeepSWE at 67.5.</p><p>On agentic evaluations, VentureBeat&#8217;s read of the launch documentation has K3 first in four of eight real-world automation benchmarks, including Automation Bench, SpreadsheetBench 2, and BrowseComp, and second to Fable 5 in most of the rest. </p><p>The reported scores include 91.2 on BrowseComp, a 95.0 F1 on DeepSearchQA, and 30.8 on Automation Bench, which tells you as much about Automation Bench as about the model. </p><p>The claim we find most interesting is methodological: Moonshot says these results came from a <strong>single-agent setup </strong>running on the raw million-token window, with no context compression or management scaffolding (<em>BrowseComp reads 90.4 in that configuration against the 91.2 headline figure, both vendor numbers</em>). </p><p>Taken at face value, that is a real data point in the context-versus-orchestration argument, raw window plus strong retrieval beating elaborate<strong> multi-agent plumbing,</strong> and it needs independent replication before it graduates from press release to argument.</p><p>Two launch-week numbers deserve a somewhat explicit skepticism. K3 jumped from rank 18 to rank 1 on<strong> LMArena&#8217;s Frontend Code Arena</strong> within hours, at a score of 1679. </p><p>The placement itself has real sourcing, the Associated Press reported K3 at the top of the frontend ranking, but arena scores during a hype cycle measure attention as much as ability; check back in a month. </p><p>BenchLM&#8217;s weighted leaderboard still lacked a K3 row at press time, an absence worth noting, but the first independent aggregate has landed: <strong>Artificial Analysis scores K3 at 57</strong> on its Intelligence Index, fourth of 189 models tracked, against an average of 31 for comparable models, and the index rolls up nine evaluations including GDPval-AA v2, Terminal-Bench 2.1, GPQA Diamond, and Humanity&#8217;s Last Exam. </p><p>Frontier company, by an outside ruler. The 2.5x scaling efficiency figure, the 90 percent cache hit rate, and the single-agent protocol are all vendor-reported. <strong>None of this means the numbers are wrong</strong>. It means the confirmations are scheduled for<strong> after July 27.</strong></p><p>One demo is worth relaying because it is verifiable in kind if not yet in fact. Moonshot reports that K3, in a single 48-hour autonomous run, designed a serving chip for a nano model on its own architecture using <strong>open-source EDA tools</strong> against the Nangate 45nm library: four square millimeters, timing closed at 100 MHz, 1.46 million standard cells, 0.277 MB of SRAM, an INT4 multiply-accumulate array with fused dequantization, and a simulated decode throughput above 8,700 tokens per second. A staged demo, obviously, and simulation is not silicon. </p><p>But as a choice of demo it is telling: the lab wanted to show long-horizon agency on exactly the kind of task this audience does for a living. </p><p>A second staged demo aims at researchers: to reproduce the I-Love-Q universal relations in computational astrophysics, K3 reportedly cross-validated <strong>more than twenty papers</strong>, evaluated over three hundred equations of state, caught inconsistencies in published formulas, and shipped three thousand lines of Python plus an interactive dashboard in about two hours, against Moonshot&#8217;s estimate of one to two weeks for an experienced researcher. </p><p>Vendor-staged, like the chip, and chosen with the same message: long horizons are the product.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Where it strains</h2><p>Moonshot names three weaknesses, and all three have operational consequences.</p><p>The first is thinking history sensitivity: agent harnesses that truncate or rewrite the model&#8217;s reasoning traces degrade quality significantly. Launch coverage adds the reason: K3 was trained in a <strong>preserved thinking history mode</strong>, so a harness that drops the thinking, or a session handed mid-flight from another model, is operating outside the training distribution. </p><p>For agent builders this is a real constraint, and it compounds with the cache economics, because the two failure modes share a root cause: mutating the transcript. </p><p>Strip the thinking and quality drops; touch anything upstream in the serialized prefix and you<strong> lose the $0.30 tier </strong>and re-prefill from the perturbation point. </p><p>The defense against both is the same discipline, an append-only history with a verifiable prefix, cheap enough to enforce in a dozen lines:</p><pre><code><code># K3 penalizes two things at once: perturbing the prompt prefix (you lose the
# $0.30 cache tier) and truncating its reasoning traces (quality degrades).
# Same defense for both: an append-only history with a verifiable prefix.
import hashlib, json

def fingerprint(messages):
    blob = json.dumps(messages, sort_keys=True, separators=(",", ":"))
    return hashlib.sha256(blob.encode()).hexdigest()[:16]

def prefix_intact(old, new):
    return len(new) &gt;= len(old) and fingerprint(new[:len(old)]) == fingerprint(old)

history = [
    {"role": "system", "content": "You are a code review agent."},
    {"role": "user", "content": "Review PR 4217."},
    {"role": "assistant", "content": "&lt;thinking&gt;...&lt;/thinking&gt; Two issues found."},
]
appended  = history + [{"role": "user", "content": "Fix the first one."}]
truncated = [history[0], history[1],
             {"role": "assistant", "content": "Two issues found."}]  # thinking stripped
reordered = [history[1], history[0], history[2]]                     # system moved

for name, h in (("append-only", appended), ("thinking stripped", truncated),
                ("system reordered", reordered)):
    print(f"{name:&lt;20} prefix intact: {prefix_intact(history, h)}")

# Client shape (not executed here; the endpoint is OpenAI-compatible):
#   client = OpenAI(base_url="https://api.moonshot.ai/v1", api_key=KEY)
#   r = client.chat.completions.create(model="kimi-k3", messages=history)
#   history.append({"role": "assistant", "content": r.choices[0].message.content})
</code></code></pre><p>Executed output:</p><pre><code><code>append-only          prefix intact: True
thinking stripped    prefix intact: False
system reordered     prefix intact: False
</code></code></pre><p>The two False rows are the two expensive mistakes. A harness that &#8220;<em>helpfully</em>&#8221; compacts old reasoning fails the quality constraint and the cache constraint in the same commit; one that <strong>rebuilds the message list per call</strong>, reordering tools or regenerating timestamps, silently pays full prefill price on every iteration. </p><p>Fingerprint the prefix in development, assert it in production, and design summarization as a new branch rather than an edit to history. On long sessions where the window genuinely fills, the <strong>append-only rule </strong>forces the honest version of the tradeoff: fork the conversation with an explicit summary and eat one cold prefill, rather than quietly degrading the model mid-branch.</p><p>The second weakness is overeagerness. In ambiguous situations K3 tends to act rather than ask. On benchmarks that rewards decisiveness. In production it is <strong>how an agent cheerfully migrates the wrong database</strong>. The mitigation is boring and necessary: explicit confirmation gates in the harness, because the model will not supply them.</p><p>The third is polish. Moonshot concedes the subjective experience trails Fable 5 and GPT-5.6 Sol even where benchmark numbers are close. Benchmark parity and product parity remain different finish lines.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share The Software Frontier&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share The Software Frontier</span></a></p><div><hr></div><h2>What July 27 decides</h2><p>The weight release is the actual event; July 16 was the trailer.</p><p>The license text comes first. Modified MIT is reported across the launch coverage and would match the K2 family precedent, and the precedent has teeth worth knowing: the <strong>modified MIT </strong>that ships with K2.7 Code adds an attribution requirement that triggers at 100 million monthly active users or 20 million dollars of monthly revenue. </p><p>If K3 inherits the clause it is a display obligation for giants rather than a commercial restriction, but the binding text is the <strong>LICENSE file</strong> that lands with the weights, and nobody should sign a deployment plan before reading it.</p><p>Day-one inference support is the second gate, and the stacks section above is the checklist: block-scaled grouped GEMM at 896 groups, retuned dispatch and combine at top-16, the two-type cache with KDA state restore, and a Hopper W4A8 path that does not bleed the break-even dry. </p><p>Moonshot says it is working with inference partners ahead of the drop, and the <strong>Mooncake integrations </strong>already sitting in those codebases suggest the relationship is real; the commits will tell.</p><p>Third, reproduction. Someone outside Moonshot needs to run the benchmark suite, and the single-agent long-context protocol specifically.</p><p>Fourth, the research surface the checkpoint opens. With 896 experts in the open, expert specialization can finally be probed at frontier scale.<strong> Pruning experiments </strong>will ask whether the pool compresses to something that fits a smaller cluster, and what the quality curve looks like on the way down.</p><p>AttnRes can be ablated at small scale to see whether it generalizes or only pays at depth. And someone will attempt the <strong>first LoRA on a natively FP4 base </strong>and discover what breaks.</p><p>Then, the field is not holding still. The DeepSeek V4 line is already live at a fraction of K3&#8217;s price, and community reporting has a V4 refresh in staged testing; <strong>GLM 5.5 and MiniMax Pro </strong>are both expected around the trillion-parameter mark, and Qwen 4 is on the horizon. Whatever lead K3 holds is measured in weeks, and DeepSeek&#8217;s answer is the one to watch.</p><p>Our read, calmly: K3 does not win the frontier. By Moonshot&#8217;s own tables it sits a step behind Fable 5 and Sol. What it breaks is the assumption that open weights trail the frontier by six months and a capability class. </p><p>The gap is now single-digit percentage points and a dated download link, at Sonnet prices, on an architecture whose every major choice, the sparsity ratio, the hybrid attention, the <strong>four-bit training</strong>, the disaggregated serving, was made with the inference bill in mind. The link goes live on July 27, and that is where the work starts.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Software Frontier is a reader-supported publication. Consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p><div><hr></div><h2>Sources and confidence</h2><p><strong>A, primary:</strong> Moonshot&#8217;s K3 launch blog and platform documentation: release dates, architecture component names, context window, pricing, stated limitations, chip design demo, benchmark tables. Published background used for parameterized math and the serving-stack analysis: the OCP Microscaling format specification, the Kimi Linear technical report (<em>KDA design, the three to one hybrid ratio, KV reduction and decode speedup figures, the open kernel release</em>), the Mooncake paper and open-source release (FAST 25 best paper; transfer engine integrated in mainstream serving stacks), the DeepSeek V3 MLA configuration (512 plus 64 latent dimensions), and the DeepEP communication library.</p><p><strong>B, independent measurement:</strong> Artificial Analysis: Intelligence Index score of 57 (fourth of 189), output token consumption on the index run (130M vs a 63M median), decode speed of 62 tokens per second against a 72.7 tier median, latency of about 2 seconds to first token and 34.24 seconds to first answer token, evaluation cost of $2,709.75, and blended cost ($2.31/M at 7:2:1). OpenRouter: listing at matching rates, and platform-reported realized prices 60 to 80 percent below list under caching.</p><p><strong>C, informed secondary:</strong> VentureBeat on benchmark placements, release timing, and corporate context. The Hugging Face community model overview for the ~50B active parameter estimate, MXFP4 and MXFP8 details, and deployment sizing. Verdent and multiple pricing guides for the comparative rate card and the modified MIT report.</p><p><strong>D, vendor-claimed, unverified:</strong> the 2.5x scaling efficiency over K2, the 90 percent cache hit rate, the single-agent no-compression benchmark protocol, launch-week arena rankings, and the chip demo results.</p><p><strong>E, our arithmetic and code:</strong> the batch amortization simulation, the quantization grid experiment, the dispatch reference implementation, the cost model, the prefix fingerprint demo, all-to-all sizing, KV cache bytes, prefill FLOPs, and the self-host break-even. Every code block in this issue was executed as printed, CPython 3.12.3 with numpy 2.4.4, with fixed seeds; outputs are shown verbatim. Assumptions are stated inline; every figure downstream of the ~50B active estimate, the assumed 7,168 hidden width, the ~60 MoE layer count, the ~1.5 GB per expert, and the assumed global-layer ratio inherits their uncertainty and will be recomputed when the technical report publishes the real dimensions.</p><p>The Kimi K3 technical report had not been published at the time of writing. Numbers in this issue may be revised when it appears, and we will follow up after the July 27 weight release.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/p/kimi-k3-28-trillion-parameters-four?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/p/kimi-k3-28-trillion-parameters-four?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><div><hr></div><h2>The &#8220;fact check&#8221; ledger</h2><p>Every load-bearing claim, its source tier, and how it was checked. Rows marked recomputed were re-derived at build time by the scripts shipped with this issue.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!K-_1!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6df8a002-7fa0-40c9-84c4-8877539496ec_2320x2122.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!K-_1!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6df8a002-7fa0-40c9-84c4-8877539496ec_2320x2122.png 424w, https://substackcdn.com/image/fetch/$s_!K-_1!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6df8a002-7fa0-40c9-84c4-8877539496ec_2320x2122.png 848w, https://substackcdn.com/image/fetch/$s_!K-_1!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6df8a002-7fa0-40c9-84c4-8877539496ec_2320x2122.png 1272w, https://substackcdn.com/image/fetch/$s_!K-_1!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6df8a002-7fa0-40c9-84c4-8877539496ec_2320x2122.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!K-_1!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6df8a002-7fa0-40c9-84c4-8877539496ec_2320x2122.png" width="1456" height="1332" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6df8a002-7fa0-40c9-84c4-8877539496ec_2320x2122.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1332,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:444890,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/206014241?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6df8a002-7fa0-40c9-84c4-8877539496ec_2320x2122.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!K-_1!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6df8a002-7fa0-40c9-84c4-8877539496ec_2320x2122.png 424w, https://substackcdn.com/image/fetch/$s_!K-_1!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6df8a002-7fa0-40c9-84c4-8877539496ec_2320x2122.png 848w, https://substackcdn.com/image/fetch/$s_!K-_1!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6df8a002-7fa0-40c9-84c4-8877539496ec_2320x2122.png 1272w, https://substackcdn.com/image/fetch/$s_!K-_1!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6df8a002-7fa0-40c9-84c4-8877539496ec_2320x2122.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>References</h2><p>Primary and background sources. <strong>arXiv identifiers</strong> for the Kimi Linear and <strong>Gated DeltaNet </strong>entries were verified against arXiv during fact-checking; the remaining identifiers are standard citations for their papers.</p><ol><li><p><em>Moonshot AI. Kimi K3 Tech Blog: Open Frontier Intelligence. kimi.com/blog/kimi-k3, July 2026.</em></p></li><li><p><em>Moonshot AI. Kimi K3 quickstart and platform pricing. platform.kimi.ai, July 2026.</em></p></li><li><p><em>Kimi Team. Kimi Linear: An Expressive, Efficient Attention Architecture. arXiv:2510.26692, 2025.</em></p></li><li><p><em>Yang, S. et al. Gated Delta Networks: Improving Mamba2 with Delta Rule. arXiv:2412.06464, 2024.</em></p></li><li><p><em>Yang, S. et al. Parallelizing Linear Transformers with the Delta Rule over Sequence Length. arXiv:2406.06484, 2024.</em></p></li><li><p><em>DeepSeek-AI. DeepSeek-V3 Technical Report. arXiv:2412.19437, 2024.</em></p></li><li><p><em>DeepSeek-AI. DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model. arXiv:2405.04434, 2024.</em></p></li><li><p><em>Qin, R. et al. Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving. arXiv:2407.00079; USENIX FAST 2025, best paper.</em></p></li><li><p><em>Rouhani, B. et al. Microscaling Data Formats for Deep Learning. arXiv:2310.10537, 2023. See also the OCP Microscaling Formats Specification v1.0.</em></p></li><li><p><em>Micikevicius, P. et al. FP8 Formats for Deep Learning. arXiv:2209.05433, 2022.</em></p></li><li><p><em>Shazeer, N. et al. Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer. arXiv:1701.06538, 2017.</em></p></li><li><p><em>Lepikhin, D. et al. GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding. arXiv:2006.16668, 2020.</em></p></li><li><p><em>Fedus, W. et al. Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity. arXiv:2101.03961, 2021.</em></p></li><li><p><em>Wang, L. et al. Auxiliary-Loss-Free Load Balancing Strategy for Mixture-of-Experts. arXiv:2408.15664, 2024.</em></p></li><li><p><em>Kimi Team. Kimi K2: Open Agentic Intelligence. arXiv:2507.20534, 2025.</em></p></li><li><p><em>Liu, J. et al. Muon is Scalable for LLM Training. arXiv:2502.16982, 2025.</em></p></li><li><p><em>Ainslie, J. et al. GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints. arXiv:2305.13245, 2023.</em></p></li><li><p><em>Kwon, W. et al. Efficient Memory Management for Large Language Model Serving with PagedAttention. arXiv:2309.06180, 2023.</em></p></li><li><p><em>Zheng, L. et al. SGLang: Efficient Execution of Structured Language Model Programs. arXiv:2312.07104, 2023.</em></p></li><li><p><em>Open-source repositories referenced: MoonshotAI/Kimi-Linear, fla-org/flash-linear-attention, deepseek-ai/DeepEP, deepseek-ai/FlashMLA, kvcache-ai/Mooncake, on GitHub.</em></p></li><li><p><em>Launch-week coverage and trackers cited in text: VentureBeat, the Hugging Face community overview, Artificial Analysis, OpenRouter, BenchLM, and the pricing and readiness guides from Verdent, eesel, aireiter, kie.ai, avenchat, wan27.org, NxCode, digitalapplied, and TECHi, all July 2026.</em></p></li></ol>]]></content:encoded></item><item><title><![CDATA[Exploring how MLIR works: the compiler rewiring the AI stack ]]></title><description><![CDATA[From tensor graphs to machine code, MLIR is quietly becoming the abstraction layer connecting modern AI frameworks to increasingly specialized hardware.]]></description><link>https://www.thesoftwarefrontier.com/p/exploring-how-mlir-works-the-compiler</link><guid isPermaLink="false">https://www.thesoftwarefrontier.com/p/exploring-how-mlir-works-the-compiler</guid><dc:creator><![CDATA[Lorenzo Bradanini]]></dc:creator><pubDate>Wed, 15 Jul 2026 17:34:55 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!auv1!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F069655fa-6217-4d2d-8889-c97eba300015_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!auv1!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F069655fa-6217-4d2d-8889-c97eba300015_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!auv1!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F069655fa-6217-4d2d-8889-c97eba300015_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!auv1!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F069655fa-6217-4d2d-8889-c97eba300015_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!auv1!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F069655fa-6217-4d2d-8889-c97eba300015_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!auv1!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F069655fa-6217-4d2d-8889-c97eba300015_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!auv1!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F069655fa-6217-4d2d-8889-c97eba300015_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/069655fa-6217-4d2d-8889-c97eba300015_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2337295,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/203533021?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F069655fa-6217-4d2d-8889-c97eba300015_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!auv1!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F069655fa-6217-4d2d-8889-c97eba300015_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!auv1!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F069655fa-6217-4d2d-8889-c97eba300015_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!auv1!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F069655fa-6217-4d2d-8889-c97eba300015_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!auv1!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F069655fa-6217-4d2d-8889-c97eba300015_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><p><strong>Qualcomm</strong> just agreed to pay roughly $3.9 billion for a 150 person compiler company. Tesla says it rebuilt the FSD compiler and runtime on the same infrastructure. NVIDIA&#8217;s newest Python kernel DSL compiles through it. That infrastructure is <strong>MLIR</strong>. </p><p>This is the full tour: the object model, the dialect system, one<strong> MLP block lowered</strong> step by step into genuine <strong>sm_90 PTX</strong>, a custom dialect built from scratch, and the economics of who now controls the narrowest point in the AI software stack.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Three headlines, one substrate</h2><p>On June 24, 2026, Qualcomm announced it would acquire <strong>Modular</strong>, the AI infrastructure company founded in 2022 by Chris Lattner and Tim Davis. </p><p>Qualcomm&#8217;s own press release gives no price; Reuters and the Wall Street Journal reported an all-stock deal worth <em>approximately $3.9 billion</em>, expected to close in the second half of 2026. Modular employs roughly 150 people, per the industry analysts at <strong>NAND Research. </strong></p><p>Do the arithmetic and Qualcomm is paying something like $26 million per person for a company whose flagship language, <strong>Mojo</strong>, reached its first 1.0 beta seven weeks before the announcement, and whose compiler is not even open source yet. <strong>Qualcomm CEO</strong> Cristiano Amon justified it in one line: developers, he said, &#8220;<em>demand a more open and modern software foundation.&#8221;</em></p><p>Two months earlier, in April 2026, <strong>Tesla shipped FSD v14.3.</strong> Buried in the release notes, as documented by the release trackers at Tesla Oracle, was the claim that the AI compiler and runtime had been rewritten &#8220;<em>from the ground up with MLIR</em>&#8221;, with Tesla attributing a <strong>20 percent improvement</strong> in vehicle reaction time to the new stack. </p><p>And a year before that, in May 2025, NVIDIA shipped CUTLASS 4 with something unprecedented for the company: <strong>CuTe DSL</strong>, a Python interface for authoring peak-performance GPU kernels which, per NVIDIA&#8217;s own documentation, are <strong>JIT compiled</strong> through MLIR and then handed to ptxas.</p><p>A $3.9 billion acquisition by a mobile silicon giant. A safety-critical automotive inference stack. The kernel library that NVIDIA itself uses to demonstrate what Blackwell can do. </p><p>Three announcements, three companies with three entirely different agendas, one common substrate underneath. That<strong> substrate is MLIR</strong>, and if you work anywhere near GPUs, inference serving, or the economics of accelerated compute, you are running on top of it whether you know it or not.</p><p>In the previous installment of this series, <em>How LLVM Works: The IR That Took Over Modern Computing</em>, we told the story of the intermediate representation that <strong>unified the CPU world</strong>: one typed SSA language in the middle, every source language on one side, every instruction set on the other. </p><p>This piece is about what happened when that model met machine learning and broke, and about the infrastructure Lattner&#8217;s team built in response. MLIR is routinely described as &#8220;<em>LLVM for AI.</em>&#8221; It is not. </p><p>It is something stranger and <strong>more consequential</strong>: not a compiler, but a machine for building compilers, and it has quietly become the coordination point for the entire AI hardware buildout.</p><h4><span>How to read the code in this piece</span></h4><p><em>Every block of IR, TableGen, and C++ below was run through real tools before publication: </em><code>mlir-opt</code><em>, </em><code>mlir-runner</code><em>, </em><code>mlir-tblgen</code><em>, and GCC against the MLIR 20 headers, all from the stock LLVM 20.1.2 packages on Ubuntu 24.04. </em></p><p>A <strong><span>verified</span></strong> badge means the snippet parses and passes the verifier; <strong><span>executed</span></strong> means we compiled it to native code and ran it, and the output you see is what the machine printed. </p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/p/exploring-how-mlir-works-the-compiler?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/p/exploring-how-mlir-works-the-compiler?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><div><hr></div><h2>The problem LLVM could not see</h2><p><strong>LLVM IR&#8217;s</strong> contract, the one that conquered the CPU world, is deliberately narrow: scalar values in <strong>SSA form</strong>, a flat control flow graph of basic blocks, loads and stores against a flat memory model, one abstraction level for everything. </p><p>That narrowness is why<em> thirty-odd frontends </em>and every serious instruction set could meet in the middle. It is also precisely what makes LLVM IR the wrong meeting point for machine learning.</p><p>Consider what a tensor compiler needs to know about a matrix multiplication. </p><ul><li><p><em>That the two input buffers are <strong>dense 2-D arrays </strong>with known strides. </em></p></li><li><p><em>That the <strong>triple loop nest </strong>around it is two parallel dimensions and one reduction. </em></p></li><li><p><em>That the accumulator is <strong>transient </strong>and could live in registers or shared memory. </em></p></li><li><p><em>That the whole operation is <strong>one node in a graph </strong>whose neighbors it might profitably fuse with. </em></p></li></ul><p>By the time a matmul has been lowered to LLVM IR, every one of those facts has been destroyed. The<strong> arrays are opaque pointers</strong>, the loop structure is a thicket of compare-and-branch, the parallelism is gone, and the fusion opportunity evaporated three layers up. </p><p>Recovering that information from LLVM IR is the <strong>auto-vectorization problem</strong>, a research area with forty years of partial credit. The lesson the ML compiler generation drew was blunt: do not recover structure. Never lose it.</p><p>The second problem was <strong>combinatorial</strong> rather than semantic. Lattner has told the origin story himself, most recently in January 2026 in &#8220;<em>Democratizing AI Compute, Part 8</em>,&#8221; on Modular&#8217;s blog: <strong>MLIR began in 2018 inside Google</strong>, in the middle of the TPU buildout, when the TensorFlow ecosystem had accumulated a zoo of graph representations, converters, and codegen paths that shared almost nothing. </p><p>Every framework wanted to reach every accelerator. Every accelerator vendor was rebuilding the same 80 percent of a compiler: parsers, printers, pass managers, verifiers, canonicalizers, testing harnesses. </p><p>With N frontends and M targets, the industry was signing up to build and staff something like <strong>N times M compilers</strong>.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!-2xT!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d0c6a4e-aaa8-4dd0-87bf-5b3ec002986f_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!-2xT!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d0c6a4e-aaa8-4dd0-87bf-5b3ec002986f_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!-2xT!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d0c6a4e-aaa8-4dd0-87bf-5b3ec002986f_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!-2xT!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d0c6a4e-aaa8-4dd0-87bf-5b3ec002986f_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!-2xT!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d0c6a4e-aaa8-4dd0-87bf-5b3ec002986f_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!-2xT!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d0c6a4e-aaa8-4dd0-87bf-5b3ec002986f_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/5d0c6a4e-aaa8-4dd0-87bf-5b3ec002986f_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Point to point lowering paths versus a shared waist&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Point to point lowering paths versus a shared waist" title="Point to point lowering paths versus a shared waist" srcset="https://substackcdn.com/image/fetch/$s_!-2xT!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d0c6a4e-aaa8-4dd0-87bf-5b3ec002986f_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!-2xT!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d0c6a4e-aaa8-4dd0-87bf-5b3ec002986f_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!-2xT!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d0c6a4e-aaa8-4dd0-87bf-5b3ec002986f_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!-2xT!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d0c6a4e-aaa8-4dd0-87bf-5b3ec002986f_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><em>Figure 1. The arithmetic that motivated MLIR. Five frontends and six target families connected point to point is 30 lowering paths, each a compiler someone must build and maintain. Routed through a shared waist it is 11. LLVM proved the hourglass works for CPUs; MLIR generalizes the hourglass itself.</em></figcaption></figure></div><p>The insight that became MLIR was to go one level more abstract than LLVM did. LLVM&#8217;s answer to fragmentation was a single <strong>fixed IR</strong> in the middle. </p><p>But ML compilation does <strong>not have one natural middle</strong>: it has a graph level, a tensor algebra level, a loop level, a vector level, a hardware intrinsic level, and programs need to descend through all of them. </p><p>So instead of shipping another fixed IR, ship the infrastructure that makes intermediate representations cheap to build, and let them all coexist in one framework, one type system, one pass manager, one textual format. </p><p>Google open sourced the project in April 2019 and donated it to the LLVM Foundation, with the code landing in the LLVM monorepo at the end of that year. </p><p>The design paper, &#8220;<em>MLIR: Scaling Compiler Infrastructure for Domain Specific Computation</em>&#8221; by Lattner, Mehdi Amini, Uday Bondhugula, Albert Cohen, and colleagues, appeared at <strong>CGO 2021</strong> and is still the best single description of the architecture.</p><p>The name expands to <strong>Multi-Level Intermediate Representation</strong>, and the plural is the entire point. One artifact can hold, simultaneously and legally, a tensor-algebra view of one function, an explicit loop nest view of another, and <strong>raw hardware intrinsics</strong> for a third, with every level checked by the same verifier and transformed by the same pass infrastructure. </p><p>The rest of this article is a tour of how that actually works, at the level of real IR, because the abstractions only make sense once you have watched them move.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share The Software Frontier&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share The Software Frontier</span></a></p><div><hr></div><h2>One idea, applied without exception</h2><p>Strip away everything else and MLIR is one data structure applied with unusual discipline. The universe contains <strong>operations</strong>. </p><p>An operation has a name, takes SSA <strong>values</strong> as operands, produces values as results, carries compile-time constants called <strong>attributes</strong>, and, this is the structural leap, may contain <strong>regions</strong>: nested lists of blocks which themselves contain operations. Values have <strong>types</strong>. </p><p>That is the whole object model. There are no statements, no expressions, no special forms. Here is the smallest interesting function, in the friendly syntax you will see in every MLIR dump:</p><p>01_axpy.mlir <strong><span>verified: mlir-opt 20.1.2</span></strong></p><pre><code><code>func.func @axpy(%a: f32, %x: f32, %y: f32) -&gt; f32 {
  %0 = arith.mulf %a, %x : f32
  %1 = arith.addf %0, %y : f32
  return %1 : f32
}</code></code></pre><p>Friendly syntax is sugar. Ask <code>mlir-opt</code> to print the same module in generic form and the uniformity becomes visible:</p><p>mlir-opt --mlir-print-op-generic 01_axpy.mlir <strong><span>verified: actual tool output</span></strong></p><pre><code><code>"builtin.module"() ({
  "func.func"() &lt;{function_type = (f32, f32, f32) -&gt; f32, sym_name = "axpy"}&gt; ({
  ^bb0(%arg0: f32, %arg1: f32, %arg2: f32):
    %0 = "arith.mulf"(%arg0, %arg1) &lt;{fastmath = #arith.fastmath&lt;none&gt;}&gt; : (f32, f32) -&gt; f32
    %1 = "arith.addf"(%0, %arg2) &lt;{fastmath = #arith.fastmath&lt;none&gt;}&gt; : (f32, f32) -&gt; f32
    "func.return"(%1) : (f32) -&gt; ()
  }) : () -&gt; ()
}) : () -&gt; ()</code></code></pre><p>Read that carefully, because it is the <strong>most important dump</strong> in this article. The module is an operation named <code>builtin.module</code> with one region. The function is an operation named <code>func.func</code> whose &#8220;<em>body</em>&#8221; is just a region and whose name and signature are ordinary attributes. </p><p>Even <code>return</code> is an operation. There are no built-in language constructs at all; the things that look built in live in a dialect that happens to be called <code>builtin</code>. Every structure you will meet from here on, loop nests, <strong>GPU kernels</strong>, entire schedules, is an operation containing regions containing operations, all the way down.</p><p>Regions are what LLVM IR never had. LLVM gives you <strong>one flat control flow graph per function</strong>, so a loop exists only as a cycle among basic blocks, something analyses must rediscover with dominator trees. </p><p>In MLIR, an operation like a loop <em>contains</em> its body as a region, so the nesting of the <strong>source program </strong>survives as nesting in the IR. A GPU kernel is an operation containing the kernel body. A module is an operation containing functions. </p><p>SSA and dominance still hold inside each region, so classic optimization theory transfers intact, but structure is now first class rather than archaeological.</p><p>Discipline this strict pays off in the machinery. Because everything is an operation, one parser, one printer, one pass manager, one multithreading strategy, and <strong>one verification framework</strong> serve every abstraction level ever built on MLIR, including ones that do not exist yet. </p><p>Extensibility is kept honest by three mechanisms. </p><ul><li><p><em><strong>Verifiers</strong> let each operation enforce its own structural invariants, so malformed IR dies at the boundary instead of corrupting a pass 40 steps later. </em></p></li><li><p><em><strong>Traits</strong> declare properties like &#8220;this op has no side effects&#8221; or &#8220;operands and results share a type.&#8221; </em></p></li><li><p><em>And <strong>interfaces</strong> are the load-bearing wall: a pass written against </em><code>LoopLikeOpInterface</code><em> or </em><code>DestinationStyleOpInterface</code><em> works on any operation implementing the interface, including operations invented years after the pass was written by people who never met the pass author. </em></p></li></ul><p>That is the property that makes what follows possible: generic infrastructure over an open universe of abstractions.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Dialects: a periodic table you can extend</h2><p>Operations are grouped into <strong>dialects</strong>: namespaces that bundle related ops, types, and attributes together with their verifiers, canonicalization patterns, and documentation. </p><p>The prefix before the dot tells you the dialect: <code>arith.mulf</code>, <code>func.func</code>, <code>linalg.matmul</code>. A <strong>dialect is not a language</strong>; it is closer to a library of IR vocabulary, and dialects mix freely in one function. </p><p>The stock distribution is already a small civilization:</p><p>the in-tree roster <strong><span>verified: actual tool output</span></strong></p><pre><code><code>$ mlir-opt --show-dialects
Available Dialects (48):
  acc, affine, amdgpu, amx, arith, arm_neon, arm_sme, arm_sve, async,
  bufferization, builtin, cf, complex, dlti, emitc, func, gpu, index,
  irdl, linalg, llvm, math, memref, mesh, ml_program, mpi, nvgpu,
  nvvm, omp, pdl, pdl_interp, polynomial, ptr, quant, rocdl, scf,
  shape, sparse_tensor, spirv, tensor, test, test_dyn, tosa,
  transform, ub, vector, x86vector, xegpu</code></code></pre><p>Forty-eight dialects ship in <strong>LLVM 20.1.2</strong>, and the count only grows: CPU vector extensions (<code>amx</code>, <code>arm_sve</code>, <code>x86vector</code>), GPU vendor exits (<code>nvvm</code>, <code>rocdl</code>, <code>xegpu</code>), parallel runtimes (<code>omp</code>, <code>async</code>), quantization, sparse tensors, hardware description. </p><p>Out of tree, the population explodes; we will meet the important settlers in a few sections. What matters is the shape of the traffic: programs enter at high abstraction and descend. A few landmarks, each verified:</p><p><code>scf</code><strong>, structured control flow.</strong> Loops and conditionals as region-holding operations rather than branch spaghetti. Loop-carried state is explicit SSA, threaded through <code>iter_args</code>:</p><p>dot product in scf <strong><span>verified: mlir-opt 20.1.2</span></strong></p><pre><code><code>// scf: structured control flow with loop-carried values
func.func @dot(%a: memref&lt;1024xf32&gt;, %b: memref&lt;1024xf32&gt;) -&gt; f32 {
  %c0 = arith.constant 0 : index
  %c1 = arith.constant 1 : index
  %n  = arith.constant 1024 : index
  %zero = arith.constant 0.0 : f32
  %sum = scf.for %i = %c0 to %n step %c1
         iter_args(%acc = %zero) -&gt; (f32) {
    %x = memref.load %a[%i] : memref&lt;1024xf32&gt;
    %y = memref.load %b[%i] : memref&lt;1024xf32&gt;
    %m = arith.mulf %x, %y : f32
    %s = arith.addf %acc, %m : f32
    scf.yield %s : f32
  }
  return %sum : f32
}</code></code></pre><p><code>affine</code><strong>, the polyhedral inheritance.</strong> Same loops, harsher rules: bounds and subscripts must be affine expressions of the induction variables. In exchange, dependence analysis becomes decidable. </p><p>The compiler can <em>prove</em> that iterations are independent, that tiling is legal, that this access pattern never aliases that one, instead of guessing:</p><p>1-D stencil in affine <strong><span>verified: mlir-opt 20.1.2</span></strong></p><pre><code><code>// affine: loops + accesses restricted to affine expressions, so the
// compiler can *prove* dependence facts instead of guessing
func.func @blur1d(%in: memref&lt;1026xf32&gt;, %out: memref&lt;1024xf32&gt;) {
  affine.for %i = 0 to 1024 {
    %l = affine.load %in[%i]     : memref&lt;1026xf32&gt;
    %c = affine.load %in[%i + 1] : memref&lt;1026xf32&gt;
    %r = affine.load %in[%i + 2] : memref&lt;1026xf32&gt;
    %s0 = arith.addf %l, %c : f32
    %s1 = arith.addf %s0, %r : f32
    affine.store %s1, %out[%i] : memref&lt;1024xf32&gt;
  }
  return
}</code></code></pre><p><code>tensor</code><strong> and </strong><code>memref</code><strong>, the two memories.</strong> This pair encodes the single most useful distinction in the whole stack. </p><p>A <code>tensor</code> is an immutable SSA <em>value</em>: no address, no aliasing, safe to reorder and fuse aggressively, the natural currency of ML graphs. </p><p>A <code>memref</code> is a <em>view of actual memory</em>: base pointer, sizes, strides, and, critically for this audience, an address space. The types are expressive enough to carry real kernel engineering:</p><p>types that will look familiar <strong><span>verified: mlir-opt 20.1.2</span></strong></p><pre><code><code>func.func private @smem_tile(%t: memref&lt;64x64xf16, strided&lt;[68, 1]&gt;, 3&gt;)
func.func private @kv_page(%p: tensor&lt;16x8x128xbf16&gt;)</code></code></pre><p>The first is a <strong>64x64 half-precision tile </strong>with a padded row stride of 68 elements, the classic plus-four skew that kills shared memory bank conflicts, living in address space 3: GPU shared memory. </p><p>The second is a bf16 paged-KV block of the kind every serving engine slings around. MLIR&#8217;s type system says these things natively; nothing here is a comment or a convention.</p><p><code>vector</code><strong>, the portable SIMD layer.</strong> Fixed-shape virtual registers and the operations that move them: <code>vector.fma</code>, transfers, contractions, shuffles. Target independent until the last moment, then peeled into NEON, AVX-512, or GPU instructions:</p><p>an 8-wide fused multiply-add <strong><span>verified: mlir-opt 20.1.2</span></strong></p><pre><code><code>// vector: SIMD/SIMT-width-explicit compute on virtual registers
func.func @vec_fma(%a: vector&lt;8xf32&gt;, %b: vector&lt;8xf32&gt;, %c: vector&lt;8xf32&gt;) -&gt; vector&lt;8xf32&gt; {
  %0 = vector.fma %a, %b, %c : vector&lt;8xf32&gt;
  return %0 : vector&lt;8xf32&gt;
}</code></code></pre><p>And beneath everything, the exits: the <code>gpu</code> dialect for launch grids and kernels, <code>nvvm</code> and <code>rocdl</code> wrapping the vendors&#8217; intrinsics, and the <code>llvm</code> dialect, <strong>which models LLVM IR</strong> itself as just another MLIR dialect so that the handoff to the old empire is an ordinary rewrite rather than a file format negotiation. </p><h4>One table to orient the descent:</h4><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!0WfX!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F71f5b704-e895-41f5-afb5-e3d0f05a62b0_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!0WfX!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F71f5b704-e895-41f5-afb5-e3d0f05a62b0_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!0WfX!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F71f5b704-e895-41f5-afb5-e3d0f05a62b0_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!0WfX!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F71f5b704-e895-41f5-afb5-e3d0f05a62b0_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!0WfX!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F71f5b704-e895-41f5-afb5-e3d0f05a62b0_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!0WfX!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F71f5b704-e895-41f5-afb5-e3d0f05a62b0_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/71f5b704-e895-41f5-afb5-e3d0f05a62b0_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:177831,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/203533021?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F71f5b704-e895-41f5-afb5-e3d0f05a62b0_1634x893.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!0WfX!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F71f5b704-e895-41f5-afb5-e3d0f05a62b0_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!0WfX!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F71f5b704-e895-41f5-afb5-e3d0f05a62b0_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!0WfX!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F71f5b704-e895-41f5-afb5-e3d0f05a62b0_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!0WfX!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F71f5b704-e895-41f5-afb5-e3d0f05a62b0_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Read the<strong> right-hand column top</strong> to bottom and you have the core operating principle of the whole system, the one we will watch in motion next:<em> information is only destroyed downward, so every optimization should run at the highest level where its enabling facts still exist. </em></p><p>Fusion happens in tensor algebra, where aliasing cannot exist. Tiling happens where loops are structured objects. Bank conflict avoidance happens where address spaces are types. </p><p><em>By the time you reach LLVM, you are not optimizing anymore; you are translating.</em></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/p/exploring-how-mlir-works-the-compiler?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/p/exploring-how-mlir-works-the-compiler?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><div><hr></div><h2>linalg: the algebra at the waist</h2><p>If MLIR has a center of gravity, it is the <code>linalg</code> dialect: the tensor-algebra layer where most serious codegen pipelines, <strong>IREE&#8217;s</strong>, Modular&#8217;s, the upstream CPU flows, do their thinking. </p><p>Its design principle is the anti-autovectorizer stance taken to its conclusion. Instead of <strong>encoding a matmul as loops</strong> and hoping to rediscover its nature, <code>linalg</code> encodes the <em>nature</em> and treats loops as one of several possible spellings, generated on demand.</p><p>The workhorse is <code>linalg.generic</code>, which describes a perfectly nested computation with three pieces of data. First, <strong>indexing maps</strong>: one affine map per operand saying which element each point in the iteration space touches. </p><p>Second, <strong>iterator types</strong>: each loop dimension is declared <code>parallel</code> or <code>reduction</code>, which is the fact autovectorizers spend their lives failing to prove. Third, a region holding the scalar payload, the body computed at every point. </p><p>Everything the dialect knows, it knows structurally. A matmul&#8217;s maps are <code>(m, n, k) -&gt; (m, k)</code>, <code>(m, n, k) -&gt; (k, n)</code>, <code>(m, n, k) -&gt; (m, n)</code> over iterators <code>[parallel, parallel, reduction]</code>; named ops like <code>linalg.matmul</code> are simply memorable spellings of specific generics. </p><p>Here is the <strong>payload </strong>we will spend the rest of the article compiling, a real inference fragment: a linear layer with fused bias and ReLU.</p><p>02_mlp.mlir: matmul + fused bias/ReLU on tensors <strong><span>verified: mlir-opt 20.1.2</span></strong></p><pre><code><code>#id2d   = affine_map&lt;(m, n) -&gt; (m, n)&gt;
#bcast  = affine_map&lt;(m, n) -&gt; (n)&gt;

func.func @mlp_block(%x: tensor&lt;64x512xf32&gt;,
                     %w: tensor&lt;512x256xf32&gt;,
                     %b: tensor&lt;256xf32&gt;) -&gt; tensor&lt;64x256xf32&gt; {
  %c0 = arith.constant 0.0 : f32

  // Materialize the destination and zero-init the accumulator.
  %empty = tensor.empty() : tensor&lt;64x256xf32&gt;
  %acc   = linalg.fill ins(%c0 : f32)
                       outs(%empty : tensor&lt;64x256xf32&gt;) -&gt; tensor&lt;64x256xf32&gt;

  // y = x @ w
  %mm = linalg.matmul ins(%x, %w : tensor&lt;64x512xf32&gt;, tensor&lt;512x256xf32&gt;)
                      outs(%acc : tensor&lt;64x256xf32&gt;) -&gt; tensor&lt;64x256xf32&gt;

  // out = max(y + b, 0), one fused elementwise op
  %out = linalg.generic
           {indexing_maps = [#id2d, #bcast, #id2d],
            iterator_types = ["parallel", "parallel"]}
           ins(%mm, %b : tensor&lt;64x256xf32&gt;, tensor&lt;256xf32&gt;)
           outs(%empty : tensor&lt;64x256xf32&gt;) {
  ^bb0(%y: f32, %bias: f32, %o: f32):
    %sum  = arith.addf %y, %bias : f32
    %relu = arith.maximumf %sum, %c0 : f32
    linalg.yield %relu : f32
  } -&gt; tensor&lt;64x256xf32&gt;

  return %out : tensor&lt;64x256xf32&gt;
}</code></code></pre><p>Three details deserve attention. The bias&#8217;s indexing map, <code>(m, n) -&gt; (n)</code>, expresses broadcasting as pure index algebra: no replicated buffer, no runtime rule, just a map that ignores <code>m</code>. </p><p>The elementwise op fuses the add and the clamp into one region, because a region can hold any scalar computation, which is how epilogue fusion is spelled at this level. </p><p>And every linalg op writes into an explicit <code>outs</code> operand it also returns: <strong>destination-passing style</strong>. On immutable tensors the destination looks redundant, and semantically it is. </p><p>It exists as a standing declaration of where results <em>could</em> live, which is exactly the hint that lets bufferization later plan in-place updates instead of solving a global aliasing puzzle. Keep an eye on that <code>%empty</code>; it is about to earn its keep.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share The Software Frontier&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share The Software Frontier</span></a></p><div><hr></div><h2>The walkthrough: one MLP block to sm_90 PTX</h2><p>Now the demonstration the whole article is built around. We will take that block and walk it down the ladder one verified step at a time, watching what each level adds and what it forgets. </p><p>This is the same journey your <strong>PyTorch graphs</strong> take through Triton, that JAX programs take through XLA, that Mojo kernels take through MAX; we are just taking it on foot, with the intermediate files open.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!osL7!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9ceb5aa2-3004-40b4-8e18-ec62404d85ca_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!osL7!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9ceb5aa2-3004-40b4-8e18-ec62404d85ca_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!osL7!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9ceb5aa2-3004-40b4-8e18-ec62404d85ca_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!osL7!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9ceb5aa2-3004-40b4-8e18-ec62404d85ca_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!osL7!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9ceb5aa2-3004-40b4-8e18-ec62404d85ca_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!osL7!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9ceb5aa2-3004-40b4-8e18-ec62404d85ca_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/9ceb5aa2-3004-40b4-8e18-ec62404d85ca_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;The lowering ladder walked in this article&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="The lowering ladder walked in this article" title="The lowering ladder walked in this article" srcset="https://substackcdn.com/image/fetch/$s_!osL7!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9ceb5aa2-3004-40b4-8e18-ec62404d85ca_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!osL7!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9ceb5aa2-3004-40b4-8e18-ec62404d85ca_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!osL7!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9ceb5aa2-3004-40b4-8e18-ec62404d85ca_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!osL7!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9ceb5aa2-3004-40b4-8e18-ec62404d85ca_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><em>Figure 2. The descent this section performs, every stage machine-verified. The annotations name the decision each level owns; the italic line at the bottom is the operating principle from section 4.</em></figcaption></figure></div><h3>Step 1. Tiling, written as IR</h3><p>The first performance decision on <strong>any matrix multiply</strong> is blocking it for the memory hierarchy. In most compilers, tile sizes hide inside a C++ heuristic. </p><p>In MLIR, the transformation itself can be scripted in IR, using the <code>transform</code> dialect: a schedule language whose operations manipulate <em>other operations</em>. We append this to the file:</p><p>03_tile.mlir, the schedule part <strong><span>verified: mlir-opt 20.1.2</span></strong></p><pre><code><code>// ---- Schedule: also IR, in the transform dialect ----
module attributes {transform.with_named_sequence} {
  transform.named_sequence @__transform_main(%root: !transform.any_op {transform.readonly}) {
    %mm = transform.structured.match ops{["linalg.matmul"]} in %root
      : (!transform.any_op) -&gt; !transform.any_op
    %tiled, %loops:3 = transform.structured.tile_using_for %mm tile_sizes [16, 32, 64]
      : (!transform.any_op) -&gt; (!transform.any_op, !transform.any_op, !transform.any_op, !transform.any_op)
    transform.yield
  }
}</code></code></pre><p>The handles typed <code>!transform.any_op</code> are <strong>SSA values</strong> whose runtime contents are payload operations: the schedule matches every <code>linalg.matmul</code>, then tiles each one 16 by 32 by 64. </p><p>Run <code>mlir-opt --transform-interpreter</code> and the payload function comes back rewritten:</p><p>output: the tiled matmul <strong><span>verified: actual tool output, constants elided as marked</span></strong></p><pre><code><code>  func.func @matmul(%arg0: tensor&lt;64x512xf32&gt;, %arg1: tensor&lt;512x256xf32&gt;, %arg2: tensor&lt;64x256xf32&gt;) -&gt; tensor&lt;64x256xf32&gt; {
    // ... index constants (0, 16, 32, 64, 256, 512) elided ...
    %0 = scf.for %arg3 = %c0 to %c64 step %c16 iter_args(%arg4 = %arg2) -&gt; (tensor&lt;64x256xf32&gt;) {
      %1 = scf.for %arg5 = %c0_0 to %c256 step %c32 iter_args(%arg6 = %arg4) -&gt; (tensor&lt;64x256xf32&gt;) {
        %2 = scf.for %arg7 = %c0_1 to %c512 step %c64_2 iter_args(%arg8 = %arg6) -&gt; (tensor&lt;64x256xf32&gt;) {
          %extracted_slice = tensor.extract_slice %arg0[%arg3, %arg7] [16, 64] [1, 1] : tensor&lt;64x512xf32&gt; to tensor&lt;16x64xf32&gt;
          %extracted_slice_3 = tensor.extract_slice %arg1[%arg7, %arg5] [64, 32] [1, 1] : tensor&lt;512x256xf32&gt; to tensor&lt;64x32xf32&gt;
          %extracted_slice_4 = tensor.extract_slice %arg8[%arg3, %arg5] [16, 32] [1, 1] : tensor&lt;64x256xf32&gt; to tensor&lt;16x32xf32&gt;
          %3 = linalg.matmul ins(%extracted_slice, %extracted_slice_3 : tensor&lt;16x64xf32&gt;, tensor&lt;64x32xf32&gt;) outs(%extracted_slice_4 : tensor&lt;16x32xf32&gt;) -&gt; tensor&lt;16x32xf32&gt;
          %inserted_slice = tensor.insert_slice %3 into %arg8[%arg3, %arg5] [16, 32] [1, 1] : tensor&lt;16x32xf32&gt; into tensor&lt;64x256xf32&gt;
          scf.yield %inserted_slice : tensor&lt;64x256xf32&gt;
        }
        scf.yield %2 : tensor&lt;64x256xf32&gt;
      }
      scf.yield %1 : tensor&lt;64x256xf32&gt;
    }
    return %0 : tensor&lt;64x256xf32&gt;
  }</code></code></pre><p>One schedule op became a three-deep <code>scf.for</code> nest. <code>tensor.extract_slice</code> carves 16x64 and 64x32 tiles from the operands, an inner <code>linalg.matmul</code>, still the full structured op, now shaped <strong>16x64 times 64x32</strong>, computes each 16x32 block, and <code>tensor.insert_slice</code> threads the result through the loop-carried tensor. </p><p>Notice what did <em>not</em> happen: we have loops, but we did not lose the algebra. The inner op is <strong>still a matmul with its maps</strong> and iterator types intact, still available for vectorization, for lowering to tensor core intrinsics, or for another round of tiling. </p><p>Structure survives the transformation because the transformation is structure-aware.</p><h3>Step 2. Fusion: FlashAttention thinking, six lines long</h3><p>Tiling one op is table stakes. The move that defines modern kernel engineering, compute the producer inside the consumer&#8217;s tile so intermediates never <strong>round-trip through HBM</strong>, is the same idea that makes FlashAttention FlashAttention. </p><p>Here it is as a schedule: tile the bias/ReLU consumer into a parallel grid, then pull the matmul into each tile.</p><p>04_fuse.mlir, the schedule part <strong><span>verified: mlir-opt 20.1.2</span></strong></p><pre><code><code>module attributes {transform.with_named_sequence} {
  transform.named_sequence @__transform_main(%root: !transform.any_op {transform.readonly}) {
    // 1. Tile the *consumer* (bias+relu) into a parallel grid of 16x32 tiles.
    %gen = transform.structured.match ops{["linalg.generic"]} in %root
      : (!transform.any_op) -&gt; !transform.any_op
    %tiled, %grid = transform.structured.tile_using_forall %gen tile_sizes [16, 32]
      : (!transform.any_op) -&gt; (!transform.any_op, !transform.any_op)

    // 2. Pull the matmul producer inside each tile of that grid.
    %mm = transform.structured.match ops{["linalg.matmul"]} in %root
      : (!transform.any_op) -&gt; !transform.any_op
    %fused, %grid2 = transform.structured.fuse_into_containing_op %mm into %grid
      : (!transform.any_op, !transform.any_op) -&gt; (!transform.any_op, !transform.any_op)
    transform.yield
  }
}</code></code></pre><p>output after --transform-interpreter -canonicalize -cse <strong><span>verified: actual tool output</span></strong></p><pre><code><code>#map = affine_map&lt;(d0) -&gt; (d0 * 16)&gt;
#map1 = affine_map&lt;(d0) -&gt; (d0 * 32)&gt;
#map2 = affine_map&lt;(d0, d1) -&gt; (d0, d1)&gt;
#map3 = affine_map&lt;(d0, d1) -&gt; (d1)&gt;
  func.func @mlp_block(%arg0: tensor&lt;64x512xf32&gt;, %arg1: tensor&lt;512x256xf32&gt;, %arg2: tensor&lt;256xf32&gt;) -&gt; tensor&lt;64x256xf32&gt; {
    %cst = arith.constant 0.000000e+00 : f32
    %0 = tensor.empty() : tensor&lt;64x256xf32&gt;
    %1 = linalg.fill ins(%cst : f32) outs(%0 : tensor&lt;64x256xf32&gt;) -&gt; tensor&lt;64x256xf32&gt;
    %2 = scf.forall (%arg3, %arg4) in (4, 8) shared_outs(%arg5 = %0) -&gt; (tensor&lt;64x256xf32&gt;) {
      %3 = affine.apply #map(%arg3)
      %4 = affine.apply #map1(%arg4)
      %extracted_slice = tensor.extract_slice %arg0[%3, 0] [16, 512] [1, 1] : tensor&lt;64x512xf32&gt; to tensor&lt;16x512xf32&gt;
      %extracted_slice_0 = tensor.extract_slice %arg1[0, %4] [512, 32] [1, 1] : tensor&lt;512x256xf32&gt; to tensor&lt;512x32xf32&gt;
      %extracted_slice_1 = tensor.extract_slice %1[%3, %4] [16, 32] [1, 1] : tensor&lt;64x256xf32&gt; to tensor&lt;16x32xf32&gt;
      %5 = linalg.matmul ins(%extracted_slice, %extracted_slice_0 : tensor&lt;16x512xf32&gt;, tensor&lt;512x32xf32&gt;) outs(%extracted_slice_1 : tensor&lt;16x32xf32&gt;) -&gt; tensor&lt;16x32xf32&gt;
      %extracted_slice_2 = tensor.extract_slice %arg2[%4] [32] [1] : tensor&lt;256xf32&gt; to tensor&lt;32xf32&gt;
      %extracted_slice_3 = tensor.extract_slice %arg5[%3, %4] [16, 32] [1, 1] : tensor&lt;64x256xf32&gt; to tensor&lt;16x32xf32&gt;
      %6 = linalg.generic {indexing_maps = [#map2, #map3, #map2], iterator_types = ["parallel", "parallel"]} ins(%5, %extracted_slice_2 : tensor&lt;16x32xf32&gt;, tensor&lt;32xf32&gt;) outs(%extracted_slice_3 : tensor&lt;16x32xf32&gt;) {
      ^bb0(%in: f32, %in_4: f32, %out: f32):
        %7 = arith.addf %in, %in_4 : f32
        %8 = arith.maximumf %7, %cst : f32
        linalg.yield %8 : f32
      } -&gt; tensor&lt;16x32xf32&gt;
      scf.forall.in_parallel {
        tensor.parallel_insert_slice %6 into %arg5[%3, %4] [16, 32] [1, 1] : tensor&lt;16x32xf32&gt; into tensor&lt;64x256xf32&gt;
      }
    }
    return %2 : tensor&lt;64x256xf32&gt;
  }</code></code></pre><p>The consumer became an <code>scf.forall</code>, a parallel-by-construction grid of 4 by 8 tiles, and inside each tile sits a private 16x512 by 512x32 matmul feeding the fused bias-and-clamp directly. </p><p>The <strong>64x256 intermediate matrix</strong> that used to exist between the two ops is gone from the program: each tile&#8217;s slice is produced, consumed, and discarded in place. <code>canonicalize</code> and <code>cse</code> then deleted the original full-size matmul as dead code. </p><p>This is a working, verified skeleton of producer fusion, the transformation entire serving stacks are built around, expressed in <strong>six lines of schedule</strong> that a performance engineer can read, diff, and review like any other code. </p><h3>Step 3. Bufferization: the value world ends here</h3><p>Everything so far lived on immutable tensors, where fusion and reordering are trivially safe because aliasing cannot be expressed. Hardware, unfortunately, sells mutable bytes. </p><p>The crossing is <strong>one-shot bufferization</strong>, a whole-module analysis that assigns every tensor a buffer while proving where it can reuse memory instead of copying. Run it on the unfused MLP and look at what comes back:</p><p>mlir-opt -one-shot-bufferize=&#8221;bufferize-function-boundaries&#8221; <strong><span>verified: actual tool output</span></strong></p><pre><code><code>#map = affine_map&lt;(d0, d1) -&gt; (d0, d1)&gt;
#map1 = affine_map&lt;(d0, d1) -&gt; (d1)&gt;
module {
  func.func @mlp_block(%arg0: memref&lt;64x512xf32, strided&lt;[?, ?], offset: ?&gt;&gt;, %arg1: memref&lt;512x256xf32, strided&lt;[?, ?], offset: ?&gt;&gt;, %arg2: memref&lt;256xf32, strided&lt;[?], offset: ?&gt;&gt;) -&gt; memref&lt;64x256xf32&gt; {
    %cst = arith.constant 0.000000e+00 : f32
    %alloc = memref.alloc() {alignment = 64 : i64} : memref&lt;64x256xf32&gt;
    linalg.fill ins(%cst : f32) outs(%alloc : memref&lt;64x256xf32&gt;)
    linalg.matmul ins(%arg0, %arg1 : memref&lt;64x512xf32, strided&lt;[?, ?], offset: ?&gt;&gt;, memref&lt;512x256xf32, strided&lt;[?, ?], offset: ?&gt;&gt;) outs(%alloc : memref&lt;64x256xf32&gt;)
    linalg.generic {indexing_maps = [#map, #map1, #map], iterator_types = ["parallel", "parallel"]} ins(%alloc, %arg2 : memref&lt;64x256xf32&gt;, memref&lt;256xf32, strided&lt;[?], offset: ?&gt;&gt;) outs(%alloc : memref&lt;64x256xf32&gt;) {
    ^bb0(%in: f32, %in_0: f32, %out: f32):
      %0 = arith.addf %in, %in_0 : f32
      %1 = arith.maximumf %0, %cst : f32
      linalg.yield %1 : f32
    }
    %cast = memref.cast %alloc : memref&lt;64x256xf32&gt; to memref&lt;64x256xf32, strided&lt;[?, ?], offset: ?&gt;&gt;
    return %alloc : memref&lt;64x256xf32&gt;
    // ...</code></code></pre><p>Tensors became <code>memref</code>s; function boundaries acquired conservative strided layouts so callers can pass any compatible view. </p><p>But the line to study is the fused elementwise: it now reads from <code>%alloc</code> <em>and writes into </em><code>%alloc</code>. The analysis proved the<strong> bias/ReLU </strong>can execute in place on the matmul&#8217;s buffer, so the second 64KB allocation that a naive translation would emit simply does not exist. </p><p>That is <strong>destination-passing style</strong> paying out: because every linalg op declared its destination back in value land, the planner had the aliasing story handed to it. </p><p>This pass, more than any other, is where <strong>MLIR&#8217;s two-world design </strong>earns its complexity: algebra above the line, memory below it, and one well-tested analysis holding the border.</p><h3>Step 4. Proof of life: the IR actually runs</h3><p>Before descending to <strong>GPU intrinsics</strong>, a checkpoint that separates this article from a whiteboard exercise. </p><p>The remaining distance to a<strong> CPU binary is mechanical</strong>: linalg to loops, loops to branches, memrefs and arithmetic to the <code>llvm</code> dialect, then JIT. </p><p>We wrapped the MLP in a tiny <code>main</code> with known inputs, chose a bias of [-10, 1] so the ReLU visibly clamps, and ran the pipeline:</p><p>the full descent, one command <strong><span>executed: mlir-runner, LLVM 20.1.2</span></strong></p><pre><code><code>$ mlir-opt mlp_run.mlir \
    -one-shot-bufferize="bufferize-function-boundaries" \
    -convert-linalg-to-loops \
    -expand-strided-metadata -lower-affine \
    -convert-scf-to-cf -finalize-memref-to-llvm \
    -convert-arith-to-llvm -convert-cf-to-llvm -convert-func-to-llvm \
    -reconcile-unrealized-casts \
  | mlir-runner -e main -entry-point-result=void \
      -shared-libs=$LLVM/lib/libmlir_runner_utils.so.20.1,$LLVM/lib/libmlir_c_runner_utils.so.20.1</code></code></pre><p>what the machine printed</p><pre><code><code>Unranked Memref base@ = 0x55b27c52f500 rank = 2 offset = 0 sizes = [2, 2] strides = [2, 1] data = 
[[0,   6], 
 [0,   12]]</code></code></pre><p>By hand: x@w is [[4, 5], [10, 11]]; add the bias to get [[-6, 6], [0, 12]]; clamp at zero for [[0, 6], [0, 12]]. The machine agrees. </p><p>Every abstraction in this article compiles to electrons, and the fused ReLU is provably doing its job on the negative entries.</p><h3>Step 5. The GPU descent, ending in genuine PTX</h3><p>Now the branch this audience came for. Same matmul, memref form, aimed at a GPU. </p><p>First, expose the parallelism that linalg has known about all along: <code>-convert-linalg-to-parallel-loops</code> turns the two parallel iterators into an <code>scf.parallel</code> over (m, n) with the reduction as an ordinary inner loop:</p><p>parallel loops <strong><span>verified: actual tool output</span></strong></p><pre><code><code>  func.func @matmul(%arg0: memref&lt;64x512xf32&gt;, %arg1: memref&lt;512x256xf32&gt;, %arg2: memref&lt;64x256xf32&gt;) {
    %c0 = arith.constant 0 : index
    %c64 = arith.constant 64 : index
    %c1 = arith.constant 1 : index
    %c256 = arith.constant 256 : index
    %c512 = arith.constant 512 : index
    scf.parallel (%arg3, %arg4) = (%c0, %c0) to (%c64, %c256) step (%c1, %c1) {
      scf.for %arg5 = %c0 to %c512 step %c1 {
        %0 = memref.load %arg0[%arg3, %arg5] : memref&lt;64x512xf32&gt;
        %1 = memref.load %arg1[%arg5, %arg4] : memref&lt;512x256xf32&gt;
        %2 = memref.load %arg2[%arg3, %arg4] : memref&lt;64x256xf32&gt;
        %3 = arith.mulf %0, %1 : f32
        %4 = arith.addf %2, %3 : f32
        memref.store %4, %arg2[%arg3, %arg4] : memref&lt;64x256xf32&gt;
      }
      scf.reduce 
    }
    return
  }
}</code></code></pre><p>Next, the mapping decision: <code>-gpu-map-parallel-loops -convert-parallel-loops-to-gpu</code> assigns parallel dimensions to the grid and materializes a launch:</p><p>a kernel is born <strong><span>verified: actual tool output</span></strong></p><pre><code><code>    gpu.launch blocks(%arg3, %arg4, %arg5) in (%arg9 = %0, %arg10 = %1, %arg11 = %c1_0) threads(%arg6, %arg7, %arg8) in (%arg12 = %c1_0, %arg13 = %c1_0, %arg14 = %c1_0) {
      %2 = affine.apply #map1(%arg3)[%c1, %c0]
      %3 = affine.apply #map1(%arg4)[%c1, %c0]
      scf.for %arg15 = %c0 to %c512 step %c1 {
        %4 = memref.load %arg0[%2, %arg15] : memref&lt;64x512xf32&gt;
        %5 = memref.load %arg1[%arg15, %3] : memref&lt;512x256xf32&gt;
        %6 = memref.load %arg2[%2, %3] : memref&lt;64x256xf32&gt;
        %7 = arith.mulf %4, %5 : f32
        %8 = arith.addf %6, %7 : f32
        memref.store %8, %arg2[%2, %3] : memref&lt;64x256xf32&gt;
      }
      // ... gpu.terminator</code></code></pre><p>Be honest about what we are looking at: a 64 by 256 grid of blocks with a single thread each, the most <strong>naive mapping</strong> a GPU has ever been insulted with. </p><p>Production pipelines tile <em>before</em> mapping, so blocks get tiles and threads get elements, then layer on shared memory staging and tensor core ops from the <code>nvgpu</code> dialect (<code>mma.sync</code> and friends). </p><p>We are deliberately taking the unoptimized path because the plumbing, not the schedule, is today&#8217;s subject; later sections covers who builds the good schedules. <code>-gpu-kernel-outlining</code> then splits host from device, the step every CUDA programmer performs mentally when writing <code>__global__</code>:</p><p>outlined device module <strong><span>verified: actual tool output, trimmed as marked</span></strong></p><pre><code><code>  gpu.module @matmul_kernel {
    gpu.func @matmul_kernel(%arg0: index, %arg1: index, %arg2: memref&lt;64x512xf32&gt;, %arg3: memref&lt;512x256xf32&gt;, %arg4: memref&lt;64x256xf32&gt;, %arg5: index) kernel attributes {known_block_size = array&lt;i32: 1, 1, 1&gt;} {
      %block_id_x = gpu.block_id  x
      %block_id_y = gpu.block_id  y
      %thread_id_x = gpu.thread_id  x
      %thread_id_y = gpu.thread_id  y
      %0 = affine.apply #map1(%block_id_x)[%arg0, %arg1]
      %1 = affine.apply #map1(%block_id_y)[%arg0, %arg1]
      // ... same loop body, now reading gpu.block_id instead of loop IVs</code></code></pre><p>The <strong>loop body is unchanged</strong>, but the block indices now arrive from <code>gpu.block_id</code> instead of loop induction variables, and the kernel records its launch invariants (<code>known_block_size</code>) as attributes for later passes to exploit. </p><p>From here, <code>-convert-gpu-to-nvvm</code> rewrites device code into the <code>llvm</code> dialect sprinkled with NVIDIA&#8217;s intrinsics:</p><p>nvvm: the vendor boundary <strong><span>verified: excerpt of actual tool output</span></strong></p><pre><code><code>gpu.module @matmul_kernel {
  // signature abbreviated: 24 pointer and index arguments
  llvm.func @matmul_kernel(%arg0: i64, ..., %arg23: i64)
      attributes {gpu.kernel, gpu.known_block_size = array&lt;i32: 1, 1, 1&gt;,
                  nvvm.kernel, nvvm.maxntid = array&lt;i32: 1, 1, 1&gt;} {
    // ...
    %24 = nvvm.read.ptx.sreg.ctaid.x : i32
    %25 = llvm.sext %24 : i32 to i64
    %26 = nvvm.read.ptx.sreg.ctaid.y : i32
    // ... address arithmetic in llvm dialect: getelementptr, mul, add ...
  }
}</code></code></pre><p><code>gpu.block_id x</code> became <code>nvvm.read.ptx.sreg.ctaid.x</code>, a name any CUDA disassembly veteran will greet like an old acquaintance: the special register file, now visible in the IR. </p><p>Finally, the packaged pipeline <code>--gpu-lower-to-nvvm-pipeline="cubin-chip=sm_90 cubin-format=isa"</code> runs the whole descent and invokes LLVM&#8217;s NVPTX backend, embedding the result in a <code>gpu.binary</code> op. Extract it and you are holding 90 lines of the real thing:</p><p>out/matmul.ptx <strong><span>generated: LLVM 20.1.2 NVPTX backend, sm_90</span></strong></p><pre><code><code>//
// Generated by LLVM NVPTX Back-End
//

.version 7.8
.target sm_90
.address_size 64

&#9;// .globl&#9;matmul_kernel
&#9;// ... prologue: bounds check, address setup ...
$L__BB0_2:
&#9;ld.global.f32 &#9;%f4, [%rd35];
&#9;ld.global.f32 &#9;%f5, [%rd34];
&#9;mul.rn.f32 &#9;%f6, %f4, %f5;
&#9;add.rn.f32 &#9;%f7, %f7, %f6;
&#9;st.global.f32 &#9;[%rd6], %f7;
&#9;add.s64 &#9;%rd36, %rd36, %rd17;
&#9;add.s64 &#9;%rd35, %rd35, %rd8;
&#9;add.s64 &#9;%rd34, %rd34, %rd10;
&#9;setp.lt.s64 &#9;%p2, %rd36, %rd19;
&#9;@%p2 bra &#9;$L__BB0_2;
&#9;// ... epilogue ... (90 lines total)</code></code></pre><p>There is a<strong> final teaching moment</strong> hiding in that inner loop, and it is the thesis of this article restated by the machine. Look at <code>st.global.f32</code>: the accumulator is written back to global memory <em>every iteration of k</em>. </p><p>The value lives happily in register <code>%f7</code>, yet the store stays, because at this level nothing can prove the output buffer does not alias the inputs, so the backend must assume a reader might observe <strong>every intermediate sum. </strong></p><p>Four levels up, in linalg on tensors, that fact was free: tensors cannot alias. We destroyed the information on the way down, exactly as the ladder predicted, and the machine code is paying for it, one redundant HBM transaction per multiply. </p><p>Every dollar of <strong>kernel engineering</strong> ever spent is, in some form, the cost of putting facts back that a lowering threw away. MLIR&#8217;s entire bet is that it is cheaper to never throw them away.</p><p>Be precise about the distance between this demonstration and a kernel you would actually deploy, because the gap is the whole subject of the next section. </p><p>A production pipeline tiles <em>before</em> it maps to hardware, so a block owns a tile and a thread owns a register-sized sliver rather than the one-thread-per-block insult we emitted. </p><p>It stages the operand tiles through the padded shared-memory <code>memref</code> from section 4, turning the plus-four skew into live bank-conflict avoidance. </p><p>And instead of scalar <code>mul.rn.f32</code> plus that redundant store, it lowers the inner tile op through the <code>nvgpu</code> dialect into <code>mma.sync</code> or, on Hopper, <code>wgmma</code>, so the loop body becomes a <strong>warp-level matrix</strong> instruction feeding tensor cores, with the accumulator held in registers across the entire reduction and written back exactly once. </p><p>Here is the point that matters for understanding the whole system: <em>every one of those improvements is a different schedule over the identical set of passes we just ran.</em> The naive descent and the peak one share their entire lowering machinery; what separates 2 percent of peak from 90 percent is the sequence of tiling, fusion, and mapping decisions applied on the way down. </p><p>Which is why the ability to write that sequence as a first-class, reviewable, searchable artifact, rather than bury it in compiler C++, is not a curiosity. It is the ballgame.</p><div><hr></div><h2>The schedule is also IR</h2><p>Steps 1 and 2 of the walkthrough smuggled in the most radical idea in modern compiler engineering, so let us stop and look at it directly. </p><p>Classically, a <strong>compiler&#8217;s transformations</strong> are opaque C++ controlled by flags and heuristics; if the heuristic tiles your attention kernel wrong, your options are a rebuild or a prayer. </p><p>Halide&#8217;s great insight, back in 2012, was to split the <em>algorithm</em> from the <em>schedule</em> and make the schedule a first-class program. MLIR&#8217;s <code>transform</code> dialect imports that split into the compiler infrastructure itself: the schedule is IR, in the same file format, checked by the same verifier, with SSA handles pointing at the payload operations it manipulates.</p><p>The consequences compound. A schedule that is data can be <strong>diffed and code reviewed</strong>: the six lines in step 2 are a reviewable artifact in a way a C++ pass never is. </p><p>It can be <strong>searched</strong>: an autotuner can emit thousands of candidate schedules, run the interpreter, and measure, without recompiling the compiler; the transform interpreter is exactly the substrate you would build a kernel search system on. </p><p>It can be <strong>shipped per workload</strong>: IREE has used transform-dialect scripts to pin codegen strategies for specific dispatches, which is how you get reproducible kernels out of a general compiler. </p><p>And it fails loudly: apply a tiling to an op that cannot be tiled and the interpreter reports a definite error against a definite handle, rather than silently generating slow code.</p><p>For a <strong>CUDA-native audience </strong>the honest framing is this: the transform dialect is the compiler conceding that <em>you</em>, the performance engineer, sometimes know better, and giving you a typed, verifiable console instead of a fork of the codebase. </p><p>The upstream vocabulary already covers the daily verbs, <code>tile_using_for</code>, <code>tile_using_forall</code>, <code>fuse_into_containing_op</code>, <code>vectorize</code>, <code>pad</code>, plus pattern application and PDL-based matching for everything else. </p><p>It is not the default path for most users and may never be; it is the expert path, and its existence changes what the expert path costs.</p><div><hr></div><h2>Build a dialect before lunch</h2><p>The claim that<strong> MLIR makes IRs cheap</strong> deserves a demonstration with a bill attached. Suppose we run inference infrastructure and want the compiler to reason about our world: quantized weights, dequantization groups, fused epilogues. </p><p>In a classical compiler, adding a first-class operation means touching the parser, the printer, the verifier, the serializer, and a dozen switch statements. </p><p>In MLIR it means describing the op once, declaratively, in TableGen&#8217;s Operation Definition Specification, and generating the rest. Here is a real op for a real concern, weight-only quantized matmul, in 28 lines:</p><p>serve_ops.td <strong><span>verified: mlir-tblgen 20.1.2</span></strong></p><pre><code><code>include "mlir/IR/OpBase.td"
include "mlir/Interfaces/SideEffectInterfaces.td"

def Serve_Dialect : Dialect {
  let name = "serve";
  let summary = "Ops for LLM serving kernels";
  let cppNamespace = "::mlir::serve";
}

class Serve_Op&lt;string mnemonic, list&lt;Trait&gt; traits = []&gt;
    : Op&lt;Serve_Dialect, mnemonic, traits&gt;;

def Serve_DequantMatmulOp : Serve_Op&lt;"dequant_matmul", [Pure]&gt; {
  let summary = "matmul with fused groupwise weight dequantization";
  let description = [{
    Computes out = act @ dequant(weights, scales) without ever
    materializing the dequantized weight tensor in memory.
  }];
  let arguments = (ins
    TensorOf&lt;[F16]&gt;:$act,
    TensorOf&lt;[I8]&gt;:$weights,
    TensorOf&lt;[F16]&gt;:$scales,
    I64Attr:$group_size
  );
  let results = (outs TensorOf&lt;[F16]&gt;:$out);
  let assemblyFormat =
    "$act `,` $weights `,` $scales attr-dict `:` functional-type(operands, results)";
}</code></code></pre><p>Feed it to <code>mlir-tblgen</code> and the machinery materializes:</p><p>the leverage, measured <strong><span>verified: actual tool output</span></strong></p><pre><code><code>$ mlir-tblgen -gen-op-decls -I /usr/lib/llvm-20/include serve_ops.td &gt; ServeOps.h.inc
$ mlir-tblgen -gen-op-defs  -I /usr/lib/llvm-20/include serve_ops.td &gt; ServeOps.cpp.inc

# 28 lines of ODS in, 259 lines of generated C++ declarations out:
class DequantMatmulOp;
    ...</code></code></pre><p>Twenty-eight declarative lines became 259 lines of generated C++ declarations, plus definitions: typed accessors (<code>getAct()</code>, <code>getGroupSize()</code>), builders, parser and printer honoring our <code>assemblyFormat</code>, and verifier scaffolding enforcing the type constraints, an f16 activation tensor, i8 weights, f16 scales, or the IR does not construct. </p><p><strong>Nine-to-one leverage</strong> on the boilerplate, and every generated line is the same battle-tested code path the 48 upstream dialects use. </p><p>This is what &#8220;<em>infrastructure</em>&#8221; means concretely: our bespoke serving op gets round-trippable syntax, verification, and pass-manager citizenship for the price of a lunch break, which is why every hardware startup in section 9 could afford a real compiler at all.</p><p>Ops are half a dialect; the other half is transformations, and those are written in C++ against the pattern infrastructure. Here is a complete, compiling pass with a pattern this audience will feel in their DRAM bills: detect a matmul whose accumulator is a fresh zero fill, and tag it so a later lowering can emit a<strong> beta-equals-zero GEMM</strong> instead of a memset followed by a GEMM, saving one full write-and-read pass over the output through HBM:</p><p>TagZeroInit.cpp <strong><span>verified: g++ -fsyntax-only against MLIR 20 headers, exit 0</span></strong></p><pre><code><code>#include "mlir/Dialect/Arith/IR/Arith.h"
#include "mlir/Dialect/Linalg/IR/Linalg.h"
#include "mlir/IR/BuiltinOps.h"
#include "mlir/IR/PatternMatch.h"
#include "mlir/Pass/Pass.h"
#include "mlir/Transforms/GreedyPatternRewriteDriver.h"

using namespace mlir;

namespace {
// Match  linalg.fill(+0.0) feeding a linalg.matmul accumulator and tag
// the matmul, so a later lowering can emit a beta = 0 GEMM instead of
// a memset followed by a GEMM (one less trip through HBM).
struct TagZeroInitMatmul : OpRewritePattern&lt;linalg::MatmulOp&gt; {
  using OpRewritePattern::OpRewritePattern;

  LogicalResult matchAndRewrite(linalg::MatmulOp op,
                                PatternRewriter &amp;rewriter) const override {
    if (op-&gt;hasAttr("serve.zero_init"))
      return failure();          // already rewritten: stop, or loop forever
    auto fill =
        op.getDpsInitOperand(0)-&gt;get().getDefiningOp&lt;linalg::FillOp&gt;();
    if (!fill)
      return failure();
    auto cst =
        fill.getInputs()[0].getDefiningOp&lt;arith::ConstantFloatOp&gt;();
    if (!cst || !cst.value().isPosZero())
      return failure();
    rewriter.modifyOpInPlace(op, [&amp;] {
      op-&gt;setAttr("serve.zero_init", rewriter.getUnitAttr());
    });
    return success();
  }
};

struct TagZeroInitPass
    : PassWrapper&lt;TagZeroInitPass, OperationPass&lt;ModuleOp&gt;&gt; {
  MLIR_DEFINE_EXPLICIT_INTERNAL_INLINE_TYPE_ID(TagZeroInitPass)
  StringRef getArgument() const final { return "serve-tag-zero-init"; }
  StringRef getDescription() const final {
    return "Tag matmuls whose accumulator is a fresh zero fill";
  }
  void runOnOperation() override {
    RewritePatternSet patterns(&amp;getContext());
    patterns.add&lt;TagZeroInitMatmul&gt;(&amp;getContext());
    if (failed(applyPatternsGreedily(getOperation(), std::move(patterns))))
      signalPassFailure();
  }
};
} // namespace</code></code></pre><p>The anatomy generalizes to every pattern you will ever write. Match structurally by walking use-def edges (<code>getDefiningOp</code> through the destination operand, courtesy of destination-passing style). </p><p>Guard exhaustively, including against your own previous rewrite, because the <strong>greedy driver</strong> reapplies patterns to fixpoint and an unguarded pattern is an infinite loop. </p><p>Mutate only through the <code>rewriter</code>, never behind its back, so the driver&#8217;s worklist stays coherent. Fifty-ish lines, and it composes with every canonicalization and lowering upstream ships.</p><p>The last pillar of the workflow is cultural: everything above gets a <strong>FileCheck test</strong>, a <code>.mlir</code> file whose <code>RUN</code> line invokes the pass and whose <code>CHECK</code> lines assert on the output IR. </p><p>Because IR is text with stable syntax, transformation tests are diffs, reviewable by humans and <strong>replayable by CI</strong>. Teams that build on MLIR inherit this discipline for free, and it shows: it is the reason a project like </p><p>Triton can refactor its entire layout system without regressing a thousand kernels silently. </p><p>The honest total cost of a production dialect is of course higher than lunch, the semantics and the <strong>verifier discipline</strong> are the real work, but the floor has moved by an order of magnitude, and floors are what determine who gets to play.</p><div><hr></div><h2>Where MLIR is hiding in your stack</h2><p>Infrastructure succeeds when it disappears. Seven years after open sourcing, MLIR has disappeared into more of the AI stack than almost anyone tracks, so let us make the map explicit.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!WIp_!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F66b9e83b-34a7-4264-9604-c5bed22fdba9_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!WIp_!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F66b9e83b-34a7-4264-9604-c5bed22fdba9_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!WIp_!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F66b9e83b-34a7-4264-9604-c5bed22fdba9_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!WIp_!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F66b9e83b-34a7-4264-9604-c5bed22fdba9_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!WIp_!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F66b9e83b-34a7-4264-9604-c5bed22fdba9_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!WIp_!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F66b9e83b-34a7-4264-9604-c5bed22fdba9_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/66b9e83b-34a7-4264-9604-c5bed22fdba9_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;MLIR timeline 2018 to 2026&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="MLIR timeline 2018 to 2026" title="MLIR timeline 2018 to 2026" srcset="https://substackcdn.com/image/fetch/$s_!WIp_!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F66b9e83b-34a7-4264-9604-c5bed22fdba9_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!WIp_!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F66b9e83b-34a7-4264-9604-c5bed22fdba9_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!WIp_!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F66b9e83b-34a7-4264-9604-c5bed22fdba9_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!WIp_!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F66b9e83b-34a7-4264-9604-c5bed22fdba9_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><em>Figure 3. Eight years from a Google whiteboard to a $3.9 billion acquisition premium on the founding team. Dates and sourcing for every event are in the dossier.</em></figcaption></figure></div><p><strong>Triton, which means PyTorch.</strong> When OpenAI rebuilt Triton on MLIR in 2022, the language stopped being a research artifact and became a compiler stack. </p><p>Python is traced into the <code>triton</code> dialect (TTIR), a target-independent tile program; conversion to the <code>triton_gpu</code> dialect (TTGIR) then makes the performance-critical decision explicit by attaching a <strong>layout encoding</strong> to every tensor, describing exactly which thread of which warp owns which element, with the <strong>LinearLayout</strong> framework giving those mappings a uniform algebra; vendor sub-dialects handle Hopper and CDNA specifics before the drop into the <code>llvm</code> dialect and PTX or AMDGCN. </p><p>Since PyTorch 2 made Inductor the default compile path and Inductor emits <strong>Triton for GPU graphs</strong>, the practical consequence is that most compiled PyTorch GPU workloads on the planet route through MLIR today, whether their owners have heard of it or not. </p><p>And in 2025 the Triton team surfaced the lower rungs of its own ladder as a product: Gluon, documented in the Triton repository, hands expert users the <strong>TTGIR-level controls</strong>, layouts, shared memory, warp specialization, while reusing the same dialect stack underneath.</p><p><strong>Which means your serving engine, too.</strong> The inference stacks this publication&#8217;s readers actually operate, vLLM and SGLang chief among them, are not usually thought of as <strong>MLIR consumers,</strong> but trace the dependency and the connection is direct. </p><p>Both lean heavily on Triton for their fused and custom kernels: fused <strong>RMSNorm</strong> and RoPE, fused MoE routing, quantized GEMM epilogues, and increasingly the attention paths through backends in the FlashInfer lineage that mix <strong>hand-written CUDA</strong> with Triton-generated variants. </p><p>Every one of those Triton kernels is a program descending through TTIR and TTGIR, which is to say through MLIR, before it reaches ptxas. The abstraction that a serving engine calls &#8220;<em>a Triton kernel</em>&#8221; is, one layer down, exactly the layout-annotated dialect IR. </p><p>And the data structure those engines sling most obsessively, the paged KV cache, is precisely the strided, address-spaced <code>memref</code> from section 4 made concrete: a block table indexing <strong>fixed-size buffers</strong>, the same view-of-memory abstraction the type system was built to name. </p><p>When you profile a <strong>vLLM deployment </strong>on a B300 and find time in a fused decode kernel, you are, whether the dashboard says so or not, looking at the output of the compiler stack this article has been dissecting. The waist is not upstream of your infrastructure; it is inside it.</p><p><strong>NVIDIA, voluntarily.</strong> CUTLASS 4&#8217;s CuTe DSL lets engineers author kernels in Python that, per NVIDIA&#8217;s documentation, JIT through MLIR into optimized code finished by ptxas, with the CuTe layout algebra as the programming model. </p><p>Read that strategically: the company with the most to gain from CUDA C++ lock-in decided its flagship kernel library&#8217;s productivity layer should be built on the industry&#8217;s shared compiler substrate. </p><p>NVIDIA is <strong>not ceding the moat</strong>, ptxas and SASS remain closed, but it has conceded where the moat is not: the language layer above the driver is now contested ground, and NVIDIA chose to contest it with MLIR rather than against it.</p><p><strong>Google, the origin.</strong> XLA remains the TPU&#8217;s compiler, and since 2022 its portable front door is StableHLO, an MLIR dialect with versioned serialization that lets JAX, TensorFlow, and PyTorch programs travel between compilers without marrying one. </p><p>IREE, under the<strong> OpenXLA umbrella</strong>, is the fully MLIR-native end-to-end: StableHLO or Torch in, linalg in the middle, compiled artifacts for CPU, CUDA, ROCm, Vulkan, and mobile out. It is the closest thing to the textbook hourglass shipped as a product, and its codegen is where much of the transform-dialect and structured-codegen machinery in this article was hardened.</p><p><strong>The silicon insurgents, which is the actual story.</strong> Look at who builds their entire software identity on this infrastructure. Tenstorrent&#8217;s tt-mlir stack, public on GitHub since 2024, compiles graphs through TTIR into its TTNN library ops for Wormhole and Blackhole silicon. </p><p>AMD ships <strong>MLIR-AIE</strong>, its 1.2 release landing in January 2026 per Phoronix, as the low-level programming layer for the NPUs in every Ryzen AI laptop. Qualcomm&#8217;s compiler group published <strong>Hexagon-MLIR </strong>on arXiv this spring: a stack that ingests, notably, both PyTorch graphs and Triton kernels and lowers them onto Hexagon NPUs. </p><p>Tesla, per the FSD v14.3 release notes, rebuilt its in-car AI compiler and runtime on MLIR and claims a <strong>20 percent reaction-time improvement </strong>from the new stack, which, whatever discount you apply to vendor claims, is a statement about MLIR&#8217;s fitness for safety-critical, latency-bounded deployment that no benchmark suite could make. </p><p>Every one of these organizations faced the same choice: a <em>decade of compiler archaeology from scratch</em>, or a running start on shared infrastructure. None of them chose archaeology.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!eHF4!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79309645-08b2-4276-ba93-5677e7da4a5b_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!eHF4!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79309645-08b2-4276-ba93-5677e7da4a5b_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!eHF4!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79309645-08b2-4276-ba93-5677e7da4a5b_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!eHF4!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79309645-08b2-4276-ba93-5677e7da4a5b_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!eHF4!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79309645-08b2-4276-ba93-5677e7da4a5b_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!eHF4!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79309645-08b2-4276-ba93-5677e7da4a5b_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/79309645-08b2-4276-ba93-5677e7da4a5b_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Production software built on MLIR as of July 2026&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Production software built on MLIR as of July 2026" title="Production software built on MLIR as of July 2026" srcset="https://substackcdn.com/image/fetch/$s_!eHF4!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79309645-08b2-4276-ba93-5677e7da4a5b_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!eHF4!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79309645-08b2-4276-ba93-5677e7da4a5b_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!eHF4!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79309645-08b2-4276-ba93-5677e7da4a5b_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!eHF4!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79309645-08b2-4276-ba93-5677e7da4a5b_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><em>Figure 4. A selection, not a census, of production systems that are MLIR inside as of July 2026, grouped by what they are to their users. Sourcing per entry in the dossier.</em></figcaption></figure></div><p><strong>And beyond AI entirely</strong>, the quieter colonization: Flang, LLVM&#8217;s Fortran frontend, is built on MLIR dialects; CIRCT applies the same infrastructure to hardware design and verification; ClangIR is the upstream effort to give C and C++ themselves a structured MLIR layer above LLVM IR. </p><p>The pattern across all of it is the one Lattner himself named in January&#8217;s &#8220;<em>Democratizing AI Compute, Part 8</em>&#8221;: MLIR the <em>infrastructure</em> won more or less totally, while the original dream of one shared AI <em>compiler</em> on top of it fractured into vigorous, incompatible ecosystems. </p><p>The waist is shared; the programs flowing through it are not. Which is precisely what makes the waist worth money.</p><div><hr></div><h2>The economics of the waist</h2><p>Regular readers know this publication&#8217;s core lens: in accelerated computing, margins pool wherever software friction is highest. </p><p>CUDA&#8217;s twenty-year lesson is that the compiler and kernel layer is not a cost center; it is the <strong>tollbooth</strong>. MLIR changes the tollbooth&#8217;s economics in three distinct ways, and the June acquisition finally put a market price on one of them.</p><p><strong>First, it collapsed the entry ticket.</strong> Before this infrastructure existed, a credible accelerator software stack meant hundreds of compiler engineers for the better part of a decade; that is what CUDA cost, and it is why so many silicon startups died with excellent chips and unusable toolchains. </p><p>Previous sections showed the mechanism by which that changed: parsing, printing, verification, <strong>pass management</strong>, testing culture, all amortized across the industry, leaving a new entrant to spend only on what is genuinely theirs, the semantics of their silicon. </p><p>The evidence is the roster in Figure 4. Tenstorrent, a company of startup scale, fields a full <strong>graph-to-silicon compiler</strong>. AMD stands up an NPU stack as a side effort. This does not guarantee anyone beats NVIDIA; it guarantees the <em>attempt</em> no longer costs a billion dollars, and lowering the cost of attempts is how oligopolies erode.</p><p><strong>Second, it standardized the engineers, not just the code.</strong> A compiler developer who learns dialects, patterns, and the linalg descent is productive at Google, AMD, Qualcomm, Tenstorrent, Modular, or any of a hundred teams, because they all speak the same infrastructure. </p><p>Before MLIR, compiler talent was siloed by proprietary IR; now there is a labor market for the waist. Every technology that has standardized its practitioners, from TCP/IP to Kubernetes, has seen the layer above it commoditize and the layer below it consolidate. Keep that in mind when reading the next paragraph.</p><p><strong>Third, it created a control point you can buy.</strong> Qualcomm did not pay approximately $3.9 billion, per the Reuters-reported terms, for Modular&#8217;s revenue. It paid for roughly 150 people, per <strong>NAND Research&#8217;s</strong> headcount estimate, who include the founding architects of both MLIR and LLVM, plus MAX and Mojo, the <em>most complete independent attempt</em> to unify the layer above everyone&#8217;s silicon. </p><p>That is on the order of $26 million per employee, ten times acquihire norms, for a company whose language hit 1.0 beta in May, per its own PyPI listing, and whose compiler remains closed source with an open sourcing commitment for later this year. </p><p>The premium is not for what Modular sells; it is for what Qualcomm was <strong>structurally unable to build alone</strong>, demonstrated by the fact that its own Hexagon-MLIR effort was already consuming the ecosystem Modular&#8217;s founders created. </p><p>In the framework of our kernel-moat analysis: when the moat migrates from a proprietary language to a shared substrate, the scarce asset becomes the people who define the substrate&#8217;s direction, and June 24 was the day that asset got marked to market.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!ufst!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F40a98fbc-dafe-4d3e-acef-d60f77bfb5fa_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!ufst!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F40a98fbc-dafe-4d3e-acef-d60f77bfb5fa_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!ufst!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F40a98fbc-dafe-4d3e-acef-d60f77bfb5fa_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!ufst!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F40a98fbc-dafe-4d3e-acef-d60f77bfb5fa_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!ufst!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F40a98fbc-dafe-4d3e-acef-d60f77bfb5fa_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!ufst!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F40a98fbc-dafe-4d3e-acef-d60f77bfb5fa_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/40a98fbc-dafe-4d3e-acef-d60f77bfb5fa_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Modular capital raised, valuation, and acquisition price&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Modular capital raised, valuation, and acquisition price" title="Modular capital raised, valuation, and acquisition price" srcset="https://substackcdn.com/image/fetch/$s_!ufst!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F40a98fbc-dafe-4d3e-acef-d60f77bfb5fa_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!ufst!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F40a98fbc-dafe-4d3e-acef-d60f77bfb5fa_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!ufst!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F40a98fbc-dafe-4d3e-acef-d60f77bfb5fa_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!ufst!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F40a98fbc-dafe-4d3e-acef-d60f77bfb5fa_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><em>Figure 5. The compiler layer, marked to market. Funding history per Modular&#8217;s own repository newsline ($380M raised, $1.6B valuation at the September 2025 round); acquisition price as reported by Reuters and the WSJ; headcount per NAND Research; the per-person figure is our derived arithmetic.</em></figcaption></figure></div><p>NVIDIA&#8217;s response is the most sophisticated play on the board: embrace the waist from above. </p><p>Contributing to Triton&#8217;s NVIDIA backends and shipping <strong>CuTe DSL</strong> on MLIR costs NVIDIA the exclusivity of its <em>language</em> layer, which was already eroding, while making its <em>hardware</em> the best-supported target of the shared stack everyone else depends on, and keeping the truly proprietary layers, ptxas, SASS scheduling, <strong>NVLink-domain libraries</strong>, exactly as closed as before. </p><p>Commoditize your complement, in textbook form. For the neoclouds and inference operators this publication serves, this is not abstract; it changes the arithmetic that turns capex into margin. </p><p>A fleet&#8217;s single largest source of pricing power against its accelerator vendor is <strong>credible multi-homing</strong>: the ability to qualify a second silicon supplier and actually move workloads. </p><p>Historically that ability died at the kernel layer, because a serving stack tuned to CUDA was a serving stack married to NVIDIA, and the cost of re-qualifying every fused kernel on a new target was prohibitive enough that &#8220;<em>we could switch</em>&#8221; was rarely true. </p><p>The <strong>shared waist</strong> is what makes the threat credible, incrementally. Every serving stack that targets the waist, vLLM and SGLang through Triton, IREE, MAX, is quietly reducing the switching cost between a B300 fleet and an MI400 fleet at the kernel layer. </p><p>Portability never arrives as a binary event, and CUDA-only paths (<em>FlashAttention&#8217;s fastest branches, NCCL, the long tail of fused op</em>s) remain very real. But price convergence per token happens at the margin, and the <strong>margin is now programmable</strong>. </p><p>Watch whether accelerator price-performance spreads narrow over the next four quarters as these stacks mature; that spread is the purest financial expression of what this article has been describing.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>The counter-narrative</h2><p>The strongest version of the case against, because a thesis untested against its critics is marketing.</p><p><strong>&#8220;The library ate the compiler.&#8221;</strong> The performance frontier on NVIDIA hardware is still held by artisanal kernels: FlashAttention lineages, cuDNN, CUTLASS templates, hand-scheduled decode paths. </p><p>Much of what shipping compilers do is orchestrate calls into those libraries, and a fair reading of the <strong>TPP-MLIR work</strong> (arXiv 2404.15204), which reached roughly 90 percent of expert-written kernel performance on CPU benchmarks using upstream MLIR, is that the last ten percent is precisely where the money lives. </p><p>The rebuttal is not that compilers beat ninjas; it is that MLIR is the substrate the ninjas themselves are migrating to, CuTe DSL and Gluon being exhibits A and B. But the point stands: the waist wins by hosting excellence, not by automating it away.</p><p><strong>&#8220;LLMs will write the kernels.&#8221;</strong><em> If models generate CUDA or PTX directly on demand, does a shared IR layer matter less?</em> Plausibly it matters <em>more</em>: generated code needs verification, search needs a space with structure, and an IR with machine-checkable semantics and a transform language is a far better substrate for automated kernel search than free-form C++. </p><p>But intellectual honesty requires flagging this as the genuine wildcard: a world of cheap, correct, model-generated kernels reshuffles every layer of this analysis, ours included.</p><p><strong>&#8220;Dialect soup.&#8221;</strong> Forty-eight upstream dialects, hundreds downstream, and no blessed end-to-end path is real fragmentation with real costs: two teams can both &#8220;<em>use MLIR</em>&#8221; and share nothing but a parser. </p><p>Lattner&#8217;s own Part 8 critique makes essentially this charge, and the upstream community&#8217;s response, charters, governance reform, the <em>StableHLO-style versioning </em>discipline spreading to more dialects, is work in progress, not a solved problem. </p><p>Relatedly, core MLIR offers no stable cross-release IR contract: downstream projects pay a real rebase tax every LLVM release, a tax Triton and IREE budget actual headcount for.</p><p><strong>&#8220;The C++ tax.&#8221;</strong> The infrastructure&#8217;s expressiveness rides on heavy C++ templates and TableGen metaprogramming; build times are punishing, the learning curve is a cliff, and debugging a misapplied pattern through the greedy driver is an acquired skill. </p><p>Python bindings and tooling soften the edges for users, but dialect <em>authors</em> live in C++, and that gates the contributor pool. Fair, true, and priced in: the roster in Figure 4 is the market&#8217;s judgment that the tax is worth the leverage. </p><p>Every one of these criticisms describes a<strong> middle-aged infrastructure project</strong> with genuine adoption. None of them describes an alternative, and in infrastructure, the alternative is the only criticism that kills.</p><div><hr></div><h2>What to do with this</h2><p>If this article did its job, <strong>MLIR stopped being a logo</strong> on other people&#8217;s architecture slides and became legible machinery. </p><p>Here is the shortest path from legible to useful, calibrated for readers who already know what a memory coalescing problem feels like.</p><p><strong>Learn to read dumps before writing anything.</strong> The single highest-leverage habit is inspecting the IR your existing tools already produce. Triton will show you its TTIR and TTGIR for any kernel you own; start with a matmul you understand and find the layout encodings. </p><p>On your own experiments, <code>mlir-opt --mlir-print-ir-after-all</code> is the x-ray: every pass, before and after. When output confuses you, <code>--mlir-print-op-generic</code> strips the sugar and shows you the honest structure, exactly as in section 3.</p><p><strong>Replay this article.</strong> Everything here ran on stock Ubuntu 24.04 packages: <code>apt install mlir-20-tools llvm-20 libmlir-20-dev</code> and you have <code>mlir-opt</code>, <code>mlir-runner</code>, <code>mlir-translate</code>, <code>mlir-tblgen</code>, and <code>mlir-reduce</code> (the last being delta-debugging for IR: feed it a crashing module and a script, get back a minimal reproducer). </p><p>The dossier below lists every command. An afternoon of replaying the walkthrough will teach you more than a month of architecture diagrams.</p><p><strong>Then climb the same ladder the ecosystem climbed.</strong> The canonical study sequence: the Toy tutorial on mlir.llvm.org for the object model; the linalg dialect rationale for the structured-ops philosophy of section 5; the transform dialect tutorial for section 7&#8217;s machinery; then the CGO 2021 paper, which reads completely differently once you have touched the IR. </p><blockquote><p>For a <strong>working CUDA engineer,</strong> a realistic 30-day arc is: week one, read IR and replay the walkthrough; week two, lower your own toy op through the same pipeline and break things on purpose; week three, build the previous section dialect and make the pattern fire on real IR; week four, open the dumps of one production Triton kernel you own and annotate every layout decision. </p></blockquote><p>At the end of that month you will be conversant in the layer where, as section 10 argued, an increasing share of this industry&#8217;s margin gets decided.</p><p>Seven years ago this was a whiteboard sketch about taming TensorFlow&#8217;s compiler zoo. Today it compiles the kernels NVIDIA brags with, the <strong>NPU stacks </strong>of three chip giants, the vision system in a million cars, and it just priced a 150 person team at nearly four billion dollars. </p><p>The previous installment of this series ended by observing that LLVM won by being the infrastructure nobody had an incentive to replace. MLIR is running the same play one level up, on the layer where the AI hardware war is actually fought. </p><p>The empire had one IR. The dominion has as many as you can define, verify, and lower, and now you know how it is done.</p><div><hr></div><h2>Verification dossier</h2><p>House methodology: every load-bearing claim is tiered, every number is sourced or explicitly derived, and for this piece, every code artifact was executed against real tools before publication. </p><p>Environment: Ubuntu 24.04, stock packages <code>mlir-20-tools</code>, <code>llvm-20</code>, <code>libmlir-20-dev</code>, all LLVM/MLIR <strong>20.1.2</strong>. </p><p>Syntax note: MLIR evolves between LLVM releases; snippets are guaranteed against 20.1.2 and may need small spelling changes on older or newer toolchains.</p><h3>Code verification manifest</h3><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Qn7C!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb0f091cf-8966-4b8d-8664-d78ff4c6e534_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Qn7C!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb0f091cf-8966-4b8d-8664-d78ff4c6e534_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!Qn7C!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb0f091cf-8966-4b8d-8664-d78ff4c6e534_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!Qn7C!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb0f091cf-8966-4b8d-8664-d78ff4c6e534_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!Qn7C!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb0f091cf-8966-4b8d-8664-d78ff4c6e534_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Qn7C!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb0f091cf-8966-4b8d-8664-d78ff4c6e534_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b0f091cf-8966-4b8d-8664-d78ff4c6e534_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:233813,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/203533021?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb0f091cf-8966-4b8d-8664-d78ff4c6e534_1634x893.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Qn7C!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb0f091cf-8966-4b8d-8664-d78ff4c6e534_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!Qn7C!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb0f091cf-8966-4b8d-8664-d78ff4c6e534_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!Qn7C!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb0f091cf-8966-4b8d-8664-d78ff4c6e534_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!Qn7C!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb0f091cf-8966-4b8d-8664-d78ff4c6e534_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h3>Claims and confidence</h3><p>Tiers: <strong><span>A</span></strong> primary source or machine-verified. <strong><span>B</span></strong> reputable secondary reporting or vendor primary for vendor facts. <strong><span>C</span></strong> reasoned inference or derived arithmetic, flagged as such in text. <strong><span>D</span></strong> interpretation and forecast: our analysis, argued not asserted.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!LZdF!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26da92fe-0d13-43df-8f2f-cd7a64f952ff_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!LZdF!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26da92fe-0d13-43df-8f2f-cd7a64f952ff_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!LZdF!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26da92fe-0d13-43df-8f2f-cd7a64f952ff_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!LZdF!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26da92fe-0d13-43df-8f2f-cd7a64f952ff_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!LZdF!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26da92fe-0d13-43df-8f2f-cd7a64f952ff_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!LZdF!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26da92fe-0d13-43df-8f2f-cd7a64f952ff_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/26da92fe-0d13-43df-8f2f-cd7a64f952ff_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:226833,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/203533021?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26da92fe-0d13-43df-8f2f-cd7a64f952ff_1634x893.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!LZdF!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26da92fe-0d13-43df-8f2f-cd7a64f952ff_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!LZdF!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26da92fe-0d13-43df-8f2f-cd7a64f952ff_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!LZdF!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26da92fe-0d13-43df-8f2f-cd7a64f952ff_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!LZdF!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26da92fe-0d13-43df-8f2f-cd7a64f952ff_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Sources</h2><ol><li><p>Qualcomm, &#8220;Qualcomm to Acquire Modular,&#8221; investor press release, June 24, 2026. investor.qualcomm.com/news-events/press-releases/news-details/2026/Qualcomm-to-Acquire-Modular/default.aspx</p></li><li><p>Modular, &#8220;Qualcomm to acquire Modular,&#8221; company blog, June 2026. modular.com/blog/qualcomm-to-acquire-modular</p></li><li><p>Quartz, deal report citing Reuters and WSJ terms, June 24, 2026. qz.com/qualcomm-acquires-modular-ai-software-stock-deal-062426</p></li><li><p>NAND Research, &#8220;Qualcomm acquires Modular for its hardware-agnostic AI software layer,&#8221; June 2026. nand-research.com</p></li><li><p>Modular repository newsline (funding history, releases). github.com/modular/modular</p></li><li><p>Chris Lattner, &#8220;Democratizing AI Compute, Part 8: What about the MLIR compiler infrastructure?&#8221;, Modular blog, January 2026. modular.com/blog/democratizing-ai-compute-part-8-what-about-the-mlir-compiler-infrastructure</p></li><li><p>Modular, &#8220;The Path to Mojo 1.0&#8221;; PyPI, mojo-compiler release history (1.0.0b1, May 7, 2026). modular.com/blog/the-path-to-mojo-1-0; pypi.org/project/mojo-compiler</p></li><li><p>NVIDIA, CUTLASS documentation: CuTe DSL introduction and overview. docs.nvidia.com/cutlass/latest/media/docs/pythonDSL/cute_dsl_general/dsl_introduction.html</p></li><li><p>Triton project, Gluon documentation and dialect sources. triton-lang.org/main/gluon</p></li><li><p>Qualcomm et al., &#8220;Hexagon-MLIR,&#8221; arXiv 2602.19762, 2026. arxiv.org/pdf/2602.19762</p></li><li><p>Phoronix, &#8220;AMD MLIR-AIE 1.2,&#8221; January 2026. phoronix.com/news/AMD-MLIR-AIE-1.2</p></li><li><p>Tesla Oracle, &#8220;Tesla rolls out FSD v14.3 (2026.2.9.6): better reaction time, rewritten AI compiler with MLIR,&#8221; April 8, 2026. teslaoracle.com</p></li><li><p>Lattner, Amini, Bondhugula, Cohen, Davis, Pienaar, Riddle, Shpeisman, Vasilache, Zinenko, &#8220;MLIR: Scaling Compiler Infrastructure for Domain Specific Computation,&#8221; CGO 2021.</p></li><li><p>&#8220;TPP-MLIR&#8221; upstream performance study, arXiv 2404.15204. arxiv.org/pdf/2404.15204</p></li><li><p>MLIR project documentation: Toy tutorial, linalg rationale, transform dialect tutorial. mlir.llvm.org</p></li><li><p>StableHLO specification and repository. github.com/openxla/stablehlo</p></li><li><p>Tenstorrent, tt-mlir repository. github.com/tenstorrent/tt-mlir</p></li></ol>]]></content:encoded></item><item><title><![CDATA[How LLVM Works: The IR That Took Over Modern Computing]]></title><description><![CDATA[A two-person grant at the University of Illinois became the compiler substrate under Apple, Android, the PlayStation, Google and Meta datacenters, and every serious AI stack shipping today.]]></description><link>https://www.thesoftwarefrontier.com/p/how-llvm-works-the-ir-that-took-over</link><guid isPermaLink="false">https://www.thesoftwarefrontier.com/p/how-llvm-works-the-ir-that-took-over</guid><dc:creator><![CDATA[Lorenzo Bradanini]]></dc:creator><pubDate>Sun, 12 Jul 2026 12:56:15 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!tKIr!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c240161-6eb0-4373-8587-5ad5c822855e_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!tKIr!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c240161-6eb0-4373-8587-5ad5c822855e_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!tKIr!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c240161-6eb0-4373-8587-5ad5c822855e_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!tKIr!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c240161-6eb0-4373-8587-5ad5c822855e_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!tKIr!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c240161-6eb0-4373-8587-5ad5c822855e_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!tKIr!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c240161-6eb0-4373-8587-5ad5c822855e_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!tKIr!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c240161-6eb0-4373-8587-5ad5c822855e_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/4c240161-6eb0-4373-8587-5ad5c822855e_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2504090,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/203533011?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c240161-6eb0-4373-8587-5ad5c822855e_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!tKIr!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c240161-6eb0-4373-8587-5ad5c822855e_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!tKIr!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c240161-6eb0-4373-8587-5ad5c822855e_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!tKIr!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c240161-6eb0-4373-8587-5ad5c822855e_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!tKIr!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c240161-6eb0-4373-8587-5ad5c822855e_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>Intro</h2><p>This is the story of <strong>how LLVM&#8217;s intermediate representation ate modern computing</strong>, told at the level of the instructions themselves, with the source numbers checked.</p><p><strong>LLVM won </strong>because of two decisions made in one semester in the year 2000: a single, strongly typed, target-independent intermediate representation in <strong>SSA form </strong>that can be shipped and executed on its own, and a modular library architecture that let anyone reuse a compiler pass without dragging the whole compiler along. </p><p>Everything else, the trillion-dollar reach across mobile, cloud, consoles and AI, is a consequence of those two decisions. </p><p>The current release, LLVM 22.1, shipped in February 2026 and still carries the exact architecture<strong> Chris Lattner </strong>sketched over a<em> winter break twenty-five years ago.</em></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>The hole that GCC left</h2><p>Why the dominant open compiler of 2000 could not become infrastructure</p><p><strong>Compilers are the one technology every other technology has to pass through.</strong> They translate every line a human writes into something a processor can run, which means a compiler that is easy to build on is a lever under the entire software economy. </p><p>By the <strong>year 2000</strong> the open lever was the GNU Compiler Collection, and it was showing its age.</p><p><strong>GCC</strong>, launched in the 1980s, had become foundational: it built every Linux system, every Apple computer of the era, and a large share of the open source world. </p><p>But as <strong>Vikram Adve</strong> and Chris Lattner recount in their <a href="https://cacm.acm.org/federal-funding-of-academic-research/the-llvm-compiler-infrastructure/">June 2026 retrospective in Communications of the ACM</a>, GCC lagged on the techniques that were becoming central to modern compilation. </p><p>It was written in C rather than a modern object-oriented language, it lacked cross-module interprocedural optimization, and it did not gain <strong>Static Single Assignment</strong> form until GCC 4.0 in 2005. </p><p>Static Single Assignment, or SSA, is the representation on which most serious dataflow optimization is built, and GCC arrived late to it.</p><p>The deeper problem was structural. <strong>GCC was monolithic</strong>. You could not lift a piece of it out, an optimizer, a code generator, an analysis, and reuse that piece without pulling most of the compiler with it. </p><p>That single fact foreclosed most of the interesting futures: load-time and just-in-time compilation of mobile code, embedded scripting languages, sandboxed browser extensions, graphics shader compilation. </p><p>The managed-language runtimes of the day, the<strong> Jikes RVM for Java</strong>, Mono and later Roslyn for the CLR, were genuinely excellent inside their domain and advanced garbage collection and JIT compilation considerably. </p><p>But each imposed a language object model and a set of runtime semantics that ruled it out for the vast majority of other languages, and for anything that needed to ship mobile code without a managed runtime underneath it.</p><p>So the opening was specific and it was large: no open system combined ahead-of-time optimizing compilation of the kind static languages need with the self-contained, shippable, <strong>just-in-time-capable code representation</strong> that managed languages have, and offered both to <em>every</em> language rather than one family. Filling that hole is the whole of what follows.</p><h3>The federal money that made it possible</h3><p>Adve had come to Illinois with a research program aimed at flexible dynamic compilation for arbitrary languages. </p><p>His CAREER proposal to the <strong>National Science Foundation</strong>&#8216;s Next Generation Software program, submitted in 2000, funded the work; the acknowledgments in the CACM piece name grant <code>NSF EIA-0093426</code>, and Adve&#8217;s own record lists the <strong>CAREER </strong>award at roughly $499,211. </p><p>That grant, <strong>Adve and Lattner</strong> write, &#8220;<em>would prove pivotal to the group&#8217;s early work on LLVM</em>,&#8221; and it supported both authors through most of the project&#8217;s early life. </p><p>The point they press, and it is worth pressing, is that this was blue-sky research with an unknowable outcome, funded on the condition that the artifacts be released as open source. The<strong> trillion-dollar downstream</strong> was not forecast. It was seeded.</p><h4>A naming footnote</h4><p>LLVM originally stood for Low Level Virtual Machine. The name outgrew the acronym years ago; the project&#8217;s own position, stated on <a href="https://llvm.org/">llvm.org</a>, is that &#8220;LLVM&#8221; is now simply the name of the umbrella project and has &#8220;little to do with traditional virtual machines.&#8221; We use it as a proper noun throughout.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>The IR is the entire argument</h2><p>Everything LLVM can do is downstream of a <strong>single design object</strong></p><p>If you understand one thing about LLVM, understand the intermediate representation, because the reach of the project is a property of the IR and not of any particular optimizer or backend. </p><p><strong>LLVM IR is a strongly typed</strong>, RISC-like virtual instruction set. It abstracts away the machine, there are no physical registers and no addressing modes, but it stays low enough to represent operating systems, kernels and application code without imposing any language&#8217;s object model on top.</p><p>Three things about it do most of the work.</p><p><strong>It is in SSA form.</strong> Every value is assigned exactly once, and each use points back to a single, unambiguous definition. Where control flow merges, an explicit <code>phi</code> instruction selects the incoming value based on which predecessor block you arrived from. </p><p>This is the representation that <strong>Ron Cytron</strong> and colleagues formalized in 1991, and it is what makes dataflow optimization tractable: because there is one definition per value, the flow from definitions to uses is explicit, and reordering transformations are freed from spurious dependencies. </p><p>Consider, for example, a trivial loop.</p><pre><code><strong><span>int</span></strong> sum_to(<strong><span>int</span></strong> n){
  <strong><span>int</span></strong> s = <span>0</span>;
  <strong><span>for</span></strong> (<strong><span>int</span></strong> i = <span>0</span>; i &lt; n; i++)
    s += i;
  <strong><span>return</span></strong> s;
}</code></pre><p>Compiled at <code>-O0</code>, <strong>Clang</strong> first emits the naive version: local variables become stack slots via <code>alloca</code>, and every read and write is an explicit <code>load</code> or <code>store</code>. </p><p>Note the pointer type is simply <code>ptr</code>, a point we return to below.</p><pre><code><strong><span>define</span></strong> <span>i32</span> <span>@sum_to</span>(<span>i32</span> %n) {
<span>entry:</span>
  %i   = <strong><span>alloca</span></strong> <span>i32</span>
  %s   = <strong><span>alloca</span></strong> <span>i32</span>
  <strong><span>store</span></strong> <span>i32</span> <span>0</span>, <span>ptr</span> %s
  <strong><span>store</span></strong> <span>i32</span> <span>0</span>, <span>ptr</span> %i
  <strong><span>br</span></strong> <strong><span>label</span></strong> %cond
  <em><span>; ... loads and stores on every iteration ...</span></em>
}</code></pre><p>Then the <code>mem2reg</code> pass (memory-to-register promotion, one of the first transformations any pipeline runs) proves those stack slots never escape and rewrites them into <strong>pure SSA values. </strong></p><p>The stack traffic disappears, and <code>phi</code> nodes appear at the loop header and at the exit merge:</p><pre><code><strong><span>define</span></strong> <span>i32</span> <span>@sum_to</span>(<span>i32</span> %n) {
<span>entry:</span>
  %pos = <strong><span>icmp</span></strong> sgt <span>i32</span> %n, <span>0</span>
  <strong><span>br</span></strong> <span>i1</span> %pos, <strong><span>label</span></strong> %loop, <strong><span>label</span></strong> %exit

<span>loop:</span>                                <em><span>; preds = %entry, %loop</span></em>
  %i = <strong><span>phi</span></strong> <span>i32</span> [ <span>0</span>, %entry ], [ %i.next, %loop ]
  %s = <strong><span>phi</span></strong> <span>i32</span> [ <span>0</span>, %entry ], [ %s.next, %loop ]
  %s.next = <strong><span>add</span></strong> nsw <span>i32</span> %s, %i
  %i.next = <strong><span>add</span></strong> nsw <span>i32</span> %i, <span>1</span>
  %done = <strong><span>icmp</span></strong> eq <span>i32</span> %i.next, %n
  <strong><span>br</span></strong> <span>i1</span> %done, <strong><span>label</span></strong> %exit, <strong><span>label</span></strong> %loop

<span>exit:</span>                                <em><span>; preds = %entry, %loop</span></em>
  %r = <strong><span>phi</span></strong> <span>i32</span> [ <span>0</span>, %entry ], [ %s.next, %loop ]
  <strong><span>ret</span></strong> <span>i32</span> %r
}</code></pre><p>The <code>nsw</code><strong> flags </strong>mark the additions as having no signed wraparound, a promise the front end makes on the compiler&#8217;s behalf so later passes can optimize more aggressively. </p><p>This is the<strong> texture of LLVM IR</strong>: low-level operations, explicit control flow, and a scattering of flags and metadata that carry higher-level facts down to where they can be exploited.</p><p><strong>It carries a real type system.</strong> Integers of arbitrary bit width (<code>i1</code>, <code>i8</code>, <code>i32</code>, <code>i64</code>), floating point types, vectors as first-class values, arrays and structures. </p><p>Vectors matter more than they look: because a <strong>vector is a first-class value</strong>, the same IR expresses Intel SSE, AVX and AVX-512, Arm Neon and SVE, and the tensor operations that machine-learning front ends emit, and the auto-vectorizer can create them from scalar loops using dependence analysis. </p><p>Aggregate access goes through <code>getelementptr</code>, the address-arithmetic instruction that computes a field or element address without loading it:</p><pre><code><em><span>; struct Point { i32 x; i32 y; };  compute &amp;p-&gt;y</span></em>
%y.ptr = <strong><span>getelementptr</span></strong> inbounds { <span>i32</span>, <span>i32</span> }, <span>ptr</span> %p, <span>i32</span> <span>0</span>, <span>i32</span> <span>1</span>
%y     = <strong><span>load</span></strong> <span>i32</span>, <span>ptr</span> %y.ptr</code></pre><p><strong>It is extensible through intrinsics.</strong> Intrinsic functions are built-ins with defined semantics that behave like ordinary calls, so a pass that does not care about them can ignore them safely, while selected passes and the backend can recognize <code>llvm.memcpy</code>, <code>llvm.masked.load</code>, the matrix and vector-predication <strong>intrinsics</strong>, atomics and the rest, and lower them specially. </p><p>This is the mechanism that let the IR absorb decades of new hardware features without a language change every time.</p><p>The complexity was not free, and the authors are candid about it.<strong> LLVM IR launched in 2003</strong> with fewer than thirty-five polymorphic operations. It now carries more than one hundred and seventy-five, including in excess of one hundred and <strong>ten vector operations </strong>defined as intrinsics, plus hundreds of further semantic intrinsics for garbage collection, memory ordering, synchronization and stack manipulation. </p><p>Most of that can be ignored when writing a front end or an optimization pass, but it makes writing a new backend, or a formal semantics for the IR, genuinely hard.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!7UdV!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6b75fb7-62b5-46ac-b48a-3ce8b04cfd0c_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!7UdV!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6b75fb7-62b5-46ac-b48a-3ce8b04cfd0c_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!7UdV!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6b75fb7-62b5-46ac-b48a-3ce8b04cfd0c_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!7UdV!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6b75fb7-62b5-46ac-b48a-3ce8b04cfd0c_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!7UdV!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6b75fb7-62b5-46ac-b48a-3ce8b04cfd0c_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!7UdV!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6b75fb7-62b5-46ac-b48a-3ce8b04cfd0c_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e6b75fb7-62b5-46ac-b48a-3ce8b04cfd0c_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Bar chart: LLVM IR operation count grew from fewer than 35 in 2003 to more than 175 in 2026, with roughly 110 vector operations defined as intrinsics.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Bar chart: LLVM IR operation count grew from fewer than 35 in 2003 to more than 175 in 2026, with roughly 110 vector operations defined as intrinsics." title="Bar chart: LLVM IR operation count grew from fewer than 35 in 2003 to more than 175 in 2026, with roughly 110 vector operations defined as intrinsics." srcset="https://substackcdn.com/image/fetch/$s_!7UdV!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6b75fb7-62b5-46ac-b48a-3ce8b04cfd0c_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!7UdV!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6b75fb7-62b5-46ac-b48a-3ce8b04cfd0c_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!7UdV!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6b75fb7-62b5-46ac-b48a-3ce8b04cfd0c_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!7UdV!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6b75fb7-62b5-46ac-b48a-3ce8b04cfd0c_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Figure 1.</strong> The count of polymorphic operations in the IR, at first release and today. The 2026 split between core and vector operations is illustrative; the totals (&#8221;fewer than 35,&#8221; &#8220;more than 175,&#8221; &#8220;in excess of 110&#8221; vector) are the figures Adve and Lattner give in the CACM retrospective.</figcaption></figure></div><h3>Three forms of one thing, and the opaque-pointer turn</h3><p>The IR exists in three isomorphic forms: a human-readable textual assembly (<code>.ll</code>), a dense binary serialization called <strong>bitcode</strong> (<code>.bc</code>), and an in-memory form the compiler manipulates directly. </p><p>They are interconvertible without loss, which is exactly what lets the IR be shipped and compiled later rather than only used in-process.</p><pre><code>clang -S -emit-llvm foo.c -o foo.ll   <em><span># textual IR</span></em>
clang -c -emit-llvm foo.c -o foo.bc   <em><span># bitcode</span></em>
llvm-as foo.ll -o foo.bc              <em><span># text  -&gt; bitcode</span></em>
llvm-dis foo.bc -o foo.ll             <em><span># bitcode -&gt; text</span></em></code></pre><p>The IR is not frozen. Two recent changes show the project still reshaping its own foundation. The larger one is <strong>opaque pointers</strong>. For most of LLVM&#8217;s life a pointer carried its pointee type, so you wrote <code>i32*</code>. </p><p>That information was redundant with the types on the loads and stores that actually used the pointer, and it forced a swarm of no-op bitcast instructions. </p><p>Per the <a href="https://llvm.org/docs/OpaquePointers.html">official documentation</a>, <strong>opaque pointers</strong>, a single untyped <code>ptr</code>, became the default in LLVM 15 (2022), and typed pointers were removed entirely in LLVM 17 (2023). Every IR sample in this article uses <code>ptr</code> for that reason. </p><p>The second, newer change landed in <strong>LLVM 22</strong>: a dedicated <code>ptrtoaddr</code> instruction that extracts a pointer&#8217;s integer address while separating it from provenance tracking, which the older <code>ptrtoint</code> conflated. </p><p>It is a small instruction with a real purpose, sharper semantics for the memory model, and it is the kind of change that keeps the IR honest as verification tools grow more demanding.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Passes, and the separation of concerns</h2><p>The modular library architecture is the part that actually won</p><p>The most cited <strong>technical novelty of LLVM is the IR</strong>. The most <em>consequential</em> design decision, the authors argue and we agree, is not technical at all in the  usual sense: it is the modular, library-based architecture, built on a strict adherence to separation of concerns. </p><p>A <strong>compiler built on LLVM</strong> is assembled from libraries with clean interfaces rather than carved out of a monolith. That is why pieces of LLVM turn up in circuit-design tools, symbolic-execution engines, browser sandboxes and quantum-computing compilers, uses the authors freely admit they never imagined.</p><p>Concretely, transformation is organized as <strong>a pipeline of passes</strong>. Analysis passes compute facts (<em>dominator trees, alias analysis, loop information, scalar evolution</em>) and cache them; transformation passes consume those facts and rewrite the IR, declaring which analyses they preserve so the manager knows what to recompute. </p><p>Since LLVM 13 (2021) the default engine is the <strong>new pass manager</strong>, a rewrite that made analysis caching explicit and pass composition cheaper.</p><p>A canonical <code>-O2</code> pipeline threads dozens of passes together. A representative slice: </p><ul><li><p><code>SROA</code> and <code>mem2reg</code> to lift memory into SSA values; </p></li><li><p><code>EarlyCSE</code> and <code>GVN</code> to eliminate redundant computation; </p></li><li><p><code>InstCombine</code> to canonicalize and simplify instruction patterns; </p></li><li><p><code>SimplifyCFG</code> to clean up control flow; the inliner to pull small callees into their callers; </p></li><li><p><code>LICM</code> to hoist loop-invariant code; </p></li><li><p><code>SCCP</code> for sparse conditional constant propagation; </p></li><li><p><code>IndVars</code> and <code>LoopUnroll</code> for loops; </p></li></ul><p>then the two vectorizers, the <strong>loop vectorizer </strong>and the<strong> SLP</strong> (superword-level parallelism) vectorizer, near the end. You can watch any of it run.</p><pre><code><em><span># run one pass and print the result</span></em>
opt -passes=<span>&#8220;mem2reg&#8221;</span> -S sum.ll -o -

<em><span># see the full default -O2 pipeline the pass builder constructs</span></em>
opt -passes=<span>&#8220;default&lt;O2&gt;&#8221;</span> -print-pipeline-passes -S sum.ll -o /dev/null</code></pre><p>Writing a pass is deliberately cheap, which is the whole point. The 2003 release announcement claimed new users could &#8220;<em>write their first LLVM pass in hours,</em>&#8221; and the modern <strong>out-of-tree plugin form </strong>holds to that. Here is a complete, loadable function pass in the new pass manager:</p><pre><code><em><span>#include &#8220;llvm/IR/PassManager.h&#8221;</span></em>
<em><span>#include &#8220;llvm/Passes/PassBuilder.h&#8221;</span></em>
<em><span>#include &#8220;llvm/Passes/PassPlugin.h&#8221;</span></em>
<strong><span>using namespace</span></strong> llvm;

<strong><span>struct</span></strong> CountAllocas : PassInfoMixin&lt;CountAllocas&gt; {
  PreservedAnalyses run(Function &amp;F, FunctionAnalysisManager &amp;) {
    <strong><span>unsigned</span></strong> n = <span>0</span>;
    <strong><span>for</span></strong> (BasicBlock &amp;BB : F)
      <strong><span>for</span></strong> (Instruction &amp;I : BB)
        <strong><span>if</span></strong> (isa&lt;AllocaInst&gt;(I)) ++n;
    errs() &lt;&lt; F.getName() &lt;&lt; <span>&#8220;: &#8220;</span> &lt;&lt; n &lt;&lt; <span>&#8220; allocas\n&#8221;</span>;
    <strong><span>return</span></strong> PreservedAnalyses::all();   <em><span>// we changed nothing</span></em>
  }
};

<strong><span>extern</span></strong> <span>&#8220;C&#8221;</span> LLVM_ATTRIBUTE_WEAK ::llvm::PassPluginLibraryInfo
llvmGetPassPluginInfo() {
  <strong><span>return</span></strong> {LLVM_PLUGIN_API_VERSION, <span>&#8220;CountAllocas&#8221;</span>, LLVM_VERSION_STRING,
    [](PassBuilder &amp;PB) {
      PB.registerPipelineParsingCallback(
        [](StringRef Name, FunctionPassManager &amp;FPM,
           ArrayRef&lt;PassBuilder::PipelineElement&gt;) {
          <strong><span>if</span></strong> (Name == <span>&#8220;count-allocas&#8221;</span>) {
            FPM.addPass(CountAllocas()); <strong><span>return</span></strong> <strong><span>true</span></strong>;
          }
          <strong><span>return</span></strong> <strong><span>false</span></strong>;
        });
    }};
}</code></pre><pre><code>clang++ -shared -fPIC CountAllocas.cpp -o CountAllocas.so \
  $(llvm-config --cxxflags)
opt -load-pass-plugin=./CountAllocas.so \
    -passes=<span>&#8220;count-allocas&#8221;</span> -disable-output sum.bc</code></pre><p><em>The IR was the innovation everyone cites. The library architecture was the innovation that let ten thousand projects reuse a compiler without inheriting one.</em></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Five capabilities no one else has together</h2><p>A twenty-year-old assertion from the CGO 2004 paper that still holds</p><p>In their 2004 paper at the <strong>International Symposium on Code Generation and Optimization</strong>, Lattner and Adve made a specific claim: LLVM provides a combination of five capabilities that no other system offers at once. </p><p>The paper won the conference&#8217;s Most Influential Paper award ten years later, and the retrospective restates the claim as still true more than twenty years on. The five:</p><p>CapabilityWhat it means<strong>Persistent program information</strong>The IR can be preserved across an application&#8217;s whole lifetime, so optimization can happen at compile time, link time, load time, run time and even idle time between runs on the user&#8217;s machine.</p><p><strong>Offline code generation</strong>Regardless of persistence, code can be lowered to efficient native machine code offline, using techniques too expensive to run at runtime.<strong>User-specific profiling</strong>Profile data can be gathered from deployed software as the end user actually runs it, then fed back into <strong>profile-guided optimization</strong>.</p><p><strong>Transparent runtime model</strong>The IR imposes no object model, exception semantics or runtime, so it serves any language or mix of languages.<strong>Uniform whole-program compilation</strong>Because it is language-independent, an entire application, including its language runtime and system libraries, can be optimized together after linking.</p><p>The argument is not that each property is unique; it is that the <em>conjunction</em> is. Managed virtual machines like the<strong> JVM or the .NET CLI</strong> deliver user-specific profiling and partially deliver persistence and whole-program compilation, but they do not deliver a transparent runtime model, which is what confines them to a language family. </p><p>And when they do offer offline code generation, the authors note, they do it at the expense of persistence and profiling. GPU instruction sets such as CUDA&#8217;s PTX and <strong>OpenCL&#8217;s SPIR-V,</strong> and the JITs behind JavaScript, Python and WebAssembly, cover some combination of persistence and profiling and none of the rest. </p><p>GCC, superb at offline code generation, was never designed to ship a language-independent representation for lifelong compilation at all.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Qrpx!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d2d4b76-b556-4b61-9477-2eaad8127164_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Qrpx!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d2d4b76-b556-4b61-9477-2eaad8127164_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!Qrpx!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d2d4b76-b556-4b61-9477-2eaad8127164_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!Qrpx!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d2d4b76-b556-4b61-9477-2eaad8127164_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!Qrpx!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d2d4b76-b556-4b61-9477-2eaad8127164_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Qrpx!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d2d4b76-b556-4b61-9477-2eaad8127164_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/5d2d4b76-b556-4b61-9477-2eaad8127164_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Matrix comparing LLVM, JVM/.NET, GPU instruction sets, browser JITs and GCC across the five capabilities. Only LLVM is 'full' on all five.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Matrix comparing LLVM, JVM/.NET, GPU instruction sets, browser JITs and GCC across the five capabilities. Only LLVM is 'full' on all five." title="Matrix comparing LLVM, JVM/.NET, GPU instruction sets, browser JITs and GCC across the five capabilities. Only LLVM is 'full' on all five." srcset="https://substackcdn.com/image/fetch/$s_!Qrpx!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d2d4b76-b556-4b61-9477-2eaad8127164_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!Qrpx!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d2d4b76-b556-4b61-9477-2eaad8127164_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!Qrpx!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d2d4b76-b556-4b61-9477-2eaad8127164_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!Qrpx!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d2d4b76-b556-4b61-9477-2eaad8127164_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Figure 2.</strong> The five-capabilities matrix, encoded from the claims in the CGO 2004 paper and the 2026 retrospective. Teal marks full support, gold partial, empty none. LLVM is the only row that is full across the board, which is the entire thesis of the paper.</figcaption></figure></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>From IR to silicon</h2><p>Instruction selection, <strong>register allocation</strong>, and the <strong>TableGen</strong> machine</p><p>Optimized IR is still target-independent. Turning it into machine code is the backend&#8217;s job, and LLVM has, notably, two backends that do it.</p><p>The mature path is <strong>SelectionDAG</strong>. Each basic block is turned into a directed acyclic graph of operations; the DAG is legalized so every node is something the target can actually do; instruction selection matches subgraphs against target patterns to pick real machine instructions; the result is scheduled into a linear order; and then register allocation and emission follow. </p><p>The newer path is <strong>GlobalISel</strong>, designed to eventually replace <strong>SelectionDAG</strong>. It works on the whole function rather than one block at a time and skips the DAG entirely: an <code>IRTranslator</code> lowers IR to generic machine instructions, a <code>Legalizer</code> makes them legal, <code>RegBankSelect</code> assigns register banks, and an <code>InstructionSelect</code> pass chooses final opcodes. </p><p>As of the current release <strong>GlobalISel</strong> is the default at <code>-O0</code> on AArch64 and is used for AMD GPUs, and, as an ongoing <a href="https://discourse.llvm.org/t/status-on-enabling-globalisel-by-default-on-clang-on-aarch64/89964">LLVM forum discussion from February 2026</a> shows, work to enable it more broadly on AArch64 is still live. Two instruction selectors, maintained in parallel for years, is itself a statement about how much the project values getting this layer right.</p><p>Register allocation is where virtual registers, of which the IR has infinitely many, meet the finite physical register file. The default at <code>-O1</code> and above is the <strong>greedy</strong> allocator, which orders live ranges by priority and spills to the stack when it must; <code>-O0</code> uses a fast allocator that trades quality for speed. </p><p>Both are among the passes now being handed to machine learning, a thread we pick up in section 10.</p><p>Almost none of this is hand-written per target. Targets are described declaratively in <strong>TableGen</strong>, a domain-specific language whose <code>.td</code> files specify registers, instructions, calling conventions, scheduling models and selection patterns; a generator turns them into the C++ tables the backend uses. </p><p>A <strong>single instruction definition</strong> ties an assembly syntax to a SelectionDAG pattern:</p><pre><code><strong><span>def</span></strong> ADDrr : Instruction {
  <strong><span>let</span></strong> Namespace       = <span>&#8220;MyTarget&#8221;</span>;
  <strong><span>let</span></strong> OutOperandList = (outs GPR:$dst);
  <strong><span>let</span></strong> InOperandList  = (ins GPR:$src1, GPR:$src2);
  <strong><span>let</span></strong> AsmString      = <span>&#8220;add $dst, $src1, $src2&#8221;</span>;
  <em><span>// match (add a, b) in the DAG and emit this instruction</span></em>
  <strong><span>let</span></strong> Pattern = [(set <span>i32</span>:$dst, (add <span>i32</span>:$src1, <span>i32</span>:$src2))];
}</code></pre><p>This is why standing up a new architecture in LLVM, <strong>RISC-V</strong> being the most visible recent example, is a matter of writing target descriptions rather than a compiler. </p><p>The LLVM 22 notes alone record new scheduling models for <strong>NVIDIA&#8217;s Olympus CPU</strong>, additional Arm C1 cores, Intel Wildcat Lake and Nova Lake, and fresh RISC-V vector extensions, all arriving as target-description work on top of the shared infrastructure.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share The Software Frontier&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share The Software Frontier</span></a></p><div><hr></div><h2>The link-time frontier, and why ThinLTO exists</h2><p>Whole-program optimization that actually scales to a datacenter build</p><p><strong>Persistent IR</strong> unlocks link-time optimization: instead of discarding each source file&#8217;s IR after compiling it, you keep the bitcode, hand all of it to the linker, and optimize across module boundaries, inlining a function defined in one file into a caller in another, propagating constants program-wide. </p><p>Classic <strong>full LTO</strong> merges every module into one giant IR module and optimizes it as a unit. The result is excellent and the cost is brutal: the whole program has to fit in one process, and the work is largely serial, which is a non-starter when the program is a datacenter binary with tens of thousands of translation units.</p><p><strong>ThinLTO</strong>, first presented at EuroLLVM 2015 by <strong>Teresa Johnson</strong>, <strong>Mehdi Amini</strong> and <strong>David Li</strong> and merged into LLVM over the following year, is the answer that made link-time optimization practical at industrial scale. Each module is compiled to bitcode alongside a compact <em>summary</em>: a thin index of its call graph and symbols. </p><p>A cheap global step reads only the summaries, decides cross-module actions such as which functions to import for inlining, and then the heavy per-module optimization runs in parallel, importing only the specific functions each module needs. It recovers<strong> most of full LTO&#8217;s benefit</strong> at a fraction of the memory, and it parallelizes. The numbers are quite stark. </p><p>Full LTO can inflate compile time two to three fold and demands that the whole program fit in a single process; ThinLTO&#8217;s per-module summaries add under one percent to bitcode size,<strong> roughly 0.8 percent for Clang </strong>built without debug information, and keep build time and peak memory close to a non-LTO build. </p><p>On some benchmarks ThinLTO even edges out full LTO, because its scalable design lets each module run a more aggressive backend pipeline than a single monolithic module could afford. </p><p>LLVM 22 is now upstreaming <strong>Distributed ThinLTO</strong> (DTLTO), which pushes those parallel backend jobs across a build cluster, with caching for incremental rebuilds. Turning it on is a flag.</p><pre><code>clang -flto=thin -O2 a.c b.c c.c -o app   <em><span># scalable cross-module LTO</span></em>
clang -flto      -O2 a.c b.c c.c -o app   <em><span># full (monolithic) LTO</span></em></code></pre><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!UmjI!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5931e02a-ba05-45f2-b860-b9884c2c7654_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!UmjI!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5931e02a-ba05-45f2-b860-b9884c2c7654_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!UmjI!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5931e02a-ba05-45f2-b860-b9884c2c7654_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!UmjI!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5931e02a-ba05-45f2-b860-b9884c2c7654_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!UmjI!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5931e02a-ba05-45f2-b860-b9884c2c7654_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!UmjI!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5931e02a-ba05-45f2-b860-b9884c2c7654_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/5931e02a-ba05-45f2-b860-b9884c2c7654_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Horizontal bar chart of compile time relative to a normal build: No LTO 1.0x, ThinLTO about 1x, Full LTO 2x to 3x.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Horizontal bar chart of compile time relative to a normal build: No LTO 1.0x, ThinLTO about 1x, Full LTO 2x to 3x." title="Horizontal bar chart of compile time relative to a normal build: No LTO 1.0x, ThinLTO about 1x, Full LTO 2x to 3x." srcset="https://substackcdn.com/image/fetch/$s_!UmjI!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5931e02a-ba05-45f2-b860-b9884c2c7654_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!UmjI!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5931e02a-ba05-45f2-b860-b9884c2c7654_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!UmjI!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5931e02a-ba05-45f2-b860-b9884c2c7654_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!UmjI!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5931e02a-ba05-45f2-b860-b9884c2c7654_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Figure 3.</strong> Why ThinLTO exists. Full LTO buys cross-module optimization at 2 to 3 times the compile time and needs the whole program in one process; ThinLTO adds under 1 percent to bitcode, stays close to a normal build, and parallelizes. Only the 2 to 3x figure is a hard number (Gentoo LTO documentation and the LLVM ThinLTO write-up); the ThinLTO bar is indicative of its near-baseline behavior.</figcaption></figure></div><p>This is the mechanism underneath one of the headline claims: that all of <strong>Google&#8217;s and Meta&#8217;s datacenter C and C++ </strong>is built with LLVM. It is not merely that Clang compiles the files; it is that link-time optimization across an entire service is tractable because <em>ThinLTO made it so.</em></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/p/how-llvm-works-the-ir-that-took-over?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/p/how-llvm-works-the-ir-that-took-over?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><div><hr></div><h2>What LLVM actually runs</h2><p>The reach, with the adoption dates checked against primary sources</p><p>The scale claims in the retrospective are large, and they hold up. LLVM long ago<strong> replaced GCC</strong> as the foundation of Apple&#8217;s software across every device line. Google uses it for Android and the Android NDK. Applications on Qualcomm&#8217;s Snapdragon processors are built in substantial part with LLVM. </p><p>The<strong> two dominant CPU vendors</strong> for PCs and servers, Arm and Intel, retired their proprietary compilers, ARMCC and ICC, in favor of LLVM-based toolchains. Google&#8217;s and Meta&#8217;s datacenter C and C++ is compiled with it. Sony&#8217;s PlayStation 4 and 5 and the Nintendo Switch use Clang as their primary application toolchain. </p><p>And the AI stack leans on it heavily, through NVIDIA&#8217;s CUDA compiler, through the Triton compiler used by PyTorch and OpenAI, and most broadly through <strong>MLIR</strong>.</p><p>The adoption did not happen at once, and the sequence matters, because it shows a single seed (<em>Apple, via Lattner&#8217;s hiring</em>) catalyzing an industry.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Kmcg!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe9e30f2d-ab4b-435a-b82b-8b5f16bf82ef_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Kmcg!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe9e30f2d-ab4b-435a-b82b-8b5f16bf82ef_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!Kmcg!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe9e30f2d-ab4b-435a-b82b-8b5f16bf82ef_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!Kmcg!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe9e30f2d-ab4b-435a-b82b-8b5f16bf82ef_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!Kmcg!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe9e30f2d-ab4b-435a-b82b-8b5f16bf82ef_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Kmcg!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe9e30f2d-ab4b-435a-b82b-8b5f16bf82ef_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e9e30f2d-ab4b-435a-b82b-8b5f16bf82ef_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Timeline of LLVM industry adoption from 2005 to 2022: Apple begins, macOS OpenGL JIT ships, Clang open-sourced, PS4 ships with Clang, Apple ships bitcode, ARM and Intel adopt LLVM, Apple deprecates bitcode.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Timeline of LLVM industry adoption from 2005 to 2022: Apple begins, macOS OpenGL JIT ships, Clang open-sourced, PS4 ships with Clang, Apple ships bitcode, ARM and Intel adopt LLVM, Apple deprecates bitcode." title="Timeline of LLVM industry adoption from 2005 to 2022: Apple begins, macOS OpenGL JIT ships, Clang open-sourced, PS4 ships with Clang, Apple ships bitcode, ARM and Intel adopt LLVM, Apple deprecates bitcode." srcset="https://substackcdn.com/image/fetch/$s_!Kmcg!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe9e30f2d-ab4b-435a-b82b-8b5f16bf82ef_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!Kmcg!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe9e30f2d-ab4b-435a-b82b-8b5f16bf82ef_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!Kmcg!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe9e30f2d-ab4b-435a-b82b-8b5f16bf82ef_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!Kmcg!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe9e30f2d-ab4b-435a-b82b-8b5f16bf82ef_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Figure 4.</strong> Adoption milestones with dates verified against primary sources. Apple began serious LLVM work around 2005; the first shipping product was the macOS OpenGL JIT in 2006. The App Store accepted app bitcode from the mid-2010s until Apple deprecated it in Xcode 14 (2022).</figcaption></figure></div><p>The bitcode chapter is the sharpest illustration of the &#8220;<em>lifelong compilation</em>&#8221; idea reaching production. For years, developers shipped iOS,<strong> watchOS </strong>and <strong>tvOS apps</strong> to Apple&#8217;s App Store not as machine code but as LLVM bitcode, and Apple compiled and specialized each app for the target device on its side. </p><p>That let one submission serve <strong>32-bit and 64-bit Arm iPhones</strong>, low-power Apple Watch variants and Apple TV, shrank downloads, and let Apple roll out new compiler optimizations without developers rebuilding. </p><p>Apple deprecated bitcode in <a href="https://developer.apple.com/documentation/xcode-release-notes/xcode-14-release-notes">Xcode 14 in 2022</a>, once its hardware had converged on Arm64 and the flexibility was no longer worth the friction. The feature retired, but it stands as the largest deployment of shippable IR the industry has seen.</p><h3>Sizing the surface area</h3><p>The authors put a figure on all this: software compiled with LLVM, they write, represents<strong> hundreds of billions of dollars</strong> in annual revenue across mobile, cloud, desktop, server, gaming, supercomputing and AI. </p><p>We think that undercounts it, and it is worth being precise about what can actually be measured, because &#8220;<em>revenue compiled by LLVM</em>&#8221; is not a single clean quantity. </p><p>The honest move is to measure the top-line revenue of the platforms whose flagship products are demonstrably built with LLVM, treat it as a floor rather than an attribution, and be explicit that this is revenue riding on <strong>LLVM-compiled software</strong>, not value the compiler captures.</p><p>Do that and the picture is not subtle. In their most recent full fiscal years three companies alone, Apple, Alphabet and NVIDIA, booked <strong>about 1.04 trillion dollars in combined revenue</strong>, and the flagship products behind every one of those dollars are built with LLVM: Apple&#8217;s entire device line through Clang and Swift, Android and Google&#8217;s hyperscale C and C++ through Clang and ThinLTO, and the AI-accelerator stack through the <strong>CUDA</strong>, <strong>Triton </strong>and <strong>MLIR compilers</strong> that lower into LLVM.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!ekQK!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd8d54fdb-6d29-4e59-9d46-2d34403a05b9_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!ekQK!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd8d54fdb-6d29-4e59-9d46-2d34403a05b9_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!ekQK!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd8d54fdb-6d29-4e59-9d46-2d34403a05b9_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!ekQK!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd8d54fdb-6d29-4e59-9d46-2d34403a05b9_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!ekQK!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd8d54fdb-6d29-4e59-9d46-2d34403a05b9_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!ekQK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd8d54fdb-6d29-4e59-9d46-2d34403a05b9_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d8d54fdb-6d29-4e59-9d46-2d34403a05b9_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Horizontal bar chart of most-recent-fiscal-year revenue for Apple (416 billion dollars), Alphabet (403 billion dollars) and NVIDIA (216 billion dollars), summing to about 1.04 trillion dollars, all with flagship products built using LLVM.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Horizontal bar chart of most-recent-fiscal-year revenue for Apple (416 billion dollars), Alphabet (403 billion dollars) and NVIDIA (216 billion dollars), summing to about 1.04 trillion dollars, all with flagship products built using LLVM." title="Horizontal bar chart of most-recent-fiscal-year revenue for Apple (416 billion dollars), Alphabet (403 billion dollars) and NVIDIA (216 billion dollars), summing to about 1.04 trillion dollars, all with flagship products built using LLVM." srcset="https://substackcdn.com/image/fetch/$s_!ekQK!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd8d54fdb-6d29-4e59-9d46-2d34403a05b9_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!ekQK!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd8d54fdb-6d29-4e59-9d46-2d34403a05b9_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!ekQK!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd8d54fdb-6d29-4e59-9d46-2d34403a05b9_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!ekQK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd8d54fdb-6d29-4e59-9d46-2d34403a05b9_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Figure 5.</strong> Three companies whose flagship products are compiled with LLVM, and their most recent full fiscal year of top-line revenue: Apple 416 billion dollars (FY2025), Alphabet 403 billion dollars (2025) and NVIDIA 216 billion dollars (fiscal 2026, ended January 25, 2026). The combined 1.04 trillion dollars is a deliberate floor: it excludes Meta, Microsoft, Qualcomm, Sony, Nintendo and the entire Android hardware market beyond Google. Figures from each company&#8217;s SEC filings.</figcaption></figure></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!J6kW!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f510ecd-d1f5-45cf-9c10-a4522e40c128_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!J6kW!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f510ecd-d1f5-45cf-9c10-a4522e40c128_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!J6kW!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f510ecd-d1f5-45cf-9c10-a4522e40c128_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!J6kW!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f510ecd-d1f5-45cf-9c10-a4522e40c128_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!J6kW!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f510ecd-d1f5-45cf-9c10-a4522e40c128_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!J6kW!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f510ecd-d1f5-45cf-9c10-a4522e40c128_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2f510ecd-d1f5-45cf-9c10-a4522e40c128_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:192104,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/203533011?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f510ecd-d1f5-45cf-9c10-a4522e40c128_1634x893.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!J6kW!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f510ecd-d1f5-45cf-9c10-a4522e40c128_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!J6kW!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f510ecd-d1f5-45cf-9c10-a4522e40c128_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!J6kW!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f510ecd-d1f5-45cf-9c10-a4522e40c128_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!J6kW!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f510ecd-d1f5-45cf-9c10-a4522e40c128_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>We <strong>do not sum the rows of that table</strong>: the smartphone market overlaps the Apple and Alphabet figures, and the point is the surface area, not a grand total. </p><p>But the<strong> three company revenues</strong> in Figure 4 are independently additive and clear a trillion dollars between them, which is why we are comfortable calling <strong>LLVM&#8217;s reach a trillion-dollar one. </strong>The precise figure is unauditable and we mark it accordingly in the dossier; the order of magnitude is not in doubt. </p><p>A large majority of end users across computing are, at some layer, running code that LLVM compiled.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Tq0r!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F27ec6d5c-6d7e-4659-845f-0cc80f75014b_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Tq0r!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F27ec6d5c-6d7e-4659-845f-0cc80f75014b_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!Tq0r!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F27ec6d5c-6d7e-4659-845f-0cc80f75014b_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!Tq0r!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F27ec6d5c-6d7e-4659-845f-0cc80f75014b_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!Tq0r!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F27ec6d5c-6d7e-4659-845f-0cc80f75014b_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Tq0r!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F27ec6d5c-6d7e-4659-845f-0cc80f75014b_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/27ec6d5c-6d7e-4659-845f-0cc80f75014b_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Horizontal bar chart of active devices running LLVM-compiled system software: Android more than 3 billion, Apple 2.5 billion, PlayStation and Switch several hundred million.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Horizontal bar chart of active devices running LLVM-compiled system software: Android more than 3 billion, Apple 2.5 billion, PlayStation and Switch several hundred million." title="Horizontal bar chart of active devices running LLVM-compiled system software: Android more than 3 billion, Apple 2.5 billion, PlayStation and Switch several hundred million." srcset="https://substackcdn.com/image/fetch/$s_!Tq0r!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F27ec6d5c-6d7e-4659-845f-0cc80f75014b_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!Tq0r!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F27ec6d5c-6d7e-4659-845f-0cc80f75014b_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!Tq0r!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F27ec6d5c-6d7e-4659-845f-0cc80f75014b_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!Tq0r!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F27ec6d5c-6d7e-4659-845f-0cc80f75014b_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Figure 6.</strong> The same reach, counted in devices rather than dollars, which is the more defensible floor. More than 3 billion active Android devices (Google&#8217;s own stated figure), 2.5 billion active Apple devices (January 2026), and several hundred million PlayStation and Switch consoles, all running system software built with LLVM. More than 5.8 billion devices, before the Android hardware Google does not build or the datacenters they run against.</figcaption></figure></div><h3><strong>The academic and language dividend</strong></h3><p>Two quieter forms of impact compound over time. LLVM underlies a generation of programming languages that an easy-to-target, modular back end made feasible: <strong>Swift</strong>, <strong>Rust</strong>, <strong>Julia</strong>, <strong>Halide</strong> and <strong>Mojo</strong> among them.</p><p> Rust&#8217;s memory-safety guarantees and Julia&#8217;s JIT-driven scientific performance both rest on LLVM code generation. </p><p>And in the academy, LLVM democratized compiler research and teaching: a graduate student can write a novel pass in a day and test it on large real programs immediately, and the authors report meeting undergraduates who arrived at college <strong>having already built compiler projects</strong> on LLVM in high school, &#8220;<em>for fun.</em>&#8221;</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/p/how-llvm-works-the-ir-that-took-over/comments&quot;,&quot;text&quot;:&quot;Leave a comment&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/p/how-llvm-works-the-ir-that-took-over/comments"><span>Leave a comment</span></a></p><div><hr></div><h2>MLIR, and the compiler layer under modern AI</h2><p>When one IR was not enough, the answer was a framework for making IRs</p><p>This is the part that matters most for anyone working on inference infrastructure, so we will spend a moment on it. <strong>LLVM IR is deliberately low-level,</strong> which is a strength for representing machines and a weakness for representing a matrix multiply. </p><p>By the time a tensor operation has been shredded into scalar loads, adds and branches, the structure a compiler most wants to exploit, the fact that this is a <strong>dense linear-algebra kernel</strong>, is gone. That loss is fine for C. It is expensive for machine learning, where the highest-value optimizations (tiling, fusion, layout, targeting a systolic array) live at exactly the level LLVM IR throws away.</p><p><strong>MLIR</strong>, the Multi-Level Intermediate Representation, introduced in 2019 and presented at<strong> CGO 2021</strong>, is the response, and it is characteristically an infrastructure answer rather than a point solution. Instead of one IR, MLIR provides the machinery to define many coexisting IRs, called <em>dialects</em>, that all share the same structural substrate of operations, regions and blocks. </p><p>A compiler keeps a computation at a high level of abstraction as long as that is useful, then <em>progressively lowers</em> it, dialect by dialect, until it reaches the <code>llvm</code> dialect and finally <strong>ordinary LLVM IR</strong>, where the mature backends take over. A tensor program can begin like this:</p><pre><code><em><span>// a matmul on tensors, before any lowering</span></em>
<strong><span>func.func</span></strong> <span>@matmul</span>(%A: <span>tensor</span>&lt;128x256xf32&gt;,
                  %B: <span>tensor</span>&lt;256x512xf32&gt;,
                  %C: <span>tensor</span>&lt;128x512xf32&gt;) -&gt; <span>tensor</span>&lt;128x512xf32&gt; {
  %0 = <strong><span>linalg.matmul</span></strong>
         ins(%A, %B : <span>tensor</span>&lt;128x256xf32&gt;, <span>tensor</span>&lt;256x512xf32&gt;)
         outs(%C   : <span>tensor</span>&lt;128x512xf32&gt;) -&gt; <span>tensor</span>&lt;128x512xf32&gt;
  <strong><span>return</span></strong> %0 : <span>tensor</span>&lt;128x512xf32&gt;
}</code></pre><p>From there the same operation descends through structured and loop dialects (<code>linalg</code>, <code>affine</code>, <code>scf</code>, <code>tensor</code>), gets its buffers assigned (<code>memref</code>, <code>bufferization</code>, <code>vector</code>), is<strong> retargeted to a device dialect</strong> (<code>gpu</code>, <code>nvvm</code>, <code>rocdl</code>, <code>spirv</code>), and finally lands in the <code>llvm</code> dialect and LLVM IR. </p><p>Each step is a rewrite in the same framework, and each step is where a tiling or fusion decision can be made while the structure is still visible.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!fL2Q!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5e84afc1-9774-4a35-9342-23999d8f60d4_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!fL2Q!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5e84afc1-9774-4a35-9342-23999d8f60d4_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!fL2Q!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5e84afc1-9774-4a35-9342-23999d8f60d4_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!fL2Q!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5e84afc1-9774-4a35-9342-23999d8f60d4_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!fL2Q!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5e84afc1-9774-4a35-9342-23999d8f60d4_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!fL2Q!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5e84afc1-9774-4a35-9342-23999d8f60d4_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/5e84afc1-9774-4a35-9342-23999d8f60d4_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Layered stack showing MLIR progressive lowering from framework graph through high-level dialects, structured and loop dialects, buffer and vector dialects, target dialects, the LLVM dialect, LLVM IR, and finally target machine code.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Layered stack showing MLIR progressive lowering from framework graph through high-level dialects, structured and loop dialects, buffer and vector dialects, target dialects, the LLVM dialect, LLVM IR, and finally target machine code." title="Layered stack showing MLIR progressive lowering from framework graph through high-level dialects, structured and loop dialects, buffer and vector dialects, target dialects, the LLVM dialect, LLVM IR, and finally target machine code." srcset="https://substackcdn.com/image/fetch/$s_!fL2Q!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5e84afc1-9774-4a35-9342-23999d8f60d4_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!fL2Q!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5e84afc1-9774-4a35-9342-23999d8f60d4_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!fL2Q!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5e84afc1-9774-4a35-9342-23999d8f60d4_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!fL2Q!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5e84afc1-9774-4a35-9342-23999d8f60d4_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Figure 7.</strong> Progressive lowering in MLIR. The compiler holds tensor-level structure at the top and converts it step by step into the same LLVM IR that Clang and Rust emit, then to PTX, x86-64, AArch64 or RISC-V. The stack narrows toward machine code to suggest the convergence onto shared infrastructure.</figcaption></figure></div><p>For an inference audience the payoff of this layering has a name: fusion. Because <strong>progressive lowering</strong> keeps tensor-level structure visible until late, the compiler can fuse a chain of operations, a matmul followed by a bias add followed by an activation, into a single kernel, so the intermediate tensors are never written out to high-bandwidth memory and read back. </p><p>That matters because decode, the<strong> token-by-token phase</strong> of LLM serving, is bound by memory bandwidth rather than arithmetic: the accelerator spends its time moving weights and activations, not multiplying them. </p><p>Collapsing those <strong>HBM round trips</strong> is therefore one of the largest levers on tokens per second per GPU, and it is a compiler transformation, made at exactly the abstraction level that raw LLVM IR throws away. </p><p>This is the concrete reason every serious inference stack now ships a compiler rather than a<strong> fixed library of kernels</strong>, and why the layer we are describing sits directly upstream of inference cost.</p><p>This is not a research curiosity. MLIR is the backbone of TensorFlow&#8217;s XLA compiler and the broader <strong>OpenXLA ecosystem</strong>, of parts of PyTorch through <strong>Torch-MLIR</strong>, of IREE, and of OpenAI&#8217;s Triton, the kernel language a great deal of contemporary GPU work is written in. Even components of NVIDIA&#8217;s own CUDA toolkit use it. </p><p>And the<strong> underlying LLVM backend</strong> still does the final, unglamorous work of register allocation and scheduling for all of them. When people say the AI stack runs on LLVM, this is the mechanism they are pointing at: not that models are<strong> written in C</strong>, but that the compilers turning tensor graphs into GPU code are, almost universally, MLIR lowering into LLVM. </p><p>The hardware-design world is following the same path through <strong>CIRCT</strong>, which extends MLIR into digital circuit design.</p><div><hr></div><h2>The one thing LLVM is bad at</h2><p>Compile speed, and the backends being built to route around it</p><p>A retrospective written by the project&#8217;s founders is not the place to look for LLVM&#8217;s weaknesses, so we will supply the missing side. </p><p>LLVM is engineered for the quality of the code it emits, and it pays for that in the time it takes to emit it. Even with optimizations disabled the <strong>pipeline is heavy</strong>, and for a developer sitting in an edit, compile, run loop the wait is a real tax. </p><p>As bjorn3, the author of Rust&#8217;s <strong>Cranelift</strong> backend, puts it, LLVM is optimized for output quality at the cost of compilation speed even when optimizations are turned off. This is the most common and most legitimate complaint about the project, and two major language communities are now building their own code generators specifically to escape it.</p><p>Rust took the incremental path. Its alternative <code>rustc_codegen_cranelift</code> <strong>backend uses Cranelift</strong>, a code generator built for speed rather than peak output and aimed squarely at debug builds. </p><p>The<strong> Rust team&#8217;s 2025 measurements</strong> on large real projects, among them Zed, Tauri and hickory-dns, show roughly a twenty percent cut in code-generation time and about a five percent speedup in total clean-build time, and getting the backend production-ready for local development is an official<strong> 2025 project goal. </strong></p><p>The bargain is stated plainly: Cranelift produces code almost as fast as LLVM with optimizations off, in exchange for compiling far more quickly.</p><p><strong>Zig</strong> went further and wrote its own backends outright. In mid-2025 it made its self-hosted x86-64 backend the default for debug builds on Linux and macOS, <strong>bypassing LLVM entirely</strong> on that path, and reported wall-clock compile improvements ranging from five to fifty percent across real projects. </p><p>The rationale is a direct indictment: because compilation time was dominated by LLVM, and because <strong>LLVM&#8217;s heavy use of shared state</strong> kept Zig&#8217;s LLVM path effectively single-threaded while its own backend parallelized code generation across cores, the only way to get materially faster was to leave. </p><p>Notably, <strong>Zig still emits LLVM bitcode </strong>when it wants LLVM&#8217;s optimizer for release builds. That is the pattern worth seeing.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!CMmQ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F05f0cce8-c325-473e-b99a-30f557519b34_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!CMmQ!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F05f0cce8-c325-473e-b99a-30f557519b34_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!CMmQ!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F05f0cce8-c325-473e-b99a-30f557519b34_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!CMmQ!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F05f0cce8-c325-473e-b99a-30f557519b34_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!CMmQ!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F05f0cce8-c325-473e-b99a-30f557519b34_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!CMmQ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F05f0cce8-c325-473e-b99a-30f557519b34_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/05f0cce8-c325-473e-b99a-30f557519b34_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Horizontal bar chart of compile-time improvement from non-LLVM backends: Rust Cranelift about 20 percent on codegen and 5 percent on clean build, Zig self-hosted backend 5 to 50 percent wall clock.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Horizontal bar chart of compile-time improvement from non-LLVM backends: Rust Cranelift about 20 percent on codegen and 5 percent on clean build, Zig self-hosted backend 5 to 50 percent wall clock." title="Horizontal bar chart of compile-time improvement from non-LLVM backends: Rust Cranelift about 20 percent on codegen and 5 percent on clean build, Zig self-hosted backend 5 to 50 percent wall clock." srcset="https://substackcdn.com/image/fetch/$s_!CMmQ!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F05f0cce8-c325-473e-b99a-30f557519b34_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!CMmQ!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F05f0cce8-c325-473e-b99a-30f557519b34_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!CMmQ!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F05f0cce8-c325-473e-b99a-30f557519b34_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!CMmQ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F05f0cce8-c325-473e-b99a-30f557519b34_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Figure 8.</strong> The cost of LLVM&#8217;s code quality, measured. Rust&#8217;s Cranelift backend cuts code-generation time by about 20 percent and clean-build time by about 5 percent (Rust project goals, 2025); Zig&#8217;s self-hosted x86-64 backend, default for debug builds since mid-2025, reports 5 to 50 percent faster wall-clock compiles (Zig devlog, 2025). The three bars measure different denominators. Both ecosystems keep LLVM for optimized release builds.</figcaption></figure></div><p><em>Keep LLVM for the optimized release build, where its code quality is still unmatched. Route around it for the fast debug build, where its latency hurts.</em></p><h3>The shape of both the Rust and the Zig response to LLVM&#8217;s one real weakness</h3><p>That shared pattern is not a threat to the empire so much as a map of its single soft border, and it is worth being honest that the border exists. There is a second, quieter tension living inside LLVM&#8217;s own success.</p><p><strong>MLIR&#8217;s openness</strong> invites a proliferation of dialects, and a representation that can be anything risks fragmenting into many incompatible somethings; the whole force of the original LLVM argument was that the IR won because it was <em>one</em> thing. </p><p><em>&#8220;The IR won because it was one representation</em>&#8221; and &#8220;<em>the future is many coexisting representations</em>&#8221; are in genuine tension, and how MLIR holds its dialects together as<strong> their number grows</strong> is one of the more interesting open questions in the field. So far the shared substrate has held. </p><p>The bet is that it keeps holding.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Where LLVM goes next</h2><p>Verified compilation, auto-generated backends, machine-learned heuristics, hardened code</p><p>The retrospective closes on active research, and several threads are worth stating precisely because they point at the next decade of the project.</p><h3>Proving the optimizer correct</h3><p>A compiler bug silently miscompiles correct source into wrong machine code, and optimizers are where these bugs breed. </p><p><strong>Alive2</strong> <em>(Nuno Lopes and colleagues, PLDI 2021)</em> attacks this with bounded translation validation: it takes an IR function before and after a transformation and uses an SMT solver to check whether the rewrite is sound for all inputs, <strong>encoding LLVM&#8217;s tricky semantics</strong>, including undefined behavior and <code>poison</code>, into the query. </p><p>Alive2 runs continuously against LLVM and has found real miscompilations in core passes such as <code>InstCombine</code>. The <code>ptrtoaddr</code> instruction we met in section 02 is part of the same trend: the <strong>project is steadily sharpening</strong> its own semantics so that this kind of verification can go further.</p><h3>Generating the compiler from the chip</h3><p>Two projects challenge the assumption that backends and peephole optimizations must be written by hand. <strong>Hydride</strong> (ASPLOS 2024) and <strong>MISAAL</strong> (<em>published in the ACM&#8217;s Proceedings on Programming Languages, 2025</em>) use program synthesis to <strong>generate compiler components</strong>, target-independent IR operations, front-end translators and backend code generators, directly from a vendor&#8217;s formal specification of an instruction set. </p><p>The reported result is striking: <strong>auto-generated compilers</strong> that match or beat a heavily hand-engineered production compiler for Halide while requiring roughly an order of magnitude less manual effort. </p><p>As instruction sets multiply, matrix and vector extensions arriving on <strong>nearly every new chip, </strong>generating the compiler from the specification rather than writing it by hand is a plausible future for the backend.</p><h3>Machine learning inside the compiler, and about the compiler</h3><p>There are two distinct AI-and-compiler stories, and they are easy to confuse. The first is <strong>machine learning inside LLVM</strong>: Google&#8217;s MLGO framework replaces hand-tuned heuristics for inlining-for-size and register-allocation eviction with policies trained by reinforcement learning, and these have shipped in production builds. </p><p>The second is <strong>LLVM as training data for models about code</strong>. Meta&#8217;s <strong>LLM Compiler</strong> (2024), built on Code Llama and published at Compiler Construction 2025, was trained on a corpus of <strong>546 billion tokens</strong> of LLVM IR and assembly, then instruction-tuned to predict the effect of optimization passes; the flag-tuning and disassembly variants add a further 164 billion tokens for 710 billion in total. </p><p>Meta reports the model reaching <strong>77 percent of the optimization benefit</strong> of an autotuning search, and a 45 percent round-trip on disassembling machine code back to IR. </p><p>The reason such a model can exist at all is the reason <strong>LLVM matters everywhere else</strong>: the sheer volume of LLVM-compiled open source makes a compiler-scale training corpus possible.</p><h4>One number, corrected</h4><p>The CACM retrospective states the LLM Compiler was fine-tuned on &#8220;<em>537 billion tokens.</em>&#8221; The figure in Meta&#8217;s paper and in the peer-reviewed Compiler Construction 2025 version is <strong>546 billion tokens</strong> of LLVM IR and assembly. We use 546 billion, which is what the primary source reports.</p><h3>Hardened code and heterogeneous parallelism</h3><p>Two more directions round out the picture. </p><p>On security, LLVM and Clang are the delivery vehicle for a generation of hardening: forward-edge control-flow integrity, shadow call stacks, and the Arm hardware defenses of pointer authentication (<em>PAC</em>), branch target identification (<em>BTI</em>) and memory tagging (<em>MTE</em>), alongside the sanitizer family (<em>AddressSanitizer, ThreadSanitizer, MemorySanitizer, UndefinedBehaviorSanitizer</em>) that made whole classes of C and C++ bugs findable. </p><p>LLVM 22 even added the <strong>AArch64 build-attribute machinery</strong> that lets linkers reason about PAC and BTI compatibility across object files. </p><p>On parallelism, the <strong>HPVM</strong> project (IEEE Micro 2022) layers a hierarchical dataflow graph over LLVM IR to target heterogeneous systems, CPUs, GPUs and accelerators, from a single representation, and national laboratories are extending LLVM for exascale machines through efforts such as <strong>Los Alamos&#8217;s Kitsune </strong>and the NNSA-backed Flang Fortran front end.</p><div class="captioned-button-wrap" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/p/how-llvm-works-the-ir-that-took-over?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="CaptionedButtonToDOM"><div class="preamble"><p class="cta-caption">Thanks for reading The Software Frontier! This post is public so feel free to share it.</p></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/p/how-llvm-works-the-ir-that-took-over?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/p/how-llvm-works-the-ir-that-took-over?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p></div><div><hr></div><h2>What the story is actually about</h2><p>Cadence, and the transferable lesson under the technology</p><p>Strip away the specifics and LLVM is a case study in how foundational infrastructure gets built, and the ingredients are unglamorous. Federal money funded <strong>early-stage research </strong>whose payoff no one could forecast, on the condition that the results be released openly. </p><p>An academic effort took known techniques, SSA, separation of concerns, retargetable code generation, interprocedural analysis, and applied all of them together more thoroughly than any production compiler had. </p><p>A commercially neutral open license, eventually <strong>Apache 2.0 with LLVM exceptions</strong>, let the result be reused, again and again, in products that competed with each other. Strong technical and community leadership, first inside Apple and then across the industry, drove adoption. </p><p>And a family of successful derivatives, Clang, Swift and MLIR chief among them, widened the reach far beyond the original core.</p><p>The result compounds on a schedule. Twenty-three years after the first release, <strong>LLVM ships a full toolchain</strong>, Clang, LLD, LLDB, libc++, compiler-rt, MLIR, Flang, Polly, BOLT, on a strict six-month train.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!bIew!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F490824b8-4388-44ed-9df1-c0f1d0380112_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!bIew!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F490824b8-4388-44ed-9df1-c0f1d0380112_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!bIew!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F490824b8-4388-44ed-9df1-c0f1d0380112_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!bIew!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F490824b8-4388-44ed-9df1-c0f1d0380112_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!bIew!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F490824b8-4388-44ed-9df1-c0f1d0380112_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!bIew!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F490824b8-4388-44ed-9df1-c0f1d0380112_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/490824b8-4388-44ed-9df1-c0f1d0380112_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Stepped timeline of LLVM major versions from 1.0 in 2003 to 22 in 2026, showing an irregular early period and a steady six-month cadence after version 4 in 2017, annotated with IR milestones.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Stepped timeline of LLVM major versions from 1.0 in 2003 to 22 in 2026, showing an irregular early period and a steady six-month cadence after version 4 in 2017, annotated with IR milestones." title="Stepped timeline of LLVM major versions from 1.0 in 2003 to 22 in 2026, showing an irregular early period and a steady six-month cadence after version 4 in 2017, annotated with IR milestones." srcset="https://substackcdn.com/image/fetch/$s_!bIew!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F490824b8-4388-44ed-9df1-c0f1d0380112_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!bIew!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F490824b8-4388-44ed-9df1-c0f1d0380112_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!bIew!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F490824b8-4388-44ed-9df1-c0f1d0380112_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!bIew!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F490824b8-4388-44ed-9df1-c0f1d0380112_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Figure 9.</strong> Every LLVM major version against its release date, with IR milestones marked: the new pass manager becoming default (v13), opaque pointers by default (v15), typed pointers removed (v17) and <strong>ptrtoaddr</strong> added (v22). The clear elbow at version 4 in 2017 is where the project settled into a predictable half-year release cadence.</figcaption></figure></div><p><strong>Adve </strong>and <strong>Lattner </strong>end their retrospective on the claim that a small, federally funded academic project produced &#8220;<em>world-changing, foundational infrastructure,</em>&#8221; and on the evidence it is not an overstatement. </p><p>The instruction set that Lattner sketched over a winter break in 2000, three isomorphic forms of one strongly typed, target-independent, shippable IR, is visibly the <strong>same architecture running underneath the compiler</strong> in every phone, console, datacenter and AI accelerator we have discussed. </p><p>The empire is real. It was always, underneath, an argument about a representation.</p><div><hr></div><h2>Verification dossier</h2><p>Every load-bearing claim above, checked against a primary or authoritative source. Confidence tiers: <strong>A</strong> primary source, direct confirmation &#183; <strong>B</strong> authoritative secondary &#183; <strong>C</strong> reasonable inference or the authors&#8217; own estimate &#183; <strong>D</strong> flagged discrepancy or figure we corrected.</p><p><strong><span>A. LLVM 1.0 was first released in October 2003</span></strong></p><p>Chris Lattner&#8217;s announcement, &#8220;The LLVM 1.0 Release is finally available!&#8221;, is dated October 24, 2003; Wikidata records the same inception date. <strong><span>Discrepancy noted:</span></strong> the CACM retrospective says &#8220;October 2003&#8221; in its introduction but &#8220;December 2003&#8221; in two later places. October 24, 2003 is correct; the December date appears to be the release-notes page&#8217;s last-modified timestamp (Dec 8, 2003), not the release.</p><p><strong><span>A. Current stable release is LLVM 22.1, from February 2026</span></strong></p><p>LLVM/Clang 22.1 was released February 24, 2026 per the project and Phoronix; the latest point release, 22.1.8, is dated June 16, 2026 on GitHub. LLVM 23 is in development as 23.0.0git. Consistent with the project&#8217;s post-2017 six-month cadence.</p><p><strong><span>A. The CGO 2004 paper won Most Influential Paper; the team won the 2012 ACM Software System Award</span></strong></p><p>&#8220;LLVM: A Compilation Framework for Lifelong Program Analysis and Transformation&#8221; (CGO 2004, pp. 75 to 86, DOI 10.1109/CGO.2004.1281665) was named Most Influential Paper of CGO 2004, awarded 2014. <strong><span>Nuance:</span></strong> the 2012 ACM Software System Award went to Adve, Lattner <strong>and Evan Cheng</strong>; the CACM retrospective names only the two authors. The &#8220;more than 8,000 citations&#8221; figure is the authors&#8217; own and is conservative against public citation indices.</p><p><strong><span>D. The Meta LLM Compiler was trained on 546 billion tokens, not 537 billion</span></strong></p><p>The CACM piece states 537 billion tokens. Meta&#8217;s paper (arXiv 2407.02524) and the peer-reviewed Compiler Construction 2025 version (DOI 10.1145/3708493.3712691) both report <strong>546 billion tokens</strong> of LLVM IR and assembly, with a further 164 billion for flag-tuning and disassembly (710 billion total). We use 546 billion.</p><p><strong><span>A. Opaque pointers became default in LLVM 15; typed pointers were removed in LLVM 17</span></strong></p><p>Confirmed by the official Opaque Pointers documentation: default in LLVM 15 (2022), typed-pointer support removed in LLVM 17 (2023). The <code>ptrtoaddr</code> instruction is a genuine LLVM 22 addition, per the 22.1 release coverage.</p><p><strong><span>A. Apple deprecated bitcode in Xcode 14 (2022)</span></strong></p><p>Xcode 14 release notes: the App Store no longer accepts bitcode submissions and Xcode no longer builds bitcode by default. This bounds the &#8220;mid-2010s through 2022&#8221; window the retrospective gives for App Store bitcode.</p><p><strong><span>B. Industry adoption: Apple, Google, Arm, Intel, Sony, Nintendo, Qualcomm</span></strong></p><p>Apple&#8217;s OpenGL/iOS SDK work began around 2005 and the macOS OpenGL JIT shipped in 2006 (Clang and LLVM histories). Intel switched its C/C++ stack to LLVM (ICX / oneAPI DPC++); Arm shipped LLVM-based Arm Compiler 6 and drove the Apache 2.0 relicensing over patent concerns. Sony uses Clang for PS4/PS5. These are well documented across primary announcements and vendor pages.</p><p><strong><span>C. &#8221;Hundreds of billions of dollars in annual revenue&#8221; compiled by LLVM</span></strong></p><p>This is the authors&#8217; order-of-magnitude estimate, not an audited figure, and we present it as such. The qualitative claim, that a large majority of end users run LLVM-compiled code at some layer, is well supported by the adoption evidence above.</p><p><strong><span>A. MLIR underpins XLA, PyTorch (Torch-MLIR), IREE and Triton; GlobalISel is default at -O0 on AArch64</span></strong></p><p>MLIR was presented at CGO 2021 (DOI 10.1109/CGO51591.2021.9370308) and is used across the named AI compilers. GlobalISel&#8217;s default-at-O0 status on AArch64 and ongoing work to broaden it are documented in current LLVM forum threads (February 2026).</p><p><strong><span>B. ThinLTO: first presented at EuroLLVM 2015; summaries add under one percent to bitcode; full LTO costs two to three times the compile time</span></strong></p><p>ThinLTO was introduced by Teresa Johnson, Mehdi Amini and David Li (LLVM Project blog, 2016; presented EuroLLVM 2015). The summary overhead of about 0.8 percent for Clang without debug info, and the design goal of build time and memory close to a non-LTO build, are from that write-up and the Clang ThinLTO documentation. The two-to-three-fold compile-time cost of full LTO is a widely reported figure (for example, the Gentoo LTO documentation). &#8220;Sometimes beats full LTO&#8221; reflects the authors&#8217; own SPEC results, where ThinLTO&#8217;s more aggressive per-module pipeline occasionally wins.</p><p><strong><span>C. LLVM&#8217;s reach can be sized as a trillion-dollar surface area</span></strong></p><p>Apple FY2025 revenue 416.16 billion dollars (Form 8-K, quarter and year ended Sept 27, 2025); Alphabet 2025 revenue 402.84 billion dollars (Form 10-K / Annual Report 2025); NVIDIA fiscal 2026 revenue 215.9 billion dollars, data center 193.7 billion, for the year ended Jan 25, 2026 (Form 8-K). Combined, about 1.04 trillion dollars. This is top-line company revenue for firms whose flagship products are built with LLVM, presented as a floor, not value attributed to or captured by the compiler. Smartphone units of about 1.25 billion for 2025 are from IDC and Omdia. We do not treat the constructed total as audited.</p><p><strong><span>A. Rust (Cranelift) and Zig are building non-LLVM backends to escape LLVM compile times</span></strong></p><p>Rust&#8217;s <code>rustc_codegen_cranelift</code> reports roughly a twenty percent reduction in code-generation time and about five percent in clean-build time on large projects, and production-readiness is a 2025 Rust project goal (rust-project-goals, 2025h2). Zig made its self-hosted x86-64 backend the default for Debug builds on Linux and macOS in mid-2025, reporting five to fifty percent wall-clock improvements and citing LLVM as the compile-time bottleneck (Zig devlog, 2025). Both retain an LLVM path for optimized release builds.</p><div><hr></div><h2><strong>Sources</strong></h2><ol><li><p>V. Adve and C. Lattner, &#8220;The LLVM Compiler Infrastructure,&#8221; <em>Communications of the ACM</em>, June 2026. <a href="https://cacm.acm.org/federal-funding-of-academic-research/the-llvm-compiler-infrastructure/">cacm.acm.org</a></p></li><li><p>C. Lattner and V. Adve, &#8220;LLVM: A Compilation Framework for Lifelong Program Analysis and Transformation,&#8221; <em>CGO 2004</em>, pp. 75 to 86. DOI 10.1109/CGO.2004.1281665.</p></li><li><p>C. Lattner, &#8220;The LLVM 1.0 Release is finally available!&#8221;, llvm-announce, October 24, 2003.</p></li><li><p>LLVM Project, &#8220;Opaque Pointers.&#8221; <a href="https://llvm.org/docs/OpaquePointers.html">llvm.org/docs/OpaquePointers.html</a></p></li><li><p>Arm, &#8220;What is new in LLVM 22,&#8221; 2026; M. Larabel, &#8220;LLVM/Clang 22.1 Released,&#8221; <em>Phoronix</em>, Feb 24, 2026.</p></li><li><p>LLVM Project, release tags through 22.1.8. <a href="https://github.com/llvm/llvm-project/releases">github.com/llvm/llvm-project</a></p></li><li><p>Apple, &#8220;Xcode 14 Release Notes,&#8221; 2022. <a href="https://developer.apple.com/documentation/xcode-release-notes/xcode-14-release-notes">developer.apple.com</a></p></li><li><p>N. P. Lopes et al., &#8220;Alive2: Bounded Translation Validation for LLVM,&#8221; <em>PLDI 2021</em>. DOI 10.1145/3453483.3454030.</p></li><li><p>C. Lattner et al., &#8220;MLIR: Scaling Compiler Infrastructure for Domain Specific Computation,&#8221; <em>CGO 2021</em>. DOI 10.1109/CGO51591.2021.9370308.</p></li><li><p>C. Cummins et al., &#8220;Meta Large Language Model Compiler,&#8221; arXiv 2407.02524, 2024; <em>CC 2025</em>, DOI 10.1145/3708493.3712691.</p></li><li><p>A. Kothari et al., &#8220;Hydride,&#8221; <em>ASPLOS 2024</em>, DOI 10.1145/3620665.3640385; A. R. Noor et al., &#8220;MISAAL,&#8221; <em>PACMPL 2025</em>.</p></li><li><p>A. Ejjeh et al., &#8220;HPVM: Hardware-agnostic programming for heterogeneous parallel systems,&#8221; <em>IEEE Micro</em> 42(5), 2022.</p></li><li><p>V. Adve, faculty biography and CV, University of Illinois. <a href="https://vikram.cs.illinois.edu/bio/">vikram.cs.illinois.edu</a></p></li><li><p>T. Johnson, M. Amini, D. Li, &#8220;ThinLTO: Scalable and Incremental LTO,&#8221; LLVM Project Blog, 2016 (presented EuroLLVM 2015); Clang &#8220;ThinLTO&#8221; documentation. <a href="https://blog.llvm.org/2016/06/thinlto-scalable-and-incremental-lto.html">blog.llvm.org</a></p></li><li><p>Rust Project Goals 2025h2, &#8220;Production-ready Cranelift backend.&#8221; <a href="https://rust-lang.github.io/rust-project-goals/2025h2/production-ready-cranelift.html">rust-lang.github.io</a></p></li><li><p>Zig, Devlog 2025 (self-hosted x86-64 backend default for Debug builds). <a href="https://ziglang.org/devlog/2025/">ziglang.org/devlog/2025</a></p></li><li><p>Apple Inc., Form 8-K, fiscal 2025 fourth-quarter and full-year results (year ended Sept 27, 2025), revenue 416 billion dollars. <a href="https://www.sec.gov/cgi-bin/browse-edgar?action=getcompany&amp;CIK=0000320193">SEC EDGAR</a></p></li><li><p>Alphabet Inc., Annual Report / Form 10-K for 2025, total revenue 402.84 billion dollars. <a href="https://www.sec.gov/cgi-bin/browse-edgar?action=getcompany&amp;CIK=0001652044">SEC EDGAR</a></p></li><li><p>NVIDIA Corp., Form 8-K, fourth quarter and fiscal 2026 (year ended Jan 25, 2026), revenue 215.9 billion dollars, data center 193.7 billion. <a href="https://www.sec.gov/cgi-bin/browse-edgar?action=getcompany&amp;CIK=0001045810">SEC EDGAR</a></p></li><li><p>IDC and Omdia, worldwide smartphone shipments for 2025, about 1.25 billion units.</p></li></ol><p><strong>Inference.Engineering</strong> &#183;<em> Independent analysis of LLM serving infrastructure, GPU economics, and the compiler layer beneath them. This piece is built on the June 2026 CACM retrospective by Vikram Adve and Chris Lattner and cross-checked against LLVM 22.1 documentation, release notes and the primary research literature. Vendor-independent; every figure sourced. Corrections and disputes welcome.</em></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/p/how-llvm-works-the-ir-that-took-over/comments&quot;,&quot;text&quot;:&quot;Leave a comment&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/p/how-llvm-works-the-ir-that-took-over/comments"><span>Leave a comment</span></a></p><p></p>]]></content:encoded></item><item><title><![CDATA[The Wafer & the Wallet]]></title><description><![CDATA[Memory costs are inflecting up as enterprise token budgets slam shut. The wafer math, the pass-through proof, and the engineering playbook that defends the spread.]]></description><link>https://www.thesoftwarefrontier.com/p/the-wafer-and-the-wallet</link><guid isPermaLink="false">https://www.thesoftwarefrontier.com/p/the-wafer-and-the-wallet</guid><dc:creator><![CDATA[Lorenzo Bradanini]]></dc:creator><pubDate>Sun, 05 Jul 2026 07:56:35 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!0y1r!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc81fedd3-e26c-4b0d-9cc1-718e1c0f97e2_1122x1402.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!0y1r!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc81fedd3-e26c-4b0d-9cc1-718e1c0f97e2_1122x1402.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!0y1r!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc81fedd3-e26c-4b0d-9cc1-718e1c0f97e2_1122x1402.png 424w, https://substackcdn.com/image/fetch/$s_!0y1r!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc81fedd3-e26c-4b0d-9cc1-718e1c0f97e2_1122x1402.png 848w, https://substackcdn.com/image/fetch/$s_!0y1r!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc81fedd3-e26c-4b0d-9cc1-718e1c0f97e2_1122x1402.png 1272w, https://substackcdn.com/image/fetch/$s_!0y1r!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc81fedd3-e26c-4b0d-9cc1-718e1c0f97e2_1122x1402.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!0y1r!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc81fedd3-e26c-4b0d-9cc1-718e1c0f97e2_1122x1402.png" width="1122" height="1402" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c81fedd3-e26c-4b0d-9cc1-718e1c0f97e2_1122x1402.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1402,&quot;width&quot;:1122,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2521188,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/203533001?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc81fedd3-e26c-4b0d-9cc1-718e1c0f97e2_1122x1402.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!0y1r!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc81fedd3-e26c-4b0d-9cc1-718e1c0f97e2_1122x1402.png 424w, https://substackcdn.com/image/fetch/$s_!0y1r!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc81fedd3-e26c-4b0d-9cc1-718e1c0f97e2_1122x1402.png 848w, https://substackcdn.com/image/fetch/$s_!0y1r!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc81fedd3-e26c-4b0d-9cc1-718e1c0f97e2_1122x1402.png 1272w, https://substackcdn.com/image/fetch/$s_!0y1r!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc81fedd3-e26c-4b0d-9cc1-718e1c0f97e2_1122x1402.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>Intro </h2><p>For three years, the cost of making a token fell every quarter while the budget to buy tokens grew without a meter. In <strong>H1 2026 both curves reversed</strong> at once: memory pass-through is raising the cost floor of every token served, while enterprises institutionalize token budgets at the ceiling. </p><p>This is the anatomy of <strong>inference&#8217;s first margin squeeze</strong>: the wafer math, the proof of pass-through, modeled cost floors marked to real market prices, and the engineering playbook that defends the spread.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Software Frontier is a reader-supported publication. Please, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!PC3u!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96bbb590-bfcc-40f5-99ac-caf2f1d5b3bb_1100x353.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!PC3u!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96bbb590-bfcc-40f5-99ac-caf2f1d5b3bb_1100x353.png 424w, https://substackcdn.com/image/fetch/$s_!PC3u!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96bbb590-bfcc-40f5-99ac-caf2f1d5b3bb_1100x353.png 848w, https://substackcdn.com/image/fetch/$s_!PC3u!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96bbb590-bfcc-40f5-99ac-caf2f1d5b3bb_1100x353.png 1272w, https://substackcdn.com/image/fetch/$s_!PC3u!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96bbb590-bfcc-40f5-99ac-caf2f1d5b3bb_1100x353.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!PC3u!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96bbb590-bfcc-40f5-99ac-caf2f1d5b3bb_1100x353.png" width="1100" height="353" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/96bbb590-bfcc-40f5-99ac-caf2f1d5b3bb_1100x353.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:353,&quot;width&quot;:1100,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:133192,&quot;alt&quot;:&quot;c00-masthead&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="c00-masthead" title="c00-masthead" srcset="https://substackcdn.com/image/fetch/$s_!PC3u!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96bbb590-bfcc-40f5-99ac-caf2f1d5b3bb_1100x353.png 424w, https://substackcdn.com/image/fetch/$s_!PC3u!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96bbb590-bfcc-40f5-99ac-caf2f1d5b3bb_1100x353.png 848w, https://substackcdn.com/image/fetch/$s_!PC3u!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96bbb590-bfcc-40f5-99ac-caf2f1d5b3bb_1100x353.png 1272w, https://substackcdn.com/image/fetch/$s_!PC3u!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96bbb590-bfcc-40f5-99ac-caf2f1d5b3bb_1100x353.png 1456w" sizes="100vw"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>Two curves reversed in the same half-year</h2><p>Every industry eventually meets the quarter when its two defining curves cross. For LLM inference, that quarter just happened. </p><p>Since late 2022, the economics of this business rested on two assumptions so reliable nobody wrote them down: <strong>the cost of producing a token falls every quarter</strong> (process nodes, kernel engineering, quantization, and a relentless price war), and <strong>the budget for buying tokens grows without a meter</strong>, the era of &#8220;tokenmaxxing,&#8221; when Meta and Salesforce were publicly urging employees to consume as many tokens as possible. In the first half of 2026, both assumptions died within months of each other.</p><p><strong>Jaw one: the cost floor turned upward.</strong> That is the core of this issue: the deepest memory supercycle in the forty-year history of DRAM, transmitted measurably into GPU rental rates with a 4-6 week lag ( Fig. 7). For the first time since the ChatGPT moment, <strong>the marginal cost of serving a token is inflecting up, not down</strong>: driven by the single most supply-constrained commodity in technology.</p><p><strong>Jaw two: the price ceiling got institutionalized.</strong> SemiAnalysis published field data on July 2 from conversations with 50+ enterprises: Uber burned through its annual Claude Code and Codex budget in four months and responded with a<strong> $1,500/month per-employee cap</strong>; an aerospace manufacturer&#8217;s $250/month caps were exhausted by power users in four days; companies are switching off premium model tiers and downgrading defaults; enterprise budgets now range from $250 to tens of thousands per employee per month, but they are <em>budgets</em>, reviewed by finance. </p><p> Read the finding precisely, because it is not a demand-collapse story: SemiAnalysis concludes there is no material risk to 2H26 AI budgets and API spend keeps compounding. <strong>The jaw is not falling demand, it is the arrival of price sensitivity.</strong> </p><p>Unbounded willingness-to-pay is over; every serving provider now negotiates against a CFO&#8217;s cap instead of an enthusiast&#8217;s curiosity.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!zbOW!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F603f7041-8239-4c16-a157-e9ccd588f95f_1100x690.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!zbOW!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F603f7041-8239-4c16-a157-e9ccd588f95f_1100x690.png 424w, https://substackcdn.com/image/fetch/$s_!zbOW!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F603f7041-8239-4c16-a157-e9ccd588f95f_1100x690.png 848w, https://substackcdn.com/image/fetch/$s_!zbOW!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F603f7041-8239-4c16-a157-e9ccd588f95f_1100x690.png 1272w, https://substackcdn.com/image/fetch/$s_!zbOW!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F603f7041-8239-4c16-a157-e9ccd588f95f_1100x690.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!zbOW!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F603f7041-8239-4c16-a157-e9ccd588f95f_1100x690.png" width="1100" height="690" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/603f7041-8239-4c16-a157-e9ccd588f95f_1100x690.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:690,&quot;width&quot;:1100,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;c-fig0&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="c-fig0" title="c-fig0" srcset="https://substackcdn.com/image/fetch/$s_!zbOW!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F603f7041-8239-4c16-a157-e9ccd588f95f_1100x690.png 424w, https://substackcdn.com/image/fetch/$s_!zbOW!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F603f7041-8239-4c16-a157-e9ccd588f95f_1100x690.png 848w, https://substackcdn.com/image/fetch/$s_!zbOW!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F603f7041-8239-4c16-a157-e9ccd588f95f_1100x690.png 1272w, https://substackcdn.com/image/fetch/$s_!zbOW!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F603f7041-8239-4c16-a157-e9ccd588f95f_1100x690.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!9PWX!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72d19ef9-e996-462d-bab6-03fab8436ba5_1100x473.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!9PWX!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72d19ef9-e996-462d-bab6-03fab8436ba5_1100x473.png 424w, https://substackcdn.com/image/fetch/$s_!9PWX!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72d19ef9-e996-462d-bab6-03fab8436ba5_1100x473.png 848w, https://substackcdn.com/image/fetch/$s_!9PWX!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72d19ef9-e996-462d-bab6-03fab8436ba5_1100x473.png 1272w, https://substackcdn.com/image/fetch/$s_!9PWX!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72d19ef9-e996-462d-bab6-03fab8436ba5_1100x473.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!9PWX!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72d19ef9-e996-462d-bab6-03fab8436ba5_1100x473.png" width="1100" height="473" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/72d19ef9-e996-462d-bab6-03fab8436ba5_1100x473.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:473,&quot;width&quot;:1100,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;c-callout0&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="c-callout0" title="c-callout0" srcset="https://substackcdn.com/image/fetch/$s_!9PWX!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72d19ef9-e996-462d-bab6-03fab8436ba5_1100x473.png 424w, https://substackcdn.com/image/fetch/$s_!9PWX!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72d19ef9-e996-462d-bab6-03fab8436ba5_1100x473.png 848w, https://substackcdn.com/image/fetch/$s_!9PWX!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72d19ef9-e996-462d-bab6-03fab8436ba5_1100x473.png 1272w, https://substackcdn.com/image/fetch/$s_!9PWX!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72d19ef9-e996-462d-bab6-03fab8436ba5_1100x473.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div><hr></div><h2>The quarter the memory market broke</h2><p>Start with the raw numbers, because they have no precedent in the forty-year history of the DRAM industry. TrendForce&#8217;s Q1 2026 contract-price revisions came in at <strong>+90-95% quarter-over-quarter for conventional DRAM</strong>: revised <em>upward</em> from an initial forecast of +55-60% that analysts had already called unprecedented. </p><p>PC DRAM was expected to more than double in a single quarter; server DDR5 rose roughly 90%; <strong>NAND contract prices </strong>climbed 33-38% in the same window. Counterpoint tracked spot DRAM up 80-90% inside the quarter. </p><p>For the full year, TrendForce projects DRAM up more than 70% <em>on top of</em> the Q1 step-change; Bank of America has DRAM industry revenue +51% YoY and NAND +45%.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!OkVP!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4fb89a05-21d2-440a-8e5c-7ad016d63e21_1100x571.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!OkVP!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4fb89a05-21d2-440a-8e5c-7ad016d63e21_1100x571.png 424w, https://substackcdn.com/image/fetch/$s_!OkVP!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4fb89a05-21d2-440a-8e5c-7ad016d63e21_1100x571.png 848w, https://substackcdn.com/image/fetch/$s_!OkVP!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4fb89a05-21d2-440a-8e5c-7ad016d63e21_1100x571.png 1272w, https://substackcdn.com/image/fetch/$s_!OkVP!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4fb89a05-21d2-440a-8e5c-7ad016d63e21_1100x571.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!OkVP!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4fb89a05-21d2-440a-8e5c-7ad016d63e21_1100x571.png" width="1100" height="571" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/4fb89a05-21d2-440a-8e5c-7ad016d63e21_1100x571.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:571,&quot;width&quot;:1100,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;c-fig1&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="c-fig1" title="c-fig1" srcset="https://substackcdn.com/image/fetch/$s_!OkVP!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4fb89a05-21d2-440a-8e5c-7ad016d63e21_1100x571.png 424w, https://substackcdn.com/image/fetch/$s_!OkVP!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4fb89a05-21d2-440a-8e5c-7ad016d63e21_1100x571.png 848w, https://substackcdn.com/image/fetch/$s_!OkVP!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4fb89a05-21d2-440a-8e5c-7ad016d63e21_1100x571.png 1272w, https://substackcdn.com/image/fetch/$s_!OkVP!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4fb89a05-21d2-440a-8e5c-7ad016d63e21_1100x571.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The supply side tells you why this is not a blip. <strong>SK Hynix announced in October 2025 that its entire 2026 production capacity was already sold out.</strong> Kioxia has said the same of its 2026 NAND output. </p><p>Micron shut down its Crucial consumer business outright (a two-decade brand, liquidated) to route every wafer toward <strong>hyperscaler and GPU-grade contracts</strong>, and its executives describe being able to cover at most two-thirds of medium-term demand for some customers.</p><p>Micron&#8217;s Boise fabs ramp in 2027-2028; the New York megafab in 2030. <strong>Jensen Huang</strong>, asked at CES 2026 whether gamers should resent AI for GPU prices, answered in effect that the world needs more memory factories. </p><blockquote><p><em>As stated even before, in the previous issues: when the largest memory buyer on Earth says the fix is on the supply side, the message downstream is: you will not negotiate your way to allocation relief.</em></p></blockquote><h3>Three structural breaks from every previous cycle</h3><p><strong>1. The demand driver is durable, not episodic.</strong> Prior shortages came from fab fires, earthquakes, supply discipline, or one-off demand pulses (crypto, pandemic PCs). This one comes from inference workloads compounding in production. IDC&#8217;s assessment calls it a <strong>&#8220;permanent reallocation&#8221; of the world&#8217;s wafer capacity, not a cyclical shortage</strong>, and IDC does not deploy the word &#8220;permanent&#8221; casually.</p><p><strong>2. The margin gradient is a one-way valve, and the wafer math is brutal.</strong> HBM carries gross margins 3-5&#215; commodity DRAM, so every rational fab reallocates cleanroom toward it. </p><p>But because of TSV formation, die thinning, stacking, and compounded assembly yield, <strong>one gigabyte of HBM consumes &#8776;4&#215; the wafer capacity of one gigabyte of standard DRAM; GDDR7 consumes &#8776;1.7&#215;</strong>. AI&#8217;s <em>effective</em> wafer claim reaches ~20% of world DRAM capacity in 2026 against total bit-supply growth of only 10-16% per year. </p><p>The arithmetic cannot close without starving someone, and &#8220;<em>someone</em>&#8221; is every non-AI buyer plus every AI buyer without a long-term agreement.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!6WWW!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3a88d9d-78ef-4e79-8bfa-b0f22bbd8935_1100x562.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!6WWW!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3a88d9d-78ef-4e79-8bfa-b0f22bbd8935_1100x562.png 424w, https://substackcdn.com/image/fetch/$s_!6WWW!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3a88d9d-78ef-4e79-8bfa-b0f22bbd8935_1100x562.png 848w, https://substackcdn.com/image/fetch/$s_!6WWW!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3a88d9d-78ef-4e79-8bfa-b0f22bbd8935_1100x562.png 1272w, https://substackcdn.com/image/fetch/$s_!6WWW!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3a88d9d-78ef-4e79-8bfa-b0f22bbd8935_1100x562.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!6WWW!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3a88d9d-78ef-4e79-8bfa-b0f22bbd8935_1100x562.png" width="1100" height="562" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a3a88d9d-78ef-4e79-8bfa-b0f22bbd8935_1100x562.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:562,&quot;width&quot;:1100,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;c-fig2&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="c-fig2" title="c-fig2" srcset="https://substackcdn.com/image/fetch/$s_!6WWW!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3a88d9d-78ef-4e79-8bfa-b0f22bbd8935_1100x562.png 424w, https://substackcdn.com/image/fetch/$s_!6WWW!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3a88d9d-78ef-4e79-8bfa-b0f22bbd8935_1100x562.png 848w, https://substackcdn.com/image/fetch/$s_!6WWW!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3a88d9d-78ef-4e79-8bfa-b0f22bbd8935_1100x562.png 1272w, https://substackcdn.com/image/fetch/$s_!6WWW!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3a88d9d-78ef-4e79-8bfa-b0f22bbd8935_1100x562.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><strong>3. The price </strong><em><strong>convergence</strong></em><strong> matters more than the price level.</strong> HBM3e historically sold at 4-5&#215; server DDR5 per bit. TrendForce expects that gap to compress to <strong>1-2&#215; by end-2026</strong>, not because HBM got cheaper, but because DDR5 inflated toward it. Sit with that. </p><p><strong>The entire tiered-memory architecture of modern serving (</strong><em>KV offload to host DRAM, paged hierarchies, DDR5-backed prefix caches, NVMe cold tiers</em><strong>) was designed in a world where the tier below HBM was 4-5&#215; cheaper per bit.</strong> </p><p>When the discount for descending a tier collapses from ~80% to ~30-50%, a large fraction of the &#8220;obvious&#8221; tiering optimizations of 2024-2025 stop being obvious. Section 5 does that math properly.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!J560!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feeefc492-5628-48c1-8f25-09ae71c75c19_1100x560.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!J560!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feeefc492-5628-48c1-8f25-09ae71c75c19_1100x560.png 424w, https://substackcdn.com/image/fetch/$s_!J560!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feeefc492-5628-48c1-8f25-09ae71c75c19_1100x560.png 848w, https://substackcdn.com/image/fetch/$s_!J560!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feeefc492-5628-48c1-8f25-09ae71c75c19_1100x560.png 1272w, https://substackcdn.com/image/fetch/$s_!J560!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feeefc492-5628-48c1-8f25-09ae71c75c19_1100x560.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!J560!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feeefc492-5628-48c1-8f25-09ae71c75c19_1100x560.png" width="1100" height="560" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/eeefc492-5628-48c1-8f25-09ae71c75c19_1100x560.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:560,&quot;width&quot;:1100,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;c-fig3&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="c-fig3" title="c-fig3" srcset="https://substackcdn.com/image/fetch/$s_!J560!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feeefc492-5628-48c1-8f25-09ae71c75c19_1100x560.png 424w, https://substackcdn.com/image/fetch/$s_!J560!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feeefc492-5628-48c1-8f25-09ae71c75c19_1100x560.png 848w, https://substackcdn.com/image/fetch/$s_!J560!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feeefc492-5628-48c1-8f25-09ae71c75c19_1100x560.png 1272w, https://substackcdn.com/image/fetch/$s_!J560!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feeefc492-5628-48c1-8f25-09ae71c75c19_1100x560.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>One supply-side note that matters for H2: <strong>HBM4 becomes the mainstream HBM in the second half of 2026</strong>, with a doubled 2,048-bit interface, ~2 TB/s per stack, and a manufacturing-complexity premium above 30% over HBM3e. NVIDIA&#8217;s Rubin ships with up to 288 GB of HBM4 per GPU, and every one of those stacks is wafer capacity that used to be laptops.</p><div><hr></div><h2>The KV cache ate the memory supply</h2><p>The lazy narrative is &#8220;AI ate the memory market.&#8221; The precise narrative is that <strong>the KV cache ate the memory market</strong>, and the industry&#8217;s shift from training-dominated to inference-dominated compute is what lit the fuse.</p><p>Training demand is large but <em>concentrated and schedulable</em>: a fixed pool of HBM for weights, activations, and optimizer states, planned quarters ahead. Inference demand is <em>elastic and per-user</em>: every concurrent conversation, agent trajectory, and <strong>200K-token codebase </strong>instantiates its own slab of state that must live somewhere in the hierarchy for the life of the request, and increasingly beyond it, as prefix caches persist context between turns. </p><p>Industry estimates via TrendForce put cloud high-speed memory consumption near <strong>3 exabytes in 2026</strong>, with core inference platforms (<em>Gemini-, Bedrock-, ChatGPT-class serving</em>) accounting for ~750 PB of <em>live</em> memory demand before redundancy roughly doubles it. </p><p>TrendForce separately projects that <strong>by 2029 inference (not training) becomes the primary driver of AI server demand outright</strong>, and North American CSPs are already bulk-buying high-capacity DDR5 RDIMMs specifically for inference fleets, the demand class that used to be the safety valve absorbing DRAM oversupply.</p><h3>The per-token physics, derived from first principles</h3><p>Everything in this issue hangs on one number: bytes of KV state per token. Take the workhorse case, a Llama-3-class 70B dense model with grouped-query attention: 80 layers, 8 KV heads, head dimension 128.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!zKAn!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faf0f9d6d-1829-4a06-b45c-6306b2c3dd3b_1100x449.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!zKAn!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faf0f9d6d-1829-4a06-b45c-6306b2c3dd3b_1100x449.png 424w, https://substackcdn.com/image/fetch/$s_!zKAn!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faf0f9d6d-1829-4a06-b45c-6306b2c3dd3b_1100x449.png 848w, https://substackcdn.com/image/fetch/$s_!zKAn!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faf0f9d6d-1829-4a06-b45c-6306b2c3dd3b_1100x449.png 1272w, https://substackcdn.com/image/fetch/$s_!zKAn!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faf0f9d6d-1829-4a06-b45c-6306b2c3dd3b_1100x449.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!zKAn!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faf0f9d6d-1829-4a06-b45c-6306b2c3dd3b_1100x449.png" width="1100" height="449" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/af0f9d6d-1829-4a06-b45c-6306b2c3dd3b_1100x449.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:449,&quot;width&quot;:1100,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;c-math0&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="c-math0" title="c-math0" srcset="https://substackcdn.com/image/fetch/$s_!zKAn!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faf0f9d6d-1829-4a06-b45c-6306b2c3dd3b_1100x449.png 424w, https://substackcdn.com/image/fetch/$s_!zKAn!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faf0f9d6d-1829-4a06-b45c-6306b2c3dd3b_1100x449.png 848w, https://substackcdn.com/image/fetch/$s_!zKAn!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faf0f9d6d-1829-4a06-b45c-6306b2c3dd3b_1100x449.png 1272w, https://substackcdn.com/image/fetch/$s_!zKAn!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faf0f9d6d-1829-4a06-b45c-6306b2c3dd3b_1100x449.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!QCjx!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fca93064a-232c-4fe8-b04b-919a5b7a5ea3_1100x625.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!QCjx!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fca93064a-232c-4fe8-b04b-919a5b7a5ea3_1100x625.png 424w, https://substackcdn.com/image/fetch/$s_!QCjx!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fca93064a-232c-4fe8-b04b-919a5b7a5ea3_1100x625.png 848w, https://substackcdn.com/image/fetch/$s_!QCjx!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fca93064a-232c-4fe8-b04b-919a5b7a5ea3_1100x625.png 1272w, https://substackcdn.com/image/fetch/$s_!QCjx!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fca93064a-232c-4fe8-b04b-919a5b7a5ea3_1100x625.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!QCjx!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fca93064a-232c-4fe8-b04b-919a5b7a5ea3_1100x625.png" width="1100" height="625" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ca93064a-232c-4fe8-b04b-919a5b7a5ea3_1100x625.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:625,&quot;width&quot;:1100,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;c-fig4&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="c-fig4" title="c-fig4" srcset="https://substackcdn.com/image/fetch/$s_!QCjx!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fca93064a-232c-4fe8-b04b-919a5b7a5ea3_1100x625.png 424w, https://substackcdn.com/image/fetch/$s_!QCjx!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fca93064a-232c-4fe8-b04b-919a5b7a5ea3_1100x625.png 848w, https://substackcdn.com/image/fetch/$s_!QCjx!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fca93064a-232c-4fe8-b04b-919a5b7a5ea3_1100x625.png 1272w, https://substackcdn.com/image/fetch/$s_!QCjx!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fca93064a-232c-4fe8-b04b-919a5b7a5ea3_1100x625.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>A single 128K sequence on the GQA-8 model pins <strong>~41 GB at FP16: half an H100&#8217;s entire HBM for one user&#8217;s context</strong>, before a single weight is stored. This is why effective batch size at long context is a memory-capacity calculation, not a compute calculation, and why <strong>every marginal concurrent user at long context is, economically, a memory purchase</strong>. </p><p>Multiply by the fleet: 10,000 concurrent long-context sequences at FP8 KV pins <strong>~200 TB of state across HBM</strong>, DDR5, and NVMe, state that produces no tokens itself; it exists purely so decode can stream it past the compute units at terabytes per second. </p><p>The <strong>2026 memory market</strong> is what happens when the industry&#8217;s aggregate KV state, growing super-linearly with agent adoption and context inflation, collides with wafer supply growing 10-16% a year.</p><h3>Two second-order effects complete the picture</h3><p><strong>NAND got dragged in.</strong> NVMe became the cold tier for prefix caches and the staging layer for weights (including GPU-direct storage paths). Kioxia now says nearly half its future NAND demand could come from AI. Meanwhile Samsung and SK Hynix <em>cut</em> NAND wafer output in 2024-2025 while chasing HBM margins. Omdia has Samsung&#8217;s NAND wafers falling 4.9M&#8594;4.68M and SK Hynix&#8217;s 1.9M&#8594;1.7M. <strong>There is no cheap tier left to hide in.</strong></p><p><strong>The host-DRAM tax on every GPU node inflated.</strong> A serious 2026 inference node pairs its GPUs with 1-2 TB of DDR5 for KV offload, CPU-side batching, and cache tiers. At Q1 2026 server-DRAM pricing, the host memory on a single node can cost what a mid-range GPU did two years ago. No 2024-vintage TCO model has this line item at anywhere near its current size.</p><div><hr></div><h2>Roofline: why prefill and decode were never the same workload</h2><p>Before the silicon story makes sense, you need the roofline argument stated with numbers, not vibes. A kernel&#8217;s attainable throughput is bounded by min(peak_FLOPs, intensity &#215; bandwidth), where <strong>arithmetic intensity</strong> is FLOPs performed per byte moved from memory. </p><p>The machine has a <em>ridge point</em>, the intensity at which it stops being bandwidth-bound and becomes compute-bound.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!4bPr!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F93058626-ab79-40c9-b56a-cf5c87b9acb2_1100x483.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!4bPr!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F93058626-ab79-40c9-b56a-cf5c87b9acb2_1100x483.png 424w, https://substackcdn.com/image/fetch/$s_!4bPr!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F93058626-ab79-40c9-b56a-cf5c87b9acb2_1100x483.png 848w, https://substackcdn.com/image/fetch/$s_!4bPr!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F93058626-ab79-40c9-b56a-cf5c87b9acb2_1100x483.png 1272w, https://substackcdn.com/image/fetch/$s_!4bPr!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F93058626-ab79-40c9-b56a-cf5c87b9acb2_1100x483.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!4bPr!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F93058626-ab79-40c9-b56a-cf5c87b9acb2_1100x483.png" width="1100" height="483" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/93058626-ab79-40c9-b56a-cf5c87b9acb2_1100x483.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:483,&quot;width&quot;:1100,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;c-math1&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="c-math1" title="c-math1" srcset="https://substackcdn.com/image/fetch/$s_!4bPr!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F93058626-ab79-40c9-b56a-cf5c87b9acb2_1100x483.png 424w, https://substackcdn.com/image/fetch/$s_!4bPr!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F93058626-ab79-40c9-b56a-cf5c87b9acb2_1100x483.png 848w, https://substackcdn.com/image/fetch/$s_!4bPr!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F93058626-ab79-40c9-b56a-cf5c87b9acb2_1100x483.png 1272w, https://substackcdn.com/image/fetch/$s_!4bPr!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F93058626-ab79-40c9-b56a-cf5c87b9acb2_1100x483.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!6tit!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa859114a-4861-45ca-8684-6e2ac283494e_1100x647.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!6tit!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa859114a-4861-45ca-8684-6e2ac283494e_1100x647.png 424w, https://substackcdn.com/image/fetch/$s_!6tit!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa859114a-4861-45ca-8684-6e2ac283494e_1100x647.png 848w, https://substackcdn.com/image/fetch/$s_!6tit!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa859114a-4861-45ca-8684-6e2ac283494e_1100x647.png 1272w, https://substackcdn.com/image/fetch/$s_!6tit!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa859114a-4861-45ca-8684-6e2ac283494e_1100x647.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!6tit!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa859114a-4861-45ca-8684-6e2ac283494e_1100x647.png" width="1100" height="647" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a859114a-4861-45ca-8684-6e2ac283494e_1100x647.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:647,&quot;width&quot;:1100,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;c-fig5&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="c-fig5" title="c-fig5" srcset="https://substackcdn.com/image/fetch/$s_!6tit!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa859114a-4861-45ca-8684-6e2ac283494e_1100x647.png 424w, https://substackcdn.com/image/fetch/$s_!6tit!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa859114a-4861-45ca-8684-6e2ac283494e_1100x647.png 848w, https://substackcdn.com/image/fetch/$s_!6tit!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa859114a-4861-45ca-8684-6e2ac283494e_1100x647.png 1272w, https://substackcdn.com/image/fetch/$s_!6tit!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa859114a-4861-45ca-8684-6e2ac283494e_1100x647.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The corollary that most cost models miss: <strong>decode throughput is not a FLOPs number, it is a bandwidth budget divided by bytes-per-step</strong>. Per decode step the device must stream the weights once plus every live sequence&#8217;s KV:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!sjTw!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faedb9217-181e-49c1-8ffb-258e768d104d_1100x381.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!sjTw!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faedb9217-181e-49c1-8ffb-258e768d104d_1100x381.png 424w, https://substackcdn.com/image/fetch/$s_!sjTw!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faedb9217-181e-49c1-8ffb-258e768d104d_1100x381.png 848w, https://substackcdn.com/image/fetch/$s_!sjTw!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faedb9217-181e-49c1-8ffb-258e768d104d_1100x381.png 1272w, https://substackcdn.com/image/fetch/$s_!sjTw!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faedb9217-181e-49c1-8ffb-258e768d104d_1100x381.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!sjTw!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faedb9217-181e-49c1-8ffb-258e768d104d_1100x381.png" width="1100" height="381" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/aedb9217-181e-49c1-8ffb-258e768d104d_1100x381.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:381,&quot;width&quot;:1100,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;c-math2&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="c-math2" title="c-math2" srcset="https://substackcdn.com/image/fetch/$s_!sjTw!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faedb9217-181e-49c1-8ffb-258e768d104d_1100x381.png 424w, https://substackcdn.com/image/fetch/$s_!sjTw!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faedb9217-181e-49c1-8ffb-258e768d104d_1100x381.png 848w, https://substackcdn.com/image/fetch/$s_!sjTw!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faedb9217-181e-49c1-8ffb-258e768d104d_1100x381.png 1272w, https://substackcdn.com/image/fetch/$s_!sjTw!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faedb9217-181e-49c1-8ffb-258e768d104d_1100x381.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Read those three lines again. Same GPU, same model, same bandwidth: <strong>16&#215; throughput difference between short-context and long-context traffic</strong>, and <strong>a clean 2&#215; recovered at 128K purely by halving KV bytes</strong>. Context length is not a feature flag; it is a cost multiplier that acts through memory. </p><p>Every<strong> KV byte you remove</strong> converts directly into either more concurrent sequences or faster steps: both of which are revenue.</p><div><hr></div><h2>Rubin CPX is a memory trade wearing a GPU costume</h2><p>Now read <strong>NVIDIA&#8217;s 2026 inference roadmap </strong>through the memory market and it snaps into focus in a way most launch coverage missed. If prefill is compute-bound and decode is bandwidth-bound (&#167;03), then serving both phases on <strong>identical HBM-stuffed GPUs</strong> means that during prefill you are paying for the most supply-constrained, margin-rich commodity on the planet (HBM bandwidth) <em>and not using it</em>. </p><p>In 2024 that was an inefficiency. In 2026, with HBM at 4&#215; wafer-equivalent cost and fully allocated into next year, <strong>it is an unforced error measured in real money</strong>.</p><p><strong>Rubin CPX is the correction.</strong> Announced at the AI Infra Summit in September 2025, detailed through CES and GTC 2026, generally available late this year: a monolithic prefill-specialized die pairing <strong>30 PFLOPS of sparse NVFP4 compute</strong><em> (20 PFLOPS dense, per SemiAnalysis)</em><strong> with 128 GB of GDDR7 at only ~2 TB/s of bandwidth</strong> (32 Gbps GDDR7 on a 512-bit bus. </p><p>That bandwidth would embarrass a decode GPU) it&#8217;s below an H100. On a prefill part it is the entire point: SemiAnalysis&#8217;s launch framing was exactly right, a chip deliberately skinny on bandwidth and fat on compute, because that is what prefill consumes. Its intensity budget sits far to the right of <strong>Fig. 5&#8217;s ridge, where GDDR7 </strong>is not a compromise but a correct sizing.</p><p>Overlay the wafer math and the arbitrage is explicit: GDDR7 costs 1.7&#215; standard DRAM wafer-equivalent; HBM costs 4&#215;, plus CoWoS packaging, interposer, and thermal budget. <strong>By building the prefill tier out of GDDR7, NVIDIA routes the fastest-growing slice of inference demand</strong><em> (long-context prompt processing) </em><strong>around the most constrained commodity in its own supply chain.</strong> </p><p>Every prefill FLOP served from GDDR7 is HBM allocation freed for decode, where it earns its keep.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!n-TM!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf8f4658-7538-401d-8414-7c00ee5969d1_1100x460.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!n-TM!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf8f4658-7538-401d-8414-7c00ee5969d1_1100x460.png 424w, https://substackcdn.com/image/fetch/$s_!n-TM!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf8f4658-7538-401d-8414-7c00ee5969d1_1100x460.png 848w, https://substackcdn.com/image/fetch/$s_!n-TM!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf8f4658-7538-401d-8414-7c00ee5969d1_1100x460.png 1272w, https://substackcdn.com/image/fetch/$s_!n-TM!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf8f4658-7538-401d-8414-7c00ee5969d1_1100x460.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!n-TM!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf8f4658-7538-401d-8414-7c00ee5969d1_1100x460.png" width="1100" height="460" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/cf8f4658-7538-401d-8414-7c00ee5969d1_1100x460.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:460,&quot;width&quot;:1100,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;c-table0&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="c-table0" title="c-table0" srcset="https://substackcdn.com/image/fetch/$s_!n-TM!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf8f4658-7538-401d-8414-7c00ee5969d1_1100x460.png 424w, https://substackcdn.com/image/fetch/$s_!n-TM!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf8f4658-7538-401d-8414-7c00ee5969d1_1100x460.png 848w, https://substackcdn.com/image/fetch/$s_!n-TM!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf8f4658-7538-401d-8414-7c00ee5969d1_1100x460.png 1272w, https://substackcdn.com/image/fetch/$s_!n-TM!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf8f4658-7538-401d-8414-7c00ee5969d1_1100x460.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>NVIDIA is pitching this with a number that deserves both scrutiny and attention : <strong>$5B of token revenue per $100M of infrastructure</strong>, 30-50&#215; platform ROI, and naming Cursor, Runway, and Magic as early partners, which tells you the target token distribution: repository-scale coding context and long-form generative media, i.e., prefill-dominated workloads. Discount the multiplier as marketing; the direction is load-bearing.</p><p>The strategic read for this audience: <strong>hardware disaggregation of prefill and decode is memory-market arbitrage instantiated in silicon, and it will not stay proprietary to NVIDIA</strong>. SRAM-based decode tiers are being positioned into the same disaggregated fabrics, AMD will be forced to answer, and every serious serving stack (<em>Dynamo, vLLM&#8217;s disagg mode, SGLang, Mooncake-style architectures</em>) has converged on <strong>KV-cache-transfer-over-fabric as the central abstraction of 2026 serving</strong>. </p><p>The clearest tell that this is permanent: per Tom&#8217;s Hardware&#8217;s platform teardown, the Vera Rubin rack&#8217;s BlueField-4 DPU ships with an <em>integrated SSD specifically to store KV cache</em>: cached context now has its own dedicated hardware tier in the reference design. </p><p>If your mental model of an &#8220;<em>inference GPU</em>&#8221; is one SKU doing both phases, you are one hardware generation behind the economics, and Q1&#8217;s memory prices just made that lag expensive.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!NaQE!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F95a817d3-72a1-4c0a-910b-a5ebef3df62f_1100x555.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!NaQE!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F95a817d3-72a1-4c0a-910b-a5ebef3df62f_1100x555.png 424w, https://substackcdn.com/image/fetch/$s_!NaQE!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F95a817d3-72a1-4c0a-910b-a5ebef3df62f_1100x555.png 848w, https://substackcdn.com/image/fetch/$s_!NaQE!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F95a817d3-72a1-4c0a-910b-a5ebef3df62f_1100x555.png 1272w, https://substackcdn.com/image/fetch/$s_!NaQE!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F95a817d3-72a1-4c0a-910b-a5ebef3df62f_1100x555.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!NaQE!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F95a817d3-72a1-4c0a-910b-a5ebef3df62f_1100x555.png" width="1100" height="555" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/95a817d3-72a1-4c0a-910b-a5ebef3df62f_1100x555.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:555,&quot;width&quot;:1100,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;c-callout1&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="c-callout1" title="c-callout1" srcset="https://substackcdn.com/image/fetch/$s_!NaQE!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F95a817d3-72a1-4c0a-910b-a5ebef3df62f_1100x555.png 424w, https://substackcdn.com/image/fetch/$s_!NaQE!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F95a817d3-72a1-4c0a-910b-a5ebef3df62f_1100x555.png 848w, https://substackcdn.com/image/fetch/$s_!NaQE!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F95a817d3-72a1-4c0a-910b-a5ebef3df62f_1100x555.png 1272w, https://substackcdn.com/image/fetch/$s_!NaQE!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F95a817d3-72a1-4c0a-910b-a5ebef3df62f_1100x555.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!quTc!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc975bc73-b912-4dc6-a8c5-9b0f5d200935_1100x433.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!quTc!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc975bc73-b912-4dc6-a8c5-9b0f5d200935_1100x433.png 424w, https://substackcdn.com/image/fetch/$s_!quTc!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc975bc73-b912-4dc6-a8c5-9b0f5d200935_1100x433.png 848w, https://substackcdn.com/image/fetch/$s_!quTc!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc975bc73-b912-4dc6-a8c5-9b0f5d200935_1100x433.png 1272w, https://substackcdn.com/image/fetch/$s_!quTc!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc975bc73-b912-4dc6-a8c5-9b0f5d200935_1100x433.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!quTc!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc975bc73-b912-4dc6-a8c5-9b0f5d200935_1100x433.png" width="1100" height="433" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c975bc73-b912-4dc6-a8c5-9b0f5d200935_1100x433.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:433,&quot;width&quot;:1100,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;c-callout2&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="c-callout2" title="c-callout2" srcset="https://substackcdn.com/image/fetch/$s_!quTc!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc975bc73-b912-4dc6-a8c5-9b0f5d200935_1100x433.png 424w, https://substackcdn.com/image/fetch/$s_!quTc!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc975bc73-b912-4dc6-a8c5-9b0f5d200935_1100x433.png 848w, https://substackcdn.com/image/fetch/$s_!quTc!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc975bc73-b912-4dc6-a8c5-9b0f5d200935_1100x433.png 1272w, https://substackcdn.com/image/fetch/$s_!quTc!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc975bc73-b912-4dc6-a8c5-9b0f5d200935_1100x433.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div><hr></div><h2>The repriced hierarchy, and what happens to cost per token</h2><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!4czk!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74d49549-f1d1-428f-b202-669ebefeea81_1100x672.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!4czk!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74d49549-f1d1-428f-b202-669ebefeea81_1100x672.png 424w, https://substackcdn.com/image/fetch/$s_!4czk!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74d49549-f1d1-428f-b202-669ebefeea81_1100x672.png 848w, https://substackcdn.com/image/fetch/$s_!4czk!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74d49549-f1d1-428f-b202-669ebefeea81_1100x672.png 1272w, https://substackcdn.com/image/fetch/$s_!4czk!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74d49549-f1d1-428f-b202-669ebefeea81_1100x672.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!4czk!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74d49549-f1d1-428f-b202-669ebefeea81_1100x672.png" width="1100" height="672" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/74d49549-f1d1-428f-b202-669ebefeea81_1100x672.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:672,&quot;width&quot;:1100,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;c-table1&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="c-table1" title="c-table1" srcset="https://substackcdn.com/image/fetch/$s_!4czk!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74d49549-f1d1-428f-b202-669ebefeea81_1100x672.png 424w, https://substackcdn.com/image/fetch/$s_!4czk!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74d49549-f1d1-428f-b202-669ebefeea81_1100x672.png 848w, https://substackcdn.com/image/fetch/$s_!4czk!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74d49549-f1d1-428f-b202-669ebefeea81_1100x672.png 1272w, https://substackcdn.com/image/fetch/$s_!4czk!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74d49549-f1d1-428f-b202-669ebefeea81_1100x672.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The uncomfortable synthesis: <strong>the price ratios between tiers compressed at the same moment the absolute levels rose</strong>, and tiering strategies earn their complexity from the <em>ratio</em>, not the level. A prefix cache offloading to DDR5 at 20-25% of HBM&#8217;s per-bit cost is an easy win. </p><p>The same cache at 50-70% of HBM&#8217;s per-bit cost must clear a much higher bar once you charge it for transfer bandwidth, PCIe hops, TTFT-on-miss, and the payroll that maintains it.</p><h3>The offload break-even, derived</h3><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!m4Pk!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F796ba650-9bfc-40aa-ac94-41d75ad0ed8a_1100x449.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!m4Pk!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F796ba650-9bfc-40aa-ac94-41d75ad0ed8a_1100x449.png 424w, https://substackcdn.com/image/fetch/$s_!m4Pk!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F796ba650-9bfc-40aa-ac94-41d75ad0ed8a_1100x449.png 848w, https://substackcdn.com/image/fetch/$s_!m4Pk!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F796ba650-9bfc-40aa-ac94-41d75ad0ed8a_1100x449.png 1272w, https://substackcdn.com/image/fetch/$s_!m4Pk!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F796ba650-9bfc-40aa-ac94-41d75ad0ed8a_1100x449.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!m4Pk!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F796ba650-9bfc-40aa-ac94-41d75ad0ed8a_1100x449.png" width="1100" height="449" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/796ba650-9bfc-40aa-ac94-41d75ad0ed8a_1100x449.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:449,&quot;width&quot;:1100,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;c-math3&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="c-math3" title="c-math3" srcset="https://substackcdn.com/image/fetch/$s_!m4Pk!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F796ba650-9bfc-40aa-ac94-41d75ad0ed8a_1100x449.png 424w, https://substackcdn.com/image/fetch/$s_!m4Pk!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F796ba650-9bfc-40aa-ac94-41d75ad0ed8a_1100x449.png 848w, https://substackcdn.com/image/fetch/$s_!m4Pk!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F796ba650-9bfc-40aa-ac94-41d75ad0ed8a_1100x449.png 1272w, https://substackcdn.com/image/fetch/$s_!m4Pk!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F796ba650-9bfc-40aa-ac94-41d75ad0ed8a_1100x449.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>System prompts, shared codebase contexts, RAG corpus headers: cache. One-shot user uploads and cold agent trajectories: recompute. <strong>Cache residency is now a metered commodity; treat admission like an underwriting decision.</strong> </p><p>One nuance cuts the other way: NAND inflated <em>less</em> than DRAM, so the relative case for demoting warm entries to NVMe actually improved even as both tiers rose.</p><h3>Three channels into cost per token</h3><p><strong>Channel 1: CapEx per node.</strong> HBM is now the <strong>single largest line item in a flagship accelerator&#8217;s package BOM</strong> (SemiAnalysis&#8217;s finding for the GB300, after rising as a share every generation since Hopper) and adding host DDR5 and NVMe puts memory at roughly half the bill of a modern inference node (author&#8217;s estimate on top of that sourced base). </p><p>HBM moves slowly through long-term contracts; host DRAM and SSD are bought near spot: which is where 2026 budgets are bleeding now. A node spec&#8217;d mid-2025 with 1.5 TB DDR5 + 30 TB NVMe carries tens of thousands of dollars of new memory cost at current prices. Amortized over four years at high utilization it&#8217;s single-digit percent per token, but it compounds with Channel 2.</p><p><strong>Channel 2: the capacity ceiling (the big one).</strong> When memory binds, its price doesn&#8217;t just raise cost : <strong>it caps revenue per node</strong>. Max concurrency = free-memory-after-weights &#247; KV-bytes-per-sequence; decode throughput scales with batch until bandwidth saturates. As traffic mix shifts long-context (it is: agents, codebases, multimodal), each node serves fewer users and each token carries more fixed cost. </p><p><strong>Cost per token rises without any component getting more expensive: pure mix shift.</strong> The shortage means you can&#8217;t fix it by bolting on DRAM at 2024 prices; the escape routes are compression and disaggregation. Both are engineering, not procurement.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!_h4r!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F16083e0e-6498-46d6-80ae-a32f3316319e_1100x647.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!_h4r!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F16083e0e-6498-46d6-80ae-a32f3316319e_1100x647.png 424w, https://substackcdn.com/image/fetch/$s_!_h4r!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F16083e0e-6498-46d6-80ae-a32f3316319e_1100x647.png 848w, https://substackcdn.com/image/fetch/$s_!_h4r!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F16083e0e-6498-46d6-80ae-a32f3316319e_1100x647.png 1272w, https://substackcdn.com/image/fetch/$s_!_h4r!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F16083e0e-6498-46d6-80ae-a32f3316319e_1100x647.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!_h4r!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F16083e0e-6498-46d6-80ae-a32f3316319e_1100x647.png" width="1100" height="647" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/16083e0e-6498-46d6-80ae-a32f3316319e_1100x647.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:647,&quot;width&quot;:1100,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;c-fig6&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="c-fig6" title="c-fig6" srcset="https://substackcdn.com/image/fetch/$s_!_h4r!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F16083e0e-6498-46d6-80ae-a32f3316319e_1100x647.png 424w, https://substackcdn.com/image/fetch/$s_!_h4r!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F16083e0e-6498-46d6-80ae-a32f3316319e_1100x647.png 848w, https://substackcdn.com/image/fetch/$s_!_h4r!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F16083e0e-6498-46d6-80ae-a32f3316319e_1100x647.png 1272w, https://substackcdn.com/image/fetch/$s_!_h4r!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F16083e0e-6498-46d6-80ae-a32f3316319e_1100x647.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><strong>Channel 3: the tier-ratio compression</strong> of Fig. 3, which quietly deprecates a generation of offloading tricks and re-rates model-architecture choices (&#167;02, Fig. 4).</p><h3>Winners, losers, and the direction of the tilt</h3><p><strong>Hyperscalers and frontier labs are insulated.</strong> Multi-year HBM/DDR5 LTAs signed pre-squeeze, first-priority allocation (Micron: &#8220;larger, strategic customers&#8221;), fleets big enough that architecture fixes amortize across <strong>trillions of tokens. </strong></p><p><strong>Databricks&#8217; disclosure</strong> that cost-aware autoscaling and &#8220;model unit&#8221; abstractions cut GPU spend over 80% versus static provisioning (on a platform serving 120 trillion tokens a month) shows where leverage lives at that scale: <em>utilization engineering</em>, because their input prices are contractually smoothed.</p><p><strong>Mid-tier providers and self-hosters absorb the shock.</strong> They buy near spot, hold no allocation priority, and their pitch (undercutting frontier APIs) sits directly on the inflating commodity. Worse, their buyers are the newly capped: under a token budget, an enterprise squeezes more work from the same spend rather than paying up, so the mid-tier faces rising input costs and a customer base structurally optimizing against price <em>at the same time</em>.</p><p> That is the <strong>textbook geometry of a margin squeeze</strong>, and it lands here first. Watch for: per-token price floors firming through H2 2026 in open-weights serving; long-context surcharges turning explicit; and prompt-caching <em>pricing</em> rebalancing: cached-token discounts exist because storage was cheap relative to recompute, and that ratio just moved. Cache-storage line items (already visible in frontier pricing) will spread.</p><p><strong>The self-hosting math shifted in a direction almost nobody has updated for.</strong> The 2025 pitch (&#8221;your $15-25K refurbished H100 workstation replaces $500/month of API spend&#8221;) has two 2026 problems. The workstation&#8217;s own BOM inflated: RAM alone on a new server build can now approach the cost of an entire refurbished platform. And the API side is partially shielded by hyperscaler contract insulation plus superior compression engineering. </p><p>The perverse result: <strong>the memory supercycle is a centralizing force, it taxes small-fleet inference more heavily than hyperscale inference, at precisely the moment open-weights quality made decentralized serving viable.</strong> </p><p>If you&#8217;re modeling a GPU buy this year, redo it with 2026 memory line items, and note that used enterprise hardware with pre-crisis RAM already installed is currently the cleanest arbitrage in the market, which is exactly why the refurb channel is having a moment.</p><div><hr></div><h2>Proof of pass-through: the March repricing, and the token bill in real dollars</h2><p>A thesis this size needs a smoking gun: evidence that memory contract prices actually reach the price <em>you</em> pay per GPU-hour, on a measurable lag. Q1 2026 provided it.</p><p>Silicon Data&#8217;s SDB200RT index (the standardized benchmark for B200 cloud rental pricing) opened 2026 at 4.40, drifted through January and February, then <strong>rose 23.6% inside March alone</strong>, crossing 5.0 on March 15 and 6.0 on March 23 before settling at 5.48: up 24.4% year-to-date.</p><p> Their attribution is the entire argument of this issue in one sentence: <strong>Samsung and SK Hynix raised HBM3e contract prices ~20% for 2026 deliveries, NVIDIA revised hardware MSRPs upward in late February citing memory component costs, and those input costs flowed into cloud hourly rates with a 4-6 week lag</strong>: landing squarely in March. </p><p>Meanwhile the two-year-old H100, whose memory was bought at pre-squeeze prices, traded in a tight, boring band all quarter. Same market, same month: <strong>the GPUs carrying new memory repriced; the GPUs carrying old memory didn&#8217;t.</strong> That divergence <em>is</em> the supercycle reaching your invoice.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!r2ZO!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e7e1680-0830-4ab0-b17f-1a7755da9205_1100x637.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!r2ZO!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e7e1680-0830-4ab0-b17f-1a7755da9205_1100x637.png 424w, https://substackcdn.com/image/fetch/$s_!r2ZO!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e7e1680-0830-4ab0-b17f-1a7755da9205_1100x637.png 848w, https://substackcdn.com/image/fetch/$s_!r2ZO!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e7e1680-0830-4ab0-b17f-1a7755da9205_1100x637.png 1272w, https://substackcdn.com/image/fetch/$s_!r2ZO!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e7e1680-0830-4ab0-b17f-1a7755da9205_1100x637.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!r2ZO!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e7e1680-0830-4ab0-b17f-1a7755da9205_1100x637.png" width="1100" height="637" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/3e7e1680-0830-4ab0-b17f-1a7755da9205_1100x637.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:637,&quot;width&quot;:1100,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;c-fig7&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="c-fig7" title="c-fig7" srcset="https://substackcdn.com/image/fetch/$s_!r2ZO!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e7e1680-0830-4ab0-b17f-1a7755da9205_1100x637.png 424w, https://substackcdn.com/image/fetch/$s_!r2ZO!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e7e1680-0830-4ab0-b17f-1a7755da9205_1100x637.png 848w, https://substackcdn.com/image/fetch/$s_!r2ZO!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e7e1680-0830-4ab0-b17f-1a7755da9205_1100x637.png 1272w, https://substackcdn.com/image/fetch/$s_!r2ZO!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e7e1680-0830-4ab0-b17f-1a7755da9205_1100x637.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h3>The token bill, priced in real mid-2026 dollars</h3><p>With real rental rates in hand (H200 working median ~$3.50/GPU-hr on-demand across neocloud trackers (cohort median ~$4.00; floor $2.30; hyperscalers to $13.78)) the bandwidth model from &#167;03 becomes an actual price sheet. </p><p>These are <strong>bandwidth-bound floors</strong>: ideal kernels, full overlap, output tokens only. Real deployments land at 50-70% of these throughputs; the <em>ratios</em> between cells are the robust result.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!-mk-!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd59ba112-c8ab-4aed-b804-076fc419af92_1100x313.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!-mk-!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd59ba112-c8ab-4aed-b804-076fc419af92_1100x313.png 424w, https://substackcdn.com/image/fetch/$s_!-mk-!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd59ba112-c8ab-4aed-b804-076fc419af92_1100x313.png 848w, https://substackcdn.com/image/fetch/$s_!-mk-!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd59ba112-c8ab-4aed-b804-076fc419af92_1100x313.png 1272w, https://substackcdn.com/image/fetch/$s_!-mk-!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd59ba112-c8ab-4aed-b804-076fc419af92_1100x313.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!-mk-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd59ba112-c8ab-4aed-b804-076fc419af92_1100x313.png" width="1100" height="313" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d59ba112-c8ab-4aed-b804-076fc419af92_1100x313.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:313,&quot;width&quot;:1100,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;c-math4&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="c-math4" title="c-math4" srcset="https://substackcdn.com/image/fetch/$s_!-mk-!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd59ba112-c8ab-4aed-b804-076fc419af92_1100x313.png 424w, https://substackcdn.com/image/fetch/$s_!-mk-!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd59ba112-c8ab-4aed-b804-076fc419af92_1100x313.png 848w, https://substackcdn.com/image/fetch/$s_!-mk-!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd59ba112-c8ab-4aed-b804-076fc419af92_1100x313.png 1272w, https://substackcdn.com/image/fetch/$s_!-mk-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd59ba112-c8ab-4aed-b804-076fc419af92_1100x313.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!sZWU!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd100fe8c-3a01-47c9-ba87-4111eac15e28_1100x689.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!sZWU!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd100fe8c-3a01-47c9-ba87-4111eac15e28_1100x689.png 424w, https://substackcdn.com/image/fetch/$s_!sZWU!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd100fe8c-3a01-47c9-ba87-4111eac15e28_1100x689.png 848w, https://substackcdn.com/image/fetch/$s_!sZWU!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd100fe8c-3a01-47c9-ba87-4111eac15e28_1100x689.png 1272w, https://substackcdn.com/image/fetch/$s_!sZWU!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd100fe8c-3a01-47c9-ba87-4111eac15e28_1100x689.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!sZWU!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd100fe8c-3a01-47c9-ba87-4111eac15e28_1100x689.png" width="1100" height="689" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d100fe8c-3a01-47c9-ba87-4111eac15e28_1100x689.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:689,&quot;width&quot;:1100,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;c-fig8&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="c-fig8" title="c-fig8" srcset="https://substackcdn.com/image/fetch/$s_!sZWU!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd100fe8c-3a01-47c9-ba87-4111eac15e28_1100x689.png 424w, https://substackcdn.com/image/fetch/$s_!sZWU!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd100fe8c-3a01-47c9-ba87-4111eac15e28_1100x689.png 848w, https://substackcdn.com/image/fetch/$s_!sZWU!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd100fe8c-3a01-47c9-ba87-4111eac15e28_1100x689.png 1272w, https://substackcdn.com/image/fetch/$s_!sZWU!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd100fe8c-3a01-47c9-ba87-4111eac15e28_1100x689.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>One honest caveat, stated before a reader states it for me: this model charges every decode step for streaming the full weights and KV and ignores prefill cost, which at 128K is substantial and shifts spend toward exactly the CPX-shaped silicon. </p><p>A production number also carries utilization (Databricks&#8217; 80% savings figure is the size of <em>that</em> lever), SLO headroom, and interconnect overhead in multi-GPU serving. <strong>The model&#8217;s job is not to predict your invoice to the cent; it is to rank your decisions, and the ranking is unambiguous: </strong>context policy first, KV dtype second, hardware price third.</p><h3>The spread, marked to market. July 2, 2026</h3><p>Now close the loop the title promises. As of July 2, flat per-token pricing for Llama-3.3-70B-class serving spans <strong>$0.31-$0.90 per million output tokens</strong> across the fifteen providers Artificial Analysis tracks (DeepInfra&#8217;s FP8 &#8220;Turbo&#8221; at $0.40, Groq at $0.79, Together AI and Fireworks at $0.88-0.90) and the rate is <em>flat across the model&#8217;s 131K context window</em>. </p><p>Set those prices against this section&#8217;s floors and the cross-subsidy stops being an inference and becomes a table:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!PvSl!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a284d54-9659-4978-82b2-8a7123776172_1100x459.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!PvSl!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a284d54-9659-4978-82b2-8a7123776172_1100x459.png 424w, https://substackcdn.com/image/fetch/$s_!PvSl!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a284d54-9659-4978-82b2-8a7123776172_1100x459.png 848w, https://substackcdn.com/image/fetch/$s_!PvSl!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a284d54-9659-4978-82b2-8a7123776172_1100x459.png 1272w, https://substackcdn.com/image/fetch/$s_!PvSl!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a284d54-9659-4978-82b2-8a7123776172_1100x459.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!PvSl!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a284d54-9659-4978-82b2-8a7123776172_1100x459.png" width="1100" height="459" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/9a284d54-9659-4978-82b2-8a7123776172_1100x459.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:459,&quot;width&quot;:1100,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;c-table2&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="c-table2" title="c-table2" srcset="https://substackcdn.com/image/fetch/$s_!PvSl!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a284d54-9659-4978-82b2-8a7123776172_1100x459.png 424w, https://substackcdn.com/image/fetch/$s_!PvSl!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a284d54-9659-4978-82b2-8a7123776172_1100x459.png 848w, https://substackcdn.com/image/fetch/$s_!PvSl!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a284d54-9659-4978-82b2-8a7123776172_1100x459.png 1272w, https://substackcdn.com/image/fetch/$s_!PvSl!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a284d54-9659-4978-82b2-8a7123776172_1100x459.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Note the confirming detail hiding in the market data: the price leader doesn&#8217;t serve the reference model at all, it serves an FP8-quantized variant. <strong>The cheapest provider in the market is this issue&#8217;s playbook, priced and shipped.</strong> </p><p>And the honest caveats, before a reader raises them: these floors are single-GPU bandwidth ideals (real serving is less efficient, which raises them), while real fleets also earn input-token revenue and batch mixed contexts (which lowers effective floors), so treat the magnitudes as directional. </p><p>The <em>sign pattern</em> is the robust result: <strong>at July 2026 flat prices, long-context traffic on 70B-class serving is sold below its modeled production floor, and budget-capped buyers remove the option of raising the flat rate to fix it.</strong> Hence the final item of &#167;07&#8217;s playbook.</p><div><hr></div><h2>Eight decisions, ranked by dollar leverage</h2><p>Everything above, compiled into action: ranked by expected cost impact per unit of engineering effort for a team serving open-weights models at scale in 2026. Under the scissors of &#167;00, read &#8220;cost impact&#8221; as what it now is: <strong>spread defense</strong>: every dollar of floor you remove is margin the ceiling can no longer take from you.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!xTts!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F193e4edc-128f-483f-8df8-7f9a299c4472_1100x405.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!xTts!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F193e4edc-128f-483f-8df8-7f9a299c4472_1100x405.png 424w, https://substackcdn.com/image/fetch/$s_!xTts!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F193e4edc-128f-483f-8df8-7f9a299c4472_1100x405.png 848w, https://substackcdn.com/image/fetch/$s_!xTts!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F193e4edc-128f-483f-8df8-7f9a299c4472_1100x405.png 1272w, https://substackcdn.com/image/fetch/$s_!xTts!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F193e4edc-128f-483f-8df8-7f9a299c4472_1100x405.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!xTts!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F193e4edc-128f-483f-8df8-7f9a299c4472_1100x405.png" width="1100" height="405" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/193e4edc-128f-483f-8df8-7f9a299c4472_1100x405.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:405,&quot;width&quot;:1100,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;c-play01&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="c-play01" title="c-play01" srcset="https://substackcdn.com/image/fetch/$s_!xTts!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F193e4edc-128f-483f-8df8-7f9a299c4472_1100x405.png 424w, https://substackcdn.com/image/fetch/$s_!xTts!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F193e4edc-128f-483f-8df8-7f9a299c4472_1100x405.png 848w, https://substackcdn.com/image/fetch/$s_!xTts!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F193e4edc-128f-483f-8df8-7f9a299c4472_1100x405.png 1272w, https://substackcdn.com/image/fetch/$s_!xTts!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F193e4edc-128f-483f-8df8-7f9a299c4472_1100x405.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!qafv!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6542ca3-f586-484f-920e-11ae05da957d_1100x378.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!qafv!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6542ca3-f586-484f-920e-11ae05da957d_1100x378.png 424w, https://substackcdn.com/image/fetch/$s_!qafv!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6542ca3-f586-484f-920e-11ae05da957d_1100x378.png 848w, https://substackcdn.com/image/fetch/$s_!qafv!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6542ca3-f586-484f-920e-11ae05da957d_1100x378.png 1272w, https://substackcdn.com/image/fetch/$s_!qafv!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6542ca3-f586-484f-920e-11ae05da957d_1100x378.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!qafv!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6542ca3-f586-484f-920e-11ae05da957d_1100x378.png" width="1100" height="378" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b6542ca3-f586-484f-920e-11ae05da957d_1100x378.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:378,&quot;width&quot;:1100,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;c-play02&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="c-play02" title="c-play02" srcset="https://substackcdn.com/image/fetch/$s_!qafv!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6542ca3-f586-484f-920e-11ae05da957d_1100x378.png 424w, https://substackcdn.com/image/fetch/$s_!qafv!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6542ca3-f586-484f-920e-11ae05da957d_1100x378.png 848w, https://substackcdn.com/image/fetch/$s_!qafv!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6542ca3-f586-484f-920e-11ae05da957d_1100x378.png 1272w, https://substackcdn.com/image/fetch/$s_!qafv!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6542ca3-f586-484f-920e-11ae05da957d_1100x378.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!OcjK!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb08315fd-80fd-475b-b82f-ecd2565bc281_1100x378.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!OcjK!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb08315fd-80fd-475b-b82f-ecd2565bc281_1100x378.png 424w, https://substackcdn.com/image/fetch/$s_!OcjK!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb08315fd-80fd-475b-b82f-ecd2565bc281_1100x378.png 848w, https://substackcdn.com/image/fetch/$s_!OcjK!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb08315fd-80fd-475b-b82f-ecd2565bc281_1100x378.png 1272w, https://substackcdn.com/image/fetch/$s_!OcjK!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb08315fd-80fd-475b-b82f-ecd2565bc281_1100x378.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!OcjK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb08315fd-80fd-475b-b82f-ecd2565bc281_1100x378.png" width="1100" height="378" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b08315fd-80fd-475b-b82f-ecd2565bc281_1100x378.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:378,&quot;width&quot;:1100,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;c-play03&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="c-play03" title="c-play03" srcset="https://substackcdn.com/image/fetch/$s_!OcjK!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb08315fd-80fd-475b-b82f-ecd2565bc281_1100x378.png 424w, https://substackcdn.com/image/fetch/$s_!OcjK!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb08315fd-80fd-475b-b82f-ecd2565bc281_1100x378.png 848w, https://substackcdn.com/image/fetch/$s_!OcjK!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb08315fd-80fd-475b-b82f-ecd2565bc281_1100x378.png 1272w, https://substackcdn.com/image/fetch/$s_!OcjK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb08315fd-80fd-475b-b82f-ecd2565bc281_1100x378.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!AJii!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4052a6f5-6c01-4e98-a857-d06cba0e22e5_1100x415.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!AJii!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4052a6f5-6c01-4e98-a857-d06cba0e22e5_1100x415.png 424w, https://substackcdn.com/image/fetch/$s_!AJii!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4052a6f5-6c01-4e98-a857-d06cba0e22e5_1100x415.png 848w, https://substackcdn.com/image/fetch/$s_!AJii!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4052a6f5-6c01-4e98-a857-d06cba0e22e5_1100x415.png 1272w, https://substackcdn.com/image/fetch/$s_!AJii!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4052a6f5-6c01-4e98-a857-d06cba0e22e5_1100x415.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!AJii!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4052a6f5-6c01-4e98-a857-d06cba0e22e5_1100x415.png" width="1100" height="415" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/4052a6f5-6c01-4e98-a857-d06cba0e22e5_1100x415.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:415,&quot;width&quot;:1100,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;c-play04&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="c-play04" title="c-play04" srcset="https://substackcdn.com/image/fetch/$s_!AJii!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4052a6f5-6c01-4e98-a857-d06cba0e22e5_1100x415.png 424w, https://substackcdn.com/image/fetch/$s_!AJii!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4052a6f5-6c01-4e98-a857-d06cba0e22e5_1100x415.png 848w, https://substackcdn.com/image/fetch/$s_!AJii!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4052a6f5-6c01-4e98-a857-d06cba0e22e5_1100x415.png 1272w, https://substackcdn.com/image/fetch/$s_!AJii!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4052a6f5-6c01-4e98-a857-d06cba0e22e5_1100x415.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!lrul!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2096bc08-9268-4de0-b7f2-3285f250e586_1100x306.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!lrul!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2096bc08-9268-4de0-b7f2-3285f250e586_1100x306.png 424w, https://substackcdn.com/image/fetch/$s_!lrul!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2096bc08-9268-4de0-b7f2-3285f250e586_1100x306.png 848w, https://substackcdn.com/image/fetch/$s_!lrul!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2096bc08-9268-4de0-b7f2-3285f250e586_1100x306.png 1272w, https://substackcdn.com/image/fetch/$s_!lrul!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2096bc08-9268-4de0-b7f2-3285f250e586_1100x306.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!lrul!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2096bc08-9268-4de0-b7f2-3285f250e586_1100x306.png" width="1100" height="306" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2096bc08-9268-4de0-b7f2-3285f250e586_1100x306.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:306,&quot;width&quot;:1100,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;c-play05&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="c-play05" title="c-play05" srcset="https://substackcdn.com/image/fetch/$s_!lrul!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2096bc08-9268-4de0-b7f2-3285f250e586_1100x306.png 424w, https://substackcdn.com/image/fetch/$s_!lrul!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2096bc08-9268-4de0-b7f2-3285f250e586_1100x306.png 848w, https://substackcdn.com/image/fetch/$s_!lrul!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2096bc08-9268-4de0-b7f2-3285f250e586_1100x306.png 1272w, https://substackcdn.com/image/fetch/$s_!lrul!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2096bc08-9268-4de0-b7f2-3285f250e586_1100x306.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!XHBY!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ca5727f-44af-4947-b152-5d01560fcf6d_1100x342.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!XHBY!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ca5727f-44af-4947-b152-5d01560fcf6d_1100x342.png 424w, https://substackcdn.com/image/fetch/$s_!XHBY!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ca5727f-44af-4947-b152-5d01560fcf6d_1100x342.png 848w, https://substackcdn.com/image/fetch/$s_!XHBY!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ca5727f-44af-4947-b152-5d01560fcf6d_1100x342.png 1272w, https://substackcdn.com/image/fetch/$s_!XHBY!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ca5727f-44af-4947-b152-5d01560fcf6d_1100x342.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!XHBY!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ca5727f-44af-4947-b152-5d01560fcf6d_1100x342.png" width="1100" height="342" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/8ca5727f-44af-4947-b152-5d01560fcf6d_1100x342.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:342,&quot;width&quot;:1100,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;c-play06&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="c-play06" title="c-play06" srcset="https://substackcdn.com/image/fetch/$s_!XHBY!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ca5727f-44af-4947-b152-5d01560fcf6d_1100x342.png 424w, https://substackcdn.com/image/fetch/$s_!XHBY!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ca5727f-44af-4947-b152-5d01560fcf6d_1100x342.png 848w, https://substackcdn.com/image/fetch/$s_!XHBY!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ca5727f-44af-4947-b152-5d01560fcf6d_1100x342.png 1272w, https://substackcdn.com/image/fetch/$s_!XHBY!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ca5727f-44af-4947-b152-5d01560fcf6d_1100x342.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!3N8x!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb7585f3d-de3d-479f-a018-ab52d6708861_1100x378.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!3N8x!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb7585f3d-de3d-479f-a018-ab52d6708861_1100x378.png 424w, https://substackcdn.com/image/fetch/$s_!3N8x!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb7585f3d-de3d-479f-a018-ab52d6708861_1100x378.png 848w, https://substackcdn.com/image/fetch/$s_!3N8x!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb7585f3d-de3d-479f-a018-ab52d6708861_1100x378.png 1272w, https://substackcdn.com/image/fetch/$s_!3N8x!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb7585f3d-de3d-479f-a018-ab52d6708861_1100x378.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!3N8x!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb7585f3d-de3d-479f-a018-ab52d6708861_1100x378.png" width="1100" height="378" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b7585f3d-de3d-479f-a018-ab52d6708861_1100x378.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:378,&quot;width&quot;:1100,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;c-play07&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="c-play07" title="c-play07" srcset="https://substackcdn.com/image/fetch/$s_!3N8x!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb7585f3d-de3d-479f-a018-ab52d6708861_1100x378.png 424w, https://substackcdn.com/image/fetch/$s_!3N8x!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb7585f3d-de3d-479f-a018-ab52d6708861_1100x378.png 848w, https://substackcdn.com/image/fetch/$s_!3N8x!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb7585f3d-de3d-479f-a018-ab52d6708861_1100x378.png 1272w, https://substackcdn.com/image/fetch/$s_!3N8x!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb7585f3d-de3d-479f-a018-ab52d6708861_1100x378.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!gfIz!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b53e289-c6b1-459b-a166-a7c2ee46298b_1100x514.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!gfIz!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b53e289-c6b1-459b-a166-a7c2ee46298b_1100x514.png 424w, https://substackcdn.com/image/fetch/$s_!gfIz!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b53e289-c6b1-459b-a166-a7c2ee46298b_1100x514.png 848w, https://substackcdn.com/image/fetch/$s_!gfIz!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b53e289-c6b1-459b-a166-a7c2ee46298b_1100x514.png 1272w, https://substackcdn.com/image/fetch/$s_!gfIz!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b53e289-c6b1-459b-a166-a7c2ee46298b_1100x514.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!gfIz!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b53e289-c6b1-459b-a166-a7c2ee46298b_1100x514.png" width="1100" height="514" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/1b53e289-c6b1-459b-a166-a7c2ee46298b_1100x514.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:514,&quot;width&quot;:1100,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;c-play08&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="c-play08" title="c-play08" srcset="https://substackcdn.com/image/fetch/$s_!gfIz!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b53e289-c6b1-459b-a166-a7c2ee46298b_1100x514.png 424w, https://substackcdn.com/image/fetch/$s_!gfIz!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b53e289-c6b1-459b-a166-a7c2ee46298b_1100x514.png 848w, https://substackcdn.com/image/fetch/$s_!gfIz!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b53e289-c6b1-459b-a166-a7c2ee46298b_1100x514.png 1272w, https://substackcdn.com/image/fetch/$s_!gfIz!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b53e289-c6b1-459b-a166-a7c2ee46298b_1100x514.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div><hr></div><h2>Five falsifiable predictions</h2><p>Scored publicly in Q1 2027, alongside the Issue 04 speculative-decoding calls. Confidence is stated so being wrong costs me something.</p><p><strong>P1 &#183; Supplyconfidence: high</strong><br>DRAM relief does not arrive before late 2027. New capacity (Micron Boise) ramps 2027-2028; suppliers themselves guide consumer relief to ~2028. Any 2027 hardware-refresh model priced at 2024 memory levels is fiction.</p><p><strong>P2 &#183; Architectureconfidence: high</strong><br>Open-weights flagships converge on compressed KV. By mid-2027 the majority of new frontier-adjacent open releases ship MLA-like latent attention, &#8804;8-head GQA, or hybrid sliding windows: marketed explicitly as serving-cost features. Model architecture is now downstream of the DRAM spot price.</p><p><strong>P3 &#183; Pricingconfidence: medium-high</strong><br>Per-token API pricing bifurcates by context residency: at least two major providers introduce or restructure explicit cached-context <em>storage</em> pricing (per-token-per-hour or equivalent) by mid-2027, decoupling state from compute. Partial confirmation is already on the books (Google Vertex bills per-hour cache storage and Anthropic bills TTL-tiered cache writes) so the live prediction is that this becomes the <em>norm in open-weights serving</em>, not just frontier practice.</p><p><strong>P4 &#183; Siliconconfidence: medium</strong><br>Prefill-specialized silicon becomes a category, not a SKU: a CPX competitor (AMD or an ASIC player) is announced within 12 months of CPX GA, and &#8220;prefill accelerator&#8221; enters the standard rack taxonomy.</p><p><strong>P5 &#183; Marketconfidence: medium</strong><br>The open-weights serving price war pauses: median $/Mtok for 70B-class serving on third-party clouds goes flat-to-up through H1 2027 (the first sustained non-decline in that series since it has existed. The March B200 repricing (Fig. 7) is the leading edge) and enterprise token budgets are the demand-side mechanism that lets long-context surcharges and cache-storage fees actually stick, because a buyer operating under a cap optimizes usage instead of switching providers.</p><h3>The dashboard: six leading indicators, with trigger levels</h3><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!ozsr!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff24d4938-9512-4e15-9d7e-41f79234be1a_1100x1033.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!ozsr!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff24d4938-9512-4e15-9d7e-41f79234be1a_1100x1033.png 424w, https://substackcdn.com/image/fetch/$s_!ozsr!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff24d4938-9512-4e15-9d7e-41f79234be1a_1100x1033.png 848w, https://substackcdn.com/image/fetch/$s_!ozsr!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff24d4938-9512-4e15-9d7e-41f79234be1a_1100x1033.png 1272w, https://substackcdn.com/image/fetch/$s_!ozsr!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff24d4938-9512-4e15-9d7e-41f79234be1a_1100x1033.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!ozsr!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff24d4938-9512-4e15-9d7e-41f79234be1a_1100x1033.png" width="1100" height="1033" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f24d4938-9512-4e15-9d7e-41f79234be1a_1100x1033.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1033,&quot;width&quot;:1100,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;c-table3&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="c-table3" title="c-table3" srcset="https://substackcdn.com/image/fetch/$s_!ozsr!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff24d4938-9512-4e15-9d7e-41f79234be1a_1100x1033.png 424w, https://substackcdn.com/image/fetch/$s_!ozsr!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff24d4938-9512-4e15-9d7e-41f79234be1a_1100x1033.png 848w, https://substackcdn.com/image/fetch/$s_!ozsr!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff24d4938-9512-4e15-9d7e-41f79234be1a_1100x1033.png 1272w, https://substackcdn.com/image/fetch/$s_!ozsr!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff24d4938-9512-4e15-9d7e-41f79234be1a_1100x1033.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!GcTh!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F32114b54-e57d-4baa-bcb8-d576d0a9d59c_1100x593.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!GcTh!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F32114b54-e57d-4baa-bcb8-d576d0a9d59c_1100x593.png 424w, https://substackcdn.com/image/fetch/$s_!GcTh!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F32114b54-e57d-4baa-bcb8-d576d0a9d59c_1100x593.png 848w, https://substackcdn.com/image/fetch/$s_!GcTh!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F32114b54-e57d-4baa-bcb8-d576d0a9d59c_1100x593.png 1272w, https://substackcdn.com/image/fetch/$s_!GcTh!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F32114b54-e57d-4baa-bcb8-d576d0a9d59c_1100x593.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!GcTh!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F32114b54-e57d-4baa-bcb8-d576d0a9d59c_1100x593.png" width="1100" height="593" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/32114b54-e57d-4baa-bcb8-d576d0a9d59c_1100x593.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:593,&quot;width&quot;:1100,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;c-fig9&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="c-fig9" title="c-fig9" srcset="https://substackcdn.com/image/fetch/$s_!GcTh!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F32114b54-e57d-4baa-bcb8-d576d0a9d59c_1100x593.png 424w, https://substackcdn.com/image/fetch/$s_!GcTh!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F32114b54-e57d-4baa-bcb8-d576d0a9d59c_1100x593.png 848w, https://substackcdn.com/image/fetch/$s_!GcTh!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F32114b54-e57d-4baa-bcb8-d576d0a9d59c_1100x593.png 1272w, https://substackcdn.com/image/fetch/$s_!GcTh!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F32114b54-e57d-4baa-bcb8-d576d0a9d59c_1100x593.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div><hr></div><h2>The spread is the product now</h2><p>For two years, inference economics was a story about FLOPs: kernel efficiency, quantized matmuls, speculative decoding, Blackwell&#8217;s FP4 throughput. </p><p>Those battles were real; this publication has spent six issues inside them. But they were fought against a background assumption so universal nobody stated it: <em>bytes are cheap and getting cheaper.</em></p><p>That assumption died in Q1 2026. <strong>The binding constraint of the inference industry is now measured in wafer starts and TSV yields</strong>: in exabytes of KV state colliding with 10-16% annual bit-supply growth, in an IDC report that uses the word &#8220;permanent,&#8221; in a rental index that repriced 24% in one month because a memory contract reset five weeks earlier. </p><p>The consequences run one direction through the entire stack: hardware specializes by phase <strong>because HBM is too precious to waste</strong> on prefill; model architectures compress their KV because fat caches became unserveable; serving stacks reorganize around shipping cached state across fabrics because holding it still got expensive; and the cost advantage tilts toward <strong>whoever holds allocation contracts </strong>and compression engineering: which is to say, toward the largest players, in an industry that spent two years congratulating itself on decentralizing.</p><p>The engineers who internalize this fastest hold a simple advantage: while the market reprices memory, they reprice their <em>need</em> for it. <strong>Every KV byte you decline to allocate is bought at 2026 prices and sold at 2026 prices, at 100% margin, with zero lead time and no allocation meeting.</strong> The wafer sets the cost floor. The wallet sets the price ceiling. </p><p>The engineering in between is the only variable you control, and in a margin regime, that engineering stops being a cost-optimization side quest and <strong>becomes the P&amp;L itself.</strong> Squeezes do not reward the biggest fleet; they reward the widest spread per byte.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!ULaf!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd001de35-a3bc-4195-a51a-e01c27856add_1100x311.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!ULaf!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd001de35-a3bc-4195-a51a-e01c27856add_1100x311.png 424w, https://substackcdn.com/image/fetch/$s_!ULaf!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd001de35-a3bc-4195-a51a-e01c27856add_1100x311.png 848w, https://substackcdn.com/image/fetch/$s_!ULaf!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd001de35-a3bc-4195-a51a-e01c27856add_1100x311.png 1272w, https://substackcdn.com/image/fetch/$s_!ULaf!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd001de35-a3bc-4195-a51a-e01c27856add_1100x311.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!ULaf!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd001de35-a3bc-4195-a51a-e01c27856add_1100x311.png" width="1100" height="311" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d001de35-a3bc-4195-a51a-e01c27856add_1100x311.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:311,&quot;width&quot;:1100,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;c-callout3&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="c-callout3" title="c-callout3" srcset="https://substackcdn.com/image/fetch/$s_!ULaf!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd001de35-a3bc-4195-a51a-e01c27856add_1100x311.png 424w, https://substackcdn.com/image/fetch/$s_!ULaf!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd001de35-a3bc-4195-a51a-e01c27856add_1100x311.png 848w, https://substackcdn.com/image/fetch/$s_!ULaf!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd001de35-a3bc-4195-a51a-e01c27856add_1100x311.png 1272w, https://substackcdn.com/image/fetch/$s_!ULaf!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd001de35-a3bc-4195-a51a-e01c27856add_1100x311.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><strong>METHOD &amp; VERIFICATION</strong>. Market figures trace to named sources below; all KV, roofline, concurrency, and $/Mtok arithmetic is the author&#8217;s, derived in-text from published model configs and device datasheets so you can check every step. Cost floors are bandwidth-bound ideals: absolute values are lower bounds, ratios are the robust result. </p><p>The Fig. 7 path between published waypoints is interpolated; Fig. 3&#8217;s intermediate points are directional; Fig. 0 is a directional synthesis whose turning points, not slopes, are the sourced claims. Pricing data was spot-checked the week of publication and moves weekly: recheck before committing capital. Corrections run at the top of the next issue, as always.</p><div><hr></div><p><strong>SOURCES:</strong> <em>TrendForce: 1Q26 DRAM/NAND contract revisions; &#8220;Memory Wall&#8221; research (Jan 2026); HBM adoption &amp; 2029 inference-demand outlook; AI wafer-equivalent consumption via Commercial Times (Dec 2025) &#183; IDC (&#8221;Global Memory Shortage Crisis&#8221; (Feb 2026) &#183; IEEE Spectrum) &#8220;AI Is a Memory Hog&#8221; (Apr 2026) &#183; Counterpoint Research (Q1&#8217;26 spot DRAM tracking &#183; CNBC) Micron/SK Hynix/Samsung allocation &amp; CES 2026 reporting (Jan 2026) &#183; Avnet/Omdia/BofA (supercycle &amp; supply-relief estimates &#183; SemiAnalysis) &#8220;Another Giant Leap: The Rubin CPX Specialized Accelerator &amp; Rack&#8221; (Sep 2025) &#183; The Register (CES 2026 Vera Rubin systems coverage (Jan 2026) &#183; Tom&#8217;s Hardware) &#8220;Nvidia&#8217;s Vera Rubin platform in depth&#8221; (Nov 2025, incl. BlueField-4 KV-cache SSD) &#183; Futurum/Ori/NADDOD (Rubin CPX technical analyses &#183; NVIDIA) AI Infra Summit &amp; GTC 2026 announcements &#183; Databricks Engineering (&#8221;Reliable LLM Inference at Scale&#8221; (May 2026) &#183; SemiAnalysis) &#8220;TokenBudgeting: Our Conversations with Enterprises on Token Spend&#8221; (Jul 2, 2026) &#183; Artificial Analysis (Llama 3.3 70B provider benchmarking &amp; cache-pricing mechanics (Jul 2026) &#183; aipricing.guru / Price Per Token) provider price pages (sourced Jul 2, 2026) &#183; Silicon Data. SDB200RT B200 index, March 2026 update &#183; aimultiple GPU Rental Price Index; getdeploying B200/H200 trackers; Jarvislabs H200 pricing (2026) &#183; DeepSeek-V3 technical report (MLA configuration &#183; vLLM) NVFP4 KV-cache PR (in progress at publication).</em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Software Frontier is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[The Router and the Wire]]></title><description><![CDATA[Mixture-of-experts promised cheaper inference by doing less arithmetic. The bill did not disappear. It moved into the network, and the entire shape of a 2026 serving stack is the receipt.]]></description><link>https://www.thesoftwarefrontier.com/p/the-router-and-the-wire</link><guid isPermaLink="false">https://www.thesoftwarefrontier.com/p/the-router-and-the-wire</guid><dc:creator><![CDATA[Lorenzo Bradanini]]></dc:creator><pubDate>Mon, 29 Jun 2026 21:18:38 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!5xDE!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb9cb9fb9-d6a7-45e1-a371-0d6456ee6809_1024x1536.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!5xDE!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb9cb9fb9-d6a7-45e1-a371-0d6456ee6809_1024x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!5xDE!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb9cb9fb9-d6a7-45e1-a371-0d6456ee6809_1024x1536.png 424w, https://substackcdn.com/image/fetch/$s_!5xDE!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb9cb9fb9-d6a7-45e1-a371-0d6456ee6809_1024x1536.png 848w, https://substackcdn.com/image/fetch/$s_!5xDE!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb9cb9fb9-d6a7-45e1-a371-0d6456ee6809_1024x1536.png 1272w, https://substackcdn.com/image/fetch/$s_!5xDE!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb9cb9fb9-d6a7-45e1-a371-0d6456ee6809_1024x1536.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!5xDE!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb9cb9fb9-d6a7-45e1-a371-0d6456ee6809_1024x1536.png" width="1024" height="1536" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b9cb9fb9-d6a7-45e1-a371-0d6456ee6809_1024x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1536,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2363576,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/199726817?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb9cb9fb9-d6a7-45e1-a371-0d6456ee6809_1024x1536.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!5xDE!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb9cb9fb9-d6a7-45e1-a371-0d6456ee6809_1024x1536.png 424w, https://substackcdn.com/image/fetch/$s_!5xDE!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb9cb9fb9-d6a7-45e1-a371-0d6456ee6809_1024x1536.png 848w, https://substackcdn.com/image/fetch/$s_!5xDE!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb9cb9fb9-d6a7-45e1-a371-0d6456ee6809_1024x1536.png 1272w, https://substackcdn.com/image/fetch/$s_!5xDE!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb9cb9fb9-d6a7-45e1-a371-0d6456ee6809_1024x1536.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>The cost did not vanish. It moved.</h2><p>Every few months the inference market tells itself a story about where the money goes, and every few months the story is wrong in the same direction. For two years the story was memory. </p><p>The <strong>key-value cache grew with context</strong>, the weights grew with parameter count, and the binding constraint on a serving deployment was how many bytes of high-bandwidth memory you could buy and how fast you could read them. </p><p>That story was true, and it is the subject of two earlier issues of this publication. It is also no longer the whole story for the models that now define the frontier.</p><p>The frontier moved to sparsity. A <strong>dense model the size of DeepSeek-V3 </strong>would activate all 671 billion of its parameters on every token it processed. The mixture-of-experts version activates roughly 37 billion. </p><p>On paper that is a reduction in arithmetic of roughly eighteen to one, and it is the single reason a model with two-thirds of a trillion parameters can be served at all without a fleet of accelerators per request. </p><p>The promise of the architecture was always framed in floating-point operations: do less math, pay less money. The <strong>vLLM and SGLang communities</strong>, NVIDIA, DeepSeek, and every serving vendor in between repeated some version of it.</p><p>The arithmetic did get cheaper. What the framing left out is that the arithmetic was never the part that was hard to scale. When you spread 256 experts across <strong>dozens or hundreds of accelerators</strong>, a token routed to eight of them has to physically travel to those eight accelerators, be processed, and travel back to be recombined. </p><p>That round trip is an all-to-all communication pattern, and it does not appear in<strong> any FLOP count</strong>. It is the part of the bill that the sparsity story quietly moved off the compute line and onto the network line, where it has been growing ever since.</p><p>DigitalOcean&#8217;s engineering writers put the distinction more bluntly than most vendors will. For a <strong>dense model</strong>, they note, cost scales with memory and is linear and predictable. For a mixture-of-experts model, cost becomes <em>a game of communication</em>. That is the thesis of this issue, stated in five words by someone selling cloud capacity. </p><p>The rest of this report is the long version: what the all-to-all actually costs in bytes and in silicon, why a rack that lists for two to three million dollars is best understood as an answer to a networking problem rather than a compute one, and whether the <strong>wide expert parallelism</strong> that everyone is now deploying actually earns its keep.</p><p>There is a seductive counterexample worth disposing of immediately, because it will come up. The <strong>KTransformers project</strong> can run the complete DeepSeek-V3 model on a single low-cost server with one consumer GPU, a machine that costs in the neighborhood of ten thousand dollars, and still produce nearly twenty tokens per second. I</p><blockquote><p><em>f a mixture-of-experts model can run on a ten-thousand-dollar box, how can the routing be expensive? </em></p></blockquote><p>The answer is that the <strong>KTransformers </strong>configuration never pays the all-to-all toll, because there is no all-to-all. With every expert resident in the memory of a single node, routing a token to an expert is a memory lookup, not a network transfer. </p><p>The economics that follow in this issue are the economics of scale, of serving thousands of concurrent users at frontier latency, and the moment you cross the <strong>boundary from one node</strong> to many, the toll switches on. The single-box demo is real, and it is exactly why the multi-box reality is so often misunderstood.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!lsqO!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F381212c9-40f7-4d9d-859e-fd3b4659c6bc_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!lsqO!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F381212c9-40f7-4d9d-859e-fd3b4659c6bc_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!lsqO!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F381212c9-40f7-4d9d-859e-fd3b4659c6bc_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!lsqO!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F381212c9-40f7-4d9d-859e-fd3b4659c6bc_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!lsqO!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F381212c9-40f7-4d9d-859e-fd3b4659c6bc_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!lsqO!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F381212c9-40f7-4d9d-859e-fd3b4659c6bc_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/381212c9-40f7-4d9d-859e-fd3b4659c6bc_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:97417,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/199726817?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F381212c9-40f7-4d9d-859e-fd3b4659c6bc_1634x893.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!lsqO!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F381212c9-40f7-4d9d-859e-fd3b4659c6bc_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!lsqO!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F381212c9-40f7-4d9d-859e-fd3b4659c6bc_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!lsqO!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F381212c9-40f7-4d9d-859e-fd3b4659c6bc_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!lsqO!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F381212c9-40f7-4d9d-859e-fd3b4659c6bc_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div><hr></div><h2>What sparsity actually buys, and what it borrows</h2><p>Begin with the structure of the model itself, because the geometry of the dispatch is dictated by it. <strong>DeepSeek-V3 carries 256 routed experts</strong> per mixture-of-experts layer plus one shared expert that processes every token. A gating network selects the top eight routed experts for each token.</p><p> The model has <strong>61 transformer layers,</strong> 58 of which are mixture-of-experts layers; the first few are dense. So for the overwhelming majority of the depth of the network, every single token triggers a routing decision, a dispatch to eight destinations, and a combine back.</p><p>The routing itself is not free, though it is cheap relative to the transfer.<strong> A gating network scores all 256 experts </strong>for every token and selects the top eight, and that scoring, the sorting, and the construction of the dispatch order add a small compute and synchronization cost before any data moves. </p><p>It is<strong> minor against the 168 kilobytes</strong> that follow, but it is one more thing the dense model never does, and at the token rates of a frontier deployment even minor per-token costs accumulate. The gating is also where the load imbalance originates, since it is the gate&#8217;s learned preferences that<strong> send too many tokens to too few experts</strong>, which makes it both the cheapest and the most consequential of the operations the router performs.</p><p>The sparsity is genuine and the savings are genuine. Switch Transformer, the architecture that popularized the modern top-k mixture, demonstrated <strong>roughly a sevenfold speedup</strong> over a dense model of equivalent quality, and that ratio has only widened as expert counts have grown. </p><p>When practitioners say mixture-of-experts reduces computation by ninety percent, they are describing the activated-parameter ratio, and they are not wrong about it. A<strong> token that touches 37 billion of 671 billion parameters</strong> is doing far less matrix multiplication than a token that touches all of them.</p><p>But the activation has to move. In a single-accelerator world, the experts a token needs are <strong>sitting in local memory</strong> and the only thing that travels is a memory read. In a serving deployment large enough to matter, the experts are spread across the accelerators by expert parallelism precisely so that each accelerator holds only a few of them and the aggregate weight footprint fits. </p><p>That is the entire point of <strong>expert parallelism</strong>: it is what lets you serve a model whose experts, summed, are far too large for any single device. And it is also what guarantees that the token and its eight chosen experts will, in general, live on different devices.</p><p>So the model does <strong>two collective</strong> operations per mixture-of-experts layer that a dense model never does. The dispatch, sometimes called the scatter, sends<strong> each token&#8217;s hidden activation</strong> to the devices holding its selected experts. The combine, the gather, takes the eight expert outputs and reduces them back into a single vector on the token&#8217;s home device. </p><p>Survey work on efficient inference serving is consistent on the consequence: this<strong> all-to-all exchange of token dispatch</strong> and output gathering is the bottleneck in large-scale mixture-of-experts inference. Not the expert math. The exchange around it.</p><p>This is the borrowing that the sparsity bargain does not advertise. You spend<strong> less on arithmetic and you take on a debt</strong> denominated in bandwidth, and the debt comes due on a part of the machine that has improved far more slowly than the compute has. </p><p>To see why that matters, you have to look at the wire.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>A rack-scale answer to a network problem</h2><p>The defining fact about modern accelerators is that their arithmetic has outrun their interconnect, and it has done so by a margin that is hard to overstate until you put the numbers on a single axis. A Blackwell B200 reads from <strong>its own high-bandwidth memory</strong> at roughly eight terabytes per second. </p><p>The <strong>fifth-generation NVLink fabric</strong> that connects it to its neighbors moves 1.8 terabytes per second per GPU, eighteen links at a hundred gigabytes per second each. That is the fast path between two accelerators, and it is already more than four times slower than the path to local memory. </p><p>Then you fall off the edge. A single four-hundred-gigabit InfiniBand network card, the scale-out path that connects one node to another, moves <strong>about fifty gigabytes per second. </strong>The cliff from on-package memory to the cross-node network is more than two orders of magnitude.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!KbLz!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a4fe2a3-24a7-436b-a1cd-cd49dbd9626b_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!KbLz!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a4fe2a3-24a7-436b-a1cd-cd49dbd9626b_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!KbLz!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a4fe2a3-24a7-436b-a1cd-cd49dbd9626b_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!KbLz!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a4fe2a3-24a7-436b-a1cd-cd49dbd9626b_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!KbLz!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a4fe2a3-24a7-436b-a1cd-cd49dbd9626b_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!KbLz!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a4fe2a3-24a7-436b-a1cd-cd49dbd9626b_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/5a4fe2a3-24a7-436b-a1cd-cd49dbd9626b_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Horizontal bar chart on a log scale comparing effective bandwidth: HBM3e on-package at 8000 GB/s, NVLink 5 fabric at 1800, DeepEP all-to-all over NVLink measured at 730, DeepEP all-to-all over RDMA internode at 90, and a single InfiniBand NIC at 50.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Horizontal bar chart on a log scale comparing effective bandwidth: HBM3e on-package at 8000 GB/s, NVLink 5 fabric at 1800, DeepEP all-to-all over NVLink measured at 730, DeepEP all-to-all over RDMA internode at 90, and a single InfiniBand NIC at 50." title="Horizontal bar chart on a log scale comparing effective bandwidth: HBM3e on-package at 8000 GB/s, NVLink 5 fabric at 1800, DeepEP all-to-all over NVLink measured at 730, DeepEP all-to-all over RDMA internode at 90, and a single InfiniBand NIC at 50." srcset="https://substackcdn.com/image/fetch/$s_!KbLz!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a4fe2a3-24a7-436b-a1cd-cd49dbd9626b_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!KbLz!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a4fe2a3-24a7-436b-a1cd-cd49dbd9626b_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!KbLz!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a4fe2a3-24a7-436b-a1cd-cd49dbd9626b_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!KbLz!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a4fe2a3-24a7-436b-a1cd-cd49dbd9626b_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong><span>FIG 1</span></strong> The bandwidth available to a token collapses as it travels outward from the chip. The all-to-all of a mixture-of-experts layer lives somewhere on the right half of this chart, and where exactly is the whole question.</figcaption></figure></div><p>This is the chart that explains the rest of the hardware industry&#8217;s behavior. If the <strong>all-to-all of a mixture-of-experts layer</strong> can be kept inside the NVLink domain, it runs at hundreds of gigabytes per second. </p><p>If it has to cross the InfiniBand fabric between racks, it runs at a fraction of that. Introl&#8217;s infrastructure analysis puts the ratio at roughly eighteen to one between scale-up bandwidth <strong>inside the NVLink domain</strong> and scale-out bandwidth between racks. For an architecture whose dominant cost is an all-to-all, that ratio is not a detail. It is the design center.</p><p>Which is what the <strong>GB200 NVL72 is.</strong> NVIDIA&#8217;s rack-scale system connects 72 Blackwell GPUs and 36 Grace CPUs into a single NVLink domain delivering 130 terabytes per second of aggregate, non-blocking, all-to-all bandwidth, with <strong>13.5 terabytes of high-bandwidth memory</strong> addressable as one pool. Before this system, the largest NVLink domain you could buy was eight GPUs on a single baseboard. </p><p>The NVL72 takes the fast interconnect and stretches it across an entire rack so that 72 accelerators can talk to each other as though they were neighbors on the same board. <strong>NVIDIA&#8217;s own materials describe the result as a single massive GPU</strong>, and for the purposes of a mixture-of-experts all-to-all, that marketing is closer to literally true than marketing usually is.</p><p>The price of that rack is two to three million dollars, it draws around a hundred and twenty kilowatts, and it is liquid-cooled because there is no other way to remove the heat. It is easy to read those numbers as a statement about <strong>compute density</strong>, and the 1.44 exaflops of four-bit tensor performance per rack invites that reading. </p><p>But the compute was never the scarce thing. You can buy <strong>720 petaflops of eight-bit compute</strong> in roughly 182 H100 accelerators for less money than an NVL72 costs. </p><p>What you cannot buy that way is a 72-way all-to-all domain. The premium on the rack is, in substantial part, the premium on the wire. It is the cost of <strong>not having to cross InfiniBand</strong> for the operation that a mixture-of-experts model performs 58 times per token.</p><blockquote><p><em><span>The premium on the rack is, in substantial part, the premium on the wire. It is the cost of not crossing InfiniBand for the operation a mixture-of-experts model performs fifty-eight times per token.</span></em></p></blockquote><p>Read this way, a great deal of the 2026 accelerator roadmap resolves into a single sentence: make the all-to-all domain larger than the problem. </p><p>At CES 2026 NVIDIA disclosed that the next-generation <strong>Vera Rubin NVL72 will roughly double per-GPU NVLink bandwidth</strong> to 3.6 terabytes per second and lift aggregate all-to-all bandwidth to 260 terabytes per second, with the explicit justification, in NVIDIA&#8217;s own words, that this is the bandwidth needed for the all-to-all communications of leading mixture-of-experts architectures. </p><p>The company has <strong>stopped being coy about it. </strong>The interconnect generation is being sold, by name, as the answer to the problem that the model generation created.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Twenty streaming multiprocessors, give or take</h2><p>Hardware sets the ceiling. Whether you reach it is a question of kernels, and the reference implementation for the <strong>mixture-of-experts all-to-all is DeepEP</strong>, the communication library DeepSeek open-sourced during its 2025 release week. </p><p>DeepEP is worth studying closely not because it is the only such library but because it is the one whose <strong>measured numbers are public</strong>, and those numbers are the closest thing the field has to a ground truth for what the all-to-all costs at the kernel level.</p><p>DeepEP provides two classes of kernel, and the split maps precisely onto the two phases of inference. The normal kernels are tuned for throughput and serve training and the <strong>prefill phase</strong>, where batches are large and the all-to-all moves a great deal of data at once. </p><p>The <strong>low-latency kernels</strong> are tuned for the decode phase, where each step generates one token per sequence, the batches are tiny, and what matters is not bandwidth but the round-trip time of the dispatch and combine. </p><p>This is the <strong>same prefill-versus-decode division</strong> that disaggregated serving exploits, examined one issue ago, now visible at the level of individual communication kernels.</p><p>The measured bandwidths tell the scale-up story in a single table. On Blackwell-class hardware, DeepEP&#8217;s dispatch kernel moves 726 gigabytes per second and<strong> its combine kernel 740 gigabytes per second</strong> when the experts are inside the NVLink domain. The same kernels, forced across the internode RDMA fabric on the same generation of hardware, move about 90 gigabytes per second each. </p><p>That is the<strong> eighteen-to-one ratio of Figure 1,</strong> reproduced at the kernel level on a real workload: the configuration DeepSeek published, with eight thousand tokens per batch, a <em>hidden dimension of 7168</em>, top-eight routing, eight-bit dispatch, and sixteen-bit combine.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!aJDM!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa35c0b14-487f-47ac-9079-b9349b576f29_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!aJDM!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa35c0b14-487f-47ac-9079-b9349b576f29_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!aJDM!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa35c0b14-487f-47ac-9079-b9349b576f29_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!aJDM!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa35c0b14-487f-47ac-9079-b9349b576f29_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!aJDM!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa35c0b14-487f-47ac-9079-b9349b576f29_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!aJDM!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa35c0b14-487f-47ac-9079-b9349b576f29_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a35c0b14-487f-47ac-9079-b9349b576f29_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Grouped bar chart comparing DeepEP dispatch and combine kernel bandwidth: NVLink intranode at 726 and 740 GB/s versus RDMA internode at 90 GB/s each, roughly eight times slower off the NVLink domain.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Grouped bar chart comparing DeepEP dispatch and combine kernel bandwidth: NVLink intranode at 726 and 740 GB/s versus RDMA internode at 90 GB/s each, roughly eight times slower off the NVLink domain." title="Grouped bar chart comparing DeepEP dispatch and combine kernel bandwidth: NVLink intranode at 726 and 740 GB/s versus RDMA internode at 90 GB/s each, roughly eight times slower off the NVLink domain." srcset="https://substackcdn.com/image/fetch/$s_!aJDM!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa35c0b14-487f-47ac-9079-b9349b576f29_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!aJDM!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa35c0b14-487f-47ac-9079-b9349b576f29_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!aJDM!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa35c0b14-487f-47ac-9079-b9349b576f29_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!aJDM!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa35c0b14-487f-47ac-9079-b9349b576f29_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong><span>FIG 2</span></strong> The same kernel, the same hardware, the same workload. The only variable is whether the experts sit inside the NVLink domain or across the network. That one boundary costs roughly a factor of eight.</figcaption></figure></div><p>The second thing DeepEP reveals is subtler and, for the economics, more important. <strong>Moving data costs compute.</strong> The all-to-all kernels do not run on dedicated networking silicon; they run on the same streaming multiprocessors that would otherwise be doing matrix multiplication. </p><p>Every SM assigned to push bytes through the fabric is an SM not computing an expert. <strong>DeepEP&#8217;s first version spent around 24 SMs</strong> on the communication for a training-scale all-to-all. </p><p>Its <strong>second version</strong>, a substantial rewrite that moved from a custom backend to a more lightweight one built on NVIDIA&#8217;s NCCL, cut that to <strong>between four and six SMs</strong> for the same work while matching or exceeding the old bandwidth. </p><p>The library&#8217;s authors describe the V2 rewrite as achieving extreme performance with <strong>several times fewer SM resources,</strong> and the measured table backs the claim: up to 1.3 times the peak bandwidth at up to four times fewer SMs.</p><p>But the decode path is hungrier than the training path, and here the numbers sharpen into a real cost. To hit maximum throughput on the NVLink decode all-to-all, DeepEP&#8217;s table shows the kernel consuming 64 streaming multiprocessors<strong>. A B200 has 148 of them.</strong> That is forty-three percent of the entire accelerator spent moving data rather than computing, in the configuration tuned for speed. </p><p>You can run the same kernel in a low-SM mode that uses 24, but you give up bandwidth to do it. The library also offers genuinely zero-SM paths for pipeline parallelism, <strong>context parallelism</strong>, and certain RDMA transfers, by offloading the movement to copy engines and the network cards directly, and a great deal of the <strong>engineering frontier in 2026 </strong>is about pushing more of the all-to-all onto those zero-SM paths. </p><p>The reason that frontier exists is <strong>Figure 3.</strong></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!jiut!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4713276c-ab2d-4df2-8213-757d8836a8bc_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!jiut!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4713276c-ab2d-4df2-8213-757d8836a8bc_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!jiut!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4713276c-ab2d-4df2-8213-757d8836a8bc_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!jiut!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4713276c-ab2d-4df2-8213-757d8836a8bc_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!jiut!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4713276c-ab2d-4df2-8213-757d8836a8bc_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!jiut!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4713276c-ab2d-4df2-8213-757d8836a8bc_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/4713276c-ab2d-4df2-8213-757d8836a8bc_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Bar chart of streaming multiprocessors consumed by communication kernels: DeepEP V1 training at 24 SMs, V2 training at 5, NVLink decode min-SM mode at 24, NVLink decode max-throughput mode at 64, against a reference line of 148 total SMs on a B200.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Bar chart of streaming multiprocessors consumed by communication kernels: DeepEP V1 training at 24 SMs, V2 training at 5, NVLink decode min-SM mode at 24, NVLink decode max-throughput mode at 64, against a reference line of 148 total SMs on a B200." title="Bar chart of streaming multiprocessors consumed by communication kernels: DeepEP V1 training at 24 SMs, V2 training at 5, NVLink decode min-SM mode at 24, NVLink decode max-throughput mode at 64, against a reference line of 148 total SMs on a B200." srcset="https://substackcdn.com/image/fetch/$s_!jiut!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4713276c-ab2d-4df2-8213-757d8836a8bc_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!jiut!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4713276c-ab2d-4df2-8213-757d8836a8bc_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!jiut!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4713276c-ab2d-4df2-8213-757d8836a8bc_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!jiut!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4713276c-ab2d-4df2-8213-757d8836a8bc_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong><span>FIG 3</span></strong> The all-to-all is not free even when the wire is fast, because it is paid for in the same currency as the math. At peak decode throughput the communication kernel claims forty-three percent of the GPU.</figcaption></figure></div><p>There is real craft in how DeepEP hides this. The library exposes a hook-based mechanism for overlapping communication with computation: a dispatch is launched,<strong> independent work proceeds on the compute stream </strong>while the data is in flight, and only when the result is needed does the kernel wait. </p><p>Done well, the all-to-all latency disappears behind the expert computation and the SM cost is the only thing left to account for. Done badly, the all-to-all stalls the pipeline and the expensive accelerators sit idle waiting for the network. The <strong>difference between those two outcomes</strong> is most of the difference between a good mixture-of-experts deployment and a wasteful one, and none of it is visible in a<strong> FLOP count.</strong></p><p>The second version pushes the SM problem harder by moving the data movement off the streaming multiprocessors entirely wherever it can. Its experimental branches<strong> expose zero-SM paths</strong> for pipeline and context parallelism, handing the transfers to the GPU&#8217;s copy engines, and a <strong>zero-SM remote-memory primitive</strong> the authors call Engram that lets one device reach into another&#8217;s memory over RDMA without spending a single SM on the transfer. </p><p>The motivation is exactly <strong>Figure 3: </strong>every SM the network gives back is an SM the experts can use. The rewrite also abandoned the custom communication backend for a <strong>lighter one built on NVIDIA&#8217;s NCCL</strong>, which let it reuse existing communicators and scale the expert-parallel domain to as many as two thousand devices, far past anything a production model currently needs. </p><p>A separate branch rebuilds the kernels around the tensor-memory-accelerator instructions on<strong> Hopper and Blackwell</strong>, shrinking SM usage again and adding native four-bit support, which is how the dispatch leg of the toll gets cheaper at the same moment the experts do.</p><p>None of this would matter if the expert computation itself were not reorganized to match. The all-to-all delivers a <strong>variable number of tokens to each expert</strong>, because the gating network does not distribute traffic evenly, and a standard batched matrix multiply assumes a fixed shape. </p><p>The answer, embodied in DeepSeek&#8217;s companion <strong>DeepGEMM library</strong>, is a grouped matrix multiply that processes each expert&#8217;s variable token count as a contiguous segment, so the expert math runs as one efficient kernel rather than a ragged collection of small ones. </p><p>The <strong>communication and the computation are co-designed</strong>: the all-to-all produces exactly the memory layout the grouped GEMM wants to consume. </p><p>Pull either apart from the other and the efficiency collapses, which is part of why a<strong> tuned mixture-of-experts stack</strong> is so much harder to assemble than the FLOP savings would suggest.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>What every token pays at the door</h2><p>The bandwidth numbers describe the pipe. The next question is how much the model tries to push through it, and that can be computed directly from the geometry, which makes it <strong>one of the few places in this analysis </strong>where the arithmetic is exact rather than measured.</p><p>Take <strong>DeepSeek-V3&#8217;s mixture-of-experts layer.</strong> Each token&#8217;s hidden activation is a vector of 7168 values. On the dispatch, those values are sent in eight-bit precision, so one byte each, and they are sent to each of the eight selected experts. That is <strong>7168 times 8, roughly 56 kilobytes of dispatch traffic</strong> per token per layer. </p><p>On the combine, each of the eight experts returns an output vector of the same width, but the combine is performed in sixteen-bit precision to preserve the accuracy of the reduction, so two bytes each. That is <strong>7168 times 2 times 8</strong>, roughly <strong>112 kilobytes of combine traffic per token</strong> per layer. </p><p>Add them and a single token, passing through a single mixture-of-experts layer, generates about 168 kilobytes of all-to-all traffic.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!2NvU!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8fe61e25-78e5-4414-864e-a3860fd9cdc0_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!2NvU!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8fe61e25-78e5-4414-864e-a3860fd9cdc0_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!2NvU!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8fe61e25-78e5-4414-864e-a3860fd9cdc0_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!2NvU!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8fe61e25-78e5-4414-864e-a3860fd9cdc0_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!2NvU!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8fe61e25-78e5-4414-864e-a3860fd9cdc0_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!2NvU!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8fe61e25-78e5-4414-864e-a3860fd9cdc0_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/8fe61e25-78e5-4414-864e-a3860fd9cdc0_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Stacked bar chart comparing bytes moved per token per layer: a dense layer at 14 KB with no routing, versus a mixture-of-experts layer at 168 KB total, split into 56 KB of eight-bit dispatch and 112 KB of sixteen-bit combine.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Stacked bar chart comparing bytes moved per token per layer: a dense layer at 14 KB with no routing, versus a mixture-of-experts layer at 168 KB total, split into 56 KB of eight-bit dispatch and 112 KB of sixteen-bit combine." title="Stacked bar chart comparing bytes moved per token per layer: a dense layer at 14 KB with no routing, versus a mixture-of-experts layer at 168 KB total, split into 56 KB of eight-bit dispatch and 112 KB of sixteen-bit combine." srcset="https://substackcdn.com/image/fetch/$s_!2NvU!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8fe61e25-78e5-4414-864e-a3860fd9cdc0_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!2NvU!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8fe61e25-78e5-4414-864e-a3860fd9cdc0_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!2NvU!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8fe61e25-78e5-4414-864e-a3860fd9cdc0_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!2NvU!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8fe61e25-78e5-4414-864e-a3860fd9cdc0_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong><span>FIG 4</span></strong> A derived figure for DeepSeek-V3 geometry. The routing multiplies per-token data movement by roughly twelve and, unlike a dense layer, all of it has to cross the fabric. The model stacks 58 of these layers.</figcaption></figure></div><p>Set that against what a dense layer moves for the same token, which is essentially nothing across the fabric: the activation stays on the device and the only traffic is the <strong>local memory read of about 14 kilobytes</strong>. The mixture-of-experts layer moves roughly twelve times as much data per token, and the crucial difference is not the multiple but the destination. </p><p>The dense traffic stays on-chip. The mixture-of-experts traffic crosses the network. And<strong> this happens 58 times</strong> as the token descends through the model.</p><p>Two things follow from the structure of those 168 kilobytes. The first is that the combine is twice the dispatch, because the combine runs in higher precision. This is not an arbitrary choice; reducing eight expert outputs in <strong>eight-bit precision degrades quality unacceptably</strong>, so the field has settled on eight-bit dispatch and sixteen-bit combine as the standard, and that asymmetry means the return trip is the more expensive leg. </p><p>Any optimization that can compress the combine, including the four-bit experiments now appearing in DeepEP&#8217;s experimental branches, attacks the larger half of the toll.</p><p>The second is that the toll is paid per token, which means the decode phase, where tokens are generated one at a time, pays it in the worst possible way. In prefill, <strong>thousands of tokens are dispatched together and </strong>the <strong>all-to-all amortizes its latency</strong> across an enormous batch; the kernel runs in its throughput regime and the bandwidth numbers of Figure 2 apply. </p><p>In decode, a single step might dispatch only a handful of tokens per sequence, the batch is tiny, the <strong>bandwidth of the pipe</strong> is irrelevant because the pipe is nearly empty, and what dominates is the fixed round-trip latency of reaching across the fabric and back. </p><p>This is why the low-latency decode kernels exist as a separate class, why they are <strong>willing to burn 64 SMs to shave microseconds</strong>, and why decode is the phase where the mixture-of-experts toll hurts most. </p><p>It is also why the entire industry serves prefill and decode on separately tuned pools of hardware, a point this publication examined at length one issue ago and which the all-to-all only sharpens.</p><p>The decode penalty is worth making concrete, because it is where the toll is most counterintuitive. At a service level of a hundred tokens per second per user, the budget for generating one token is ten milliseconds, and into that budget the model must <strong>fit 58 mixture-of-experts layers</strong>, each with a dispatch and a combine that reach across the fabric. </p><p>Inside the NVLink domain a round trip is measured in microseconds and 58 of them fit with room to spare; across the <strong>InfiniBand fabric</strong> the same round trips, with their higher fixed latency, begin to eat the budget directly. </p><p>That is why the decode all-to-all spends 64 SMs to shave microseconds, and why a decode deployment forced to leave the <strong>NVLink domain for its all-to-all can miss its latency target</strong> even when its aggregate bandwidth looks adequate on paper. In decode, latency is the currency, and the fabric boundary is where it gets spent.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share The Software Frontier&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share The Software Frontier</span></a></p><div><hr></div><h2>The hottest expert sets the clock</h2><p>There is a failure mode hiding inside the all-to-all that the bandwidth numbers do not capture at all, and it is the one that most often separates a deployment hitting its theoretical throughput from one falling well short of it. An <strong>all-to-all is a synchronization barrier. </strong></p><p>The combine cannot complete until every expert has returned its outputs, which means the slowest expert on the most overloaded device sets the pace for the entire operation. If the gating network sends a disproportionate share of tokens to a <strong>handful of popular experts</strong>, the devices holding those experts become stragglers, and every other device in the domain waits on them.</p><p>Expert load is not uniform in practice, and it is not even stable. Certain experts specialize in patterns that appear frequently in real traffic, and the imbalance shifts with the workload. Survey work documents the consequence plainly: <strong>imbalanced token distribution</strong> causes device underutilization, and the whole expensive all-to-all runs at the speed of its hottest path. </p><p>A mixture-of-experts deployment can have perfectly adequate aggregate bandwidth and still bleed throughput because the load is lumpy.</p><p>DeepSeek&#8217;s answer in production is an expert-parallel load balancer that the community has <strong>reproduced under the name EPLB</strong>. The mechanism is to identify the high-load experts from live deployment statistics and replicate them: a hot expert is duplicated onto multiple devices so that the tokens destined for it can be spread, flattening the straggler. This is a direct trade of memory for balance. </p><p>You spend extra capacity holding redundant copies of the popular experts in order to keep the all-to-all from stalling on them. It works, and it is now standard, but it is<strong> another line on the bill that the sparsity story did not mention</strong>, and it interacts with the deployment topology in a way that is worth seeing concretely.</p><p>DeepSeek runs the same model checkpoint as two physically different machines, one for each phase, and the contrast is the clearest illustration in the field of how the all-to-all reshapes a deployment. According to DeepSeek&#8217;s own published inference overview and the <strong>CloudMatrix serving analysis</strong> that reconstructs it, the prefill machine groups four nodes, 32 GPUs, into a single unit running 32-way expert parallelism alongside 32-way data parallelism. </p><p>Across those 32 GPUs the routed experts are distributed nine to a device once the redundant copies of the popular experts are counted, with the shared expert and the <strong>attention mechanism replicated on every one.</strong> The raw figure would be eight; the ninth is the load balancer at work. </p><p>The decode machine expands the same model to 18 nodes, 144 GPUs, running 144-way expert parallelism and <strong>144-way data parallelism</strong>, where each device holds only about two routed experts.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!5OdG!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fba346f04-17d0-4e74-8005-8d1018faff52_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!5OdG!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fba346f04-17d0-4e74-8005-8d1018faff52_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!5OdG!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fba346f04-17d0-4e74-8005-8d1018faff52_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!5OdG!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fba346f04-17d0-4e74-8005-8d1018faff52_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!5OdG!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fba346f04-17d0-4e74-8005-8d1018faff52_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!5OdG!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fba346f04-17d0-4e74-8005-8d1018faff52_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ba346f04-17d0-4e74-8005-8d1018faff52_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Two-panel chart. Left panel: GPUs in one expert-parallel domain, 32 for prefill versus 144 for decode. Right panel: routed experts per GPU, 8 for prefill versus 1.78 for decode, each plus one shared expert.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Two-panel chart. Left panel: GPUs in one expert-parallel domain, 32 for prefill versus 144 for decode. Right panel: routed experts per GPU, 8 for prefill versus 1.78 for decode, each plus one shared expert." title="Two-panel chart. Left panel: GPUs in one expert-parallel domain, 32 for prefill versus 144 for decode. Right panel: routed experts per GPU, 8 for prefill versus 1.78 for decode, each plus one shared expert." srcset="https://substackcdn.com/image/fetch/$s_!5OdG!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fba346f04-17d0-4e74-8005-8d1018faff52_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!5OdG!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fba346f04-17d0-4e74-8005-8d1018faff52_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!5OdG!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fba346f04-17d0-4e74-8005-8d1018faff52_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!5OdG!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fba346f04-17d0-4e74-8005-8d1018faff52_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong><span>FIG 5</span></strong> One checkpoint, two machines. The decode deployment spreads the experts across more than four times as many GPUs, which is partly about latency and partly about leaving room to replicate the hot experts.</figcaption></figure></div><blockquote><p><em>Why spread the same 256 experts across 144 devices for decode when 32 sufficed for prefill?</em> </p><p><strong>Two reasons</strong>, and both come back to the all-to-all. </p></blockquote><ul><li><p>The first is latency: with<strong> fewer experts resident per device,</strong> each device does less work per step and the decode latency target is easier to hit. </p></li><li><p>The second is precisely the <strong>straggler problem</strong>. Spreading thin leaves headroom to replicate the popular experts without overflowing any device&#8217;s memory, so the load balancer has somewhere to put the redundant copies. </p></li></ul><p>The decode machine is wider not because the math demands it but because the communication and the balance do. <strong>The shape of the deployment is dictated by the toll,</strong> not the FLOPs.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>The toolchain the toll demanded</h2><p>The all-to-all did not only reshape the hardware and the kernels. It pulled an entire toolchain into being around itself, and the size of that toolchain is the clearest measure of how far the cost migrated from the math. </p><p>A <strong>2026 mixture-of-experts</strong> serving stack at frontier scale is not a model and a runtime. It is a model, a communication library, a grouped-GEMM library, an expert load balancer, a disaggregation layer, and an overlap scheduler, each of which exists to manage some facet of the routing tax. The FLOP count described <strong>one of those six boxes</strong>.</p><p>Consider the overlap problem at the level of an entire forward pass rather than a single layer. <strong>Hiding the all-to-all behind computation</strong> works within a layer, but the decode phase is so latency-sensitive that the field has gone further and split each batch in two, running the communication of one half against the computation of the other in a continuous pipeline. </p><p><strong>SGLang&#8217;s two-batch overlap </strong>and the analogous schemes in other runtimes exist for one reason: to keep the expensive accelerators busy with expert math while the all-to-all of a different microbatch is in flight. It is the same instinct as the <strong>kernel-level hooks</strong>, lifted to the level of the request scheduler, and it is now a standard part of large-scale deployments rather than an exotic optimization.</p><p>Disaggregation adds a second communication problem on top of the all-to-all. Once prefill and decode run on separate pools of hardware, the <strong>key-value cache </strong>computed during prefill<strong> has to be shipped</strong> to the decode pool before generation can begin, and at frontier scale that transfer is large enough and frequent enough to need its own engine. </p><p>The <strong>Mooncake transfer engine</strong> and the equivalent layers inside vLLM and SGLang exist to move key-value caches across the network efficiently, overlapping the transfer with computation so the handoff does not stall the pipeline. This is a network tax distinct from the all-to-all, and it is the price of the <strong>prefill-decode split</strong> that the all-to-all economics make worthwhile in the first place. </p><p>The <strong>two taxes are siblings</strong>: both are consequences of spreading one model&#8217;s inference across many devices, and both are paid down by the same instinct of overlapping transfer with compute.</p><p>The lesson in the length of that list is that the sparsity bargain did not merely move the cost to the network. It moved the cost to a place where <strong>extracting good performance</strong> requires assembling and tuning half a dozen interacting systems, any one of which, misconfigured, hands the savings back. </p><p>The vLLM and SGLang playbooks both carry warnings to this effect, and <strong>AMD&#8217;s ROCm guide to the vLLM mixture-of-experts options </strong>is blunt that the wrong combination of tensor, data, pipeline, and expert parallelism can duplicate the key-value cache many times over and consume far more memory than expected. </p><p>The <strong>FLOP count said the model got cheaper</strong>. The operations manual says it got more complicated, and the complication is where a large part of the real cost now lives.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Does wide expert parallelism pay for itself?</h2><p>All of this is overhead, and the natural reaction to a catalogue of overhead is to minimize it. If the <strong>all-to-all is the cost</strong>, why not keep the expert-parallel domain small, so the all-to-all stays inside a tight, fast group of devices? </p><p>The answer is that narrowing the domain trades one cost for another, and the trade does not run in the obvious direction. </p><p>Wider expert parallelism, counterintuitively, often <strong>produces more throughput per GPU</strong>, not less, and understanding why is the crux of whether the whole approach earns its keep.</p><p>The mechanism is expert packing. When experts are spread across more devices, each device holds fewer of them, which means <strong>more of each device&#8217;s memory and compute</strong> can be devoted to the batch of tokens currently being processed rather than to holding a large slice of the model. </p><p>Larger effective batches per device improve the arithmetic intensity of the expert matrix multiplications, the kernels run closer to the hardware&#8217;s peak, and the <strong>per-GPU throughput rises</strong>, provided the all-to-all overhead can be kept hidden behind that larger computation. The question is always whether the communication grows faster than the packing benefit, and up to a point, on the right interconnect, it does not.</p><p><strong>NVIDIA&#8217;s measurements on the GB200 NVL72 quantify the dividend directly</strong>. Moving from an eight-way expert-parallel configuration to a 32-way one delivers up to 1.8 times the output token throughput per GPU, at a fixed service level of a hundred tokens per second per user, with disaggregated serving and multi-token prediction in both cases. </p><p>Same hardware, same latency target, nearly double the per-GPU output, purely from going wider on expert parallelism.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!CaiQ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10a0a69c-89d2-406f-944f-319e09edd0a7_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!CaiQ!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10a0a69c-89d2-406f-944f-319e09edd0a7_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!CaiQ!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10a0a69c-89d2-406f-944f-319e09edd0a7_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!CaiQ!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10a0a69c-89d2-406f-944f-319e09edd0a7_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!CaiQ!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10a0a69c-89d2-406f-944f-319e09edd0a7_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!CaiQ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10a0a69c-89d2-406f-944f-319e09edd0a7_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/10a0a69c-89d2-406f-944f-319e09edd0a7_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Bar chart of output tokens per second per GPU, normalized to EP8 at 100: EP8 at 100, EP32 at 180, showing 1.8 times the per-GPU throughput from wider expert parallelism.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Bar chart of output tokens per second per GPU, normalized to EP8 at 100: EP8 at 100, EP32 at 180, showing 1.8 times the per-GPU throughput from wider expert parallelism." title="Bar chart of output tokens per second per GPU, normalized to EP8 at 100: EP8 at 100, EP32 at 180, showing 1.8 times the per-GPU throughput from wider expert parallelism." srcset="https://substackcdn.com/image/fetch/$s_!CaiQ!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10a0a69c-89d2-406f-944f-319e09edd0a7_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!CaiQ!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10a0a69c-89d2-406f-944f-319e09edd0a7_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!CaiQ!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10a0a69c-89d2-406f-944f-319e09edd0a7_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!CaiQ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10a0a69c-89d2-406f-944f-319e09edd0a7_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong><span>FIG 6</span></strong> NVIDIA&#8217;s Wide-EP figures on the NVL72. Going wider improves per-GPU throughput, because the packing benefit outweighs the added all-to-all, as long as the all-to-all stays inside the NVLink domain.</figcaption></figure></div><p>The decisive qualifier is the last clause. The 1.8 times holds because the 32-way all-to-all stays inside the NVL72&#8217;s NVLink domain, where Figure 2 says it runs at <strong>726 gigabytes per second. </strong></p><p>The dividend exists because the wire is fast enough that going wider does not push the communication off the cliff. Try the same widening on a cluster where<strong> 32-way expert parallelism forces the all-to-all across InfiniBand</strong>, and the calculus inverts: the packing benefit is swamped by the eightfold bandwidth penalty of leaving the domain, and wider becomes worse. </p><p>This is the same fact from a different angle. The reason the <strong>rack-scale NVLink domain is worth its price </strong>is that it is what makes the wide-EP dividend positive instead of negative.</p><p>There is a second lever working alongside the width, and it appears in <strong>nearly every published wide-EP result</strong>: multi-token prediction. Rather than generating one token per forward pass, the model proposes several and verifies them together, which raises the number of tokens flowing through each all-to-all and pushes the decode kernel out of its worst, smallest-batch regime toward something the bandwidth can amortize. </p><p>Multi-token prediction and wide expert parallelism are complementary for the same underlying reason: <strong>both increase the work done </strong>per round trip across the fabric, and the all-to-all rewards anything that makes its fixed latency a smaller fraction of the whole. </p><p>The dividend in <strong>Figure 6 is partly a multi-token-prediction dividend</strong>, which is why NVIDIA and SGLang report the two together. They are deployed together because they solve the same problem from two directions.</p><p>So the answer to whether wide expert parallelism pays for itself is conditional, and the condition is the interconnect. Inside a sufficiently large fast domain, wider is genuinely better and the measurements prove it. Outside one, wider is a trap. The<strong> crossover sits exactly at the boundary of the NVLink domain</strong>, which is why the size of that domain, 8 GPUs yesterday, 72 today, the same 72 at higher bandwidth tomorrow, is the number that determines how far the dividend extends. </p><p><strong>Expert parallelism</strong> and the interconnect are not two separate decisions. They are one decision, and the hardware vendor has been making half of it for you.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>How much is silicon, and how much is numerics</h2><p>It is tempting to attribute the throughput of a Blackwell mixture-of-experts deployment to the silicon, and the marketing encourages it, but the <strong>public measurements </strong>let us decompose the uplift, and the decomposition is instructive about where the real leverage sits.</p><p>The <strong>LMSYS and SGLang teams</strong> have published a careful progression of DeepSeek serving results on the GB200 NVL72, and the numbers are specific. </p><p>With disaggregated prefill and decode, large-scale expert parallelism, and the conservative numeric configuration of sixteen-bit attention and eight-bit experts, <strong>SGLang reaches 18,471 input tokens per second per GPU</strong> on prefill and 9,087 output tokens per second per GPU on decode, for two-thousand-token sequences. </p><p>Switch to the aggressive configuration, eight-bit attention and four-bit <strong>NVFP4 experts</strong>, and the same system reaches <strong>26,156 input and 13,386 output tokens per second per GPU.</strong> Against the H100 baseline the teams report, those aggressive numbers represent a 3.8 times prefill and 4.8 times decode improvement.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!t5VC!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F58fdb759-c5b7-4999-a918-b5332633c973_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!t5VC!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F58fdb759-c5b7-4999-a918-b5332633c973_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!t5VC!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F58fdb759-c5b7-4999-a918-b5332633c973_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!t5VC!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F58fdb759-c5b7-4999-a918-b5332633c973_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!t5VC!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F58fdb759-c5b7-4999-a918-b5332633c973_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!t5VC!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F58fdb759-c5b7-4999-a918-b5332633c973_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/58fdb759-c5b7-4999-a918-b5332633c973_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Grouped bar chart of tokens per second per GPU for prefill and decode across three configurations: H100 baseline at 6883 prefill and 2789 decode, GB200 with BF16 attention and FP8 experts at 18471 and 9087, and GB200 with FP8 attention and NVFP4 experts at 26156 and 13386, marked as 3.8 times and 4.8 times the baseline.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Grouped bar chart of tokens per second per GPU for prefill and decode across three configurations: H100 baseline at 6883 prefill and 2789 decode, GB200 with BF16 attention and FP8 experts at 18471 and 9087, and GB200 with FP8 attention and NVFP4 experts at 26156 and 13386, marked as 3.8 times and 4.8 times the baseline." title="Grouped bar chart of tokens per second per GPU for prefill and decode across three configurations: H100 baseline at 6883 prefill and 2789 decode, GB200 with BF16 attention and FP8 experts at 18471 and 9087, and GB200 with FP8 attention and NVFP4 experts at 26156 and 13386, marked as 3.8 times and 4.8 times the baseline." srcset="https://substackcdn.com/image/fetch/$s_!t5VC!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F58fdb759-c5b7-4999-a918-b5332633c973_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!t5VC!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F58fdb759-c5b7-4999-a918-b5332633c973_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!t5VC!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F58fdb759-c5b7-4999-a918-b5332633c973_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!t5VC!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F58fdb759-c5b7-4999-a918-b5332633c973_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong><span>FIG 7</span></strong> The Blackwell uplift, decomposed. A large share of the gain over the conservative GB200 configuration comes from dropping the experts to four-bit NVFP4, not from the silicon alone.</figcaption></figure></div><p>The decomposition is the point. The jump from the H100 baseline to the conservative GB200 configuration is the hardware: faster tensor cores, the NVLink domain, more memory bandwidth. But the further jump from the conservative to the aggressive <strong>GB200 configuration, from 18,471 to 26,156 on prefill and from 9,087 to 13,386 on decode</strong>, is numerics. </p><p>It comes from running the experts in four-bit NVFP4 rather than eight-bit. That is a software-and-format change applied to the same rack, and it accounts for a substantial fraction of the total uplift over H100.</p><p>NVFP4 earns its own treatment, and it is a strong candidate for a future issue, but the relevant fact here is why it interacts so favorably with the all-to-all. <strong>Four-bit experts are half the bytes of eight-bit experts</strong>, which directly shrinks the dispatch leg of the toll, and they double the tensor-core throughput of the expert math itself, so the computation that hides the all-to-all gets faster at the same time the all-to-all gets smaller. </p><p>NVIDIA&#8217;s format reportedly holds accuracy within about one percent of the higher-precision baseline on large models through a two-level scaling scheme, and the <strong>accuracy holds up best precisely on the large mixture-of-experts models </strong>where it matters most. The format is, in effect, a second lever on the same toll that the interconnect attacks, and the two compound. </p><p>This is also <strong>why NVIDIA can credibly claim a fivefold reduction in cost per token</strong> from software optimization alone in the two months after Blackwell&#8217;s launch, with no hardware change: a large part of that was kernel and format work on exactly these operations.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>One dollar, or twenty cents</h2><p>The throughput numbers are engineering. The reason they matter is that they convert, almost directly, into the <strong>only number a serving operator actually cares about</strong>, which is dollars per million tokens. And here the all-to-all moves from being a technical concern to being the dominant line item in the unit economics.</p><p>The cleanest demonstration in the public record is the <strong>LMSYS deployment of DeepSeek on 96 H100 GPUs</strong>, twelve nodes of eight, using prefill-decode disaggregation and large-scale expert parallelism with the full DeepEP, DeepGEMM, and EPLB stack. </p><p>That deployment reached 52,300 input tokens per second and 22,300 output tokens per second per node, and when the team translated the throughput into cost, it came to twenty cents per million output tokens. That figure is <strong>roughly one-fifth of what DeepSeek&#8217;s own public API charged at the time</strong>, achieved on rented hardware by an outside team reproducing the architecture.</p><p>The comparison that matters most, though, is the one against the naive alternative on identical hardware. The same report states that the optimized expert-parallel strategy improved output throughput by up to five times over <strong>vanilla tensor parallelism</strong> using the same resources. Five times the throughput on the same GPUs is five times lower cost per token. </p><p>The all-to-all engineering, getting the dispatch and combine to run efficiently inside the fast domain, hiding the latency behind computation, balancing the hot experts, is the entire difference between a deployment at twenty cents and a deployment at a dollar.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!OqS9!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe045fa0-2808-4ee2-8e81-4476ee597b4c_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!OqS9!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe045fa0-2808-4ee2-8e81-4476ee597b4c_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!OqS9!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe045fa0-2808-4ee2-8e81-4476ee597b4c_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!OqS9!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe045fa0-2808-4ee2-8e81-4476ee597b4c_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!OqS9!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe045fa0-2808-4ee2-8e81-4476ee597b4c_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!OqS9!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe045fa0-2808-4ee2-8e81-4476ee597b4c_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/fe045fa0-2808-4ee2-8e81-4476ee597b4c_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Horizontal bar chart of US dollars per million output tokens: vanilla tensor parallel on 96 H100 at one dollar, official DeepSeek API as a reference at one dollar, and PD plus large-scale expert parallelism self-hosted on 96 H100 at twenty cents, five times cheaper on identical hardware.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Horizontal bar chart of US dollars per million output tokens: vanilla tensor parallel on 96 H100 at one dollar, official DeepSeek API as a reference at one dollar, and PD plus large-scale expert parallelism self-hosted on 96 H100 at twenty cents, five times cheaper on identical hardware." title="Horizontal bar chart of US dollars per million output tokens: vanilla tensor parallel on 96 H100 at one dollar, official DeepSeek API as a reference at one dollar, and PD plus large-scale expert parallelism self-hosted on 96 H100 at twenty cents, five times cheaper on identical hardware." srcset="https://substackcdn.com/image/fetch/$s_!OqS9!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe045fa0-2808-4ee2-8e81-4476ee597b4c_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!OqS9!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe045fa0-2808-4ee2-8e81-4476ee597b4c_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!OqS9!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe045fa0-2808-4ee2-8e81-4476ee597b4c_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!OqS9!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe045fa0-2808-4ee2-8e81-4476ee597b4c_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong><span>FIG 8</span></strong> Same 96 GPUs, two ways of organizing them. The five-fold gap between vanilla tensor parallelism and tuned expert parallelism is, almost entirely, the all-to-all done well versus done naively.</figcaption></figure></div><p>Put that five-fold against the backdrop of where inference pricing has gone, and the stakes of the routing tax become clear. The price of <strong>frontier-class inference has fallen by something close to fifty times in three years</strong>, from around twenty dollars per million tokens for GPT-4-class output in late 2022 to roughly forty cents in early 2026. </p><p>Public trackers attribute the collapse to four compounding forces, and mixture-of-experts together with expert parallelism is explicitly one of them, alongside hardware efficiency, kernel and compiler optimization, and low-precision formats.<strong> Inference now consumes roughly two-thirds of all AI compute</strong>, having crossed over from a minority of it only a couple of years ago. </p><p>In that environment a five-fold cost difference is not a margin to be optimized later. It is the difference between a viable serving business and an unviable one.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!CvQD!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa7a5f99e-3225-4cd6-9684-5d6e2eb8c25c_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!CvQD!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa7a5f99e-3225-4cd6-9684-5d6e2eb8c25c_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!CvQD!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa7a5f99e-3225-4cd6-9684-5d6e2eb8c25c_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!CvQD!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa7a5f99e-3225-4cd6-9684-5d6e2eb8c25c_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!CvQD!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa7a5f99e-3225-4cd6-9684-5d6e2eb8c25c_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!CvQD!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa7a5f99e-3225-4cd6-9684-5d6e2eb8c25c_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a7a5f99e-3225-4cd6-9684-5d6e2eb8c25c_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Line chart on a log scale of US dollars per million tokens for GPT-4-class output from 2022 to 2026: 20 dollars in late 2022, 5 in 2023, 2 in 2024, 0.8 in 2025, and 0.4 in early 2026, about a fifty-fold decline, with drivers listed as hardware, kernels, mixture-of-experts plus expert parallelism, and four-bit formats.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Line chart on a log scale of US dollars per million tokens for GPT-4-class output from 2022 to 2026: 20 dollars in late 2022, 5 in 2023, 2 in 2024, 0.8 in 2025, and 0.4 in early 2026, about a fifty-fold decline, with drivers listed as hardware, kernels, mixture-of-experts plus expert parallelism, and four-bit formats." title="Line chart on a log scale of US dollars per million tokens for GPT-4-class output from 2022 to 2026: 20 dollars in late 2022, 5 in 2023, 2 in 2024, 0.8 in 2025, and 0.4 in early 2026, about a fifty-fold decline, with drivers listed as hardware, kernels, mixture-of-experts plus expert parallelism, and four-bit formats." srcset="https://substackcdn.com/image/fetch/$s_!CvQD!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa7a5f99e-3225-4cd6-9684-5d6e2eb8c25c_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!CvQD!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa7a5f99e-3225-4cd6-9684-5d6e2eb8c25c_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!CvQD!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa7a5f99e-3225-4cd6-9684-5d6e2eb8c25c_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!CvQD!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa7a5f99e-3225-4cd6-9684-5d6e2eb8c25c_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong><span>FIG 9</span></strong> The price floor that makes a routing tax of cents per token worth a flagship. Expert parallelism is one of the four named drivers of this curve, not a footnote to it.</figcaption></figure></div><div><hr></div><h2>Making the domain bigger than the problem</h2><p>Step back from the individual numbers and a single strategic motion organizes all of them. The <strong>mixture-of-experts architecture</strong> created a communication problem. </p><p>The hardware industry&#8217;s response has been to make the <strong>fast communication domain large enough</strong> to swallow the problem whole, and the trajectory of that response is the most reliable predictor of where serving economics go next.</p><p>DeepSeek&#8217;s own engineers, in their published reflections on the <strong>hardware lessons of training V3</strong>, frame the future in exactly these terms. They call for the convergence of scale-up and scale-out, for precise low-precision compute units, and for innovations in low-latency communication fabrics.</p><p> Read against this issue, that is a wish list written by the people paying the all-to-all toll, addressed to the people who can make the domain bigger. The <strong>scale-up and scale-out convergence</strong> they ask for is<strong> precisely the elimination of the cliff in Figure 1:</strong> a world where crossing from one node to the next does not cost a factor of eight, because the fast domain has grown to encompass both.</p><p>NVIDIA is building toward exactly that, and is increasingly explicit that it is doing so for this reason. The<strong> NVL72 took the NVLink domain from 8 to 72.</strong> The NVLink Switch architecture is specified to reach 576 GPUs in a single non-blocking fabric. The Rubin generation lifts the per-GPU bandwidth again and ties the increase directly, in NVIDIA&#8217;s own framing, to the all-to-all needs of mixture-of-experts models. </p><p>Each step is sold, more openly than the last, as a larger container for the communication problem that sparsity created. The <strong>architecture and the interconnect are co-evolving</strong>, and the direction is set: the domain keeps growing, the cliff keeps receding, and the toll keeps shrinking as a fraction of the work, without ever quite reaching zero.</p><p>The domain cannot grow without limit, and the constraints on how far it can stretch are physical. NVLink at rack scale runs over copper, which is cheap and reliable but <strong>reaches only a couple of meters</strong>; pushing the domain past a single rack toward the <strong>576-GPU fabric the switch silicon</strong> can address means either optical interconnect, with its added cost, power draw, and failure modes, or denser and hotter racks than the current design. </p><p><strong>Power and cooling </strong>are already near the edge of what a standard data center hall delivers per rack, which is why the NVL72 is<strong> liquid-cooled</strong> and why each new generation leans harder on liquid. And the fault domain grows with the fabric, because a larger coherent domain is a larger blast radius for a single failure. </p><p>The trajectory is set toward bigger domains, but each <strong>expansion buys less headroom than the last</strong> against a wall of copper reach, power density, and fault tolerance that the all-to-all cannot argue its way past.</p><p>What this does not resolve is the dependency it creates. An operator who builds a serving business on wide expert parallelism is building on the assumption that the<strong> fast domain will keep growing</strong>, and that assumption ties the economics of the model layer to the roadmap of a single interconnect vendor. </p><p>The <strong>wide-EP dividend is real</strong>, but it is contingent on hardware that one company predominantly supplies, and the contingency is worth naming. The cheapest way to serve a frontier mixture-of-experts model in 2026 runs through a rack that is, for now, effectively sole-sourced. </p><p>That is a strategic fact about the inference market as much as a technical one, and it is the part of the story most likely to matter in the issues to come.</p><blockquote><p><em><span>The cheapest way to serve a frontier mixture-of-experts model in 2026 runs through a rack that is, for now, effectively sole-sourced. That is a strategic fact as much as a technical one.</span></em></p></blockquote><p>The <strong>dependency has not gone unanswered</strong>. An industry that has watched a single vendor&#8217;s interconnect become the determinant of mixture-of-experts economics has begun to organize alternatives. </p><p>The <strong>UALink consortium</strong> and the <strong>Ultra Ethernet </strong>effort are both attempts to build an open scale-up fabric that could host the all-to-all without routing through one company&#8217;s switches, and <strong>AMD&#8217;s serving stack </strong>now carries its own expert-parallel communication path, a port of the DeepEP ideas onto its accelerators. </p><p>None of these has yet demonstrated the rack-scale all-to-all bandwidth of an NVL72 in production, and the gap is real, but the direction of the effort is itself a <strong>measure of how much the all-to-all matters</strong>. An entire alternative-hardware ecosystem is organizing around the single operation that this issue is about.</p><p>There is also a cost that none of the throughput numbers capture, which is reliability. A 144-GPU decode deployment is one coordinated system, and the all-to-all is a <strong>synchronization barrier across all of it</strong>, which means a fault or a slowdown on any single device degrades the whole. </p><p>The larger the expert-parallel domain, the more devices have to stay healthy and in lockstep for the all-to-all to complete on time, and the operational burden of keeping a domain of that size running at frontier latency is substantial. </p><p><strong>DeepSeek&#8217;s own diagnostic tooling</strong> for locating slow ranks in a DeepEP deployment exists because, at this scale, finding the one straggling device in a domain of hundreds is a routine and necessary operation. </p><p>The wide-EP dividend is real, but it is collected by operators who can keep a very large, very tightly coupled machine running, and that capability is a cost the smaller-domain alternatives never have to pay.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>What to actually do</h2><p>The analysis <strong>resolves into a handful of decisions</strong> that an operator faces in practice, and they follow from the structure rather than from any single benchmark.</p><p>The first decision is whether to use expert parallelism at all, and the honest answer is that it depends entirely on whether your all-to-all can be kept inside a fast domain.</p><p>If you are serving a <strong>frontier mixture-of-experts model at scale</strong> and you have access to a rack-scale NVLink domain, wide expert parallelism is the right tool and the measurements say to go as wide as the domain allows, because the packing dividend is positive inside the fast fabric. </p><p>If your all-to-all would have to cross InfiniBand to go wider, stop widening before it does, because the cliff inverts the dividend. The boundary of the NVLink domain is the boundary of the decision.</p><p>The second decision is how to split the phases. Prefill and decode want different all-to-all kernels, <strong>different expert-parallel widths</strong>, and in DeepSeek&#8217;s production case different physical machines entirely. The decode machine should be wider, both to hit latency targets and to leave room for the load balancer to replicate hot experts. </p><p>If you cannot afford to disaggregate, the decode phase is where the toll will hurt, and the low-latency kernels are where to spend your tuning effort. The <strong>vLLM and SGLang playbooks</strong> both warn, correctly, that the wrong parallelism strategy can duplicate key-value caches across the domain and consume many times the memory you expected, so the parallelism decision is not only about the all-to-all but about what else it forces to be replicated.</p><p>The third decision is precision, and it is mostly free throughput if you are on Blackwell.<strong> Four-bit NVFP4 experts shrink the dispatch leg </strong>of the toll and double the expert math throughput at an accuracy cost that, on large models, is small. The aggressive configuration in Figure 7 is not a marginal tuning; it is a large fraction of the total uplift, and it attacks the same toll the interconnect attacks. If <strong>your hardware supports it and your accuracy budget allows it,</strong> it is among the highest-leverage changes available.</p><p>And the fourth decision is whether you need any of this at all. If your workload is single-user or small-scale, the<strong> KTransformers lesson stands</strong>: a mixture-of-experts model on a single node never pays the toll, and the entire apparatus of expert parallelism is overhead you can decline. </p><p>The all-to-all economics in this issue are the economics of serving at frontier scale and <strong>frontier latency.</strong> Below that scale, the right move is to keep the experts local and let the toll switch stay off.</p><p>The deeper lesson is the one the sparsity story obscured for two years. Mixture-of-experts did not make inference cheaper by doing less work. It moved the work from a <strong>place that was easy to scale</strong>, the arithmetic, to a place that was hard, the network, and then the hardware industry spent two product generations and a great deal of money making the network easy to scale too. </p><p>The bargain was always real. It was just never free, and the bill was always going to come due on the wire.<strong> Knowing where it comes due</strong>, and how much, is most of what it takes to serve these models without overpaying. </p><p> The router decides which experts a token needs. The wire decides what that decision costs. For the models that now define the frontier, the <strong>wire is the more expensive </strong>of the two.</p><div><hr></div><h2>What we are confident about, and what we estimated</h2><p><strong>A</strong></p><p><em>NVLink 5 delivers 1.8 TB/s per GPU; the GB200 NVL72 provides 130 TB/s aggregate all-to-all bandwidth across 72 GPUs, with 13.5 TB of unified HBM3e.</em></p><p><em>NVIDIA GB200 NVL72 datasheet; NVIDIA multi-node NVLink tuning guide; Introl and Spheron interconnect analyses.</em></p><p><strong>A</strong></p><p><em>DeepEP measures dispatch and combine at 726 and 740 GB/s inside the NVLink domain on Blackwell, versus about 90 GB/s each across internode RDMA, on the published V3 workload.</em></p><p><em>DeepEP V2 performance table, deepseek-ai/DeepEP repository.</em></p><p><strong>A</strong></p><p><em>DeepEP&#8217;s decode all-to-all consumes up to 64 SMs at peak throughput; the V2 rewrite cut training all-to-all SM use from 24 to between 4 and 6. A B200 has 148 SMs.</em></p><p><em>DeepEP V2 performance table and release notes; Blackwell architecture specifications.</em></p><p><strong>A</strong></p><p><em>SGLang on the GB200 NVL72 reaches 26,156 prefill and 13,386 decode tokens/sec/GPU with eight-bit attention and NVFP4 experts, reported as 3.8x and 4.8x over H100; the conservative configuration reaches 18,471 and 9,087.</em></p><p><em>LMSYS Org, GB200 NVL72 Part II, September 2025.</em></p><p><strong>A</strong></p><p><em>An LMSYS 96-GPU H100 deployment reached 52.3k input and 22.3k output tokens/sec/node and translated to $0.20 per 1M output tokens, about one-fifth the official API price, and up to 5x the throughput of vanilla tensor parallelism on the same hardware.</em></p><p><em>LMSYS Org, large-scale EP on 96 H100, May 2025.</em></p><p><strong>B</strong></p><p><em>Moving from EP8 to EP32 yields up to 1.8x output throughput per GPU at a fixed 100 tok/s/user SLA on the NVL72, with disaggregated serving and multi-token prediction.</em></p><p><em>NVIDIA, Wide Expert Parallelism on NVL72, January 2026.</em></p><p><strong>B</strong></p><p><em>DeepSeek-V3 runs DP32+EP32 across 32 GPUs for prefill (nine routed experts per GPU plus one shared, including one redundant) and DP144+EP144 across 144 GPUs for decode (about two routed experts per GPU plus one shared).</em></p><p><em>DeepSeek Open Source Week inference system overview (Day 6); CloudMatrix serving analysis (arXiv 2506.12708). The V3 technical report describes a different decode configuration (EP320, one expert per GPU).</em></p><p><strong>B</strong></p><p><em>NVFP4 holds accuracy within roughly one percent of the higher-precision baseline on large models via two-level scaling, and accuracy recovery is strongest on the largest dense and MoE models.</em></p><p><em>NVIDIA NVFP4 technical blogs; Red Hat AI NVFP4 evaluation.</em></p><p><strong>C</strong></p><p><em>A DeepSeek-V3 mixture-of-experts layer moves about 56 KB of dispatch (FP8, top-8) and 112 KB of combine (BF16, top-8) per token, roughly 168 KB total, against about 14 KB for a dense layer.</em></p><p><em>Derived from V3 geometry (hidden 7168, top-8, FP8 dispatch, BF16 combine). Excludes the shared expert and any local-rank optimization.</em></p><p><strong>C</strong></p><p><em>The vanilla-tensor-parallel and official-API reference points of roughly $1.00 per 1M output tokens are derived from the LMSYS statements (optimized $0.20 figure at one-fifth of API, and 5x over vanilla TP).</em></p><p><em>Derived from LMSYS 96-GPU report figures.</em></p><p><strong>C</strong></p><p><em>The H100 baseline in Figure 7 (6,883 prefill, 2,789 decode tokens/sec/GPU) is back-calculated from the reported 3.8x and 4.8x speedups, not independently measured.</em></p><p><em>Derived from LMSYS GB200 Part II reported multipliers.</em></p><p><strong>D</strong></p><p><em>Frontier-class inference pricing has fallen roughly fifty-fold from about $20 to about $0.40 per 1M tokens from late 2022 to early 2026, with MoE plus expert parallelism among four named drivers.</em></p><p><em>Public inference price trackers, 2022 to 2026. Order-of-magnitude trend across vendors, not a single price series.</em></p><p><strong>D</strong></p><p><em>Vera Rubin NVL72 is specified for roughly 3.6 TB/s per GPU and 260 TB/s aggregate, framed by NVIDIA as serving MoE all-to-all needs.</em></p><p><em>NVIDIA NVLink product page and CES 2026 disclosures; pre-release specification subject to change.</em></p><p><em>A = primary or measured | B = single strong vendor or operator source | C = derived by us from sourced inputs | D = directional, treat as trend not point estimate.<br>Character scan: this issue contains zero em dashes and zero en dashes, verified programmatically against the rendered text.</em></p><div><hr></div><h2>Bibliography</h2><ol><li><p><span>DeepSeek-AI. DeepEP: an efficient expert-parallel communication library. GitHub repository, 2025. Performance table, V2 release notes, decode and prefill kernel interfaces.github.com/deepseek-ai/DeepEP</span></p></li><li><p><span>DeepSeek-AI. Insights into DeepSeek-V3: Scaling Challenges and Reflections on Hardware for AI Architectures. arXiv 2505.09343, 2025.arxiv.org/abs/2505.09343</span></p></li><li><p><span>LMSYS Org. Deploying DeepSeek with PD Disaggregation and Large-Scale Expert Parallelism on 96 H100 GPUs. May 2025.lmsys.org/blog/2025-05-05-large-scale-ep</span></p></li><li><p><span>LMSYS Org. Deploying DeepSeek on GB200 NVL72 with PD and Large Scale EP, Part I: 2.7x Higher Decoding Throughput. June 2025.lmsys.org/blog/2025-06-16-gb200-part-1</span></p></li><li><p><span>LMSYS Org. Deploying DeepSeek on GB200 NVL72 with PD and Large Scale EP, Part II: 3.8x Prefill, 4.8x Decode Throughput. September 2025.lmsys.org/blog/2025-09-25-gb200-part-2</span></p></li><li><p><span>LMSYS Org. SGLang and NVIDIA Accelerating SemiAnalysis InferenceMAX and GB200 Together. October 2025.lmsys.org/blog/2025-10-14-sa-inference-max</span></p></li><li><p><span>NVIDIA. Scaling Large MoE Models with Wide Expert Parallelism on NVL72 Rack-Scale Systems. NVIDIA Technical Blog, January 2026.developer.nvidia.com/blog</span></p></li><li><p><span>NVIDIA. GB200 NVL72 product page and datasheet. 130 TB/s NVLink domain, 72-GPU rack specifications.nvidia.com/en-us/data-center/gb200-nvl72</span></p></li><li><p><span>NVIDIA. Inside NVIDIA Blackwell Ultra: The Chip Powering the AI Factory Era. NVIDIA Technical Blog, January 2026.developer.nvidia.com/blog</span></p></li><li><p><span>NVIDIA. Introducing NVFP4 for Efficient and Accurate Low-Precision Inference. NVIDIA Technical Blog, 2025.developer.nvidia.com/blog</span></p></li><li><p><span>NVIDIA. The Economic Value of Inference Software Optimization at the Datacenter Level. April 2026. Fivefold cost-per-token reduction via software.perspectives.nvidia.com</span></p></li><li><p><span>NVIDIA. Multi-Node NVLink Systems Tuning Guide and NVLink / NVLink Switch product documentation. Fifth-generation NVLink and NVSwitch specifications.docs.nvidia.com; nvidia.com/en-us/data-center/nvlink</span></p></li><li><p><span>Microsoft. Achieving Optimal Performance for DeepSeek Expert Parallelism (DeepEP) on Azure. Azure HPC Blog, May 2025.techcommunity.microsoft.com</span></p></li><li><p><span>AMD ROCm. The vLLM MoE Playbook: A Practical Guide to TP, DP, PP and Expert Parallelism. November 2025.rocm.blogs.amd.com</span></p></li><li><p><span>Taming the Titans: A Survey of Efficient LLM Inference Serving. arXiv 2504.19720, 2025. All-to-all as the MoE bottleneck; expert load balancing.arxiv.org/abs/2504.19720</span></p></li><li><p><span>Serving Large Language Models on Huawei CloudMatrix384. arXiv 2506.12708, 2025. DeepSeek DP32+EP32 prefill and DP144+EP144 decode topology.arxiv.org/abs/2506.12708</span></p></li><li><p><span>DeepSeek-AI. DeepSeek-V3/R1 Inference System Overview (Open Source Week, Day 6). February 2025. Production prefill EP32 (9 experts/GPU) and decode EP144 (2 experts/GPU) topology.github.com/deepseek-ai/open-infra-index</span></p></li><li><p><span>Introl. NVLink and Scale-Up Networking. 2026. Scale-up versus scale-out bandwidth ratio; NVL72 physical architecture.introl.com/blog</span></p></li><li><p><span>DigitalOcean. The LLM Inference Trilemma: Throughput, Latency, Cost. April 2026. MoE cost as a game of communication.digitalocean.com/blog</span></p></li><li><p><span>GPUnex. AI Inference Economics: The 1,000x Cost Collapse Reshaping GPUs. February 2026. Inference price trend and drivers.gpunex.com/blog</span></p></li><li><p><span>NVIDIA. NVLink and NVLink Switch, Vera Rubin NVL72 and NVLink 6 disclosures. CES 2026. 260 TB/s aggregate, MoE all-to-all framing.nvidia.com/en-us/data-center/nvlink</span></p></li></ol>]]></content:encoded></item><item><title><![CDATA[Decode Is Memory-Bound. Speculation Is the Arbitrage]]></title><description><![CDATA[Speculative decoding is the only inference optimization that turns idle silicon into tokens without changing a single output. Whether that lands on your bill as a discount or a surcharge is not a prop]]></description><link>https://www.thesoftwarefrontier.com/p/decode-is-memory-bound-speculation</link><guid isPermaLink="false">https://www.thesoftwarefrontier.com/p/decode-is-memory-bound-speculation</guid><dc:creator><![CDATA[Lorenzo Bradanini]]></dc:creator><pubDate>Thu, 25 Jun 2026 10:19:56 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!1i2p!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff2f8e194-f28b-4508-aa07-1a6406746102_1122x1122.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!1i2p!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff2f8e194-f28b-4508-aa07-1a6406746102_1122x1122.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!1i2p!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff2f8e194-f28b-4508-aa07-1a6406746102_1122x1122.png 424w, https://substackcdn.com/image/fetch/$s_!1i2p!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff2f8e194-f28b-4508-aa07-1a6406746102_1122x1122.png 848w, https://substackcdn.com/image/fetch/$s_!1i2p!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff2f8e194-f28b-4508-aa07-1a6406746102_1122x1122.png 1272w, https://substackcdn.com/image/fetch/$s_!1i2p!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff2f8e194-f28b-4508-aa07-1a6406746102_1122x1122.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!1i2p!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff2f8e194-f28b-4508-aa07-1a6406746102_1122x1122.png" width="1122" height="1122" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f2f8e194-f28b-4508-aa07-1a6406746102_1122x1122.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1122,&quot;width&quot;:1122,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2300883,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/199710199?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff2f8e194-f28b-4508-aa07-1a6406746102_1122x1122.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!1i2p!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff2f8e194-f28b-4508-aa07-1a6406746102_1122x1122.png 424w, https://substackcdn.com/image/fetch/$s_!1i2p!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff2f8e194-f28b-4508-aa07-1a6406746102_1122x1122.png 848w, https://substackcdn.com/image/fetch/$s_!1i2p!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff2f8e194-f28b-4508-aa07-1a6406746102_1122x1122.png 1272w, https://substackcdn.com/image/fetch/$s_!1i2p!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff2f8e194-f28b-4508-aa07-1a6406746102_1122x1122.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>Introduction</h2><p><strong>Rent a B200 for an hour</strong> and you are paying for roughly four and a half thousand trillion floating-point operations per second. Ask it to generate text from a seventy-billion-parameter model one user at a time, and for most of that hour the tensor cores do almost nothing. </p><p>The reason is not a bug, a bad kernel, or a scheduling failure. It is arithmetic. To <strong>produce a single token</strong>, the machine must read every weight in the model out of high-bandwidth memory, and reading is the slow part. </p><p>The multiply that follows the read is nearly free, and <strong>there is almost nothing to multiply</strong>, because a single decode step touches one token&#8217;s worth of activations against the entire weight matrix. <mark>You are paying for a fleet of trucks and using them to deliver one envelope per trip.</mark></p><p>Put numbers on it. A <strong>seventy-billion-parameter model</strong> in the eight-bit precision typical of modern serving is seventy gigabytes of weights. </p><p>On a B200 with eight terabytes per second of memory bandwidth, sweeping those weights once takes just under nine milliseconds, and that single sweep yields exactly one token for one user. </p><p>The tensor cores that could have executed thousands of trillions of operations in that window execute a few billion. The arithmetic intensity of <strong>single-stream decode,</strong> the ratio of compute performed to bytes moved, sits at roughly one to two floating-point operations per byte. </p><p>The hardware does not break even until that ratio reaches several hundred. The gap between those two numbers is the entire subject of this issue, because <strong>that gap is compute you have already paid</strong> for and are not using.</p><p><mark>Speculative decoding is the one technique in the</mark><strong><mark> </mark></strong><mark>inference toolbox that spends that idle compute on tokens</mark>, and, in its exact formulations, does so without altering the model&#8217;s output by a single logit. Every other lever trades something visible. </p><p><strong>Quantization trades precision</strong>. Pruning trades capacity. Distillation trades a different model entirely. Speculation, done correctly, trades nothing the user can observe; it simply reorganizes when the weight reads happen so that one read can validate several tokens at once. That is <strong>what makes it unusual</strong>, and it is why every major laboratory shipped a version of it over the last eighteen months.</p><p>And yet the operator folklore says to turn it off above a certain batch size, and the operator folklore is correct, as far as it goes. The resolution of that apparent contradiction is the thesis of this piece. </p><p>The value of speculative decoding is not a number you can quote. <mark>It is a position on a plane whose axes are</mark><strong><mark> batch size and context length,</mark></strong><mark> measured against the ridge of the roofline.</mark> In one region it cuts your cost per token roughly in half. </p><p>In the adjacent region it raises your cost per token by a fifth. The technique never changed. The regime did. The job of this issue is to draw the plane, mark the line that divides it, and show<strong> why the workload that came to dominate 2026</strong>, long-form reasoning, walked straight into the half of the plane where speculation pays.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>What speculation actually does, stated precisely</h2><p>The <strong>mechanism is worth stating exactly</strong>, because almost every confusion about the economics traces back to a loose mental model of it. </p><p>A small, cheap model, the draft, proposes a <strong>short run of candidate tokens</strong>, say four or five of them, by generating them autoregressively in the ordinary way. Because the draft is small, those proposals are fast. </p><p>The large model, the target, then performs a single forward pass that scores all of the candidate positions at once. This is the move that matters: <strong>verifying four candidate tokens</strong> costs the target essentially one weight load, the same memory sweep it would have spent producing one token on its own, because the candidates are processed in parallel across the sequence dimension rather than one step at a time.</p><p>The target then walks the candidates left to right and applies a modified rejection-sampling test at each position. </p><p>It keeps the <strong>longest prefix of candidates </strong>that agrees with what it would have sampled itself, discards the first disagreement and everything after it, and emits one additional bonus token drawn from its own distribution at the point of divergence. </p><p>So a step that began with a draft of length K returns the number of accepted candidates, call it n, plus one. <strong>If the draft proposed five tokens and the target accepted three</strong>, the step produced four tokens for the price of one memory sweep. If the target accepted all five, it produced six. If it accepted none, it produced one, the bonus token, and you paid the draft&#8217;s cost for nothing.</p><p>This is the first thing the folklore gets right and the economics must respect: the speedup is governed by <strong>how many tokens the target accepts </strong>per step, and specifically by the <em>accept length</em>, the mean size of that accepted run plus the bonus. </p><p>It is not governed by the raw acceptance rate in isolation, and it is not governed by how clever the draft sounds. A <strong>draft that is right ninety percent of the time </strong>on the next token but falls apart by the third token buys you less than a draft that is right seventy percent of the time but stays coherent for five. </p><p>The lever is the length of the run, because each run, however long, costs exactly one expensive weight load of the target.</p><h3><span>The lossleness property</span></h3><p>For the <strong>rejection-sampling formulations </strong>introduced by <em>Leviathan and colleagues in 2023</em> and independently by <em>Chen and colleagues</em> the same year, the output distribution is provably identical to standard autoregressive sampling from the target. </p><p>The <strong>modified rejection test</strong> is constructed precisely so that the accepted-token statistics match the target&#8217;s own. EAGLE preserves this exactly, as Hugging Face&#8217;s engineering writeup states plainly. The user cannot tell, from the output alone, that speculation was used.</p><p>That property deserves a caveat stated in the same breath, because vendors are <strong>not always careful about it</strong>. The losslessness holds for the exact rejection-sampling rule. </p><p>There are faster variants, relaxed acceptance, typical acceptance, and several aggressive tree-acceptance schemes, that raise the acceptance rate by loosening the test, and these do change the output distribution. They are often worth it. </p><p>But a quoted speedup that came from a relaxed acceptance rule is not the same artifact as a quoted speedup from <strong>exact rejection sampling</strong>, and an honest ledger keeps them in separate columns. When this issue later cites a four-times number, it will say which rule produced it.</p><p>The methods themselves form a clean lineage, and the direction of travel tells you what the field decided mattered. The original formulation used a <em>separate</em> draft model, <strong>a smaller member</strong> of the same family, which is simple but means maintaining and serving two models. </p><p>Medusa removed the second model by attaching several prediction heads to the target itself, each guessing a future position in parallel. EAGLE, in its <strong>first and second versions</strong>, moved the autoregression down a level, drafting in the target&#8217;s own feature space rather than in token space, which made the draft both cheaper and better aligned. </p><p><strong>EAGLE-3, presented at NeurIPS 2025</strong> and described in arXiv:2503.01840, pushed further: it fuses features from early, middle, and late layers of the target, predicts tokens directly rather than through an intermediate feature-regression step, removes a constraint that had limited how much training data helped, and uses a dynamic draft tree that expands the most promising candidates. </p><p>The endpoint of the lineage is to fold the draft into the target entirely, which is <strong>what DeepSeek&#8217;s multi-token prediction does</strong>, and which the next sections will show has economic consequences beyond mere convenience.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>The roofline and the speculation budget</h2><p>The previous issue, <em>The Split and the Seam</em>, derived the roofline for LLM serving in detail and split inference into its prefill and decode phases on exactly these grounds. </p><p>This issue assumes that derivation rather than repeating it, and reuses its house figures. The roofline says that for any kernel there is a ridge point, an <strong>arithmetic intensity above</strong> which you are limited by the chip&#8217;s compute throughput and below which you are limited by its memory bandwidth. </p><p>The ridge is simply peak compute divided by peak bandwidth. For an H100 SXM running FP8, that is one thousand nine hundred and seventy-<strong>nine teraFLOPS of dense tensor throughput </strong>against three and thirty-five hundredths terabytes per second of HBM3, which puts the ridge at five hundred and ninety-one FLOP per byte. </p><p>The H200 keeps the same compute but <strong>raises bandwidth to four and eight tenths terabytes per second</strong>, dropping the ridge to four hundred and twelve. A B200 at roughly four thousand five hundred teraFLOPS against eight terabytes per second sits near five hundred and sixty-two.</p><p>Single-stream decode operates at one to two FLOP per byte. Hold those two numbers next to each other. The operating point is two to nearly three orders of magnitude below the ridge. </p><p>That distance, expressed as a ratio, is the factor by which you could <strong>multiply the compute performed per byte</strong> <strong>moved </strong>before you would hit the memory ceiling and start paying for it in latency. Call it the <em>speculation budget</em>. On a single stream it is somewhere between three hundred and nearly six hundred times. </p><p>It is, very precisely, the ceiling on what any decode-side technique could reclaim from idle compute, and the headroom that speculative decoding draws on.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!9x7-!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb3d8c768-f443-486c-b423-9078bc10d614_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!9x7-!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb3d8c768-f443-486c-b423-9078bc10d614_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!9x7-!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb3d8c768-f443-486c-b423-9078bc10d614_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!9x7-!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb3d8c768-f443-486c-b423-9078bc10d614_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!9x7-!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb3d8c768-f443-486c-b423-9078bc10d614_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!9x7-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb3d8c768-f443-486c-b423-9078bc10d614_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b3d8c768-f443-486c-b423-9078bc10d614_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Roofline chart showing the speculation budget as the vertical gap between the single-stream decode operating point and the compute ridge for H100, H200, and B200.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Roofline chart showing the speculation budget as the vertical gap between the single-stream decode operating point and the compute ridge for H100, H200, and B200." title="Roofline chart showing the speculation budget as the vertical gap between the single-stream decode operating point and the compute ridge for H100, H200, and B200." srcset="https://substackcdn.com/image/fetch/$s_!9x7-!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb3d8c768-f443-486c-b423-9078bc10d614_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!9x7-!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb3d8c768-f443-486c-b423-9078bc10d614_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!9x7-!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb3d8c768-f443-486c-b423-9078bc10d614_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!9x7-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb3d8c768-f443-486c-b423-9078bc10d614_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Figure 1.</strong> The speculation budget is the vertical distance from the decode operating point to the roofline ridge. Single-stream decode runs two to three orders of magnitude below the point where compute becomes the limit, which is the headroom speculation cashes in. Ridge points use Issue 03 house figures: H100 SXM FP8 1,979 TFLOPS / 3.35 TB/s, H200 same compute / 4.8 TB/s, B200 ~4,500 TFLOPS / 8 TB/s.</figcaption></figure></div><p>A reader should immediately ask why, if the budget is several hundred times, speculation delivers only two or three. </p><p>The answer is that no single technique spends the whole budget, and speculation in particular spends only a sliver of it. </p><p><strong>Its yield is capped by accept length</strong>: each verification step still costs one weight load and returns at most the accepted run plus a bonus, which in practice is two to five tokens, so the multiple is bounded there no matter how much idle compute waits unused. </p><p>The draft is not free either, and its own forward passes consume part of the budget before any of it reaches the output. The rest of the headroom is what <em>batching</em> claims, the other and larger way to <strong>raise arithmetic intensity</strong>, and whatever neither mechanism reaches simply sits idle under the latency ceiling. </p><p><mark>So the budget is the size of the prize, not the size of the winnings.</mark> Speculation is the instrument that collects the part of it that batching cannot, which, as the rest of this issue argues, is exactly the part that matters when a<strong> latency SLA forbids batching in the first place</strong>.</p><p>This budget is not an accident of one chip generation. It is the accumulated result of a divergence that has run for a decade. </p><p>Across the <strong>span from V100 to B200</strong>, tensor compute throughput grew by roughly thirty-six times, while HBM bandwidth over the same generations grew by only about nine times, a gap documented in the systems literature and discussed at length in this publication&#8217;s earlier piece on the memory wall. </p><p><strong>Compute outran memory</strong> by a factor of four across those generations, and every factor of that divergence widened the speculation budget, because it pushed the ridge further above the place where decode actually runs. </p><p>The technique gets structurally more attractive with each generation of hardware that <strong>improves compute faster than bandwidth</strong>, which is to say, with each generation.</p><h3><span>The budget is real, and finite</span></h3><p>The roofline guarantees the<strong> headroom exists on a single stream</strong>. It does not guarantee the headroom survives batching, or long context. The next two sections are the story of what spends the budget down, and they reach opposite conclusions depending on which axis you move along.</p><p>The framing to carry forward is that speculation is, mechanically, a<strong> way of converting roofline headroom into tokens</strong>. When the headroom is large, the conversion is cheap and the tokens are nearly free. </p><p>When the headroom has been consumed by something else, there is nothing left to convert, and the draft&#8217;s compute becomes pure overhead. </p><p>Everything downstream is a question about <strong>how much headroom is actually available</strong> in your serving regime, and the surprising part, the part the folklore half-misses, is that the answer depends on two independent variables, not one.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Why batching is said to kill it</h2><p>Here is the story every production guide tells, and it is the right place to start because <strong>it is true within its assumptions.</strong> As you raise the batch size, packing more concurrent sequences into each forward pass, the target&#8217;s arithmetic intensity rises. </p><p>The reason is that <strong>the weights are read once per step</strong> regardless of how many sequences are in the batch, so the cost of that read amortizes across the batch. One sequence pays the full one hundred and forty gigabyte sweep for one token. </p><p><strong>Thirty-two sequences</strong> split the same sweep across thirty-two tokens. The bytes-per-token falls, the FLOP-per-byte rises, and at some batch size the target crosses its ridge and becomes compute-bound. </p><p>Past that crossing, the free headroom is gone, because the compute is now the scarce resource, and the draft model&#8217;s extra forward passes are competing for it against real work.</p><p>The crossing is commonly placed around a batch of thirty-two. <strong>Spheron&#8217;s production guide</strong> from March 2026 and<strong> E2E Networks&#8217; </strong>engineering notes both put the practical break-even in that neighborhood, with the qualification that it moves with model size, quantization, and sequence length. </p><p>Below a draft acceptance of roughly one half, the guides agree, speculation hurts at any batch, because too few candidates survive verification to cover the draft&#8217;s cost. </p><p>The operational rule that falls out is blunt and widely repeated: <strong>disable speculation when batch sizes climb past the low tens</strong>, when outputs are short, when generation is high-entropy, or when you are memory-constrained on weights to the point that the draft displaces batch capacity.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!b_nj!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F62ed2ac3-10a2-40a0-84d1-0ad7180b8940_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!b_nj!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F62ed2ac3-10a2-40a0-84d1-0ad7180b8940_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!b_nj!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F62ed2ac3-10a2-40a0-84d1-0ad7180b8940_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!b_nj!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F62ed2ac3-10a2-40a0-84d1-0ad7180b8940_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!b_nj!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F62ed2ac3-10a2-40a0-84d1-0ad7180b8940_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!b_nj!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F62ed2ac3-10a2-40a0-84d1-0ad7180b8940_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/62ed2ac3-10a2-40a0-84d1-0ad7180b8940_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Chart showing decode speedup from speculation decaying from over 3x at batch 1 toward break-even near batch 32, with EAGLE 3.1 measured points overlaid.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Chart showing decode speedup from speculation decaying from over 3x at batch 1 toward break-even near batch 32, with EAGLE 3.1 measured points overlaid." title="Chart showing decode speedup from speculation decaying from over 3x at batch 1 toward break-even near batch 32, with EAGLE 3.1 measured points overlaid." srcset="https://substackcdn.com/image/fetch/$s_!b_nj!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F62ed2ac3-10a2-40a0-84d1-0ad7180b8940_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!b_nj!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F62ed2ac3-10a2-40a0-84d1-0ad7180b8940_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!b_nj!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F62ed2ac3-10a2-40a0-84d1-0ad7180b8940_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!b_nj!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F62ed2ac3-10a2-40a0-84d1-0ad7180b8940_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Figure 1.</strong> The conventional envelope. As the batch fills, the spare compute that speculation feeds on disappears, and the speedup decays toward break-even near a batch of 32. Measured points are EAGLE 3.1 on Kimi-K2.6-NVFP4, vLLM tensor-parallel 4, GB200, SPEED-Bench, published by the vLLM team in May 2026: <strong>2.03x</strong> at concurrency 1, <strong>1.71x</strong> at 4, <strong>1.66x</strong> at 16. The break-even location and the 0.5-acceptance floor are from E2E Networks and the Spheron production guide. The envelope is illustrative; the points are measured.</figcaption></figure></div><p>The measured points anchor the shape. <strong>EAGLE 3.1, released jointly by the EAGLE, vLLM, and TorchSpec teams </strong>in May 2026 and benchmarked in the vLLM team&#8217;s own writeup running on Kimi-K2.6 in NVFP4 under vLLM with tensor parallelism of four on a GB200, delivered a per-user throughput multiple of two and three hundredths at concurrency one, one and<strong> seventy-one hundredths </strong>at concurrency four, and one and sixty-six hundredths at concurrency sixteen, on the SPEED-Bench suite. </p><p>The curve is unmistakable: the benefit is largest when the machine is emptiest, and it erodes as the batch fills. This is the empirical backbone of the folklore, and nothing in this issue disputes it on its own terms.</p><p>There is a sharper version of the same point that the practitioner Tian Pan has called the critical inversion. At<strong> low concurrency</strong> the draft runs in compute the target was wasting anyway, so it is free. </p><p>At high concurrency the draft&#8217;s forward passes contend with queued real requests for the same saturated compute, so the draft is no longer free; it is actively stealing throughput from work you could otherwise be doing. </p><p>Under that framing, speculation is fundamentally a low-concurrency latency optimization, and treating it as a <strong>throughput optimization at scale </strong>is a category error. This is good guidance. It is also, and this is the whole turn of the issue, an argument that silently assumes short context.</p><p>The <strong>amortization story is entirely about weights.</strong> It says the weight read, which dominates single-stream decode, gets cheaper per token as the batch grows. That is true. But the weight read is not the only thing decode reads from memory on every step, and the other thing it reads does not amortize across the batch at all. </p><p>The conventional wisdom is not wrong. It is two-thirds of a three-variable problem, and the missing variable is the one that 2026&#8217;s workloads turned up to eleven.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>The KV cache re-opens the budget</h2><p>Every token a transformer has already produced leaves behind a key and a value vector in every attention layer, and every future token must read all of them. <strong>That is the KV cache</strong>, and it is the second great memory cost of decode. Crucially, it behaves nothing like the weights. </p><p>The weights are shared across the batch, so their read amortizes. The KV cache is private to each sequence and grows with that sequence&#8217;s length, so its read scales with the <strong>batch size </strong>and with the context length simultaneously. </p><p>Doubling the batch does not split the KV read across more tokens; it doubles the total KV that must be read. Doubling the context length doubles it again.</p><p>The consequence is the result at the center of the <strong>MagicDec work, described in arXiv:2408.11049</strong> and in Together AI&#8217;s analysis of it. There is a critical sequence length, call it S-star, beyond which the per-step KV read dominates the per-step weight read even at large batch. </p><p>Past S-star, decode is memory-bound <em>again</em>, not because the weights are unamortized, but because the KV cache is enormous and unamortizable. The free compute the conventional wisdom said batching had consumed comes back, because batching only consumed the part of the memory bill that the weights were responsible for. The <strong>KV part grew instead of shrinking.</strong></p><p>This changes the geometry of the entire question. The compute-bound region is not the half-plane &#8220;<em>batch greater than thirty-two</em>.&#8221; It is a wedge: compute-bound requires high batch <em>and</em> short context, both at once. Move to small batch and you are <strong>memory-bound on weights</strong>. Move to long context and you are memory-bound on KV. </p><p>Only in the corner where the batch is large and the sequences are short does the target actually saturate its compute. Everywhere else, on a single stream, on long documents, on <strong>extended reasoning traces</strong>, the headroom is open and speculation has something to convert. </p><p>The short-context side of that corner has a hard edge worth naming. Because the KV cache caps how high arithmetic intensity can climb, there is a context length, roughly <strong>a thousand tokens on a B200</strong> and closer to eleven hundred on an H200, past which no batch size reaches the compute-bound ridge at all. </p><p>That ceiling is the dashed wall in the diagram below, and the compute-bound wedge lives entirely to its left. Every reasoning trace sits far to its right.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!eB9V!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97470b63-02ca-454b-add0-20561267e0be_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!eB9V!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97470b63-02ca-454b-add0-20561267e0be_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!eB9V!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97470b63-02ca-454b-add0-20561267e0be_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!eB9V!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97470b63-02ca-454b-add0-20561267e0be_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!eB9V!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97470b63-02ca-454b-add0-20561267e0be_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!eB9V!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97470b63-02ca-454b-add0-20561267e0be_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/97470b63-02ca-454b-add0-20561267e0be_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Phase diagram on axes of sequence length and batch size showing the compute-bound region as a high-batch short-context wedge and the memory-bound region everywhere else, with workload markers.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Phase diagram on axes of sequence length and batch size showing the compute-bound region as a high-batch short-context wedge and the memory-bound region everywhere else, with workload markers." title="Phase diagram on axes of sequence length and batch size showing the compute-bound region as a high-batch short-context wedge and the memory-bound region everywhere else, with workload markers." srcset="https://substackcdn.com/image/fetch/$s_!eB9V!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97470b63-02ca-454b-add0-20561267e0be_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!eB9V!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97470b63-02ca-454b-add0-20561267e0be_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!eB9V!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97470b63-02ca-454b-add0-20561267e0be_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!eB9V!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97470b63-02ca-454b-add0-20561267e0be_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Figure 2.</strong> The two-axis map, drawn from the roofline condition. Decode is compute-bound, where speculation taxes throughput, only in the small wedge at high batch and short context: it needs batch above roughly the ridge over two (around 250 for a B200) and context below the dashed wall. The wall sits where even infinite batch cannot lift arithmetic intensity to the ridge, at a sequence length of about 2C/R, near a thousand tokens on a B200 and eleven hundred on an H200. Everything else is memory-bound, where speculation pays. Interactive chat, single-user reasoning, and the long-document MagicDec regime all sit in speculation-pays territory; only short-prompt high-batch offline serving sits in the tax. With a speculative draft tree the effective batch reaches the wedge nearer nominal batch 32, which is the conventional break-even. House calculation; the KV-versus-weight crossover is a distinct curve, in Figure 4.</figcaption></figure></div><h3>Where the line actually sits</h3><p>The boundary is not abstract; you can locate it with the model&#8217;s own dimensions, and where it lands is the punchline. Take a seventy-billion-parameter model of the Llama-3-70B shape: <strong>eighty layers, grouped-query attention</strong> with eight key-value heads of head-dimension one hundred and twenty-eight. </p><p>The key-value cache that must be read per token is two vectors, key and value, times eight heads, times one hundred and <strong>twenty-eight dimensions</strong>, times eighty layers, which is one hundred and sixty-three thousand eight hundred and forty elements per token. </p><p>In a<strong> sixteen-bit KV cache</strong> that is about three tenths of a megabyte for every token already in the sequence, per sequence. The weights, in an eight-bit serving format, are seventy gigabytes, read once per step and shared across the whole batch.</p><p>The crossover, the point where the <strong>per-step key-value</strong> read equals the per-step weight read, is therefore where batch size times sequence length reaches roughly seventy gigabytes divided by three tenths of a megabyte, which is about two hundred and twenty thousand. </p><p>That locus, batch times sequence held constant, is a hyperbola: it is the line drawn in <strong>Figure 3 below</strong>, and it marks where the KV read overtakes the weight read, which is to say where adding more batch stops reducing the bytes paid per token. </p><p>At a <strong>batch of thirty-two</strong> it puts that amortization crossover near seven thousand tokens; at a batch of sixty-four, near three thousand five hundred; at a batch of one hundred and twenty-eight, near one thousand seven hundred. </p><p>An <strong>eight-bit key-value cache</strong> roughly doubles all of those. This crossover is a finer fact than the compute-bound wall of the previous figure, and the two should not be confused: the wall is the context length past which no batch reaches the ridge, while the crossover is the point at a given batch where batching has stopped buying amortization. </p><p>The reasoning workload clears both at once. Hold it against the MLPerf numbers: a mean output of three thousand eight hundred and eighty tokens, a maximum of twenty thousand, <strong>AIME traces running to twenty-three thousand. </strong></p><p>Those sequences run far past the roughly one-thousand-token compute-bound wall, so no batch reaches the ridge, and at any batch an operator can realistically run under a <strong>latency SLA </strong>they are past the amortization crossover as well. This is not a near miss. </p><p><mark>The workload that came to define 2026 lives deep in the memory-bound region by a wide margin, which is the entire reason speculation pays there.</mark> (Both lines are <strong>house order-of-magnitude calculations</strong> from the stated architecture, graded in the dossier; the precise coefficients move with KV precision, head count, ridge, and serving format, but the order of magnitude, and therefore the conclusion, is robust.)</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!HWH4!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F382010e2-3271-4faa-8148-bbf4387b0154_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!HWH4!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F382010e2-3271-4faa-8148-bbf4387b0154_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!HWH4!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F382010e2-3271-4faa-8148-bbf4387b0154_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!HWH4!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F382010e2-3271-4faa-8148-bbf4387b0154_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!HWH4!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F382010e2-3271-4faa-8148-bbf4387b0154_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!HWH4!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F382010e2-3271-4faa-8148-bbf4387b0154_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/382010e2-3271-4faa-8148-bbf4387b0154_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Log-log chart of per-step memory read versus aggregate tokens in flight, showing a flat weight-read line crossed by rising KV-read lines for dense GQA and MLA, with crossover points marked.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Log-log chart of per-step memory read versus aggregate tokens in flight, showing a flat weight-read line crossed by rising KV-read lines for dense GQA and MLA, with crossover points marked." title="Log-log chart of per-step memory read versus aggregate tokens in flight, showing a flat weight-read line crossed by rising KV-read lines for dense GQA and MLA, with crossover points marked." srcset="https://substackcdn.com/image/fetch/$s_!HWH4!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F382010e2-3271-4faa-8148-bbf4387b0154_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!HWH4!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F382010e2-3271-4faa-8148-bbf4387b0154_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!HWH4!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F382010e2-3271-4faa-8148-bbf4387b0154_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!HWH4!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F382010e2-3271-4faa-8148-bbf4387b0154_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Figure 3.</strong> The crossover, drawn. The weight read is flat because it amortizes across the batch; the KV read rises linearly with aggregate tokens in flight because it does not. For a dense-GQA 70B model the two meet near 220,000 aggregate tokens; a reasoning workload at batch 32 and 8,000 tokens of context already sits past it. Compressed attention moves the crossover, it does not remove it. House calculation; the order of magnitude is the point.</figcaption></figure></div><p>One honest qualification belongs here, because it is the first thing a careful reader will raise. The arithmetic above is for dense grouped-query attention, where the KV cache is large. </p><p>Architectures that compress the cache move the crossover to the right. <strong>DeepSeek&#8217;s Multi-head Latent Attention</strong>, by its hardware paper&#8217;s account, holds the KV cache to roughly seventy kilobytes per token, about a fifth of the dense-GQA figure, which pushes the crossover out toward a million aggregate tokens, the shallower line in the chart above. </p><p><strong>DeepSeek-V4 goes further still</strong>: its model card reports that at a one-million-token context, V4-Pro spends about ten percent of V3.2&#8217;s KV cache and twenty-seven percent of its per-token compute, with V4-Flash at seven percent and ten percent, through a compressed sparse-attention stack. </p><p>This does not rescue the throughput regime. It relocates the line, and it does so precisely in service of making very long contexts affordable, which <strong>keeps sequences long</strong>, which keeps the <strong>budget open</strong>. The compression buys context length, and context length is what holds decode in the memory-bound region. The two facts point the same way.</p><p>MagicDec turns this into a working technique with one additional move: the <strong>draft itself must be light on KV,</strong> not just light on weights, or it reintroduces the very bottleneck it is trying to relieve. </p><p>With a draft that uses a fixed sparse or short-window KV footprint, MagicDec reports up to roughly two times on both throughput and latency together in the<strong> large-batch long-context regime</strong> on eight A100s, a regime where the conventional wisdom predicts speculation should be dead. </p><p>The reported draft-to-target memory ratio for a Llama-3.1-70B target with an <strong>eight-billion-parameter draft</strong> stays near four tenths and, importantly, stays constant as the batch grows, because the draft&#8217;s KV is bounded by design while the target&#8217;s KV grows. </p><p>That constancy is what keeps the draft cheap exactly where the conventional analysis assumed it would become expensive.</p><p>The corrected physics is therefore a single sentence with three clauses. <strong>Small batch is memory-bound </strong>because weights dominate. Long context is memory-bound because the KV read dominates. </p><p>Compute-bound is only the high-batch corner below the thousand-token wall, and that corner is smaller than the folklore implies. The &#8220;<em>disable above batch thirty-two</em>&#8221; rule is not wrong; it is a short-context rule wearing the costume of a general one. </p><p>And the moment your workload develops long sequences, whether from large documents or from long generations,<strong> the rule inverts</strong>, and speculation comes back to life precisely where you had been told to switch it off.</p><div><hr></div><h2>What actually shows on the ledger</h2><p><strong>Physics is not the bill</strong>. To get from the roofline to dollars, you have to know how the operator is allowed to set the batch size, and that is a question about service-level agreements, not about chips. </p><p>There are <strong>two serving regimes</strong>, and they read the same technique with opposite signs.</p><p>In <em>throughput-maximizing</em> service, the operator is free to batch all the way to the compute-bound point, because the <strong>only objective is cost per token</strong> and the way to minimize it is to amortize the weight read across as many sequences as possible. In that regime the machine is, by construction, saturated. </p><p><strong>There is no idle compute</strong>. Speculation adds the draft&#8217;s forward passes to a chip that has nothing spare to run them in, so the cost per token rises. This is the regime the folklore is built for, and in it the folklore&#8217;s advice is exactly right.</p><p>In <em>latency-capped</em> service, the operator may <em>not</em> batch to the compute-bound point, because there is a ceiling on how long each token may take, and raising the batch raises per-token latency. The operator batches only until the <strong>latency SLA binds</strong>, and then stops, often well short of saturation. </p><p>The machine therefore runs with idle compute by design, not by accident, because the SLA forbids filling it. That idle compute is the speculation budget, and <strong>speculation converts it into tokens</strong>, cutting the cost per token. <mark>Same technique, opposite sign, and the only thing that changed was whether a latency ceiling capped the batch.</mark></p><p><mark>These two signs are not asserted; they fall out of a one-line cost model.</mark> Cost per token is the <strong>rental rate of the GPU</strong> divided by the tokens it delivers each second, so anything that multiplies throughput divides cost by the same factor. </p><p>In the latency-capped regime the wasted verification compute is free, because the chip sat idle under the SLA anyway, so throughput scales with the accept length discounted only by the draft&#8217;s own overhead: an accept length of about two and a half against a draft overhead near a fifth gives a <strong>throughput multiple close to two</strong>, which is the cut of roughly half the chart shows. </p><p>In the throughput-maximized regime the chip has no spare compute, so the draft&#8217;s wasted work bites directly. If the draft proposes three tokens and <strong>two and a half clear verification on average</strong>, the target spends three positions of compute to yield two and a half tokens, a throughput multiple near five sixths, which is the cost rise of about a fifth the chart shows. </p><p>The same two numbers, an accept length near two and a half and a draft length near three, generate both bars, and they are the same numbers behind the one-and-a-half to two-and-a-half times production speedups.</p><p>Push the draft length above the accept length and the throughput-regime penalty grows, which is precisely why <strong>over-drafting is the classic way to lose money</strong> on speculation in a saturated cluster.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!T4i-!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F832809e8-ea41-4db9-be51-1a2a5a28cc53_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!T4i-!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F832809e8-ea41-4db9-be51-1a2a5a28cc53_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!T4i-!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F832809e8-ea41-4db9-be51-1a2a5a28cc53_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!T4i-!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F832809e8-ea41-4db9-be51-1a2a5a28cc53_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!T4i-!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F832809e8-ea41-4db9-be51-1a2a5a28cc53_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!T4i-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F832809e8-ea41-4db9-be51-1a2a5a28cc53_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/832809e8-ea41-4db9-be51-1a2a5a28cc53_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Bar chart comparing cost per million tokens for autoregressive versus speculative decoding under latency-capped serving and throughput-maximizing serving.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Bar chart comparing cost per million tokens for autoregressive versus speculative decoding under latency-capped serving and throughput-maximizing serving." title="Bar chart comparing cost per million tokens for autoregressive versus speculative decoding under latency-capped serving and throughput-maximizing serving." srcset="https://substackcdn.com/image/fetch/$s_!T4i-!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F832809e8-ea41-4db9-be51-1a2a5a28cc53_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!T4i-!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F832809e8-ea41-4db9-be51-1a2a5a28cc53_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!T4i-!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F832809e8-ea41-4db9-be51-1a2a5a28cc53_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!T4i-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F832809e8-ea41-4db9-be51-1a2a5a28cc53_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Figure 4.</strong> The sign of the ledger is set by the regime. Under a latency SLA that caps the batch below saturation, speculation converts idle compute into tokens and cuts cost per million output tokens (here roughly <strong>-48%</strong>). In throughput-maximizing service batched to the compute-bound point, there is no idle compute for the draft to use, and the draft&#8217;s overhead raises cost (here roughly <strong>+19%</strong>). Both magnitudes follow from the cost model in the text (accept length near 2.5, draft length near 3) and are consistent with the production speedups in Figure 8; the absolute dollar levels still move with rate, model, and quantization, anchored here to 2026 neocloud figures (Spheron, getdeploying).</figcaption></figure></div><p>The reason latency-capped service is the common case in 2026, rather than a corner case, is written directly into the benchmark SLAs. </p><p>MLPerf Inference v5.1, published by MLCommons in September 2025, sets for its DeepSeek-R1 reasoning workload a time-to-first-token ninety-ninth-percentile threshold of two seconds and a<strong> time-per-output-token ninety-ninth-percentile threshold of eighty milliseconds</strong>, against a mean input of around eight hundred tokens and a mean output of three thousand eight hundred and eighty, with a maximum output of twenty thousand, the highest the benchmark has ever specified. </p><p>An eighty-millisecond ceiling on per-token latency, applied to sequences thousands of tokens long, caps the batch far below the compute-bound point, because<strong> long sequences mean large KV reads</strong> and large KV reads mean each added unit of batch costs latency you do not have. </p><p>The SLA traps the GPU in the <strong>memory-bound regime</strong>. <mark>The trap is the opportunity: a memory-bound GPU has idle compute, and idle compute is what speculation eats.</mark></p><h3><span>The benchmark concedes the point</span></h3><p>The argument stops being a thesis and becomes a rule when the benchmark authority writes it into the rules. In March 2026, <strong>MLPerf Inference v6.0 added an interactive reasoning scenario</strong> for DeepSeek-R1 with the ceiling pulled tighter still, a 1.5-second TTFT and a 15-millisecond TPOT at the ninety-ninth percentile. </p><p>To make that scenario achievable at all, MLCommons mandates speculative decoding for it: implementations must run the official <strong>DeepSeek-R1 MTP head with EAGLE-style decoding. </strong></p><p>The independent body that defines how inference is measured decided that, past a certain latency target on reasoning traffic, speculation is not an optional optimization but a requirement of entry.</p><p>The dollar figures that frame the chart are anchored to 2026 market rates and published per-token costs, kept deliberately conservative. Neocloud H100 capacity runs around two dollars an hour, with <strong>Spheron listing two dollars and one cent</strong>; B200 on-demand sits in the five-to-six-dollar range across getdeploying and aimultiple&#8217;s trackers. </p><p>Published serving costs land near forty-two cents per million tokens on a B200 and<strong> forty-seven cents on an H100 PCIe.</strong> The point of the chart is not to nail a single deployment&#8217;s economics to the cent, which would be dishonest given how much rate, model, and quantization move the number. </p><p>The point is the asymmetry: the same forty-something cents per million can become a discount or a surcharge depending solely on which side of the saturation line your SLA puts you.</p><p>The most honest evidence for this whole framing comes, unexpectedly, from the vendor with the most incentive to claim an unqualified win. <strong>DeepSeek&#8217;s hardware paper, arXiv:2505.09343</strong>, states that its multi-token prediction module can slightly hurt raw throughput while significantly improving end-to-end generation latency. </p><p>Read that again in the context of the two regimes. DeepSeek is reporting, in print, that in a <strong>throughput accounting MTP </strong>can cost a little, and in a latency accounting it helps a lot, and that they ship it because latency is the product. </p><p>They add a second-order point that sharpens it further: MTP raises the effective batch size, which in their <strong>mixture-of-experts architecture </strong>increases expert-parallel arithmetic intensity, partially offsetting the throughput cost. </p><p>A company could have quoted the latency win alone and called it a free lunch. Instead they <strong>documented the tradeoff in both directions</strong>, which is precisely the shape of the real ledger this issue is arguing for.</p><p>When the vendor with the strongest incentive to claim a pure throughput win instead publishes that the technique &#8220;<em>slightly hurts throughput while significantly improving latency</em>,&#8221; that is not a weakness in the technique. </p><p>It is the ledger showing its true two-sided shape, in the vendor&#8217;s own numbers.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/p/decode-is-memory-bound-speculation?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/p/decode-is-memory-bound-speculation?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><div><hr></div><h2>Why reasoning moved the bill into decode</h2><p>Cost per token has a denominator, and the denominator is dominated by decode steps, because <strong>prefill is a single parallel pass</strong> over the prompt while decode is a long sequence of memory-bound steps, one per output token. </p><p>Anything that multiplies the number of output tokens multiplies the share of the bill that lives in decode, which is exactly the share speculation can attack. This is <strong>why 2026 is a different economic environment</strong> for speculative decoding than 2023 was, even though the technique is largely the same. The traffic changed.</p><p>Reasoning models emit output on a different scale entirely. A conventional chat reply is a few hundred tokens. A reasoning trace runs to thousands, and the trend within the model generation has been sharply upward: <strong>BentoML&#8217;s deployment guide </strong>notes that DeepSeek-R1-0528 nearly doubled its reasoning length over the prior R1, from around twelve thousand to around twenty-three thousand tokens on a single hard math question. </p><p>MLPerf&#8217;s DeepSeek-R1 workload puts the mean output at three thousand eight hundred and eighty and the <strong>maximum at twenty thousand</strong>. Agentic systems then chain many such traces into a single user-visible task, so the effective output length per task can be larger still. </p><p>The bill, which used to be split between a substantial prefill and a modest decode, has tilted hard toward decode.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!hi0v!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F340c3881-d019-43c5-acf6-e5a668d8f89a_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!hi0v!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F340c3881-d019-43c5-acf6-e5a668d8f89a_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!hi0v!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F340c3881-d019-43c5-acf6-e5a668d8f89a_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!hi0v!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F340c3881-d019-43c5-acf6-e5a668d8f89a_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!hi0v!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F340c3881-d019-43c5-acf6-e5a668d8f89a_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!hi0v!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F340c3881-d019-43c5-acf6-e5a668d8f89a_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/340c3881-d019-43c5-acf6-e5a668d8f89a_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Horizontal bar chart of output tokens per request growing from a few hundred for chat to tens of thousands for reasoning workloads.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Horizontal bar chart of output tokens per request growing from a few hundred for chat to tens of thousands for reasoning workloads." title="Horizontal bar chart of output tokens per request growing from a few hundred for chat to tens of thousands for reasoning workloads." srcset="https://substackcdn.com/image/fetch/$s_!hi0v!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F340c3881-d019-43c5-acf6-e5a668d8f89a_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!hi0v!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F340c3881-d019-43c5-acf6-e5a668d8f89a_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!hi0v!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F340c3881-d019-43c5-acf6-e5a668d8f89a_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!hi0v!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F340c3881-d019-43c5-acf6-e5a668d8f89a_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Figure 5.</strong> Reasoning moved the bill into decode. A reasoning request emits one to two orders of magnitude more output tokens than a chat reply, and every one of them is a memory-bound decode step that must clear the latency SLA. MLPerf Inference v5.1 DeepSeek-R1 (MLCommons, September 2025) reports a mean output of 3,880 tokens and a maximum of 20,000; R1 and R1-0528 AIME usage runs from roughly 12,000 to 23,000 tokens per question (BentoML guide). The chat baseline is a round-number reference.</figcaption></figure></div><p><strong>Two facts</strong> about reasoning traffic place it squarely in the regime where speculation pays. The first is the one just shown: it is <strong>decode-heavy</strong>, so the part of the bill speculation can lower is the dominant part. The second is subtler and follows from the previous sections. </p><p>Long traces mean long sequences in flight, which means<strong> large KV reads</strong>, which means decode is memory-bound even when the operator manages to batch, both because the <strong>latency SLA</strong> caps the batch and because the KV cost re-opens the budget past S-star. </p><p>The two mechanisms reinforce each other. The workload is in the memory-bound regime by virtue of its output length, and it is held there by <strong>virtue of its latency SLA</strong>. There is idle compute on the machine for both reasons at once, and speculation is the technique that turns idle compute into tokens.</p><p>The architectural direction of travel keeps the budget open rather than closing it. <strong>DeepSeek-V4&#8217;s sparse-attention work,</strong> with V4-Pro reportedly using about twenty-seven percent of the FLOPs and ten percent of the KV of V3.2 at a one-million-token context through DeepSeek Sparse Attention, is explicitly aimed at making very long contexts affordable.</p><p>Cheaper long context means more long context, which means more memory-bound decode, which means a wider speculation budget, not a narrower one. The <strong>hardware trend</strong> widens the budget by improving compute faster than bandwidth; the model trend widens it by pushing context length up. Both vectors point the same way.</p><p>And this is where accept length, the lever from the first section, cashes in <strong>directly against the denominator</strong>. The decode discount is, to first order, the accept length: a method that lands two and a half accepted tokens per step is doing roughly two and a half times the decode work per expensive weight load. </p><p>The measured accept lengths of the shipped 2026 methods, two and fifty-five hundredths for <strong>DeepSeek-V3.2&#8217;s MTP,</strong> two and seventy-six hundredths for GLM-5&#8217;s shared-MTP design, and four and seven tenths for EAGLE-3 on coding and reasoning workloads, are therefore not abstract quality scores. </p><p>They are multipliers on the largest line item in the reasoning-era bill.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!U_FB!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F410e3d7e-698b-4812-8847-93486f5c88d5_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!U_FB!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F410e3d7e-698b-4812-8847-93486f5c88d5_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!U_FB!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F410e3d7e-698b-4812-8847-93486f5c88d5_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!U_FB!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F410e3d7e-698b-4812-8847-93486f5c88d5_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!U_FB!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F410e3d7e-698b-4812-8847-93486f5c88d5_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!U_FB!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F410e3d7e-698b-4812-8847-93486f5c88d5_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/410e3d7e-698b-4812-8847-93486f5c88d5_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Bar chart of mean accepted tokens per verification step across vanilla drafting, DeepSeek-V3.2 MTP, GLM-5 shared-MTP, and EAGLE-3.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Bar chart of mean accepted tokens per verification step across vanilla drafting, DeepSeek-V3.2 MTP, GLM-5 shared-MTP, and EAGLE-3." title="Bar chart of mean accepted tokens per verification step across vanilla drafting, DeepSeek-V3.2 MTP, GLM-5 shared-MTP, and EAGLE-3." srcset="https://substackcdn.com/image/fetch/$s_!U_FB!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F410e3d7e-698b-4812-8847-93486f5c88d5_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!U_FB!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F410e3d7e-698b-4812-8847-93486f5c88d5_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!U_FB!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F410e3d7e-698b-4812-8847-93486f5c88d5_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!U_FB!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F410e3d7e-698b-4812-8847-93486f5c88d5_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Figure 6.</strong> The lever is accept length, not raw acceptance. Tokens produced per step equal the accepted run plus one bonus token, and each step costs exactly one target weight load, so accept length is the decode discount. Reported figures: DeepSeek-V3.2 MTP <strong>2.55</strong> and GLM-5 shared-MTP <strong>2.76</strong> from the GLM-5 technical report (arXiv:2602.15763); EAGLE-3 <strong>4.5 to 5.0</strong> on coding and reasoning from E2E Networks. The vanilla-draft figure is a representative reference.</figcaption></figure></div><div><hr></div><h2>Where speculation loses</h2><p>A technique that only ever helps does not need an issue written about when to use it. Speculation has real failure modes, and an analysis that buries them is worth less than one that lists them, so here is the full debit column, without hedging.</p><p><strong>The throughput regime.</strong> In saturated high-batch short-context serving, the offline-batch corner of the phase diagram, <mark>speculation is a tax and should be disabled.</mark> The compute is fully employed, the draft has nothing free to run in, and its forward passes displace real work. The Spheron and Tian Pan guidance is correct here without qualification. If your job is to push the maximum number of short completions through a fleet of GPUs at minimum cost per token, speculation is the wrong lever.</p><p><strong>Acceptance collapse.</strong> <mark>Below roughly one-half acceptance, speculation hurts at any batch</mark>, because too few candidates survive to cover the draft&#8217;s cost. Acceptance is not a constant; it falls with high sampling temperature, with out-of-distribution inputs the draft was never trained on, and with the kind of high-entropy generation where the next token is genuinely uncertain. A draft trained against one target distribution and then serving a drifted or fine-tuned target degrades silently, the acceptance rate sliding without any error being raised. Monitoring accept length in production is not optional; it is the only way to notice that your discount has quietly become a surcharge.</p><p><strong>VRAM pressure.</strong> The draft model and its KV cache occupy memory you could otherwise spend on a larger batch or a longer context. A Llama-3.3-70B target in FP8 alongside a one-billion-parameter draft consumes roughly seventy-five to seventy-eight gigabytes on an eighty-gigabyte H100, per Spheron&#8217;s figures, leaving very little headroom. On a memory-constrained deployment, the draft can cost you more in lost batch capacity than it returns in accept length, and that tradeoff has to be measured, not assumed.</p><p><strong>No help for time-to-first-token.</strong> Speculation accelerates decode, and only decode. It does nothing for prefill, which means it does nothing for time-to-first-token. Under the MLPerf two-second TTFT ceiling, that is a separate problem requiring separate techniques, which is precisely the prefill-decode disaggregation argument of Issue 03. Speculation and prefill optimization are complementary, not substitutes, and a serving stack that needs both will not get the first from the second.</p><p><strong>Draft maintenance.</strong> A separate draft model is a second training, evaluation, and deployment surface that must be kept aligned as the target evolves. Every target update risks degrading a draft that was tuned against the previous version. EAGLE-style heads and built-in MTP layers reduce this by coupling the draft to the target&#8217;s own features or parameters, but they do not eliminate the obligation to retrain and revalidate. GLM-5&#8217;s choice to share parameters across three MTP layers, described in arXiv:2602.15763, is partly an answer to exactly this maintenance cost: fewer independent parameters to train and keep aligned.</p><p><strong>The losslessness caveat, restated.</strong> The provable equivalence to standard sampling holds for the exact rejection-sampling rule. Relaxed acceptance, typical acceptance, and aggressive tree-acceptance schemes raise throughput by changing the output distribution. They are frequently worth it, but a four-times figure obtained under a relaxed rule is not interchangeable with a four-times figure under exact sampling, and a serving team quoting a speedup owes itself, and its users, clarity about which rule produced it.</p><p><strong>Draft-length tuning.</strong> The number of tokens the draft proposes per step, often written gamma, is a workload-dependent knob with a real optimum. Set it too long and the draft burns compute generating candidates that will be rejected; set it too short and you leave accept length on the table. The optimum moves with acceptance rate and with batch size, so a value tuned on one workload can be wrong on another, and dynamic schemes that adjust it per request exist precisely because no single value is right everywhere.</p><p><strong>The lab-versus-production gap.</strong> The EAGLE-3 paper reports speedups of up to six and a half times, but those are temperature-zero academic measurements on Vicuna-13B, Llama-3.1-8B, and Llama-3.3-70B. Production reports cluster instead around two to three times: LMSYS and Vertex describe two-to-three-times figures for EAGLE-3 on SGLang, E2E Networks reports two and three tenths on Llama-3.1-8B at a batch of four, and a Gemma-4 EAGLE3 draft head is documented at one and seventy-two hundredths at batch one on conversational traffic. The gap between the lab number and the production number is itself one of the most important facts in this space, because <mark>quoting the former as if it were the latter is the single most common honesty failure in vendor material on speculative decoding.</mark></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!IE6Z!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1992dc1-fdc3-410b-af49-2787e536fe46_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!IE6Z!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1992dc1-fdc3-410b-af49-2787e536fe46_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!IE6Z!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1992dc1-fdc3-410b-af49-2787e536fe46_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!IE6Z!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1992dc1-fdc3-410b-af49-2787e536fe46_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!IE6Z!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1992dc1-fdc3-410b-af49-2787e536fe46_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!IE6Z!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1992dc1-fdc3-410b-af49-2787e536fe46_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f1992dc1-fdc3-410b-af49-2787e536fe46_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Horizontal bar chart contrasting a 6.5x lab speedup against a cluster of production speedups between 1.66x and 2.3x.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Horizontal bar chart contrasting a 6.5x lab speedup against a cluster of production speedups between 1.66x and 2.3x." title="Horizontal bar chart contrasting a 6.5x lab speedup against a cluster of production speedups between 1.66x and 2.3x." srcset="https://substackcdn.com/image/fetch/$s_!IE6Z!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1992dc1-fdc3-410b-af49-2787e536fe46_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!IE6Z!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1992dc1-fdc3-410b-af49-2787e536fe46_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!IE6Z!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1992dc1-fdc3-410b-af49-2787e536fe46_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!IE6Z!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1992dc1-fdc3-410b-af49-2787e536fe46_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Figure 7.</strong> The headline academic figure (6.5x, EAGLE-3 paper, temperature 0) sits far above the production cluster, which lands between roughly 1.66x and 2.3x across the EAGLE 3.1 vLLM benchmark, DeepSeek&#8217;s vendor-reported MTP TPS, a Gemma-4 EAGLE3 draft head, and E2E Networks. The bars use different targets and conditions and are not strictly comparable; they are shown to convey the range, and the distance between the gold bar and the teal cluster is the point.</figcaption></figure></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>A decision on the plane</h2><p>The verdict is not a yes or a no. It is a lookup. Given a workload&#8217;s batch size, its context length, its <strong>latency SLA</strong>, and its measured acceptance, the phase diagram tells you which regime you are in, and the regime tells you the sign of the ledger. The table below collapses the analysis into that lookup.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!wYrw!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37830286-9e44-4ed2-a815-f14199ca909b_1710x1168.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!wYrw!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37830286-9e44-4ed2-a815-f14199ca909b_1710x1168.png 424w, https://substackcdn.com/image/fetch/$s_!wYrw!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37830286-9e44-4ed2-a815-f14199ca909b_1710x1168.png 848w, https://substackcdn.com/image/fetch/$s_!wYrw!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37830286-9e44-4ed2-a815-f14199ca909b_1710x1168.png 1272w, https://substackcdn.com/image/fetch/$s_!wYrw!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37830286-9e44-4ed2-a815-f14199ca909b_1710x1168.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!wYrw!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37830286-9e44-4ed2-a815-f14199ca909b_1710x1168.png" width="1456" height="995" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/37830286-9e44-4ed2-a815-f14199ca909b_1710x1168.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:995,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:248170,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/199710199?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37830286-9e44-4ed2-a815-f14199ca909b_1710x1168.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!wYrw!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37830286-9e44-4ed2-a815-f14199ca909b_1710x1168.png 424w, https://substackcdn.com/image/fetch/$s_!wYrw!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37830286-9e44-4ed2-a815-f14199ca909b_1710x1168.png 848w, https://substackcdn.com/image/fetch/$s_!wYrw!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37830286-9e44-4ed2-a815-f14199ca909b_1710x1168.png 1272w, https://substackcdn.com/image/fetch/$s_!wYrw!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37830286-9e44-4ed2-a815-f14199ca909b_1710x1168.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>State the thesis cleanly now that the machinery is in place. <strong>Speculative decoding is the only inference technique</strong> that converts decode&#8217;s idle compute into tokens losslessly, and the sign of its effect on your bill is set entirely by whether that compute was actually idle. </p><p>Whether it was idle is a question about the roofline, and <strong>where you sit on the roofline</strong> is a question about batch size and context length, the two axes of the phase diagram. </p><p>There is no universal answer because the inputs are not universal. <mark>There is a correct answer for every point on the plane, and the table is how you read it.</mark></p><p>The reason this matters more now than it did three years ago is that the median workload moved. In <strong>2023 the prototypical request</strong> was a short chat completion at modest context, which lives near the compute-bound corner once you batch it, where speculation is at best neutral. </p><p>In 2026 the prototypical high-value request is a long reasoning trace under a <strong>tight per-token latency SLA</strong>, which lives deep in the memory-bound region for two independent reasons, its output length and its SLA, and which is therefore exactly where speculation pays. <mark>The technique did not move toward the workload. The workload moved toward the technique.</mark></p><p>That migration is why the shipping decisions of the major laboratories converged. <strong>DeepSeek built multi-token prediction into V3</strong> and carried it through V3.2, documenting the latency win and the throughput cost honestly. </p><p>GLM-5 shipped a shared-parameter three-layer MTP design with a measured accept length near two and three-quarters. <strong>NVIDIA&#8217;s NeMo RL work applied EAGLE-3 to reinforcement-learning rollouts</strong> and reported a one-and-eight-tenths-times generation speedup at the eight-billion scale, with validation accuracy on AIME-2024 evolving identically under autoregressive and speculative decoding, a clean confirmation that the lossless guarantee holds across training. </p><p><strong>EAGLE-3 landed across vLLM, SGLang, and TensorRT-LLM</strong>, the three serving stacks that matter. These are not independent fashions. They are the same bet, placed by everyone who looked at the same plane and saw that reasoning traffic had walked into the half where the ledger reads in your favor.</p><p><em>The technique never changed. The regime did. Speculation lowers your cost per token exactly where batching cannot help you, and reasoning is the workload that made that region the center of the map.</em></p><div><hr></div><h2>Confidence tiers and external-audit read</h2><p>Every load-bearing claim in this issue is scored below against a four-tier confidence scale, with its source named inline. </p><p>The <strong>scale is applied as an external auditor would apply it</strong>, crediting primary and measured sources, discounting derived and illustrative ones, and flagging the weakest links explicitly rather than hiding them in the prose.</p><p><strong>Tier A</strong> <em>primary or measured: peer-reviewed papers, vendor hardware disclosures, MLPerf-published SLAs and benchmark statistics.</em><br><strong>Tier B</strong> <em>secondary, with method: vendor or practitioner reports that state their configuration and measurement conditions.</em><br><strong>Tier C</strong> <em>derived or stylized: house figures and curves built from the cited physics, presented as illustrative renderings, not measurements.</em><br><strong>Tier D</strong> <em>illustrative or round-number: reference values chosen for scale, not claimed as measured.</em></p><p><strong>A</strong></p><p><strong>Exact rejection-sampling speculative decoding is output-distribution lossless.</strong></p><p>Leviathan et al. 2023; Chen et al. 2023; EAGLE losslessness per Hugging Face engineering writeup. Provable equivalence to standard sampling under the exact rule.</p><p><strong>A</strong></p><p><strong>EAGLE-3 mechanism: training-time test, multi-layer feature fusion, direct token prediction, dynamic draft tree.</strong></p><p>EAGLE-3 paper, NeurIPS 2025, arXiv:2503.01840. Headline lab speedups up to 6.5x at temperature 0 on Vicuna-13B, Llama-3.1-8B, Llama-3.3-70B (explicitly a best-case lab figure; production lands far lower, see Figure 8).</p><p><strong>A</strong></p><p><strong>DeepSeek MTP slightly hurts throughput while significantly improving end-to-end latency; raises effective batch and expert-parallel intensity.</strong></p><p>DeepSeek hardware paper, arXiv:2505.09343. The two-sided tradeoff is stated in the vendor&#8217;s own text, which is the strongest single piece of evidence in this issue.</p><p><strong>A</strong></p><p><strong>MLPerf DeepSeek-R1 SLAs: TTFT 99p 2s, TPOT 99p 80ms; mean input 800, mean output 3,880, max output 20,000.</strong></p><p>MLPerf Inference v5.1, MLCommons, September 2025. These SLAs are the anchor for why latency-capped serving traps the GPU in the memory-bound regime.</p><p><strong>A</strong></p><p><strong>MLPerf Inference v6.0 added a DeepSeek-R1 interactive scenario (TTFT 1.5s, TPOT 15ms) and mandates speculative decoding (official MTP head, EAGLE-style) to meet it.</strong></p><p>MLCommons, March 2026. The benchmark authority requiring speculation for tight-latency reasoning is the strongest external corroboration of the thesis.</p><p><strong>A</strong></p><p><strong>NVIDIA NeMo RL: EAGLE-3 gives ~1.8x rollout generation speedup at 8B; AIME-2024 accuracy identical under autoregressive and speculative decoding throughout training.</strong></p><p>NVIDIA NeMo RL research, May 2026. Independent empirical confirmation that the lossless guarantee holds in practice across training.</p><p><strong>A</strong></p><p><strong>Critical-sequence-length result: past S*, decode is memory-bound even at large batch via KV read; KV-light draft delivers up to ~2x throughput and latency.</strong></p><p>MagicDec, arXiv:2408.11049; Together AI long-context analysis. Draft-to-target memory ratio ~0.4, constant at large batch, for Llama-3.1-70B with 8B draft.</p><p><strong>A</strong></p><p><strong>Roofline ridge points: H100 SXM FP8 591, H200 412, B200 ~562 FLOP/byte; compute outgrew bandwidth ~36x vs ~9x V100 to B200.</strong></p><p>Carried from Issue 03, The Split and the Seam, and Issue on the memory wall, The Wall and the Stack. Hardware specifications and systems-literature divergence figures.</p><p><strong>A</strong></p><p><strong>DeepSeek-V4 sparse attention: V4-Pro 27% FLOPs and 10% KV of V3.2 at 1M context; V4-Flash 10% and 7%.</strong></p><p>Official DeepSeek-V4-Pro / V4-Flash model cards (Compressed Sparse Attention + Heavily Compressed Attention). Verified figures. Direction-of-travel evidence that long context is getting cheaper, widening the budget.</p><p><strong>A</strong></p><p><strong>EAGLE 3.1 per-user throughput: 2.03x at concurrency 1, 1.71x at 4, 1.66x at 16.</strong></p><p>Primary: vLLM team blog, May 2026 (EAGLE / vLLM / TorchSpec joint release). Kimi-K2.6-NVFP4, tensor-parallel 4, GB200, non-disagg, SPEED-Bench coding. Verified against the primary engineering writeup.</p><div><hr></div><p><strong>B</strong></p><p><strong>Practical break-even near batch 32; below ~0.5 acceptance speculation hurts at any batch.</strong></p><p>Spheron production guide, March 2026; E2E Networks. Practitioner guidance with stated qualifications on model, quantization, and sequence length.</p><p><strong>B</strong></p><p><strong>The lab-versus-production speedup comparison in Figure 8 (6.5x lab vs a 1.66 to 2.3x production cluster).</strong></p><p>Compiled from verified primary and vendor sources (EAGLE-3 paper, EAGLE 3.1 vLLM blog, DeepSeek hardware paper, Gemma-4 EAGLE3 card, E2E Networks). Bars use different targets and conditions and are not strictly comparable; shown to convey range, not to rank.</p><p><strong>B</strong></p><p><strong>Production speedups cluster at 2 to 3x; E2E reports 2.3x on Llama-3.1-8B at batch 4, accept length 4.5 to 5.0.</strong></p><p>LMSYS and Vertex on SGLang; E2E Networks. Multiple independent practitioner reports converging on the same range.</p><p><strong>B</strong></p><p><strong>Accept lengths: DeepSeek-V3.2 MTP 2.55, GLM-5 shared-MTP 2.76; GLM-5 shares parameters across 3 MTP layers.</strong></p><p>GLM-5 technical report, arXiv:2602.15763. EAGLE-3 accept length 4.5 to 5.0 from E2E Networks.</p><p><strong>B</strong></p><p><strong>2026 GPU rates and per-token costs: H100 ~$2/hr, B200 ~$5 to $6/hr on-demand; ~$0.42/M (B200), ~$0.47/M (H100 PCIe).</strong></p><p>Spheron ($2.01/hr H100), getdeploying, aimultiple. Market trackers; rates move with provider, commitment, and region.</p><p><strong>B</strong></p><p><strong>Reasoning length growth: R1-0528 nearly doubled to ~23K tokens per AIME question vs ~12K for prior R1.</strong></p><p>BentoML DeepSeek deployment guide, 2026. R1 token pricing ~$0.55/M in, ~$2.19/M out; output billed higher and dominates cost.</p><p><strong>C</strong></p><p><strong>The KV-versus-weight amortization crossover (Figure 4): batch x sequence near 220,000 for a 70B model, where the KV read overtakes the weight read (~7,000 tokens at batch 32, ~1,750 at batch 128).</strong></p><p>House order-of-magnitude calculation from Llama-3-70B architecture (80 layers, 8 KV heads, head-dim 128) and the weight-versus-KV read balance, drawn in Figure 4. The coefficient moves with KV precision and serving format; the order of magnitude, and the conclusion that reasoning traces sit past the crossover, is robust.</p><p><strong>C</strong></p><p><strong>The compute-bound boundary and ~1,000-token wall in Figure 3.</strong></p><p>Derived, not stylized: the boundary is the roofline condition AI = 2B/(1+B*S/C) exceeding the ridge, with the wall at S = 2C/R (C ~ 220,000; R = 562 FP8 for B200). A house calculation with standard simplifying assumptions (GEMM-dominated FLOPs, dense GQA KV); exact coordinates shift with precision and ridge, the structure does not.</p><p><strong>C</strong></p><p><strong>The -48% / +19% cost magnitudes in Figure 5.</strong></p><p>Derived from the cost model in the text (accept length ~2.5, draft length ~3, draft overhead ~0.2) and consistent with the production speedups in Figure 8. Absolute dollar levels still vary with rate, model, and quantization; the regime-dependent sign and rough magnitude are the load-bearing claim.</p><p><strong>C</strong></p><p><strong>The shape of the speedup-decay envelope in Figure 2.</strong></p><p>House curve illustrating the conventional decay toward break-even; the measured EAGLE 3.1 points on it are Tier A and the break-even location is sourced, but the connecting envelope is schematic, not fitted.</p><p><strong>D</strong></p><p><strong>The ~300-token chat-reply baseline and the vanilla-draft 2.1 accept-length reference.</strong></p><p>Round-number references chosen to set scale against the measured reasoning and method figures, not claimed as measured values.</p><h3>External-audit simulation</h3><p>Audited as a whole, the issue rests on a spine of Tier A primary sources, and every load-bearing empirical claim in it was checked against its primary source: the <strong>EAGLE-3 paper </strong>(arXiv:2503.01840),<strong> MagicDec </strong>(arXiv:2408.11049), the <strong>DeepSeek hardware disclosure</strong> (arXiv:2505.09343), <strong>the GLM-5 report </strong>(arXiv:2602.15763), the <strong>MLPerf v5.1 and v6.0 </strong>specifications, the <strong>DeepSeek-V4 model cards</strong>, and the EAGLE 3.1 vLLM <strong>release </strong>each confirmed the figures attributed to them. </p><p>The<strong> central argument</strong>, that the sign of the ledger is set by serving regime and that reasoning traffic sits in the favorable regime, follows from those sources rather than from the house figures, which is the property an audit most wants to see. </p><p>Two pieces of evidence are doing disproportionate work and both survive scrutiny: DeepSeek&#8217;s own statement that <strong>MTP slightly hurts throughput </strong>while significantly <strong>improving latency</strong>, a direct vendor quote, and MLPerf v6.0&#8217;s decision to mandate speculative decoding for its tight-latency reasoning scenario, which is the measuring authority writing the thesis into the rules.</p><p>The weakest links are named rather than hidden, and after this revision they are narrow. <strong>Boundary and Crossover</strong> are now derivations rather than stylizations: the first is the roofline condition with the wall at twice the bytes-ratio over the ridge, the second is the<strong> weight-versus-KV balance</strong>, both house calculations carried out with standard simplifying assumptions (<em>GEMM-dominated compute, dense grouped-query KV</em>) whose exact coordinates move with precision and ridge while the structure holds. </p><p>The cost magnitudes are likewise derived from an explicit model, an accept length near two and a half and a draft length near three, and <strong>cross-checked against the production speedups</strong>; what remains genuinely soft there is the absolute dollar level, which varies too much with rate, model, and quantization to pin down. </p><p>The <strong>decay envelope is a schematic shape</strong>, though the measured points on it and its break-even location are sourced. The chat and vanilla-draft baselines are round numbers for scale. A handful of practitioner figures, the break-even batch, the <strong>VRAM envelope</strong>, and the production speedup band, come from individual engineering guides rather than independently reproduced benchmarks, and are tiered B accordingly. </p><p>None of these elements carries the conclusion: a reader who accepts only the Tier A claims, and works the two house calculations independently, arrives at the same verdict table.</p><p><strong>Overall confidence in the thesis is high</strong>, because the thesis is a statement about regimes and signs that the primary sources support directly, and it is deliberately not a statement that speculation yields any specific universal multiple, which the evidence would not support. </p><p>The <strong>quantitative illustrations</strong> are held at lower confidence by design, and labeled as such, so that the argument does not borrow credibility it has not earned. A reader who accepts only the Tier A claims still arrives at the same verdict table; the lower tiers furnish the texture, not the conclusion.</p><div><hr></div><h2>Sources and useful informations</h2><p>Primary and secondary sources for the load-bearing claims, with arXiv identifiers and venues where applicable. The text above attributes each source at its point of use; this is the consolidated record. </p><p>Papers appear first, then benchmark specifications, vendor and practitioner writeups, and model cards.</p><ol><li><p><span>Y. Leviathan, M. Kalman, and Y. Matias. </span><em><span>Fast Inference from Transformers via Speculative Decoding.</span></em><span> International Conference on Machine Learning (ICML), 2023. arXiv:2211.17192.</span></p></li><li><p><span>C. Chen, S. Borgeaud, G. Irving, J.-B. Lespiau, L. Sifre, and J. Jumper. </span><em><span>Accelerating Large Language Model Decoding with Speculative Sampling.</span></em><span> arXiv:2302.01318, 2023.</span></p></li><li><p><span>T. Cai, Y. Li, Z. Geng, H. Peng, J. D. Lee, D. Chen, and T. Dao. </span><em><span>Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads.</span></em><span> International Conference on Machine Learning (ICML), 2024. arXiv:2401.10774.</span></p></li><li><p><span>Y. Li, F. Wei, C. Zhang, and H. Zhang. </span><em><span>EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty.</span></em><span> International Conference on Machine Learning (ICML), 2024. arXiv:2401.15077.</span></p></li><li><p><span>Y. Li, F. Wei, C. Zhang, and H. Zhang. </span><em><span>EAGLE-2: Faster Inference of Language Models with Dynamic Draft Trees.</span></em><span> Conference on Empirical Methods in Natural Language Processing (EMNLP), 2024. arXiv:2406.16858.</span></p></li><li><p><span>Y. Li, F. Wei, C. Zhang, and H. Zhang. </span><em><span>EAGLE-3: Scaling up Inference Acceleration of Large Language Models via Training-Time Test.</span></em><span> Conference on Neural Information Processing Systems (NeurIPS), 2025. arXiv:2503.01840.</span></p></li><li><p><span>R. Sadhukhan, J. Chen, Z. Chen, V. Tiwari, A. May, T. Chen, and B. Chen. </span><em><span>MagicDec: Breaking the Latency-Throughput Tradeoff for Long Context Generation with Speculative Decoding.</span></em><span> arXiv:2408.11049, 2024.</span></p></li><li><p><span>DeepSeek-AI. </span><em><span>DeepSeek-V3 Technical Report.</span></em><span> arXiv:2412.19437, 2024.</span></p></li><li><p><span>DeepSeek-AI. </span><em><span>Insights into DeepSeek-V3: Scaling Challenges and Reflections on Hardware for AI Architectures.</span></em><span> International Symposium on Computer Architecture (ISCA), 2025. arXiv:2505.09343.</span></p></li><li><p><span>DeepSeek-AI. </span><em><span>DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.</span></em><span> arXiv:2501.12948, 2025.</span></p></li><li><p><span>Z.ai (Zhipu AI). </span><em><span>GLM-5 Technical Report.</span></em><span> arXiv:2602.15763, 2026.</span></p></li><li><p><span>MLCommons. </span><em><span>MLPerf Inference: Datacenter, v5.1 (DeepSeek-R1 reasoning workload).</span></em><span> Benchmark rules and results, 2025.</span></p></li><li><p><span>MLCommons. </span><em><span>MLPerf Inference: Datacenter, v6.0 (DeepSeek-R1 Interactive scenario, mandated speculative decoding).</span></em><span> Benchmark rules, 2026.</span></p></li><li><p><span>EAGLE Team, vLLM Team, and TorchSpec. </span><em><span>EAGLE 3.1: release and SPEED-Bench results on Kimi-K2.6.</span></em><span> vLLM Blog, May 2026.</span></p></li><li><p><span>NVIDIA. </span><em><span>Speculative Decoding for Reinforcement-Learning Rollouts in NeMo RL.</span></em><span> NVIDIA Developer technical writeup, 2026.</span></p></li><li><p><span>Together AI. </span><em><span>Speculative decoding for high-throughput long-context inference (analysis of MagicDec).</span></em><span> Together AI Blog, 2024.</span></p></li><li><p><span>BentoML. </span><em><span>The Complete Guide to DeepSeek Models: V3, R1, V4 and Beyond.</span></em><span> BentoML Blog, 2026.</span></p></li><li><p><span>Spheron Network. </span><em><span>Speculative decoding in production: a practitioner&#8217;s guide.</span></em><span> Engineering guide, 2026.</span></p></li><li><p><span>E2E Networks. </span><em><span>Speculative decoding performance on Llama-3.1 serving.</span></em><span> Engineering notes, 2026.</span></p></li><li><p><span>DeepSeek-AI. </span><em><span>DeepSeek-R1-0528.</span></em><span> Model card, Hugging Face, 2025.</span></p></li><li><p><span>DeepSeek-AI. </span><em><span>DeepSeek-V4-Pro and DeepSeek-V4-Flash.</span></em><span> Model cards, Hugging Face, 2026.</span></p><p></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share The Software Frontier&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share The Software Frontier</span></a></p><p></p></li></ol>]]></content:encoded></item></channel></rss>