<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[The Software Frontier]]></title><description><![CDATA[Where abstraction ends. Essays on GPU execution, kernel internals, and distributed systems at scale.]]></description><link>https://www.thesoftwarefrontier.com</link><image><url>https://substackcdn.com/image/fetch/$s_!SAY7!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F54550d86-2756-4131-8818-956604f6749d_608x608.png</url><title>The Software Frontier</title><link>https://www.thesoftwarefrontier.com</link></image><generator>Substack</generator><lastBuildDate>Sat, 22 Aug 2026 07:39:22 GMT</lastBuildDate><atom:link href="https://www.thesoftwarefrontier.com/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Lorenzo Bradanini]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[softwarefrontier@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[softwarefrontier@substack.com]]></itunes:email><itunes:name><![CDATA[Lorenzo Bradanini]]></itunes:name></itunes:owner><itunes:author><![CDATA[Lorenzo Bradanini]]></itunes:author><googleplay:owner><![CDATA[softwarefrontier@substack.com]]></googleplay:owner><googleplay:email><![CDATA[softwarefrontier@substack.com]]></googleplay:email><googleplay:author><![CDATA[Lorenzo Bradanini]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[DeepSeek V4-Flash: The Cost of Deciding What to Read]]></title><description><![CDATA[284 billion parameters rebuilt from the published constants, a million-token cache in 3.37 GiB, and the arithmetic showing that 4/5 of the attention budget is spent choosing what to attend to.]]></description><link>https://www.thesoftwarefrontier.com/p/deepseek-v4-flash-the-cost-of-deciding</link><guid isPermaLink="false">https://www.thesoftwarefrontier.com/p/deepseek-v4-flash-the-cost-of-deciding</guid><dc:creator><![CDATA[Lorenzo Bradanini]]></dc:creator><pubDate>Mon, 10 Aug 2026 06:45:30 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!i_aj!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F58602cdd-1158-491a-827b-650143b365c9_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!i_aj!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F58602cdd-1158-491a-827b-650143b365c9_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!i_aj!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F58602cdd-1158-491a-827b-650143b365c9_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!i_aj!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F58602cdd-1158-491a-827b-650143b365c9_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!i_aj!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F58602cdd-1158-491a-827b-650143b365c9_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!i_aj!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F58602cdd-1158-491a-827b-650143b365c9_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!i_aj!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F58602cdd-1158-491a-827b-650143b365c9_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/58602cdd-1158-491a-827b-650143b365c9_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2550191,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/209988515?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F58602cdd-1158-491a-827b-650143b365c9_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!i_aj!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F58602cdd-1158-491a-827b-650143b365c9_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!i_aj!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F58602cdd-1158-491a-827b-650143b365c9_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!i_aj!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F58602cdd-1158-491a-827b-650143b365c9_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!i_aj!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F58602cdd-1158-491a-827b-650143b365c9_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>I&#8217;ll show you seven numbers this piece derives</h2><div class="callout-block" data-callout="true"><p><strong>284.202 B</strong> ; reconstructed backbone against a published 284 B</p><p><strong>39.3 %</strong> ; of activated parameters are attention, which is 1.74 percent of the weights</p><p><strong>3.37 GiB</strong> ; of KV for a 1,048,576-token sequence, 1.96 percent of a GQA-8 baseline</p><p><strong>79 %</strong> ; of attention FLOPs at 1M are the selector, not the attention</p><p><strong>4096</strong> ; FLOPs per byte Flash needs from its interconnect, 1.5x Pro&#8217;s demand</p><p><strong>5,504</strong> ; tokens of recompute restore a million-token prefix, a 191x saving</p><p><strong>$2.92</strong> ; per GPU-hour is what 1M context pays; owners clear it, renters do not</p></div><p>I didn&#8217;t set out to write about <strong>DeepSeek V4-Flash,</strong> but i just wanted to check a precise number. The technical report gives the model&#8217;s constants in a paragraph on page 25 and its headline size, 284B total and 13B activated, on page 4, and I wanted to know whether the two agreed before I trusted anything else in the document. </p><p>They agree to seven hundredths of one percent, but only once you notice that the headline quietly leaves out the multi-token prediction module and that<strong> Heavily Compressed Attention</strong> carries half the key-value projections that Compressed Sparse Attention does. </p><p>Doing the raw math anyway shows something that the report never states: attention is 1.7 percent of this model&#8217;s weights and 39 percent of the weights it touches per token.</p><p>That is the kind of fact that changes what you build. It means the sparsity everyone is now discussing on reddit, lives <strong>entirely in the expert bank</strong>, that the attention stack is as dense as it has ever been, and that there is a hard floor under how cheap a model in this family can get. It also doesnt appear in any of the two dozen write-ups of this model I read before starting.</p><p>So this is more a <strong>reconstruction </strong>rather than a summary. Every architectural number below was just rebuilt from the published constants and checked against something DeepSeek or the <strong>vLLM team</strong> published independently. </p><p>Where the reconstruction disagrees with what is currently written about V4-Flash on the open web, I say so. Where it disagrees with the report, I say that too.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>What shipped, and what the number on the card means</h2><p>As of now, there are<strong> two DeepSeek-V4-Flash releases</strong> and they are the same weights. The preview landed on 24 April 2026 with the 58 page technical report and DeepSeek-V4-Pro. </p><p>The official release, tagged 0731, landed on 31 July 2026 as a public beta of the API. DeepSeek&#8217;s changelog states that <strong><span>DeepSeek-V4-Flash-0731</span></strong> has the same model structure and the same size as the preview and that only the post-training was rerun. I&#8217;ve not yet seen a new architecture, or a new parameter count; even the price hasn&#8217;t changed.</p><p>That is convenient for a piece like this, because every architectural fact in the April report still describes the model you can call today. It also means the <strong>agent scores DeepSeek published with the 0731 release</strong>, Terminal Bench 2.1 at 82.7 and DeepSWE at 54.4, are alignment results rather than efficiency results. </p><p>They were measured with DeepSeek&#8217;s own harness in what the changelog calls minimal mode at the max reasoning tier, temperature 1.0, top_p 0.95. The harness has not shipped. </p><p>Agent scores move by ten points on harness changes. Just treat them as a <strong>statement of intent.</strong></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!10Gx!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcc9d9f9d-0547-44ef-b260-8398c5da5a95_1600x832.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!10Gx!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcc9d9f9d-0547-44ef-b260-8398c5da5a95_1600x832.png 424w, https://substackcdn.com/image/fetch/$s_!10Gx!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcc9d9f9d-0547-44ef-b260-8398c5da5a95_1600x832.png 848w, https://substackcdn.com/image/fetch/$s_!10Gx!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcc9d9f9d-0547-44ef-b260-8398c5da5a95_1600x832.png 1272w, https://substackcdn.com/image/fetch/$s_!10Gx!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcc9d9f9d-0547-44ef-b260-8398c5da5a95_1600x832.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!10Gx!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcc9d9f9d-0547-44ef-b260-8398c5da5a95_1600x832.png" width="1456" height="757" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/cc9d9f9d-0547-44ef-b260-8398c5da5a95_1600x832.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:757,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;DeepSeek-V4-Flash, published configuration. Report section 4.2.1, stated directly.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="DeepSeek-V4-Flash, published configuration. Report section 4.2.1, stated directly." title="DeepSeek-V4-Flash, published configuration. Report section 4.2.1, stated directly." srcset="https://substackcdn.com/image/fetch/$s_!10Gx!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcc9d9f9d-0547-44ef-b260-8398c5da5a95_1600x832.png 424w, https://substackcdn.com/image/fetch/$s_!10Gx!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcc9d9f9d-0547-44ef-b260-8398c5da5a95_1600x832.png 848w, https://substackcdn.com/image/fetch/$s_!10Gx!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcc9d9f9d-0547-44ef-b260-8398c5da5a95_1600x832.png 1272w, https://substackcdn.com/image/fetch/$s_!10Gx!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcc9d9f9d-0547-44ef-b260-8398c5da5a95_1600x832.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>DeepSeek-V4-Flash, published configuration. Report section 4.2.1, stated directly.</em></p><p>I want to focus on one correction, before we go further and deeper. A widely cited <strong>third-party guide</strong> states that DeepSeek publishes the layer arrangement for Pro but not for Flash, and advises readers to treat Flash&#8217;s layer count as unconfirmed. </p><p>That is<strong> totally wrong. </strong>The report gives Flash 43 layers and hidden dimension 4096 in the same paragraph that gives it 284B parameters. Pro gets 61 layers and 7168. </p><p>The arrangements differ in one way that matters: Pro&#8217;s first two layers are HCA, Flash&#8217;s first two are pure sliding window attention with no compression at all.</p><div><hr></div><h2>Rebuilding the model from its constants</h2><p>Every expert is a <strong>SwiGLU block</strong>, so three matrices of 4096 by 2048, which is precisely 25.17M parameters. With 256 routed experts plus one shared expert in all 43 blocks, the expert bank alone is 278.108B. </p><p>That is 97.8 percent of the model and it takes one line of arithmetic. Everything interesting is (in my opinion) in the <strong>remaining 2.2 percent.</strong></p><p>The attention shapes are not what you would guess from a normal transformer, and <strong>CSA</strong> and <strong>HCA </strong>are not the same size. CSA computes two independent key-value streams, so four projection matrices from equations 9 and 10 of the report. </p><p>HCA computes one, so two matrices, from equations 20 and 21. Getting this wrong is the difference between a reconstruction that lands and one that does not.</p><pre><code>CSA layer
  W^aKV, W^bKV, W^aZ, W^bZ    4 x (4096 x 512)   =  8.389 M   two overlapped KV streams
  B^a, B^b                     2 x (4 x 512)      =  0.004 M   learnable positional bias
  W^DQ                         4096 x 1024        =  4.194 M   query down-projection
  W^UQ                         1024 x (512 x 64)  = 33.554 M   query up-projection
  W^IUQ                        1024 x (128 x 64)  =  8.389 M   indexer query up-projection
  W^w                          4096 x 64          =  0.262 M   per-head indexer gate
  grouped output, 8 groups     8 x (4096 x 1024)  = 33.554 M
  final output projection      8192 x 4096        = 33.554 M
  attention sink logits        64                 =  0.000 M
                                                    ---------
                                                    121.901 M

HCA layer   one KV stream, no indexer, bias over m&#8217; = 128         109.117 M
SWA layer   uncompressed KV, no compression weights, no indexer   106.954 M</code></pre><p>Manifold-constrained hyper-connections add 393K per residual junction. <strong>Router gates add 45M</strong> across the model. Embeddings, using the DeepSeek-V3 tokenizer with a handful of added context-construction tokens, add 530M on each side. </p><p>The report says the vocabulary remains 128K, which I read as the 129,280 entries V3 used.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!sWAa!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc88d9bcd-3f58-4e01-957b-c952ded8e67b_1600x619.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!sWAa!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc88d9bcd-3f58-4e01-957b-c952ded8e67b_1600x619.png 424w, https://substackcdn.com/image/fetch/$s_!sWAa!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc88d9bcd-3f58-4e01-957b-c952ded8e67b_1600x619.png 848w, https://substackcdn.com/image/fetch/$s_!sWAa!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc88d9bcd-3f58-4e01-957b-c952ded8e67b_1600x619.png 1272w, https://substackcdn.com/image/fetch/$s_!sWAa!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc88d9bcd-3f58-4e01-957b-c952ded8e67b_1600x619.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!sWAa!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc88d9bcd-3f58-4e01-957b-c952ded8e67b_1600x619.png" width="1456" height="563" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c88d9bcd-3f58-4e01-957b-c952ded8e67b_1600x619.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:563,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Where 284 billion parameters actually sit. Reconstructed from the published constants; the residual against the published figure is 0.07 percent.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Where 284 billion parameters actually sit. Reconstructed from the published constants; the residual against the published figure is 0.07 percent." title="Where 284 billion parameters actually sit. Reconstructed from the published constants; the residual against the published figure is 0.07 percent." srcset="https://substackcdn.com/image/fetch/$s_!sWAa!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc88d9bcd-3f58-4e01-957b-c952ded8e67b_1600x619.png 424w, https://substackcdn.com/image/fetch/$s_!sWAa!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc88d9bcd-3f58-4e01-957b-c952ded8e67b_1600x619.png 848w, https://substackcdn.com/image/fetch/$s_!sWAa!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc88d9bcd-3f58-4e01-957b-c952ded8e67b_1600x619.png 1272w, https://substackcdn.com/image/fetch/$s_!sWAa!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc88d9bcd-3f58-4e01-957b-c952ded8e67b_1600x619.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>Where 284 billion parameters actually sit. Reconstructed from the published constants; the residual against the published figure is 0.07 percent.</em></p><p>284.202B against a published 284B,<em> a 0.07 percent residual</em> on a reconstruction with no free parameters. </p><p>The <strong>MTP head is real</strong>, it is 6.6B, vLLM will use it for speculative decoding, and it is not in the number on the model card. DeepSeek did the same with V3, whose Hugging Face repository showed 685B against a stated 671B.</p><p>There is a second confirmation of the reconstruction hiding in an unlikely place. In the <strong>determinism section</strong>, discussing why they cannot use split-k for <em>one particular GEMM</em>, the report mentions in passing that &#8220;<em>mHC involves a matrix multiplication with an output dimension of only 24.</em>&#8221; </p><p>With n_hc = 4, the dynamic parameterisation generates A in R^4, B in R^{4x4} and C in R^4 from one flattened input. Four plus sixteen plus four is twenty-four. </p><p>The three mappings are produced by a single GEMM, and a throwaway sentence in a section about floating-point associativity confirms the shape.</p><p>The <strong>activated count</strong> follows. Seven experts fire per token, six routed plus the shared one, which is 7.575B. Everything else in the forward pass is dense: all 4.956B of attention, all of mHC, all of the router gates. </p><p>That is <strong>12.610B</strong> excluding embeddings and 13.140B counting the output head, against a published 13B.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!zTVu!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe778cefd-9d87-4c3f-9dff-dad878a474c0_1600x806.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!zTVu!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe778cefd-9d87-4c3f-9dff-dad878a474c0_1600x806.png 424w, https://substackcdn.com/image/fetch/$s_!zTVu!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe778cefd-9d87-4c3f-9dff-dad878a474c0_1600x806.png 848w, https://substackcdn.com/image/fetch/$s_!zTVu!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe778cefd-9d87-4c3f-9dff-dad878a474c0_1600x806.png 1272w, https://substackcdn.com/image/fetch/$s_!zTVu!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe778cefd-9d87-4c3f-9dff-dad878a474c0_1600x806.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!zTVu!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe778cefd-9d87-4c3f-9dff-dad878a474c0_1600x806.png" width="1456" height="733" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e778cefd-9d87-4c3f-9dff-dad878a474c0_1600x806.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:733,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Attention weights are a rounding error in the checkpoint and a plurality of the arithmetic at decode. Reconstructed from the constants in section 4.2.1.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Attention weights are a rounding error in the checkpoint and a plurality of the arithmetic at decode. Reconstructed from the constants in section 4.2.1." title="Attention weights are a rounding error in the checkpoint and a plurality of the arithmetic at decode. Reconstructed from the constants in section 4.2.1." srcset="https://substackcdn.com/image/fetch/$s_!zTVu!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe778cefd-9d87-4c3f-9dff-dad878a474c0_1600x806.png 424w, https://substackcdn.com/image/fetch/$s_!zTVu!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe778cefd-9d87-4c3f-9dff-dad878a474c0_1600x806.png 848w, https://substackcdn.com/image/fetch/$s_!zTVu!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe778cefd-9d87-4c3f-9dff-dad878a474c0_1600x806.png 1272w, https://substackcdn.com/image/fetch/$s_!zTVu!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe778cefd-9d87-4c3f-9dff-dad878a474c0_1600x806.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>Attention weights are a rounding error in the checkpoint and a plurality of the arithmetic at decode. Reconstructed from the constants in section 4.2.1.</em></p><p>Attention is 1.74 percent of the weight bank and 39.3 percent of what gets touched per token. That&#8217;s most load-bearing fact about serving this model and it<strong> appears nowhere</strong> in the report, because the report presents attention as the thing being optimised away rather than as a fixed cost that survives the optimisation.</p><blockquote><p><em>The sparsity everyone is discussing lives entirely in the expert bank. The attention stack is as dense as it has ever been.</em></p></blockquote><p>The consequence is a floor. <em>Push MoE sparsity as far as you like</em>, drop from six routed experts to four, go <strong>from 256 experts to 512 </strong>at the same activation count, and the activated parameter count will not fall below roughly 5.0B, because the attention stack is dense and it is 4.956B of it.</p><p>Every decode step of every request reads all of it. At 12.6B activated you are already 40 percent of the way to that floor. A hypothetical<strong> V4-Flash-Nano </strong>with two routed experts would be a 10.1B-activated model, not a 4B one.</p><p>It also reframes what the hybrid attention is for. It is not there to make attention cheap in absolute terms. It is there to stop attention&#8217;s <em>state</em> from growing without bound, which is a <strong>memory problem </strong>rather than a FLOPs problem, and the FLOPs bill arrives anyway. </p><p>We will come back later to how large it gets.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Compressed Sparse Attention</h2><p>CSA is <strong>five separate ideas stacked</strong>, and it is often described as one. Taken apart, each of them is simple, and two of them are the sort of thing you only invent after you have been beaten by the alternative.</p><h3>Key and value are the same vector</h3><p>V4 stores one cached entry per compressed position, not two. Here, the key and the value are literally the <strong>same tensor</strong>, used in both roles by a multi-query attention. That is an immediate halving of the cache before any compression happens at all.</p><p>It should not work. Attention output is normally translation invariant: rotate the query by <strong>R(t)</strong> and the key by <strong>R(s)</strong> and the score depends on <em>R(t-s), which is relative.</em> Values carry no rotation, so the output carries none. </p><p>Share the key and the value, and the value inherits the key&#8217;s rotation, the output acquires an <strong>absolute position</strong> through R(s), and shifting the whole sequence changes the answer.</p><p>The fix is in section 2.3.3 and it is one line: apply RoPE with position minus i to the last 64 dimensions of each head&#8217;s output. Because R is orthogonal and <strong>R(t)<sup>-1</sup>R(s) = R(s-t)</strong>, rotating the output backwards restores relative positioning, and the contribution of each cached entry ends up depending on its distance from the query again. </p><p>The vLLM team&#8217;s write-up derives the same result from the other direction and calls it <strong>inverse RoPE</strong>. Their implementation fuses it into the FP8 quantisation ahead of the output projection, worth two to three times over doing the two separately.</p><p><strong>Two times the cache</strong>, bought with one elementwise kernel. This is the cheapest trick in the model and it is the one that has attracted the least attention.</p><h3>The compressor has two streams and they overlap</h3><p>CSA does not average four tokens into one. It computes two independent projections of the <strong>hidden state, C<sup>a</sup> and C<sup>b</sup></strong>, each with its own learned compression weights <strong>Z<sup>a</sup> and Z<sup>b</sup></strong>, then takes a softmax across the concatenation of 2m weights and forms a weighted sum over both streams. </p><p>Compressed entry i draws C<sup>a</sup> from positions [mi, m(i+1)] and C<sup>b</sup> from positions [m(i-1), mi].</p><p>With m = 4 that makes every compressed entry a data-dependent weighted sum of eight consecutive tokens taken at a stride of four. <strong>vLLM names this attention type <span>c4a</span> </strong>and documents it as a weighted sum of 8 uncompressed tokens with a stride of 4, which is exactly equations 11 and 12.</p><p>The overlap is the point. A hard boundary every four tokens cuts arbitrary spans in arbitrary places and the model is blind across the seam. Overlapping means every token appears in <strong>two compressed entries </strong>under different weights, while the sequence still shrinks by exactly four because the stride is four. The receptive field of an entry is eight; the compression ratio is four.</p><p>The compression weights are per dimension. The softmax runs over <strong>2m elements independently</strong> for each of the 512 channels, so each channel picks its own mixture of the eight tokens. </p><p>This is a learned per-channel pooling rather than a summarisation step, and describing it as &#8220;remembering the paragraph&#8217;s key point&#8221; undersells it by a wide margin.</p><h3>The Lightning Indexer</h3><p>After compression a one-million-token context still holds 262,144 entries in every CSA layer. Attending densely to a quarter of a million entries is not obviously better than attending densely to a million, so <strong>CSA runs DeepSeek Sparse Attention</strong> over the compressed stream: a cheap scoring pass picks the top 512, and the real attention runs only over those.</p><p>The scorer builds 64 low-rank query heads of dimension 128 from the same compressed query latent the main attention uses, computes a rectified dot product against a separately compressed indexer key for each block, and sums across heads with a learned per-head gate:</p><pre><code>c^Q_t   = h_t &#183; W^DQ                                   shared with the main attention queries
q^I_t   = c^Q_t &#183; W^IUQ                                64 indexer heads of dimension 128
w^I_t   = h_t &#183; W^w                                    one gate per indexer head, from the hidden state
I(t,s)  = sum_h  w^I(t,h) &#183; ReLU( q^I(t,h) &#183; K^IComp_s )</code></pre><p>Two choices there are worth stopping on. <strong>ReLU rather than softmax leaves the score unnormalised</strong>, so a head can contribute nothing to a block rather than merely down-weighting it, and heads can veto. And the gate w is produced from the hidden state by a 4096 by 64 matrix, so the model decides per token which of its 64 scoring heads to trust. </p><p>That is a router in everything but name, sitting in front of the attention, and it is trained end to end with it.</p><p><strong>Flash&#8217;s top-k is 512</strong>. Pro&#8217;s is 1024. V3.2&#8217;s was 2048. The report is direct about the reason: a smaller top-k greatly improves efficiency on short and medium texts, which is where the traffic is.</p><h3>Grouped output projection</h3><p>Sixty-four heads at head dimension 512 produce <strong>32,768 values per token. </strong>A conventional output projection would be 32,768 by 4096, which is 134M parameters per layer, more than everything else in the attention block put together. </p><p><strong>V4 splits the heads into g = 8 groups of eight</strong>, projects each group&#8217;s 4096 values to 1024, concatenates the eight results into 8192, and projects that to 4096. Total 67.1M, half the naive cost, and the bottleneck at 8192 acts as a constraint on how freely heads can mix.</p><p>Pro uses <strong>g = 16 at the same d_g = 1024</strong>, so a 16,384-wide concatenation from 128 heads. The report also applies RMSNorm per head on the queries and on the single head of the compressed KV entries just before the core attention, which is the same numerical hygiene MLA needed, for the same reason, and which turns out to matter for the optimiser.</p><h3>Attention sink, and heads that abstain</h3><p>Both CSA and HCA carry a set of learnable sink logits, one per head. For head h, <strong>Exp(z&#8217;<sub>h</sub>) is added to the denominator </strong>of the softmax and to nothing else:</p><pre><code>s(h,i,j) = Exp(z(h,i,j)) / ( sum_k Exp(z(h,i,k)) + Exp(z&#8217;_h) )</code></pre><p>The report&#8217;s description of what this buys is unusually blunt. It allows each query head to make its total attention score not equal to one, &#8220;<em>and even to be near 0.</em>&#8221; A head with a large sink logit contributes almost nothing to the output no matter what is in the context. </p><p>Given that <strong>this model asks 64 heads per layer </strong>to attend to a top-512 selection out of a quarter of a million compressed blocks, giving heads a principled way to decline is not decoration. </p><p>It is what stops a head from being forced to spend its mass on whichever blocks the indexer happened to hand it.</p><h3>The sliding window is not an optimisation</h3><p>Every compressed layer also keeps <em>128 uncompressed tokens in a window</em>, concatenated with the selected compressed entries before the softmax. This is a correctness requirement and the reason is causality.</p><p>A <strong>compressed entry i in <span>c128a</span></strong> summarises positions 128i through 128(i+1)-1. A query at position t may only use information derived from positions at or before t, so it cannot use entry i unless 128(i+1)-1 is at or before t. </p><p>A query sitting anywhere inside the current block therefore has no compressed entry it is permitted to read. Without the window, tokens 1 through 127 of every block would attend to no local context whatsoever. </p><p>The window covers the distance between the query and the most recent legal compression boundary, and 128 is exactly m&#8217; for that reason.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Heavily Compressed Attention</h2><p>HCA is the same compressor at m&#8217; = 128, with two differences that both simplify it. There is<strong> one key-value stream</strong> instead of two, so no overlap and no seam handling. And there is no indexer, so no top-k. It attends densely over everything it holds.</p><p>The arithmetic explains the second choice. A <strong>one-million-token context under 128x compression</strong> yields 8,192 entries, and eight thousand keys is an ordinary attention problem, shorter than most models&#8217; native context. Sparsity would buy nothing; vLLM implements it as a sparse-attention call with top-k set to 8192, a selection that selects everything, purely so one kernel serves both paths.</p><p>Dropping the overlap is defensible because HCA is not responsible for local detail. The sliding window handles that and the interleaved CSA layers handle the middle range. <strong>HCA&#8217;s job is coarse global memory</strong>, and boundary precision on a 128-token block does not matter much when the question is what the document was about.</p><p>What the pair produces is a genuinely two-rate memory. In Flash&#8217;s stack, two sliding-window layers are followed by 41 alternating layers, giving<strong> 21 CSA and 20 HCA</strong>. </p><p>Every token&#8217;s representation passes through 21 layers that can retrieve 512 four-token spans from anywhere in the history and 20 layers that see all 8,192 coarse summaries at once. Neither works alone. </p><p>Sparse selection over fine blocks has recall problems on diffuse queries; dense attention over coarse blocks cannot resolve a specific line of code.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>The router, and the three ways V4 changed it</h2><p>Every write-up of this model spends its length on the attention and treats the mixture of experts as inherited furniture. It is not. <strong>Section 2.1 makes four changes to routing</strong>, and one of them removes a whole class of layer that every DeepSeek model before this one had.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!kL1g!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83871c94-9dd8-471d-91b4-d1b1b9cae9af_1600x450.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!kL1g!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83871c94-9dd8-471d-91b4-d1b1b9cae9af_1600x450.png 424w, https://substackcdn.com/image/fetch/$s_!kL1g!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83871c94-9dd8-471d-91b4-d1b1b9cae9af_1600x450.png 848w, https://substackcdn.com/image/fetch/$s_!kL1g!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83871c94-9dd8-471d-91b4-d1b1b9cae9af_1600x450.png 1272w, https://substackcdn.com/image/fetch/$s_!kL1g!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83871c94-9dd8-471d-91b4-d1b1b9cae9af_1600x450.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!kL1g!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83871c94-9dd8-471d-91b4-d1b1b9cae9af_1600x450.png" width="1456" height="409" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/83871c94-9dd8-471d-91b4-d1b1b9cae9af_1600x450.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:409,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Forty-three layers, two independent schedules. Reconstructed from report sections 2.1 and 4.2.1; the CSA and HCA assignment within the interleave is inferred, and the split is either 21 to 20 or the reverse.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Forty-three layers, two independent schedules. Reconstructed from report sections 2.1 and 4.2.1; the CSA and HCA assignment within the interleave is inferred, and the split is either 21 to 20 or the reverse." title="Forty-three layers, two independent schedules. Reconstructed from report sections 2.1 and 4.2.1; the CSA and HCA assignment within the interleave is inferred, and the split is either 21 to 20 or the reverse." srcset="https://substackcdn.com/image/fetch/$s_!kL1g!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83871c94-9dd8-471d-91b4-d1b1b9cae9af_1600x450.png 424w, https://substackcdn.com/image/fetch/$s_!kL1g!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83871c94-9dd8-471d-91b4-d1b1b9cae9af_1600x450.png 848w, https://substackcdn.com/image/fetch/$s_!kL1g!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83871c94-9dd8-471d-91b4-d1b1b9cae9af_1600x450.png 1272w, https://substackcdn.com/image/fetch/$s_!kL1g!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83871c94-9dd8-471d-91b4-d1b1b9cae9af_1600x450.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>Forty-three layers, two independent schedules. Reconstructed from report sections 2.1 and 4.2.1; the CSA and HCA assignment within the interleave is inferred, and the split is either 21 to 20 or the reverse.</em></p><p><strong>There are no dense FFN layers.</strong> V3 kept ordinary dense feed-forward blocks in its first three Transformer layers, on the standard argument that early layers do generic work and routing them wastes capacity. </p><p>V4 replaces those with MoE layers that use <em>hash routing</em>: the target experts for a token are fixed by a hash of its token ID, following Roller&#8217;s Hash Layers work from 2021. The Hugging Face implementation makes the mechanism concrete. </p><p><strong>Routing type</strong> is set per layer through <span>mlp_layer_types</span>, and a hash layer resolves its experts through a frozen <span>tid2eid</span> lookup shipped inside the checkpoint.</p><p>The details that make this more than a curiosity is that only the <em>selection</em> is static. The learned gate still produces the per-expert scores that weight the chosen experts. So a <strong>hash layer </strong>is not an un-routed layer; it is a layer where the router has been told which experts to consider and gets to decide how much to trust each one. </p><p>That converts the hardest part of early-layer routing, an <strong>unstable argmax over 256 options</strong> while the model knows nothing, into a fixed assignment with a learnable mixture on top.</p><p>It also <strong>does something useful</strong> for the infrastructure. A hash of the token ID is known before the forward pass reaches the layer, which means dispatch for the first three layers can be planned as soon as the tokens are known rather than after the previous block finishes. </p><h4>three layers where the overlap is free</h4><ul><li><p><strong>The affinity function changed.</strong> <em>V3 scored expert affinity with a sigmoid. V4 uses the square root of a softplus. Both are positive and monotone, so the ranking behaviour is similar, but the tails are not. A sigmoid saturates at one, so once an expert is clearly the best its score stops responding and the gradient through it vanishes. Softplus does not saturate, and the square root damps its growth to sublinear without ever flattening. The practical effect is that a strongly preferred expert keeps receiving gradient signal instead of going quiet, which is exactly the failure mode you would expect to precede the routing-driven loss spikes described two sections later.</em></p></li><li><p><strong>The routing target cap is gone.</strong> <em>V3 constrained how many nodes a token&#8217;s experts could be spread across, through the <span>n_group</span> and <span>topk_group</span> parameters, because unconstrained routing means a token&#8217;s six experts can live on six different machines and the all-to-all cost is set by the worst case. V4 drops the constraint entirely and says the parallelism strategy was redesigned to pay for it. Which is the same trade appearing again in a different costume. The wave-partitioned mega-kernel in section 3.1 exists so that communication hides under computation. Once it does, capping communication to protect throughput stops being necessary, and the model gets its routing freedom back. Every efficiency result in this report buys an architectural freedom somewhere else, and the report never quite says so.</em></p></li><li><p><strong>Load balancing stays auxiliary-loss-free, with one addition.</strong> <em>V4 keeps V3&#8217;s scheme, where a per-expert bias is added to the score for the purpose of top-k selection and excluded from the gating weight, so balance is enforced without a gradient term competing with the language modelling objective. The Hugging Face implementation keeps it as an <span>e_score_correction_bias</span> buffer that shifts the argmax without carrying gradients. On top of that V4 adds a mild sequence-wise balance loss whose only job is to stop a single sequence collapsing onto a handful of experts, with a weight of 0.0001 and a bias update speed of 0.001. Global balance from a bias, local balance from a loss, and the loss is small enough to be a guardrail rather than an objective.</em></p></li></ul><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Manifold-Constrained Hyper-Connections</h2><p>Hyper-connections widen the residual stream from one vector to n_hc vectors and let the model learn the mixing, which decouples residual width from hidden size. </p><p><strong>The update is X<sub>l+1</sub> = B<sub>l</sub>X<sub>l</sub> + C<sub>l</sub>F<sub>l</sub>(A<sub>l</sub>X<sub>l</sub>),</strong> with A projecting the widened stream down to the layer input, C projecting the layer output back up, and B mixing the residual lanes among themselves.</p><p>DeepSeek&#8217;s stated problem with plain hyper-connections is that stacking them is numerically unstable. B is applied at every junction, so across 86 junctions the model computes a product of <strong>86 learned matrices. </strong>If the spectral norm of B exceeds 1 by any margin, the product diverges. If it falls below 1, the signal dies.</p><p>The fix is to constrain B to the <strong>Birkhoff polytope</strong>, the set of doubly stochastic matrices: nonnegative, every row and column summing to one. Two properties make this the right set rather than a convenient one. </p><p>A <strong>doubly stochastic matrix</strong> has spectral norm exactly 1, so the residual mapping is non-expansive by construction and neither the forward nor the backward pass can blow up through it. And the set is closed under multiplication, so a product of 86 of them is still doubly stochastic. The stability is structura-l, not empirical.</p><p>Projection onto the polytope uses <strong>Sinkhorn-Knopp:</strong> exponentiate the raw matrix for positivity, then alternate row and column normalisation, twenty times. A and C get a sigmoid, with C scaled by two so it can express amplification up to a factor of two while staying nonnegative, which rules out lanes cancelling one another.</p><p>One implementation detail worth recording because the paper skips it: the widened stream has to collapse before the output. A final hyper-head folds the <strong>four residual lanes</strong> back into a single sequence just ahead of the model norm, so the language modelling head sees an ordinary hidden state and nothing downstream needs to know the residual was ever four vectors wide.</p><p><strong>At n_hc = 4 this costs 393K parameters per junction and 34M across the model</strong>, about a hundredth of one percent of the weights. The cost is elsewhere. The residual stream is four times wider in activation memory, twenty Sinkhorn iterations sit on the critical path of every junction, and pipeline communication between stages grows. </p><p>DeepSeek&#8217;s answer is obviously <strong>fused kernels</strong>, a recomputation strategy that checkpoints most inter-layer hidden states and all normalised layer inputs while leaving compute-intensive operations alone, and an adjustment to the DualPipe 1F1B overlap so parts of mHC run concurrently with the pipeline. </p><p>The number they report for the whole apparatus is 6.7 percent of the overlapped 1F1B stage.</p><p><strong>SGLang found the other end</strong> of the same problem at serving time. In low-latency decode the batch is small, the <strong>pre-GEMM</strong> that feeds the Sinkhorn normalisation has almost no parallelism, and it becomes the bottleneck. Their answer was to split the K dimension of that GEMM across CTAs. </p><p>Which is where the determinism section&#8217;s remark about an output dimension of 24 comes from: the<strong> GEMM is small enough that split-k is compulsory</strong> and split-k is non-deterministic, so they emit each split separately and reduce in a following kernel.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Muon, and the thing it did not break</h2><p>V4 trains with Muon for most parameters and AdamW for the embedding, the prediction head, the static biases and gating factors of mHC, and all <strong>RMSNorm weights</strong>. Momentum 0.95, weight decay 0.1, update RMS rescaled to 0.18 so the AdamW learning rate schedule could be reused unchanged.</p><p>The orthogonalisation is a hybrid Newton-Schulz, ten iterations in two stages. The first eight use coefficients (<em>3.4445, -4.7750, 2.0315</em>), which converge fast and overshoot; the final two use (<em>2, -1.5, 0.5</em>), which are gentler and settle the singular values precisely at one. </p><p>Splitting the schedule this way is the practical answer to a known tension in <strong>Newton-Schulz</strong>: aggressive coefficients reach the neighbourhood quickly and oscillate there, conservative ones land cleanly but slowly.</p><p>The detail worth flagging is a negative result. Muon has a documented pathology where orthogonalised updates keep singular values near uniform, query and key norms drift up together, and pre-softmax logits reach values low precision cannot hold. Moonshot&#8217;s answer in the Kimi work was<strong> QK-Clip</strong>.</p><p>DeepSeek states plainly that they do not use it, because the attention architecture<strong> already applies RMSNorm</strong> to the queries and the KV entries, which prevents the logits from exploding in the first place.</p><p>So the RMSNorm in section 2.3.3, which reads like routine hygiene, is load-bearing for the optimiser choice. That is a co-design decision presented as a footnote.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>How the sparsity was actually taught</h2><p>You can&#8217;t train a top-k selector from scratch. If the indexer is random at initialisation, the <strong>attention sees 512 random blocks</strong>, the gradient signal for choosing better blocks is buried, and the model learns to ignore the compressed path entirely.</p><p>The schedule in section 4.2.2 solves this in stages, and it is worth reading as a recipe rather than a list. Flash starts at <strong>sequence length 4K and extends through 16K and 64K to 1M. </strong>Attention is dense for the first 1T tokens. Sparse attention is introduced at the 64K stage, and before it is switched on there is a short stage that warms up the lightning indexer alone. Then sparse attention runs for the rest of training, which is most of 32T tokens.</p><p>The ordering is the interesting part. Dense first so the model learns what to attend to; then an <strong>indexer warmup </strong>so the selector learns to imitate the dense attention&#8217;s choices; then sparsity, at a sequence length long enough that selection matters and short enough that dense supervision was still affordable to produce. </p><p><strong>Pro gets a longer dense stage than Flash</strong>, which is what you would expect if the dense phase is the expensive part and the larger model needs more of it.</p><p>Two other schedule details are worth having. Batch size ramps to 75.5M tokens and stays there. Learning rate warms over 2000 steps to 2.7e-4, holds, then decays to 2.7e-5 on a cosine near the end. </p><p><strong>MTP loss weight is 0.3</strong> for most of training and drops to 0.1 when the learning rate starts decaying, which reads as a decision to stop letting the speculative head pull on the backbone once the model is being finished.</p><h3>The instability, and the two things that fixed it</h3><p>The report is unusually candid here. Training was unstable, rollbacks did not prevent recurrence, and the spikes were <strong>consistently traced to outliers in the MoE layers</strong>, with the routing mechanism itself appearing to make the outliers worse. </p><p>Two techniques fixed it and <strong>DeepSeek </strong>says openly that they do not have a theory for why.</p><div class="callout-block" data-callout="true"><p><strong>Anticipatory routing</strong>. It decouples the routing decision from the backbone update. At step t the model computes features with current parameters but routes with parameters from step t minus delta, and to avoid loading weights twice they fetch step t&#8217;s data early and cache the routing indices during the earlier step&#8217;s forward pass. That costs about 20 percent of wall time, so they do not run it continuously: an automatic detector triggers a short rollback and switches the mode on when a spike occurs, then reverts after a period. Amortised, the overhead is close to nothing. The mechanism is worth thinking about. If routing and features update together, a token that starts going to a bad expert gets a gradient that makes both the expert worse and the routing decision more confident, which is a positive feedback loop with no damping. Freezing the routing for a few steps breaks the loop by making the router a fixed target that the experts have to fit rather than a moving one that co-adapts.</p></div><div class="callout-block" data-callout="true"><p><strong>SwiGLU clamping</strong> is the blunt half. The linear component is clamped to [-10, 10] and the gate component is capped at 10, throughout the training of both models. Clamping an activation is what you reach for after watching a run diverge, and DeepSeek reports it eliminates outliers without compromising performance.</p></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>What the cache costs, checked four ways</h2><p>The report gives ratios against DeepSeek-V3.2 and<strong> vLLM gives absolute figures for Pro</strong>. Nobody publishes the absolute figure for Flash. It is derivable, and the derivation can be validated before it is used.</p><p>Build a byte model from the published dimensions. Section 2.3.3 fixes the rotary dimension at exactly 64, and section 2.3.4 specifies bf16 for the <strong>RoPE dimensions</strong> and FP8 for the rest, so a shared key-value entry of head dimension 512 costs 64 times 2 plus 448, which is 576 bytes. </p><p>The <strong>indexer cache runs in FP4</strong> under quantization-aware training, so a 128-dimension indexer key costs 64 bytes.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!yAdk!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F343d20eb-08c4-4d61-9833-c0699949033e_1600x760.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!yAdk!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F343d20eb-08c4-4d61-9833-c0699949033e_1600x760.jpeg 424w, https://substackcdn.com/image/fetch/$s_!yAdk!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F343d20eb-08c4-4d61-9833-c0699949033e_1600x760.jpeg 848w, https://substackcdn.com/image/fetch/$s_!yAdk!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F343d20eb-08c4-4d61-9833-c0699949033e_1600x760.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!yAdk!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F343d20eb-08c4-4d61-9833-c0699949033e_1600x760.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!yAdk!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F343d20eb-08c4-4d61-9833-c0699949033e_1600x760.jpeg" width="1456" height="692" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/343d20eb-08c4-4d61-9833-c0699949033e_1600x760.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:692,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:74905,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/209988515?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F343d20eb-08c4-4d61-9833-c0699949033e_1600x760.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!yAdk!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F343d20eb-08c4-4d61-9833-c0699949033e_1600x760.jpeg 424w, https://substackcdn.com/image/fetch/$s_!yAdk!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F343d20eb-08c4-4d61-9833-c0699949033e_1600x760.jpeg 848w, https://substackcdn.com/image/fetch/$s_!yAdk!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F343d20eb-08c4-4d61-9833-c0699949033e_1600x760.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!yAdk!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F343d20eb-08c4-4d61-9833-c0699949033e_1600x760.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>Two against figures vLLM published independently, two against ratios stated in the report itself.</em></p><p>V3.2 caches an MLA latent of 512 dimensions plus 64 RoPE and a 128-dimension indexer key. In bf16 that is 1,408 bytes per token per layer, which over 61 layers at 1,048,576 tokens is 83.88 GiB<strong>. vLLM publishes 83.9. V4 at Pro&#8217;s 30 CSA</strong> and 31 HCA layers gives 9.62 GiB in bf16. vLLM publishes 9.62. Applying the same model at production precision, V3.2 comes to 49.12 GB and Flash to 3.62 GB, a ratio of 13.57 to one, against the 13.7x the report prints on <strong>Figure 1. </strong></p><p>And section 3.6.2 says the uncompressed sliding-window state would be roughly eight times the volume of the compressed state; the byte model says 7.2.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!tIaC!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6bd4b030-c6d5-4bbd-89ea-d81f308c65a9_1600x880.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!tIaC!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6bd4b030-c6d5-4bbd-89ea-d81f308c65a9_1600x880.png 424w, https://substackcdn.com/image/fetch/$s_!tIaC!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6bd4b030-c6d5-4bbd-89ea-d81f308c65a9_1600x880.png 848w, https://substackcdn.com/image/fetch/$s_!tIaC!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6bd4b030-c6d5-4bbd-89ea-d81f308c65a9_1600x880.png 1272w, https://substackcdn.com/image/fetch/$s_!tIaC!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6bd4b030-c6d5-4bbd-89ea-d81f308c65a9_1600x880.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!tIaC!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6bd4b030-c6d5-4bbd-89ea-d81f308c65a9_1600x880.png" width="1456" height="801" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6bd4b030-c6d5-4bbd-89ea-d81f308c65a9_1600x880.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:801,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;One 1,048,576-token sequence. The model reproduces vLLM's published V3.2 and V4-Pro figures before being applied to Flash.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="One 1,048,576-token sequence. The model reproduces vLLM's published V3.2 and V4-Pro figures before being applied to Flash." title="One 1,048,576-token sequence. The model reproduces vLLM's published V3.2 and V4-Pro figures before being applied to Flash." srcset="https://substackcdn.com/image/fetch/$s_!tIaC!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6bd4b030-c6d5-4bbd-89ea-d81f308c65a9_1600x880.png 424w, https://substackcdn.com/image/fetch/$s_!tIaC!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6bd4b030-c6d5-4bbd-89ea-d81f308c65a9_1600x880.png 848w, https://substackcdn.com/image/fetch/$s_!tIaC!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6bd4b030-c6d5-4bbd-89ea-d81f308c65a9_1600x880.png 1272w, https://substackcdn.com/image/fetch/$s_!tIaC!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6bd4b030-c6d5-4bbd-89ea-d81f308c65a9_1600x880.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>One 1,048,576-token sequence. The model reproduces vLLM&#8217;s published V3.2 and V4-Pro figures before being applied to Flash.</em></p><p>So: <strong>3.372 GiB per one-million-token sequence</strong>, or 3,453 bytes for each token of context across the whole 43-layer stack. Against the baseline the report chooses, bf16 grouped-query attention with 8 KV heads at head dimension 128, which is 176,128 bytes per token and 172 GiB for the same sequence, Flash comes in at 1.96 percent. </p><p>The report claims <strong>approximately 2 percent</strong>. Derived independently, it holds.</p><p>Three multipliers produce that and only one of them is the headline. Fifty-one times from sequence-axis compression and the interleave, two times from sharing key and value, and slightly under two times from mixed FP8 and FP4 storage. </p><p><strong>Top-k selection, the mechanism everyone names when they describe this model, saves no memory at all</strong>. It saves bandwidth at read time. The memory win is compression, sharing and precision, in that order.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>The indexer is O(n), and that is the ceiling</h2><p>Here is what the efficiency section does not say. Top-k selection bounds the cost of attending. It does not bound the cost of selecting. To pick the <strong>best 512 of 262,144 compressed entries</strong>, the indexer scores all of them. That scan is linear in context length and it never becomes sparse.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!oHYg!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff9f3b59e-3f9b-4c48-bc29-9819d9b45f64_1600x922.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!oHYg!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff9f3b59e-3f9b-4c48-bc29-9819d9b45f64_1600x922.png 424w, https://substackcdn.com/image/fetch/$s_!oHYg!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff9f3b59e-3f9b-4c48-bc29-9819d9b45f64_1600x922.png 848w, https://substackcdn.com/image/fetch/$s_!oHYg!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff9f3b59e-3f9b-4c48-bc29-9819d9b45f64_1600x922.png 1272w, https://substackcdn.com/image/fetch/$s_!oHYg!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff9f3b59e-3f9b-4c48-bc29-9819d9b45f64_1600x922.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!oHYg!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff9f3b59e-3f9b-4c48-bc29-9819d9b45f64_1600x922.png" width="1456" height="839" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f9f3b59e-3f9b-4c48-bc29-9819d9b45f64_1600x922.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:839,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;The matmul side is flat by construction. The attention side is linear in context because of the indexer scan and the HCA dense pass.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="The matmul side is flat by construction. The attention side is linear in context because of the indexer scan and the HCA dense pass." title="The matmul side is flat by construction. The attention side is linear in context because of the indexer scan and the HCA dense pass." srcset="https://substackcdn.com/image/fetch/$s_!oHYg!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff9f3b59e-3f9b-4c48-bc29-9819d9b45f64_1600x922.png 424w, https://substackcdn.com/image/fetch/$s_!oHYg!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff9f3b59e-3f9b-4c48-bc29-9819d9b45f64_1600x922.png 848w, https://substackcdn.com/image/fetch/$s_!oHYg!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff9f3b59e-3f9b-4c48-bc29-9819d9b45f64_1600x922.png 1272w, https://substackcdn.com/image/fetch/$s_!oHYg!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff9f3b59e-3f9b-4c48-bc29-9819d9b45f64_1600x922.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>The matmul side is flat by construction. The attention side is linear in context because of the indexer scan and the HCA dense pass.</em></p><p>At 4K of context, attention is 9 percent of per-token arithmetic and the model behaves like a cheap 13B. At 32K it is 18 percent. At 128K it is 39 percent. <strong>Around 213K tokens</strong> the attention overtakes the entire 284B expert bank, and at 1M it is 82 percent of the work.</p><p><strong> The model is linear in context, not sub-linear.</strong> Compression changed the constant by roughly fifty. It did not change the exponent.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!rOt3!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F86ee4412-8b4d-4163-9535-4c522b82f00f_1600x867.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!rOt3!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F86ee4412-8b4d-4163-9535-4c522b82f00f_1600x867.png 424w, https://substackcdn.com/image/fetch/$s_!rOt3!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F86ee4412-8b4d-4163-9535-4c522b82f00f_1600x867.png 848w, https://substackcdn.com/image/fetch/$s_!rOt3!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F86ee4412-8b4d-4163-9535-4c522b82f00f_1600x867.png 1272w, https://substackcdn.com/image/fetch/$s_!rOt3!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F86ee4412-8b4d-4163-9535-4c522b82f00f_1600x867.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!rOt3!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F86ee4412-8b4d-4163-9535-4c522b82f00f_1600x867.png" width="1456" height="789" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/86ee4412-8b4d-4163-9535-4c522b82f00f_1600x867.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:789,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;At 32K of context, half the attention arithmetic is spent deciding what to attend to. At 1M, four fifths of it is.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="At 32K of context, half the attention arithmetic is spent deciding what to attend to. At 1M, four fifths of it is." title="At 32K of context, half the attention arithmetic is spent deciding what to attend to. At 1M, four fifths of it is." srcset="https://substackcdn.com/image/fetch/$s_!rOt3!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F86ee4412-8b4d-4163-9535-4c522b82f00f_1600x867.png 424w, https://substackcdn.com/image/fetch/$s_!rOt3!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F86ee4412-8b4d-4163-9535-4c522b82f00f_1600x867.png 848w, https://substackcdn.com/image/fetch/$s_!rOt3!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F86ee4412-8b4d-4163-9535-4c522b82f00f_1600x867.png 1272w, https://substackcdn.com/image/fetch/$s_!rOt3!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F86ee4412-8b4d-4163-9535-4c522b82f00f_1600x867.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>At 32K of context, half the attention arithmetic is spent deciding what to attend to. At 1M, four fifths of it is.</em></p><p><strong>Half the attention FLOPs at 32K are the indexer. Four fifths at 1M.</strong> The thing that makes attention sparse is the dominant cost of the attention.</p><blockquote><p><em>Four fifths of the cost of attention in this model is the cost of deciding what to attend to.</em><strong><span>Derived from the published dimensions</span></strong></p></blockquote><p>DeepSeek clearly knew. The indexer runs in FP4 while the main attention runs in FP8, which is the more aggressive precision going to the larger consumer. <strong>Flash&#8217;s top-k came down to 512 while Pro&#8217;s is 1024</strong>, because reducing k shrinks the attention pass and does nothing at all to the scan, and Flash needs the attention pass small relative to its smaller matmul side. </p><p>And the <strong>compression rate m = 4 </strong>is not primarily a memory optimisation: the scan runs over n/m keys, so compressing by four is a four times discount on selection, with the cache saving arriving as a side effect.</p><p>If someone finds a sub-linear selector, this architecture gets a second life. There is already a paper trying, from a group that fine-tuned V4-Flash with a <strong>Neural Memory Indexer </strong>that predicts and prefetches only the query-critical KV chunks, reporting comparable benchmark scores at 13.5 percent of the GPU memory. </p><p>Its authors are explicit that the work was constrained by resources and cut short, with the indexer trained on frozen keys and no end-to-end optimisation against the backbone. As a result it is <strong>a direction rather than a result.</strong> It is the right direction.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>FP4, down to the index scores</h2><p>The so-called<strong> Quantization-aware training</strong> is applied during post-training to two things: the MoE expert weights, and the query-key path of the CSA indexer, where activations are cached, loaded and multiplied entirely in FP4.</p><p>The expert-weight scheme has a property worth stating because it explains why the whole thing was affordable. Master weights are held in FP32, quantised to MXFP4, then dequantised back to FP8 for the actual computation, and <em><strong>the FP4 to FP8 dequantisation is lossless</strong></em>. </p><p><strong>FP8 in E4M3</strong> has two more exponent bits than FP4 in E2M1, so as long as the ratio between the largest and smallest scale factors of the FP4 sub-blocks, which are 1 by 32 tiles, inside a given FP8 quantisation block, which is 128 by 128, stays under a threshold, the finer scale information is absorbed entirely by the wider dynamic range. </p><p>DeepSeek verified their weights satisfy the condition. The consequence is that the <strong>entire QAT pipeline</strong> reuses the existing FP8 training framework without modification, with a straight-through estimator carrying gradients back to the FP32 masters, and no need to requantise transposed weights.</p><p>Then there is one number in that section that deserves its own paragraph. They also quantise the index scores themselves, the output of the lightning indexer, from FP32 to BF16. </p><p>That gives a <strong>two times speedup on the top-k selector while preserving a 99.7 percent recall rate of KV entries.</strong> Given that the selector is the dominant term in attention cost at long context, halving it for three tenths of a percent of recall is the highest-leverage line in the report.</p><p>During rollouts and any inference-only forward pass, including teachers and reference models,<em> real FP4 weights are used rather than simulated quantisation</em>, so sampling behaviour during RL is identical to deployment behaviour. </p><p>That is a correctness argument dressed as an efficiency one, and it matters: a policy trained against a simulated-quantisation rollout is optimising a model that will never be served.</p><div><hr></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><h2>The mega-kernel, and the balance point DeepSeek wants hardware to hit</h2><p>Expert parallelism needs an all-to-all dispatch and an all-to-all combine per MoE layer, and the <strong>conventional implementation</strong> runs communication and computation as separate serial kernels, which leaves both the interconnect and the SMs idle half the time. </p><p>DeepSeek&#8217;s answer fuses them into one pipelined kernel and then partitions the experts into waves. As soon as every expert in a wave has its tokens, that wave computes, while the next wave&#8217;s tokens are still in flight and the previous wave&#8217;s results are being sent back. In steady state all three proceed at once.</p><p>The <strong>reported gains are 1.50 to 1.73 times against strong non-fused baselines</strong> for general inference, and up to 1.96 times for latency-sensitive work like RL rollouts and high-speed agent serving, where batches are small and long-tailed and the pipeline has the most idle time to recover. </p><p>The report&#8217;s own figure puts the theoretical ceiling of the wave scheme at 1.92 times against 1.42 for Comet, which overlaps dispatch with the first linear and the second linear with combine but not at wave granularity, and it evaluates both in the V4-Flash configuration specifically. The implementation is open, as <strong>MegaMoE inside DeepGEMM</strong>, and it was validated on both NVIDIA GPUs and Huawei Ascend NPUs.</p><p>Then comes the paragraph I think is the most economically consequential in the entire report, and it is addressed to hardware vendors rather than to users.</p><p>Communication hides under computation when<strong> C/B is at most V<sub>comp</sub>/V<sub>comm</sub>,</strong> where C is peak compute and B is interconnect bandwidth. For a DeepSeekMoE layer each token-expert pair costs 6hd<sub>ff</sub> FLOPs across the gate, up and down projections, and 3h bytes of traffic, being h bytes of FP8 dispatch and 2h bytes of BF16 combine. </p><p>The h cancels. The condition collapses to:</p><pre><code>C / B  &lt;=  2 * d_ff</code></pre><p>The report evaluates this for Pro, whose <strong>expert intermediate dimension is 3072</strong>, and gets 6144 FLOPs per byte, then observes that once bandwidth clears that threshold it stops being the bottleneck and further silicon spent on it brings diminishing returns. Their recommendation to hardware designers is to target the balance point rather than scale bandwidth unconditionally.</p><blockquote><p><em>The report contains the equation that says the cheaper model is the harder one to host. It does not evaluate it for the cheaper model.</em><strong><span>Report section 3.1</span></strong></p></blockquote><p>Flash&#8217;s expert intermediate dimension is 2048. Run the same derivation and Flash&#8217;s balance point is 4096 FLOPs per byte, which is two thirds of Pro&#8217;s.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!5DGw!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc3fc06f6-13ba-4789-95b3-0f2ff5cec30c_1600x748.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!5DGw!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc3fc06f6-13ba-4789-95b3-0f2ff5cec30c_1600x748.png 424w, https://substackcdn.com/image/fetch/$s_!5DGw!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc3fc06f6-13ba-4789-95b3-0f2ff5cec30c_1600x748.png 848w, https://substackcdn.com/image/fetch/$s_!5DGw!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc3fc06f6-13ba-4789-95b3-0f2ff5cec30c_1600x748.png 1272w, https://substackcdn.com/image/fetch/$s_!5DGw!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc3fc06f6-13ba-4789-95b3-0f2ff5cec30c_1600x748.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!5DGw!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc3fc06f6-13ba-4789-95b3-0f2ff5cec30c_1600x748.png" width="1456" height="681" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c3fc06f6-13ba-4789-95b3-0f2ff5cec30c_1600x748.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:681,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Derived from the report's own condition at 4.5 PFLOP/s of dense FP8 per GPU. Smaller experts do less arithmetic per byte moved.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Derived from the report's own condition at 4.5 PFLOP/s of dense FP8 per GPU. Smaller experts do less arithmetic per byte moved." title="Derived from the report's own condition at 4.5 PFLOP/s of dense FP8 per GPU. Smaller experts do less arithmetic per byte moved." srcset="https://substackcdn.com/image/fetch/$s_!5DGw!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc3fc06f6-13ba-4789-95b3-0f2ff5cec30c_1600x748.png 424w, https://substackcdn.com/image/fetch/$s_!5DGw!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc3fc06f6-13ba-4789-95b3-0f2ff5cec30c_1600x748.png 848w, https://substackcdn.com/image/fetch/$s_!5DGw!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc3fc06f6-13ba-4789-95b3-0f2ff5cec30c_1600x748.png 1272w, https://substackcdn.com/image/fetch/$s_!5DGw!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc3fc06f6-13ba-4789-95b3-0f2ff5cec30c_1600x748.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>Derived from the report&#8217;s own condition at 4.5 PFLOP/s of dense FP8 per GPU. Smaller experts do less arithmetic per byte moved.</em></p><p><strong>The smaller model is the harder one to interconnect.</strong> Flash needs 1.5 times more bandwidth per unit of compute than Pro does, because its experts do less arithmetic for the same number of bytes dispatched and combined. </p><p>At a B200&#8217;s dense FP8 throughput, <strong>Flash wants about 1.10 TB/s per GPU</strong> to hide its all-to-all and Pro wants about 0.73. NVLink 5 covers both comfortably. A PCIe Gen5 box does not cover either, and misses Flash by a factor of seventeen.</p><p>Anyone sizing a deployment on the assumption that the cheaper model is the easier one to host has the relationship backwards, and the report contains the equation that says so.</p><p>Three other proposals in that section are worth recording because they are a roadmap. DeepSeek asks for more power headroom, on the grounds that <strong>extreme fusion drives compute</strong>, memory and network to high load simultaneously and power throttling becomes the limiter, which is a real and underdiscussed consequence of fusing everything. </p><p>They use pull-based communication, where each GPU reads from remote GPUs, because fine-grained push carries too much notification latency, and they ask for <strong>lower-latency cross-GPU signalling</strong> so push becomes viable. </p><p>And they <strong>propose replacing SwiGLU</strong> with a cheap elementwise activation with no exponential and no division, because that lightens post-GEMM work and, under a fixed parameter budget, removing the gate projection lets d<sub>ff</sub> grow, which pushes the balance point up and relaxes the bandwidth requirement further.</p><p>That last one is a description of the next model.<br></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>TileLang, and a theorem prover in the compiler</h2><p>Section 3.2 is the part of this report that a compiler person should read twice. DeepSeek&#8217;s architecture, written naively, <strong>decomposes into hundreds of fine-grained Torch ATen operators</strong>, and they replaced most of them with fused kernels written in TileLang, a tile-level DSL, rather than in CUDA.</p><p>They did two things to TileLang along the way are more interesting than the choice itself.</p><div class="callout-block" data-callout="true"><p><strong>Host codegen.</strong> As accelerators get faster, CPU-side orchestration becomes the ceiling for small kernels, and the usual source is host-side logic such as runtime contract checks written in Python for flexibility. DeepSeek co-generates the device kernel and a lightweight host launcher at the IR level, embedding data types, rank and shape constraints and stride and layout assumptions parsed from the frontend, then lowers the launcher to host source on TVM-FFI, whose compact calling convention and zero-copy tensor interop keep the overhead small. Validation and argument marshalling happen in generated C rather than in Python. Their measurement: <em>CPU-side validation drops from tens or hundreds of microseconds per invocation to under one</em>.</p><p>That is a two-order-of-magnitude reduction in a cost that most people do not measure at all, and it is the sort of thing that only shows up when your model has enough small kernels for launch overhead to dominate. Which this one does, by construction.</p></div><div class="callout-block" data-callout="true"><p><strong>Z3 in the algebraic system.</strong> TileLang kernels are full of complex tensor index arithmetic, and passes like layout inference, memory hazard detection and bound analysis all need to prove properties of integer expressions before they are allowed to fire. Weak integer reasoning means conservative passes means slower kernels. DeepSeek integrated the Z3 SMT solver into TileLang&#8217;s algebraic system, translating integer expressions into quantifier-free non-linear integer arithmetic, which handles the ordinary linear index algebra through ILP and the harder cases, such as vectorising over variable tensor shapes, through genuine non-linear reasoning. They report a few seconds of added compilation time and improvements across vectorisation, barrier insertion and simplification.</p></div><p><strong>A production LLM shipped with an SMT solver inside its kernel compiler. </strong>This is the direction I have argued the field goes: the bottleneck in kernel performance is not the language, it is how much the compiler can prove, and buying proving power off the shelf is cheaper than hand-writing the kernel.</p><p>The numerics policy in the same section is equally deliberate. <strong>Fast-math is disabled at the compiler level</strong> by default, precision-affecting approximations are opt-in frontend operators, and IEEE-compliant intrinsics with explicit rounding modes are available when strict semantics are required. </p><p>They also align TileLang&#8217;s algebraic simplification and lowering rules with <em>NVCC </em>so that kernels can be validated<strong> bit-for-bit against hand-written CUDA baselines</strong>, with layout annotations available to pin down lowering decisions and hold accumulation order constant. Accuracy by default, speed by opt-in, which is the opposite of the usual arrangement.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share&quot;,&quot;text&quot;:&quot;Share The Software Frontier&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share The Software Frontier</span></a></p><div><hr></div><h2>Bitwise batch invariance, and why they wanted it</h2><p>Batch invariance means a given token&#8217;s output is bitwise identical regardless of where it sits in a batch. Almost nobody ships this, because it costs performance, and <strong>DeepSeek&#8217;s reasoning for paying</strong> is that they wanted bitwise alignment across pre-training, post-training and inference, which makes loss spikes diagnosable and post-training behaviour consistent with what gets served.</p><p>Getting it required giving up two standard optimisations and then engineering the loss back out.</p><blockquote><p><strong>Attention.</strong> Split-KV, which spreads one sequence&#8217;s attention across many SMs to balance load, is not batch invariant. Abandoning it causes wave quantisation, where the final partially filled wave of thread blocks leaves most of the GPU idle. DeepSeek&#8217;s answer is a dual-kernel decode: a first kernel computes an entire sequence&#8217;s attention inside a single SM, giving throughput on fully occupied waves, and a second kernel spreads one sequence across multiple SMs to shorten the trailing partial wave. The two are engineered to have the <em>same accumulation order</em> so their outputs are bit-identical, and the second uses distributed shared memory within thread-block clusters to exchange partial results across SMs at speed. The reported overhead of batch-invariant decoding after this is negligible.</p></blockquote><blockquote><p><strong>Matrix multiplication.</strong> cuBLAS cannot be made batch invariant, so it is replaced end to end by DeepGEMM. Split-k, which is how you get performance at very small batch, is also not batch invariant, so it is dropped in most scenarios and the resulting loss is recovered by other means.</p></blockquote><p>Determinism is a separate problem from batch invariance and it comes from accumulation order, usually via atomic addition in the backward pass. </p><p>Sparse attention backward normally uses atomicAdd to accumulate KV gradients, so they <strong>allocate a separate accumulation buffer per SM</strong> and do a global deterministic summation afterwards. MoE backward is non-deterministic because SMs from different ranks negotiate write positions into the same receiving buffer, so they pre-process token order within each rank and isolate buffers across ranks. </p><p>And the <strong>mHC GEMM with its output dimension of 24 is small enough that split-k is unavoidable</strong>, so each split is emitted separately and reduced deterministically in a following kernel.</p><p>There is a payoff for this that shows up two sections later, in the rollout service. Because generation can be preempted at any time on their cluster, they keep a token-granular write-ahead log per request and resume from it. </p><p>The report explains why <strong>they cannot simply regenerate interrupted requests </strong>from scratch, and the argument is a good one: shorter responses are more likely to survive an interruption, so regenerating the survivors biases the training distribution towards short outputs. </p><p>They note that a batch-invariant deterministic stack could fix this instead by regenerating with a consistent sampler seed, but that this still costs a <strong>full re-decode</strong>, so the log wins. Batch invariance is what makes that alternative even expressible.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Two cache hierarchies, and lcm(4, 128)</h2><p>The <strong>hybrid attention breaks the assumption PagedAttention is built on</strong>, which is that every layer&#8217;s KV state has the same shape and the same eviction policy. </p><p>In V4 the compression ratios differ per layer, the indexer carries its own embedding size, the sliding-window layers have their own hit and eviction rules, and there is a <strong>rolling residual of tokens</strong> not yet numerous enough to compress that has to live somewhere.</p><p>DeepSeek splits it in two. A classical paged KV cache holds the compressed CSA and HCA entries. A separate <em>state cache</em> holds the sliding-window entries and the uncompressed tail, on the argument that both are a function only of the current position, which makes them a<strong> state-space model rather than a growing history</strong>, so a fixed-size pool can be pre-allocated and assigned per sequence.</p><p>The block geometry falls out of the two compression rates. A block has to cover a whole number of compressed entries in every layer, so it must span a multiple of the least common multiple of m and m&#8217;. </p><p><strong>For Flash that is lcm(4, 128) = 128 original tokens</strong>, giving 32 CSA entries and exactly one HCA entry per block. vLLM independently chose 256 native positions, which is two of DeepSeek&#8217;s minimum blocks, giving 64 c4a entries and 2 c128a entries.</p><p><strong>vLLM&#8217;s implementation notes</strong> are the best serving document published on this model, and their three decisions are worth having next to DeepSeek&#8217;s. </p><p>One logical block size in native token positions for every compressed layer, so slot mapping, scheduler accounting and prefix-hit detection use one unit instead of branching on the compression ratio. The compressor&#8217;s rolling residual registered as<strong> sliding-window KV with sliding_window set to the compression stride</strong>, rather than as a side buffer, so prefix caching lands on block boundaries and disaggregated prefill ships it through the existing SWA transfer path instead of a second one. </p><p>And a page-size argument: <strong>page size is block_size times compress_ratio times entry_size, all three are controllable</strong>, and chosen carefully the five cache kinds collapse into three buckets, each backed by one pool, sized once at load, with no runtime repartitioning and no cross-kind fragmentation.</p><p>On the kernel side vLLM fuses three groups: compressor with RMSNorm, RoPE and cache insertion, all elementwise, for 1.4 to 3 times; inverse RoPE with <strong>FP8 quantisation ahead of the output projection, for 2 to 3 times</strong>; and a horizontal fusion of query normalisation, KV RoPE and sliding-window key insertion using static warp-ID dispatch, each warp working independently on a query head or a key head with no cross-warp communication, for 10 to 20 times over the naive version. </p><p>The <strong>indexer then runs on its own CUDA stream </strong>alongside KV compression and window insertion, worth 5 to 6 percent end to end at low batch.</p><p>SGLang went at the FP4 weights instead, pairing MXFP8 activations with MXFP4 expert weights through <strong>FlashInfer&#8217;s TRTLLM-Gen fused MoE backend, and splitting K across CTAs</strong> in the mHC pre-GEMM for exactly the small-batch parallelism problem described earlier.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/p/deepseek-v4-flash-the-cost-of-deciding?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/p/deepseek-v4-flash-the-cost-of-deciding?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><div><hr></div><h2>What a cache hit actually costs</h2><p>DeepSeek stores all compressed CSA and HCA entries to disk. </p><p>When a request hits a stored prefix, <strong>those entries are read back rather than recomputed</strong>, up to the last complete compression block; the tail of an incomplete block still has to be recomputed because uncompressed entries are not stored.</p><p>The sliding-window state is the problem. It is uncompressed and it exists in every layer, so storing it for every token would be roughly eight times the volume of everything else. </p><p>The report gives <strong>precisely three strategies</strong> with different trade-offs, and the third is the one that makes the economics work.</p><ol><li><p><strong>Full SWA caching</strong> stores everything, so a hit reads the last n_win tokens of the prefix and recomputes nothing. Zero redundancy, but only a sliver of what was written is ever read, which is a write-heavy unbalanced access pattern that SSDs handle badly.</p></li><li><p><strong>Periodic checkpointing</strong> saves the window state every p tokens, loads the nearest checkpoint on a hit and recomputes the tail, with p tuning the storage against compute trade.</p></li><li><p><strong>Zero SWA caching</strong> stores none of it. Here is the argument, and it is the neatest piece of reasoning in the report. Each token&#8217;s sliding-window entry in a given layer depends only on the window entries of the previous layer, which span n_win tokens. So the dependency cone going back through L layers is exactly n_win times L tokens wide. Recompute that many tokens and the entire window state is restored.</p></li></ol><blockquote><p><em>A million-token prefix comes back for the price of five and a half thousand.</em><strong><span>Report section 3.6.2, zero SWA caching</span></strong></p></blockquote><p>For Flash, n_win times L is 128 times 43, which is 5,504 tokens. Restoring the full sliding-window state of a one-million-token prefix costs a 5,504-token recompute, about 140 TFLOP, against the 26.6 PFLOP a cold prefill of that prefix would cost. <strong>That is a factor of 191, and it is why a cache hit can be sold for a fiftieth of a cache miss.</strong></p><p>The whole thing only works because the compressed state is small. Storing 172 GiB per session on disk is possible, but reading it back at request time competes with the prefill it was meant to replace. At 3.62 GB it does not. </p><p>Compressing the sequence axis is <strong>what turns disk from a bad idea into the cheapest tier in the hierarchy</strong>, and the $0.0028 line on the rate card is the direct commercial expression of that.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/p/deepseek-v4-flash-the-cost-of-deciding?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/p/deepseek-v4-flash-the-cost-of-deciding?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><div><hr></div><h2>Post-training, and the parts of it that are cost mechanisms</h2><p>The pipeline broadly follows <em>DeepSeek-V3.2 with one substitution the report calls critical</em>: the mixed reinforcement learning stage is replaced entirely by on-policy distillation. </p><p>Domain specialists are trained separately, each through supervised fine-tuning then<strong> GRPO against domain-specific rewards</strong>, and then more than ten of them are merged into one student by having the student sample its own trajectories and minimise reverse KL against the relevant teacher.</p><p>They use full-vocabulary logit distillation rather than the usual token-level KL estimate, on the grounds that the cheap estimator has high gradient variance and destabilises training. Making that affordable took two tricks. </p><p>Teacher weights live in centralised distributed storage and are loaded on demand with <strong>ZeRO-like sharding. </strong></p><p>And rather than materialising logits for a vocabulary above 100k across more than ten teachers, they cache only the last-layer teacher hidden states and <strong>reconstruct logits on the fly through the prediction head at training time</strong>, ordering training samples by teacher index so exactly one teacher head is resident on device at a time. The KL itself is a TileLang kernel.</p><p>Three things from the post-training section have direct consequences for what you pay.</p><h3>The reasoning ladder is a context ladder</h3><p>Three modes, trained as separate RL configurations with distinct length penalties and context windows, then unified. Non-think evaluates at 8K of context,<strong> Think High at 128K, Think Max at 384K.</strong> Non-think emits an empty reasoning block and goes straight to the summary. </p><p>Think Max additionally prepends a system instruction, printed verbatim in the report, which tells the model that shortcuts are not permitted and that it must document every intermediate step, considered alternative and rejected hypothesis.</p><p>That instruction is <strong>why Artificial Analysis</strong> measures this model generating 210M output tokens across its index against a 100M median. The verbosity is not a training accident. </p><p><strong>It is an instruction, in the system prompt, that DeepSeek wrote and that you are billed for at $0.28 per million. </strong>Anyone running Flash at max effort and complaining about token burn is paying for a behaviour they asked for by name.</p><h3>The tool schema is XML on purpose</h3><p>V4 introduces a tool-call format built on a dedicated <span>|DSML|</span> token with XML-shaped invocations rather than JSON. </p><p>String parameters go through as-is with an explicit <span>string=&#8221;true&#8221;</span> flag; <strong>everything else is JSON-encoded with the flag false. </strong>The stated reason is that XML mitigates escaping failures and reduces tool-call errors.</p><p>This is a small decision with a large downstream effect. Every escaping failure in a JSON tool call is a wasted turn, and a wasted turn in an agent loop costs a full round of prefill plus decode. Reducing tool-call error rate is a cost reduction that never appears on a rate card.</p><h3>Interleaved thinking, and a warning inside it</h3><p>V3.2 kept reasoning traces across tool-result rounds but discarded them when a new user message arrived. V4 keeps everything, across user message boundaries, for <strong>tool-calling conversations</strong>, so a long-horizon agent maintains one cumulative chain of thought instead of reconstructing its state each turn. </p><p>General conversation keeps the old discarding behaviour, on the reasonable grounds that persistent traces buy little there and cost context.</p><p>The <strong>warning is in the same paragraph </strong>and it is easy to miss. Agent frameworks that simulate tool interactions through user messages, and the report names Terminus, will not trigger the tool-calling context path and therefore will not get the persistence. </p><p>DeepSeek&#8217;s own recommendation for those frameworks is to use non-think models. If you are <strong>benchmarking Flash inside a harness </strong>that fakes tools as user turns, you are measuring the wrong path.</p><h3>Quick Instruction, which is an inference-economics feature wearing a post-training costume</h3><p>In a chat product, a handful of auxiliary decisions run before the real response: <strong>whether to trigger a web search</strong>, what the query should be, how authoritative a source needs to be, what domain the request belongs to, whether a pasted URL should be fetched. </p><p>The standard answer is a separate small model, which means a second prefill of the same prompt because it cannot share the big model&#8217;s KV cache.</p><p>DeepSeek trained special tokens for each of those tasks and appends them to the input sequence directly. The<strong> auxiliary task runs on the already-computed KV cache. </strong>There is no second prefill, several of the tasks run in parallel, the user-perceived time to first token drops, and there is no small model to maintain.</p><p>The published tokens are <strong><span>|action|</span>, <span>|title|</span>, <span>|query|</span>, <span>|authority|</span>, <span>|domain|</span>, <span>|extracted_url|</span> and <span>|read_url|</span></strong>. Read the list and it is obvious this is DeepSeek&#8217;s own chat product spilling into the model card, which is exactly what makes it interesting. </p><p>It is a vertical integration of the router into the weights, and it deletes a whole class of serving infrastructure that everyone else runs.</p><h3>The sandbox, briefly</h3><p>Agentic RL needs somewhere to execute, and DeepSeek built a platform they call DSec:<strong> three Rust components</strong> on top of their 3FS distributed filesystem, running hundreds of thousands of concurrent sandbox instances per cluster. </p><ul><li><p>One Python SDK abstracts four execution substrates behind one API, switchable by a parameter. </p></li><li><p>Function Call dispatches stateless invocations to a pre-warmed pool with no cold start. </p></li><li><p>Container is Docker-compatible with EROFS on-demand image loading. MicroVM is Firecracker for security-sensitive high-density work. </p></li><li><p>FullVM is QEMU for arbitrary guest operating systems. </p></li><li><p>Base images sit on 3FS as read-only layers shared across instances, writes go to a local copy-on-write layer, and snapshots chain, which gets them millisecond-scale resumption.</p></li></ul><p>Each sandbox keeps a globally ordered trajectory log of every command and result, which serves three purposes: fast-forwarding after a preemption by <strong>replaying cached results rather than re-executing non-idempotent commands</strong>, provenance for every state change, and deterministic replay of any historical session.</p><p>None of this is in the model. All of it is why the model has agent scores.</p><div class="community-chat" data-attrs="{&quot;url&quot;:&quot;https://open.substack.com/pub/softwarefrontier/chat?utm_source=chat_embed&quot;,&quot;subdomain&quot;:&quot;softwarefrontier&quot;,&quot;pub&quot;:{&quot;id&quot;:3575776,&quot;name&quot;:&quot;The Software Frontier&quot;,&quot;author_name&quot;:&quot;Lorenzo Bradanini&quot;,&quot;author_photo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!ACM6!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F18342bf1-31cb-404a-b9e1-998a38d299bf_1200x1200.jpeg&quot;}}" data-component-name="CommunityChatRenderPlaceholder"></div><div><hr></div><h2>What a node holds</h2><p>Now the economics, and they start with capacity rather than speed.</p><p><strong>Four B200 at 180 GB usable is 720 GB of HBM</strong>. The weights ship natively mixed: FP4 for routed experts, FP8 for attention, norms and router. That is 139 GB plus 6 GB, call it 145 GB resident. </p><p>At 85 percent HBM utilisation, leaving room for activations, the four-times-wider mHC residual stream and the compressor states, the KV pool is roughly 467 GB.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!yz23!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F375bb791-df0e-4820-9b0a-07b8547e8d77_1600x882.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!yz23!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F375bb791-df0e-4820-9b0a-07b8547e8d77_1600x882.png 424w, https://substackcdn.com/image/fetch/$s_!yz23!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F375bb791-df0e-4820-9b0a-07b8547e8d77_1600x882.png 848w, https://substackcdn.com/image/fetch/$s_!yz23!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F375bb791-df0e-4820-9b0a-07b8547e8d77_1600x882.png 1272w, https://substackcdn.com/image/fetch/$s_!yz23!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F375bb791-df0e-4820-9b0a-07b8547e8d77_1600x882.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!yz23!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F375bb791-df0e-4820-9b0a-07b8547e8d77_1600x882.png" width="1456" height="803" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/375bb791-df0e-4820-9b0a-07b8547e8d77_1600x882.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:803,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;The same node, the same money, the same power draw. The difference is which axis was compressed.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="The same node, the same money, the same power draw. The difference is which axis was compressed." title="The same node, the same money, the same power draw. The difference is which axis was compressed." srcset="https://substackcdn.com/image/fetch/$s_!yz23!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F375bb791-df0e-4820-9b0a-07b8547e8d77_1600x882.png 424w, https://substackcdn.com/image/fetch/$s_!yz23!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F375bb791-df0e-4820-9b0a-07b8547e8d77_1600x882.png 848w, https://substackcdn.com/image/fetch/$s_!yz23!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F375bb791-df0e-4820-9b0a-07b8547e8d77_1600x882.png 1272w, https://substackcdn.com/image/fetch/$s_!yz23!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F375bb791-df0e-4820-9b0a-07b8547e8d77_1600x882.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>The same node, the same money, the same power draw. The difference is which axis was compressed.</em></p><p><strong>129 concurrent one-million-token sessions on one four-GPU node</strong>. A V3.2-style stack on the same hardware holds five. At 128K, the working length of most real agent traffic, it is 1,031.</p><p>This is the actual product. Not the million-token window as a marketing number, but the ability to keep a thousand long-lived agent sessions warm on one node without evicting anyone. </p><p>Eviction is what makes long-context serving expensive, because every eviction is a <strong>re-prefill a</strong>nd a re-prefill of 128K tokens costs more than the entire conversation that followed it. Compressing the sequence axis turns that from a scheduling problem into a non-problem.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>The roofline, and what $0.28 buys</h2><p>Decode on a large MoE is bandwidth-bound. At any batch deep enough to hit every expert, the node streams the full expert bank once per step: <strong>278.1B parameters at FP4 is 139 GB</strong>, plus the dense stack replicated across four data-parallel ranks, giving 164 GB of weight traffic per step against 32 TB/s of aggregate bandwidth. A 5.12 ms floor, or 195 steps per second.</p><p>Artificial Analysis measures 122.7 output tokens per second per user on DeepSeek&#8217;s own API, which is 8.15 ms per step. <strong>The model achieves 63 percent of peak HBM bandwidth.</strong> That is a good number for something with this much elementwise work between matmuls, and it is a direct vindication of the fusion work in both vLLM and DeepSeek&#8217;s own kernels. </p><p>Holding that efficiency and adding KV read traffic gives node throughput, and node throughput times $0.28 per million gives revenue.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!EU1M!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2848795e-f386-486d-b1ba-eadeb90eb0c5_1600x862.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!EU1M!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2848795e-f386-486d-b1ba-eadeb90eb0c5_1600x862.png 424w, https://substackcdn.com/image/fetch/$s_!EU1M!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2848795e-f386-486d-b1ba-eadeb90eb0c5_1600x862.png 848w, https://substackcdn.com/image/fetch/$s_!EU1M!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2848795e-f386-486d-b1ba-eadeb90eb0c5_1600x862.png 1272w, https://substackcdn.com/image/fetch/$s_!EU1M!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2848795e-f386-486d-b1ba-eadeb90eb0c5_1600x862.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!EU1M!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2848795e-f386-486d-b1ba-eadeb90eb0c5_1600x862.png" width="1456" height="784" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2848795e-f386-486d-b1ba-eadeb90eb0c5_1600x862.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:784,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Break-even GPU rate from output revenue alone. The gold band is the rentable B200 market, roughly $3.35 reserved to $6.35 median as of August 2026.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Break-even GPU rate from output revenue alone. The gold band is the rentable B200 market, roughly $3.35 reserved to $6.35 median as of August 2026." title="Break-even GPU rate from output revenue alone. The gold band is the rentable B200 market, roughly $3.35 reserved to $6.35 median as of August 2026." srcset="https://substackcdn.com/image/fetch/$s_!EU1M!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2848795e-f386-486d-b1ba-eadeb90eb0c5_1600x862.png 424w, https://substackcdn.com/image/fetch/$s_!EU1M!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2848795e-f386-486d-b1ba-eadeb90eb0c5_1600x862.png 848w, https://substackcdn.com/image/fetch/$s_!EU1M!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2848795e-f386-486d-b1ba-eadeb90eb0c5_1600x862.png 1272w, https://substackcdn.com/image/fetch/$s_!EU1M!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2848795e-f386-486d-b1ba-eadeb90eb0c5_1600x862.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>Break-even GPU rate from output revenue alone. The gold band is the rentable B200 market, roughly $3.35 reserved to $6.35 median as of August 2026.</em></p><p>At 32K of context you need a sustained decode batch of about 256 before <strong>$0.28 per million covers a mid-market B200 at $5.50 per GPU-hour.</strong> At 512 you clear 57 percent gross margin, at 1024 you clear 77. </p><p>At 1M of context it never clears: the KV pool caps the batch at 129 sequences and the break-even rate there is $2.92 per GPU-hour, below the cheapest reserved B200 anyone publishes.</p><p>That is the analysis I would&#8217;ve published if I had stopped there, and it is wrong in two ways that point in opposite directions. It prices only output tokens, which understates the revenue. And <strong>it prices GPUs at rental rates</strong>, which overstates the cost for the only company that matters here.</p><h3>Nobody buys only output tokens</h3><p>Real traffic has a shape. A coding agent sends tens of thousands of input tokens for every thousand it gets back, most of the input is a prefix it sent before, and DeepSeek charges for all three streams at three different prices. </p><p>A <strong>node has one time budget </strong>and has to split it between prefilling input and decoding output, so the sustainable output rate falls as the input ratio rises while total revenue climbs.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!q8Ak!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd8d8f28e-d3d1-4f3e-89a1-58f899f5d4e0_1600x894.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!q8Ak!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd8d8f28e-d3d1-4f3e-89a1-58f899f5d4e0_1600x894.png 424w, https://substackcdn.com/image/fetch/$s_!q8Ak!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd8d8f28e-d3d1-4f3e-89a1-58f899f5d4e0_1600x894.png 848w, https://substackcdn.com/image/fetch/$s_!q8Ak!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd8d8f28e-d3d1-4f3e-89a1-58f899f5d4e0_1600x894.png 1272w, https://substackcdn.com/image/fetch/$s_!q8Ak!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd8d8f28e-d3d1-4f3e-89a1-58f899f5d4e0_1600x894.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!q8Ak!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd8d8f28e-d3d1-4f3e-89a1-58f899f5d4e0_1600x894.png" width="1456" height="814" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d8d8f28e-d3d1-4f3e-89a1-58f899f5d4e0_1600x894.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:814,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;32K context, batch 1024, node time split between prefill and decode. Cache hits consume neither prefill nor decode, only a disk read.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="32K context, batch 1024, node time split between prefill and decode. Cache hits consume neither prefill nor decode, only a disk read." title="32K context, batch 1024, node time split between prefill and decode. Cache hits consume neither prefill nor decode, only a disk read." srcset="https://substackcdn.com/image/fetch/$s_!q8Ak!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd8d8f28e-d3d1-4f3e-89a1-58f899f5d4e0_1600x894.png 424w, https://substackcdn.com/image/fetch/$s_!q8Ak!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd8d8f28e-d3d1-4f3e-89a1-58f899f5d4e0_1600x894.png 848w, https://substackcdn.com/image/fetch/$s_!q8Ak!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd8d8f28e-d3d1-4f3e-89a1-58f899f5d4e0_1600x894.png 1272w, https://substackcdn.com/image/fetch/$s_!q8Ak!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd8d8f28e-d3d1-4f3e-89a1-58f899f5d4e0_1600x894.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>32K context, batch 1024, node time split between prefill and decode. Cache hits consume neither prefill nor decode, only a disk read.</em></p><p>Output-only pricing understates this card by roughly a factor of two. At a 20-to-1 input ratio with no caching, break-even is $55 per GPU-hour rather than $28. </p><p><strong>At 200-to-1 with a 90 percent hit rate, which is what a long-running agent on a stable repository actually looks like, it is $64</strong>. Every one of those numbers is an order of magnitude above the rental market.</p><p>Cache hits are the mechanism. They consume no prefill and no decode, only a disk read of <strong>compressed entries</strong> plus a 5,504-token recompute, so raising the hit rate raises sustainable throughput without raising cost. </p><p>The $0.0028 price is low because the marginal cost is low, and it is worth having because it makes the traffic denser.</p><h3>DeepSeek is not renting these GPUs</h3><p>A rental rate contains a lessor&#8217;s margin, a scarcity premium that has been visibly volatile all year, and the lessor&#8217;s own financing cost. DeepSeek owns its fleet, and an owner&#8217;s cost is amortisation plus power.</p><p>Take a B200 at roughly $38,000 of capex, three years of life, 80 percent utilisation: $1.81 per GPU-hour. Add a kilowatt at PUE 1.3 and eight cents a kilowatt-hour: ten cents. </p><p>Call it $1.91 per GPU-hour all in, against a rental band of $3.35 to $6.35. <strong>An owner&#8217;s floor sits somewhere between 1.8 and 3.3 times below the price of renting the same silicon</strong><mark data-color="rgb(244, 229, 189)" style="background-color: rgb(244, 229, 189); color: rgb(0, 0, 0);">.</mark></p><p>Note what power is and is not in that sum. A four-GPU node draws about 5.2 kW including overhead, which costs 42 cents an hour, which is under two percent of what the same node costs to rent. Power is not a meaningful line in the financial model. </p><p>It is a meaningful line in the physical one, and the report says so: the same paragraph that asks hardware vendors for a balance point also asks for more power headroom, because <strong>extreme kernel fusion drives compute</strong>, memory and network to high load simultaneously and power throttling becomes the limiter. Fusing everything is how you become thermally bound rather than bandwidth bound.</p><h3>Which changes the verdict on the million-token window</h3><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!fBIg!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc68cd709-ac92-4feb-8935-aab8909feb74_1600x838.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!fBIg!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc68cd709-ac92-4feb-8935-aab8909feb74_1600x838.png 424w, https://substackcdn.com/image/fetch/$s_!fBIg!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc68cd709-ac92-4feb-8935-aab8909feb74_1600x838.png 848w, https://substackcdn.com/image/fetch/$s_!fBIg!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc68cd709-ac92-4feb-8935-aab8909feb74_1600x838.png 1272w, https://substackcdn.com/image/fetch/$s_!fBIg!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc68cd709-ac92-4feb-8935-aab8909feb74_1600x838.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!fBIg!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc68cd709-ac92-4feb-8935-aab8909feb74_1600x838.png" width="1456" height="763" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c68cd709-ac92-4feb-8935-aab8909feb74_1600x838.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:763,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;The break-even line is what $0.28 per million pays at 1M context with the batch capped at 129 by the KV pool.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="The break-even line is what $0.28 per million pays at 1M context with the batch capped at 129 by the KV pool." title="The break-even line is what $0.28 per million pays at 1M context with the batch capped at 129 by the KV pool." srcset="https://substackcdn.com/image/fetch/$s_!fBIg!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc68cd709-ac92-4feb-8935-aab8909feb74_1600x838.png 424w, https://substackcdn.com/image/fetch/$s_!fBIg!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc68cd709-ac92-4feb-8935-aab8909feb74_1600x838.png 848w, https://substackcdn.com/image/fetch/$s_!fBIg!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc68cd709-ac92-4feb-8935-aab8909feb74_1600x838.png 1272w, https://substackcdn.com/image/fetch/$s_!fBIg!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc68cd709-ac92-4feb-8935-aab8909feb74_1600x838.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>The break-even line is what $0.28 per million pays at 1M context with the batch capped at 129 by the KV pool.</em></p><p>Serving a million-token context at $0.28 per million output tokens needs a cost basis under $2.92 per GPU-hour. Nobody renting Blackwell has one. DeepSeek does, with room.</p><blockquote><p><em>The million-token window is not a loss leader. It is a moat, and a precisely shaped one. </em><strong><span>Owner economics against rental economics</span></strong></p></blockquote><p>So the million-token window is not a loss leader. It is a moat, and a precisely shaped one. The twenty providers currently reselling the preview weights on <strong>OpenRouter at 37 percent below DeepSeek&#8217;s</strong> own card can do that at short context, where deep batches make the arithmetic work on rented hardware. </p><p>They <strong>structurally cannot do it at a million tokens</strong>, at that price, on rented Blackwell, because the KV pool caps the batch and the capped batch does not clear the rent. Which is presumably why none of them advertise it.</p><p>An open-weights MIT model whose most expensive capability is only economic for an owner-operator is a strange and rather elegant object. DeepSeek gave away the architecture, the <strong>mega-kernel and the kernel library,</strong> and kept the one thing that does not fit in a repository: a fleet, a request volume large enough to make the on-disk cache pay, and a cost basis a third of what anyone else can rent.</p><h3>What the whole card is betting on</h3><p>Put the three prices next to the three cost structures and the strategy is legible. Prefill is compute-bound and enormously profitable: at<strong> 35 percent MFU a node prefills roughly 485,000 tokens per second</strong>, which at $0.14 per million is $244 an hour against a node that costs $22 to rent and $8 to own. Decode is bandwidth-bound and needs deep batches. </p><p>Cache hits cost almost nothing and are priced almost at nothing, which makes them a throughput multiplier rather than a revenue line.</p><p>The card rewards exactly one traffic shape: high input-to-output ratio, high prefix reuse, deep concurrency, moderate context. </p><p>That is coding-agent and tool-calling traffic, which is what the 0731 post-training pass targeted, what the<em> native Responses API and Codex adaptation are for, and what Quick Instruction and interleaved thinking were built to make cheaper.</em> The architecture, the post-training, the serving stack and the price list are one artefact pointed at one workload.</p><p>Point a different workload at it, <strong>long single-turn document analysis with no prefix reuse</strong> and shallow concurrency, and the margin thins toward nothing even for DeepSeek. The card does not distinguish, yet. I do not expect that to last.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Reading the benchmarks in full</h2><p>The report&#8217;s summary of its own results is accurate and selective, and the difference is worth spending a section on.</p><p>Start with the base models, which are <strong>the cleanest comparison in the document </strong>because all three ran in one internal harness with identical settings. </p><p>V4-Flash-Base carries 13B activated against V3.2-Base&#8217;s 37B, and 284B total against 671B. It should lose. It mostly does not.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!GNaW!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10f7e37b-1aa1-4f20-893d-5c23df20b66a_1600x864.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!GNaW!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10f7e37b-1aa1-4f20-893d-5c23df20b66a_1600x864.png 424w, https://substackcdn.com/image/fetch/$s_!GNaW!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10f7e37b-1aa1-4f20-893d-5c23df20b66a_1600x864.png 848w, https://substackcdn.com/image/fetch/$s_!GNaW!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10f7e37b-1aa1-4f20-893d-5c23df20b66a_1600x864.png 1272w, https://substackcdn.com/image/fetch/$s_!GNaW!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10f7e37b-1aa1-4f20-893d-5c23df20b66a_1600x864.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!GNaW!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10f7e37b-1aa1-4f20-893d-5c23df20b66a_1600x864.png" width="1456" height="786" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/10f7e37b-1aa1-4f20-893d-5c23df20b66a_1600x864.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:786,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Report Table 1. Same harness, same settings. Flash carries 35 percent of V3.2's activated parameters and 42 percent of its total.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Report Table 1. Same harness, same settings. Flash carries 35 percent of V3.2's activated parameters and 42 percent of its total." title="Report Table 1. Same harness, same settings. Flash carries 35 percent of V3.2's activated parameters and 42 percent of its total." srcset="https://substackcdn.com/image/fetch/$s_!GNaW!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10f7e37b-1aa1-4f20-893d-5c23df20b66a_1600x864.png 424w, https://substackcdn.com/image/fetch/$s_!GNaW!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10f7e37b-1aa1-4f20-893d-5c23df20b66a_1600x864.png 848w, https://substackcdn.com/image/fetch/$s_!GNaW!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10f7e37b-1aa1-4f20-893d-5c23df20b66a_1600x864.png 1272w, https://substackcdn.com/image/fetch/$s_!GNaW!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10f7e37b-1aa1-4f20-893d-5c23df20b66a_1600x864.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>Report Table 1. Same harness, same settings. Flash carries 35 percent of V3.2&#8217;s activated parameters and 42 percent of its total.</em></p><blockquote><p>Plus 6.8 on FACTS Parametric, plus 6.7 on HumanEval, plus 4.5 on LongBench-V2, plus 3.5 on MultiLoKo, plus 2.8 on MMLU-Pro. </p></blockquote><p>Getting more world knowledge out of fewer total parameters is the surprising one, because knowledge retention is supposed to scale with parameter count and Flash has 42 percent of them.</p><p>And then two regressions that nobody discussing this model has mentioned. <strong>BigCodeBench falls 7.1 points, from 63.9 to 56.8</strong>. MATH falls 3.1, from 60.5 to 57.4. Those are not noise, they are the two largest deltas in the table after FACTS, and one of them is a coding benchmark on a model being marketed for coding agents. </p><p><strong>Post-training clearly recovers a great deal of it</strong>, since the instructed model&#8217;s code-agent scores are strong. But the base model is worse at BigCodeBench than its predecessor and the report does not discuss why.</p><p>My guess, and it is a guess: BigCodeBench is a three-shot benchmark on library-heavy code, which is a knowledge task about API surfaces rather than a reasoning task, and it is the <strong>sort of long-tail recall that a smaller expert bank should hurt. </strong></p><p>The world-knowledge gains going the other way argue against that, which is why it stays a guess.</p><p>Then the long-context claim. The report&#8217;s own summary says V4-Pro-Max delivers strong results with a one-million-token window, <em>&#8220;surpassing even Gemini-3.1-Pro on academic benchmarks.&#8221; </em>That is totally true.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!q6i3!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a4cc0b7-6824-44c9-a006-558630822962_1600x718.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!q6i3!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a4cc0b7-6824-44c9-a006-558630822962_1600x718.png 424w, https://substackcdn.com/image/fetch/$s_!q6i3!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a4cc0b7-6824-44c9-a006-558630822962_1600x718.png 848w, https://substackcdn.com/image/fetch/$s_!q6i3!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a4cc0b7-6824-44c9-a006-558630822962_1600x718.png 1272w, https://substackcdn.com/image/fetch/$s_!q6i3!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a4cc0b7-6824-44c9-a006-558630822962_1600x718.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!q6i3!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a4cc0b7-6824-44c9-a006-558630822962_1600x718.png" width="1456" height="653" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/9a4cc0b7-6824-44c9-a006-558630822962_1600x718.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:653,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Report Table 6, DeepSeek's own numbers, standardised configuration across models.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Report Table 6, DeepSeek's own numbers, standardised configuration across models." title="Report Table 6, DeepSeek's own numbers, standardised configuration across models." srcset="https://substackcdn.com/image/fetch/$s_!q6i3!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a4cc0b7-6824-44c9-a006-558630822962_1600x718.png 424w, https://substackcdn.com/image/fetch/$s_!q6i3!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a4cc0b7-6824-44c9-a006-558630822962_1600x718.png 848w, https://substackcdn.com/image/fetch/$s_!q6i3!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a4cc0b7-6824-44c9-a006-558630822962_1600x718.png 1272w, https://substackcdn.com/image/fetch/$s_!q6i3!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a4cc0b7-6824-44c9-a006-558630822962_1600x718.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>Report Table 6, DeepSeek&#8217;s own numbers, standardised configuration across models.</em></p><p>It is also the smaller half of the story. On LongMRCR at 1M, Claude Opus 4.6 scores 92.9, V4-Pro-Max 83.5, Gemini-3.1-Pro 76.3. On CorpusQA at 1M, Opus 71.7, V4-Pro-Max 62.0, Gemini 53.8. DeepSeek beats Gemini on both, comfortably, and <strong>loses to Opus by 9.4 and 9.7 points, in their own table, measured with their own harness</strong>. </p><p>Choosing Gemini as the comparison in the summary is a defensible framing decision and it is a framing decision.</p><p>The report is honest in other places where it did not have to be. It <strong>reports Terminal-Bench 2.0 at 67.9</strong> on the original dataset while noting environment issues raised by another lab and disclosing that on the Verified subset V4-Pro scores about 72.0, which is higher. </p><p>It leaves cells blank for K2.6 and GLM-5.1 rather than filling them, saying those APIs were too busy to answer. It states that <strong>its reasoning performance trails the frontier by roughly three to six months.</strong> </p><p>That last sentence is in the report&#8217;s own summary of its own results and it is a more useful number than most third-party analysis of the same question.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/p/deepseek-v4-flash-the-cost-of-deciding?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/p/deepseek-v4-flash-the-cost-of-deciding?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><div><hr></div><h2>What this does to everyone else&#8217;s floor</h2><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!9O26!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F63f0fd58-5b62-4abd-9e19-078a2c22b67d_1600x864.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!9O26!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F63f0fd58-5b62-4abd-9e19-078a2c22b67d_1600x864.png 424w, https://substackcdn.com/image/fetch/$s_!9O26!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F63f0fd58-5b62-4abd-9e19-078a2c22b67d_1600x864.png 848w, https://substackcdn.com/image/fetch/$s_!9O26!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F63f0fd58-5b62-4abd-9e19-078a2c22b67d_1600x864.png 1272w, https://substackcdn.com/image/fetch/$s_!9O26!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F63f0fd58-5b62-4abd-9e19-078a2c22b67d_1600x864.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!9O26!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F63f0fd58-5b62-4abd-9e19-078a2c22b67d_1600x864.png" width="1456" height="786" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/63f0fd58-5b62-4abd-9e19-078a2c22b67d_1600x864.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:786,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;List rates, August 2026. The gap to the cheapest Western flagship is more than two orders of magnitude.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="List rates, August 2026. The gap to the cheapest Western flagship is more than two orders of magnitude." title="List rates, August 2026. The gap to the cheapest Western flagship is more than two orders of magnitude." srcset="https://substackcdn.com/image/fetch/$s_!9O26!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F63f0fd58-5b62-4abd-9e19-078a2c22b67d_1600x864.png 424w, https://substackcdn.com/image/fetch/$s_!9O26!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F63f0fd58-5b62-4abd-9e19-078a2c22b67d_1600x864.png 848w, https://substackcdn.com/image/fetch/$s_!9O26!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F63f0fd58-5b62-4abd-9e19-078a2c22b67d_1600x864.png 1272w, https://substackcdn.com/image/fetch/$s_!9O26!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F63f0fd58-5b62-4abd-9e19-078a2c22b67d_1600x864.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>List rates, August 2026. The gap to the cheapest Western flagship is more than two orders of magnitude.</em></p><p>Reuters, reporting <strong>Artificial Analysis figures,</strong> put V4-Flash at roughly 3 cents to complete the Intelligence Index battery, against 86 cents for Kimi K3, $1.86 for GPT-5.6 Sol and $3.15 for Claude Fable 5. </p><p>On the index itself Flash scores 50, tying Gemini 3.6 Flash, one point behind <strong>GLM-5.2 and Muse Spark 1.1</strong>, seven behind Kimi K3, and nine or more behind Opus 5, Fable 5 and GPT-5.6.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!egWI!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6e2f797-4468-4f8c-bf30-47125bab71f8_1600x836.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!egWI!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6e2f797-4468-4f8c-bf30-47125bab71f8_1600x836.png 424w, https://substackcdn.com/image/fetch/$s_!egWI!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6e2f797-4468-4f8c-bf30-47125bab71f8_1600x836.png 848w, https://substackcdn.com/image/fetch/$s_!egWI!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6e2f797-4468-4f8c-bf30-47125bab71f8_1600x836.png 1272w, https://substackcdn.com/image/fetch/$s_!egWI!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6e2f797-4468-4f8c-bf30-47125bab71f8_1600x836.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!egWI!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6e2f797-4468-4f8c-bf30-47125bab71f8_1600x836.png" width="1456" height="761" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b6e2f797-4468-4f8c-bf30-47125bab71f8_1600x836.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:761,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Artificial Analysis, August 2026. The vertical axis is nine points wide. The horizontal axis is two orders of magnitude.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Artificial Analysis, August 2026. The vertical axis is nine points wide. The horizontal axis is two orders of magnitude." title="Artificial Analysis, August 2026. The vertical axis is nine points wide. The horizontal axis is two orders of magnitude." srcset="https://substackcdn.com/image/fetch/$s_!egWI!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6e2f797-4468-4f8c-bf30-47125bab71f8_1600x836.png 424w, https://substackcdn.com/image/fetch/$s_!egWI!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6e2f797-4468-4f8c-bf30-47125bab71f8_1600x836.png 848w, https://substackcdn.com/image/fetch/$s_!egWI!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6e2f797-4468-4f8c-bf30-47125bab71f8_1600x836.png 1272w, https://substackcdn.com/image/fetch/$s_!egWI!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6e2f797-4468-4f8c-bf30-47125bab71f8_1600x836.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>Artificial Analysis, August 2026. The vertical axis is nine points wide. The horizontal axis is two orders of magnitude.</em></p><p>The standard rebuttal to a cheap model is that it burns more tokens, so cost per task closes the gap that cost per token opens. It does not work here and the numbers say so precisely. </p><p><strong>Flash generated 210M output tokens on the index against a 100M median</strong>, so it is 2.1 times more verbose than the typical model on identical work, for the documented reason that its Think Max system prompt instructs it to be. It is still 105 times cheaper per task than Fable 5, whose output token costs 179 times more. </p><p>Doubling token count against a 179x price advantage leaves 89x. The measured 105x is in that neighbourhood, and <strong>the verbosity tax is nowhere near large enough to matter</strong>.</p><p>What follows for the market is narrower than the headline suggests.</p><p>The nine-point index gap is not a rounding error. It is the <em>difference between a model that finishes a hard task and one that plausibly fails it, </em>and on hard agentic work a failed run costs more than the token price of a successful one. </p><p>At the top of the market price is not the binding constraint, and a 105x discount on a wrong answer is not a discount.</p><blockquote><p><em>A 105x discount on a wrong answer is not a discount.</em><strong><span>On where the price pressure actually lands</span></strong></p></blockquote><p>The pressure is on the middle. Every workload running on a flagship because nobody bothered to route it, every classification and extraction and summarisation and <strong>first-draft call</strong>, is now paying somewhere between 60 and 105 times more than it needs to. </p><p>Routing is the mechanism that transfers that value and routing is engineering work most teams have not done. The models that should be nervous are not Opus 5 and Fable 5. </p><p>They are Haiku, Sonnet, the Flash and Terra and Luna tiers, <strong>everything between $1 and $5 per million output tokens</strong> where the capability gap to V4-Flash is small or negative and the price gap is 20x or more.</p><p><strong>Two second-order effects </strong>are worth flagging. OpenRouter currently lists twenty providers serving the preview weights at $0.088 in and $0.176 out, 37 percent below DeepSeek&#8217;s own card. </p><p>That is an MIT-licensed model being resold below the price its author charges. Given the<em> break-even analysis above</em>, those hosts are running deeper batches, or cheaper capacity, or buying share. Only the first is durable.</p><p>And the weights are MIT, the mega-kernel is open inside DeepGEMM, the kernel library is open, and the fine-grained expert-parallel scheme was validated on Huawei Ascend as well as NVIDIA. <strong>The architecture is not the moat. The traffic shape is</strong>, and so is the on-disk cache that only pays off at DeepSeek&#8217;s request volume.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Where we might be wrong</h2><p>The CSA and HCA layer split for Flash is inferred. The report says the first two layers are pure sliding window and the rest interleave, which for 41 remaining layers gives either 21 CSA and 20 HCA or the reverse. </p><p><strong>We assumed the interleave opens with CSA.</strong> If it opens with HCA the cache figure moves from 3.372 to 3.216 GiB, about 5 percent, the parameter reconstruction moves by 13M, and nothing in the argument changes.</p><p>The B200 configuration is nameplate: 180 GB usable per GPU, 8 TB/s, four GPUs, 85 percent HBM utilisation, and a 63 percent achieved bandwidth fraction calibrated against a single third-party throughput measurement. </p><p>A production deployment with prefill and decode disaggregated across separate node pools has <strong>different and probably better economics than the single-pool model we used</strong>, and DeepSeek describes that disaggregation themselves.</p><p>The owner-cost figure of $1.91 per GPU-hour is built from a $38,000 capex assumption, a three-year life and 80 percent utilisation, none of which DeepSeek publishes. Capex at $30,000 gives $1.53 and at $45,000 gives $2.24, so the <strong>conclusion that an owner clears the 1M break-even survives the range</strong>, but only just at the top of it. The utilisation assumption is the fragile one: at 50 percent utilisation the figure is $2.99 and the verdict flips.</p><p>The blended-revenue model splits node time between prefill and decode on a single pool. Real serving disaggregates them across separate node pools with different shapes, which changes the split and generally improves it. </p><p><strong>It also assumes cache hits consume no node time beyond the disk read</strong>, which understates their cost at high hit rates where storage bandwidth starts to bind.</p><p>The FLOPs reconstruction lands at 12.2 percent of V3.2 at 1M against the report&#8217;s 10 percent. The report measures in equivalent FP8 FLOPs while <strong>V4&#8217;s experts run FP4</strong>, which has identical peak throughput on current silicon but which the report says could be a third cheaper on future hardware. </p><p>That accounting difference is the likely source. My V3.2 indexer model may also be too generous.</p><p>The prefill MFU of 35 percent is an assumption, not a measurement, and both prefill revenue and the blended break-even scale linearly with it. At <strong>20 percent the prefill figure drops from $244 an hour to $140</strong> and every blended break-even in the traffic-mix chart falls by roughly a third.</p><p>Our reading of <strong>BigCodeBench&#8217;s regression as an API-recall effect</strong> is speculation and I have flagged it as such in the text. I would drop it entirely if the world-knowledge results did not cut the other way.</p><p>And the largest one. Every agent benchmark in the 0731 release is vendor-reported, measured with a <strong>harness DeepSeek has announced but not shipped</strong>, at maximum reasoning effort, with sampling parameters DeepSeek chose. Two of the nine suites are DeepSeek&#8217;s own internal sets. I have reproduced none of them and neither has anyone else.</p><h2>Seven predictions, dated</h2><ol><li><p><strong>By June 2027</strong>, at least one major Western lab ships a production model that compresses the KV cache along the sequence axis rather than the head axis. The mechanism is cheap, the memory win is fifty times, and it has now been demonstrated at 32T tokens of pre-training rather than in an ablation.</p></li><li><p><strong>By 31 December 2026</strong>, someone publishes a sub-linear top-k selector for compressed attention, most likely hierarchical or learned-hash, and demonstrates it on V4&#8217;s open weights. The indexer being four fifths of attention cost at long context is too visible a target, and the FlashMemory work has already aimed at it from the prefetch side.</p></li><li><p><strong>By March 2027</strong>, DeepSeek introduces context-length-tiered pricing, a peak-hour surcharge that actually activates, or both. A flat rate across a thirty times span of serving cost is not stable, and a surcharge has already been announced without being switched on.</p></li><li><p><strong>By September 2027</strong>, at least one of Anthropic, OpenAI or Google cuts a mid-tier model&#8217;s output price by more than 50 percent without a corresponding capability release. The squeeze is on the middle of the ladder, not the top.</p></li><li><p><strong>By the end of December 2027</strong>, DeepSeek-V5 or its equivalent replaces SwiGLU with a gate-free elementwise activation. The report asks for this explicitly in its hardware proposals, gives the reason (removing the gate projection lets the intermediate dimension grow under a fixed budget, which raises the interconnect balance point), and labs that publish that kind of request are usually describing work already underway.</p></li><li><p><strong>By late June 2027</strong>, an SMT or ILP solver appears in the compilation pipeline of at least one other major inference stack. Z3 inside TileLang&#8217;s algebraic system is the first production instance I know of, the payoff is a few seconds of compile time for stronger vectorisation and bound analysis, and the idea travels.</p></li><li><p><strong>By 31 December 2027</strong>, no model with a one-million-token context window is profitable at that context length on <em>rented</em> NVIDIA hardware at published rates, while remaining profitable for owner-operators. The KV pool caps concurrency, concurrency is what pays for decode, and the gap between owning and renting is wider than the margin. This is the prediction I expect to age worst, and it fails if HBM capacity per package jumps faster than I think.</p></li></ol><div class="directMessage button" data-attrs="{&quot;userId&quot;:201922174,&quot;userName&quot;:&quot;Lorenzo Bradanini&quot;,&quot;canDm&quot;:null,&quot;dmUpgradeOptions&quot;:null,&quot;isEditorNode&quot;:true}" data-component-name="DirectMessageToDOM"></div><div><hr></div><h2>Confidence dossier</h2><p><em><strong>Tier A &#183; Stated in a primary source &#183; 13 claims</strong></em></p><ul><li><p><em><strong>43 layers, d=4096, m=4, m&#8217;=128, top-k 512, 256+1 experts, 6 activated</strong><br><span>Report section 4.2.1, stated directly</span></em></p></li><li><p><em><strong>Partial RoPE on exactly the last 64 dimensions; inverse RoPE at position -i on the output</strong><br><span>Report section 2.3.3, stated directly</span></em></p></li><li><p><em><strong>HCA has one KV stream and no indexer; CSA has two overlapping streams</strong><br><span>Report equations 9 to 12 versus 20 to 23</span></em></p></li><li><p><em><strong>Attention sink logits per head, added to the softmax denominator</strong><br><span>Report equation 27</span></em></p></li><li><p><em><strong>Hybrid Newton-Schulz, 8 steps at (3.4445, -4.7750, 2.0315) then 2 at (2, -1.5, 0.5)</strong><br><span>Report section 2.4</span></em></p></li><li><p><em><strong>No QK-Clip, because RMSNorm on queries and KV entries already bounds the logits</strong><br><span>Report section 2.4, stated as a deliberate omission</span></em></p></li><li><p><em><strong>Dense attention for the first 1T tokens; sparsity introduced at 64K with an indexer warmup</strong><br><span>Report section 4.2.2</span></em></p></li><li><p><em><strong>Index scores quantised FP32 to BF16: 2x on top-k, 99.7 percent KV recall preserved</strong><br><span>Report section 3.4</span></em></p></li><li><p><em><strong>Interconnect condition C/B &lt;= 2 d_ff; 6144 FLOPs per byte for Pro</strong><br><span>Report section 3.1, derived and evaluated there</span></em></p></li><li><p><em><strong>Three on-disk SWA strategies; zero-caching needs n_win x L tokens of recompute</strong><br><span>Report section 3.6.2</span></em></p></li><li><p><em><strong>Quick Instruction tokens reuse the existing KV cache to avoid a second prefill</strong><br><span>Report section 5.1.1, Table 5</span></em></p></li><li><p><em><strong>0731 is the same structure and size as the preview; post-training only</strong><br><span>DeepSeek API changelog, 31 July 2026</span></em></p></li><li><p><em><strong>Rate card $0.14 / $0.0028 / $0.28 per million</strong><br><span>Artificial Analysis and DeepSeek changelog, mutually consistent</span></em></p></li></ul><p><em><strong>Tier B &#183; Derived here and cross-checked &#183; 8 claims</strong></em></p><ul><li><p><em><strong>Backbone reconstructs to 284.20B; MTP module excluded from the headline</strong><br><span>Derived here from published constants, 0.07 percent residual</span></em></p></li><li><p><em><strong>Attention is 1.74 percent of weights and 39.3 percent of activated</strong><br><span>Derived here; follows from the reconstruction</span></em></p></li><li><p><em><strong>mHC GEMM output dimension 24 = n_hc + n_hc squared + n_hc</strong><br><span>Derived here; matches the figure quoted in report section 3.3</span></em></p></li><li><p><em><strong>3.372 GiB of KV per 1M sequence at production precision</strong><br><span>Derived here; four independent cross-checks, worst 10 percent</span></em></p></li><li><p><em><strong>Indexer is 79 percent of attention FLOPs at 1M, 50 percent at 32K</strong><br><span>Derived here from published dimensions</span></em></p></li><li><p><em><strong>Attention overtakes the expert bank at roughly 213K tokens</strong><br><span>Derived here; sensitive to the CSA/HCA split assumption</span></em></p></li><li><p><em><strong>Flash&#8217;s balance point is 4096 FLOPs per byte, 1.5x more demanding than Pro</strong><br><span>Derived here by applying the report&#8217;s own condition to Flash&#8217;s d_ff</span></em></p></li><li><p><em><strong>5,504-token recompute restores the full SWA state, a 191x saving on prefill</strong><br><span>Derived here from n_win, L and the activated parameter count</span></em></p></li></ul><p><em><strong>Tier C &#183; Model output, assumptions named &#183; 7 claims</strong></em></p><ul><li><p><em><strong>129 concurrent 1M sessions on one 4xB200 node</strong><br><span>Model; depends on 180 GB usable and 85 percent utilisation</span></em></p></li><li><p><em><strong>63 percent of peak HBM bandwidth achieved at low batch</strong><br><span>Inferred from one third-party throughput measurement</span></em></p></li><li><p><em><strong>1M context clears at owner cost and not at any rental rate</strong><br><span>Model output; hinges on the $38k / 3yr / 80 percent capex assumption</span></em></p></li><li><p><em><strong>Blended break-even is 32 to 64 dollars per GPU-hour on agentic traffic mixes</strong><br><span>Model output; single-pool prefill and decode, 35 percent prefill MFU</span></em></p></li><li><p><em><strong>Owner cost near $1.91 per GPU-hour against a $3.35 to $6.35 rental band</strong><br><span>Model output; capex and utilisation assumed, power from published TDP</span></em></p></li><li><p><em><strong>Break-even needs batch 256 at 32K against a $5.50 GPU-hour</strong><br><span>Model output; single-pool assumption, no disaggregation</span></em></p></li><li><p><em><strong>Prefill grosses $244 an hour against a $22 node</strong><br><span>Model output; 35 percent MFU is assumed, not measured</span></em></p></li></ul><p><em><strong>Tier D &#183; Unreproduced or speculative &#183; 3 claims</strong></em></p><ul><li><p><em><strong>BigCodeBench regression is a long-tail API recall effect</strong><br><span>Speculation, contradicted by the world-knowledge results</span></em></p></li><li><p><em><strong>Agent benchmark figures for the 0731 release</strong><br><span>Vendor reported, unshipped harness, two internal suites</span></em></p></li><li><p><em><strong>OpenRouter hosts undercutting DeepSeek by 37 percent are unprofitable</strong><br><span>Speculative; their batch depth and capacity costs are unknown</span></em></p></li></ul><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/p/deepseek-v4-flash-the-cost-of-deciding?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/p/deepseek-v4-flash-the-cost-of-deciding?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><div><hr></div><h2>Reproducing this</h2><p>Everything numerical above comes from three short scripts with no dependencies beyond the standard library. None of it needs a GPU, because none of it is a measurement of a running model. </p><p>It is arithmetic over published constants, validated against figures published independently by vLLM and against ratios stated in the report itself.</p><p>The parameter reconstruction. Note the asymmetry between CSA and HCA in the projection count, which is the thing that is easy to get wrong:</p><pre><code>L, d = 43, 4096
n_h, c, d_c = 64, 512, 1024          # query heads, head dim, query compression dim
n_hI, c_I   = 64, 128                # indexer heads, indexer head dim
g, d_g      = 8, 1024                # output projection groups
m, mp       = 4, 128                 # CSA and HCA compression rates
n_exp, n_sh, n_act, d_ff = 256, 1, 6, 2048
n_hc, V     = 4, 129280
n_swa, n_csa, n_hca = 2, 21, 20      # 2 SWA + 41 interleaved

per_expert = 3 * d * d_ff                        # SwiGLU: gate, up, down
moe_tot, moe_act = (n_exp+n_sh)*per_expert, (n_act+n_sh)*per_expert

def attn(kind):
    if   kind == &#8216;csa&#8217;: p = 4*d*c + 2*m*c        # two KV streams + positional biases
    elif kind == &#8216;hca&#8217;: p = 2*d*c + mp*c         # one KV stream  + positional bias
    else:               p = 1*d*c                # pure SWA, uncompressed
    p += d*d_c + d_c*(c*n_h)                     # query down then up
    if kind == &#8216;csa&#8217;:
        p += d_c*(c_I*n_hI) + d*n_hI             # indexer queries and per-head gate
    p += g*((n_h//g)*c)*d_g + (g*d_g)*d          # grouped output projection
    return p + n_h                               # attention sink logits

mhc = 2*(n_hc*d)*n_hc + (n_hc*d)*(n_hc**2)       # W_pre, W_post, W_res
att = n_swa*attn(&#8217;swa&#8217;) + n_csa*attn(&#8217;csa&#8217;) + n_hca*attn(&#8217;hca&#8217;)
backbone = L*moe_tot + att + 2*L*mhc + L*d*n_exp + 2*V*d
active   = L*moe_act + att + 2*L*mhc + L*d*n_exp

print(backbone/1e9, active/1e9)                  # 284.202  12.610
print(n_hc + n_hc**2 + n_hc)                     # 24, matching the mHC GEMM in section 3.3</code></pre><p>The KV byte model, with all four calibration checks it has to pass before being used:</p><pre><code>GiB, N = 1024**3, 1_048_576
c, c_I = 512, 128

def layer(kind, ent, idx, m=4, mp=128, n_win=128):
    if kind == &#8216;c4a&#8217;:   return (N//m)  * (ent + idx)
    if kind == &#8216;c128a&#8217;: return (N//mp) * ent
    if kind == &#8216;swa&#8217;:   return n_win * ent

# check 1: V3.2 bf16, MLA 512 latent + 64 rope, indexer 128, 61 layers
assert abs(61*(576*2 + 128*2)*N/GiB - 83.9) &lt; 0.1              # vLLM publishes 83.9

# check 2: V4-Pro bf16, 30 c4a + 31 c128a, key and value shared
pro = 30*layer(&#8217;c4a&#8217;, c*2, c_I*2) + 31*layer(&#8217;c128a&#8217;, c*2, 0)
assert abs(pro/GiB - 9.62) &lt; 0.01                              # vLLM publishes 9.62

# production precision: 64 rope dims bf16 + 448 fp8 = 576 B; fp4 indexer = 64 B
ent, idx = 64*2 + (c-64)*1, c_I//2
flash = 21*layer(&#8217;c4a&#8217;, ent, idx) + 20*layer(&#8217;c128a&#8217;, ent, 0) + 43*layer(&#8217;swa&#8217;, ent, 0)
v32   = 61*(64*2 + 512 + 128)*N

# check 3: the report&#8217;s own Figure 1 ratio
print(v32/flash)                       # 13.57  vs the report&#8217;s &#8220;13.7x smaller&#8221;

# check 4: the report&#8217;s claim that uncompressed SWA state is ~8x the compressed state
print(43*N*ent / (flash - 43*layer(&#8217;swa&#8217;, ent, 0)))            # 7.2  vs &#8220;approximately 8&#8221;

print(flash/GiB, flash/N)              # 3.372 GiB, 3453 bytes per context token
print(100*flash / (43*(2*8*128*2)*N))  # 1.96 percent of bf16 GQA-8, report says ~2</code></pre><p>The decode roofline and break-even, for substituting your own hardware and rates:</p><pre><code>NG, BW = 4, 8.0e12                                   # four B200, 8 TB/s each
w_step  = 278.1e9*0.5 + NG*6.2e9                     # FP4 experts + FP8 dense per rank
t_floor = w_step / (NG*BW)                           # 5.12 ms
eff     = t_floor / (1/122.7)                        # 0.63, from the measured 122.7 tok/s

def kv_bytes(ctx, ent=576, idx=64):
    return 21*((512+128)*ent + (ctx//4)*idx) + 20*(max(1, ctx//128)*ent + 128*ent)

def breakeven(batch, ctx, price=0.28):
    t = ((w_step + batch*kv_bytes(ctx)) / (NG*BW)) / eff
    return (batch/t) * 3600/1e6 * price / NG         # dollars per GPU-hour

print(breakeven(512, 32768))      # 14.76  clears the market comfortably
print(breakeven(256, 32768))      #  7.64  clears a mid-market B200
print(breakeven(129, 1048576))    #  2.93  clears nothing you can rent

# the report&#8217;s interconnect condition, applied to Flash
for name, d_ff in ((&#8221;Flash&#8221;, 2048), (&#8221;Pro&#8221;, 3072)):
    print(name, 2*d_ff, &#8220;FLOP/Byte -&gt;&#8221;, 4.5e15/(2*d_ff)/1e12, &#8220;TB/s at 4.5 PFLOP/s FP8&#8221;)</code></pre><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Sources</h2><ul><li><p><em>DeepSeek-AI. DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence. arXiv:2606.19348, 26 April 2026. Sections 2.2 to 2.4 for architecture, 3.1 for expert parallelism and the interconnect balance point, 3.2 for TileLang, 3.3 for batch invariance and determinism, 3.4 for FP4 QAT, 3.5 for the training framework, 3.6 for KV cache management and on-disk storage, 4.2 for the constants and the training schedule, 5.1 for post-training, 5.2 for RL infrastructure and DSec, 5.3 for evaluation.</em></p></li><li><p><em>vLLM Team. DeepSeek V4 in vLLM: Efficient Long-context Attention, 24 April 2026. Appendix contains the KV arithmetic used here for calibration, the derivation of why inverse RoPE is needed when key and value are shared, and the exact top-k values for c4a and c128a.</em></p></li><li><p><em>LMSYS. DeepSeek-V4 on Day 0: From Fast Inference to Verified RL with SGLang and Miles, 25 April 2026, for the MXFP8 by MXFP4 fused MoE path and the split-K mHC pre-GEMM kernel.</em></p></li><li><p><em>DeepSeek API changelog, 31 July 2026, for the 0731 release scope, the agent benchmark table and the harness settings.</em></p></li><li><p><em>Artificial Analysis, model page for DeepSeek V4 Flash 0731, accessed 5 August 2026, for Intelligence Index v4.1, output speed, time to first token, token volume and the rate card.</em></p></li><li><p><em>Reuters, via Quartz and Business Standard, 3 August 2026, for the cost-per-task comparison across V4-Flash, Kimi K3, GPT-5.6 Sol and Claude Fable 5.</em></p></li><li><p><em>Hugging Face model cards for <span>deepseek-ai/DeepSeek-V4-Flash</span> and <span>deepseek-ai/DeepSeek-V4-Pro</span>, for the mixed FP4 and FP8 weight format and the reference inference implementation.</em></p></li><li><p><em>DeepGEMM pull request 304, for the open-sourced MegaMoE fused expert-parallel kernel.</em></p></li><li><p><em>FlashMemory-DeepSeek-V4: Lightning Index Ultra-Long Context via Lookahead Sparse Attention, arXiv:2606.09079, for the Neural Memory Indexer follow-up and its own account of its limitations.</em></p></li><li><p><em>getdeploying.com B200 index (4 August 2026), gpuprice.fyi B200 index (31 July 2026) and published neocloud rate cards, for the $3.35 to $6.35 band.</em></p></li><li><p><em>Anthropic, OpenAI, Moonshot and Z.ai published rate cards as of 4 August 2026, for the price ladder.</em></p></li></ul>]]></content:encoded></item><item><title><![CDATA[Invite your friends to read The Software Frontier]]></title><description><![CDATA[A warm thank you]]></description><link>https://www.thesoftwarefrontier.com/p/invite-your-friends-to-read-the-software</link><guid isPermaLink="false">https://www.thesoftwarefrontier.com/p/invite-your-friends-to-read-the-software</guid><dc:creator><![CDATA[Lorenzo Bradanini]]></dc:creator><pubDate>Tue, 04 Aug 2026 06:26:25 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!SAY7!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F54550d86-2756-4131-8818-956604f6749d_608x608.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2>A warm thank you</h2><p>Thanks so much for reading <a href="https://www.thesoftwarefrontier.com/">The Software Frontier</a> ! Your huge support keeps us motivated to do this work.</p><p>If you do really enjoy <strong>The Software Frontier</strong>, we would be really happy if you invited friends to subscribe and read with us. If you refer friends, you will receive a few benefits that give you special access to The Software Frontier.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>How to participate </h2><p><strong>1. Share The Software Frontier. </strong>When you use the referral link below, or the &#8220;Share&#8221; button on any post, you'll get credit for any new subscribers. Simply send the link in a text, email, or share it on social media with friends.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/leaderboard?&amp;utm_source=post&quot;,&quot;text&quot;:&quot;Refer a friend&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/leaderboard?&amp;utm_source=post"><span>Refer a friend</span></a></p><p>2.<strong> Earn benefits.</strong> When more friends use your referral link to subscribe (free or paid), you&#8217;ll receive special benefits.</p><ul><li><p>Get a 1 month comp for 3 referrals</p></li><li><p>Get a 3 month comp for 5 referrals</p></li><li><p>Get a 6 month comp for 25 referrals</p></li></ul><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/leaderboard?&amp;utm_source=post&quot;,&quot;text&quot;:&quot;Visit the leaderboard&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/leaderboard?&amp;utm_source=post"><span>Visit the leaderboard</span></a></p><p>To learn more, check out <a href="https://support.substack.com/hc/en-us/articles/16142857300372">Substack&#8217;s FAQ</a>.</p><p>Thank you for helping get the word out about The Software Frontier! Without you, none of this could have been made possible. </p><p>Lorenzo Bradanini and Lorenzo Tettamanti. </p>]]></content:encoded></item><item><title><![CDATA[How Blackwell’s Tensor Memory Actually Works ]]></title><description><![CDATA[Blackwell's largest matrix instruction needs 256 registers per thread. The ceiling is 255.]]></description><link>https://www.thesoftwarefrontier.com/p/how-blackwells-tensor-memory-actually</link><guid isPermaLink="false">https://www.thesoftwarefrontier.com/p/how-blackwells-tensor-memory-actually</guid><dc:creator><![CDATA[Lorenzo Bradanini]]></dc:creator><pubDate>Mon, 03 Aug 2026 15:07:50 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!zfIO!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4a8ed67-c0a5-4192-8a9c-b5a9b3edd44c_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!zfIO!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4a8ed67-c0a5-4192-8a9c-b5a9b3edd44c_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!zfIO!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4a8ed67-c0a5-4192-8a9c-b5a9b3edd44c_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!zfIO!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4a8ed67-c0a5-4192-8a9c-b5a9b3edd44c_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!zfIO!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4a8ed67-c0a5-4192-8a9c-b5a9b3edd44c_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!zfIO!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4a8ed67-c0a5-4192-8a9c-b5a9b3edd44c_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!zfIO!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4a8ed67-c0a5-4192-8a9c-b5a9b3edd44c_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e4a8ed67-c0a5-4192-8a9c-b5a9b3edd44c_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2322458,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/208027653?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4a8ed67-c0a5-4192-8a9c-b5a9b3edd44c_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!zfIO!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4a8ed67-c0a5-4192-8a9c-b5a9b3edd44c_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!zfIO!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4a8ed67-c0a5-4192-8a9c-b5a9b3edd44c_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!zfIO!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4a8ed67-c0a5-4192-8a9c-b5a9b3edd44c_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!zfIO!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4a8ed67-c0a5-4192-8a9c-b5a9b3edd44c_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>For nine years the register file per SM didn&#8217;t grow of a byte, while tensor throughput per clock doubled almost four times. Blackwell&#8217;s answer was to take out the accumulator out of the register file, and to put it in an address space with its own allocator, its own barrier, and no coherence with anything. This is what that memory is, what the compiler actually emits for it, and the invariance that explains why it had to exist.</em></p><h2>CUDA Mastery 2026</h2><p>This article is a narrow part of one architecture. The <strong>actual guide</strong> is the wide one: 34 chapters and thousands words on CUDA 13.x, from memory model up through <strong>Hopper</strong> and <strong>Blackwell tensor cores</strong>, written the same way as everything here. Compile it, disassemble it, show the command, then explain what happened.</p><p>It also has a corrections page at the end. An early edition listed H100 sparse tensor rates as dense ones, and that single mislabel propagated into three roofline figures before a reader caught it. </p><p>That figures are fixed and the mistake is documented rather than quietly removed, which is <strong>the standard </strong>we would want from anyone selling us a technical book.</p><p><a href="https://lorenzobrada.gumroad.com/l/cuda_mastery">Get CUDA Mastery guide </a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share The Software Frontier&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share The Software Frontier</span></a></p><div><hr></div><h2>Introduction</h2><p>The very first thing I did with Blackwell was to compile something for it on a machine that has <strong>no GPU</strong> in it at all. That is not a counter sense. </p><p><code>ptxas</code>, <code>nvdisasm</code> and <code>cuobjdump</code> are ordinary x86 programs, they ship as Python wheels on PyPI, and they will <strong>lower PTX </strong>for compute capability 10.0 on a laptop with a 40 mb download and no driver. You can&#8217;t run the result, but what you can do is to read every instruction, which for my own research, turned out to be the more useful half.</p><p>I wanted to know what a <code>tcgen05.alloc</code> costs, first was a small quest, but it got way bigger. It&#8217;s important because that&#8217;s the instruction that reserves <strong>Tensor Memory on Blackwell</strong>, it&#8217;s in every CUTLASS kernel for the architecture, and every explanation out there said the same <em>three things</em>: it takes a column count, it must be issued by one warp, and the result comes back through shared memory. </p><p>Nobody explained in full what the machine does. So, I wrote a <strong>small PTX kernel</strong> around it, assembled it for <code>sm_100a</code>, and disassembled the result.</p><p>What I saw was not an allocation in the sense a systems programmer means: it&#8217;s a <strong>uniform-datapath atomic</strong> against a per-SM pool, wrapped by the compiler in a spin loop with a <code>NANOSLEEP</code> backoff, guarded by three distinct trap handlers that <code>ptxas</code> injects on your behalf, with names like <code>__cuda_sm10x_tcgen05_guardrail_trap_unallocated_columns_being_dealloced</code>. </p><p>The tensor cores now have a memory allocator, with contention, with a retry path, and with a compiler-inserted runtime safety checker for use-after-free.</p><p>That&#8217;s the shape of the thing this article is about. Blackwell <strong>added an address space</strong> with its own instruction family, not just a cache with its own bus. On top of it we have its own allocator, a barrier, its own access-permission model, and zero coherence with anything else on the chip.</p><p>Almost everything written about Blackwell treats Tensor Memory as an implementation detail of the new MMA, but I found out it&#8217;s the opposite.</p><p>The <strong>MMA changed</strong> because the memory changed, which consequentially changed due to an arithmetic problem that had been building since 2017, and the consequences run outward through occupancy, epilogue design, quantization format choice, kernel portability, and eventually the cost per million tokens of anything you serve on hardware.</p><p>Neither of us owns a B200, unfortunately. Everything here that is measured was measured with a toolchain and a disassembler and is <strong>reproducible from the appendix</strong>. Derived things are in the open with the arithmetic on the page. All data taken from somebody else&#8217;s hardware is attributed and tiered in a dossier at the end. </p><p>Where the published literature and our disassembly disagree,(<em>spoiler alert: in two places they do</em>) both are shown.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>The number that did not move</h2><p>We write this very often: every <strong>streaming multiprocessor</strong> NVIDIA has launched since Volta has exactly 65,536 registers of 32 bits. That is 262,144 bytes of register file per SM, and NVIDIA&#8217;s own tuning guides give the same 64 K figure for Ampere and, in the Blackwell tuning guide, for compute capability 10.0. </p><p>The same documents states the<strong> 255 register per thread ceiling</strong>, which we&#8217;ll see in use. what we have is: five architectures, three process nodes, four HBM generations, and the largest piece of storage in the SM has not changed capacity by one byte in nine years.</p><p>Over the same period <strong>tensor throughput per SM</strong> per clock went up eight times, and the doubling is exact rather than approximate; NVIDIA&#8217;s own numbers make this easily checkable without trusting anybody&#8217;s marketing. </p><p>A Volta tensor core does <strong>64 fused multiply-adds</strong> per clock, eight of them per SM, so 1,024 FP16 FLOPs per clock per SM. Ampere doubled it to 2,048, then Hopper doubled it again to 4,096. </p><p>Multiply those out and the datasheets fall out to three digits: 108 A100 SMs times 2,048 times 1.410 GHz is 312 teraflops, which is exactly the published A100 figure. 132 H100 SXM SMs times 4,096 times 1.830 GHz is 989.7 teraflops against a published 989.4.</p><p>Run the same identity backwards on Blackwell. The B200 is two dies of 80 SMs with 74 enabled on each, so 148, and 2.25 petaflops dense FP16. That requires 8,192 FLOPs per clock per SM at 1.86 GHz. </p><p>Another doubling, and not only ours: the authors of the <strong>JAX scaling book </strong>run the same division and land on 2,048 FLOPs per tensor core per cycle across four tensor cores per SM, which is the same figure. So the ratio that actually matters, register file bytes per FP16 FLOP per clock, has fallen from 256 on Volta to 32 on Blackwell.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!vClk!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf0eac48-a9af-4255-afb6-0b6d447f9cd8_1800x800.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!vClk!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf0eac48-a9af-4255-afb6-0b6d447f9cd8_1800x800.png 424w, https://substackcdn.com/image/fetch/$s_!vClk!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf0eac48-a9af-4255-afb6-0b6d447f9cd8_1800x800.png 848w, https://substackcdn.com/image/fetch/$s_!vClk!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf0eac48-a9af-4255-afb6-0b6d447f9cd8_1800x800.png 1272w, https://substackcdn.com/image/fetch/$s_!vClk!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf0eac48-a9af-4255-afb6-0b6d447f9cd8_1800x800.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!vClk!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf0eac48-a9af-4255-afb6-0b6d447f9cd8_1800x800.png" width="1456" height="647" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/cf0eac48-a9af-4255-afb6-0b6d447f9cd8_1800x800.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:647,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Chart showing register file per SM flat at 262144 bytes from Volta to Blackwell while FP16 tensor FLOPs per clock per SM double each generation from 1024 to 8192&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Chart showing register file per SM flat at 262144 bytes from Volta to Blackwell while FP16 tensor FLOPs per clock per SM double each generation from 1024 to 8192" title="Chart showing register file per SM flat at 262144 bytes from Volta to Blackwell while FP16 tensor FLOPs per clock per SM double each generation from 1024 to 8192" srcset="https://substackcdn.com/image/fetch/$s_!vClk!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf0eac48-a9af-4255-afb6-0b6d447f9cd8_1800x800.png 424w, https://substackcdn.com/image/fetch/$s_!vClk!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf0eac48-a9af-4255-afb6-0b6d447f9cd8_1800x800.png 848w, https://substackcdn.com/image/fetch/$s_!vClk!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf0eac48-a9af-4255-afb6-0b6d447f9cd8_1800x800.png 1272w, https://substackcdn.com/image/fetch/$s_!vClk!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf0eac48-a9af-4255-afb6-0b6d447f9cd8_1800x800.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em><strong><sup>Figure 1.</sup></strong><sup> Register file capacity per SM against FP16 tensor throughput per SM per clock. The bars are flat by architectural fact. The line doubles every generation. Register file bytes per FLOP per clock: 256, 128, 64, 32.</sup></em></p><p>You can absorb a gap like that for one generation by being clever, and NVIDIA did, twice. Ampere added <strong>asynchronous copy</strong> so global to shared traffic stopped passing through registers. </p><p>Hopper added the Tensor Memory Accelerator so the address arithmetic for a tiled copy stopped consuming a warp&#8217;s registers, and added <code>setmaxnreg</code> so a<strong> producer warpgroup</strong> could donate its register budget to a consumer at runtime. Both are the same move: find something sitting in registers for no good reason and evict it.</p><p>By Hopper the evictable things were gone. What remained was the one thing that genuinely belongs to the math: the accumulator.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>The tile that cannot exist</h2><p>Here is the constraint that decided Blackwell&#8217;s design, and it is a single line of arithmetic against a limit you can make the assembler confirm.</p><p>On Hopper, <code>wgmma.mma_async</code> accumulates into registers owned by the 128 threads of a warpgroup. The largest shape is <code>m64n256k16</code>. That accumulator is 16,384 FP32 values, which over 128 threads is 128 registers per thread. </p><p>As stated in the subtitle, the architectural ceiling is 255, and <code>ptxas</code> will tell you so directly:</p><pre><code><code>$ ptxas -arch=sm_100a -maxrregcount=255 probe.ptx -o /dev/null
$ ptxas -arch=sm_100a -maxrregcount=256 probe.ptx -o /dev/null
ptxas warning : Too big maxrregcount value specified 256, will be ignored</code></code></pre><p>So Hopper&#8217;s largest MMA already spends half of every thread&#8217;s addressable register space on the output tile, before any addressing, predication, loop state or <strong>epilogue arithmetic.</strong></p><p>Now&#8230; let&#8217;s double it, which is what <code>tcgen05.mma</code> does. The largest single-CTA UMMA atom is <code>m128n256k16</code>, twice the area of the largest WGMMA atom. Its accumulator is 32,768 FP32 values. Over a warpgroup that is <strong>256 registers per thread</strong>, against a ceiling of 255.</p><p>Blackwell&#8217;s headline matrix instruction produces a result that is, by one register, unrepresentable in the programming model of every NVIDIA GPU that came before it.</p><p>This is a representability problem and not a turning tradeoff. There is no register allocation, no spill policy, <strong>no compiler heroics</strong> that make a 128 by 256 FP32 tile live in the fragments of a warpgroup, because the fragment model tops out below the tile. </p><p>If you want that instruction to exist, its output must go somewhere that is not the register file. Everything else about Tensor Memory follows from that sentence.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!ygwX!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6477b278-49a2-430f-8aae-b9bffd458da2_1800x760.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!ygwX!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6477b278-49a2-430f-8aae-b9bffd458da2_1800x760.png 424w, https://substackcdn.com/image/fetch/$s_!ygwX!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6477b278-49a2-430f-8aae-b9bffd458da2_1800x760.png 848w, https://substackcdn.com/image/fetch/$s_!ygwX!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6477b278-49a2-430f-8aae-b9bffd458da2_1800x760.png 1272w, https://substackcdn.com/image/fetch/$s_!ygwX!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6477b278-49a2-430f-8aae-b9bffd458da2_1800x760.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!ygwX!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6477b278-49a2-430f-8aae-b9bffd458da2_1800x760.png" width="1456" height="615" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6477b278-49a2-430f-8aae-b9bffd458da2_1800x760.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:615,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Bar chart of registers per thread required to hold FP32 accumulator tiles, with the 255 register ceiling crossed by the two largest Blackwell tiles&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Bar chart of registers per thread required to hold FP32 accumulator tiles, with the 255 register ceiling crossed by the two largest Blackwell tiles" title="Bar chart of registers per thread required to hold FP32 accumulator tiles, with the 255 register ceiling crossed by the two largest Blackwell tiles" srcset="https://substackcdn.com/image/fetch/$s_!ygwX!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6477b278-49a2-430f-8aae-b9bffd458da2_1800x760.png 424w, https://substackcdn.com/image/fetch/$s_!ygwX!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6477b278-49a2-430f-8aae-b9bffd458da2_1800x760.png 848w, https://substackcdn.com/image/fetch/$s_!ygwX!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6477b278-49a2-430f-8aae-b9bffd458da2_1800x760.png 1272w, https://substackcdn.com/image/fetch/$s_!ygwX!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6477b278-49a2-430f-8aae-b9bffd458da2_1800x760.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em><strong><sup>Figure 2.</sup></strong><sup> The forcing function. Hopper&#8217;s largest MMA already spent half the addressable register space per thread on the output tile. Blackwell&#8217;s largest single-CTA MMA needs one register more than the architecture allows.</sup></em></p><p>There is a second argument in the same direction, less absolute but more expensive in practice. A register is <strong>thread-private</strong>, so an MMA that accumulates into registers is an operation the owning threads must be present for. </p><p>On Hopper this is a scheduling tax: a warpgroup issues <code>wgmma</code>, and although the <em>instruction is asynchronous</em>, that warpgroup cannot go and do something else with those registers, because the registers <em>are</em> the accumulator. </p><p><strong>Warp specialization </strong>on Hopper is largely a set of arrangements to ensure the warps holding accumulators are not the warps doing anything interesting. Move the accumulator out and the whole class of tricks becomes unnecessary.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/p/how-blackwells-tensor-memory-actually?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/p/how-blackwells-tensor-memory-actually?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><div><hr></div><h2>What Tensor Memory actually is</h2><p>Tensor Memory is 256 kilobytes per SM, organized as <strong>128 lanes by 512 columns</strong> of 32 bit cells. A TMEM address is a 32 bit word whose high 16 bits are the lane index and low 16 bits are the column index. It is not a linear address space with a base and an offset, cause it&#8217;s a coordinate. </p><p>One consequence worth internalising before reading any <strong>CUTLASS TMEM layout</strong>: a stride of 65,536 in a TMEM tensor is not a large jump in memory, it is a step of exactly one lane.</p><p><em>Four properties</em> matter more than the capacity, and none are properties of any other memory on the chip.</p><ul><li><p><strong>It is allocated, not addressed.</strong> You call <code>tcgen05.alloc</code> with a column count, which must be a power of two and at least 32, and the hardware returns a base address which it writes into shared memory for you. Allocation is by column, and a column is all 128 lanes: there is no way to reserve part of one. You free it with <code>tcgen05.dealloc</code>, from the same warp that allocated it, or the columns stay claimed.</p></li><li><p><strong>Nothing computes on it.</strong> The only instructions that touch TMEM are the <code>tcgen05</code> family. No ALU op, no <code>ld.shared</code>, no <code>ldmatrix</code>, no <code>cp.async</code>, no atomic, no texture path. Every pre-processing step happens before data enters and every post-processing step after it leaves. TMEM is the one memory on a Blackwell SM that a general purpose instruction cannot see.</p></li><li><p><strong>Access is partitioned by warp, in hardware.</strong> When threads read or write TMEM explicitly, warp 0 of a warpgroup reaches only lanes 0 to 31, warp 1 only lanes 32 to 63, and so on. This is the access model, not a guideline. One warp physically cannot read a full 128 lane accumulator tile, so draining one requires a whole warpgroup by construction, and CUTLASS&#8217;s <code>make_tmem_copy</code> is hardcoded to four warps for exactly that reason.</p></li><li><p><strong>It is a per-SM pool with a hard ceiling.</strong> 512 columns, shared by every CTA resident on that SM, and the largest UMMA accumulator occupies exactly 256 of them.</p></li></ul><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!yMB9!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4faa0c3f-8b88-495a-9c08-417abcab5cf8_1800x920.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!yMB9!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4faa0c3f-8b88-495a-9c08-417abcab5cf8_1800x920.png 424w, https://substackcdn.com/image/fetch/$s_!yMB9!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4faa0c3f-8b88-495a-9c08-417abcab5cf8_1800x920.png 848w, https://substackcdn.com/image/fetch/$s_!yMB9!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4faa0c3f-8b88-495a-9c08-417abcab5cf8_1800x920.png 1272w, https://substackcdn.com/image/fetch/$s_!yMB9!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4faa0c3f-8b88-495a-9c08-417abcab5cf8_1800x920.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!yMB9!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4faa0c3f-8b88-495a-9c08-417abcab5cf8_1800x920.png" width="1456" height="744" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/4faa0c3f-8b88-495a-9c08-417abcab5cf8_1800x920.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:744,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Diagram of the Tensor Memory grid, 128 lanes by 512 columns, showing a 128 by 256 accumulator allocation, scale factor columns, and the warp lane partition&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Diagram of the Tensor Memory grid, 128 lanes by 512 columns, showing a 128 by 256 accumulator allocation, scale factor columns, and the warp lane partition" title="Diagram of the Tensor Memory grid, 128 lanes by 512 columns, showing a 128 by 256 accumulator allocation, scale factor columns, and the warp lane partition" srcset="https://substackcdn.com/image/fetch/$s_!yMB9!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4faa0c3f-8b88-495a-9c08-417abcab5cf8_1800x920.png 424w, https://substackcdn.com/image/fetch/$s_!yMB9!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4faa0c3f-8b88-495a-9c08-417abcab5cf8_1800x920.png 848w, https://substackcdn.com/image/fetch/$s_!yMB9!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4faa0c3f-8b88-495a-9c08-417abcab5cf8_1800x920.png 1272w, https://substackcdn.com/image/fetch/$s_!yMB9!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4faa0c3f-8b88-495a-9c08-417abcab5cf8_1800x920.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em><strong><sup>Figure 3</sup></strong><sup>. The ledger. Every Blackwell GEMM kernel is, underneath, a plan for carving up 512 columns. The largest accumulator takes half, block scale factors take a few more, and what is left is what you have for pipelining or for a second resident CTA.</sup></em></p><p>The instruction that uses all this is <code>tcgen05.mma</code>, which <strong>CUTLASS calls UMMA</strong>. Its operand rules invert everything before them. Operand A may be in shared memory or in Tensor Memory. Operand B has to be in shared memory. The accumulator must be in Tensor Memory. Registers appear nowhere in that sentence.</p><p>And it is issued by <strong>one thread</strong>. Not a warp, not a warpgroup: a single elected thread on behalf of the whole CTA, or of a pair of CTAs under <code>cta_group::2</code>, where two SMs sharing a texture processing cluster cooperate on one logical tile. </p><p>The consequence is visible in CUTLASS: the CuTe atom&#8217;s <code>ThrID</code>, which was <code>Layout&lt;_32&gt;</code> for <strong>warp-level MMA </strong>and <code>Layout&lt;_128&gt;</code> for Hopper&#8217;s warpgroup MMA, is now <code>Layout&lt;_1&gt;</code>, and the thread layouts have been repurposed as layouts of the CTAs collaborating on the instruction.</p><p>The abstraction the entire programming model is named after has been vacated at the top of the pipeline. There is still a thread. It does not do the math, does not own the inputs, and does not own the result.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>What the compiler actually emits</h2><p>This is the part we did ourselves, so it is the part we trust most. The toolchain is<strong> three PyPI wheels</strong> and no GPU: <code>ptxas</code> 12.9.86 from <code>nvidia-cuda-nvcc-cu12</code>, and <code>nvdisasm</code> and <code>cuobjdump</code> 13.3.73 from their standalone packages. The full command sequence is in the appendix.</p><p>The probe does the minimum honest thing: reserve 128 columns of <strong>Tensor Memory</strong>, read back the base address, issue one UMMA into it, commit through an mbarrier, drain one fragment to registers, store it, free the columns.</p><pre><code><code>// probe.ptx, assembled with ptxas 12.9.86 for sm_100a
.version 8.6
.target sm_100a
.address_size 64

.visible .entry umma_probe(.param .u64 p_out, .param .u64 p_adesc, .param .u64 p_bdesc)
{
    .reg .b32  %r&lt;16&gt;;   .reg .b64 %rd&lt;8&gt;;   .reg .pred %p&lt;2&gt;;
    .shared .align 16 .b32 tmem_slot[4];
    .shared .align 8  .b64 mbar[1];

    ld.param.u64 %rd1, [p_out];
    ld.param.u64 %rd2, [p_adesc];
    ld.param.u64 %rd3, [p_bdesc];

    mov.u32 %r1, 128;
    tcgen05.alloc.cta_group::1.sync.aligned.shared::cta.b32 [tmem_slot], %r1;
    tcgen05.relinquish_alloc_permit.cta_group::1.sync.aligned;
    ld.shared.b32 %r2, [tmem_slot];

    mov.u32 %r3, 0;
    setp.eq.u32 %p1, %r3, 0;
    tcgen05.mma.cta_group::1.kind::f16 [%r2], %rd2, %rd3, %r3, %p1;

    mbarrier.init.shared::cta.b64 [mbar], 1;
    tcgen05.commit.cta_group::1.mbarrier::arrive::one.shared::cluster.b64 [mbar];

    tcgen05.ld.sync.aligned.32x32b.x1.b32 {%r10}, [%r2];
    tcgen05.wait::ld.sync.aligned;
    st.global.u32 [%rd1], %r10;

    mov.u32 %r11, 128;
    tcgen05.dealloc.cta_group::1.sync.aligned.b32 %r2, %r11;
    ret;
}</code></code></pre><p><em>Three facts fall out before the disassembler is even involved.</em></p><ol><li><p><strong>The minimum PTX ISA version is 8.6</strong>, which is CUDA 12.8. Below that the assembler names the requirement precisely: <code>Feature 'tcgen05.alloc' requires PTX ISA .version 8.6 or later</code>. We mention this because at least one widely linked public write-up puts it at 8.4, and 8.4 does not assemble.</p></li><li><p><code>tcgen05.commit</code><strong> requires </strong><code>.shared::cluster</code>, and rejects <code>.shared::cta</code> with <code>State space incorrect for instruction 'tcgen05.commit'</code>, even in a kernel with a trivial cluster and <code>cta_group::1</code>. The tensor core completion path is a cluster level mechanism whether or not you asked for a cluster.</p></li><li><p><code>ptxas</code><strong> does not validate the column count, even when it is a compile time constant.</strong> We assembled the probe with 16, 48, 96 and 1,024 columns, all of which violate the documented power-of-two and minimum-32 rules, and all of which assembled without a warning. The rule is enforced by hardware at runtime, not by the compiler. Given that the failure mode is a trap handler, this is a class of bug that only exists on a machine you may not own.</p></li></ol><h3>The allocator</h3><p>Here is what <code>tcgen05.alloc</code> becomes. Address arithmetic trimmed, control flow still intact.</p><pre><code><code>// nvdisasm -c probe_sm100a.cubin, excerpt
        ELECT P0, URZ, PT ;
   @!P0 BRA `(.L_x_2) ;
        DEPBAR.LE SB0, 0x36 ;
        UTCATOMSWS.FIND_AND_SET.ALIGN UP0, UR4, UR4 ;      // claim an aligned run of columns
        PLOP3.LUT P0, PT, PT, PT, UP0, 0x80, 0x8 ;
        SEL R0, RZ, 0xffffffff, !P0 ;
        ISETP.NE.AND P0, PT, R0, RZ, PT ;
    @P0 BRA `(.L_x_3) ;                                    // claimed, continue
.L_x_4:
        NANOSLEEP 0x64 ;                                    // back off, then retry
        UTCATOMSWS.FIND_AND_SET.ALIGN UP0, UR4, UR4 ;
        ...
   @!P0 BRA `(.L_x_4) ;                                    // spin
.L_x_3:
        ATOMS.OR RZ, [UR4], R2 ;                           // record the claimed mask in SMEM
        STS [UR6], R0 ;                                    // publish the base address
        ...
        UVIRTCOUNT.DEALLOC.SMPOOL 0x80 ;                   // SM pool virtual counter</code></code></pre><p>The claim is a <code>UTCATOMSWS.FIND_AND_SET.ALIGN</code>, a uniform-datapath atomic that searches a per-SM pool for a free, aligned run of columns and sets them. </p><p>On failure the compiler emits a <code>NANOSLEEP</code> of 0x64 units and retries indefinitely. So a Blackwell kernel that allocates Tensor Memory has a <strong>spin loop</strong> on its critical path that no source line asked for, and the pool is genuinely contended, otherwise the retry path would not be there.</p><p>Then there are the <strong>guardrails</strong>. <code>ptxas</code> injects three named trap handlers and fifteen references to them, and it does so identically at every optimisation level from <code>-O0</code> to <code>-O3</code>:</p><pre><code><code>$__internal_0_$__cuda_sm10x_tcgen05_guardrail_trap_col_being_dealloced_not_returned_by_alloc
$__internal_1_$__cuda_sm10x_tcgen05_guardrail_trap_phase_invalid_during_alloc
$__internal_2_$__cuda_sm10x_tcgen05_guardrail_trap_unallocated_columns_being_dealloced</code></code></pre><p>Read those names as a bug taxonomy. Freeing columns you did not allocate. Freeing a column that came from a different allocation. Allocating from an invalid phase. </p><p>That is a <strong>use-after-free checker</strong>, a double-free checker and a state machine assertion, compiled into every kernel, unconditionally. NVIDIA does not do this casually, and the fact that they did it tells you what the failure modes look like in practice.</p><blockquote><p><em><strong>Why this matters beyond trivia</strong></em></p><p><em>Every mental model of a GPU kernel assumes resources are assigned at launch. Registers and shared memory are fixed by the compiler and the launch configuration, and occupancy is computed from them before a single instruction runs. Tensor Memory is the first first-class SM resource that is acquired at runtime, can block, can fail, and is arbitrated by an atomic. It moves part of the occupancy calculation out of the launch and into the kernel body, where no static tool can see it.</em></p></blockquote><p>One number for scale. The probe above, whose actual work is one matrix multiply and one 32 bit drain, compiles to <strong>152 SASS instructions at </strong><code>-O3</code> and 400 at <code>-O0</code>, using 14 registers and 24 bytes of shared memory. </p><p>Two of those 152 are the tensor core; the rest is allocation, election, barrier phase tracking, address reconstruction and guardrails.</p><h3>The instruction</h3><pre><code><code>// dense f16
UTCHMMA      gdesc[UR14], gdesc[UR16], tmem[UR9], tmem[URZ], idesc[URZ], UPT ;
// same source with .sp: the sparsity metadata takes the fourth tmem slot
UTCHMMA      gdesc[UR14], gdesc[UR16], tmem[UR7], tmem[UR4], idesc[UR5], UPT ;
// same source with .cta_group::2
UTCHMMA.2CTA gdesc[UR12], gdesc[UR14], tmem[UR8], tmem[URZ], idesc[URZ], UPT ;
// NVFP4, block16 scaling: scale factors are a fifth tmem operand
UTCOMMA.4X   gdesc[UR14], gdesc[UR16], tmem[UR8], tmem[URZ], idesc[URZ], tmem[UR8], UPT ;</code></code></pre><p>Every operand is a descriptor or a Tensor Memory coordinate, and every one lives in a <strong>uniform register</strong>, the UR file, not the per-thread register file. The <code>U</code> prefix is the same <code>U</code> as in <code>UMOV</code>, <code>ULEA</code> and <code>UIADD3</code>: the scalar, warp-uniform datapath NVIDIA added in Turing for address arithmetic. Blackwell&#8217;s matrix multiply runs entirely on it. There is no vector register in the instruction at all.</p><p>That is the<strong> cleanest evidence</strong> we have found that the CTA, not the thread, is now the unit of tensor computation. It is not an abstraction in CUTLASS or a convenience in PTX. It is visible in which register file the opcode reads.</p><p>One clarification the SASS supports and the secondary literature often gets wrong: under <code>cta_group::2</code> the instruction is not issued by both CTAs in lock step. CUTLASS&#8217;s own <strong>two-SM tutorial states</strong> that only one of the two peer CTAs executes it, and names that one the leader. A single thread, in a single CTA, drives a matrix multiply spanning two SMs.</p><p>The dense form passes <code>tmem[URZ]</code>, the zero register, in the fourth slot; the sparse form fills it with a real address. </p><p>So <strong>structured sparsity metadata</strong> also lives in Tensor Memory, alongside the accumulator and, for block scaled kinds, alongside the scale factors. Three different kinds of state, one pool of 512 columns.</p><h3>The opcode family</h3><p>We compiled the probe once per qualifier and disassembled each result. This mapping is measured.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!TFo_!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15f09d28-0351-4bd9-9668-8393d162d188_2700x2244.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!TFo_!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15f09d28-0351-4bd9-9668-8393d162d188_2700x2244.png 424w, https://substackcdn.com/image/fetch/$s_!TFo_!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15f09d28-0351-4bd9-9668-8393d162d188_2700x2244.png 848w, https://substackcdn.com/image/fetch/$s_!TFo_!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15f09d28-0351-4bd9-9668-8393d162d188_2700x2244.png 1272w, https://substackcdn.com/image/fetch/$s_!TFo_!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15f09d28-0351-4bd9-9668-8393d162d188_2700x2244.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!TFo_!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15f09d28-0351-4bd9-9668-8393d162d188_2700x2244.png" width="1456" height="1210" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/15f09d28-0351-4bd9-9668-8393d162d188_2700x2244.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1210,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:89069,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/208027653?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15f09d28-0351-4bd9-9668-8393d162d188_2700x2244.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!TFo_!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15f09d28-0351-4bd9-9668-8393d162d188_2700x2244.png 424w, https://substackcdn.com/image/fetch/$s_!TFo_!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15f09d28-0351-4bd9-9668-8393d162d188_2700x2244.png 848w, https://substackcdn.com/image/fetch/$s_!TFo_!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15f09d28-0351-4bd9-9668-8393d162d188_2700x2244.png 1272w, https://substackcdn.com/image/fetch/$s_!TFo_!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15f09d28-0351-4bd9-9668-8393d162d188_2700x2244.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>This <strong>disagrees with the published literature</strong> in one place. The Delaware microbenchmarking paper, which is otherwise the most useful public measurement of this hardware and which we lean on later, reports in its Table IV that <code>tcgen05.mma</code> lowers to <code>HMMA</code>, <code>QMMA</code>, <code>OMMA</code> and <code>IMMA</code>, the same opcode names Volta through Hopper used. </p><p>On <strong>our disassembly</strong> it does not. The distinction is not cosmetic: <code>HMMA</code> reads and writes vector registers, <code>UTCHMMA</code> touches none. The most likely explanation is that the table was written from the family names rather than from a fresh disassembly.</p><h3>Which chips can run any of this</h3><p>Compiling against every target the assembler accepts produces a matrix sharper than the marketing, and the sharpest row is the one nobody mentions.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!StRx!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F551c9e65-c5d4-4eff-96b6-763c7bbb8028_900x592.svg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!StRx!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F551c9e65-c5d4-4eff-96b6-763c7bbb8028_900x592.svg 424w, https://substackcdn.com/image/fetch/$s_!StRx!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F551c9e65-c5d4-4eff-96b6-763c7bbb8028_900x592.svg 848w, https://substackcdn.com/image/fetch/$s_!StRx!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F551c9e65-c5d4-4eff-96b6-763c7bbb8028_900x592.svg 1272w, https://substackcdn.com/image/fetch/$s_!StRx!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F551c9e65-c5d4-4eff-96b6-763c7bbb8028_900x592.svg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!StRx!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F551c9e65-c5d4-4eff-96b6-763c7bbb8028_900x592.svg" width="1456" height="958" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/551c9e65-c5d4-4eff-96b6-763c7bbb8028_900x592.svg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:958,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:7800,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/svg+xml&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/208027653?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F551c9e65-c5d4-4eff-96b6-763c7bbb8028_900x592.svg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!StRx!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F551c9e65-c5d4-4eff-96b6-763c7bbb8028_900x592.svg 424w, https://substackcdn.com/image/fetch/$s_!StRx!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F551c9e65-c5d4-4eff-96b6-763c7bbb8028_900x592.svg 848w, https://substackcdn.com/image/fetch/$s_!StRx!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F551c9e65-c5d4-4eff-96b6-763c7bbb8028_900x592.svg 1272w, https://substackcdn.com/image/fetch/$s_!StRx!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F551c9e65-c5d4-4eff-96b6-763c7bbb8028_900x592.svg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><pre><code><code>$ ptxas -arch=sm_100  probe.ptx  -o /dev/null
ptxas error : Instruction 'tcgen05.alloc' not supported on .target 'sm_100'

$ ptxas -arch=sm_100a legacy.ptx -o /dev/null
ptxas error : Instruction 'wgmma.fence' not supported on .target 'sm_100a'

$ ptxas -arch=sm_120a legacy.ptx -o /dev/null
ptxas error : Instruction 'wgmma.fence' not supported on .target 'sm_120a'</code></code></pre><p>Three things to take from that table.</p><p>First, <code>wgmma</code> is not deprecated on Blackwell. It is <strong>removed</strong>, and it is removed from every Blackwell target including the consumer one. A Hopper kernel built on warpgroup MMA does not run slower on Blackwell, it fails at assembly time on all of <code>sm_100a</code>, <code>sm_103a</code> and <code>sm_120a</code>. </p><p>This is a <strong>harder break than any</strong> NVIDIA has shipped in the tensor core era, and it is worth stating precisely because the commonly repeated version of this claim, that consumer Blackwell falls back to <code>mma.sync</code> and <code>wgmma</code>, is wrong on the second half. There is no <code>wgmma</code> to fall back to.</p><p>Second, <code>sm_100</code> without the trailing <code>a</code>, the forward compatible target whose PTX a future driver may recompile for a future chip, has neither <code>tcgen05</code> nor <code>wgmma</code>. </p><p>The <strong>entire tensor path</strong> that defines Blackwell is available only on architecture specific targets, which by NVIDIA&#8217;s documented rules are not forward compatible with anything.</p><p>Third, the only instruction family that spans Hopper, datacenter Blackwell and consumer Blackwell is <code>mma.sync</code>, the warp-level, register-resident, <code>m16n8k16</code> class of instruction that predates all of this. That is the portable subset now. It is also the one whose accumulator lives in the register file, which is the constraint we said the architecture outgrew.</p><p>The <strong>portable path</strong> and the fast path have fully separated. What runs everywhere is the instruction whose limits forced Tensor Memory into existence.</p><p>The practical consequence is a development loop, not a benchmark. You cannot write, debug or profile a <code>tcgen05</code> kernel on a workstation. Not slowly, not at reduced fidelity, not at all. The instructions do not exist on hardware you can buy without a data center behind it. </p><p>In our compiler moat piece we argued the <strong>durable advantage sits at the </strong><code>ptxas</code><strong> and SASS layer</strong> and that the counter-technology is research rather than a new language. This adds a cruder second mechanism that has nothing to do with compilers: iteration on the instructions that matter now requires an allocation of scarce hardware.</p><p>We should immediately weaken that. The barrier is money and patience rather than access. B200 time is rentable by the hour, and the author of the best <code>tcgen05</code> tutorial we have read, writing as <strong>gau-nernst</strong>, reports reaching 98 percent of cuBLAS on a 4096 cubed problem using rented capacity. </p><p>A single motivated person got there in a blog series. That is way a much weaker moat than a first reading suggests.</p><h3>The cost of getting the answer back</h3><p><code>tcgen05.ld</code> takes a shape and a repetition count, and the count determines how many 32 bit values land in each thread&#8217;s registers. </p><p>We swept the count from 1 to 128, <strong>forced every loaded value</strong> to stay live by storing all of them, and read the allocation from <code>ptxas -v</code>.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!pZEC!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72920abc-baa3-4c1e-9133-956eea449577_1800x760.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!pZEC!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72920abc-baa3-4c1e-9133-956eea449577_1800x760.png 424w, https://substackcdn.com/image/fetch/$s_!pZEC!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72920abc-baa3-4c1e-9133-956eea449577_1800x760.png 848w, https://substackcdn.com/image/fetch/$s_!pZEC!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72920abc-baa3-4c1e-9133-956eea449577_1800x760.png 1272w, https://substackcdn.com/image/fetch/$s_!pZEC!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72920abc-baa3-4c1e-9133-956eea449577_1800x760.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!pZEC!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72920abc-baa3-4c1e-9133-956eea449577_1800x760.png" width="1456" height="615" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/72920abc-baa3-4c1e-9133-956eea449577_1800x760.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:615,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Measured register usage per thread as a function of the tcgen05.ld repetition count, rising from 14 registers at x1 to 134 registers at x128&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Measured register usage per thread as a function of the tcgen05.ld repetition count, rising from 14 registers at x1 to 134 registers at x128" title="Measured register usage per thread as a function of the tcgen05.ld repetition count, rising from 14 registers at x1 to 134 registers at x128" srcset="https://substackcdn.com/image/fetch/$s_!pZEC!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72920abc-baa3-4c1e-9133-956eea449577_1800x760.png 424w, https://substackcdn.com/image/fetch/$s_!pZEC!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72920abc-baa3-4c1e-9133-956eea449577_1800x760.png 848w, https://substackcdn.com/image/fetch/$s_!pZEC!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72920abc-baa3-4c1e-9133-956eea449577_1800x760.png 1272w, https://substackcdn.com/image/fetch/$s_!pZEC!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72920abc-baa3-4c1e-9133-956eea449577_1800x760.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em><strong><sup>Figure 4</sup></strong><sup> The epilogue tax, measured. A single ldtm.x128 costs 134 registers per thread. Across a warpgroup that is 67 KiB of the 256 KiB register file, just over a quarter of it, checked out purely to hold data on its way from one on-chip memory to another.</sup></em></p><p>The curve is N plus six, and <code>ptxas</code> emits one <code>LDTM.xN</code> rather than N loads. So the <strong>register file did not stop being the bottleneck</strong>. It stopped being the bottleneck for the multiply and became the bottleneck for the epilogue.</p><p>On a Hopper kernel the accumulator is already in registers when you want to scale it, add a bias, apply an activation and cast down. On Blackwell you must<strong> first pay to bring it back</strong>, and the wider you pay the less room remains to do anything with the result. </p><p>Every fused epilogue on this architecture is a negotiation between drain width and register headroom, and no source-level construct expresses it.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Why it had to be a separate memory</h2><p><strong>previous sections </strong>explained why the accumulator left the register file. It does not explain why the destination is a new addressable memory rather than a hidden internal buffer. </p><p>For that you need <strong>the traffic number</strong>, and the traffic number turns out to be more interesting than we expected when we first ran it.</p><p>Start with the shape rule, which is the load bearing fact. In CUTLASS&#8217;s Blackwell MMA traits the K extent of a UMMA atom is not a free parameter. It is computed:</p><pre><code><code>// cute/atom/mma_traits_sm100.hpp
// Logical shape-K is always 256bits, transform to units of elements
static constexpr int K = 256 / cute::sizeof_bits&lt;ValTypeA&gt;::value;</code></code></pre><p>K is always 256 bits, 32 bytes, of operand. FP16 gives K equal to 16. FP8 gives 32. FP4 gives 64. The K dimension scales as the inverse of element width, exactly.</p><p>Now compute the accumulator traffic. <strong>One UMMA</strong> over an M by N tile with FP32 accumulation reads the tile and writes it back, which is 8MN bytes, and performs 2MNK floating point operations. So:</p><pre><code><code>accumulator bytes per FLOP  =  8MN / (2MNK)  =  4 / K  =  (bits per element) / 64

FP16   K=16   0.2500 B/FLOP   x  2.25 PFLOP/s  =  562.5 TB/s
FP8    K=32   0.1250 B/FLOP   x  4.50 PFLOP/s  =  562.5 TB/s
FP4    K=64   0.0625 B/FLOP   x  9.00 PFLOP/s  =  562.5 TB/s

per SM (148):                                       3.80 TB/s
B200 HBM3e:                                         8.00 TB/s
ratio, chip accumulator traffic to HBM:               70 x</code></code></pre><p>The tile shape cancels, the clock cancels, and so does the precision. <strong>Accumulator traffic on a B200 is 562.5 terabytes per second at every precision the tensor cores support</strong>, because every time NVIDIA doubled the math rate by halving the element width, they simultaneously doubled K, which halved the accumulator traffic per FLOP. The two effects cancel to the digit.</p><p>That is not a coincidence and it is not a rounding artifact. It is the constraint the 256 bit K rule exists to satisfy. </p><p>Whatever structure holds the accumulator has to sustain a fixed bandwidth, and the format ladder was designed so that <strong>going from FP16 to FP4 buys four times the math</strong> without asking the accumulator path for a single additional byte per second.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!1r1K!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc8dd318-a701-4316-a238-1956fb4462b8_1800x800.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!1r1K!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc8dd318-a701-4316-a238-1956fb4462b8_1800x800.png 424w, https://substackcdn.com/image/fetch/$s_!1r1K!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc8dd318-a701-4316-a238-1956fb4462b8_1800x800.png 848w, https://substackcdn.com/image/fetch/$s_!1r1K!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc8dd318-a701-4316-a238-1956fb4462b8_1800x800.png 1272w, https://substackcdn.com/image/fetch/$s_!1r1K!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc8dd318-a701-4316-a238-1956fb4462b8_1800x800.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!1r1K!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc8dd318-a701-4316-a238-1956fb4462b8_1800x800.png" width="1456" height="647" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/bc8dd318-a701-4316-a238-1956fb4462b8_1800x800.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:647,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Chart showing math throughput rising four times from FP16 to FP4 while accumulator traffic stays constant at 562.5 terabytes per second&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Chart showing math throughput rising four times from FP16 to FP4 while accumulator traffic stays constant at 562.5 terabytes per second" title="Chart showing math throughput rising four times from FP16 to FP4 while accumulator traffic stays constant at 562.5 terabytes per second" srcset="https://substackcdn.com/image/fetch/$s_!1r1K!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc8dd318-a701-4316-a238-1956fb4462b8_1800x800.png 424w, https://substackcdn.com/image/fetch/$s_!1r1K!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc8dd318-a701-4316-a238-1956fb4462b8_1800x800.png 848w, https://substackcdn.com/image/fetch/$s_!1r1K!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc8dd318-a701-4316-a238-1956fb4462b8_1800x800.png 1272w, https://substackcdn.com/image/fetch/$s_!1r1K!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc8dd318-a701-4316-a238-1956fb4462b8_1800x800.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em><strong><sup>Figure 5.</sup></strong><sup> The invariance. This is the strongest single argument for why Tensor Memory is a fixed size, fixed geometry, precision-agnostic array: the bandwidth it must sustain does not depend on what numbers you put in it.</sup></em></p><h3>The other half, which explains where each operand lives</h3><p>Run the same calculation on the inputs and the architecture stops looking like a set of choices and starts looking like a single constraint solved twice.</p><p>One <strong>UMMA</strong> reads an M by K slab of A and a K by N slab of B, so operand bytes are (M + N) times K times the element width, against the same 2MNK operations. The K cancels here too:</p><pre><code><code>operand bytes per FLOP  =  (M+N)K b / (2MNK)  =  (M+N) b / (2MN)     // b = bytes per element

for the m128n256 tile:
FP16   b=2      0.01172 B/FLOP   x  2.25 PFLOP/s  =  26.4 TB/s
FP8    b=1      0.00586 B/FLOP   x  4.50 PFLOP/s  =  26.4 TB/s
FP4    b=0.5    0.00293 B/FLOP   x  9.00 PFLOP/s  =  26.4 TB/s

accumulator traffic / operand traffic  =  562.5 / 26.4  =  21.3 x</code></code></pre><p>Invariant again, and for the same reason: halving the element width halves the operand bytes per FLOP at exactly the rate it doubles the FLOPs. So the precision ladder from FP16 down to FP4 is <strong>bandwidth neutral at both ends of the datapath</strong>. </p><p>Blackwell quadrupled its peak math without asking either the operand path or the accumulator path for one additional byte per second.</p><p>That is also the answer to a question the operand rules raise and never explain. <em>Why must B sit in shared memory while D must sit somewhere else entirely?</em> Because the accumulator moves twenty one times the traffic. </p><p>Operands are read once per instruction and are narrow by construction; the <strong>accumulator is read and written in full</strong> every single time, in FP32, no matter how few bits the inputs have. A general purpose memory can serve the first workload. Nothing general purpose serves the second.</p><p>It also explains a fact that otherwise looks like an oversight: Blackwell&#8217;s shared memory did not grow. <strong>228 kilobytes per SM</strong>, identical to Hopper, in a generation that doubled math per SM per clock. It did not need to. </p><p>And where operand pressure did rise, NVIDIA answered with sharing rather than capacity: under <code>cta_group::2</code> two SMs in a <strong>texture processing cluster </strong>consume the same operands for one logical tile, which halves the per-SM operand traffic instead of doubling the memory that carries it.</p><p>562.5 terabytes per second across the chip is seventy times the entire HBM bandwidth of a B200, and 3.8 terabytes per second per SM sustained. Neither a cache nor a register file survives that. </p><p>What survives is a small, <strong>banked array</strong> physically adjacent to the consumer, addressed in the coordinate system the datapath already uses, and free of every general purpose obligation: no coherence, no cache tags, no arbitrary indexing, no participation in the memory model, no ability to be read by an ALU. </p><p>Which is precisely the <strong>list of things TMEM cannot do.</strong> The restrictions are not a first generation compromise to be relaxed later. They are the reason the number is achievable.</p><p>It is worth putting our figure next to the one measurement that exists. The Delaware group reports roughly 16 terabytes per second of TMEM read bandwidth on a B200. That is <strong>about four times </strong>the sustained accumulator requirement we derive, which is the right shape of answer: a peak port figure with headroom for drains overlapping accumulation, not a number that contradicts ours. </p><p>We would rather show both than pretend they measure the same thing.</p><p><strong>A caution about that paper&#8217;s peaks</strong>The same paper reports achieved throughputs as percentages of theoretical peak: FP4 at 7,700 TFLOPS being 96.2 percent, FP16 at 1,929.6 being 96.5 percent. Those imply peaks of about 8,004 and 1,999 TFLOPS, where the B200 datasheet says 9,000 and 2,250 dense. </p><p>Both of their implied peaks are <strong>11.1 percent below the datasheet</strong>, consistently, which is what you get from assuming a clock about 11 percent lower than the one the datasheet figures use. Their measurements are probably fine. </p><p>Their percentages are relative to a different baseline than NVIDIA&#8217;s, and should not be read as 96 percent of the number on the box.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Occupancy stops meaning what it meant</h2><p>Occupancy is the<strong> oldest performance heuristic</strong> in CUDA: resident warps per SM, limited by registers and shared memory, more of them meaning more latency to hide. </p><p>On Blackwell tensor kernels it is close to useless, and the reason is that a resource nobody&#8217;s occupancy calculator models now binds before the ones it does.</p><p>Start with what<strong> NVIDIA documents </strong>for compute capability 10.0, in the Blackwell tuning guide. </p><ul><li><p><em>Register file: 64 K 32-bit registers per SM, unchanged. </em></p></li><li><p><em>Maximum concurrent warps per SM: 64, unchanged since Volta. </em></p></li><li><p><em>Maximum thread blocks per SM: 32. </em></p></li><li><p><em>Shared memory capacity per SM: 228 kilobytes, the same as Hopper, with 227 addressable by a single block after CUDA&#8217;s 1 kilobyte reservation. </em></p></li><li><p><em>Combined L1, texture and shared memory: 256 kilobytes, also the same as Hopper.</em></p></li></ul><p>Read that list again with the previous sections in mind. Across a generation that doubled tensor throughput per SM per clock, <strong>not one of the classical occupancy resources grew</strong>. </p><p>The only capacity Blackwell added to the SM is the 256 kilobytes of Tensor Memory, and Tensor Memory is the one resource with an allocation rule that quantizes hard.</p><p>Columns come in powers of two, minimum 32, from a pool of 512, and <strong>every column carries all 128 lanes</strong>. So the pool admits 16 concurrent allocations at the smallest legal size, 8 at 64 columns, 4 at 128, 2 at 256, and 1 if you take the whole thing. There is no middle. And 16, the best case, is already half the hardware ceiling of 32 blocks per SM. </p><p>The moment your tile needs more than the minimum allocation, which any tile worth issuing a UMMA for does, <strong>Tensor Memory</strong> is the binding constraint and nothing else is close.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!v2Yk!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcacc8382-819f-430a-8d28-4985ba5a559f_1800x800.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!v2Yk!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcacc8382-819f-430a-8d28-4985ba5a559f_1800x800.png 424w, https://substackcdn.com/image/fetch/$s_!v2Yk!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcacc8382-819f-430a-8d28-4985ba5a559f_1800x800.png 848w, https://substackcdn.com/image/fetch/$s_!v2Yk!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcacc8382-819f-430a-8d28-4985ba5a559f_1800x800.png 1272w, https://substackcdn.com/image/fetch/$s_!v2Yk!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcacc8382-819f-430a-8d28-4985ba5a559f_1800x800.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!v2Yk!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcacc8382-819f-430a-8d28-4985ba5a559f_1800x800.png" width="1456" height="647" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/cacc8382-819f-430a-8d28-4985ba5a559f_1800x800.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:647,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Chart of concurrent tensor memory allocations permitted by a 512 column pool, falling from sixteen at thirty two columns to one at five hundred and twelve, against a hardware ceiling of thirty two thread blocks per SM&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Chart of concurrent tensor memory allocations permitted by a 512 column pool, falling from sixteen at thirty two columns to one at five hundred and twelve, against a hardware ceiling of thirty two thread blocks per SM" title="Chart of concurrent tensor memory allocations permitted by a 512 column pool, falling from sixteen at thirty two columns to one at five hundred and twelve, against a hardware ceiling of thirty two thread blocks per SM" srcset="https://substackcdn.com/image/fetch/$s_!v2Yk!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcacc8382-819f-430a-8d28-4985ba5a559f_1800x800.png 424w, https://substackcdn.com/image/fetch/$s_!v2Yk!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcacc8382-819f-430a-8d28-4985ba5a559f_1800x800.png 848w, https://substackcdn.com/image/fetch/$s_!v2Yk!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcacc8382-819f-430a-8d28-4985ba5a559f_1800x800.png 1272w, https://substackcdn.com/image/fetch/$s_!v2Yk!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcacc8382-819f-430a-8d28-4985ba5a559f_1800x800.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em><strong><sup>Figure 6.</sup></strong><sup> What the pool admits. This is an upper bound from the allocation rule, not a measurement of scheduler behaviour. Whether the hardware will actually co-resident that many CTAs is a separate question, and one we cannot settle without a B200.</sup></em></p><p>That caveat is not decoration. Colfax&#8217;s post on the workstation part, contrasting it with the datacenter part, states flatly that on SM10x <code>tcgen05.mma</code> is locked to one CTA per SM. <strong>Colfax&#8217;s own tutorial kernel</strong> allocates all 512 columns for a single 128 by 256 tile and never revisits the question. </p><p>And the <strong>PTX manual </strong>describes <code>tcgen05.relinquish_alloc_permit</code> as a promise that the CTA will make no further allocations, which Colfax glosses as allowing future CTAs to queue up for the same SM. The word queue is doing a lot of work in that sentence. </p><p>It is <em>consistent with a scheduler </em>that admits a new CTA only when the pool can serve it, which would collapse figure 6 toward 1 for any realistic tile regardless of the arithmetic.</p><p>We cannot resolve that without hardware, and we would rather show the bound and name the uncertainty than assert a residency number we have not seen. What survives either reading is the <strong>shape of the problem</strong>: the resource that limits parallelism on a Blackwell SM is acquired at runtime, quantizes in powers of two, and does not appear in any static occupancy model.</p><p>Now the part that makes low occupancy survivable, which is the more interesting half.</p><p>What occupancy hid was <em>memory latency</em>: a warp stalls on a load, the scheduler runs another warp. In a <strong>Blackwell GEMM</strong> the loads are done by the Tensor Memory Accelerator, asynchronously, into shared memory, signalled by an mbarrier. </p><p>The math is done by a <strong>single elected thread</strong> issuing an asynchronous instruction that reads shared memory and writes Tensor Memory. Neither heavy operation is a thread stalling on anything. Latency is hidden by <em>pipelining within one CTA</em>, using barriers and multiple buffers, rather than by <em>switching between CTAs</em>. </p><p>The author writing as gau-nernst describes exactly this in a working kernel: multiple <code>tcgen05.mma</code> in flight, one mbarrier per stage, so different MMA stages can be waited on independently.</p><p>The <strong>Delaware measurements</strong> make the same point from the other side. Single instruction latency for <code>tcgen05.mma</code> is 11.0 to 11.4 clocks and nearly flat across tile shapes from <code>m64n64k16</code> to <code>m256n256k16</code>. Hopper&#8217;s <code>wgmma</code> scales linearly with tile width, 32 clocks at <code>m64n64k16</code> and 128 at <code>m64n256k16</code>. </p><p>A flat latency across a <strong>sixteen fold range</strong> of tile area is the signature of a spatial array, not a deeper pipeline.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!2tlO!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89cac211-672c-4619-8795-02efdb7d63b2_1800x720.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!2tlO!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89cac211-672c-4619-8795-02efdb7d63b2_1800x720.png 424w, https://substackcdn.com/image/fetch/$s_!2tlO!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89cac211-672c-4619-8795-02efdb7d63b2_1800x720.png 848w, https://substackcdn.com/image/fetch/$s_!2tlO!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89cac211-672c-4619-8795-02efdb7d63b2_1800x720.png 1272w, https://substackcdn.com/image/fetch/$s_!2tlO!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89cac211-672c-4619-8795-02efdb7d63b2_1800x720.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!2tlO!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89cac211-672c-4619-8795-02efdb7d63b2_1800x720.png" width="1456" height="582" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/89cac211-672c-4619-8795-02efdb7d63b2_1800x720.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:582,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Grouped bar chart comparing single instruction latency of Hopper wgmma, which scales with tile width, against Blackwell tcgen05 which stays flat near eleven cycles&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Grouped bar chart comparing single instruction latency of Hopper wgmma, which scales with tile width, against Blackwell tcgen05 which stays flat near eleven cycles" title="Grouped bar chart comparing single instruction latency of Hopper wgmma, which scales with tile width, against Blackwell tcgen05 which stays flat near eleven cycles" srcset="https://substackcdn.com/image/fetch/$s_!2tlO!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89cac211-672c-4619-8795-02efdb7d63b2_1800x720.png 424w, https://substackcdn.com/image/fetch/$s_!2tlO!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89cac211-672c-4619-8795-02efdb7d63b2_1800x720.png 848w, https://substackcdn.com/image/fetch/$s_!2tlO!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89cac211-672c-4619-8795-02efdb7d63b2_1800x720.png 1272w, https://substackcdn.com/image/fetch/$s_!2tlO!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89cac211-672c-4619-8795-02efdb7d63b2_1800x720.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em><strong><sup>Figure 7.</sup></strong><sup> Hopper&#8217;s MMA latency grows with the tile. Blackwell&#8217;s does not. Combined with single thread issue and a TMEM resident accumulator, this is what makes very low CTA counts survivable.</sup></em></p><p>So the tuning knobs invert. On Hopper you asked how many warpgroups fit and how to specialize them. On Blackwell you ask how many columns your accumulator needs, how many buffers fit in what remains, and whether the tile is <strong>worth a CTA pair</strong>. </p><p>The Nsight metric that used to matter, achieved occupancy, tells you almost nothing. The metric that matters has no counter.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Quantization became an allocation problem</h2><p>Here is the part we think is genuinely underappreciated, and it is the point where <strong>this architecture stops being a matter for kernel authors</strong> and starts being a matter for anyone who chooses a serving format.</p><p>Blackwell implements block scaling in hardware. A block scaled MMA computes <em>D = C + (A x SFA) times (B x SFB)</em>, where SFA and SFB are vectors of scale factors, one per group of 16 or 32 elements along K. </p><p>The scale factors are not folded in beforehand and they are not applied afterwards. They are consumed by the tensor core as it runs. And, per the PTX ISA, <strong>they are consumed from Tensor Memory</strong>.</p><p>My disassembly shows this directly. The block scaled variants take a third and fourth <code>tmem[]</code> operand:</p><pre><code><code>// kind::mxf4nvf4.block_scale.scale_vec::4X, NVFP4 with block16 scaling
UTCOMMA.4X gdesc[UR14], gdesc[UR16], tmem[UR8], tmem[URZ], idesc[URZ], tmem[UR8], UPT ;

// kind::mxf4.block_scale.scale_vec::2X, MXFP4 with block32 scaling
UTCOMMA    gdesc[UR14], gdesc[UR16], tmem[UR8], tmem[URZ], idesc[URZ], tmem[UR8], UPT ;</code></code></pre><p>Read the consequence carefully. Your choice of numerical format now consumes the same scarce, power of two, 32 column granular, 512 column per SM resource that your accumulator consumes. </p><p>Finer scaling is not just more metadata bandwidth from HBM. It is columns you cannot use for the accumulator, for double buffering, or for a <strong>second CTA</strong>.</p><p>The two formats differ exactly where it hurts. MXFP4 follows the Open Compute microscaling specification: blocks of 32, scale in E8M0, a bare power of two exponent. </p><p><strong>NVFP4</strong> is NVIDIA&#8217;s own: blocks of 16, scale in E4M3, a real floating point number with a mantissa, plus a second level FP32 tensor-wide scale. Per the PTX rules, block32 must pair with E8M0, while block16 may use E8M0 or E4M3.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!MHla!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ea35e78-0dc7-4686-ba58-ac915d1957db_1800x760.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!MHla!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ea35e78-0dc7-4686-ba58-ac915d1957db_1800x760.png 424w, https://substackcdn.com/image/fetch/$s_!MHla!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ea35e78-0dc7-4686-ba58-ac915d1957db_1800x760.png 848w, https://substackcdn.com/image/fetch/$s_!MHla!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ea35e78-0dc7-4686-ba58-ac915d1957db_1800x760.png 1272w, https://substackcdn.com/image/fetch/$s_!MHla!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ea35e78-0dc7-4686-ba58-ac915d1957db_1800x760.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!MHla!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ea35e78-0dc7-4686-ba58-ac915d1957db_1800x760.png" width="1456" height="615" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/1ea35e78-0dc7-4686-ba58-ac915d1957db_1800x760.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:615,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Chart of effective bits per parameter and scale factor overhead for MXFP8, MXFP4 and NVFP4&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Chart of effective bits per parameter and scale factor overhead for MXFP8, MXFP4 and NVFP4" title="Chart of effective bits per parameter and scale factor overhead for MXFP8, MXFP4 and NVFP4" srcset="https://substackcdn.com/image/fetch/$s_!MHla!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ea35e78-0dc7-4686-ba58-ac915d1957db_1800x760.png 424w, https://substackcdn.com/image/fetch/$s_!MHla!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ea35e78-0dc7-4686-ba58-ac915d1957db_1800x760.png 848w, https://substackcdn.com/image/fetch/$s_!MHla!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ea35e78-0dc7-4686-ba58-ac915d1957db_1800x760.png 1272w, https://substackcdn.com/image/fetch/$s_!MHla!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ea35e78-0dc7-4686-ba58-ac915d1957db_1800x760.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em><strong><sup>Figure 8</sup></strong><sup> NVFP4 halves the block size and puts a mantissa in the scale, which is why it is more accurate than MXFP4. It also doubles the scale factor footprint, and that footprint lives in the same 512 column pool as the accumulator.</sup></em></p><p>This is a co-design decision hiding inside a numerics decision. NVIDIA defined a format whose <strong>accuracy advantage</strong> over the open standard comes from finer blocks and richer scales, then built the only silicon where those scales are consumed directly out of a dedicated on-chip memory rather than being unpacked into registers first. </p><p>On hardware without that memory, the same format is a software dequantization problem with a register cost. On Blackwell it is an operand.</p><p>The Delaware group reports the accuracy side: FP8 costs about 2 percent perplexity on Mistral 7B and Mixtral 8x7B, FP4 costs 8 to 9 percent, and their FP4 throughput is 2.5 times FP16 on the dense model and 2.7 times on the mixture of experts model. </p><p>Whether <strong>8 percent perplexity</strong> is acceptable is a per-layer question and always was. My point is narrower: on this architecture the answer is also a per-SM capacity planning question, because the scale factors and the accumulator compete for the same 512 columns.</p><p>Numerics stopped being a property of the model and became a property of the memory allocator.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>What this does to serving economics</h2><p>Everything above is architecture. Here is the part that shows up on an invoice.</p><p><strong>Two rules</strong> from earlier sections collide here. The UMMA shape table offers exactly two values of M for a single CTA, 64 and 128; there is nothing smaller. </p><p>And allocation is by whole column, so the accumulator&#8217;s footprint is fixed by the tile you chose, not by the rows you filled.</p><p>During prefill neither rule bites. M is the <strong>token count of a chunk</strong>, thousands of rows, and the tile is full. During decode both bite at once. A decode step is a GEMM whose M is the number of sequences you are batching, and it is skinny by construction, so you pick the 64-row tile and leave most of it empty.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!20ki!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9031d001-8aac-4140-8df0-e8d90d0da36c_1800x800.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!20ki!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9031d001-8aac-4140-8df0-e8d90d0da36c_1800x800.png 424w, https://substackcdn.com/image/fetch/$s_!20ki!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9031d001-8aac-4140-8df0-e8d90d0da36c_1800x800.png 848w, https://substackcdn.com/image/fetch/$s_!20ki!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9031d001-8aac-4140-8df0-e8d90d0da36c_1800x800.png 1272w, https://substackcdn.com/image/fetch/$s_!20ki!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9031d001-8aac-4140-8df0-e8d90d0da36c_1800x800.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!20ki!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9031d001-8aac-4140-8df0-e8d90d0da36c_1800x800.png" width="1456" height="647" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/9031d001-8aac-4140-8df0-e8d90d0da36c_1800x800.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:647,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Chart showing what fraction of the smallest legal UMMA accumulator tile carries a real sequence during decode, rising from 1.6 percent at batch one to 100 percent at batch sixty four&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Chart showing what fraction of the smallest legal UMMA accumulator tile carries a real sequence during decode, rising from 1.6 percent at batch one to 100 percent at batch sixty four" title="Chart showing what fraction of the smallest legal UMMA accumulator tile carries a real sequence during decode, rising from 1.6 percent at batch one to 100 percent at batch sixty four" srcset="https://substackcdn.com/image/fetch/$s_!20ki!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9031d001-8aac-4140-8df0-e8d90d0da36c_1800x800.png 424w, https://substackcdn.com/image/fetch/$s_!20ki!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9031d001-8aac-4140-8df0-e8d90d0da36c_1800x800.png 848w, https://substackcdn.com/image/fetch/$s_!20ki!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9031d001-8aac-4140-8df0-e8d90d0da36c_1800x800.png 1272w, https://substackcdn.com/image/fetch/$s_!20ki!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9031d001-8aac-4140-8df0-e8d90d0da36c_1800x800.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em><strong><sup>Figure 9</sup></strong><sup> The decode penalty, stated against the most favourable tile the instruction set offers. At a batch of eight, an ordinary steady state for a latency sensitive endpoint, an eighth of the reserved accumulator carries anything. The rest is allocated, idle, and unavailable to any other block on that SM.</sup></em></p><p>Be precise about what is and is not new, because it is easy to overstate. The 64-row floor is not new: Hopper&#8217;s <code>wgmma</code> had exactly the same minimum M, and skinny GEMMs have underused tensor cores since Volta. </p><p>That is why <strong>prefill and decode</strong> <strong>get disaggregated</strong> onto separately provisioned pools in the first place, which we argued at length in a previous issue.</p><p>What is new is where the waste is recorded. Previously it was purely a <em>throughput</em> loss: you issued an MMA whose M dimension was mostly zeros and you got a fraction of peak flops, and the moment the instruction retired the machine was free again. </p><p>Now the same tile also holds a <em>capacity</em> reservation, in a 512 column pool, acquired through an atomic that other CTAs may be spinning on, held for the tile&#8217;s lifetime, and invisible to every static occupancy model. Underutilisation stopped being a transient and became an allocation.</p><p><strong>Three practical consequences</strong> follow, and we hold them with decreasing confidence.</p><ol><li><p><strong>The first is that Blackwell widens the gap between prefill and decode economics</strong> rather than narrowing it, in a generation whose marketing is entirely about inference. The FP4 tensor cores are a prefill and large-batch story, and section 05 is the reason: the precision ladder is bandwidth neutral on chip, so what FP4 buys you is math you can only spend if you have rows to fill. Decode has neither. It remains bound by HBM bandwidth for weight movement, and now carries an on-chip capacity reservation on top. This is consistent with the Delaware measurements from a different angle: as precision drops from FP16 to FP4 their measured memory bandwidth utilization <em>falls</em> from 67 percent to 48 percent, which is what it looks like when a workload stops being bandwidth bound and starts being bound by something else.</p></li><li><p><strong>The second</strong> <strong>is that this pushes harder</strong> <strong>toward every technique that manufactures M</strong>. Speculative decoding turns one sequence into k candidate tokens per step, and on Blackwell it is not only amortizing weight reads across more rows, it is filling rows of a tile that was reserved whether or not you filled them. </p><blockquote><p><em><strong>Figure 9</strong>, read the other way, is a chart of how much accumulator a speculative draft gets for free. We modelled speculation as a throughput question in a previous issue. There is a second term in that model now, and it points the same way.</em></p></blockquote></li><li><p><strong>The third, and the one we are least sure of, is that grouped GEMM for mixture of experts becomes a harder allocation problem</strong> than it was. Each expert&#8217;s tile has its own M, determined by routing, varying per step. If your kernel allocates for the worst case it wastes columns on every expert that got fewer tokens; if it allocates per expert it pays the atomic repeatedly. </p></li></ol><p>The one worklog we have found on optimizing NVFP4 grouped GEMM on Blackwell, by Mufeez Amjad, lands on <code>cta_group::2</code> with a shared TMEM allocation across the CTA pair, and describes the shared allocation as one of the main benefits rather than the wider math tile. That is a hint about where the pressure actually is.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>What Rubin says about whether this was a one-off</h2><p>A reasonable objection to everything above is that <strong>Tensor Memory might be a Blackwell specific hack</strong>, a way to ship a doubled MMA without redesigning the register file, and that the next architecture folds it back into something more general. </p><p>NVIDIA published enough about <strong>Rubin on 21 July 2026 </strong>to answer that. It is not going away, and every disclosed change points at the numbers in section 05.</p><p>Three details in NVIDIA&#8217;s own post are load bearing.</p><ol><li><p><strong>Tensor Memory gained a consumer that is not the MMA.</strong> In the long context attention path, the intermediate scores from the dense QK transpose are, in NVIDIA&#8217;s words, loaded from Tensor Memory into a structured 2 to 4 sparse compressed form, generating both the nonzero values and the metadata. TMEM is no longer only where accumulators land. It is a staging tier that a hardware compression path reads from. That is the same direction our disassembly already showed on Blackwell, where sparsity metadata and block scale factors are already TMEM operands.</p></li><li><p><strong>Rubin doubles the K dimension per tensor core instruction.</strong> NVIDIA frames this as fewer K loop iterations and less loop overhead, which is true and is the reason a reader would care. Run it through section 05 and it is also something else. Accumulator bytes per FLOP is 4 over K. Doubling K halves it. If Rubin&#8217;s NVFP4 K goes from 64 to 128, accumulator traffic per unit of math drops by half at exactly the moment the math rate goes up. That is the same invariance trick applied across a generation instead of across a precision ladder, and it is the strongest evidence we have that accumulator bandwidth is a first order design constraint at NVIDIA rather than a consequence.</p></li><li><p><strong>Softmax became the bottleneck, so they widened it.</strong> Rubin raises exponential throughput per clock per SM by 2 times for FP32 and 4 times for BF16 against Blackwell, with Blackwell Ultra at 2 times for both. You only build that if the matrix path has already pulled far enough ahead that the transcendental path is what is left.</p></li></ol><p>Alongside those, Rubin adds inline descriptor updates for the Tensor Memory Accelerator, so a <strong>mixture of experts kernel </strong>keeps one descriptor and overrides the pointer and stride fields in the instruction instead of rewriting a descriptor in memory per expert, and counted writes for <strong>device initiated NVLink transfers</strong> so the receiver tracks completion without the barrier, acknowledgement and atomic flag sequence.</p><p>Every one of those is the same move: take a coordination cost that was paid in <strong>general purpose instructions and registers</strong>, and pay it in a dedicated mechanism instead. Tensor Memory was the first large instance. Rubin is the pattern applied to descriptors, to synchronization, to sparsity metadata and to the softmax path.</p><p>One sentence for what NVIDIA is doing to the SM across these two generations: they are disassembling the<strong> general purpose core</strong> into special purpose engines connected by dedicated memories and barriers, and leaving the threads to do bookkeeping. </p><p>The execution model is being hollowed out from the inside while its surface syntax stays the same.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Where this could be wrong</h2><ul><li><p><strong>The accumulator traffic derivation assumes a full read and write per instruction.</strong> If the tensor core retains partial accumulator state internally across the K extent, or streams only the sub-tile currently in flight, then 562.5 terabytes per second is an upper bound. What makes us reasonably confident is not the absolute number but the invariance: the cancellation between element width and K is exact, holds at three precisions, and falls out of a rule visible in CUTLASS source. A different traffic model would change the constant and probably not the invariance. Note that this is a revision of an earlier draft of ours, which held K fixed at 16 across precisions and therefore produced a per-SM figure four times too large at FP4.</p></li><li><p><strong>We are reading the SASS of a probe, not of a real kernel.</strong> A kernel with one MMA gives <code>ptxas</code> no scheduling problem to solve. The allocation spin loop may be scheduled very differently, hoisted, or dominated by something else in a CUTLASS mainloop with a producer warpgroup and four pipeline stages. We are confident the instruction sequence exists and that the guardrails are unconditional. We are not confident about its cost in situ, and the 152 instruction figure is a property of our probe, not of production kernels.</p></li><li><p><strong>Figure 6 is an upper bound, not a residency measurement.</strong> The allocation arithmetic is exact, and the documented block and warp ceilings for compute capability 10.0 are exact. What we cannot verify is whether the scheduler will actually co-resident the CTAs the pool arithmetically admits. One published source states that SM10x is locked to one CTA per SM outright, and the PTX description of <code>relinquish_alloc_permit</code> hints at queueing rather than co-residency. If that reading is right, the correct version of figure 6 is a flat line at 1, the argument gets stronger rather than weaker, and our chart is still wrong.</p></li><li><p><strong>Figure 9 measures reservation, not throughput, and only for the tile we chose.</strong> The 64 row floor is documented and the arithmetic against it is exact, but a real decode kernel may prefer a different shape, may reuse one allocation across many tiles, or may batch several matrices into a single call in ways that change what fraction of the pool sits idle at any instant. The claim we will defend is narrow: below 64 sequences the accumulator is reserved for rows that do not exist. What that costs in dollars depends on a serving stack we have not profiled.</p></li><li><p><strong>The Blackwell per-SM per-clock figure of 8,192 is inferred, not published.</strong> Volta, Ampere and Hopper rates come from NVIDIA whitepapers and reproduce the published teraflops to three digits. For Blackwell we ran the identity backwards from 2.25 petaflops over 148 SMs, which requires 8,192 FLOPs per clock at 1.86 GHz. If the real SM count in the shipping part differs, or if the datasheet figure assumes a different clock, the doubling claim survives but the exact number moves.</p></li><li><p><strong>The portability finding is solid; the moat reading of it is not.</strong> Section 04 already discounts it and we would discount it further rather than less. The instruction family genuinely does not exist outside datacenter Blackwell. What follows from that commercially is much weaker than it first sounds.</p></li></ul><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share The Software Frontier&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share The Software Frontier</span></a></p><div><hr></div><h2>Four predictions, each with its failure condition</h2><p>Deliberately conservative, and each one checkable by a specific date against a specific artifact.</p><ol><li><p><strong>NVIDIA will not expose Tensor Memory to a general purpose instruction through the end of 2028.</strong> No <code>ld</code>, <code>st</code>, atomic or ALU operand outside the <code>tcgen05</code> family, in Rubin or its successor. The invariance in section 05 is the reason: general purpose addressability is incompatible with the bandwidth. <em>Wrong if</em> a PTX ISA revision adds any TMEM access outside the dedicated opcode family.</p></li><li><p><strong>By 31 December 2027, no open source compiler will generate a </strong><code>tcgen05</code><strong> GEMM within 10 percent of CUTLASS on a mainstream shape without hand written PTX in its lowering path.</strong> Reaching the instruction is already happening. Reaching it through a general lowering that models column allocation, the warp-lane access partition and the drain width tradeoff is a different problem. <em>Wrong if</em> a mainline release of Triton, Mojo or tinygrad hits the bar with a pure compiler path.</p></li><li><p><strong>Nsight Compute will ship a Tensor Memory occupancy or column pressure section before the end of 2028.</strong> The hardware is already tracking the pool, since the allocator does a find-and-set against it, and there is currently no static way to reason about a resource acquired at runtime. <em>Wrong if</em> no NVIDIA profiler release by then reports TMEM allocation state.</p></li><li><p><strong>NVFP4 will remain the default four bit format in NVIDIA&#8217;s own inference libraries through 2027, and MXFP4 will not displace it there.</strong> Not an accuracy argument: the block16 path is the one with a dedicated opcode modifier and a hardware operand path on the silicon that matters. <em>Wrong if</em> TensorRT-LLM or NVIDIA&#8217;s quantization tooling defaults to block32 for four bit weights.</p></li></ol><p>We have dropped two predictions that appeared in an earlier draft, on model architectures being shaped around column boundaries and on serving share ratios between formats, because neither had a failure condition we could actually check.</p><div class="directMessage button" data-attrs="{&quot;userId&quot;:201922174,&quot;userName&quot;:&quot;Lorenzo Bradanini&quot;,&quot;canDm&quot;:null,&quot;dmUpgradeOptions&quot;:null,&quot;isEditorNode&quot;:true}" data-component-name="DirectMessageToDOM"></div><div><hr></div><h2>Confidence dossier</h2><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!uiBL!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1d544e02-60d4-4544-88ef-dd831cd8be82_2700x2472.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!uiBL!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1d544e02-60d4-4544-88ef-dd831cd8be82_2700x2472.png 424w, https://substackcdn.com/image/fetch/$s_!uiBL!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1d544e02-60d4-4544-88ef-dd831cd8be82_2700x2472.png 848w, https://substackcdn.com/image/fetch/$s_!uiBL!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1d544e02-60d4-4544-88ef-dd831cd8be82_2700x2472.png 1272w, https://substackcdn.com/image/fetch/$s_!uiBL!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1d544e02-60d4-4544-88ef-dd831cd8be82_2700x2472.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!uiBL!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1d544e02-60d4-4544-88ef-dd831cd8be82_2700x2472.png" width="1456" height="1333" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/1d544e02-60d4-4544-88ef-dd831cd8be82_2700x2472.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1333,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:192773,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/208027653?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1d544e02-60d4-4544-88ef-dd831cd8be82_2700x2472.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!uiBL!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1d544e02-60d4-4544-88ef-dd831cd8be82_2700x2472.png 424w, https://substackcdn.com/image/fetch/$s_!uiBL!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1d544e02-60d4-4544-88ef-dd831cd8be82_2700x2472.png 848w, https://substackcdn.com/image/fetch/$s_!uiBL!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1d544e02-60d4-4544-88ef-dd831cd8be82_2700x2472.png 1272w, https://substackcdn.com/image/fetch/$s_!uiBL!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1d544e02-60d4-4544-88ef-dd831cd8be82_2700x2472.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Reproduce every tier A row on a machine with no GPU</h2><p>Everything in tier A came out of the following, in a clean container, in under five minutes, with no CUDA installation and no driver.</p><pre><code><code># 1. the toolchain, as ordinary x86 binaries from PyPI
pip download nvidia-cuda-nvcc-cu12 --no-deps -d /tmp/w            # ptxas 12.9.86
pip download nvidia-cuda-nvdisasm nvidia-cuda-cuobjdump --no-deps -d /tmp/w   # 13.3.73
mkdir -p /tmp/tk &amp;&amp; cd /tmp/tk &amp;&amp; for f in /tmp/w/*.whl; do unzip -oq "$f"; done
PTXAS=/tmp/tk/nvidia/cuda_nvcc/bin/ptxas
DIS=/tmp/tk/nvidia/cu13/bin/nvdisasm

# A1, A2, A3, A4: assemble the probe and read the machine code
$PTXAS -arch=sm_100a probe.ptx -o probe.cubin &amp;&amp; $DIS -c probe.cubin | less
$DIS -c probe.cubin | grep -oE '__cuda_sm10x_tcgen05_[a-z_]+' | sort -u
for o in 0 1 2 3; do $PTXAS -O$o -arch=sm_100a probe.ptx -o g.cubin
  echo "-O$o $($DIS -c g.cubin | grep -c guardrail) refs, \
$($DIS -c g.cubin | grep -cE '^\s+/\*[0-9a-f]{4}\*/') instructions"; done

# A1 continued: one build per .kind qualifier
for k in f16 tf32 f8f6f4 i8; do sed "s/kind::f16/kind::$k/" probe.ptx &gt; k.ptx
  $PTXAS -arch=sm_100a k.ptx -o k.cubin &amp;&amp; $DIS -c k.cubin | grep -oE 'UTC[A-Z0-9.]+'; done

# A5: three instruction families across five targets
for a in sm_90a sm_100a sm_103a sm_100 sm_120a; do
  for f in probe legacy_wgmma legacy_mmasync; do
    printf '%-9s %-16s ' $a $f; $PTXAS -arch=$a $f.ptx -o /dev/null 2&gt;&amp;1 | head -1; echo; done; done

# A6: minimum PTX ISA version
for v in 8.3 8.4 8.5 8.6 8.7 8.8; do sed "s/^.version 8.6/.version $v/" probe.ptx &gt; v.ptx
  echo -n "$v "; $PTXAS -arch=sm_100a v.ptx -o /dev/null 2&gt;&amp;1 | head -1; echo; done

# A7: the epilogue register curve. widen tcgen05.ld to .xN and store every register
for n in 1 2 4 8 16 32 64 128; do $PTXAS -arch=sm_100a -v ld_$n.ptx -o /dev/null 2&gt;&amp;1 | grep Used; done

# A8, A9: the ceilings the compiler will and will not enforce
$PTXAS -arch=sm_100a -maxrregcount=256 probe.ptx -o /dev/null
for c in 16 48 96 1024; do sed "s/mov.u32 %r1, 128;/mov.u32 %r1, $c;/" probe.ptx &gt; a.ptx
  $PTXAS -arch=sm_100a a.ptx -o /dev/null &amp;&amp; echo "$c columns: accepted"; done</code></code></pre><p><strong>Two practical notes</strong>. <code>ptxas</code> from the 12.9 wheel caps at PTX ISA 8.8, which covers <code>sm_103a</code> and nothing above it, so pull a newer <code>nvidia-cuda-nvcc</code> for later targets. And <code>nvdisasm</code> from the CUDA 13 wheels reads CUDA 12 cubins, which is convenient because the CUDA 13 nvcc wheel did not build in our container while the standalone disassembler wheels did.</p><p>The <strong>value of this workflow</strong> is not that it replaces a GPU. It is that it separates two questions that get conflated constantly in GPU writing: what the machine <em>does</em>, which needs hardware, and what the compiler <em>emits</em>, which does not. </p><p>A surprising share of public claims about the CUDA moat are <strong>claims of the second kind</strong>, and the second kind is checkable by anyone with forty megabytes of disk.</p><div><hr></div><h2>Bibliography</h2><ol><li><p>NVIDIA, <em>Parallel Thread Execution ISA</em>: tensor memory addressing, tcgen05 MMA and its kind shapes, shared memory descriptors, instruction descriptors, data path layout organization, and the tcgen05 memory consistency model. The primary source for everything structural here.</p></li><li><p>NVIDIA, <em>Blackwell Tuning Guide</em> and <em>Ampere Tuning Guide</em>. Register file size, the register per thread ceiling, warp and thread block limits per SM, and shared memory capacities for compute capabilities 8.0, 10.0 and 12.0.</p></li><li><p>NVIDIA Volta, Ampere and Hopper architecture whitepapers. Source for 1,024, 2,048 and 4,096 FP16 tensor FLOPs per clock per SM, each of which reproduces the corresponding published teraflops figure to three digits.</p></li><li><p>NVIDIA, <em>HGX B200 datasheet</em>. Dense and sparse rates per precision, with the explicit note that dense is half of sparse, plus memory capacity and bandwidth.</p></li><li><p>Ryo, <em>CUTLASS Tutorial: Writing GEMM Kernels Using Tensor Memory For NVIDIA Blackwell GPUs</em>, Colfax Research, April 2025, updated November 2025. The clearest published account of TMEM allocation, UMMA operand rules and the CuTe abstractions over both.</p></li><li><p>Colfax Research, <em>CUTLASS Tutorial: Hardware-supported Block-scaling with NVIDIA Blackwell GPUs</em>. Scale vector rules and TMEM layouts of scale factors.</p></li><li><p>Colfax Research, <em>NVFP4 Blockscaled GEMM on NVIDIA RTX Pro Blackwell GPUs (SM12x)</em>, June 2026. The explicit statement that SM12x has neither tcgen05 nor TMEM, and that SM10x is locked to one CTA per SM.</p></li><li><p>NVIDIA, <em>CUTLASS</em> source, <code>include/cute/atom/mma_traits_sm100.hpp</code> and <code>copy_traits_sm100.hpp</code>. The 256 bit K rule, the TMEM copy atoms, and the ThrID change.</p></li><li><p>Aaron Jarmusch and Sunita Chandrasekaran, University of Delaware, <em>Microbenchmarking NVIDIA&#8217;s Blackwell Architecture: An in-depth Architectural Analysis</em>, arXiv 2512.02189. The only systematic public measurement of B200 tensor core latency, TMEM behaviour and FP4 accuracy we are aware of.</p></li><li><p>NVIDIA, <em>Inside NVIDIA Rubin GPU Architecture: Powering the Era of Agentic AI</em>, 21 July 2026.</p></li><li><p>NVIDIA, <em>CUTLASS documentation, Blackwell SM100 functionality</em>. The seven tcgen05.mma instructions and the block scaled data types.</p></li><li><p>SemiAnalysis, <em>Dissecting NVIDIA Blackwell: Tensor Cores, PTX Instructions, SASS, Floorsweep, Yield</em>. The CuTe ThrID observation and the TPC scoped CTA pair framing.</p></li><li><p>gau-nernst, <em>tcgen05 for dummies</em>, December 2025. A working tutorial in plain CUDA and inline PTX.</p></li><li><p>Mufeez Amjad, <em>Optimizing NVFP4 Grouped GEMM on Blackwell</em>, worklog, March 2026. The CTA pair and shared TMEM allocation observation in section 08.</p></li><li><p>Rouhani et al., <em>Microscaling Data Formats for Deep Learning</em>, arXiv 2310.10537. The MX specification behind MXFP4 and MXFP8.</p></li><li><p>NVIDIA developer forums, thread on computing tensor core FP16 throughput per SM per clock from whitepaper figures. The identity used in section 01.</p></li><li><p>Chips and Cheese, <em>Nvidia&#8217;s B200: Keeping the CUDA Juggernaut Rolling</em>, December 2025. Die level SM counts: 74 enabled of 80 per die.</p></li><li><p>Austin et al., <em>How To Scale Your Model</em>, GPU chapter. Independent derivation of Blackwell&#8217;s per-SM per-clock tensor rate from the same datasheet figures.</p></li><li><p>Our own previous work: <em>How the NVIDIA Compiler Moat Actually Works</em>, <em>How CUDA Binaries Actually Work</em>, <em>The Split and the Seam</em> on prefill and decode disaggregation, and <em>The Draft and the Ledger</em> on speculative decoding economics.</p><p></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share The Software Frontier&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share The Software Frontier</span></a></p><p></p></li></ol>]]></content:encoded></item><item><title><![CDATA[How CUDA Binaries Actually Work]]></title><description><![CDATA[A byte-level surgical dissection of the cubin and fatbin formats, and the second encoding nobody documents.]]></description><link>https://www.thesoftwarefrontier.com/p/how-cuda-binaries-actually-work</link><guid isPermaLink="false">https://www.thesoftwarefrontier.com/p/how-cuda-binaries-actually-work</guid><dc:creator><![CDATA[Lorenzo Bradanini]]></dc:creator><pubDate>Fri, 31 Jul 2026 12:45:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!xYOt!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F68806f07-c2e9-44a7-a9db-9dd6ff086db4_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!xYOt!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F68806f07-c2e9-44a7-a9db-9dd6ff086db4_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!xYOt!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F68806f07-c2e9-44a7-a9db-9dd6ff086db4_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!xYOt!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F68806f07-c2e9-44a7-a9db-9dd6ff086db4_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!xYOt!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F68806f07-c2e9-44a7-a9db-9dd6ff086db4_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!xYOt!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F68806f07-c2e9-44a7-a9db-9dd6ff086db4_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!xYOt!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F68806f07-c2e9-44a7-a9db-9dd6ff086db4_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/68806f07-c2e9-44a7-a9db-9dd6ff086db4_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2689230,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/207127260?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F68806f07-c2e9-44a7-a9db-9dd6ff086db4_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!xYOt!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F68806f07-c2e9-44a7-a9db-9dd6ff086db4_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!xYOt!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F68806f07-c2e9-44a7-a9db-9dd6ff086db4_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!xYOt!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F68806f07-c2e9-44a7-a9db-9dd6ff086db4_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!xYOt!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F68806f07-c2e9-44a7-a9db-9dd6ff086db4_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>Intro</h2><p>I spent a good part of last year writing an <strong>x86-64 assembler</strong> by hand, in C, for no more reason than wanting to know what an object file really is. The thing that surprised me was not the instruction encoding, but was the metadata. </p><p>A <strong>modern object file</strong> is essentially a negotiation between the compiler and the loader about few things the machine code itself cannot say properly: where the arguments are, how much stack to reserve, which symbols are entry points, what has to be patched before anything runs.</p><p>So when I started reading GPU binaries, the question I kept asking more and more was not &#8220;<em>what does this instruction do.</em>&#8221; It was &#8220;<em>what is this file promising the driver</em>.&#8221;</p><p>The answer turns out to be&#8230; well, a lot, and almost none of it is written down. NVIDIA just documents the container in one sentence and the disassembler in twelve pages. And that&#8217;s basically it. </p><p>The <a href="https://docs.nvidia.com/cuda/cuda-binary-utilities/">CUDA Binary Utilities</a> manual says a cubin is &#8220;<em>an ELF-formatted file which consists of CUDA executable code sections as well as other sections containing symbols, relocators, debug info, etc.</em>&#8221; That is the specification, and the word &#8220;etc.&#8221; is where the whole platform lives.</p><p>I want you to imagine this piece as a medical dissection, if it&#8217;s possible to say so. Every single number, hex dump and section listing below, came out of a container on <strong>my own machine</strong>, and the exact commands are in the appendix so you can disagree with me precisely. </p><p>Two things made it possible: first thing is that you do not need a GPU to compile, link or disassemble CUDA binaries: <code>ptxas</code>, <code>fatbinary</code>, <code>nvdisasm</code> and <code>cuobjdump</code> are a bunch of ordinary x86 programs. The second is that the entire toolchain is on PyPI, so getting a byte-exact CUDA 13.3.73 install is actually just one <code>pip install</code>.</p><h3>Methodology</h3><p>Everything measured here is from CUDA 13.3.73 (built 9 June 2026) installed from the <code>nvidia-cuda-nvcc</code>, <code>nvidia-cuda-cuobjdump</code> and <code>nvidia-cuda-nvdisasm</code> wheels on x86-64 Linux, plus the cuBLAS 13.6.0.2 wheel for the library dissection. </p><p>I do not currently have a GPU, which means every claim here is about what the toolchain <em>emits</em> and not about what silicon does with it. There are a few parts where I&#8217;m reading structure out of bytes instead of documentation, dont&#8217;t worry: I&#8217;ll say so and I give it like a confidence tier at the end. </p><p>NVIDIA publishes no specification for the cubin section layout, the <code>.nv.info</code> attribute encoding, the fatbin entry header or the Mercury sections, and all the interpretations are mine, unless specified.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Seven processes to compile one file</h2><p>I want you to introduce the format and the pipeline now, because the shape of the output is derived from it. Keep in mind that <code>nvcc</code> is not a compiler, it&#8217;s just a driver that<strong> orchestrates other programs</strong>, and it will show you its plan.</p><pre><code>$ nvcc -arch=sm_90 -c -o k.o saxpy.cu --dryrun
gcc      -D__CUDA_ARCH_LIST__=900 -E -x c++ -D__CUDACC__ ...        # host preprocess
<strong>cudafe++</strong> --c++17 --static-host-stub --device-hidden-visibility ...  # split host/device
gcc      -D__CUDA_ARCH__=900 -E -x c++ -DCUDA_DOUBLE_MATH_FUNCTIONS # device preprocess
<strong>cicc</strong>     ... saxpy.cudafe1.gpu -o saxpy.ptx                        # C++ front end -&gt; PTX
<strong>ptxas</strong>    -arch=sm_90 -m64 saxpy.ptx -o saxpy.sm_90.cubin           # PTX -&gt; SASS
<strong>fatbinary</strong> -64 --cicc-cmdline=... --image3=...  -o saxpy.fatbin.c    # package
gcc      -c -x c++ saxpy.cudafe1.cpp -o k.o                        # host compile</code></pre><p>We have <strong>two preprocessor passes</strong> over the same source with &#8220;different macros&#8221;, then a source-to-source splitter, a front end that emits PTX, an assembler that (you know) emits machine code, a packager that turns the machine code into a C array, and a host compiler that swallows the result. </p><p>The <em>artifacts in the middle</em> are real files you can save with <code>--keep</code>, and each one is a place where the format is decided.</p><p>The only thing that matters is where the close part actually is. We alredy know that <code>cicc</code> is the <strong>NVVM-based front end</strong> and it produces PTX, which is a published virtual instruction set, with all docs that you can read. But for <code>ptxas,</code>which takes PTX and produces the thing this article is about, we don&#8217;t know nothing. </p><p>Everything that <strong>works below PTX </strong>is the vendor&#8217;s private business, and everything upstream has alredy been made public: Clang, Triton, Julia, Mojo and tinygrad all emit PTX, and NVIDIA itself ships <code>libNVVM</code>, so they can do it. </p><p>The interesting feature in the CUDA stack is just one process wide.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>The container is an ELF file that lies about being one</h2><p>Now, let&#8217;s compile a kernel to a standalone cubin and we will see that the <code>file</code> recognizes it immediately.</p><pre><code>$ nvcc -arch=sm_90 -cubin -o saxpy.sm90.cubin saxpy.cu
$ file saxpy.sm90.cubin
saxpy.sm90.cubin: ELF 64-bit LSB executable, <strong>NVIDIA CUDA architecture</strong>, version 1,
                  statically linked, not stripped</code></pre><p>It is a <strong>real ELF64</strong>, little endian, and <code>readelf</code> will walk its section table happily. Three fields in the 64-byte header carry the CUDA-specific part.</p><pre><code>e_ident:  7f 45 4c 46 02 01 01 <strong>41</strong> <strong>08</strong> 00 00 00 00 00 00 00
                              |  |
                              |  +-- EI_ABIVERSION = 8
                              +----- EI_OSABI      = 0x41
e_type    = 2        (ET_EXEC; ET_REL when compiled with -rdc=true)
e_machine = 190      (0xbe, EM_CUDA)
e_flags   = <strong>0x06005a04</strong></code></pre><p><code>EM_CUDA = 190</code> is the one part of this that is available to everybody: it sits in LLVM&#8217;s <code>BinaryFormat/ELF.h</code> alongside every other machine type, because LLVM needs to read these files. </p><p>The <strong>OS ABI byte</strong> is where things get strange. That same header defines two CUDA values, <code>ELFOSABI_CUDA = 51</code> and <code>ELFOSABI_CUDA_V2 = 41</code>, both written in decimal. Every cubin my 13.3 toolchain produces carries 65, which is <code>0x41</code>. 51 is <code>0x33</code>, so the first constant is the older value written in decimal, and the second&#8230;. looks like <code>0x41</code> transcribed as if it were decimal. </p><p>Obv.  I don&#8217;t know where is the mistake, but if you are writing a tool that <strong>sniffs cubins</strong>, sniff for the byte <code>0x41</code> and don&#8217;t trust either constant. The ABI version byte is 8.</p><p> I know that<code> e_flags</code> is where the architecture lives, and the field names are recoverable from the disassembler: <code>EF_CUDA_SM</code>, <code>EF_CUDA_PTX_SM</code>, <code>EF_CUDA_64BIT_ADDRESS</code> and <code>EF_CUDA_ACCELERATORS</code>, alongside an <code>EF_CUDA_SM10</code> through <code>EF_CUDA_SM121</code> enumeration that covers fifteen years of hardware in one list. </p><p>The layout is easy peasy to recover by<strong> sweeping every target</strong> the compiler supports.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!36uk!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F05834540-c5d9-49a8-a0a9-f60a401bb806_1939x933.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!36uk!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F05834540-c5d9-49a8-a0a9-f60a401bb806_1939x933.png 424w, https://substackcdn.com/image/fetch/$s_!36uk!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F05834540-c5d9-49a8-a0a9-f60a401bb806_1939x933.png 848w, https://substackcdn.com/image/fetch/$s_!36uk!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F05834540-c5d9-49a8-a0a9-f60a401bb806_1939x933.png 1272w, https://substackcdn.com/image/fetch/$s_!36uk!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F05834540-c5d9-49a8-a0a9-f60a401bb806_1939x933.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!36uk!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F05834540-c5d9-49a8-a0a9-f60a401bb806_1939x933.png" width="1456" height="701" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/05834540-c5d9-49a8-a0a9-f60a401bb806_1939x933.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:701,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Left: the e_flags word of nine cubins broken into four bytes, showing the compute capability in byte 1 and a byte 0 that changes at Blackwell. Right: cubin size for each target.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Left: the e_flags word of nine cubins broken into four bytes, showing the compute capability in byte 1 and a byte 0 that changes at Blackwell. Right: cubin size for each target." title="Left: the e_flags word of nine cubins broken into four bytes, showing the compute capability in byte 1 and a byte 0 that changes at Blackwell. Right: cubin size for each target." srcset="https://substackcdn.com/image/fetch/$s_!36uk!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F05834540-c5d9-49a8-a0a9-f60a401bb806_1939x933.png 424w, https://substackcdn.com/image/fetch/$s_!36uk!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F05834540-c5d9-49a8-a0a9-f60a401bb806_1939x933.png 848w, https://substackcdn.com/image/fetch/$s_!36uk!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F05834540-c5d9-49a8-a0a9-f60a401bb806_1939x933.png 1272w, https://substackcdn.com/image/fetch/$s_!36uk!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F05834540-c5d9-49a8-a0a9-f60a401bb806_1939x933.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>The same four-line kernel, one target per row. Bits 8 to 15 hold the compute capability as a plain integer. Byte 0 flips from 0x04 to 0x02 at the Blackwell boundary, and the file grows by 44% across that same boundary for identical source.</em></p><p>This is actually very good, but at least three things don&#8217;t match reality: </p><blockquote><p>First, the <strong>virtual architecture</strong> is not in <code>e_flags</code> anymore, despite the name in the field. Compile the same kernel three ways, with the <em>PTX architecture</em> set to <code>compute_75</code>, <code>compute_80</code> and <code>compute_90</code> but the real target fixed at <code>sm_90</code>, and all three cubins carry <code>e_flags = 0x06005a04</code>. </p><p>The virtual architecture actually moved into a <strong>note section,</strong> where <code>cuobjdump</code> reports it as <code>CUDA Virtual SM: sm_75</code> and so on. This used to be visible: CUDA 8-era disassembly printed a line reading <code>.headerflags @"EF_CUDA_SM20 EF_CUDA_PTX_SM(EF_CUDA_SM20)"</code>, and current <code>nvdisasm</code> prints a plain <code>.target sm_90</code> instead. If you have old tooling that reads the <strong>PTX architecture</strong> out of the flags word, you have seen zero for some years.</p></blockquote><blockquote><p>Second, the <strong>architecture-conditional</strong> suffixes do not show up here either. <code>sm_90</code> and <code>sm_90a</code> produce identical <code>e_flags</code>, and so do all three of <code>sm_100</code>, <code>sm_100a</code> and <code>sm_100f</code>. For a kernel that uses no special features the <code>sm_100</code> and <code>sm_100f</code> cubins are the same program <strong>byte for byte</strong>, and the 210 bytes that do differ between the two files are all downstream of one string: the note section records <code>-arch sm_100f</code> instead of <code>-arch sm_100</code>, which is one byte longer and shifts everything after it. </p><p>That is in line with <strong>what NVIDIA says</strong> about the feature, which is that <code>code=sm_100</code> and <code>code=sm_100f</code> are aliases producing the same cubin when no family-specific features are in play. The suffix is recorded, but the recording is somewhere else and is &#8220;triggered&#8221; only when the kernel actually uses something.</p></blockquote><blockquote><p>Third, CUDA 13&#8217;s supported target list is very short, and that surprised even me. Turing is now the floor. <code>ptxas</code> in 13.3 accepts exactly <code>sm_75, sm_80, sm_86, sm_87, sm_88, sm_89, sm_90, sm_90a</code>, then the three-digit generation: <code>sm_100, sm_103, sm_110, sm_120, sm_121</code> each with <code>a</code> and <code>f</code> variants, plus the matching <code>lto_*</code> targets. </p><p>Maxwell, Pascal and Volta were removed in CUDA 13.0, which NVIDIA&#8217;s release notes state directly: <strong>offline compilation</strong> and library support for those architectures are gone completely, and 12.x is the last line that can target them.</p></blockquote><p>You have to know at least two numbers if you ever have to explain a build matrix. <code>sm_110</code> is Jetson AGX Thor, renamed from <code>sm_101</code> when CUDA 13.0 shipped, which means build scripts that hardcoded <code>101</code> broke on a rename rather than on a chip. NVIDIA states the rename in the<strong> nvCOMPDx release notes</strong> rather than anywhere prominent. </p><p>And <code>sm_88</code>, compute capability 8.8, is said by the community compatibility references to be the Nintendo Switch 2, which I cannot verify from the toolchain itself: <code>ptxas</code> will happily build for it and tells you nothing about what it is. </p><p>If that claim is right, the CUDA binary format has a target for a games console sitting in the same list as your H100.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share The Software Frontier&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share The Software Frontier</span></a></p><div><hr></div><h2>Fifteen sections, but only one is code</h2><p>Below I will show you the full section table of a cubin containing one trivial kernel, a <strong>SAXPY with four parameters</strong>, compiled for Hopper.</p><pre><code>idx name                                   type        off   size  align
  1 .shstrtab                              STRTAB       64    295      1
  2 .strtab                                STRTAB      406    369      1
  3 .symtab                                SYMTAB      776    240      8
  4 .debug_frame                           PROGBITS   1016    104      1
  5 .note.nv.tkinfo                        NOTE       1120    164      4
  6 .note.nv.cuinfo                        NOTE       1284     32      4
  7 .nv.info                               CUDA_INFO  1316     36      4
  8 .nv.compat                             CUDA_COMPAT_INFO
                                                      1352     28      4
  9 .nv.info._Z5saxpyifPKfPf               CUDA_INFO  1380    132      4
 10 .nv.callgraph                          CUDA_CALLGRAPH
                                                      1512     32      4
 11 .rela.debug_frame                      RELA       1544     24      8
 12 <strong>.text._Z5saxpyifPKfPf</strong>                   PROGBITS   1664    512    128
 13 .nv.shared.reserved.0                   NOBITS     2176      0      1
 14 .nv.constant0._Z5saxpyifPKfPf           PROGBITS   2176    552      4</code></pre><p><em>We have 512 bytes of machine code inside a 3,968 byte file. The italicised sections are the CUDA-specific ones, with section types in the 0x70000000 range that generic ELF tools print as &#8220;unknown&#8221;.</em></p><p>The <strong>code section</strong> is said to be &#8220;<em>per kernel</em>&#8221; and named after the mangled symbol, which is why a file with ten kernels has ten <code>.text.*</code> sections and ten of most other things too. Alignment is 128 bytes, one instruction cache line&#8217;s worth of paranoia.</p><p>Then, in a random order of how much they surprised me:</p><h3>.note.nv.tkinfo records how the binary was built</h3><p>This is a standard ELF note with owner string <code>NVIDIA Corp</code>, and <code>cuobjdump</code> will print it for you.</p><pre><code>$ cuobjdump -elf saxpy.sm90.cubin | grep -A5 &#8220;CUDA Toolkit Information&#8221;
  NVIDIA Corp   140   NVIDIA CUDA Toolkit Information
    Note Version: 2
    Tool Name: <strong>ptxas</strong>
    Tool Version: Cuda compilation tools, release 13.3, V13.3.73
    Tool Branch: Build cuda_13.3.r13.3/compiler.38244171_0
    Tool Command Line Arguments: <strong>-arch sm_90 -m 64</strong></code></pre><p>Every cubin has the <strong>exact compiler build</strong> that produced it and the flags it was invoked with. That's a provenance record, and it&#8217;s sitting in every shipped binary on every machine learning platform you have ever installed and used. </p><p>If you want to know which toolkit built the kernels inside somebody&#8217;s wheel, no need to ask them. There is also a <code>-verbose-tkinfo</code> mode in the assembler that emits, in its own words, &#8220;<em>object name and command line arguments which contains all arguments having file format,</em>&#8221; meaning paths. I havent seen it on by default, and I&#8217;d look before shipping a binary built with it.</p><h3>.nv.constant0 is the launch frame, and it keeps growing</h3><p>Kernel parameters don&#8217;t come in registers, but in<strong> constant bank 0</strong>, and the cubin declares a per-kernel section for it. On Hopper we said this section is 552 bytes for a kernel whose parameters occupy 24. The user arguments start at offset <code>0x210</code>; everything below that is owned by the driver and the ABI.</p><p>You can watch the exact boundary move across generations. It is at <code>0x160</code> for Turing through Ada, <code>0x210</code> for Hopper, and <code>0x380</code> for Blackwell and later, with the section growing from 380 to 556 to 924 bytes for the same kernel. </p><p>This is the sort of constant that <strong>reverse-engineering</strong> projects hardcode and then rediscover every two years.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!sEJJ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb71d2985-8ddd-4a60-a6a0-989bc0ca09c0_1503x878.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!sEJJ!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb71d2985-8ddd-4a60-a6a0-989bc0ca09c0_1503x878.png 424w, https://substackcdn.com/image/fetch/$s_!sEJJ!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb71d2985-8ddd-4a60-a6a0-989bc0ca09c0_1503x878.png 848w, https://substackcdn.com/image/fetch/$s_!sEJJ!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb71d2985-8ddd-4a60-a6a0-989bc0ca09c0_1503x878.png 1272w, https://substackcdn.com/image/fetch/$s_!sEJJ!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb71d2985-8ddd-4a60-a6a0-989bc0ca09c0_1503x878.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!sEJJ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb71d2985-8ddd-4a60-a6a0-989bc0ca09c0_1503x878.png" width="1456" height="851" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b71d2985-8ddd-4a60-a6a0-989bc0ca09c0_1503x878.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:851,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Two bar charts: kernel parameter base offset in constant bank 0 and total .nv.constant0 size, both by architecture, showing steps at Hopper and Blackwell.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Two bar charts: kernel parameter base offset in constant bank 0 and total .nv.constant0 size, both by architecture, showing steps at Hopper and Blackwell." title="Two bar charts: kernel parameter base offset in constant bank 0 and total .nv.constant0 size, both by architecture, showing steps at Hopper and Blackwell." srcset="https://substackcdn.com/image/fetch/$s_!sEJJ!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb71d2985-8ddd-4a60-a6a0-989bc0ca09c0_1503x878.png 424w, https://substackcdn.com/image/fetch/$s_!sEJJ!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb71d2985-8ddd-4a60-a6a0-989bc0ca09c0_1503x878.png 848w, https://substackcdn.com/image/fetch/$s_!sEJJ!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb71d2985-8ddd-4a60-a6a0-989bc0ca09c0_1503x878.png 1272w, https://substackcdn.com/image/fetch/$s_!sEJJ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb71d2985-8ddd-4a60-a6a0-989bc0ca09c0_1503x878.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">The driver&#8217;s half of the launch. The implicit state a kernel launch carries has grown 2.4x since Turing while the user-visible argument list stayed the same size. Clusters, tensor memory, distributed shared memory and grid constants all have to be described somewhere, and this is where.<span>Measured. Base offset read from the EIATTR_PARAM_CBANK attribute; 24 bytes of user arguments in every case. The space below the base belongs to the driver and the ABI.</span></figcaption></figure></div><p>You can see the effect in the disassembly, which is the easiest way I know to make constant bank 0 sort of&#8230;. real. Here is the same kernel on Hopper and on Thor.</p><pre><code>sm_90                                    sm_110
LDC   R1, c[0x0][0x28]                   LDC   R1, c[0x0][0x37c]
S2R   R0, SR_TID.X                       S2R   R0, SR_TID.X
S2UR  UR4, SR_CTAID.X                    S2UR  UR4, SR_CTAID.X
LDC   R7, c[0x0][RZ]                     <strong>LDCU</strong>  UR5, c[0x0][0x380]
IMAD  R7, R7, UR4, R0                    LDC   R7, c[0x0][0x360]
ULDC  UR4, c[0x0][0x210]                 IMAD  R7, R7, UR4, R0
ISETP.GE.AND P0, PT, R7, UR4, PT         ISETP.GE.AND P0, PT, R7, UR5, PT
@P0   EXIT                               @P0   EXIT</code></pre><p><em>Left: the first parameter is loaded from 0x210. Right: from 0x380, using <strong>LDCU</strong>, an instruction that loads a constant straight into a uniform register and that does not exist in the Hopper instruction set reference. Even the stack pointer moved, from 0x28 to 0x37c.</em></p><h3>The symbol table has its own vocabulary</h3><p>Kernels are ordinary <strong>global function symbols </strong>with two CUDA-specific decorations, but if you look carefully, ther's one symbol in a cubin that has nothing to do with your code.</p><pre><code>index  value   size   info  other  shndx  name
  0x3      0      0    0x3      0    0xc  .text._Z5saxpyifPKfPf
  0x4      0    0x4   0x21      0      0  .nv.reservedSmem.offset0
  0x5      0      0   0x20   0xa0    0xd  __nv_reservedSMEM_offset_0_alias
  0x8      0  0x200   0x12   0x10    0xc  <strong>_Z5saxpyifPKfPf</strong>
  0x9      0      0    0x3      0    0xe  .nv.constant0._Z5saxpyifPKfPf</code></pre><p>The kernel symbol carries <code>st_other = 0x10</code>, which <code>nvdisasm</code> prints as <code>STO_CUDA_ENTRY STV_DEFAULT</code>: a visibility byte reused to mark launchable entry points, which is how the driver tells a kernel from a device function without parsing names. </p><p>The reserved shared memory symbols are the interesting thing here. Even a kernel that declares &#8220;no shared memory&#8221; unexpectedly gets a <strong>reservation record</strong>, and on Blackwell that reservation has a <strong>non-zero size:</strong> somehw you see 64 bytes of shared memory that belong to the platform, not to you, aliased under a name beginning with <code>__nv_</code>. Anyone computing occupancy from source is computing it from the wrong number.</p><h3>The same source is not the same program</h3><p>I decided to write a whole paragraph to the architecture sweep, because it shows how much the<strong> cubin is influenced</strong> by the <strong>specific chip</strong>, rather than by your code. </p><p>The four-line SAXPY compiles to 24 instruction slots on <code>sm_86</code> and 48 on <code>sm_87</code>, Orin, and the difference is not optimization.</p><pre><code>sm_86 opcode histogram:   8 NOP   2 S2R   2 MOV   2 LDG.E   2 IMAD.WIDE  ...
sm_87 opcode histogram:  15 <strong>BMOV.32.CLEAR</strong>  12 NOP   2 S2R   2 LDG.E  ...</code></pre><p>We see<strong> fifteen instructions</strong> clearing barrier state at kernel entry, on one embedded part, for a kernel that uses no barriers. </p><p>In other words, that piece of code inside the kernel isn&#8217;t there because your program desperately needs it. It&#8217;s here because NVIDIA is paying, through software, the price of a <strong>silicon limitation</strong>. And that proves that SASS  is dependent not only on the PTX, but also on the hardware, the drivers and the workarounds embedded into the toolchain. </p><h3>.nv.callgraph exists because device code can call things</h3><p>Fixed 8-byte entries, and for a leaf kernel it is four rows of negative sentinels:</p><pre><code><code>.nv.callgraph
&lt;0,-1&gt;   &lt;0,-2&gt;   &lt;0,-3&gt;   &lt;0,-4&gt;</code></code></pre><p>The sentinels are the interesting part: a leaf kernel does not get an empty section, but four explicit &#8220;<em>no callee of kind N</em>&#8221; markers. This suggests that <code>.nv.callgraph</code> is not simply a list of callees, but a structured description of call relationships understood by the runtime.</p><p>That metadata matters because features such as indirect calls, dynamic parallelism, and <strong>separately compiled device functions</strong> require the driver to know the worst-case call depth before launch so it can size the per-thread stack correctly. </p><p> Even a kernel that calls nothing still participates in this contract, which is why absence is encoded explicitly rather than by omitting the section entirely.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>.nv.info is where the kernel describes itself</h2><p>This is the section that let you see a loadable cubin, rather than merely readable. The format is a simple <strong>tag-length-value stream</strong>: one byte of format, one byte of attribute, then either nothing, one byte, two bytes, or a 16-bit length followed by that many bytes. </p><p>And&#8230;.thirty-six bytes of it, at file level, look like this:</p><pre><code>04 <strong>2f</strong> 08 00 | 08 00 00 00  0a 00 00 00      EIATTR_REGCOUNT
04 <strong>11</strong> 08 00 | 08 00 00 00  00 00 00 00      EIATTR_FRAME_SIZE
04 <strong>12</strong> 08 00 | 08 00 00 00  00 00 00 00      EIATTR_MIN_STACK_SIZE
 ^  ^  ^
 |  |  +-- value length (EIFMT_SVAL)
 |  +----- attribute code
 +-------- format: 1 = none, 2 = one byte, 3 = two bytes, 4 = length-prefixed</code></pre><p>The first value that you read is a symbol table index, the second is the payload. The symbol 8 in each column is the kernel, and it needs 10 registers, no stack frame, no minimum stack. You can check that decode against the tool, because <code>cuobjdump -elf</code> knows the names.</p><pre><code>$ cuobjdump -elf saxpy.sm90.cubin
.nv.info
  Attribute: <strong>EIATTR_REGCOUNT</strong>       Value: function: _Z5saxpyifPKfPf(0x8)  register count: 10
  Attribute: <strong>EIATTR_FRAME_SIZE</strong>     Value: function: _Z5saxpyifPKfPf(0x8)  frame size: 0x0
  Attribute: <strong>EIATTR_MIN_STACK_SIZE</strong> Value: function: _Z5saxpyifPKfPf(0x8)  min stack size: 0x0</code></pre><p>The per-kernel section is way richer. Let&#8217;s see a four-parameter SAXPY:</p><pre><code>.nv.info._Z5saxpyifPKfPf
  EIATTR_LANGUAGE            PTX
  EIATTR_CUDA_API_VERSION    0x85                      (133 = CUDA 13.3)
  EIATTR_KPARAM_INFO         Ordinal 0x3  Offset 0x10  Size 0x8  Space CBANK
  EIATTR_KPARAM_INFO         Ordinal 0x2  Offset 0x8   Size 0x8  Space CBANK
  EIATTR_KPARAM_INFO         Ordinal 0x1  Offset 0x4   Size 0x4  Space CBANK
  EIATTR_KPARAM_INFO         Ordinal 0x0  Offset 0x0   Size 0x4  Space CBANK
  EIATTR_SPARSE_MMA_MASK     0x0
  EIATTR_MAXREG_COUNT        0xff
  <strong>EIATTR_MERCURY_ISA_VERSION 1.1</strong>
  EIATTR_EXIT_INSTR_OFFSETS  0x70  0x120
  EIATTR_CBANK_PARAM_SIZE    0x18                      (24 bytes of parameters)
  EIATTR_PARAM_CBANK         0x9   0x180210            (symbol 9, size 0x18 at 0x210)
  EIATTR_SW_WAR              0x8
  EIATTR_NVSAL_SW_WAR        0x1</code></pre><p><strong>Kernel parameters</strong> are described entirely through size-and-offset metadata, with entries stored in something called &#8220;<em>reverse ordinal order</em>&#8221;. </p><p>This means the host runtime doesn&#8217;t have to understand the original programming language or the function signature. It can construct the kernel&#8217;s parameter buffer directly from the metadata, placing each argument at the correct location before launch. The compiled binary therefore exposes a<strong> language-independent ABI</strong> that the driver and runtime can consume.</p><p><code>EIATTR_EXIT_INSTR_OFFSETS</code> contains the byte offset of every <code>EXIT</code> instruction generated for the kernel. At first glance this is considered to be like a minor detail, but it&#8217;s very useful for tooling. </p><p>A profiler, tracer, or instrumentation framework can find every<strong> kernel return point</strong> immediately, without performing a full disassembly or reconstructing the control-flow graph. If a tool have to inject timing code, coverage probes, or custom telemetry before kernel termination, these offsets provide the exact insertion points.</p><p><code>EIATTR_SW_WAR</code> is a different kind of purpose here. It encodes software workarounds for known hardware errata as a bitmask. Modern GPUs sometimes require<strong> compiler- or driver-level mitigations</strong> to avoid wrong behavior in specific corner cases. </p><p>Rather than hard-coding these decisions somewhere else, the compiler do a record of the workarounds that are needed, in the binary metadata. The driver then inspects the flags at load time and enables the correct mitigation paths for the architecture in place.</p><p>All these attributes show you that <strong>CUDA metadata</strong> is not  just a bunch of descriptions; it forms actively some parts of the contract between the compiler, runtime, driver, profiling tools, and the GPU itself, carrying the information needed to launch kernels, instrument execution, and safely navigate hardware-specific quirks.</p><p>Now a question arises naturally: <em>how big is this vocabulary?</em> The disassembler contains the enum names, so you can just ask it, and you&#8217;ll find out.</p><pre><code>$ strings -a nvdisasm | grep -o &#8220;EIATTR_[A-Z0-9_]*&#8221; | sort -u | wc -l
<strong>113</strong></code></pre><p>Now you see it: 113 attribute kinds. A few of them are a decent map of what the hardware has learned to do since 2010. </p><ul><li><p><code>EIATTR_TCGEN05_1CTA_USED</code> and <code>EIATTR_TCGEN05_2CTA_USED</code> flag use of Blackwell&#8217;s fifth-generation tensor cores, and the fact that there are separate flags for the one-CTA and two-CTA forms tells you the driver has to care which. </p></li><li><p><code>EIATTR_CTA_PER_CLUSTER</code>, <code>EIATTR_MAX_CLUSTER_RANK</code>, <code>EIATTR_EXPLICIT_CLUSTER</code> and <code>EIATTR_BLOCKS_ARE_CLUSTERS</code> are the Hopper cluster model. </p></li><li><p><code>EIATTR_STACK_CANARY_TRAP_OFFSETS</code> means device code has stack canaries now, and indeed <code>nvcc</code> has a flag for them. </p></li><li><p><code>EIATTR_COROUTINE_RESUME_ID_OFFSETS</code> is there for something that has not been announced. </p></li><li><p>And a small family of attributes are named after bug numbers: <code>EIATTR_SW1850030_WAR</code>, <code>EIATTR_SW2393858_WAR</code>, <code>EIATTR_SW2861232_WAR</code>, <code>EIATTR_WAR5829587_NEEDED</code>. </p></li></ul><p>There is  also an internal ticket for each somewhere, and the workaround is load-bearing enough to have its own slot in the binary format.</p><blockquote><p><em>The format is not a description of a program. It is a contract between a compiler and a driver that ship on different schedules and are written by the same company.</em></p></blockquote><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Twenty-one bits per instruction that the hardware does not check</h2><p>Starting with Volta, SASS instructions are 128 bits, and a piece of every instruction word is not the instruction at all, because it&#8217;s the <strong>scheduling metadata</strong> that the assembler computes and the hardware obeys without verifying. </p><p>You will find plenty of documentation about the structure of the field in the <strong>microbenchmarking literature</strong>, rather than by the vendor: a paper by Jia and colleagues describes stall counts, a yield flag, separate read and write dependency barriers, a wait mask that is a bitmask because an instruction can wait on several barriers at once, and reuse flags. </p><p>The exact bit positions below are the ones I decoded, and I trust them because the result is coherent with the docs: stall count in bits 105 to 108, yield flag at 109, write barrier index at 110 to 112, read barrier index at 113 to 115, a six-bit wait mask at 116 to 121, and four reuse flags at 122 to 125. Bits 126 and 127 come out zero on every instruction of every kernel I looked at, on five architectures.</p><p>You can decode it yourself with twenty lines of Python, and I think everyone who works on inference should do it at least once, because the mechanism explains more about the CUDA moat than any benchmark. Here is a real <strong>Hopper GEMM prologue</strong> with the fields pulled out.</p><pre><code>off   stall  y  wr  rd    wait  reuse   instruction
0000      1  1   -   -  000000   0000   LDC R1, c[0x0][0x28]
0010      1  1   <strong>0</strong>   -  000000   0000   S2R R9, SR_CTAID.Y
0020      1  1   -   -  000000   0000   ULDC UR4, c[0x0][0x228]
0030      1  1   -   -  000000   0000   ULDC.64 UR6, c[0x0][0x208]
0040      1  1   -   -  000000   0000   MOV R0, UR4
0050      1  1   <strong>0</strong>   -  000000   0000   S2R R13, SR_TID.Y
0060      2  1   -   -  000000   0000   HFMA2.MMA R16, -RZ, RZ, 0, 0
0070      1  1   -   -  000000   0000   ISETP.GE.AND P0, PT, R0, 0x1, PT
0080      1  1   <strong>1</strong>   -  000000   0000   S2R R17, SR_CTAID.X
0090      1  1   <strong>1</strong>   -  000000   0000   S2R R7, SR_TID.X
00a0      2  0   -   -  <strong>000001</strong>   0000   LEA R9, R9, R13, 0x5
00b0      8  0   -   -  <strong>000010</strong>   0000   LEA R11, R17, R7, 0x5</code></pre><p>Two special register reads allocate barrier 0, and the<strong> LEA</strong> that consumes them waits on barrier 0. Two more allocate barrier 1, and the next LEA waits on barrier 1. The dependency is not discovered at run time, but written into the binary.</p><p>That is the whole argument, in twelve instructions. For <strong>variable-latency instructions</strong> the assembler allocates one of six barriers and makes consumers wait on a bitmask. </p><p>For <strong>fixed-latency instructions</strong> it does not use barriers at all, it just writes a number of cycles into the stall field and the scheduler holds off that long. There is no interlock checking this: if the number is wrong, you do not get a stall, you get wrong answers. </p><p>Across a <strong>full tiled GEMM kernel</strong> the distribution looks like this.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!wUyx!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa8e2d257-93b7-427a-9a55-163592501b79_1522x854.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!wUyx!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa8e2d257-93b7-427a-9a55-163592501b79_1522x854.png 424w, https://substackcdn.com/image/fetch/$s_!wUyx!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa8e2d257-93b7-427a-9a55-163592501b79_1522x854.png 848w, https://substackcdn.com/image/fetch/$s_!wUyx!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa8e2d257-93b7-427a-9a55-163592501b79_1522x854.png 1272w, https://substackcdn.com/image/fetch/$s_!wUyx!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa8e2d257-93b7-427a-9a55-163592501b79_1522x854.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!wUyx!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa8e2d257-93b7-427a-9a55-163592501b79_1522x854.png" width="1456" height="817" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a8e2d257-93b7-427a-9a55-163592501b79_1522x854.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:817,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Histogram of encoded stall counts and a bar chart of how often each control field is used across one kernel.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Histogram of encoded stall counts and a bar chart of how often each control field is used across one kernel." title="Histogram of encoded stall counts and a bar chart of how often each control field is used across one kernel." srcset="https://substackcdn.com/image/fetch/$s_!wUyx!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa8e2d257-93b7-427a-9a55-163592501b79_1522x854.png 424w, https://substackcdn.com/image/fetch/$s_!wUyx!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa8e2d257-93b7-427a-9a55-163592501b79_1522x854.png 848w, https://substackcdn.com/image/fetch/$s_!wUyx!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa8e2d257-93b7-427a-9a55-163592501b79_1522x854.png 1272w, https://substackcdn.com/image/fetch/$s_!wUyx!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa8e2d257-93b7-427a-9a55-163592501b79_1522x854.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">One 32x32 tiled float GEMM for sm_90, 168 instruction slots. Most instructions issue back to back with a one cycle stall. Write barriers appear on 26% of slots, read barriers on none, and the yield flag on 71%. The same decode works unchanged on sm_100, sm_110 and sm_120, which is a small piece of evidence that Blackwell did not change the scheduling contract. <span>Measured: one 32x32 float tiled GEMM compiled for sm_90, 168 instruction slots including padding. Counts are slots, decoded by direct parse of the .text section.</span></figcaption></figure></div><h3>The linker can rewrite the control bits</h3><p>The relocation namespace makes the point better than I can. Pull the type names out of <code>nvlink</code> and there are 119 of them, and most encode a bit position in the name.</p><pre><code>$ strings -a nvlink | grep -o &#8220;R_CUDA_[A-Z0-9_]*&#8221; | sort -u | wc -l
<strong>119</strong>

R_CUDA_ABS32_23        patch 32 bits at bit offset 23
R_CUDA_ABS32_HI_32     high half of a 64-bit address, at bit 32
R_CUDA_CONST_FIELD19_20  19-bit constant bank field at bit 20
R_CUDA_ABS55_16_34     55-bit value split across bits 16 and 34
R_CUDA_INSTRUCTION128  a whole 128-bit instruction word
R_CUDA_PCREL_IMM24_23  24-bit program counter relative branch at bit 23
<strong>R_CUDA_YIELD_CLEAR_PRED4_87</strong>   clear 4 bits at bit 87
<strong>R_CUDA_YIELD_OPCODE9_0</strong>        rewrite a 9-bit opcode field at bit 0</code></pre><p>On a CPU, relocations patch whole bytes because operands are byte-aligned. In this case, the operands are bit fields inside a <strong>128-bit word</strong>, so the relocation type has to name the field, and there is a separate type for the same width at each position it can occur. </p><p>The last two are the ones that really stopped me. A relocation that clears the yield predicate, and one that <strong>rewrites an opcode</strong>, mean the device linker&#8217;s contract includes editing scheduling and instruction selection after the assembler has finished. </p><p>The<strong> 21 control bits</strong> are not a compiler-internal detail, which is nested at the end of <code>ptxas</code>, because they are a key part of the object format, and a later stage is allowed to change them.</p><p>I have argued many times before that this field, not the language and not PTX, is the <strong>load-bearing part</strong> of the CUDA moat, and taking cubins apart has not changed my mind. It has sharpened one point though. </p><p>The reason the assembler is fully close is not that NVIDIA wants to hide optimizations and reverse-engineering, at least not the first point. What we also need to keep in mind is that the <code>ptxas</code> is<strong> correctness-critical</strong> infrastructure. </p><p>An external tool that emits SASS is not competing with a compiler on quality of code, it is taking over responsibility for a hazard-avoidance protocol with no runtime check and <strong>no error reporting</strong>, which is extremely dangerous. </p><p>NVIDIA&#8217;s own programming guide is blunt about the consequence: you read that binary compatibility is promised only for binaries created by <strong>NVIDIA tools:</strong> manual editing or generating binary code is not supported, and compatibility promises are invalidated if binaries are modified in any way. </p><p>Read that as a description of the <strong>support boundary</strong> rather than a threat, and the closure looks less like strategy and more like the only sentence a vendor can write once it has moved the interlock into software.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share The Software Frontier&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share The Software Frontier</span></a></p><div><hr></div><h2>Blackwell cubins carry a second copy of every kernel</h2><p>When I thought about writing this part, I didn&#8217;t even know what to expect. It was a complete surprise to me. Let&#8217;s do a concrete example: compile the same <strong>trivial kernel</strong> for Hopper and for Blackwell and count sections: 15 for <code>sm_90</code>, 22 for <code>sm_100</code>. The seven extra sections are not debug info.</p><pre><code>idx name                                   type              flags     size
 12 .text._Z5saxpyifPKfPf                  PROGBITS              6      512
 13 .nv.shared.reserved.0                  NOBITS                3       64
 14 .nv.constant0._Z5saxpyifPKfPf          PROGBITS             42      920
 15 <strong>.nv.capmerc.text._Z5saxpyifPKfPf</strong>       0x70000016     10000000      258
 16 <strong>.nv.merc.debug_frame</strong>                   PROGBITS       10000000      112
 17 <strong>.nv.merc.nv.info</strong>                       0x70000083     10000000       36
 18 <strong>.nv.merc.nv.info._Z5saxpyifPKfPf</strong>       0x70000083     10000040      144
 19 <strong>.nv.merc.rela.debug_frame</strong>              0x70000082     10000040       24
 20 <strong>.nv.merc.nv.shared.reserved.0</strong>          0x70000015     10000003        0
 21 <strong>.nv.merc.symtab</strong>                        0x70000085     10000000      216</code></pre><p>You are now staring at a <strong>complete parallel object</strong>: its own symbol table, its own info sections, its own relocations, its own shared memory reservation, and a &#8220;<em>capmerc</em>&#8221; text section. Section flag bit 0x10000000 marks the whole family.</p><p>The prefix is <code>merc</code>, and the name shows up in three other places. It is in the per-kernel info as <code>EIATTR_MERCURY_ISA_VERSION 1.1</code>. It is in a new section called <code>.nv.compat</code>, which exists on every target but says more on Blackwell, and it is all over the assembler binary.</p><pre><code>$ cuobjdump -elf saxpy.sm100a.cubin | sed -n &#8216;/nv.compat/,/nv.info\._/p&#8217;
.nv.compat
  <strong>EICOMPAT_ATTR_CUDA_ACCELERATOR_TARGET</strong>                 0x1
  EICOMPAT_ATTR_ISA_CLASS                                0x1
  <strong>EICOMPAT_ATTR_INST_TCGEN05_MMA</strong>                        0x5
  EICOMPAT_ATTR_MERCURY_ISA_MAJOR_MINOR_VERSION_V2       1.1
  EICOMPAT_ATTR_MERCURY_ISA_MAJOR_MINOR_VERSION_V1       1.1
  EICOMPAT_ATTR_INST_TENSORMAP_V1                        0x0
  <strong>EICOMPAT_ATTR_CAN_FASTPATH_FINALIZE</strong>                   0x9 0x0</code></pre><p>There is the missing suffix. <code>EICOMPAT_ATTR_CUDA_ACCELERATOR_TARGET</code> is 1 for <code>sm_100a</code> and 0 for both <code>sm_100</code> and <code>sm_100f</code>, and the accelerated build declares <code>EICOMPAT_ATTR_INST_TCGEN05_MMA</code>, a capability it was permitted to use even though this kernel does not use it.</p><p>To really find out what the block really tracks, you need to build a kernel that needs a family-specific instruction. One line of <strong>inline PTX</strong> will do: <code>tcgen05.fence::before_thread_sync</code>, part of Blackwell&#8217;s fifth-generation tensor core interface. </p><p>Compiling it for <strong>five targets </strong>gives the feature tiers as a straight experiment rather than as a diagram.</p><pre><code>$ for a in sm_100 sm_100f sm_100a sm_103f sm_120a; do nvcc -arch=$a -cubin -o t.cubin tc.cu; done

sm_100    <strong>Instruction &#8216;tcgen05.fence&#8217; not supported on .target &#8216;sm_100&#8217;</strong>
sm_100f   OK  (5136 bytes)
sm_100a   OK  (5136 bytes)
sm_103f   OK  (5136 bytes)
sm_120a   <strong>Instruction &#8216;tcgen05.fence&#8217; not supported on .target &#8216;sm_120a&#8217;</strong></code></pre><p>The baseline target refuses the instruction, but the<strong> family target</strong> accepts it, and so does the family target for the other member of the same family. </p><p>And the fully accelerated target for consumer Blackwell refuses it, because <code>sm_120</code> is a different family and doesn&#8217;t have the hardware, which is the concrete answer to the frequently asked question of why a 5090 is not a small B200.</p><p>Now look at what the successful builds wrote into the object.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!nssc!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fffb50e64-a1b5-490f-926d-d21b1414ccb0_1939x1111.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!nssc!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fffb50e64-a1b5-490f-926d-d21b1414ccb0_1939x1111.png 424w, https://substackcdn.com/image/fetch/$s_!nssc!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fffb50e64-a1b5-490f-926d-d21b1414ccb0_1939x1111.png 848w, https://substackcdn.com/image/fetch/$s_!nssc!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fffb50e64-a1b5-490f-926d-d21b1414ccb0_1939x1111.png 1272w, https://substackcdn.com/image/fetch/$s_!nssc!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fffb50e64-a1b5-490f-926d-d21b1414ccb0_1939x1111.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!nssc!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fffb50e64-a1b5-490f-926d-d21b1414ccb0_1939x1111.png" width="1456" height="834" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ffb50e64-a1b5-490f-926d-d21b1414ccb0_1939x1111.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:834,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:158227,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/207127260?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fffb50e64-a1b5-490f-926d-d21b1414ccb0_1939x1111.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!nssc!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fffb50e64-a1b5-490f-926d-d21b1414ccb0_1939x1111.png 424w, https://substackcdn.com/image/fetch/$s_!nssc!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fffb50e64-a1b5-490f-926d-d21b1414ccb0_1939x1111.png 848w, https://substackcdn.com/image/fetch/$s_!nssc!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fffb50e64-a1b5-490f-926d-d21b1414ccb0_1939x1111.png 1272w, https://substackcdn.com/image/fetch/$s_!nssc!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fffb50e64-a1b5-490f-926d-d21b1414ccb0_1939x1111.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>That is the piece I was missing when I first wrote this section, infact I had to rebuild it. The compat block is not a record of your compiler flags; it is a <strong>graded statement</strong> about what the machine code inside requires, computed from the code, and the flags only set the ceiling on what the assembler was allowed to reach for. </p><p>A resolver reading <code>ISA_CLASS = 1</code> knows this object is safe anywhere in the generation; reading 2, it knows to check the family. This is the <strong>minimum viable metadata</strong> for the compatibility promise NVIDIA made in CUDA 12.9, and it is sitting in a 32-byte section nobody documents.</p><p>What reads it becomes obvious once you look inside <code>ptxas</code>. The strings are unambiguous.</p><pre><code>$ strings -a ptxas | grep -iE &#8220;merc|finaliz&#8221; | sort -u
[Finalizer] fastpath optimization applied for <strong>off-target %u -&gt; %u finalization</strong>
Generate Capsule Mercury
Specify the type of target ELF binary kind. Default on sm100+ is <strong>capmerc</strong>
Self check for capsule mercury (capmerc)
Specify the opportunistic finalization level. 0=default, 1=no opportunistic
  finalization, <strong>2=intra family finalization only, or 3=intra and inter family
  finalization</strong>
Turns off the fast-path finalization optimization (allows normal refinaization)
Specify the &#8216;sm_&#8217; name of the target architecture. If not specified, default
  behavior is <strong>on-target finalization</strong>
R_MERCURY_NONE  R_MERCURY_G64  R_MERCURY_ABS64  R_MERCURY_ABS32  R_MERCURY_ABS16 ...</code></pre><p>Read together with the sections, this describes a two-stage back end. Mercury is an encoding of a kernel that is <strong>below PTX, </strong>but above final machine code. </p><p>A capsule wraps it with the metadata a finalizer needs: its own symbol table, register and barrier counts, shared memory usage, and its own relocation types. &#8220;<em>Finalization</em>&#8221; is the step that turns a capsule into <strong>executable SASS</strong> for a specific chip. On-target finalization is the ordinary case. </p><p><strong>Off-target finalization</strong> is the same capsule being finalized for a different chip than the one it was compiled for, with a fast path when the assembler can prove the existing encoding is already valid. And the levels of &#8220;<em>opportunistic finalization</em>&#8221; go from none, to within a family, to <em>across</em> families.</p><p>That is what the <code>f</code> targets are made of. NVIDIA introduced family-conditional compilation in CUDA 12.9 and actually described it in terms of feature sets: an <code>f</code> binary may use the <strong>architecture-specific features</strong> that are common to a whole family, and will run on later members of that family. </p><p>The public framing is a compatibility promise. The mechanism, as far as I can see it in the binary, is that the cubin ships a <strong>re-finalizable representation </strong>next to the SASS, and the driver can produce correct machine code for a family member the compiler never saw.</p><h4>Where I am reading, not knowing</h4><p>Unfortunately, I havent observed finalization happen. At the moment, don&#8217;t own a Blackwell GPU, and none of this machinery has the <code>ptxas</code> command line: passing <code>--binary-kind</code> gets you <code>Unknown option</code>, so the option table I am quoting belongs to an internal or driver-side entry point. </p><p>What I am confident about is the <strong>presence and shape of the artifacts</strong>. An independent reverse-engineering effort on <code>ptxas</code> 13.0 reached the same reading and adds detail I could not verify, including a 328-byte capsule descriptor and a compilation-knob snapshot inside it. Treat the purpose as tier C, so a speculative claim.</p><p>There is more of it in the section type table than any single cubin shows. <code>cuobjdump</code> has to be able to print a name for every section type it might meet, so the names are in the binary, and the Mercury family is large.</p><pre><code>$ strings -a cuobjdump | grep -oE &#8220;CUDA_[A-Z_]{3,}&#8221; | sort -u | grep -i merc
CUDA_CAPMERC
CUDA_MERCURY
CUDA_MERCURY_CONSTANT_DRIVER      CUDA_MERCURY_CONSTANT_PARAMS
CUDA_MERCURY_CONSTANT_IMGHDR      CUDA_MERCURY_CONSTANT_PIC
CUDA_MERCURY_CONSTANT_OPT         CUDA_MERCURY_CONSTANT_TOOLS
                                  CUDA_MERCURY_CONSTANT_USER
CUDA_MERCURY_RESOLVED_RELA
<strong>CUDA_MERCURY_SASS_MAP</strong></code></pre><p><strong>Seven </strong>distinct <strong>constant bank</strong> <strong>section types</strong>, split by who owns the data: driver, image header, optimizer, parameters, position independent code, tools, user. A resolved-relocation type. And a section type called <code>CUDA_MERCURY_SASS_MAP</code>, which is hard to read as anything other than a correspondence between Mercury entities and SASS. </p><p>A format that needs a<strong> map to SASS</strong> is not an annotation on SASS, because it&#8217;s a representation of the program in its own right, and the map exists so that debuggers, profilers and the finalizer can move between the two.</p><p>All of which makes the size behaviour the really awkward part, because it argues against the simplest reading of what I just described.</p><p>If the capsule were a second copy of the instruction stream, in whatever denser encoding, its size would have to <strong>follow the SASS.</strong> Longer kernel, longer capsule. That is not what happens. Across seven kernels the SASS spans 384 to 2,048 bytes, a factor of 5.3, while the capsules span only 130 to 350, a factor of 2.7. </p><p>Correlations are carried almost entirely by a single long kernel: the coefficient across all seven is 0.76, and dropping that one point takes it to 0.33. Now, if you rank the kernels by <strong>capsule size</strong> and you get an ordering with no obvious relationship to how much code they contain.</p><p>The cleanest way to see it is a controlled pair. Two of these kernels compile to exactly 384 bytes of SASS. One stores a float, the other multiplies and adds in double precision. <strong>Same code size</strong>, same instruction count, and their capsules are 130 and 194 bytes, a 49% difference. Meanwhile a kernel using <code>sqrtf</code> compiles to 896 bytes of SASS, more than twice the double-precision kernel, and carries a <em>smaller</em> capsule at 184.</p><p>I want to be careful here, because when I first wrote this section I claimed the ordering tracked the variety of operations a kernel performs, and then I checked. It does not. Correlating capsule size against the <strong>number of distinct opcodes</strong> in each kernel gives 0.03, which is nothing at all. I could not find anything that predicts capsule size, and with seven kernels I would not trust a pattern even if I had found one.</p><p>What the double-precision pair does suggest is that whatever the capsule enumerates is closer to a set of requirements than to a program, and <strong>double precision</strong> is a plausible thing to find in such a set. Double-precision throughput is one of the sharpest differences within a Blackwell family: the datacenter parts and the consumer parts are not the same machine in that respect. </p><p>A representation whose job is to let a driver produce correct code for a family member the compiler never saw would need to record exactly this sort of thing, and would not need to grow much when you add another hundred instructions of the same kind.</p><p>So the sizes push against &#8220;<em>a second copy of every instruction</em>&#8221; and toward &#8220;<em>a description of what this code needs.</em>&#8221; They cannot separate the two readings that remain. A compact whole-kernel encoding in a much denser format, and a <strong>fixup table</strong> over the existing SASS naming only the sites a re-finalizer would have to revisit, would both be small and would both fail to scale with instruction count. </p><p>The section type called <code>CUDA_MERCURY_SASS_MAP</code> is the one piece of evidence that leans toward the first, since a table of fixups would not obviously need a map. I cannot settle it from sizes alone, and I doubt anyone can without a driver to watch.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!QrF2!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe20c3f10-bfea-46a3-b840-90d90891b8a3_1503x884.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!QrF2!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe20c3f10-bfea-46a3-b840-90d90891b8a3_1503x884.png 424w, https://substackcdn.com/image/fetch/$s_!QrF2!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe20c3f10-bfea-46a3-b840-90d90891b8a3_1503x884.png 848w, https://substackcdn.com/image/fetch/$s_!QrF2!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe20c3f10-bfea-46a3-b840-90d90891b8a3_1503x884.png 1272w, https://substackcdn.com/image/fetch/$s_!QrF2!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe20c3f10-bfea-46a3-b840-90d90891b8a3_1503x884.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!QrF2!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe20c3f10-bfea-46a3-b840-90d90891b8a3_1503x884.png" width="1456" height="856" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e20c3f10-bfea-46a3-b840-90d90891b8a3_1503x884.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:856,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Scatter plot of Mercury capsule size against SASS size for seven kernels, showing a nearly flat relationship.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Scatter plot of Mercury capsule size against SASS size for seven kernels, showing a nearly flat relationship." title="Scatter plot of Mercury capsule size against SASS size for seven kernels, showing a nearly flat relationship." srcset="https://substackcdn.com/image/fetch/$s_!QrF2!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe20c3f10-bfea-46a3-b840-90d90891b8a3_1503x884.png 424w, https://substackcdn.com/image/fetch/$s_!QrF2!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe20c3f10-bfea-46a3-b840-90d90891b8a3_1503x884.png 848w, https://substackcdn.com/image/fetch/$s_!QrF2!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe20c3f10-bfea-46a3-b840-90d90891b8a3_1503x884.png 1272w, https://substackcdn.com/image/fetch/$s_!QrF2!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe20c3f10-bfea-46a3-b840-90d90891b8a3_1503x884.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Seven kernels on sm_100. If the capsule were a compact second encoding of the instruction stream, these points would sit on a line through the origin. They do not. A double-precision kernel with 384 bytes of SASS carries a 194 byte capsule; a sqrtf kernel with 896 bytes of SASS carries 184. Whatever the capsule enumerates, it is closer to a set of properties than to a program.<span>Measured on sm_100. The double-precision kernel has the smallest SASS and nearly the largest capsule; the sqrtf kernel has the largest SASS and one of the smallest.</span></figcaption></figure></div><p>What it costs is easier to state. On the ten-kernel workload, Mercury sections add 4.2 KB to a 57.8 KB cubin, about 7%, and Blackwell cubins are 2.0x the size of Turing cubins for identical source.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Flhn!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0cce32f6-7c9d-4089-8db0-8fe09f17b5fc_1487x952.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Flhn!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0cce32f6-7c9d-4089-8db0-8fe09f17b5fc_1487x952.png 424w, https://substackcdn.com/image/fetch/$s_!Flhn!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0cce32f6-7c9d-4089-8db0-8fe09f17b5fc_1487x952.png 848w, https://substackcdn.com/image/fetch/$s_!Flhn!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0cce32f6-7c9d-4089-8db0-8fe09f17b5fc_1487x952.png 1272w, https://substackcdn.com/image/fetch/$s_!Flhn!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0cce32f6-7c9d-4089-8db0-8fe09f17b5fc_1487x952.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Flhn!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0cce32f6-7c9d-4089-8db0-8fe09f17b5fc_1487x952.png" width="1456" height="932" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0cce32f6-7c9d-4089-8db0-8fe09f17b5fc_1487x952.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:932,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Stacked bar chart of cubin size composition across twelve architectures, showing text, constant bank, Mercury sections, metadata, debug and ELF scaffolding.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Stacked bar chart of cubin size composition across twelve architectures, showing text, constant bank, Mercury sections, metadata, debug and ELF scaffolding." title="Stacked bar chart of cubin size composition across twelve architectures, showing text, constant bank, Mercury sections, metadata, debug and ELF scaffolding." srcset="https://substackcdn.com/image/fetch/$s_!Flhn!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0cce32f6-7c9d-4089-8db0-8fe09f17b5fc_1487x952.png 424w, https://substackcdn.com/image/fetch/$s_!Flhn!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0cce32f6-7c9d-4089-8db0-8fe09f17b5fc_1487x952.png 848w, https://substackcdn.com/image/fetch/$s_!Flhn!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0cce32f6-7c9d-4089-8db0-8fe09f17b5fc_1487x952.png 1272w, https://substackcdn.com/image/fetch/$s_!Flhn!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0cce32f6-7c9d-4089-8db0-8fe09f17b5fc_1487x952.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Ten kernels, twelve targets, same source file. The SASS roughly doubles from Turing to Blackwell for reasons that have nothing to do with the format, the launch frame more than doubles, and the Mercury sections appear from sm_100 onwards. Nothing here is compressed: this is the cubin as ptxas writes it.<span>Measured with nvcc -cubin. Section sizes read from the ELF section headers; &#8216;ELF scaffolding&#8217; is the file remainder (string tables, symtab, headers, padding).</span></figcaption></figure></div><p>There is <strong>one more hint </strong>about direction of travel, in the disassembler rather than the assembler. <code>nvdisasm</code> in 13.3 has an option called <code>--no-vliw</code>, described as &#8220;<em>conventional mode; disassemble paired instructions in normal syntax, instead of VLIW syntax.</em>&#8221; </p><p>Nothing I compiled for any current target produced paired output. An option to <em>turn off</em> VLIW syntax implies a target where VLIW syntax is the default. </p><p>I would not build a thesis on a <strong>command-line flag,</strong> but if a future architecture issues statically paired instructions, the tooling has already been taught the word.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>The fatbin is a filesystem, and the driver is the resolver</h2><p>A cubin runs on exactly <strong>one compute capability</strong>. Shipping software means shipping several, plus PTX for the ones that do not exist yet, and the container for that is the fatbin. </p><p>It starts with a 16-byte header whose magic is <code>0xBA55ED50</code>, and it is followed by a run of entries, each with its own header and payload.</p><pre><code>$ nvcc -gencode arch=compute_90,code=sm_90 \
       -gencode arch=compute_100,code=[sm_100,compute_100] -fatbin -o saxpy.fatbin saxpy.cu

kind     arch  hdrsz  payload   compressed  uncompressed  flags
CUBIN      90     64     3968            0             0  0x11
CUBIN     100    112     5712            0             0  0x1000011
PTX       100     80      400          399           922  0x8011</code></pre><p>The entry header fields are recoverable by differential compilation: change one thing about the build, diff the bytes, name the field.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!kd9n!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb4c40fd0-9a96-48aa-a305-15247265966b_2163x1460.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!kd9n!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb4c40fd0-9a96-48aa-a305-15247265966b_2163x1460.png 424w, https://substackcdn.com/image/fetch/$s_!kd9n!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb4c40fd0-9a96-48aa-a305-15247265966b_2163x1460.png 848w, https://substackcdn.com/image/fetch/$s_!kd9n!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb4c40fd0-9a96-48aa-a305-15247265966b_2163x1460.png 1272w, https://substackcdn.com/image/fetch/$s_!kd9n!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb4c40fd0-9a96-48aa-a305-15247265966b_2163x1460.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!kd9n!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb4c40fd0-9a96-48aa-a305-15247265966b_2163x1460.png" width="1456" height="983" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b4c40fd0-9a96-48aa-a305-15247265966b_2163x1460.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:983,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:285278,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/207127260?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb4c40fd0-9a96-48aa-a305-15247265966b_2163x1460.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!kd9n!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb4c40fd0-9a96-48aa-a305-15247265966b_2163x1460.png 424w, https://substackcdn.com/image/fetch/$s_!kd9n!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb4c40fd0-9a96-48aa-a305-15247265966b_2163x1460.png 848w, https://substackcdn.com/image/fetch/$s_!kd9n!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb4c40fd0-9a96-48aa-a305-15247265966b_2163x1460.png 1272w, https://substackcdn.com/image/fetch/$s_!kd9n!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb4c40fd0-9a96-48aa-a305-15247265966b_2163x1460.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>Reconstructed by construction, not from documentation. Every row was fixed by building the same kernel with one thing changed and diffing the header bytes. Fields I did not manage to move are omitted.</em></p><h3>PTX is not the cheap option</h3><p>The standard advice is to always ship PTX for the newest virtual architecture so future GPUs have something to JIT. The advice generally is right, and the reason people give for it is usually wrong. </p><p><strong>PTX is not a compact representation</strong>. For the ten-kernel workload the PTX is 37.3 KB against a 35.9 KB Hopper cubin, and on Blackwell it is 48.0 KB against 57.8 KB. It is the same order of magnitude in both directions, and after compression the gap narrows further because PTX is text and compresses beautifully.</p><p>What PTX buys is <strong>coverage of chips</strong> that do not exist yet, at the cost of a JIT compile on first use, per process, unless the driver&#8217;s compute cache is warm. </p><p>What it does not buy is what most people think: <strong>forward compatibility </strong>does not extend to the instructions anyone actually cares about, because PTX built for an architecture-conditional target is not forward compatible at all. </p><p>If your kernel needs <code>wgmma</code> or <code>tcgen05</code>, it needs <code>sm_90a</code> or <code>sm_100a</code>, and the compatibility story ends there. The <code>f</code> targets exist precisely because that cliff was too steep, and the Mercury capsule is how they made it less steep.</p><p>That extended header is the good part. The <code>sm_100</code> entry above has a 112-byte header instead of 64, and the extra 48 bytes are a length-prefixed copy of the <code>.nv.compat</code> capability block from inside the cubin.</p><pre><code>header bytes at offset 0x40:
48 00 00 00  20 00 00 00     -- payload at 0x48, length 0x20
02 09 00 00  02 02 01 00  03 0d 01 01  03 07 01 01
02 03 00 00  04 0b 08 00  09 00 00 00  00 00 00 00
   ^^ the same TLV stream that .nv.compat holds inside the ELF</code></pre><p>The capability claim is duplicated into the index so that whatever picks an entry never has to open the ELF. That is a loader optimization, and it is also a small confirmation of what the compat block is for: it is the thing a resolver matches against a device before committing to a payload.</p><h3>Compression is off by default for the part that matters</h3><p>CUDA has a <code>--compress-mode</code> flag with values <code>none</code>, <code>speed</code>, <code>balance</code>, <code>size</code> and <code>default</code>. Running all five over the same two-cubin-plus-PTX fatbin gives a result worth knowing.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Pt5x!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6f49747e-ba08-4693-89fe-0fb22883c10c_1699x753.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Pt5x!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6f49747e-ba08-4693-89fe-0fb22883c10c_1699x753.png 424w, https://substackcdn.com/image/fetch/$s_!Pt5x!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6f49747e-ba08-4693-89fe-0fb22883c10c_1699x753.png 848w, https://substackcdn.com/image/fetch/$s_!Pt5x!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6f49747e-ba08-4693-89fe-0fb22883c10c_1699x753.png 1272w, https://substackcdn.com/image/fetch/$s_!Pt5x!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6f49747e-ba08-4693-89fe-0fb22883c10c_1699x753.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Pt5x!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6f49747e-ba08-4693-89fe-0fb22883c10c_1699x753.png" width="1456" height="645" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6f49747e-ba08-4693-89fe-0fb22883c10c_1699x753.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:645,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Horizontal bar chart of fatbin size under four compression modes, with the size mode dramatically smaller.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Horizontal bar chart of fatbin size under four compression modes, with the size mode dramatically smaller." title="Horizontal bar chart of fatbin size under four compression modes, with the size mode dramatically smaller." srcset="https://substackcdn.com/image/fetch/$s_!Pt5x!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6f49747e-ba08-4693-89fe-0fb22883c10c_1699x753.png 424w, https://substackcdn.com/image/fetch/$s_!Pt5x!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6f49747e-ba08-4693-89fe-0fb22883c10c_1699x753.png 848w, https://substackcdn.com/image/fetch/$s_!Pt5x!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6f49747e-ba08-4693-89fe-0fb22883c10c_1699x753.png 1272w, https://substackcdn.com/image/fetch/$s_!Pt5x!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6f49747e-ba08-4693-89fe-0fb22883c10c_1699x753.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">In default, speed and balance mode, only the PTX gets compressed and both cubins are stored raw. Only -compress-mode=size compresses SASS, and on this input it takes the fatbin from 10,352 bytes to 3,000.<span>Measured. In default, speed and balance mode the two cubins are stored raw and only the PTX is compressed; -compress-mode=size compresses the cubins as well.</span></figcaption></figure></div><p>The flag values behave like two different algorithms rather than one algorithm at three settings: <code>speed</code> sets flag bit 13 and gets 504 bytes out of 922, while <code>balance</code> and <code>size</code> set bit 15 and get 400. </p><p>On the larger <strong>ten-kernel workload</strong> across twelve targets, <code>size</code> mode takes 537 KB down to 75 KB, an 86% reduction, for no measurable build time cost at this scale. </p><p>If you ship CUDA binaries and have not tried it, that is the cheapest win in this article.</p><h3>How it gets into your executable</h3><p>The host side is one of the few places NVIDIA actually publishes the format, in <code>fatbinary_section.h</code> in the toolkit include directory.</p><pre><code>#define FATBINC_MAGIC   0x466243B1
#define FATBINC_VERSION 1
typedef struct {
  int magic;
  int version;
  const unsigned long long* data;
  void *filename_or_fatbins;
} __fatBinC_Wrapper_t;

#define FATBIN_CONTROL_SECTION_NAME  &#8220;.nvFatBinSegment&#8221;
#define FATBIN_DATA_SECTION_NAME     &#8220;.nv_fatbin&#8221;</code></pre><p>Compile a <code>.cu</code> file to an object and you get exactly that: a <strong>24-byte wrapper </strong>in <code>.nvFatBinSegment</code> holding the magic and a relocation pointing at the fatbin bytes in <code>.nv_fatbin</code>, plus a section called <code>__nv_module_id</code>, plus a constructor.</p><pre><code>$ readelf -x .nvFatBinSegment saxpy.o
  0x00000000 <strong>b1436246</strong> 01000000 00000000 00000000 .CbF............
$ readelf -rW saxpy.o | grep nvFatBinSegment
  R_X86_64_64  .nv_fatbin + 0

$ cat saxpy.cudafe1.stub.c        # generated by nvcc --keep
static void __nv_cudaEntityRegisterCallback(void **__T0) {
  __cudaRegisterEntry(__T0, (void(*)(int,float,const float*,float*))saxpy,
                      _Z5saxpyifPKfPf, (-1)); }
static void <strong>__sti____cudaRegisterAll</strong>(void) __attribute__((__constructor__));
static void __sti____cudaRegisterAll(void) {
  <strong>__cudaRegisterBinary</strong>(__nv_cudaEntityRegisterCallback); }

void __device_stub__Z5saxpyifPKfPf(int p0, float p1, const float *p2, float *p3){
  __cudaLaunchPrologue(4);
  __cudaSetupArgSimple(p0, 0UL);  __cudaSetupArgSimple(p1, 4UL);
  __cudaSetupArgSimple(p2, 8UL);  __cudaSetupArgSimple(p3, 16UL);
  __cudaLaunch((char *)saxpy); }</code></pre><p>So a CUDA program registers its device code from an ELF constructor, before <code>main</code>. The <strong>argument offsets in the stub</strong>, 0, 4, 8 and 16, are the same offsets the cubin declared in <code>EIATTR_KPARAM_INFO</code>. </p><p>Host and device agree on the layout because both were generated from the same front end, and the binary carries the agreement in two places.</p><p>Two variants are worth noting because they trip people up. With <code>-rdc=true</code>, the device code is relocatable, the cubin becomes <code>ET_REL</code> with <code>.rela.text.*</code> sections and an <code>EIATTR_EXTERNS</code> attribute, and it lands in a differently named section, <code>__nv_relfatbin</code>, until the device link step turns it into a normal <code>.nv_fatbin</code>. The cost shows up in the machine code, and it is not subtle:</p><pre><code>$ nvcc -arch=sm_90 -rdc=true -c a.cu b.cu &amp;&amp; nvcc -arch=sm_90 -dlink -o dl.o a.o b.o
$ cuobjdump -sass dl.o
    Function : _Z5applyPfPKffi
        /*0100*/  <strong>CALL.ABS.NOINC</strong> 0x0 ;
    Function : _Z5scaleff
        /*0000*/  FFMA R4, R4, R5, 1 ;
        /*0010*/  <strong>RET.ABS.NODEC</strong> R20 0x0 ;</code></pre><p>A one-instruction device function became a real call, with a stack frame, a return address register and a separate <code>.text</code> section, where whole-program compilation would have inlined it into a single FFMA. </p><p>That is what <code>-dlto</code> exists to undo: with LTO the fatbin entry is kind 8 and <strong>the instruction selection</strong> is deferred to link time so the inlining can happen after all translation units are visible. </p><p>In my sandbox the LTO device link fails for want of a driver-side link library, so I can report the container shape (1,944 compressed bytes expanding to 2,548) but not the linked output.</p><h3>What the driver does with all this</h3><p>The loading path is the reason the format looks the way it does, and it is worth stating plainly because the layering is easy to lose. The constructor calls <code>__cudaRegisterBinary</code> with a pointer to the wrapper. </p><p>The runtime <strong>hands the fatbin to the driver</strong>, which reads the index and selects one payload for the device it is about to use: an exact-architecture cubin if there is one, otherwise a compatible one, otherwise PTX to be compiled on the spot. The selected payload becomes a module. </p><p>The module&#8217;s <code>.nv.info</code> tells the driver how many registers each kernel needs, how much shared memory to reserve, how big a stack to allocate given the <strong>call graph</strong>, and where in constant bank 0 to write the arguments. Only then can a launch happen, and a launch is essentially a memcpy of the argument block into the frame the cubin described, followed by a grid dispatch.</p><p>Everything in the binary that is not machine code exists to let that sequence happen without the driver consulting anything else. That is why the <strong>parameter layout is in the object</strong> rather than in a header, why the exit offsets are precomputed, why the capability claim is duplicated into the fatbin index, and why the reserved shared memory has a symbol. A loader that had to infer any of it would be slower and more fragile.</p><p><strong>Two changes to this path</strong> are worth knowing. Since CUDA 12.0 there is a second module abstraction, the library and kernel API (<code>cuLibraryLoadData</code>, <code>cuKernel</code>), which decouples a loaded artifact from a specific context, and it exists largely because the old model made lazy loading awkward. </p><p>And since 12.2 on Linux and 12.3 everywhere, lazy loading is the default: modules load on first use of a symbol from them, and kernels within a module load on first <code>cuModuleGetFunction</code>. </p><p>The requirement is a runtime of 11.7 or later, statically linked into whatever built the library, which means an old dependency in your stack can quietly opt you out. <strong>NVIDIA&#8217;s own documentation</strong> puts the requirement plainly: libraries compiled against older runtimes load all modules eagerly.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>What the format costs, in bytes and seconds</h2><p>All of the above has a price, and it is paid by everyone who ships a wheel.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Lwpg!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa0373bfe-1212-465b-8b0b-289abb835e94_1487x885.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Lwpg!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa0373bfe-1212-465b-8b0b-289abb835e94_1487x885.png 424w, https://substackcdn.com/image/fetch/$s_!Lwpg!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa0373bfe-1212-465b-8b0b-289abb835e94_1487x885.png 848w, https://substackcdn.com/image/fetch/$s_!Lwpg!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa0373bfe-1212-465b-8b0b-289abb835e94_1487x885.png 1272w, https://substackcdn.com/image/fetch/$s_!Lwpg!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa0373bfe-1212-465b-8b0b-289abb835e94_1487x885.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Lwpg!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa0373bfe-1212-465b-8b0b-289abb835e94_1487x885.png" width="1456" height="867" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a0373bfe-1212-465b-8b0b-289abb835e94_1487x885.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:867,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Two panels: fatbin size versus number of targets, and nvcc wall clock versus number of targets.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Two panels: fatbin size versus number of targets, and nvcc wall clock versus number of targets." title="Two panels: fatbin size versus number of targets, and nvcc wall clock versus number of targets." srcset="https://substackcdn.com/image/fetch/$s_!Lwpg!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa0373bfe-1212-465b-8b0b-289abb835e94_1487x885.png 424w, https://substackcdn.com/image/fetch/$s_!Lwpg!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa0373bfe-1212-465b-8b0b-289abb835e94_1487x885.png 848w, https://substackcdn.com/image/fetch/$s_!Lwpg!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa0373bfe-1212-465b-8b0b-289abb835e94_1487x885.png 1272w, https://substackcdn.com/image/fetch/$s_!Lwpg!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa0373bfe-1212-465b-8b0b-289abb835e94_1487x885.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Ten kernels. Each additional real target is a full independent compilation and a full independent copy. The relationship is linear in both size and time, which is the whole problem: there is no sharing between targets, because there is nothing to share.<span>Measured: CUDA 13.3.73, 10 templated kernels (4 GEMM, 4 elementwise, 2 reductions), single-threaded nvcc, no GPU present.</span></figcaption></figure></div><p>This is why PyTorch&#8217;s release engineering discussions read the way they do. When the<strong> CUDA 12.8 build matrix</strong> had to add Blackwell, the team explicitly rejected keeping the older architectures because of binary size. The trade is not subtle: adding a target adds a copy of every kernel in the library, and machine learning libraries have a great many kernels.</p><p>You can see the endpoint of that logic by taking apart what NVIDIA itself ships. I pulled the <strong>cuBLAS 13.6.0.2 wheel</strong> from PyPI and parsed the <code>.nv_fatbin</code> sections directly.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!uff2!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F529ff6e2-6245-42e5-a852-82e3382b2846_2276x1486.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!uff2!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F529ff6e2-6245-42e5-a852-82e3382b2846_2276x1486.png 424w, https://substackcdn.com/image/fetch/$s_!uff2!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F529ff6e2-6245-42e5-a852-82e3382b2846_2276x1486.png 848w, https://substackcdn.com/image/fetch/$s_!uff2!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F529ff6e2-6245-42e5-a852-82e3382b2846_2276x1486.png 1272w, https://substackcdn.com/image/fetch/$s_!uff2!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F529ff6e2-6245-42e5-a852-82e3382b2846_2276x1486.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!uff2!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F529ff6e2-6245-42e5-a852-82e3382b2846_2276x1486.png" width="1456" height="951" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/529ff6e2-6245-42e5-a852-82e3382b2846_2276x1486.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:951,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:234858,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/207127260?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F529ff6e2-6245-42e5-a852-82e3382b2846_2276x1486.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!uff2!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F529ff6e2-6245-42e5-a852-82e3382b2846_2276x1486.png 424w, https://substackcdn.com/image/fetch/$s_!uff2!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F529ff6e2-6245-42e5-a852-82e3382b2846_2276x1486.png 848w, https://substackcdn.com/image/fetch/$s_!uff2!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F529ff6e2-6245-42e5-a852-82e3382b2846_2276x1486.png 1272w, https://substackcdn.com/image/fetch/$s_!uff2!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F529ff6e2-6245-42e5-a852-82e3382b2846_2276x1486.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Two libraries, 6,860 fatbin entries, 178 MB on disk expanding to 1.8 GB of device code. And <strong>every single entry</strong> is compressed, which tells you that NVIDIA does not ship its own libraries with the default compression mode.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!FkGO!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa6869b1f-f7ef-48eb-97aa-14c055cc5911_1530x952.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!FkGO!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa6869b1f-f7ef-48eb-97aa-14c055cc5911_1530x952.png 424w, https://substackcdn.com/image/fetch/$s_!FkGO!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa6869b1f-f7ef-48eb-97aa-14c055cc5911_1530x952.png 848w, https://substackcdn.com/image/fetch/$s_!FkGO!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa6869b1f-f7ef-48eb-97aa-14c055cc5911_1530x952.png 1272w, https://substackcdn.com/image/fetch/$s_!FkGO!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa6869b1f-f7ef-48eb-97aa-14c055cc5911_1530x952.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!FkGO!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa6869b1f-f7ef-48eb-97aa-14c055cc5911_1530x952.png" width="1456" height="906" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a6869b1f-f7ef-48eb-97aa-14c055cc5911_1530x952.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:906,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Left: device code per architecture for the two cuBLAS libraries. Right: section sizes inside libcublasLt.so.13.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Left: device code per architecture for the two cuBLAS libraries. Right: section sizes inside libcublasLt.so.13." title="Left: device code per architecture for the two cuBLAS libraries. Right: section sizes inside libcublasLt.so.13." srcset="https://substackcdn.com/image/fetch/$s_!FkGO!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa6869b1f-f7ef-48eb-97aa-14c055cc5911_1530x952.png 424w, https://substackcdn.com/image/fetch/$s_!FkGO!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa6869b1f-f7ef-48eb-97aa-14c055cc5911_1530x952.png 848w, https://substackcdn.com/image/fetch/$s_!FkGO!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa6869b1f-f7ef-48eb-97aa-14c055cc5911_1530x952.png 1272w, https://substackcdn.com/image/fetch/$s_!FkGO!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa6869b1f-f7ef-48eb-97aa-14c055cc5911_1530x952.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Hopper is still the largest single target in cuBLASLt at 37.4 MB of compressed SASS, more than consumer Blackwell and datacenter Blackwell individually. The right panel is the part I did not expect: device code is only 27% of the file.<span>Measured on the cuBLAS 13.6.0.2 wheel from PyPI. Right panel counts only file-resident sections, which sum to 491.4 MB of the 492.7 MB file; .bss and .lbss add another 99 MB of runtime memory but zero bytes on disk. cuBLASLt holds 5,599 fatbin entries (5,309 cubins, 290 PTX), 132.3 MB compressed, expanding to 1,361 MB.</span></figcaption></figure></div><p>The right-hand panel is worth sitting with. In cuBLASLt, 101 MB is host code, and 78 MB is a section called <code>.cask_resource</code>. </p><p>One detail worth noting before anyone adds those numbers up: the library also declares 99 MB of <code>.bss</code> and <code>.lbss</code>, and those are NOBITS sections, which means they cost 99 MB of memory at run time and zero bytes on disk. </p><p>Counting only the sections that actually occupy file space gets you 491.4 MB of the 492.7 MB file, and the rest is headers and padding. I cannot tell you <strong>what CASK stands for,</strong> because NVIDIA has never said, and the name only surfaces publicly in cuDNN and TensorRT diagnostics. </p><p>What I can tell you is that the strings in that section include 2,429 mangled C++ names beginning <code>N11cask</code>, so it is a namespace, and that the rest of the section is a catalog. Pull strings out of it and you get entries like this:</p><pre><code>cutlass3x_sm103_bstensorop_s256x256x96gemm_block_scaled_ue4m3xe2m1_ue4m3xe2m1_
  f32_bf16_ue4m3xe2m1_256x256x768_0_tnn_align...</code></pre><p>A CUTLASS 3.x kernel for <code>sm_103</code> doing a block-scaled GEMM with FP8 and FP4 operand types at a 256x256x96 tile. </p><p>Counting distinct kernel-shaped names across the resource and read-only sections gives <strong>25,401</strong>, with the generation prefixes still visible in the histogram: 5,329 mentioning <code>sm90</code>, 5,015 <code>cutlass3x</code>, 1,956 <code>ampere</code>, 808 <code>volta</code>, 274 <code>turing</code>.</p><p>That number is the moat, expressed as an artifact count rather than an argument. <strong>Not one clever kernel</strong>. Twenty-five thousand parameterised ones, plus 78 MB of metadata describing which to pick, plus 101 MB of host code doing the picking. A competitor can match the code generator. </p><p>Matching the catalog and the selection heuristics is a different kind of work, and it is the kind that does not benefit much from being smart.</p><p>NVIDIA is visibly feeling the weight of it. cuDNN 9.5 shipped a build configuration called <code>GRAPH_JIT_ONLY</code> that keeps the runtime kernel generation engines and drops the <strong>precompiled ones </strong>specifically to cut binary size, which is the vendor arriving at the same trade its users have been making by hand: generate at run time or ship the archive.</p><p>It also explains why lazy loading became the default. If a process eagerly loaded <strong>all 5,309 cubins</strong> in cuBLASLt it would spend its startup budget on kernels it will never call. </p><p>Lazy loading arrived as opt-in in CUDA 11.7, became the Linux default in 12.2 and the universal default in 12.3, and defers module and kernel loading until first use. </p><p>In a world where a <strong>single library</strong> holds thousands of independently loadable objects, that stops being an optimization and becomes a requirement.</p><h3>Debug information is not free either</h3><p>For the same ten kernels on <code>sm_90</code>: 35.9 KB release, 75.6 KB with <code>-lineinfo</code>, 205.9 KB with <code>-G</code>. The last one is not just extra sections. </p><p>It changes code generation, taking <code>.text</code> from 16.0 KB to 57.5 KB, because full device debug turns off most optimization. <code>-lineinfo</code> is the option you want in production builds: it doubles the cubin and leaves the code alone.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/p/how-cuda-binaries-actually-work?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/p/how-cuda-binaries-actually-work?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><div><hr></div><h2>Why any of this matters if you are not writing a compiler</h2><p>Three reasons, in increasing order of how much I care about them.</p><p>The first is operational. Binary size and load time are real costs in inference deployments, container images and cold starts, and they are <strong>almost entirely determined by choices</strong> in the format: how many targets, PTX or not, which compression mode, lazy or eager. Those are four flags. Most teams have never touched them.</p><p>The second is diagnostic. A cubin tells you its register count, its shared memory usage, its stack frame, its parameter layout, the toolkit that built it and the <strong>flags it was built with</strong>, and you can read all of that out of a shipped wheel without running anything. </p><p>If you are trying to work out why somebody else&#8217;s kernel spills, or which architectures a vendor actually optimized for versus which they merely support, the <strong>binary is a better source</strong> than the changelog. The <code>-res-usage</code> flag on <code>cuobjdump</code> is the fastest occupancy audit I know.</p><h4>Auditing somebody else&#8217;s wheel</h4><pre><code>cuobjdump -lelf libfoo.so          # which architectures, how many cubins
cuobjdump -lptx libfoo.so          # is there PTX for anything newer
cuobjdump -res-usage libfoo.so     # registers, shared memory, stack, per kernel
cuobjdump -elf libfoo.so | grep -A5 &#8220;Toolkit Information&#8221;   # which ptxas built it
readelf -SW libfoo.so | grep nv_fatbin                      # how much of the file is device code</code></pre><p>Five commands, no GPU, no source. The last one is the one that surprises people: if <code>.nv_fatbin</code> is 85% of a library, you are shipping a kernel archive with a thin API on top, and any conversation about image size has to start there.</p><p>The third is strategic, and it is why I went looking in the first place. </p><p>There is a long line of work reverse-engineering these formats: <code>decuda</code> for the earliest chips, <code>asfermi</code> for Fermi, <code>KeplerAs</code>, Scott Gray&#8217;s <code>maxas</code> for <strong>Maxwell </strong>(the toolchain behind the fast convolution kernels of the deep learning era), <code>turingas</code>, <code>CuAssembler</code> for everything through Ampere, NVBit for dynamic instrumentation, and now work lifting SASS back into typed compiler IR. </p><p>Every one of these projects <strong>had to rediscover the container </strong>before it could touch the code, and every one of them targets an architecture at least two generations behind current silicon. That lag is the moat&#8217;s actual shape. </p><p>It is <strong>not that the format is unknowable</strong>, instead it&#8217;s that knowing it is a full-time job that resets on a yearly cadence, and the number of people doing it is small enough to name.</p><p>Mercury is the part of this that I think changes the picture, and not in the direction I expected. Publishing an interface and keeping the lowering is NVIDIA&#8217;s oldest move: <strong>PTX is public</strong>, <code>ptxas</code> is not. </p><p>What the capsule does is insert a second, private layer at exactly the point where a competitor would want to attach, and it does it for a reason customers asked for: one binary that keeps working on the next chip in the family. The <em>compatibility benefit</em> is real. </p><p>The side effect is that the interesting <strong>representation of a Blackwell kernel </strong>is no longer the SASS you can disassemble, and the format that carries it has no public name, no documented encoding, and its own relocation types.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Four things I would change, and one I would not</h2><p>Having read a few hundred of these files, I have opinions about the format that have nothing to do with the moat.</p><p><strong>Publish the metadata, keep the encoding.</strong> <em>There is no competitive information in </em><code>EIATTR_KPARAM_INFO</code><em>. It is a calling convention. Documenting the </em><code>.nv.info</code><em> attribute codes and the fatbin entry header would cost NVIDIA nothing and would stop every profiler, virtualization layer, checkpointer and build tool from rediscovering the same twenty structs. The instruction encoding is a different argument and I understand the refusal. The container is not.</em></p><p><strong>Make </strong><code>size</code><strong> the default compression mode.</strong> <em>The current default compresses PTX and leaves the SASS raw, which is the wrong way round: SASS is the bulk and the part that never gets read by a human. NVIDIA already ships its own libraries with everything compressed. Every framework wheel on PyPI is paying for a default its vendor does not use.</em></p><p><strong>Give the toolkit a way to strip PTX.</strong> <em>There is </em><code>nvprune</code><em> for removing architectures, but the common case in a container build is &#8220;I know exactly which GPUs this image runs on, remove the JIT fallback.&#8221; Doing that today means unpacking and repacking fatbins with </em><code>nvFatbin</code><em>, which is a strange amount of work for what should be a link flag.</em></p><p><strong>Version the format visibly.</strong><em> Between the ABI version byte, the OS ABI byte that changed value, the toolkit note, the API version attribute and the Mercury ISA version, there are five different notions of &#8220;which format is this&#8221; in one file, and no single field that a tool can check to know whether it will understand what follows. A cubin that a future tool cannot parse should say so in its first sixteen bytes.</em></p><p>What <strong>I would not change personally</strong> is the decision to make the binary self-describing rather than pushing the information into headers or a side-channel. </p><p>It is the reason <strong>NVBit can instrument a kernel</strong> it has never seen, the reason <code>cuobjdump -res-usage</code> works on a stranger&#8217;s wheel, and the reason this article was possible on a machine with no GPU in it. Verbose, redundant, self-describing binaries are a gift to everyone downstream, including the people trying to work out what you shipped. </p><p>Most of what is wrong with the CUDA binary format is that it is undocumented. Very little of it is that it is badly designed.</p><h4><span>Where I might be wrong</span></h4><ul><li><p><strong>The capsule might be much less than I think.</strong> If it turns out to be a small fixup table rather than a re-encodable program, then &#8220;second copy of every kernel&#8221; is the wrong phrase and the right one is &#8220;compatibility annotations.&#8221; The sublinear size scaling in the capsule plot above is evidence for the smaller reading; the existence of a section type called <code>CUDA_MERCURY_SASS_MAP</code> is evidence for the larger one. I could not settle it.</p></li><li><p><strong>I have no Blackwell hardware.</strong> Nothing here was executed. A claim about what a driver does with a capsule is, from my position, a claim about what an artifact appears designed for. Everything about run time behaviour, including JIT cost and the actual effect of lazy loading, is cited rather than measured.</p></li><li><p><strong>The OS ABI byte is a mess and I may be reading the mess wrong.</strong> LLVM lists 51 and 41; cubins carry 65. My reading is that the second constant is <code>0x41</code> written as decimal, but it is equally possible NVIDIA changed the value again and LLVM has not caught up. I only checked 13.3 output.</p></li><li><p><strong>The fatbin flag bits are inferred from a handful of data points.</strong> I set them by construction using the compression modes and the LTO path; other bits in that word certainly mean things I did not exercise. Byte 0 of <code>e_flags</code>, which flips from 4 to 2 at Blackwell, I could not explain at all.</p></li><li><p><strong>My library numbers are one version of one vendor library.</strong> cuBLAS is the extreme case for kernel count. Generalizing &#8220;27% device code&#8221; to other libraries would be wrong.</p></li><li><p><strong>The feature-tier experiment used one instruction.</strong> <code>tcgen05.fence</code> behaves as described, and I have no reason to think the mechanism differs for other family-specific instructions, but one probe is one probe.</p></li></ul><h3><span>Predictions</span></h3><ol><li><p><strong>By the end of 2027</strong>, at least one open-source project will publish a working parser for the Mercury capsule format, and it will be motivated by GPU virtualization or checkpointing rather than by performance.</p></li><li><p><strong>NVIDIA will not document the capsule encoding</strong>, on the same terms it has never documented SASS encodings, through at least CUDA 15.</p></li><li><p><strong>By the end of 2027</strong>, <code>-compress-mode=size</code> or its equivalent will be the default in at least one major framework&#8217;s release build, driven by wheel size limits rather than by anyone reading the flag documentation.</p></li><li><p><strong>The next architecture after Blackwell will ship a SASS encoding with statically paired instructions</strong>, and <code>nvdisasm</code>&#8216;s VLIW syntax will be why we find out.</p></li><li><p><strong>Family-conditional targets will absorb the architecture-conditional ones in practice</strong>: by the end of 2028, the fraction of shipped cubins carrying <code>a</code> suffixes will fall as libraries move to <code>f</code>, because one binary per family is a better deal than one per chip.</p></li></ol><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Dossier</h2><p><strong>Tier A, measured directly and reproducible from the appendix. </strong></p><ul><li><p><em>ELF header field values and the </em><code>e_flags</code><em> architecture encoding, including the absence of the virtual architecture from that word. </em></p></li><li><p><em>Section inventories and sizes for all twelve targets. </em></p></li><li><p><em>The </em><code>.nv.info</code><em> and </em><code>.nv.compat</code><em> attribute contents as decoded by </em><code>cuobjdump</code><em>, including the </em><code>ISA_CLASS</code><em> transition when a family-specific instruction is used. </em></p></li><li><p><em>The rejection of </em><code>tcgen05.fence</code><em> on </em><code>sm_100</code><em> and </em><code>sm_120a</code><em> and its acceptance on </em><code>sm_100f</code><em>, </em><code>sm_100a</code><em> and </em><code>sm_103f</code><em>. </em></p></li><li><p><em>The 113 attribute names and the 119 relocation names. Parameter base offsets and constant bank sizes. </em></p></li><li><p><em>Control field bit positions and their per-kernel distribution. </em></p></li><li><p><em>Fatbin entry field offsets, compression flag bits and all size and timing measurements. </em></p></li><li><p><em>The </em><code>CALL.ABS.NOINC</code><em> cost of relocatable device code. </em></p></li><li><p><em>The cuBLAS section and entry statistics, the 25,401 kernel names and the 2,429 </em><code>cask</code><em> mangled names. </em></p></li><li><p><em>The presence, sizes and scaling of the Mercury sections.</em></p></li></ul><p><strong>Tier B, documented by NVIDIA, in shipped headers, or in third-party source.</strong><span> </span></p><ul><li><p><em>The fatbin wrapper struct and section names (</em><code>fatbinary_section.h</code><em>). </em></p></li><li><p><em>The family and architecture-conditional target semantics, including the statement that </em><code>sm_100</code><em> and </em><code>sm_100f</code><em> are aliases absent family features (Programming Guide, PTX ISA, the CUDA 12.9 blog post). </em></p></li><li><p><em>Lazy loading history and defaults. </em></p></li><li><p><em>Dropped architecture support in CUDA 13.0 and the Thor renumbering. </em></p></li><li><p><code>EM_CUDA</code><em>, </em><code>ELFOSABI_CUDA</code><em> and </em><code>ELFOSABI_CUDA_V2</code><em> in LLVM. </em></p></li><li><p><em>cuDNN&#8217;s </em><code>GRAPH_JIT_ONLY</code><em> configuration. </em></p></li><li><p><em>The structure of the control field, which is documented in the literature rather than by the vendor.</em></p></li></ul><p><strong>Tier C, inferred from artifacts.</strong></p><ul><li><p><em>Mercury is a re-finalizable intermediate representation and that the capsule exists to serve family-conditional compatibility. </em></p></li><li><p><em>That the duplicated compat block in the fatbin entry header is a loader fast path. </em></p></li><li><p><code>ptxas</code><em>&#8216;s hidden finalizer options correspond to a driver-side code path. </em></p></li><li><p><code>.cask_resource</code><em> is a kernel selection catalog. </em></p></li><li><p><em>That the </em><code>ELFOSABI_CUDA_V2</code><em> constant is a hexadecimal value written as decimal.</em></p></li></ul><p><strong>Tier D, speculation, flagged as such in the text.</strong></p><ul><li><p><code>--no-vliw</code><em> implies a future paired-issue ISA. </em></p></li><li><p><code>EIATTR_COROUTINE_RESUME_ID_OFFSETS</code><em> points at an unannounced feature. </em></p></li><li><p><code>sm_88</code><em> is the Nintendo Switch 2, which comes from community references and not from anything I could check.</em></p></li></ul><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Reproduce this</h2><p>No GPU required. On <strong>x86-64 Linux:</strong></p><pre><code>pip install nvidia-cuda-nvcc==13.3.73 nvidia-cuda-cuobjdump==13.3.73 \
            nvidia-cuda-nvdisasm==13.3.73 nvidia-cuda-crt==13.3.73 \
            nvidia-cuda-runtime nvidia-cuda-cccl
export PATH=$(python3 -c &#8220;import sysconfig,os;print(os.path.join(sysconfig.get_paths()[&#8217;purelib&#8217;],&#8217;nvidia/cu13/bin&#8217;))&#8221;):$PATH

cat &gt; saxpy.cu &lt;&lt;&#8217;EOF&#8217;
#include &lt;cuda_runtime.h&gt;
__global__ void saxpy(int n, float a, const float* x, float* y) {
    int i = blockIdx.x * blockDim.x + threadIdx.x;
    if (i &lt; n) y[i] = a * x[i] + y[i];
}
EOF

# the container and its metadata
nvcc -arch=sm_90  -cubin -o k90.cubin  saxpy.cu
nvcc -arch=sm_100 -cubin -o k100.cubin saxpy.cu
cuobjdump -elf k90.cubin | less
diff &lt;(cuobjdump -elf k90.cubin | sed -n &#8216;/^Sections:/,/^$/p&#8217;) \
     &lt;(cuobjdump -elf k100.cubin | sed -n &#8216;/^Sections:/,/^$/p&#8217;)

# the capability claim, and the accelerator bit
for a in sm_100 sm_100a sm_100f; do nvcc -arch=$a -cubin -o c.cubin saxpy.cu
  echo &#8220;== $a&#8221;; cuobjdump -elf c.cubin | grep -A1 EICOMPAT | grep -E &#8220;Attr|Value&#8221;; done

# the attribute namespace and the finalizer options
strings -a $(command -v nvdisasm) | grep -o &#8220;EIATTR_[A-Z0-9_]*&#8221; | sort -u
strings -a $(command -v ptxas)    | grep -iE &#8220;capmerc|finaliz|R_MERCURY&#8221;

# the feature tiers, as an experiment rather than a diagram
cat &gt; tc.cu &lt;&lt;&#8217;EOF&#8217;
#include &lt;cuda_runtime.h&gt;
__global__ void k(unsigned* o){
  asm volatile(&#8221;tcgen05.fence::before_thread_sync;&#8221; ::: &#8220;memory&#8221;);
  o[0]=1;
}
EOF
for a in sm_100 sm_100f sm_100a sm_103f sm_120a; do printf &#8220;%-9s &#8220; $a
  nvcc -arch=$a -cubin -o t.cubin tc.cu 2&gt;&amp;1 | head -1; done
for a in sm_100f sm_100a; do nvcc -arch=$a -cubin -o t_$a.cubin tc.cu
  echo &#8220;== $a&#8221;; cuobjdump -elf t_$a.cubin | grep -A2 EICOMPAT | grep -E &#8220;Attribute|Value&#8221;; done

# is the virtual architecture in e_flags? (no)
for v in 75 80 90; do nvcc -gencode arch=compute_$v,code=sm_90 -cubin -o v.cubin saxpy.cu
  python3 -c &#8220;import struct;print(&#8217;compute_$v -&gt; 0x%08x&#8217;%struct.unpack_from(&#8217;&lt;I&#8217;,open(&#8217;v.cubin&#8217;,&#8217;rb&#8217;).read(),48)[0])&#8221;
  cuobjdump -elf v.cubin | grep &#8220;Virtual SM&#8221;; done

# the field names nvdisasm knows
strings -a $(command -v nvdisasm) | grep -o &#8220;EF_CUDA_[A-Z0-9_]*&#8221; | sort -u

# scheduling control bits, 21 bits starting at 105 of each 128-bit word
nvdisasm -c -hex k90.cubin | head -40
python3 - &lt;&lt;&#8217;PY&#8217;
import struct
d=open(&#8217;k90.cubin&#8217;,&#8217;rb&#8217;).read()
o,=struct.unpack_from(&#8217;&lt;Q&#8217;,d,0x28); es,n,si=struct.unpack_from(&#8217;&lt;HHH&#8217;,d,0x3a)
sh=[struct.unpack_from(&#8217;&lt;IIQQQQIIQQ&#8217;,d,o+i*es) for i in range(n)]
st=sh[si][4]; nm=lambda x: d[st+x:d.index(b&#8217;\0&#8217;,st+x)].decode()
for s in sh:
    if not nm(s[0]).startswith(&#8217;.text.&#8217;): continue
    b=d[s[4]:s[4]+s[5]]
    for i in range(0,len(b),16):
        lo,hi=struct.unpack_from(&#8217;&lt;QQ&#8217;,b,i); w=(hi&lt;&lt;64)|lo
        print(f&#8221;{i:04x} stall={(w&gt;&gt;105)&amp;0xf} yield={(w&gt;&gt;109)&amp;1} &#8220;
              f&#8221;wr={(w&gt;&gt;110)&amp;7} rd={(w&gt;&gt;113)&amp;7} wait={(w&gt;&gt;116)&amp;0x3f:06b} &#8220;
              f&#8221;reuse={(w&gt;&gt;122)&amp;0xf:04b}&#8221;)
PY

# the fatbin container and the cost of coverage
nvcc -gencode arch=compute_90,code=sm_90 \
     -gencode arch=compute_100,code=[sm_100,compute_100] -fatbin -o f.fatbin saxpy.cu
for m in none speed balance size default; do
  nvcc -compress-mode=$m -gencode arch=compute_90,code=sm_90 \
       -gencode arch=compute_100,code=[sm_100,compute_100] \
       -fatbin -o f_$m.fatbin saxpy.cu; ls -l f_$m.fatbin; done

# the host side
nvcc -arch=sm_90 -c -o k.o saxpy.cu
readelf -SW k.o | grep -E &#8220;nv_fatbin|nvFatBinSegment|module_id&#8221;
readelf -x .nvFatBinSegment k.o
nvcc -arch=sm_90 --keep -c -o /dev/null saxpy.cu &amp;&amp; cat saxpy.cudafe1.stub.c</code></pre><p>The fatbin entry parser used for the library dissection is thirty lines: walk from the <code>0xBA55ED50</code> header, then for each entry read the kind at offset 0, header size at 4, padded payload size at 8, compressed size at 0x10, architecture at 0x1c, flags at 0x28, and skip <code>header_size + payload_size</code> to the next one.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2></h2>
      <p>
          <a href="https://www.thesoftwarefrontier.com/p/how-cuda-binaries-actually-work">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[How the NVIDIA Compiler Moat Actually Works]]></title><description><![CDATA[Inside the NVIDIA Compiler Moat: ptxas, SASS, and the 21 Bits Nobody Else Can Write]]></description><link>https://www.thesoftwarefrontier.com/p/how-the-nvidia-compiler-moat-actually</link><guid isPermaLink="false">https://www.thesoftwarefrontier.com/p/how-the-nvidia-compiler-moat-actually</guid><dc:creator><![CDATA[Lorenzo Bradanini]]></dc:creator><pubDate>Mon, 27 Jul 2026 11:08:36 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!D2wR!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdc74089c-cb20-4762-9ab1-4acb811e1150_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!D2wR!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdc74089c-cb20-4762-9ab1-4acb811e1150_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!D2wR!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdc74089c-cb20-4762-9ab1-4acb811e1150_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!D2wR!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdc74089c-cb20-4762-9ab1-4acb811e1150_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!D2wR!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdc74089c-cb20-4762-9ab1-4acb811e1150_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!D2wR!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdc74089c-cb20-4762-9ab1-4acb811e1150_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!D2wR!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdc74089c-cb20-4762-9ab1-4acb811e1150_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/dc74089c-cb20-4762-9ab1-4acb811e1150_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2763373,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/206014254?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdc74089c-cb20-4762-9ab1-4acb811e1150_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!D2wR!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdc74089c-cb20-4762-9ab1-4acb811e1150_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!D2wR!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdc74089c-cb20-4762-9ab1-4acb811e1150_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!D2wR!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdc74089c-cb20-4762-9ab1-4acb811e1150_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!D2wR!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdc74089c-cb20-4762-9ab1-4acb811e1150_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>Intro</h2><p>I spent, more or less, about a year writing an <strong>x86-64 assembler</strong> in C. Did Lexer, expression evaluator, instruction encoder, ELF writer, relocations, symbol resolution, eventually a macro preprocessor. The thing I keep coming back to from that year is how little information the assembler needed to know.</p><p>x86-64 encoding is somwthing baroque. <strong>ModRM</strong> and SIB are a whole mess, REX prefixes leak into everything, and the exact same mnemonic can have six encodings depending on operand width and register numbering. </p><p>But none of that is semantic. I emit the bytes and the machine figures out the rest in some ways. If I put two dependent instructions back to back, the processor detects the hazard, stalls, and my program suddently becomes slow. It is almost never wrong. The whole apparatus of scoreboarding, register renaming and<strong> out-of-order</strong> issue exists so that the assembler can behave like a dumb table lookup with a symbol table stapled on.</p><p>Then I started reading SASS, and that assumption stopped holding.</p><p>On modern NVIDIA hardware, if the compiler gets the scheduling metadata wrong, you do not get a slow kernel, you get directly a wrong answer. The <strong>dependency interlock</strong> for fixed-latency instructions is not in the hardware. </p><p>It is basically in a 21-bit field that <code>ptxas</code> writes into every instruction, and that field is undocumented, unstable across various architectures, and produced by exactly one program on earth.</p><p>Just that, and <strong>not the C++ dialect</strong>, and not <code>cudaMalloc</code>, is where I believe the compiler moat actually lives. What follows is a personal attempt to be precise about which layer is load-bearing and which layers people keep mistaking for it, because the public version of this argument is usually pitched at the wrong altitude.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Five moats wearing one name</h2><p>&#8220;<strong>CUDA</strong>&#8221; is, in various conversations, a giant bag holding at least five separate things, with wildly different levels of defensibility. We have:</p><p><strong>1. The source language.</strong> <code>__global__</code>, <code>threadIdx</code>, <code>&lt;&lt;&lt;grid, block&gt;&gt;&gt;</code>, a C++ dialect with a few extensions. This is the" &#8220;easiest&#8221; thing everyone points at and, by design, it is the least defensible layer. Clang has compiled CUDA for more than a decade. The NVPTX backend is just upstream LLVM. <code>libNVVM</code> is a documented public API and NVVM IR is a specified subset of LLVM IR. NVIDIA gave the frontend away years ago, precisely because giving away the frontend costs them nothing.</p><p><strong>2. The virtual ISA and the JIT distribution channel.</strong> PTX plus the compiler that ships inside <code>libcuda.so</code>. This is a genuine moat, one that I&#8217;ve personally studied in the past few months, but it is a <em>distribution</em> moat, not a performance one. It is the reason a fatbinary built in 2019 still runs on Blackwell.</p><p><strong>3. </strong><code>ptxas</code><strong> and SASS.</strong> This is interesting. We&#8217;re talking about a closed optimizing compiler that does register allocation, scheduling, and the encoding of hardware control state, pointing to an instruction set with no public specification and no official assembler. </p><p><strong>4. The libraries and the layout algebra.</strong> cuBLAS, cuDNN, CUTLASS, CuTe, NCCL, TensorRT-LLM. Thousands of pre-tuned kernel variants plus a template metaprogramming layer that encodes the tensor core&#8217;s operand layout requirements. It&#8217;s partly open, but the real thing is that it is entirely co-designed with silicon that is not public.</p><p><strong>5. The numerics contract and the install base.</strong> Now we&#8217;re serious. Reduction orders, TF32 defaults, FP8 and NVFP4 scaling recipes, atomics behaviour, the accumulated set of hyperparameter recipes that were tuned on top of all of it. I don&#8217;t know why yobody writes about this one, because it&#8217;s worth more than you can think.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!CrYh!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc9c0c1c8-c242-4531-b00c-dd379e1c5245_1634x912.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!CrYh!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc9c0c1c8-c242-4531-b00c-dd379e1c5245_1634x912.png 424w, https://substackcdn.com/image/fetch/$s_!CrYh!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc9c0c1c8-c242-4531-b00c-dd379e1c5245_1634x912.png 848w, https://substackcdn.com/image/fetch/$s_!CrYh!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc9c0c1c8-c242-4531-b00c-dd379e1c5245_1634x912.png 1272w, https://substackcdn.com/image/fetch/$s_!CrYh!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc9c0c1c8-c242-4531-b00c-dd379e1c5245_1634x912.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!CrYh!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc9c0c1c8-c242-4531-b00c-dd379e1c5245_1634x912.png" width="1456" height="813" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c9c0c1c8-c242-4531-b00c-dd379e1c5245_1634x912.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:813,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Figure 3. The five things people mean when they say CUDA, scored on whether the specificat&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 3. The five things people mean when they say CUDA, scored on whether the specificat" title="Figure 3. The five things people mean when they say CUDA, scored on whether the specificat" srcset="https://substackcdn.com/image/fetch/$s_!CrYh!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc9c0c1c8-c242-4531-b00c-dd379e1c5245_1634x912.png 424w, https://substackcdn.com/image/fetch/$s_!CrYh!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc9c0c1c8-c242-4531-b00c-dd379e1c5245_1634x912.png 848w, https://substackcdn.com/image/fetch/$s_!CrYh!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc9c0c1c8-c242-4531-b00c-dd379e1c5245_1634x912.png 1272w, https://substackcdn.com/image/fetch/$s_!CrYh!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc9c0c1c8-c242-4531-b00c-dd379e1c5245_1634x912.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong><span>Figure 1 </span></strong>The five things people mean when they say CUDA, scored on whether the specification is public, whether the source is available, whether anyone has independently reimplemented it, and whether it survives a new architecture unchanged.<span>Copy imageDownload PNG</span></figcaption></figure></div><p>Layer 1 is defintely gone. Layer 2 is a legal and logistical lock, not a technical one. Layers 3, 4 and 5 are &#8220;<em>the moat</em>&#8221;, and they defend against completely different attacks. </p><p>The vast majority &#8220;<em>CUDA alternative</em>&#8221; projects pick a layer, do good work, and then basically die on a different layer they did not budget for.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/p/how-the-nvidia-compiler-moat-actually?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/p/how-the-nvidia-compiler-moat-actually?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><div><hr></div><h2>The frontend left the building a long time ago</h2><p><code>nvcc</code> is not a compiler. It is a driver script, and you can watch it work:</p><pre><code><code>nvcc --dryrun -arch=sm_90 kernel.cu 2&gt;&amp;1 | grep -v '^#\$ *[A-Z_]*=' 
</code></code></pre><p>What you get back is a pipeline: <code>cudafe++</code> splits host from device, <code>cicc</code> (<em>the NVVM-based device compiler</em>) turns device code into PTX, the <code>ptxas</code> compiler turns PTX into a <code>cubin</code>, <code>fatbinary</code> staples the cubins and the PTX together, <code>nvlink</code> handles<strong> device linking</strong>, and your host compiler of choice gently handles all the rest. Six things doing 6 separate tasks. </p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Oo-2!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbcceed06-56f0-4264-a99a-71142732f51e_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Oo-2!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbcceed06-56f0-4264-a99a-71142732f51e_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!Oo-2!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbcceed06-56f0-4264-a99a-71142732f51e_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!Oo-2!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbcceed06-56f0-4264-a99a-71142732f51e_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!Oo-2!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbcceed06-56f0-4264-a99a-71142732f51e_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Oo-2!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbcceed06-56f0-4264-a99a-71142732f51e_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/bcceed06-56f0-4264-a99a-71142732f51e_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Figure 1. Five independent frontends converge on one virtual ISA and one closed backend. T&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 1. Five independent frontends converge on one virtual ISA and one closed backend. T" title="Figure 1. Five independent frontends converge on one virtual ISA and one closed backend. T" srcset="https://substackcdn.com/image/fetch/$s_!Oo-2!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbcceed06-56f0-4264-a99a-71142732f51e_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!Oo-2!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbcceed06-56f0-4264-a99a-71142732f51e_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!Oo-2!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbcceed06-56f0-4264-a99a-71142732f51e_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!Oo-2!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbcceed06-56f0-4264-a99a-71142732f51e_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong><span>Figure 2 </span></strong>Five independent frontends converge on one virtual ISA and one closed backend. The driver carries a second copy of that backend, which is why even a fully open toolchain has an NVIDIA compiler in its runtime path.<span>Copy imageDownload PNG</span></figcaption></figure></div><p>One of those, <code>cicc,</code> is an LLVM. NVIDIA has been explicit about this for the last few years: <strong>NVVM IR </strong>is just LLVM IR with a documented set of intrinsics and address space conventions, and <code>libNVVM</code> is a shipped, documented library you can use. </p><p>If you don&#8217;t want to use it, Clang&#8217;s own CUDA support will be happy to emit PTX for you, and it has been extreme quality, at least since Google put it there to build <strong>TensorFlow</strong>.</p><p>So the frontend nowadays is a commodity. That is why Triton, Mojo, tinygrad, Julia etc&#8230; can emit PTX, all without hurting NVIDIA&#8217;s position. Every one of those projects walks up to the same closed door at the end of the hallway.</p><p>I just want to be careful here, because &#8220;<em>the frontend is commoditized</em>&#8221; is often stated as if it were an indictment of NVIDIA&#8217;s openness, but it&#8217;s not. Ceding the frontend was correct and (probably) a deliberate choice. </p><p>Various frontends just increase the <strong>number of paths into PTX</strong>, and every path into PTX eventually terminates in a program NVIDIA controls.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>The 21-bit field</h2><p>Here is the part that changed how I think about this.</p><p><strong>Kepler</strong> version removed hardware dependency checking for fixed-latency instructions and moved the responsibility into the compiler. Then we had <strong>Maxwell </strong>refining it. </p><p><strong>Volta</strong> made it permanent by widening the instruction word from 64 bits to 128 bits and embedding the control payload directly in every instruction rather than in separate scheduling words interleaved every few instructions.</p><p>The reconstructed layout of that payload, for Volta and later, looks more or less like this:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!WZA9!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8107d78d-1c19-4c0c-ba5e-d0c60ab675d6_1634x836.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!WZA9!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8107d78d-1c19-4c0c-ba5e-d0c60ab675d6_1634x836.png 424w, https://substackcdn.com/image/fetch/$s_!WZA9!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8107d78d-1c19-4c0c-ba5e-d0c60ab675d6_1634x836.png 848w, https://substackcdn.com/image/fetch/$s_!WZA9!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8107d78d-1c19-4c0c-ba5e-d0c60ab675d6_1634x836.png 1272w, https://substackcdn.com/image/fetch/$s_!WZA9!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8107d78d-1c19-4c0c-ba5e-d0c60ab675d6_1634x836.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!WZA9!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8107d78d-1c19-4c0c-ba5e-d0c60ab675d6_1634x836.png" width="1456" height="745" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/8107d78d-1c19-4c0c-ba5e-d0c60ab675d6_1634x836.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:745,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Figure 2. The scheduling payload ptxas writes into every instruction. The stall count is t&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 2. The scheduling payload ptxas writes into every instruction. The stall count is t" title="Figure 2. The scheduling payload ptxas writes into every instruction. The stall count is t" srcset="https://substackcdn.com/image/fetch/$s_!WZA9!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8107d78d-1c19-4c0c-ba5e-d0c60ab675d6_1634x836.png 424w, https://substackcdn.com/image/fetch/$s_!WZA9!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8107d78d-1c19-4c0c-ba5e-d0c60ab675d6_1634x836.png 848w, https://substackcdn.com/image/fetch/$s_!WZA9!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8107d78d-1c19-4c0c-ba5e-d0c60ab675d6_1634x836.png 1272w, https://substackcdn.com/image/fetch/$s_!WZA9!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8107d78d-1c19-4c0c-ba5e-d0c60ab675d6_1634x836.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong><span>Figure 3 </span></strong>The scheduling payload ptxas writes into every instruction. The stall count is the field that turns a compiler into a component of the machine: for fixed latency instructions there is no hardware interlock behind it.<span>Copy imageDownload PNG</span></figcaption></figure></div><p>First. four bits of operand reuse cache hints. Then, a six-bit mask of <strong>scoreboard barriers </strong>this instruction waits on. Now on, two three-bit indices naming the barriers it sets for variable-latency results. A yield bit that tells the warp scheduler whether to prefer this warp or switch away. </p><p>And finally we have four bits of static stall count: the number of cycles the scheduler must wait before issuing the next instruction from this warp.</p><p>That last field is the one that matters for real. For fixed-latency instructions, the hardware does not check read-after-write hazards at all. The microarchitecture research on this is mostly unambiguous: if you set the stall counter wrong and the program produces incorrect results, because there is <strong>no scoreboard behind </strong>you to catch it. </p><p>The register status table and the wires from the issue logic to it were purposefully deleted from the design, and the area and energy they would have cost were spent on something else. The compiler is the scoreboard.</p><p>On x86, the ISA is a contract about <em>meaning</em> and the microarchitecture is free to reorder underneath it as long as the contract holds. On NVIDIA hardware since Kepler, a <strong>meaningful piece </strong>of the microarchitecture&#8217;s correctness lives in bits the compiler writes. </p><p><code>ptxas</code> is not a tool that targets the machine, because it is a component of the machine that happens to run on your workstation. Those are the obvious consequences. </p><p>You cannot open source that boundary without publishing the timing characteristics of every functional unit in <strong>every SM variant</strong> of every SKU, including the ones you have not announced. And if you publish it, you have frozen it, because now third-party code depends on it and any change basically breaks correctness, not just performance. </p><p>The closedness is not primarily a business decision. It is downstream of an architectural decision that was made for <strong>area and power reasons </strong>around 2012, and the strategic value came along for the ride.</p><p>I find that way interesting than the general story, even thoigh it&#8217;s more depressing, because it means the moat is not a policy anyone can be lobbied to reverse.</p><h3>Occupancy is a compiler policy, not a hardware property</h3><p>The other thing <code>ptxas</code> decides on your behalf, and the one you can actually measure without a disassembler, is <strong>how many registers</strong> each thread gets.</p><p>We alredy know that an H100 SM has a <strong>256 KB register file</strong>: 65,536 registers of 32 bits, allocated to warps at a granularity of eight registers per thread. The SM can hold at most 64 resident warps. Those three numbers set the whole occupancy calculation, and the only free variable is the register count <code>ptxas</code> chose.</p><p>Take a kernel where <code>ptxas -v</code> reports 168 registers per thread. Each warp consumes <code>168 x 32 = 5,376</code> registers, so the SM can hold <code>floor(65536 / 5376) = 12</code> warps, which is 12 of 64, or 18.75 percent occupancy. Now suppose you pass <code>-maxrregcount=128</code>. </p><p>The same warp now costs 4,096 registers, 16 warps fit, occupancy goes to roughly 25 percent, and <code>ptxas</code> pays for the difference by <strong>pushing the excess to local memory</strong>, which is to say to L1 and then to L2 and then, if you are unlucky and you have a giant working set, to HBM.</p><p>Neither number is right on its own. It depends on what the kernel spends its time waiting for. A kernel that <strong>waits on memory</strong> wants more warps. When one warp stalls on a load, the scheduler switches to another one, and if enough warps are resident, the waiting never shows up in the runtime. Occupancy is what keeps that kernel fed.</p><p>On the opposite side, a kernel that lives in the tensor cores usually wants the registers. It already hides its own waiting, by fetching the next tile while the current one is being multiplied. <strong>It has a pipeline</strong>, so it has nothing to gain from having other warps to switch to, and it will run happily at <em>eighteen percent occupancy.</em></p><p>That tradeoffs are the single most consequential decision in GPU performance work, and they are made by a <strong>heuristic inside a binary </strong>that I believe is impossible to read, tune, or replace. <code>-maxrregcount</code> and <code>__launch_bounds__</code> are the two knobs you get, and they are blunt.</p><p>I want to highlight this because it is the everyday version of the argument. You don&#8217;t need to care about control bits to be affected by the closed backend. If you have ever watched a <strong>two-line source change move register count</strong> from 168 to 176 and cost you a resident warp, you have already been on the wrong side of this wall.</p><p>There is no official SASS assembler. Instead, there&#8217;s a disassembler, <code>nvdisasm</code>, and <code>cuobjdump --dump-sass</code>, and the documentation for the instruction set amounts to a table of mnemonics with one-line descriptions. The opcode encodings, the<strong> control field layout</strong>, and the exact latency table are reverse engineered by a small group of people, generation by generation.</p><p>The lineage of that work is short enough to list: <code>asfermi</code> for Fermi, Scott Gray&#8217;s <code>maxas</code> for Maxwell, <code>KeplerAs</code>, <code>TuringAs</code>, <code>CuAssembler</code> for Pascal all the way through Ampere, <code>openptxas</code> under the <strong>gpuocelot umbrella</strong> for Maxwell and Pascal. Every one of these is a heroic, partial, architecture-locked effort that goes stale the moment a new chip make it to the public.</p><p>The payoff of that is real, which is the frustrating part in my opinion. Gray&#8217;s <code>maxas</code> SGEMM reached figures close to theoretical peak on GM204, ahead of the cuBLAS of that era, purely by scheduling better than <code>ptxas</code> did. </p><p>More recently there is paper work applying reinforcement learning to SASS schedules directly (<code>CuAsmRL</code>), reordering instructions inside <code>cubin</code>s that <code>ptxas</code> already produced and finding wins. If <code>ptxas</code> were optimal, that entire research direction would return zeros.</p><p>I have personally enjoyed one single detail, that I&#8217;m gonna show you below:</p><pre><code><code>/*0080*/ LDG.E.128 R8,  desc[UR6][R2.64]    ;  /* [----:B----:R-:W2:-:S01] */
/*0090*/ LDG.E.128 R12, desc[UR6][R4.64]    ;  /* [----:B----:R-:W3:-:S01] */
/*00a0*/ IMAD.WIDE R2,  R0, R7, c[0x0][0x168] ;/* [----:B----:R-:-:-:S02] */
/*00b0*/ HMMA.16816.F32 R20, R8, R12, R20   ;  /* [--23:B----:R-:-:-:S04] */
                          ^          ^
                          |          waits on write barriers 2 and 3
                          |          set by the two loads above
                          stall 4 cycles before the next issue
</code></code></pre><p><em>This is a reconstructed fragment in the shape that </em><code>cuobjdump --dump-sass</code><em> prints, not literal output from a specific compilation, but the bracketed control payload follows the documented conventions. </em></p><p><br>Please <strong>take a deep look</strong> at the two loads first. Each one sets a write barrier, 2 and 3 respectively, because a global load has variable latency and the compiler cannot know statically when the data lands. </p><p>Then the <code>IMAD</code> in between carries no barrier at all <em>(only a stall count) </em>because integer multiply-add is fixed latency and the compiler knows exactly how long it takes. The tensor core instruction waits on both barriers before it issues.</p><p>Now notice what is doing the work. The <code>S02</code> on the <code>IMAD</code> is not an optimization hint. It is the compiler asserting the <strong>pipeline depth of the integer unit</strong> on this specific SM. Change the SM, keep the number, and the machine reads a register before the previous instruction has written it. </p><p>One last thing: <strong>NVIDIA Research&#8217;s </strong>own SASS instrumentation framework, SASSI, was distributed as a <em>closed-source fork of </em><code>ptxas</code><em> binaries</em>. Their own researchers, inside the company, shipped a patched binary rather than source. That suggests you something about how the artifact is treated internally.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share The Software Frontier&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share The Software Frontier</span></a></p><div><hr></div><h2>PTX is a treaty, not an ISA</h2><p>I often read posts about people describing PTX as &#8220;<em>NVIDIA&#8217;s assembly language,</em>&#8221; which is so wrong in a way that hides the actual mechanism. <strong>PTX is a virtual ISA</strong> with unbounded virtual registers, no notion of the physical register file, no scheduling, and no control bits. </p><p>It is much closer to LLVM IR than to x86 assembly. Writing PTX by hand does not get you anywhere near the metal; it gets you a <em>slightly lower-level input</em> to the same closed optimizer.</p><p><strong>DeepSeek&#8217;s V3</strong> work is the canonical example, and it is routinely misreported. They partitioned twenty of the H800&#8217;s 132 SMs to handle inter-node communication, used warp specialization, and wrote custom PTX to reduce L2 pressure and interference with the compute kernels. </p><p>That is obv. excellent engineering under export-control constraints, but is not an escape from CUDA. It is a <em>deeper</em> commitment to CUDA: PTX pins you to NVIDIA <strong>harder than CUDA C++</strong> does, simply because CUDA C++ at least has clean-room reimplementations and PTX has one consumer.</p><p>What PTX actually buys NVIDIA is the forward compatibility story, and it is a genuinely excellent piece of platform engineering.<strong> Ship PTX inside your fatbinary</strong>, and when your binary lands on hardware it has never seen, the driver compiles it at load time. <code>libcuda.so</code> contains a compiler. Every NVIDIA GPU in the world is shipped with a copy of the last-mile toolchain baked into the driver.</p><p>The <strong>strategic consequences</strong> of that are larger than the technical ones:</p><ul><li><p><em>Any stack that emits PTX, no matter how open, has a closed NVIDIA compiler in its runtime path. Triton is open source and ships PTX to </em><code>ptxas</code><em>. tinygrad in </em><code>PTX=1</code><em> mode does the same. Open at the top, closed at the bottom.</em></p></li><li><p><em>Compatibility is a one-way ratchet. Old PTX runs on new hardware. New PTX never runs on old hardware. Every generation, the set of instructions you need to hit peak throughput moves into a </em><code>.target sm_XXa</code><em> gate that the previous generation&#8217;s driver cannot parse.</em></p></li><li><p><em>The world&#8217;s binaries contain PTX. That corpus is a compatibility obligation NVIDIA has taken on, and it is also an asset nobody else can serve.</em></p></li></ul><h3>The compatibility promise does not cover the instructions that matter</h3><p>This is the part of the PTX story I have not seen written down anywhere, but I do believe it changes the conclusion.</p><p>PTX <strong>forward compatibility</strong> applies to base architecture targets. <code>sm_90</code>, <code>sm_100</code>, and so on. But the instructions that reach peak throughput do not live there. <code>wgmma</code> requires <code>sm_90a</code>. The <code>tcgen05</code> family requires <code>sm_100a</code> or <code>sm_103a</code>. </p><p>That &#8220;<code>a</code> suffix&#8221; means architecture-specific, and code compiled for an <code>a</code> target is explicitly <em>not</em> forward compatible: it <strong>runs on that architecture</strong> and nowhere else, ever. NVIDIA later added <code>f</code> family targets to soften this within a family, which is an admission that the problem is real.</p><p>Put the two facts next to each other. The forward compatibility that everyone cites as CUDA&#8217;s great platform virtue covers the instruction set you use when you <strong>do not care about performance</strong>. The moment you write a kernel that actually saturates the tensor cores, you have opted out of it, and you are shipping architecture-locked binaries exactly like everyone else.</p><p>So the<strong> JIT compatibility story</strong> is not a technical moat at all; it is a convenience for the long tail, and the long tail is not where the money is. What it does buy NVIDIA is that the world&#8217;s software ships PTX, and PTX has one consumer.</p><p>And then there is the legal layer, which people mention and then move past, maybe too quickly. The <strong>CUDA EULA</strong> has, since 2021 online and since the 11.6 installed files, carried a restriction on reverse engineering, decompiling or disassembling the output generated using SDK elements for the purpose of translating those artifacts to target a non-NVIDIA platform. </p><p>Read it as written: it is a restriction on what you may do with <em>the output of the compiler</em>, not just the compiler. Whatever its enforceability, and I am not a lawyer, its practical effect is a chilling one on exactly the class of project that would otherwise attack layer 2.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>The libraries are an admission, not a flex</h2><p>The standard framing is that <strong>cuBLAS</strong> and <strong>cuDNN</strong> are evidence of NVIDIA&#8217;s software superiority. I want you to read them the opposite way, like I do.</p><p>If <code>nvcc</code> and <code>ptxas</code> could reliably generate peak code from readable source, cuBLAS would be a thin wrapper over a few templated kernels. Instead it is a multi-hundred-megabyte binary containing<strong> thousands of pre-compiled kernel variants</strong>, selected at runtime by heuristics, many of them tuned by hand at the SASS level by people who work at NVIDIA. </p><p>cuDNN is even worse. The size of those <code>.so</code> files is a direct measurement of the gap between what the compiler can do and what the hardware can do.</p><p>That <strong>gap is the real product. </strong>NVIDIA is not selling you a compiler that makes your code fast. It is selling you the output of a large, expensive, permanently-employed team of kernel engineers, wrapped in an API, in a form you cannot fork.</p><h3>And the gap is widening, on purpose</h3><p>The<strong> tensor core</strong> programming model has changed shape three times in six years, and each change moved way further away from anything a general compiler can target from ordinary source.</p><p>Ampere gave us <code>mma</code>, warp-level, operands in registers, with layout requirements you could hold in your head if you tried. Hopper introduced <code>wgmma</code>, warpgroup-scoped and asynchronous, plus TMA for <strong>bulk asynchronous copies</strong> driven by descriptors, plus <code>mbarrier</code> for the synchronization that async now requires. </p><p>Blackwell deprecated <code>wgmma</code> entirely and introduced the <code>tcgen05</code> family: a dedicated <strong>on-chip Tensor Memory</strong> for accumulators, a <em>single thread</em> issuing the MMA on behalf of everyone, and <code>cta_group::2</code> letting two CTAs on a TPC cooperate on one MMA with shared operands.</p><p>Read that again as a compiler person. The CuTe MMA atoms for Blackwell use a thread ID layout of one, where Hopper used 128. At peak performance, the <strong>SIMT model is a fiction.</strong> </p><p>What is actually happening is that one elected lane fires a descriptor at an accelerator, and the rest of the code is a software pipeline built out of barriers, async copies and a memory space whose allocation you manage explicitly.</p><p>This is why the ceiling moved from the compiler to the <strong>layout algebra. </strong>You will see some numbers, because the abstraction hides how physical this is. </p><blockquote><p><em>Tensor Memory on Blackwell is 256 KB per SM, organised as 128 lanes of 512 columns of 32 bits. You allocate it explicitly, in columns, in power-of-two units with a minimum of 32. </em></p><p><em>A single warp can only reach the 32 lanes matching its position in the warpgroup, which is why the copy instruction has multicast modes to duplicate a tile across all four lane quadrants. </em></p><p><em>Block-scaled narrow formats add their own geometry on top: MXFP8 carries one E8M0 scale per block of 32 elements, NVFP4 carries an E4M3 scale per block of 16 plus a second-level FP32 tensor scale, and the scale factors themselves have to be staged into tensor memory in a specific layout before the MMA will consume them.</em></p></blockquote><p>None of that is expressible as &#8220;<em>the compiler will figure it out</em>&#8221;. It is a data layout problem with <strong>hardware-imposed</strong> alignment constraints that propagate back into your choice of tile shape, which propagates back into your choice of pipeline depth, which propagates back into your register budget.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!22VA!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9626b009-bde1-49eb-8d33-56bab7959aec_1634x912.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!22VA!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9626b009-bde1-49eb-8d33-56bab7959aec_1634x912.png 424w, https://substackcdn.com/image/fetch/$s_!22VA!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9626b009-bde1-49eb-8d33-56bab7959aec_1634x912.png 848w, https://substackcdn.com/image/fetch/$s_!22VA!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9626b009-bde1-49eb-8d33-56bab7959aec_1634x912.png 1272w, https://substackcdn.com/image/fetch/$s_!22VA!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9626b009-bde1-49eb-8d33-56bab7959aec_1634x912.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!22VA!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9626b009-bde1-49eb-8d33-56bab7959aec_1634x912.png" width="1456" height="813" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/9626b009-bde1-49eb-8d33-56bab7959aec_1634x912.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:813,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Figure 4. Ampere issued an MMA from a warp, Hopper from a warpgroup, Blackwell from a sing&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 4. Ampere issued an MMA from a warp, Hopper from a warpgroup, Blackwell from a sing" title="Figure 4. Ampere issued an MMA from a warp, Hopper from a warpgroup, Blackwell from a sing" srcset="https://substackcdn.com/image/fetch/$s_!22VA!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9626b009-bde1-49eb-8d33-56bab7959aec_1634x912.png 424w, https://substackcdn.com/image/fetch/$s_!22VA!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9626b009-bde1-49eb-8d33-56bab7959aec_1634x912.png 848w, https://substackcdn.com/image/fetch/$s_!22VA!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9626b009-bde1-49eb-8d33-56bab7959aec_1634x912.png 1272w, https://substackcdn.com/image/fetch/$s_!22VA!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9626b009-bde1-49eb-8d33-56bab7959aec_1634x912.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong><span>Figure 4 </span></strong>Ampere issued an MMA from a warp, Hopper from a warpgroup, Blackwell from a single elected lane. Meanwhile the operands and the accumulator migrated out of the register file entirely.<span>Copy imageDownload PNG</span></figcaption></figure></div><p>This is why the ceiling moved from the compiler to the layout algebra. You cannot express a <strong>competitive Blackwell GEMM</strong> by writing loops and hoping. </p><p>You express it as a set of layouts, and the algebra of those layouts is <strong>CuTe</strong>, and CuTe is co-designed with the silicon by the people who designed the silicon, eighteen months before you can buy it.</p><p>The single most clarifying data point I have seen on this: FlashAttention 4, the first version written in <strong>CuTe DSL</strong> rather than C++ templates, reportedly hits about 1613 TFLOP/s on B200 for bf16 head-dim 128 causal attention at 8K sequence length, roughly 71 percent of peak, about 1.3 times faster than cuDNN 9.13, and around <em>2.7 times faster</em> than Triton on the same hardware.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!vR4q!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2987cdf1-5bb5-4e21-ade3-583932f9909b_1634x836.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!vR4q!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2987cdf1-5bb5-4e21-ade3-583932f9909b_1634x836.png 424w, https://substackcdn.com/image/fetch/$s_!vR4q!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2987cdf1-5bb5-4e21-ade3-583932f9909b_1634x836.png 848w, https://substackcdn.com/image/fetch/$s_!vR4q!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2987cdf1-5bb5-4e21-ade3-583932f9909b_1634x836.png 1272w, https://substackcdn.com/image/fetch/$s_!vR4q!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2987cdf1-5bb5-4e21-ade3-583932f9909b_1634x836.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!vR4q!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2987cdf1-5bb5-4e21-ade3-583932f9909b_1634x836.png" width="1456" height="745" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2987cdf1-5bb5-4e21-ade3-583932f9909b_1634x836.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:745,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Figure 5. One benchmark, three conclusions: the closed library is not the ceiling, the til&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 5. One benchmark, three conclusions: the closed library is not the ceiling, the til" title="Figure 5. One benchmark, three conclusions: the closed library is not the ceiling, the til" srcset="https://substackcdn.com/image/fetch/$s_!vR4q!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2987cdf1-5bb5-4e21-ade3-583932f9909b_1634x836.png 424w, https://substackcdn.com/image/fetch/$s_!vR4q!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2987cdf1-5bb5-4e21-ade3-583932f9909b_1634x836.png 848w, https://substackcdn.com/image/fetch/$s_!vR4q!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2987cdf1-5bb5-4e21-ade3-583932f9909b_1634x836.png 1272w, https://substackcdn.com/image/fetch/$s_!vR4q!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2987cdf1-5bb5-4e21-ade3-583932f9909b_1634x836.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong><span>Figure 5 </span></strong>One benchmark, three conclusions: the closed library is not the ceiling, the tile level DSL is the productive authoring layer, and the portable tile language is a long way off peak on the newest tensor cores.<span>Copy imageDownload PNG</span></figcaption></figure></div><p>Three conclusions fall out of one benchmark. NVIDIA&#8217;s own closed, hand-tuned attention library was not the ceiling. </p><p>The<strong> tile-level DSL</strong>, not the C++ template library, is now the productive authoring layer. And Triton, the great portable hope, is nowhere near peak on the newest tensor cores, because its abstractions do not yet expose what <code>tcgen05</code> and TMA require.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>What is actually in a cubin</h2><p>I want to discuss and be clear about one thing worth opening up, because it is the part that felt most familiar to me coming from ELF.</p><p>A <code>cubin</code> is an <strong>ELF64 object.</strong> S ame header layout you already know, with a CUDA-specific machine type. Inside, you get a <code>.text</code> section per kernel holding the SASS, a <code>.nv.constant0</code> section per kernel holding the kernel parameters and grid constants (<em>this is why parameters arrive via </em><code>c[0x0][...]</code><em> in the disassembly rather than in registers</em>), <code>.nv.info</code> sections carrying metadata like register counts and parameter layout, plus the usual relocation sections. </p><p>The fatbinary that <code>nvcc</code> embeds in your host object is itself a container of these, in a <code>.nv_fatbin</code> section, with a small descriptor segment telling the runtime what is inside.</p><p>You can walk all of it here:</p><pre><code><code>cuobjdump -elf a.out            # fatbin contents, section by section
nvdisasm -elf kernel.cubin      # ELF structure plus disassembly
readelf -S kernel.cubin         # it really is just ELF
</code></code></pre><p>The greatest absent, once you have looked, is the same one as before. <strong>Every part of this container is inspectable</strong>: the relocations and the metadata are readable, the instruction stream is disassemblable. What you can&#8217;t do is produce a valid one from modified SASS, because the only program that writes the <code>.text</code> section is the one you do not have.</p><p>Coming from x86, that asymmetry is the strangest part of the whole stack. I can hand-assemble an ELF object, link it, and run it. Here I can read everything and write nothing.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>What is not being told</h2><p>Most of the commentary stops at &#8220;<em>CUDA is sticky.</em>&#8221; Here are the parts I think are underweighted, roughly in order of how much they change the conclusion.</p><h3>The compiler is correctness-critical, not just performance-critical</h3><p>Again, sorry for that, but I have not seen this stated plainly anywhere outside the microarchitecture literature. Every discussion of open SASS toolchains treats the problem as &#8220;<em>we would generate slower code</em>&#8221;. </p><p>The actual problem is that a <strong>third-party SASS assembler</strong> that mis-models the latency of one functional unit on one SKU generates <em>silently incorrect</em> results, in a data-dependent way, on a subset of hardware. </p><p>That is a categorically harder engineering problem than performance parity, and it is the reason no serious commercial effort has attempted it. The risk profile is closer to writing a memory controller than to writing a backend.</p><h3>The moat is a labor market</h3><p>Strip away the branding and ask what actually<strong> cannot be replicated</strong> and you&#8217;ll find out that is not the compiler source. </p><p>It is the few hundred people worldwide who can look at a <code>cuobjdump</code> listing, reason about the control bits, and know from experience that a particular shared memory access pattern will hit a bank conflict on this generation but not the last. </p><p>A large fraction of them work at NVIDIA. Most of the rest work at four or five labs and two or three GPU startups. That reframing matters because it tells you what the actual counter-technology is. Many have attemped to address with it, but no, it&#8217;s not a new language. </p><p>It is just search: autotuners, superoptimizers, and increasingly <strong>LLM-driven kernel generation</strong>, all of which convert scarce expert labour into abundant compute. If the moat is a hiring problem, then the thing that erodes it is the thing that makes kernel expertise reproducible.</p><h3>There is a telemetry loop nobody else has</h3><p>Nsight, <strong>DCGM</strong>, the driver, customer escalations from every large training run on the planet. NVIDIA sees the performance pathologies of the entire industry&#8217;s kernels, continuously, at a scale no competitor approaches. </p><p>Then it fixes the top ones in the next library release and, where the fix has to be in hardware, in the next architecture. Every &#8220;<em>CUDA maturity</em>&#8221; argument is really an argument about the number of laps this loop has run.</p><h3>The collectives are compiled kernels, and they eat your SMs</h3><p>I listed the interconnect <strong>under layer 5 </strong>and then almost walked past it, which would have been a mistake, because NCCL is where the compiler moat and the network moat turn out to be exactly the same thing.</p><p>A collective is not a library call that hands work to a NIC, it&#8217;s a set of CUDA kernels. NCCL builds rings and trees over the <strong>discovered topology</strong>, splits each collective across some number of channels, and every channel is a resident thread block. </p><p>Which means bandwidth on the wire is bought with SMs that are then not available for compute. It also means the protocol matters: <strong>NCCL </strong>picks between a low latency mode that trades payload efficiency for skipping memory fences, a 128 byte variant that recovers most of the bandwidth but wants <strong>NVLink underneath</strong>, and a plain mode that gets full bandwidth at higher latency. </p><p>The choice is made by heuristics against measured topology, and the tuning knobs are environment variables. We see probably at least two consequences that are worth holding onto.</p><p>First, this is exactly what DeepSeek was doing when they carved<strong> twenty SMs</strong> out of 132 for communication. Their work was quite revolutionary because they were taking<strong> manual control</strong> of a tradeoff that NCCL normally makes for you, because on export-limited H800s with halved inter-GPU bandwidth, NCCL&#8217;s defaults were wrong for their shape. All of that, as stated before, working with CUDA, not around it or without it.</p><p>Second, <strong>in-network reduction</strong> changes the arithmetic entirely. When the switch itself can perform the reduction, the collective stops being a kernel that moves the same bytes several times and becomes a kernel that ships bytes once. </p><p>That capability lives in the NVSwitch and InfiniBand silicon, it is exposed through the same library, and there is no version of &#8220;<em>port your kernels</em>&#8221; that gets you it. AMD&#8217;s answer, <strong>RCCL</strong>, is a real port of the same design, and the vLLM tuning guidance for it involves pushing the channel count up sharply on <strong>MI300X</strong>, which tells you how much of the performance is in the topology-specific constants rather than the algorithm.</p><p>If you are keeping a scorecard of which moats are compiler-shaped, this one is a hybrid, and it is the least portable thing in the stack.</p><p>Port a model from <strong>cuBLAS </strong>to hipBLASLt and the arithmetic is not bit-identical for a few reasons. Split-k strategies differ, even reduction orders do not behave the same. </p><blockquote><p><em>TF32 is on by default in one place and not another. </em></p><p><em>FP8 scaling granularity, per-tensor versus per-block versus the MX and NVFP4 formats, is a semantic choice baked into the kernels.</em></p></blockquote><p>For inference this usually shows up as a small quality wobble that people can get along with. For training, it shows up as a recipe that diverges at step 40,000 for <strong>some random reasons</strong> nobody can bisect, on a piece of hardware where you have less tooling. </p><p>The cost of that risk almost never appears in the total-cost-of-ownership spreadsheets that compare dollars per FLOP, and it is a real reason large labs pay the NVIDIA premium.</p><h3>Watch which artifact they do not publish</h3><p>There is a pattern in NVIDIA&#8217;s open sourcing, and once you see it you can predict the next move. They publish the specification and the dialect, while keeping the lowering.</p><p><strong>NVVM IR</strong>: specified, with a public library. <code>cicc</code>: closed. PTX: exhaustively documented, hundreds of pages, updated every release. <code>ptxas</code>: closed. CUTLASS: open, permissively licensed, genuinely excellent. </p><p>The SASS it lowers to: closed. And now<strong> CUDA Tile IR</strong>: the MLIR dialect, the Python bindings, the bytecode format and the conformance suite are on GitHub. The compiler that turns tile bytecode into <code>cubin</code> is not.</p><p>This is a coherent, repeated strategy. Publishing the interface grows the number of producers. Keeping the lowering means every producer terminates in your binary. </p><p>It is the same move IBM made with channel architecture and the same move ARM makes with the ISA versus the implementations, and NVIDIA runs it more cleanly than either.</p><h3>AMD is the counterexample that breaks the simple story</h3><p>Here is the fact that should discipline every &#8220;<em>closed compiler is the moat</em>&#8221; argument: AMD publishes its ISA documents. The <strong>AMDGPU </strong>backend is upstream LLVM. ROCm is open source, top to bottom, including the compiler and the runtime. </p><p>You can read the machine encodings for CDNA3 and CDNA4 in a PDF on AMD&#8217;s website But (unexpectedly) even this has not collapsed the moat.</p><p>So openness is not the mechanism. If it were, AMD would have won already a few years ago. The mechanism is the <strong>co-design loop</strong>, the library labour, the numerics contract, the install base and the interconnect, and NVIDIA&#8217;s closed compiler is a symptom of the same architectural strategy that produces the rest of it. </p><p>Anyone whose plan is &#8220;<em>make an open CUDA</em>&#8221; has already been run as an experiment, by a company with $25 billion of revenue, for at least a decade.</p><h3>The moat is thinning at the top and thickening at the bottom</h3><p><strong>Two things are true at once</strong> and they point in opposite directions, which is again interesting because it sounds counterintuitive and understandable at the same time.</p><p>At the top, the default path <strong>from PyTorch to the GPU</strong> is now a portable tile language. <code>torch.compile</code> emits Triton, not CUDA C++. Liger, FlexAttention and a large fraction of production fused kernels are&#8230;. yes, again, Triton. That layer is vendor-neutral by construction, and it is where most new kernel code is being written.</p><p>At the bottom, the instructions you need to reach peak on the newest silicon are <strong>less reachable from a generic compiler </strong>than at any point since 2007. TMA descriptors, tensor memory allocation, CTA-pair MMA, block-scaled narrow formats with alignment constraints that constrain your tile shapes. </p><p>The gap between &#8220;<em>compiles and runs</em>&#8221; and &#8220;<em>hits peak</em>&#8221; is the widest it has ever been in 19 years.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!q8BO!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6bd9d02f-7d35-46fd-a5a8-f9d1c3fb3f46_1634x931.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!q8BO!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6bd9d02f-7d35-46fd-a5a8-f9d1c3fb3f46_1634x931.png 424w, https://substackcdn.com/image/fetch/$s_!q8BO!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6bd9d02f-7d35-46fd-a5a8-f9d1c3fb3f46_1634x931.png 848w, https://substackcdn.com/image/fetch/$s_!q8BO!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6bd9d02f-7d35-46fd-a5a8-f9d1c3fb3f46_1634x931.png 1272w, https://substackcdn.com/image/fetch/$s_!q8BO!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6bd9d02f-7d35-46fd-a5a8-f9d1c3fb3f46_1634x931.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!q8BO!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6bd9d02f-7d35-46fd-a5a8-f9d1c3fb3f46_1634x931.png" width="1456" height="830" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6bd9d02f-7d35-46fd-a5a8-f9d1c3fb3f46_1634x931.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:830,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Figure 9. The binding constraint has moved four times in nineteen years. It sat on the sou&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 9. The binding constraint has moved four times in nineteen years. It sat on the sou" title="Figure 9. The binding constraint has moved four times in nineteen years. It sat on the sou" srcset="https://substackcdn.com/image/fetch/$s_!q8BO!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6bd9d02f-7d35-46fd-a5a8-f9d1c3fb3f46_1634x931.png 424w, https://substackcdn.com/image/fetch/$s_!q8BO!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6bd9d02f-7d35-46fd-a5a8-f9d1c3fb3f46_1634x931.png 848w, https://substackcdn.com/image/fetch/$s_!q8BO!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6bd9d02f-7d35-46fd-a5a8-f9d1c3fb3f46_1634x931.png 1272w, https://substackcdn.com/image/fetch/$s_!q8BO!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6bd9d02f-7d35-46fd-a5a8-f9d1c3fb3f46_1634x931.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong><span>Figure 6 </span></strong>The binding constraint has moved four times in nineteen years. It sat on the source language only during the first five, and the two most recent moves went in opposite directions: down into the tensor core, and up into a new virtual ISA.<span>Copy imageDownload PNG</span></figcaption></figure></div><p>Now the interesting question is not whether the moat holds. It is more which layer the industry ends up standardizing on, and NVIDIA has just made a very large move to answer that question in its own favour.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Tile IR: the moat gets rebuilt one level up</h2><p>The most important compiler event of the last twelve months got almost no coverage outside compiler circles and a restricted group of GPU programmers.</p><p>CUDA 13.1 introduced <strong>CUDA Tile</strong>, a tile-based programming model with its own intermediate representation. Tile IR is an MLIR dialect. The dialect, the Python bindings, the bytecode serializer and a conformance test suite are open source at <code>NVIDIA/cuda-tile</code>. </p><p>The bytecode format is explicitly <strong>specified as stable</strong>, versioned, and forward and backward compatible: bytecode emitted by an older compiler is readable by a newer compiler or driver, and a driver accepts bytecode up to its supported version.</p><p>That paragraph should sound familiar, because it is the <strong>PTX contract</strong>, restated one abstraction level higher, in MLIR, with a published binary encoding. Tile IR is the new PTX.</p><p>The user-facing language is cuTile, a <strong>Python DSL</strong>, with a Julia binding already shipping at near parity on simple kernels. And critically, NVIDIA wrote a Triton backend that targets Tile IR, so Triton programs can be compiled through the new path instead of through PTX. </p><p>Meta has filed an RFC to add a cuTile backend to <strong>PyTorch Inductor</strong>, with the stated motivation that Tile IR gives bytecode portability, performance portability across GPU generations, and access to a tile-specific optimizing compiler, and that it is a prerequisite for pointing Helion at Tile IR as well.</p><p>Pro tip: read the sequence of moves rather than the individual announcements.</p><p>Triton became the default codegen target of PyTorch. That made the tile abstraction, the layer where new kernels get written, and that layer is portable across vendors. It was the most serious <strong>structural threat</strong> to the moat in fifteen years, because it made the frontend irrelevant <em>and</em> gave AMD and Intel a credible path to the same source.</p><p>NVIDIA&#8217;s response was not to fight Triton: it was to build a <strong>better tile IR,</strong> open source the dialect and the spec, write the Triton backend themselves, and offer everyone a <strong>faster path</strong> through their own infrastructure. If it works, the portable tile layer that was supposed to route around NVIDIA becomes a frontend for NVIDIA&#8217;s IR, and the optimizing compiler underneath, the one that turns tiles into <code>cubin</code>, is closed, just like <code>ptxas</code>.</p><p>It is the cleanest execution of the &#8220;<em>publish the interface, keep the lowering</em>&#8221; strategy I have seen so far, and it happened while everyone was arguing about whether ROCm had caught up.</p><p>I want to flag my uncertainty honestly: it is early. Tile IR shipped in December 2025, the <strong>Inductor RFC</strong> is in progress, and I have not seen independent third-party benchmarks of cuTile against hand-written CuTe on Blackwell at production shapes. </p><p>NVIDIA&#8217;s own claim is that cuTile is competitive with cuBLAS for large GEMMs and cuDNN for large attention. Competitive with cuDNN is a lower bar than it used to be, given FA4.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>The alternatives, honestly</h2><p>Every project below is real engineering by capable people. What I want to do is say which layer each one attacks and where it runs out of road, because most of them are described by their advocates as if they attack all five.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!vtGj!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbfdc9857-fe72-401d-8dd7-9f39f30d08d0_1634x988.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!vtGj!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbfdc9857-fe72-401d-8dd7-9f39f30d08d0_1634x988.png 424w, https://substackcdn.com/image/fetch/$s_!vtGj!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbfdc9857-fe72-401d-8dd7-9f39f30d08d0_1634x988.png 848w, https://substackcdn.com/image/fetch/$s_!vtGj!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbfdc9857-fe72-401d-8dd7-9f39f30d08d0_1634x988.png 1272w, https://substackcdn.com/image/fetch/$s_!vtGj!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbfdc9857-fe72-401d-8dd7-9f39f30d08d0_1634x988.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!vtGj!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbfdc9857-fe72-401d-8dd7-9f39f30d08d0_1634x988.png" width="1456" height="880" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/bfdc9857-fe72-401d-8dd7-9f39f30d08d0_1634x988.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:880,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Figure 7. Where each project actually pushes, and where it stops. Most of them are pitched&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 7. Where each project actually pushes, and where it stops. Most of them are pitched" title="Figure 7. Where each project actually pushes, and where it stops. Most of them are pitched" srcset="https://substackcdn.com/image/fetch/$s_!vtGj!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbfdc9857-fe72-401d-8dd7-9f39f30d08d0_1634x988.png 424w, https://substackcdn.com/image/fetch/$s_!vtGj!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbfdc9857-fe72-401d-8dd7-9f39f30d08d0_1634x988.png 848w, https://substackcdn.com/image/fetch/$s_!vtGj!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbfdc9857-fe72-401d-8dd7-9f39f30d08d0_1634x988.png 1272w, https://substackcdn.com/image/fetch/$s_!vtGj!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbfdc9857-fe72-401d-8dd7-9f39f30d08d0_1634x988.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong><span>Figure 7 </span></strong>Where each project actually pushes, and where it stops. Most of them are pitched as if they cover the whole stack.<span>Copy imageDownload PNG</span></figcaption></figure></div><h3>HIP and ROCm</h3><blockquote><p><strong>Attacks:</strong> layer 1 and layer 4. A CUDA-shaped API, a source translation tool, and a library set with matching names.</p></blockquote><p>ROCm in 2026 is a genuinely different product from ROCm in 2023, and people who formed their opinion during the Frontier era are working from stale data. </p><p>ROCm 7.2 ships tuned hipBLASLt GEMM kernels for FP8, BF16 and FP16 on MI350 and MI355X. AITER provides fused kernels that, inside vLLM, deliver large multiples over the legacy ROCm attention path. vLLM now carries seven attention backends on ROCm. </p><p>MI355X ships 288GB of HBM3E, and capacity per package is a real architectural advantage that removes tensor parallelism from model sizes where NVIDIA needs it.</p><blockquote><p><strong>Where it runs out of road:</strong> the API-clone strategy structurally guarantees you arrive second. <code>hipify</code> ignores inline PTX, which means exactly the kernels that were worth hand-optimizing are the ones that do not port. Matrix core layouts (<code>MFMA</code>) differ enough from <code>mma</code>, <code>wgmma</code> and <code>tcgen05</code> that a &#8220;<em>ported</em>&#8221; kernel is a rewrite at the layout level. And the performance story has historically been per-model tuning: a specific configuration of a specific model at a specific batch size is fast because someone at AMD made it fast, which is a different asset from a compiler that makes your unseen kernel fast.</p><p><strong>Honest read:</strong> on standard transformer inference, the gap is small enough that the decision is now economic rather than technical. On training, on novel architectures, and on anything requiring bespoke kernels, it is not close.</p></blockquote><h3>SCALE (Spectral Compute)</h3><blockquote><p><strong>Attacks:</strong> layer 1 and layer 3 simultaneously, which makes it the most technically interesting project in this space.</p></blockquote><p>SCALE is a clean-room, Clang and LLVM based drop-in replacement for <code>nvcc</code> that compiles nvcc-dialect CUDA source, including inline PTX, directly to <strong>AMD machine code</strong>. This is not translation, not even middleware: an actual second implementation of the CUDA language. </p><p>The company is small, London based, founded 2018, and funded its own compiler work off consulting. Published benchmarks claim roughly <strong>5.94 times</strong> over HIP on AMD hardware, and as of May 2026 they are claiming speedups over <code>nvcc</code> itself on B300 with CUDA 13.</p><p>The detail I keep turning over is that they joined NVIDIA Inception in June 2026. A company whose product is CUDA portability, partnering with NVIDIA. Which makes sense if you believe NVIDIA&#8217;s interest is in CUDA remaining the standard more than in CUDA remaining exclusive, but it is a strange shape in my view.</p><blockquote><p><strong>Where it runs out of road:</strong> the CUDA-X library surface. There are hundreds of libraries above the language, and a compiler that perfectly compiles the language still leaves you needing cuDNN, cuTENSOR, cuDF, NCCL and the rest. Spectral is working on delegating those to ROCm equivalents, which means SCALE inherits ROCm&#8217;s library ceiling at exactly the point where it matters most.</p></blockquote><h3>ZLUDA</h3><blockquote><p><strong>Attacks:</strong> layer 2. Binary and PTX level translation, no recompilation.</p></blockquote><p>ZLUDA is the project people cite most and understand least. It gained ROCm 7 support in December 2025 and shipped v6 in June 2026. </p><p>That release also announced that commercial funding had ended again and the project is back to being one person&#8217;s weekend work, which is why v6&#8217;s headline features are <strong>32-bit PhysX</strong> and Blender textures rather than anything relevant to inference.</p><blockquote><p><strong>Where it runs out of road:</strong> everywhere, and the reasons are structural rather than technical. Binary translation is a treadmill against a vendor who ships a new virtual ISA extension every generation and a whole new IR every few years. The EULA restriction points directly at this approach. It has now lost funding twice, from two different backers. I would treat the binary translation path as closed, not because the engineering is bad, but because the economics have been tested twice and failed twice.</p></blockquote><h3>Triton</h3><blockquote><p><strong>Attacks:</strong> layer 1 and, at least partially, layer 4.</p></blockquote><p>Triton is the most consequential thing on this list because of where it sits: it is what <code>torch.compile</code> emits, which means it is the default kernel authoring layer for most of the industry whether they know it or not. It runs on NVIDIA, AMD and Intel from the same source.</p><blockquote><p><strong>Where it runs out of road:</strong> two places. On NVIDIA it emits PTX, so it never escapes <code>ptxas</code>; it relocates the frontend, it does not cross the wall. And on the newest hardware it is a long way from peak, because the abstractions do not expose TMEM, CTA-pair MMA and the descriptor-driven async machinery that Blackwell needs. FlashAttention 4 leaving Triton for CuTe DSL is the load-bearing evidence there.</p></blockquote><p>And now NVIDIA is offering to compile it through Tile IR, which resolves the second problem in exchange for reintroducing the first at a higher altitude.</p><h3>Helion, TileLang, Pallas, and the search-based approach</h3><blockquote><p><strong>Attacks:</strong> the labour market, which as argued above is the real moat.</p></blockquote><p><strong>Helion </strong>is a Python DSL from Meta&#8217;s PyTorch compiler team that sits above Triton: you write PyTorch-shaped code with tile loops and it autotunes over hundreds of generated Triton implementations, spending roughly ten minutes of compute per kernel to do it. </p><p>The reported numbers are the interesting part. Around 1.05 times over <code>torch.compile</code> and 1.44 times over hand-written Triton on AMD on average, with individual kernels much higher, and on H100 roughly 1.2 to 1.85 times over Triton.</p><p><strong>Beating hand-written Triton </strong>on average, by searching. That is the anti-moat technology, and it is not a language feature. It is a compute substituting for expertise, and it gets better every time GPUs get cheaper, which is a nice recursion.</p><p>The same logic explains AMD&#8217;s agentic kernel work (<strong>GEAK</strong>), LLM kernel generators, and the growing body of RL-over-schedules research including the SASS-level work mentioned earlier. If your competitor&#8217;s advantage is several hundred irreplaceable engineers, the correct response is not to hire several hundred engineers. It is to make the engineers less necessary.</p><blockquote><p><strong>Where it runs out of road:</strong> search needs a correct, expressive, fast-to-evaluate space to search over, and on Blackwell the space that contains peak performance is only reachable through the vendor&#8217;s layout algebra. You can search brilliantly within Triton&#8217;s expressible set and still be 2.7 times off.</p></blockquote><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!xQrV!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2630a6f8-94c7-4c08-b0a0-6a3d8ca99a29_1634x912.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!xQrV!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2630a6f8-94c7-4c08-b0a0-6a3d8ca99a29_1634x912.png 424w, https://substackcdn.com/image/fetch/$s_!xQrV!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2630a6f8-94c7-4c08-b0a0-6a3d8ca99a29_1634x912.png 848w, https://substackcdn.com/image/fetch/$s_!xQrV!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2630a6f8-94c7-4c08-b0a0-6a3d8ca99a29_1634x912.png 1272w, https://substackcdn.com/image/fetch/$s_!xQrV!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2630a6f8-94c7-4c08-b0a0-6a3d8ca99a29_1634x912.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!xQrV!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2630a6f8-94c7-4c08-b0a0-6a3d8ca99a29_1634x912.png" width="1456" height="813" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2630a6f8-94c7-4c08-b0a0-6a3d8ca99a29_1634x912.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:813,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Figure 8. Reported multiples from vendors and projects, on different hardware, at differen&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 8. Reported multiples from vendors and projects, on different hardware, at differen" title="Figure 8. Reported multiples from vendors and projects, on different hardware, at differen" srcset="https://substackcdn.com/image/fetch/$s_!xQrV!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2630a6f8-94c7-4c08-b0a0-6a3d8ca99a29_1634x912.png 424w, https://substackcdn.com/image/fetch/$s_!xQrV!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2630a6f8-94c7-4c08-b0a0-6a3d8ca99a29_1634x912.png 848w, https://substackcdn.com/image/fetch/$s_!xQrV!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2630a6f8-94c7-4c08-b0a0-6a3d8ca99a29_1634x912.png 1272w, https://substackcdn.com/image/fetch/$s_!xQrV!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2630a6f8-94c7-4c08-b0a0-6a3d8ca99a29_1634x912.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong><span>Figure 8 </span></strong>Reported multiples from vendors and projects, on different hardware, at different shapes, against different baselines. Useful as a map of who is claiming what, not as a comparison.<span>Copy imageDownload PNG</span></figcaption></figure></div><h3>tinygrad</h3><blockquote><p><strong>Attacks:</strong> layers 1, 2 and 3, by refusing the premise.</p></blockquote><p>tinygrad generates its own kernels, can render directly to PTX, and has experimental backends that talk to AMD and NVIDIA hardware without the vendor userspace runtime at all. </p><p>Whatever you think of the framework&#8217;s ambitions, the AMD path is an existence proof that the vendor compiler and userspace stack are not physically required.</p><blockquote><p><strong>Where it runs out of road:</strong> peak tensor core utilization on the newest silicon, and the entire ecosystem surface. It is a demonstration of what is possible for a small team, not a production alternative for a lab training frontier models. Its real contribution is epistemic: it proves the stack is thinner than the marketing implies.</p></blockquote><h3>Mojo and MAX</h3><blockquote><p><strong>Attacks:</strong> layers 1 and 4, with an MLIR-native language designed for kernel authoring across vendors.</p></blockquote><p>Modular&#8217;s own writing on this is refreshingly unromantic: they note that write-once-run-everywhere in GPU land has usually meant write-once-run-slowly-everywhere, that CUTLASS does not attempt portability beyond NVIDIA and is often locked within a generation, and that Triton&#8217;s performance degrades off NVIDIA. </p><p>Their bet is that a properly designed <strong>compile-time metaprogramming language</strong> can express hardware-specific structure without forking the source.</p><blockquote><p><strong>Where it runs out of road:</strong> it is a proprietary language asking developers to adopt it on faith, competing against a Python DSL layer that PyTorch already emits by default. The technology is pretty good. The distribution problem is brutal.</p></blockquote><h3>The vertical integrators: TPU, Trainium, and friends</h3><blockquote><p><strong>Attacks:</strong> the entire column, by not participating in the compatibility game at all.</p></blockquote><p>Google&#8217;s TPU plus XLA is the only stack that has demonstrably trained frontier models at scale outside NVIDIA, sustained, for years. It did not do that by cloning CUDA. </p><p>It did it by owning the chip, the interconnect, the compiler, one framework, and one enormous first-party customer whose workloads shape the hardware roadmap. <strong>Trainium plus NKI </strong>is the same play at an earlier stage.</p><p>The lesson is uncomfortable for everyone selling portability: the proven counter to vertical integration is vertical integration. Compatibility layers have a zero-for-many record. Owning the column has a one-for-one record.</p><blockquote><p><strong>Where it runs out of road:</strong> you cannot buy this. It requires a decade, a captive workload, and a willingness to eat several generations of worse hardware.</p></blockquote><h3>The ones I am not covering properly, and why</h3><p>Three more categories exist and deserve at least an honest pointer rather than silence.</p><p><strong>Huawei&#8217;s CANN</strong> with the Ascend NPUs is the most complete non-Western vertical stack I&#8217;ve ever seen, with its own kernel language, its own graph compiler, and a captive domestic customer base that gives it the one thing compatibility layers never have, which is a workload that shapes the roadmap. </p><p>I do not have enough independent benchmark data to say anything useful about where it sits on peak utilization, and I would rather say that than guess randomly. The same caveat applies to<strong> Moore Threads</strong> and Biren, both of which ship CUDA-adjacent programming models.</p><p>On the standards side, chipStar compiles HIP and CUDA to SPIR-V for OpenCL and Level Zero targets, and Intel&#8217;s SYCLomatic migrates CUDA source to SYCL with a reported majority of code converting automatically and a manual remainder. Both are layer 1 plays with the layer 4 problem intact.</p><p>And there is a fourth category I have deliberately excluded from the comparison table because it is not a competitor at all: NVIDIA&#8217;s own Python surface. <code>cuda.core</code>, <strong>Numba</strong>, Warp, <code>nvmath</code>, and now cuTile are a coordinated effort to make sure that when kernel authoring moves fully into Python, which it is doing, it moves into NVIDIA&#8217;s Python rather than someone else&#8217;s.</p><h3>The cautionary tale: OpenCL</h3><p>Worth remembering, because the failure mode repeats. OpenCL standardized the API and the language and left the libraries, the tuning, and the <strong>last-mile compiler</strong> quality to each vendor. </p><p>Portable source, unportable performance, no cuBLAS equivalent, no tensor core story, committee cadence against NVIDIA&#8217;s annual co-design cadence. </p><p>SYCL is a better designed successor with the same structural problem: it is a standard for the layer that was never the moat.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>What replacing the kernel layer actually costs</h2><p>Let me put a rough number on the library moat, since that is the part people assume is unbuyable.</p><p>Take LLM inference specifically, not the whole CUDA-X surface. The kernels that matter for a modern decoder-only model at serving time are: <strong>dense GEMM </strong>in several shapes, grouped or batched GEMM for MoE, prefill attention, decode attention with paged KV, an MLA variant if you serve DeepSeek-shaped models, quantize and dequantize paths for FP8 and FP4, <strong>RMSNorm</strong>, RoPE, fused SwiGLU, sampling, KV cache gather and copy, speculative verification, and the collectives. </p><p>Call it 30 to 50 kernels that appear in profiles, of which eight or so carry most of the time.</p><p>Assume a fully loaded peak-quality kernel engineer at somewhere between $350,000 and $600,000, call it $450,000 to be some kind of &#8220;average&#8221;. Assume a top-tier attention or<strong> GEMM kernel</strong> for a new architecture takes three to nine engineer-months including autotuning, numerics validation and the long tail of shape coverage, and the secondary kernels take one to two.</p><ul><li><p><em>Eight critical kernels at six months each: roughly $1.8M</em></p></li><li><p><em>Forty secondary kernels at 1.5 months each: roughly $2.2M</em></p></li><li><p><em>Test infrastructure, numerics validation, CI across shapes and dtypes: call it 50 percent overhead</em></p></li></ul><p>That lands somewhere around $6M to $15M per hardware generation to build a competitive <em>inference</em> kernel layer from scratch. For a company selling accelerators, that is a rounding error. It is roughly one percent of a single generation&#8217;s mask set and tape-out cost.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!MXjK!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2a4139cb-ecc9-4d45-a7a6-7eca5dfc9eea_1634x912.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!MXjK!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2a4139cb-ecc9-4d45-a7a6-7eca5dfc9eea_1634x912.png 424w, https://substackcdn.com/image/fetch/$s_!MXjK!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2a4139cb-ecc9-4d45-a7a6-7eca5dfc9eea_1634x912.png 848w, https://substackcdn.com/image/fetch/$s_!MXjK!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2a4139cb-ecc9-4d45-a7a6-7eca5dfc9eea_1634x912.png 1272w, https://substackcdn.com/image/fetch/$s_!MXjK!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2a4139cb-ecc9-4d45-a7a6-7eca5dfc9eea_1634x912.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!MXjK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2a4139cb-ecc9-4d45-a7a6-7eca5dfc9eea_1634x912.png" width="1456" height="813" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2a4139cb-ecc9-4d45-a7a6-7eca5dfc9eea_1634x912.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:813,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Figure 6. A stated model, not a measurement. The inference kernel surface is small enough &quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 6. A stated model, not a measurement. The inference kernel surface is small enough " title="Figure 6. A stated model, not a measurement. The inference kernel surface is small enough " srcset="https://substackcdn.com/image/fetch/$s_!MXjK!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2a4139cb-ecc9-4d45-a7a6-7eca5dfc9eea_1634x912.png 424w, https://substackcdn.com/image/fetch/$s_!MXjK!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2a4139cb-ecc9-4d45-a7a6-7eca5dfc9eea_1634x912.png 848w, https://substackcdn.com/image/fetch/$s_!MXjK!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2a4139cb-ecc9-4d45-a7a6-7eca5dfc9eea_1634x912.png 1272w, https://substackcdn.com/image/fetch/$s_!MXjK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2a4139cb-ecc9-4d45-a7a6-7eca5dfc9eea_1634x912.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong><span>Figure 9 </span></strong>A stated model, not a measurement. The inference kernel surface is small enough for any accelerator vendor to buy outright, which is exactly why AMD is close on inference and not close on training.<span>Copy imageDownload PNG</span></figcaption></figure></div><p>Which explains something that otherwise looks strange: AMD is actually close on inference and not close on training. The inference kernel surface is small enough to buy. </p><p>The training surface, plus the full CUDA-X library set, plus the numerics validation across a decade of recipes, is somewhere between ten and fifty times that, and the <strong>co-design loop</strong> that makes the kernels good on day one instead of month nine is not purchasable at any price.</p><p>Treat these figures as a stated model with stated assumptions, not as a measurement. The point is the ratio, not the absolute value: the part everyone talks about is the cheap part.</p><h2>Where I might be wrong</h2><p>Four ways this argument could be badly calibrated, in descending order of how much they would cost me.</p><p><strong>The backend gap may be small.</strong> My case for layer 3 rests partly on the <code>maxas</code> era, which was Maxwell, and partly on recent reinforcement learning work over SASS schedules that finds wins by reordering <code>ptxas</code> output. If the residual there is three to five percent on typical kernels rather than the double digits the Maxwell results implied, then <code>ptxas</code> is close enough to optimal that the backend is not the moat at all, the libraries are, and I have spent a lot of words on the wrong floor. </p><p>I have not seen a rigorous modern measurement of that residual across a representative kernel set, and I would very much like to.</p><p><strong>The labour market framing may be a category error.</strong> I argue the scarce resource is people who can read control bits, and therefore that search and autotuning erode it. </p><p>But the advantage might be organisational rather than headcount: eighteen months of pre-silicon access, a hardware team down the hall, and the ability to change the chip when the kernel is awkward. If that is the real mechanism, then no amount of autotuning compute closes it, because the loop being run is design, not optimization.</p><p><strong>Tile IR may be commoditization rather than capture.</strong> I read the open dialect, the published bytecode format and the conformance suite as a way of owning the layer Triton was about to own. The more generous reading is that an MLIR dialect with a stable binary encoding and a test suite is precisely the artifact another vendor needs in order to target the same source, and NVIDIA has just handed it over. Both readings are consistent with the facts as of today. Which one is right will be visible in whether anyone outside NVIDIA ships a Tile IR backend.</p><p><strong>And the boring one.</strong> I do not have a B200. Everything in this piece at the Blackwell layer comes from documentation, from other people&#8217;s benchmarks, and from reading disassembly conventions rather than fresh disassembly. The control field layout is reverse engineered by researchers, not published. </p><p>Where I have marked things Tier B or C, that is what the marking means, and I would treat the Hopper and Blackwell specifics as directionally right rather than precisely right until someone with the hardware checks them.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Five dated calls</h2><p>Predictions with dates, so this piece can be scored rather than admired.</p><p><strong>1. No third party ships a production-quality SASS assembler for Hopper or Blackwell before 2030.</strong> Production quality meaning it passes a public correctness suite across a non-trivial kernel set on more than one SKU. Falsified by any project that does.</p><p><strong>2. NVIDIA never open sources the Tile IR to cubin lowering.</strong> The dialect, the bytecode, the spec and the conformance suite stay open. The optimizing backend does not. Falsified trivially and publicly if wrong.</p><p><strong>3. Before the end of 2027, an automated kernel generator beats a vendor library on a headline GEMM or attention shape on non-NVIDIA hardware, with reproducible numbers.</strong> The search-versus-labour thesis stands or falls on this one.</p><p><strong>4. AMD reaches within fifteen percent of NVIDIA on tokens per second per dollar for standard transformer inference on a like-for-like generation by end of 2027, and does not close the equivalent gap on training in the same window.</strong> The asymmetry is the prediction, not the number.</p><p><strong>5. At least one more CUDA compatibility project loses its funding before the end of 2027.</strong> ZLUDA has now done it twice. The economics of chasing a virtual ISA that moves every eighteen months have been tested and they do not work, and I expect the next attempt to discover this independently.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/p/how-the-nvidia-compiler-moat-actually?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/p/how-the-nvidia-compiler-moat-actually?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><div><hr></div><h2>What would actually work (for real)</h2><p>I am sorry to say it, but I do not think the moat gets breached. I think it gets routed around, in specific places, for specific workloads, and here is where I would put effort if that were my job.</p><p><strong>Own the tile layer, and make it a spec with teeth.</strong> Not a language, not a framework: a specification with a versioned binary encoding and a conformance test suite that vendors must pass. NVIDIA just told you this layer is the strategic one by shipping exactly that artifact. The mistake would be to let Tile IR become the only one.</p><p><strong>Make numerics a specification too.</strong> This is the gap nobody is filling. A conformance suite that pins reduction order tolerances, accumulation types, scaling granularity for narrow formats, and determinism guarantees, with reference outputs. Right now &#8220;<em>matches cuBLAS</em>&#8221; is folklore transmitted through issue threads. Turning it into a testable contract would remove a real and unpriced switching cost.</p><p><strong>Attack at inference, not at parity.</strong> The kernel surface is small, the workloads are known, the customers are cost-sensitive, and memory capacity per package is a lever NVIDIA does not always win. Trying to be CUDA-complete is the losing strategy; being excellent at forty kernels is the winning one.</p><p><strong>Spend compute instead of headcount.</strong> Every dollar into autotuning, superoptimization and learned scheduling attacks the actual scarce resource. Helion beating hand-written Triton on average is the proof of concept. This is the only lever in the list that gets cheaper over time.</p><p><strong>And keep the compiler argument in proportion.</strong> The binding constraint on who gets to serve AI workloads in 2026 is HBM supply, advanced packaging capacity, rack-scale networking and power. A perfect compiler on a chip you cannot buy, in a rack you cannot cool, connected by a fabric that stalls at 72 GPUs, wins nothing. The compiler moat is real and it is roughly the fourth most important moat NVIDIA has.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Check it yourself</h2><p>Nothing above requires trusting me. The relevant artifacts are all inspectable with the toolkit you already have installed.</p><p><strong>See the driver script decompose:</strong></p><pre><code><code>nvcc --dryrun -arch=sm_90 kernel.cu
</code></code></pre><p>Every stage, in order, with the real command lines. <code>cicc</code>, <code>ptxas</code>, <code>fatbinary</code>, <code>nvlink</code>.</p><p><strong>Watch the virtual ISA:</strong></p><pre><code><code>nvcc -arch=sm_90 -ptx kernel.cu -o kernel.ptx
</code></code></pre><p>Note the unbounded <code>%r</code> virtual registers and the total absence of scheduling information.</p><p><strong>Watch the last mile do the work:</strong></p><pre><code><code>ptxas -arch=sm_90 -v kernel.ptx -o kernel.cubin
</code></code></pre><p>The <code>-v</code> output tells you registers used, spill stores, spill loads, shared memory. That register count is a policy decision <code>ptxas</code> made on your behalf, and it sets your occupancy.</p><p><strong>Look at the control bits:</strong></p><pre><code><code>nvdisasm -c -g kernel.cubin
cuobjdump --dump-sass kernel.cubin
</code></code></pre><p>The bracketed fields next to each instruction are the scheduling payload. Compare the same kernel compiled with and without <code>-maxrregcount</code> and watch the stall counts and barrier usage change.</p><p><strong>Confirm the frontend is a commodity:</strong></p><pre><code><code>clang++ -x cuda --cuda-gpu-arch=sm_90 --cuda-device-only -S kernel.cu -o clang.ptx
</code></code></pre><p>Diff that against <code>nvcc</code>&#8216;s PTX. They differ. Then push both through <code>ptxas</code> and compare the SASS. That comparison is the whole argument of this piece in two commands: two independent frontends, one backend, and the backend is where the differences stop mattering.</p><p><strong>Find the boundary:</strong> try to turn a modified SASS listing back into a <code>cubin</code> using only NVIDIA tooling. There is no such command. That absence is the moat.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Confidence dossier</h2><p><strong><span>Tier A</span></strong>, directly verifiable from primary sources or first-party documentation</p><ul><li><p><code>nvcc</code> is a driver invoking <code>cicc</code>, <code>ptxas</code>, <code>fatbinary</code> and <code>nvlink</code>; observable via <code>--dryrun</code></p></li><li><p>NVVM IR is a documented LLVM IR subset with a public <code>libNVVM</code> API; Clang&#8217;s NVPTX backend is upstream</p></li><li><p>PTX is a virtual ISA; the driver JIT-compiles it, providing forward compatibility</p></li><li><p>There is no NVIDIA-provided SASS assembler; <code>nvdisasm</code> and <code>cuobjdump</code> disassemble only</p></li><li><p>Blackwell deprecated <code>wgmma</code> and introduced <code>tcgen05</code>, Tensor Memory, single-thread MMA issue and <code>cta_group::2</code></p></li><li><p>CUDA Tile IR shipped in CUDA 13.1; the MLIR dialect, bytecode format and conformance suite are open source; the bytecode is specified as stable and versioned</p></li><li><p>A Triton to Tile IR backend exists, authored by NVIDIA</p></li><li><p>The CUDA EULA restricts translating SDK-generated output artifacts to non-NVIDIA platforms</p></li><li><p><code>wgmma</code> requires the <code>sm_90a</code> architecture-specific target and <code>tcgen05</code> requires <code>sm_100a</code> or <code>sm_103a</code>; code for <code>a</code> targets is not forward compatible</p></li><li><p>A <code>cubin</code> is an ELF64 object with per-kernel <code>.text</code>, <code>.nv.constant0</code> and <code>.nv.info</code> sections</p></li><li><p>Blackwell Tensor Memory is 256 KB per SM, 128 lanes by 512 columns of 32 bits, allocated in power-of-two column counts</p></li><li><p>MXFP8 uses one E8M0 scale per 32 elements; NVFP4 uses an E4M3 scale per 16 elements plus a tensor-level FP32 scale</p></li></ul><p><strong><span>Tier B</span></strong>, well-supported by multiple independent secondary sources or peer-reviewed measurement</p><ul><li><p>The Volta-and-later control field layout: reuse, wait mask, read and write barrier indices, yield, stall count, packed into the instruction word</p></li><li><p>Fixed-latency instructions lack hardware RAW interlocks; incorrect stall counts produce incorrect results</p></li><li><p>Instruction encoding widened from 64 to 128 bits at Volta</p></li><li><p>H100 register file of 65,536 32-bit registers per SM, 64 maximum resident warps, eight-register allocation granularity, and the occupancy arithmetic that follows from them</p></li><li><p>FlashAttention 4 in CuTe DSL outperforming cuDNN 9.13 and Triton by the margins cited</p></li><li><p>ROCm 7.2 library and AITER performance improvements on MI300 through MI355X</p></li><li><p>Helion&#8217;s reported speedups over Triton and <code>torch.compile</code></p></li><li><p>NCCL&#8217;s channel-per-thread-block structure, its protocol selection, and the fact that collectives consume SMs that are then unavailable for compute</p></li><li><p>In-network reduction in NVSwitch and InfiniBand switch silicon</p></li><li><p>SCALE&#8217;s benchmark claims against HIP and <code>nvcc</code></p></li></ul><p><strong><span>Tier C</span></strong>, single-source, historical, or reported rather than independently verified</p><ul><li><p><code>maxas</code> SGEMM efficiency figures on GM204 versus contemporaneous cuBLAS</p></li><li><p>The annotated SASS fragment, which is reconstructed in the documented format rather than captured from a specific compilation</p></li><li><p>DeepSeek&#8217;s specific SM partitioning and custom PTX details, as described in the V3 paper and subsequent reporting</p></li><li><p>Spectral Compute&#8217;s NVIDIA Inception membership and its strategic reading</p></li><li><p>cuTile&#8217;s competitiveness with cuBLAS and cuDNN at large shapes, which is a first-party claim awaiting third-party benchmarks</p></li></ul><p><strong><span>Tier D</span></strong>, my model or my opinion, flagged as such</p><ul><li><p>The $6M to $15M per-generation estimate for an inference kernel layer, and the 10x to 50x multiplier for the full surface</p></li><li><p>The claim that layer 3 is the load-bearing wall rather than layer 1</p></li><li><p>The reading of Tile IR as a strategic response to Triton&#8217;s position in PyTorch</p></li><li><p>The assertion that the moat is best modelled as a labour market</p></li><li><p>The ranking of the compiler moat as roughly fourth behind supply, packaging and networking</p></li><li><p>All five dated calls, which are forecasts and should be read as such</p></li><li><p>The four self-criticisms, which are my own estimate of where this argument is weakest</p></li></ul><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Bibliography</h2><p><strong>Primary documentation</strong></p><ul><li><p>NVIDIA, <em>Parallel Thread Execution ISA</em>. https://docs.nvidia.com/cuda/parallel-thread-execution/</p></li><li><p>NVIDIA, <em>CUDA Binary Utilities</em> (<code>cuobjdump</code>, <code>nvdisasm</code>, SASS instruction set tables). https://docs.nvidia.com/cuda/cuda-binary-utilities/</p></li><li><p>NVIDIA, <em>NVVM IR Specification</em>. https://docs.nvidia.com/cuda/nvvm-ir-spec/</p></li><li><p>NVIDIA, <em>CUDA Compiler Driver NVCC</em>. https://docs.nvidia.com/cuda/cuda-compiler-driver-nvcc/</p></li><li><p>NVIDIA, <em>License Agreement for NVIDIA Software Development Kits</em>. https://docs.nvidia.com/cuda/eula/index.html</p></li><li><p>NVIDIA, <em>CUDA Tile</em>. https://developer.nvidia.com/cuda/tile</p></li><li><p>NVIDIA, <em>Tile IR: Binary Format</em>. https://docs.nvidia.com/cuda/tile-ir/latest/sections/bytecode.html</p></li><li><p>NVIDIA, <em>cuTile Python documentation</em>. https://docs.nvidia.com/cuda/cutile-python/</p></li><li><p>NVIDIA, <em>CUTLASS documentation: tcgen05 MMA Programming Guide</em>. https://docs.nvidia.com/cutlass/latest/media/docs/pythonDSL/mma_docs/tcgen05_programming.html</p></li><li><p>NVIDIA, <em>CUTLASS: Blackwell SM100 functionality</em>. https://docs.nvidia.com/cutlass/4.2.1/media/docs/cpp/blackwell_functionality.html</p></li><li><p>AMD, <em>ROCm documentation: vLLM inference optimization</em>. https://rocm.docs.amd.com/en/latest/how-to/rocm-for-ai/inference-optimization/vllm-optimization.html</p></li></ul><p><strong>Microarchitecture and reverse engineering</strong></p><ul><li><p>Jia, Maggioni, Staiger, Scarpazza, <em>Dissecting the NVIDIA Volta GPU Architecture via Microbenchmarking</em>, arXiv:1804.06826. https://arxiv.org/pdf/1804.06826</p></li><li><p>Huerta, Abaie et al., <em>Analyzing Modern NVIDIA GPU cores</em>, arXiv:2503.20481. https://arxiv.org/pdf/2503.20481</p></li><li><p><em>Dissecting and Modeling the Architecture of Modern GPU Cores</em>, MICRO 58, 2025. https://dl.acm.org/doi/10.1145/3725843.3756041</p></li><li><p>Luo et al., <em>Benchmarking and Dissecting the Nvidia Hopper GPU Architecture</em>, arXiv:2402.13499</p></li><li><p><em>CuAsmRL: Optimizing GPU SASS Schedules via Deep Reinforcement Learning</em>, arXiv:2501.08071. https://arxiv.org/html/2501.08071v1</p></li><li><p><em>GPA: A GPU Performance Advisor Based on Instruction Sampling</em>, arXiv:2009.04061</p></li><li><p>Finn, <em>What happens when you run a CUDA kernel</em>. https://fergusfinn.com/blog/what-happens-when-you-run-a-gpu-kernel/</p></li></ul><p><strong>Assembler and toolchain projects</strong></p><ul><li><p>Gray, <em>maxas</em>, Maxwell assembler and control code notes. https://github.com/NervanaSystems/maxas/wiki/Control-Codes</p></li><li><p><em>CuAssembler</em>. https://github.com/cloudcores/CuAssembler</p></li><li><p><em>openptxas</em>, gpuocelot. https://github.com/gpuocelot/openptxas</p></li><li><p>NVlabs, <em>SASSI: Flexible GPGPU instrumentation</em>, distributed as a closed-source <code>ptxas</code> fork. https://github.com/NVlabs/SASSI</p></li><li><p>NVIDIA, <em>cuda-tile</em>. https://github.com/NVIDIA/cuda-tile</p></li></ul><p><strong>Kernel authoring layers</strong></p><ul><li><p>Colfax Research, <em>CUTLASS Tutorial: Writing GEMM Kernels Using Tensor Memory for NVIDIA Blackwell GPUs</em>. https://research.colfax-intl.com/cutlass-tutorial-writing-gemm-kernels-using-tensor-memory-for-nvidia-blackwell-gpus/</p></li><li><p>Colfax Research, <em>CUTLASS Tutorial: Hardware-supported Block-scaling with NVIDIA Blackwell GPUs</em>. https://research.colfax-intl.com/cutlass-tutorial-hardware-supported-block-scaling-with-nvidia-blackwell-gpus/</p></li><li><p>SemiAnalysis, <em>Dissecting Nvidia Blackwell: Tensor Cores, PTX Instructions, SASS, Floorsweep, Yield</em>. https://newsletter.semianalysis.com/p/dissecting-nvidia-blackwell-tensor</p></li><li><p>PyTorch, <em>Helion: A High-Level DSL for Performant and Portable ML Kernels</em>. https://pytorch.org/blog/helion/</p></li><li><p>pytorch/helion repository. https://github.com/pytorch/helion</p></li><li><p>PyTorch RFC, <em>Inductor cuTile Backend</em>, issue 175311. https://github.com/pytorch/pytorch/issues/175311</p></li><li><p>NVIDIA Developer Blog, <em>Advancing GPU Programming with the CUDA Tile IR Backend for OpenAI Triton</em>. https://developer.nvidia.com/blog/?p=112268</p></li><li><p>Red Hat Emerging Technologies, <em>From hand-tuned to generated: a reproducible Triton GPU kernel benchmark across different vendors</em>. https://next.redhat.com/2026/02/12/from-hand-tuned-to-generated-a-reproducible-triton-gpu-kernel-benchmark-across-different-vendors/</p></li><li><p>Modular, <em>Structured Mojo Kernels Part 4: Portability and the Road Ahead</em>. https://www.modular.com/blog/structured-mojo-kernels-part-4-portability-and-the-road-ahead</p></li><li><p><em>Mojo: MLIR-Based Performance-Portable HPC Science Kernels on GPUs for the Python Ecosystem</em>, arXiv:2509.21039</p></li></ul><p><strong>Alternatives and portability</strong></p><ul><li><p>Spectral Compute, <em>SCALE</em>. https://scale-lang.com/ and https://github.com/spectral-compute/scale-docs</p></li><li><p>HPCwire, <em>Spectral Compute Aims to Set CUDA Free. Will It Succeed?</em>, July 2026. https://www.hpcwire.com/2026/07/09/spectral-compute-aims-to-set-cuda-free-will-it-succeed/</p></li><li><p><em>SCALE: Ahead-Of-Time Compilation of CUDA for AMD GPUs</em>, Middleware 2024 demo track. https://dl.acm.org/doi/10.1145/3704440.3704782</p></li><li><p>ZLUDA, <em>Update Q1 and Q2 2026: back to the roots</em>. https://vosen.github.io/ZLUDA/blog/zluda-update-q1q2-2026/</p></li><li><p>Phoronix, <em>ZLUDA For CUDA On Non-NVIDIA GPUs Enables AMD ROCm 7 Support</em>. https://www.phoronix.com/news/ZLUDA-ROCm-7</p></li><li><p>Tom&#8217;s Hardware, <em>Nvidia bans using translation layers for CUDA software</em>, March 2024. https://www.tomshardware.com/pc-components/gpus/nvidia-bans-using-translation-layers-for-cuda-software-to-run-on-other-chips-new-restriction-apparently-targets-zluda-and-some-chinese-gpu-makers</p></li><li><p>vLLM Blog, <em>Beyond Porting: How vLLM Orchestrates High-Performance Inference on AMD ROCm</em>. https://vllm.ai/blog/2026-02-27-rocm-attention-backend</p></li><li><p>AMD, <em>Speed is the Moat: Inference Performance on AMD GPUs</em>. https://www.amd.com/en/developer/resources/technical-articles/2026/inference-performance-on-amd-gpus.html</p></li><li><p>tinygrad repository and release notes. https://github.com/tinygrad/tinygrad</p></li><li><p>Pall et al., <em>GROMACS on AMD GPU-Based HPC Platforms: Using SYCL for Performance and Portability</em>, arXiv:2405.01420</p></li></ul><p><strong>Model and systems work referenced</strong></p><ul><li><p>DeepSeek-AI, <em>DeepSeek-V3 Technical Report</em>, arXiv:2412.19437</p></li><li><p>Tom&#8217;s Hardware, <em>DeepSeek&#8217;s AI breakthrough bypasses industry-standard CUDA for some functions</em>. https://www.tomshardware.com/tech-industry/artificial-intelligence/deepseeks-ai-breakthrough-bypasses-industry-standard-cuda-uses-assembly-like-ptx-programming-instead</p></li><li><p>Dao et al., FlashAttention series, most recently FlashAttention 4 on Blackwell in CuTe DSL</p></li><li><p>Bauer et al., <em>Singe: Leveraging Warp Specialization for High Performance on GPUs</em>, PPoPP 2014</p></li></ul><p><em>Corrections and disagreements are welcome and get published. If you have measured </em><code>ptxas</code><em> output against hand-scheduled SASS on Hopper or Blackwell, I would like to see the numbers.</em></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share The Software Frontier&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share The Software Frontier</span></a></p><p></p>]]></content:encoded></item><item><title><![CDATA[Kimi K3: 2.8 Trillion Parameters, Four Bits at a Time]]></title><description><![CDATA[Moonshot just announced the largest open-weight model ever built. The parameter count is the headline. The serving stack is the story.]]></description><link>https://www.thesoftwarefrontier.com/p/kimi-k3-28-trillion-parameters-four</link><guid isPermaLink="false">https://www.thesoftwarefrontier.com/p/kimi-k3-28-trillion-parameters-four</guid><dc:creator><![CDATA[Lorenzo Bradanini]]></dc:creator><pubDate>Wed, 22 Jul 2026 09:53:06 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!B2ac!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F73922128-9535-44e5-b887-ad4a9f5ccea0_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!B2ac!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F73922128-9535-44e5-b887-ad4a9f5ccea0_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!B2ac!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F73922128-9535-44e5-b887-ad4a9f5ccea0_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!B2ac!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F73922128-9535-44e5-b887-ad4a9f5ccea0_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!B2ac!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F73922128-9535-44e5-b887-ad4a9f5ccea0_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!B2ac!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F73922128-9535-44e5-b887-ad4a9f5ccea0_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!B2ac!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F73922128-9535-44e5-b887-ad4a9f5ccea0_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/73922128-9535-44e5-b887-ad4a9f5ccea0_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2860025,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/206014241?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F73922128-9535-44e5-b887-ad4a9f5ccea0_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!B2ac!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F73922128-9535-44e5-b887-ad4a9f5ccea0_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!B2ac!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F73922128-9535-44e5-b887-ad4a9f5ccea0_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!B2ac!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F73922128-9535-44e5-b887-ad4a9f5ccea0_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!B2ac!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F73922128-9535-44e5-b887-ad4a9f5ccea0_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>Intro</h2><p>Moonshot AI released <strong>Kimi K3 </strong>precisely on July 16. The headline number is 2.8 trillion parameters, which makes it the largest open-weight model ever announced. </p><p>We spent the launch week reading almost everything published about it, and the more we read, the less the parameter count felt like the point. The point is that a <strong>2.8T model</strong> is being served today, at Claude Sonnet prices, with a flat rate across a one million token context window. </p><p>The reasons that is possible sit exactly at the layer this newsletter cares about: attention design, expert routing, <strong>quantization format</strong>, and cache economics.</p><p>A caveat before the numbers. The technical report is not out yet. Moonshot&#8217;s launch blog states that architecture, training, and <strong>evaluation details</strong> will arrive alongside the report, and several figures below come from community analysis of the launch documentation rather than from a paper. </p><p>Where we compute something ourselves, the assumptions are stated inline and marked as ours, and there is a confidence appendix at the end. </p><p>This is a <strong>launch-week analysis</strong>, not a postmortem.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>What shipped</h2><p>The facts first, because launch weeks always tend to generate a lot of useless fog. The model went live on July 16 through kimi.com, the Kimi apps, Kimi Work, Kimi Code, and the <strong>Moonshot API</strong> at api.moonshot.ai under the <strong>model ID kimi-k3. </strong></p><p>It is also listed on <strong>OpenRouter</strong> at the same rates as the direct API. What did not ship is the checkpoint: Moonshot says the full weights land by July 27, and until the files appear on Hugging Face, K3 is a hosted model you can call but not download. </p><p>Verdent&#8217;s <strong>launch guide </strong>reports the weights are coming under a modified MIT license, which would basically match the K2 line, but the license text is unconfirmed until the release itself.</p><p>The specification, per Moonshot: 2.8 trillion total parameters, a mixture of experts <strong>activating 16 of 896 experts per token</strong>, native vision input rather than a bolted-on adapter, and a context window of 1,048,576 tokens. The model reasons before every answer. At launch, reasoning effort is locked to max, with lower effort modes promised in later updates. That detail matters more than it sounds, and we will get to it in the pricing section.</p><p>The timing is not subtle either. VentureBeat notes the release landed just a few days before the <strong>World Artificial Intelligence Conference</strong> in Shanghai, and frames it as a comeback move for a company whose position had eroded badly during DeepSeek&#8217;s rise. </p><p>Moonshot is backed by <strong>Alibaba</strong>, which also builds the <strong>Qwen line</strong>, so the Chinese open-model field is now crowded with players who share an investor and compete anyway. Moonshot&#8217;s own blog claims that for nine of the past twelve months, Kimi models have held the frontier of open-source scale. </p><p>Whatever you think of that framing, the release cadence behind it is real: K2 in July 2025 at one trillion total parameters, K2 Thinking in November with <strong>quantization-aware INT4</strong>, and now K3 at nearly three times the K2 total.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/p/kimi-k3-28-trillion-parameters-four?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/p/kimi-k3-28-trillion-parameters-four?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><div><hr></div><h2>The shape of the model</h2><p>Three important architectural pieces define K3: the <strong>sparsity of the expert layer</strong>, the hybrid attention stack, and a modified residual path. Each one is a serving decision as much as a modeling decision.</p><p>My advice is to start with sparsity, because the trend line is the clearest signal of where frontier MoE design is going. K2 activated 8 of 384 experts, roughly 32B active parameters out of 1T total, an activation ratio around 3.1 percent. K3 activates 16 of 896. </p><p>Moonshot has not published the active parameter count; the community model card analysis on Hugging Face estimates <strong>roughly 50B active equivalent</strong>, which against 2.8T total is circa 1.8 percent. Total capacity nearly tripled while active compute grew by maybe half. </p><p>That is the whole game: parameters are cheap to store and expensive to move, so you scale what sits in memory and hold nearly flat what has to cross the datapath every token. </p><p><strong>DeepSeek</strong> formalized this direction at <strong>671B total </strong>and <strong>5.5 percent active</strong>; Moonshot is now running it harder than anyone with an open checkpoint. </p><p>The lineage is worth one sentence: the sparse expert layer goes back to Shazeer&#8217;s 2017 outrageously large networks paper, was made trainable at datacenter scale by <strong>GShard </strong>and <strong>Switch</strong>, and was pushed toward fine-grained experts with a shared trunk by the DeepSeek line; K3 extends that same axis to 896, and the references at the end walk the chain to the primary sources.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!QV-V!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ec1f9f7-2bc0-4286-a8c1-9616aea80ab6_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!QV-V!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ec1f9f7-2bc0-4286-a8c1-9616aea80ab6_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!QV-V!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ec1f9f7-2bc0-4286-a8c1-9616aea80ab6_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!QV-V!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ec1f9f7-2bc0-4286-a8c1-9616aea80ab6_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!QV-V!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ec1f9f7-2bc0-4286-a8c1-9616aea80ab6_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!QV-V!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ec1f9f7-2bc0-4286-a8c1-9616aea80ab6_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6ec1f9f7-2bc0-4286-a8c1-9616aea80ab6_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;The sparsity frontier: open flagship MoE models&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="The sparsity frontier: open flagship MoE models" title="The sparsity frontier: open flagship MoE models" srcset="https://substackcdn.com/image/fetch/$s_!QV-V!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ec1f9f7-2bc0-4286-a8c1-9616aea80ab6_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!QV-V!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ec1f9f7-2bc0-4286-a8c1-9616aea80ab6_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!QV-V!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ec1f9f7-2bc0-4286-a8c1-9616aea80ab6_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!QV-V!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ec1f9f7-2bc0-4286-a8c1-9616aea80ab6_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 1. The sparsity frontier: open flagship MoE models.</figcaption></figure></div><p>The expert layer runs on what Moonshot calls <strong>Stable LatentMoE</strong>. According to the launch documentation, routing happens in a latent space rather than directly on token representations, load is managed by a mechanism called <strong>Quantile Balancing,</strong> and overflow tokens are soft-dropped rather than hard-bounced. </p><p>All three claims obviously await the technical report, but the intent is legible: at 896 experts, classical <em>routing instability</em> and load skew are the failure modes that kill both training and serving, and everything named here is aimed at them. </p><p>If you have ever watched a MoE deployment where two hot experts saturate their devices while the other <strong>few hundred idle</strong>, you know why a lab would lead with the word Stable.</p><p>The attention stack is a hybrid, and it has a paper trail. K3 is built on <strong>Kimi Delta Attention</strong>, KDA, which Moonshot introduced with the Kimi Linear report in late 2025. </p><p>KDA is essentially a gated refinement of DeltaNet: a delta-rule state update with fine-grained per-channel decay instead of a scalar forget gate, implemented with chunked parallel kernels so the recurrence trains at practical speed. </p><p>In the Kimi Linear configuration these layers were interleaved with full attention at a <em>three to one ratio</em>, the global layers used multi-head latent attention, and the paper reported on the order of a <strong>75 percent KV cache reduction</strong> and up to a reported 6x decode throughput gain at million-token lengths against a full-attention baseline. </p><p>K3&#8217;s global layers use Gated MLA, a gated variant of the same latent attention. Moonshot has not confirmed K3&#8217;s layer ratio, but the lineage tells you the design logic: KDA layers carry a fixed-size recurrent state that costs the same at position one million as at position one hundred, and only the global layers accumulate <strong>position-indexed KV</strong>, compressed into latents at that. </p><p>The million-token window is not priced flat out of generosity. It is priced flat because most of the stack does not pay for length.</p><p>For readers who want the update rule, KDA in the Kimi Linear formulation maintains a matrix-valued state S, rewritten each step as</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;plaintext&quot;,&quot;nodeId&quot;:&quot;0007b882-8309-414c-9f06-2ae450d85604&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-plaintext">S_t = (I - beta_t k_t k_t^T) Diag(alpha_t) S_{t-1} + beta_t k_t v_t^T</code></pre></div><p>read right to left: <strong>decay the previous state channel</strong> by channel through Diag(alpha_t), erase the stale association along the incoming key direction, then write the new key-value pair at strength beta_t. </p><p>The per-channel diagonal gate is the refinement over <strong>Gated DeltaNet&#8217;s scalar decay</strong>, restoring some of the selectivity full attention gets for free, and the chunked kernel that makes the recurrence trainable at scale, a specialized diagonal-plus-low-rank formulation, is the same computational shape the prefill discussion below returns to.</p><p>The third piece is <strong>Attention Residuals</strong>, AttnRes, described as a drop-in replacement for standard residual connections. Instead of every layer adding onto one uniformly accumulated stream, each layer can selectively retrieve representations from arbitrary earlier layers. </p><p>The claimed motivation is depth: in <strong>very deep MoE stacks</strong>, different experts fire at different depths, and a selective residual path keeps early information reachable without forcing it through every intermediate transformation. </p><p>It is the kind of change that sounds small and touches everything, and it is the single item on this list we most want to see ablated in the report.</p><p>Around the edges: a custom activation Moonshot calls SiTU, a sigmoid tanh unit replacing the usual <strong>SwiGLU family</strong>, and a per-head variant of the Muon optimizer, scheduling learning rates at the level of individual attention heads. </p><p>Muon is a K2-era inheritance; per-head scheduling is new. Moonshot&#8217;s aggregate claim is that the architectural and data changes together give K3 roughly <em>2.5 times the scaling efficiency of K2</em>, meaning more capability per unit of training compute. That number is <strong>vendor-measured</strong>, marked as unverified, and should be treated accordingly.</p><p>The vision path deserves more than a scope note, because it is half the modality story and it touches everything above. What is known: <strong>vision is native </strong>rather than adapter-based per the community model card reading, continuing the <strong>multimodal line</strong> Moonshot opened with K2.5, which means image tokens enter the same transformer stack and the same 896-expert router as text, and the multimodal scores ship on day one rather than in a point release. </p><p>The API mechanics are documented even where the economics are not: images go in as <strong>base64 payloads </strong>or platform file references, public image URLs are not accepted, and OpenRouter lists image input at the standard rates. </p><p>The agent framing leans on vision explicitly, a model pitched for iterating against screenshots, logs, and runtime feedback rather than captioning, and Moonshot&#8217;s own multimodal numbers, <strong>81.6 on MMMU-Pro </strong>and 94.3 on MathVision, sit in the launch tables.</p><p>What is not known at all is the <strong>exchange rate</strong>. The rate card is a single line, and whether an image bills at a multiplier over the <em>$3 text rate</em> or simply lands as however many tokens the tokenizer emits is unconfirmed, a gap the pricing trackers have flagged since day one. </p><p>The serving consequences do not wait for the answer. Whatever the patch geometry turns out to be, images are prefill load: a screenshot-heavy agent loop is a long-prompt workload wearing a different costume, and it collides with <strong>prefix caching </strong>in a way text does not, because a refreshed screenshot mid-transcript invalidates every cached token downstream of it. </p><p>The append-only discipline from the harness section applies doubly to media: attach new frames at the tail, never edit them in place.</p><p>And there is a broad research question waiting in the checkpoint. Early fusion means text and vision tokens share the router and the expert pool. </p><p>Whether the 896 experts partition by modality, some effectively becoming vision specialists, or blend, is exactly the kind of question the July 27 weights let anyone answer with an afternoon of <strong>routing statistics,</strong> and the answer feeds straight back into pruning, because a modality-partitioned pool compresses very differently from a blended one.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Four bits from the start</h2><p>The decision that makes 2.8T servable is not in the attention stack. It&#8217;s in the number format. K3 was trained with quantization-aware training from the <strong>supervised fine-tuning stage</strong> onward: MXFP4 for weights, MXFP8 for activations. </p><p>This is not a post-training quantization pass applied to a bf16 checkpoint. The model learned to live inside <strong>four-bit weights</strong>, compensating for quantization error during training instead of absorbing it afterward.</p><p>Terms, briefly, because the acronym is doing a lot of work. MXFP4 is the four-bit member of the <strong>OCP Microscaling family</strong>: weights are grouped into blocks of 32, each element is an E2M1 float, one sign bit, two exponent bits, one mantissa bit, and every block shares a single 8-bit power-of-two scale. </p><p>The element grid has eight magnitudes, 0, 0.5, 1, 1.5, 2, 3, 4, 6, spaced logarithmically rather than uniformly. Activations ride one tier up as <strong>MXFP8</strong>, eight-bit elements under the same block-scale scheme, where the precision matters for numerical stability. </p><p>The migration from K2 Thinking, which shipped INT4 quantization-aware training in November 2025, to floating-point four bits here is a change of grid, and grids have consequences you can measure. So we measured, on a synthetic but realistic weight distribution, a <strong>Gaussian with and without </strong>a sprinkle<strong> </strong>of<strong> 30-sigma outliers</strong>, block size 32, both formats implemented per spec:</p><pre><code><code># MXFP4 (E2M1 elements, shared E8M0 block scale) vs INT4 groupwise, block size 32.
# Question: what does each grid do to a realistic weight distribution?
import numpy as np
rng = np.random.default_rng(7)

E2M1 = np.array([0.0, 0.5, 1.0, 1.5, 2.0, 3.0, 4.0, 6.0])   # magnitudes

def quant_mxfp4(w, block=32):
    out = np.empty_like(w)
    for i in range(0, w.size, block):
        b = w[i:i+block]
        amax = np.abs(b).max()
        if amax == 0:
            out[i:i+block] = 0; continue
        shared_exp = np.floor(np.log2(amax)) - 2          # emax of E2M1 is 2
        scale = 2.0 ** shared_exp                          # E8M0: power of two only
        x = np.clip(np.abs(b) / scale, 0, 6.0)
        idx = np.abs(x[:, None] - E2M1[None, :]).argmin(1) # round to nearest level
        out[i:i+block] = np.sign(b) * E2M1[idx] * scale
    return out

def quant_int4(w, block=32):
    out = np.empty_like(w)
    for i in range(0, w.size, block):
        b = w[i:i+block]
        amax = np.abs(b).max()
        if amax == 0:
            out[i:i+block] = 0; continue
        scale = amax / 7.0                                 # fp16 scale, uniform grid
        out[i:i+block] = np.clip(np.round(b / scale), -7, 7) * scale
    return out

def report(name, w):
    for label, q in (("mxfp4", quant_mxfp4(w)), ("int4-g32", quant_int4(w))):
        rel = np.linalg.norm(w - q) / np.linalg.norm(w)
        small = np.abs(w) &lt; np.quantile(np.abs(w), 0.99)   # error on the 99% mass
        rel_small = np.linalg.norm(w[small] - q[small]) / np.linalg.norm(w[small])
        dead = (q[small] == 0).mean() * 100                # small weights crushed to zero
        print(f"{name:18s} {label:9s} relRMSE {rel:.4f}  relRMSE(99% mass) {rel_small:.4f}  zeroed {dead:5.2f}%")

w_gauss = rng.normal(0, 0.02, 1 &lt;&lt; 20)
report("gaussian", w_gauss)

w_out = w_gauss.copy()                                     # 0.2% outliers at 30x sigma
hit = rng.choice(w_out.size, w_out.size // 500, replace=False)
w_out[hit] = rng.choice([-1, 1], hit.size) * 0.6
report("gaussian+outliers", w_out)
</code></code></pre><p>Executed output:</p><pre><code><code>gaussian           mxfp4     relRMSE 0.1141  relRMSE(99% mass) 0.1102  zeroed  8.19%
gaussian           int4-g32  relRMSE 0.0970  relRMSE(99% mass) 0.1012  zeroed 13.37%
gaussian+outliers  mxfp4     relRMSE 0.1919  relRMSE(99% mass) 0.2353  zeroed 13.01%
gaussian+outliers  int4-g32  relRMSE 0.1499  relRMSE(99% mass) 0.2585  zeroed 18.40%
</code></code></pre><p>Read the table honestly and neither format dominates. </p><p>INT4 with a <strong>floating-point scale per block</strong> really wins global reconstruction error in both regimes, because its fifteen uniform levels and exact scale beat eight logarithmic levels under a power-of-two scale on Gaussian mass. </p><p>What the FP4 grid buys is essentially at the bottom of the distribution: it zeroes meaningfully fewer small weights, 8.2 versus 13.4 percent on clean data, 13.0 versus 18.4 with outliers, and it carries the bulk of the mass with lower error when an <strong>outlier inflates a block&#8217;s scale</strong>, because the log-spaced levels keep resolution near zero exactly where a uniform grid goes coarse. </p><p>Small weights are where accumulated damage shows up at scale, so this is not a cosmetic difference, but it is also not the headline reason for the format. </p><p>The headline reasons are that the exponent-only scale is free to apply in hardware, that <strong>Blackwell tensor cores</strong> execute block-scaled MXFP4 natively at full rate, with the Hugging Face community analysis reporting the same for<em> AMD&#8217;s MI400 class</em>, and that quantization-aware training exists precisely to claw back whatever the grid loses. </p><p>A lab that wanted a research artifact would release <strong>bf16</strong> and let the community fight over quants. A lab that wants its model actually served releases the four-bit weights it trained, in the format the current accelerator generation runs at full rate.</p><p>The arithmetic consequence: 2.8 trillion parameters at four bits is <strong>1.4 TB of element payload</strong>, and closer to 1.5 TB once the per-block scales ride along, since a 32-element block carries 136 bits, 4.25 effective bits per weight. </p><p>The same model in FP16 would be about 5.6 TB. That factor of nearly four compounds through the memory system: fewer devices to hold the model, and a<strong> quarter of the bandwidth per token </strong>spent reading weights. Be precise about what the four does and does not buy on the wire, though. </p><p>Expert dispatch and combine move activations, and activations are MXFP8, so all-to-all traffic halves relative to a bf16 model rather than quartering. The<strong> full 4x applies to weight movemen</strong>t: initial loading, host offload, expert migration and rebalancing. </p><p>Weights and activations shrink by different factors, and conflating them is how serving estimates go wrong by 2x.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Qfcd!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F25ef7b1b-c40c-4cb5-98f5-5822efaae7fe_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Qfcd!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F25ef7b1b-c40c-4cb5-98f5-5822efaae7fe_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!Qfcd!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F25ef7b1b-c40c-4cb5-98f5-5822efaae7fe_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!Qfcd!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F25ef7b1b-c40c-4cb5-98f5-5822efaae7fe_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!Qfcd!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F25ef7b1b-c40c-4cb5-98f5-5822efaae7fe_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Qfcd!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F25ef7b1b-c40c-4cb5-98f5-5822efaae7fe_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/25ef7b1b-c40c-4cb5-98f5-5822efaae7fe_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Weight storage by precision&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Weight storage by precision" title="Weight storage by precision" srcset="https://substackcdn.com/image/fetch/$s_!Qfcd!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F25ef7b1b-c40c-4cb5-98f5-5822efaae7fe_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!Qfcd!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F25ef7b1b-c40c-4cb5-98f5-5822efaae7fe_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!Qfcd!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F25ef7b1b-c40c-4cb5-98f5-5822efaae7fe_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!Qfcd!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F25ef7b1b-c40c-4cb5-98f5-5822efaae7fe_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 2. Weight storage by precision.</figcaption></figure></div><p>There is a subtler consequence for the open release. When the weights drop on July 27, they will be born four-bit. </p><p>There will be no ambiguity about which community quant is faithful, because the <strong>FP4 checkpoint is the reference</strong>, not a lossy derivative. The flip side is that the fine-tuning story gets strange: LoRA and QLoRA on top of an already-quantized MoE base is thinly explored territory, a point the Hugging Face overview raises as an open research question. </p><p>The first <strong>serious K3 fine-tunes</strong> will be experiments in method, not just in data. One loose thread already dangles here. One launch tracker describes OpenRouter&#8217;s route as Moonshot&#8217;s hosted INT4 endpoint. </p><p>That may be nothing more than four-bit shorthand for the MXFP4 checkpoint, or it may mean the serving fleet runs a different quantization than the training format, and the difference matters to anyone who plans to <strong>benchmark API behavior</strong> against the weights they download. File it with the questions the report owes.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>What it takes to serve it</h2><p>Napkin arithmetic first, ours and labeled as such, all of it downstream of the ~50B active estimate. </p><p>A single decode token&#8217;s forward pass touches 16 experts per MoE layer plus the shared trunk, on the order of 25 GB of weight reads at four bits. Against the <strong>roughly 8 TB/s of HBM bandwidth</strong> on a current Blackwell-class part, that is about three milliseconds of pure weight traffic, an upper bound near<strong> 320 tokens per second </strong>even if one device could hold the whole model, which it cannot. </p><p>The ratio is the point of the exercise: at FP16 the same pass would read 100 GB per token and the model would be bandwidth-strangled on any realistic hardware. Four-bit weights are <em>not an optimization here</em>. They are the enabling condition.</p><p>That 25 GB figure is also a batch-size-one number, and batch size one is the worst case this architecture has. Here is the mechanism, because it is the actual <strong>thesis of extreme sparsity</strong>. Dense-layer weights are read once per forward pass and amortize cleanly across every token in the batch.</p><p> Expert weights amortize only when tokens share experts, and with 16 of 896 routing, sharing takes scale. </p><p>Rather than assert the curve, we simulated it, sixteen trials per point, with two popularity models: uniform routing, which is what Quantile Balancing is trying to buy, and a moderately skewed distribution with a <strong>coefficient of variation around 0.7</strong>, which is what real routers produce when you let them:</p><pre><code><code># How expert weight traffic per token falls with batch size, and what routing
# skew does to it. E=896, k=16, ~1.5 GB per expert at MXFP4 (assumption, tier E).
import numpy as np
rng = np.random.default_rng(3)

E, K, GB_PER_EXPERT, TRIALS = 896, 16, 1.5, 16
batches = [1, 8, 64, 256, 1024, 4096]

def run(pop):                       # pop: expert popularity distribution, sums to 1
    uniq_gb, imb = [], []
    for B in batches:
        u, m = [], []
        for _ in range(TRIALS):
            draws = np.concatenate([rng.choice(E, K, replace=False, p=pop)
                                    for _ in range(B)])
            counts = np.bincount(draws, minlength=E)
            u.append((counts &gt; 0).sum())
            m.append(counts.max() / (B * K / E))          # max load over mean load
        uniq_gb.append(np.mean(u) * GB_PER_EXPERT / B)
        imb.append(np.mean(m))
    return uniq_gb, imb

uniform = np.full(E, 1 / E)
skewed = rng.dirichlet(np.full(E, 2.0))                    # cv ~0.7, mild hotness
g_u, i_u = run(uniform)
g_s, i_s = run(skewed)

print(f"{'batch':&gt;6} {'uniform GB/tok':&gt;15} {'imbalance':&gt;10} {'skewed GB/tok':&gt;14} {'imbalance':&gt;10}")
for j, B in enumerate(batches):
    print(f"{B:&gt;6} {g_u[j]:&gt;15.2f} {i_u[j]:&gt;10.1f} {g_s[j]:&gt;14.2f} {i_s[j]:&gt;10.1f}")
print(f"floor at large batch: {E * GB_PER_EXPERT:.0f} GB / B")
print("EP sizing at MXFP4: EP64 -&gt;", E // 64, "experts/GPU,", E // 64 * GB_PER_EXPERT,
      "GB | EP128 -&gt;", E // 128, "experts/GPU,", E // 128 * GB_PER_EXPERT, "GB")
</code></code></pre><p>Executed output:</p><pre><code><code> batch  uniform GB/tok  imbalance  skewed GB/tok  imbalance
     1           24.00       56.0          24.00       56.0
     8           22.72       14.9          21.82       18.8
    64           14.42        4.9          12.65        7.4
   256            5.20        2.7           4.80        5.5
  1024            1.31        1.8           1.30        4.6
  4096            0.33        1.4           0.33        4.6
floor at large batch: 1344 GB / B
EP sizing at MXFP4: EP64 -&gt; 14 experts/GPU, 21.0 GB | EP128 -&gt; 7 experts/GPU, 10.5 GB
</code></code></pre><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!RkV6!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F77019cbc-c703-4f89-8cff-24420681c437_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!RkV6!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F77019cbc-c703-4f89-8cff-24420681c437_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!RkV6!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F77019cbc-c703-4f89-8cff-24420681c437_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!RkV6!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F77019cbc-c703-4f89-8cff-24420681c437_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!RkV6!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F77019cbc-c703-4f89-8cff-24420681c437_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!RkV6!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F77019cbc-c703-4f89-8cff-24420681c437_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/77019cbc-c703-4f89-8cff-24420681c437_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Why extreme sparsity demands batch&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Why extreme sparsity demands batch" title="Why extreme sparsity demands batch" srcset="https://substackcdn.com/image/fetch/$s_!RkV6!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F77019cbc-c703-4f89-8cff-24420681c437_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!RkV6!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F77019cbc-c703-4f89-8cff-24420681c437_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!RkV6!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F77019cbc-c703-4f89-8cff-24420681c437_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!RkV6!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F77019cbc-c703-4f89-8cff-24420681c437_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 3. Why extreme sparsity demands batch.</figcaption></figure></div><p>Three things in that table are worth internalizing. First, the amortization cliff: per-token expert traffic barely moves from batch 1 to batch 8, because <strong>eight tokens drawing 128 expert slots</strong> land on about 120 distinct experts, which is to say almost no sharing at all. </p><p>It falls to 5 GB per token at batch 256, where the batch has touched nearly the whole 1,344 GB expert pool once, and approaches the 1,344 divided by B floor beyond a thousand concurrent tokens. </p><p>Second, the <strong>imbalance column,</strong> max expert load over mean, is the one your step time actually obeys, because the slowest device sets the pace of every synchronous layer.</p><p>Under uniform routing it decays toward 1.4 as the law of large numbers kicks in. Under skew it plateaus at about 4.6 and never recovers: the <strong>hot expert stays hot</strong>, no matter how big the batch gets. Third, notice the trade hiding between the columns: skew slightly reduces traffic, popular experts get reused, while wrecking balance. </p><p>Quantile Balancing, read through this table, is Moonshot choosing the uniform column, paying the <strong>worst-case traffic </strong>to buy the converging imbalance, because a 4.6x hot spot at expert parallelism width is a 4.6x tax on every token. </p><p>The structural conclusion stands either way: a 1.8 percent activation ratio is only cheap at large batch under wide expert parallelism, where every device holds a<strong> slice of the pool</strong>, reads its residents once per step, and serves whichever tokens the all-to-all delivers. </p><p>At EP64 that slice is 14 experts and 21 GB per GPU; at EP128, 7 experts and 10.5 GB, numbers that fit comfortably beside a trunk shard and a KV allocation. </p><p><strong>Extreme sparsity</strong> does not just permit big-batch serving. It demands it, which is why models shaped like this favor operators with deep pools of concurrent traffic and punish anyone trying to run them hot for three users.</p><p>The all-to-all itself can be sized on the back of the same napkin. Per MoE layer, each token&#8217;s hidden state ships to its 16 expert devices and 16 partial outputs ship back: roughly 2 x k x d_model bytes at FP8. </p><p><strong>K2&#8217;s hidden width was 7,168</strong> and K3&#8217;s is unpublished, so treat this as parameterized: 2 x 16 x 7,168 is about 229 KB per token per MoE layer, double K2&#8217;s 8-way routing at the same width. </p><p>Across an assumed 60 MoE layers that is roughly 14 MB of fabric traffic per generated token, and at <strong>15,000 tokens per second</strong> of aggregate throughput the expert-parallel group is moving on the order of 200 GB/s. Inside an NVLink domain at 1.8 TB/s per GPU that is background noise. </p><p>Stretched across racks on 400G links at 50 GB/s each, it is the budget. We have written before about the all-to-all tax and the rack-scale networking built to pay it; K3 is precisely the class of model that hardware exists for.</p><p>If the 2 x k x d term feels abstract, here is the entire MoE data path, route, dispatch, <strong>grouped GEMM</strong>, combine, in one page of numpy, checked against a dense reference:</p><pre><code><code># The whole MoE data path in one page: route, dispatch, grouped GEMM, combine.
# Toy sizes so it prints; the shapes are the ones a serving engine juggles.
import numpy as np
rng = np.random.default_rng(0)

T, D, E, K = 8, 16, 32, 4                  # tokens, hidden, experts, top-k
x = rng.normal(size=(T, D)).astype(np.float32)
W = rng.normal(size=(E, D, D)).astype(np.float32) * 0.1    # one matrix per expert
router = rng.normal(size=(D, E)).astype(np.float32)

logits = x @ router                                        # [T, E]
topk = np.argsort(-logits, axis=1)[:, :K]                  # [T, K] expert ids
gates = np.exp(logits[np.arange(T)[:, None], topk])
gates /= gates.sum(1, keepdims=True)                       # softmax over the k winners

# dispatch: replicate each token k times, then sort rows by destination expert
flat_expert = topk.ravel()                                 # [T*K]
flat_token  = np.repeat(np.arange(T), K)                   # [T*K]
order = np.argsort(flat_expert, kind="stable")             # the permutation
xp = x[flat_token[order]]                                  # [T*K, D] permuted buffer

# grouped GEMM: one segment per expert, ragged sizes
counts = np.bincount(flat_expert, minlength=E)
offsets = np.concatenate([[0], np.cumsum(counts)])
yp = np.empty_like(xp)
for e in range(E):                                         # in CUDA: one grouped kernel
    s, t = offsets[e], offsets[e + 1]
    if t &gt; s:
        yp[s:t] = xp[s:t] @ W[e]

# combine: unpermute and gate-weighted sum back to [T, D]
y = np.zeros_like(x)
np.add.at(y, flat_token[order], yp * gates.ravel()[order][:, None])

# reference: dense loop over tokens and their experts
y_ref = np.zeros_like(x)
for t in range(T):
    for j in range(K):
        y_ref[t] += gates[t, j] * (x[t] @ W[topk[t, j]])
print("matches dense loop:", np.allclose(y, y_ref, atol=1e-5))
print("permuted buffer rows:", xp.shape[0], f"({K}x the token count)")
bytes_per_tok = 2 * K * D * xp.itemsize
print(f"wire bytes per token, this layer: 2 x k x d x {xp.itemsize} = {bytes_per_tok}")
print("tokens per expert:", counts.tolist())
</code></code></pre><p>Executed output:</p><pre><code><code>matches dense loop: True
permuted buffer rows: 32 (4x the token count)
wire bytes per token, this layer: 2 x k x d x 4 = 512
tokens per expert: [2, 0, 1, 1, 0, 2, 0, 0, 2, 0, 1, 1, 1, 2, 0, 1, 2, 1, 1, 1, 0, 2, 0, 0, 2, 0, 0, 2, 2, 3, 1, 1]
</code></code></pre><p>Every serving-engine complication is already visible in the toy. The permuted buffer holds k times the token count, so top-16 doubles the activation memory and <strong>permutation traffic of a top-8 model</strong> before a single expert FLOP is spent. </p><p>The per-expert segments are ragged, 0 to 3 tokens here, thousands in production, which is why the <strong>expert matmul </strong>is a grouped GEMM with runtime-sized groups rather than a clean batched one. And the scatter-add in the combine is the reduction the fabric has to carry back. </p><p>Scale the token count by six orders of magnitude and the expert count by 28, and the sort, the offsets, and the segment loop become the router kernel, the dispatch layout, and the<strong> grouped GEMM</strong> that the serving stacks will spend the next quarter optimizing.</p><p>Memory next. Weights at ~1.49 TB fit inside a single 8x B200 node, 1,536 GB, with about <strong>50 GB to spare</strong>, and that sentence is a trap: the spare must hold every KV block, activation buffer, routing table, and graph on the box, so a one-node K3 is a demo, not a deployment. </p><p>Moonshot&#8217;s own launch guidance recommends a supernode of at least 64 accelerators, which community sizing reads as eight nodes of eight 80 GB GPUs, around<strong> 5 TB aggregate</strong>, a floor with headroom for cache and parallelism. </p><p>On a GB200 NVL72 the weights occupy roughly a tenth of the rack&#8217;s 13.4 TB of HBM. The DeepSeek 671B generation already made rack-scale expert parallelism a normal conversation; K3 raises the <strong>resident-byte requirement</strong> by roughly 4x&#8230;. and makes the full NVL72 look like the natural unit for a single replica rather than an extravagance.</p><p>The long-context economics deserve their own bytes, because the flat 1M pricing rests on them. In the DeepSeek V3 configuration that MLA descends from, each token stores a compressed latent per global layer: a <strong>512-wide KV latent</strong> plus a 64-wide decoupled positional key, 576 elements, 576 bytes at FP8. </p><p>If K3 carries something like the Kimi Linear ratio, call it 15 to 20 global layers out of the stack, a full million-token sequence holds roughly <strong>9 to 12 GB of KV</strong>, ours again and resting on the assumed ratio. A comparable dense model, 60 layers of grouped-query attention with 8 KV heads of dimension 128 at FP16, holds about 246 KB per token, 246 GB at the same window, per sequence. </p><p>That is a factor of about 25, and it is the difference between a million-token session being a scheduling problem and being a physical impossibility. The KDA layers contribute a fixed-size state regardless of length, which is a rounding error next to either number. </p><p>This is the <strong>architecture showing</strong> through the rate card: Moonshot can price the window flat because most of the model does not pay for it.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!M_21!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb59ad582-3fe1-44b7-8cb6-7ca45f4ab3f0_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!M_21!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb59ad582-3fe1-44b7-8cb6-7ca45f4ab3f0_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!M_21!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb59ad582-3fe1-44b7-8cb6-7ca45f4ab3f0_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!M_21!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb59ad582-3fe1-44b7-8cb6-7ca45f4ab3f0_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!M_21!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb59ad582-3fe1-44b7-8cb6-7ca45f4ab3f0_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!M_21!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb59ad582-3fe1-44b7-8cb6-7ca45f4ab3f0_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b59ad582-3fe1-44b7-8cb6-7ca45f4ab3f0_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;KV cache at the full window&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="KV cache at the full window" title="KV cache at the full window" srcset="https://substackcdn.com/image/fetch/$s_!M_21!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb59ad582-3fe1-44b7-8cb6-7ca45f4ab3f0_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!M_21!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb59ad582-3fe1-44b7-8cb6-7ca45f4ab3f0_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!M_21!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb59ad582-3fe1-44b7-8cb6-7ca45f4ab3f0_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!M_21!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb59ad582-3fe1-44b7-8cb6-7ca45f4ab3f0_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 4. KV cache at the full window.</figcaption></figure></div><p>Prefill is the other half, and it is why the serving system is disaggregated. Filling the full window on a cache miss costs $3.15 of input at the sticker rate, $0.31 on a hit, and the compute behind the miss is brutal. </p><p>Linear-layer FLOPs alone run <strong>about 2 x 50B x 1M tokens</strong>, on the order of 10^17, several seconds on a full H100 node at perfect efficiency and the better part of half a minute at realistic utilization, before the attention term, which grows quadratically on the global layers and dominates at this length even confined to a quarter of the stack. </p><p>The <strong>KDA layers</strong> prefill differently again: a chunked scan, parallel within each chunk with recurrent state carried across chunk boundaries, a kernel shape with nothing in common with FlashAttention. </p><p><em>Prefill at 1M</em> is a compute event; decode is a bandwidth event; the two want different hardware shapes and different batching. Moonshot serves K3 on Mooncake, its <strong>KVCache-centric</strong> disaggregated architecture, which splits prefill and decode across separate node pools and treats the cache as a first-class pooled resource between them. </p><p>Mooncake is not opaque vendor infrastructure: the design won the <strong>best paper award at FAST 25</strong>, the code is open-sourced, and its transfer engine has been integrated into the mainstream serving stacks, including vLLM and SGLang. </p><p>The company reports cache hit rates above 90 percent on coding workloads. That figure is vendor-reported, but it is plausible for agent traffic, where the <strong>same system prompt</strong>, repository context, and conversation prefix recur on every loop iteration, and the published architecture is engineered to produce exactly that number. </p><p>The hit rate is what makes the pricing model work, which brings us to the bill by way of the software that has to exist first.</p><blockquote><p><em>Extreme sparsity does not just permit big-batch serving. It demands it.</em></p></blockquote><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>What has to change in the serving stacks</h2><p>Self-hosting K3 is not a config file away, and it is worth being specific about where the engineering actually lands, because the July 27 story will be written in pull requests.</p><p>Start at the router. Every MoE layer produces logits of shape tokens by 896, takes a top-16, and must land tokens in per-expert segments fast enough <strong>not to shadow the GEMMs</strong>. </p><p>The grouped GEMM behind it now has up to 896 runtime-sized groups per layer per rank, and the tokens-per-expert variance the toy above showed becomes <strong>tile quantization waste</strong> at CUDA scale: an expert with 37 tokens still occupies whole tensor-core tiles. </p><p>The mature answers are persistent grouped kernels that pull ragged work from a queue instead of launching per group, and the CUTLASS grouped GEMM machinery most stacks already wrap. </p><p>What changes for K3 is scale, 896 groups against the 256 the current kernels were tuned around, and precision, because the groups now <strong>multiply block-scaled FP4 weights </strong>against FP8 activations.</p><p>Precision is where the hardware generations split. Blackwell&#8217;s fifth-generation tensor cores execute block-scaled MXFP4 natively, scales applied in the MMA itself, so B200-class serving runs the checkpoint as stored. </p><p>Hopper has no FP4 MMA path at all: an H100 deployment dequantizes weights to FP8 or BF16 in registers on the way into the matmul, the <strong>Marlin </strong>and <strong>Machete family </strong>of mixed-input kernels, which exist and are fast for dense weight-only GEMMs but are immature for 896-group grouped MoE at launch. </p><p>That dequant tax is not cosmetic. It is compute spent on format conversion inside the innermost loop, and it lands directly on the self-hosting ledger below: the 37 tokens per second per GPU that beats the sticker price assumes the<strong> grouped W4A8 path</strong> works well on Hopper, and today that is an assumption.</p><p>The communication layer has a shelf to raid but not a finished product. DeepSeek open-sourced DeepEP, NVSHMEM-based dispatch and combine kernels with FP8 payloads, <strong>IBGDA for the internode path</strong>, and a low-latency mode for decode, and it is the obvious starting point. </p><p>It was also built and tuned around the 256-expert, top-8 shape of the V3 generation. K3 doubles the per-token payload with top-16 and multiplies the destination fan-out with 896 experts, so the kernel launch geometry, the <strong>SM budget </strong>reserved <strong>for communication</strong>, and the overlap schedule that hides all-to-all latency behind expert compute all need retuning. </p><p>The dual-microbatch overlap trick, one microbatch&#8217;s communication hidden under another&#8217;s GEMMs, matters more here, not less, because the fabric term we sized above grows with k while the compute term holds near flat.</p><p>The subtlest problem is the cache manager, and it is the one we would watch first. Serving frameworks assume a homogeneous per-layer paged KV cache; the hybrid allocators built for the <strong>Mamba-generation models </strong>already broke that assumption, and K3 stresses it further: paged 576-byte latents for the global layers next to fixed-size recurrent state for the KDA layers. </p><p>Decode is straightforward. Prefix caching is not, because the entire cache-hit economy assumes you can resume from a stored prefix, and for a recurrent layer that means storing the state at the reuse boundary, not just the tokens. </p><p>A KV pool like Mooncake&#8217;s holds paged latents naturally; checkpointed KDA states at chunk boundaries are a second object type with different lifetime and size semantics, and <strong>how Moonshot handles state-restore</strong> for the linear layers is, to us, the single most interesting undisclosed detail in the serving stack. </p><p>Whoever implements it in the open stacks will be making a real design decision, not just a port.</p><p>The good news is how much is already on the shelf. The chunked delta-rule kernels for KDA shipped in the open with the Kimi Linear release through the flash-linear-attention line. </p><p>FlashMLA covers the global layers on Hopper. DeepEP covers the fabric, pending retuning. Mooncake&#8217;s transfer engine is merged where it needs to be. </p><p>What does not exist yet, anywhere public, is the integration: one engine that schedules block-scaled grouped GEMMs, hybrid attention with two kernel families, a two-type cache, and disaggregated state transfer, for a 1.5 TB checkpoint. </p><p>That is the <strong>artifact to watch</strong> for in the weeks after July 27, and its arrival date, more than the weights themselves, decides when open K3 is real.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>The bill</h2><p>The rate card, from Moonshot&#8217;s platform pages:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Xn2c!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa7260193-a4e1-4c92-a46a-995dc07fd591_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Xn2c!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa7260193-a4e1-4c92-a46a-995dc07fd591_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!Xn2c!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa7260193-a4e1-4c92-a46a-995dc07fd591_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!Xn2c!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa7260193-a4e1-4c92-a46a-995dc07fd591_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!Xn2c!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa7260193-a4e1-4c92-a46a-995dc07fd591_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Xn2c!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa7260193-a4e1-4c92-a46a-995dc07fd591_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a7260193-a4e1-4c92-a46a-995dc07fd591_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:81860,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/206014241?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa7260193-a4e1-4c92-a46a-995dc07fd591_1634x893.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Xn2c!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa7260193-a4e1-4c92-a46a-995dc07fd591_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!Xn2c!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa7260193-a4e1-4c92-a46a-995dc07fd591_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!Xn2c!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa7260193-a4e1-4c92-a46a-995dc07fd591_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!Xn2c!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa7260193-a4e1-4c92-a46a-995dc07fd591_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>All figures per million tokens. Web search tool calls bill separately at $0.004 each, though the tool ships flagged as being updated and one guide advises against production use for now. Vision input pricing is unconfirmed, per the vision section above.</em></p><p>The first thing the sticker tells you is positioning. K2 built its reputation as the cheap frontier model. K3 abandons that identity: at $3 and $15 it matches <strong>Claude Sonnet 5 </strong>to the cent, undercuts Opus 4.8 at $5 and $25 and GPT-5.6 Sol at $5 and $30, and costs several times its own K2.6 sibling.</p><p>Within the Chinese open cohort it is the expensive option, with DeepSeek V4 Pro at roughly a sixth of its input price and GLM 5.2 around half. Moonshot is telling you it believes K3 competes on capability, and it has priced away the discount narrative to prove it.</p><p>The second thing is that the sticker is not the effective price. The cache tier is. At a <strong>90 percent hit rate</strong>, effective input cost is 0.9 times $0.30 plus 0.1 times $3.00, about $0.57 per million tokens. </p><p>Agent workloads with stable prefixes live near that floor. Caching is automatic, with <strong>no cache ID or TTL management</strong>, which removes the usual engineering excuse for missing it, and OpenRouter&#8217;s traffic view says the same thing from the demand side: average realized K3 prices run 60 to 80 percent below list once caching is counted. </p><p>Two more levers are worth knowing. K3 ships with no batch tier, the 60 percent batch discount on Moonshot&#8217;s menu covers only the K2 line, and the API defaults max_completion_tokens to 131,072, configurable up to the full window, a cap worth leaving in place until you have watched what max-effort thinking does to a bill.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!3nL3!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe5513bb3-2038-4118-b06b-ec2249018314_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!3nL3!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe5513bb3-2038-4118-b06b-ec2249018314_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!3nL3!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe5513bb3-2038-4118-b06b-ec2249018314_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!3nL3!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe5513bb3-2038-4118-b06b-ec2249018314_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!3nL3!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe5513bb3-2038-4118-b06b-ec2249018314_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!3nL3!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe5513bb3-2038-4118-b06b-ec2249018314_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e5513bb3-2038-4118-b06b-ec2249018314_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;K3 effective input price vs cache hit rate&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="K3 effective input price vs cache hit rate" title="K3 effective input price vs cache hit rate" srcset="https://substackcdn.com/image/fetch/$s_!3nL3!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe5513bb3-2038-4118-b06b-ec2249018314_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!3nL3!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe5513bb3-2038-4118-b06b-ec2249018314_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!3nL3!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe5513bb3-2038-4118-b06b-ec2249018314_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!3nL3!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe5513bb3-2038-4118-b06b-ec2249018314_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 5. K3 effective input price vs cache hit rate.</figcaption></figure></div><p>The third thing cuts the other way. K3 always reasons, at max effort, and reasoning tokens bill as output at $15. <strong>Artificial Analysis measured 130 million output tokens </strong>to run its Intelligence Index evaluation, against a median near 63 million for comparable reasoning models. </p><p>The same measurements put decode speed at 62 tokens per second, under the 72.7 median for the price tier, and split latency into the two numbers an always-thinking model needs: about 2 seconds to the first streamed token, 34.24 seconds to the first answer token, which at the measured decode rate is roughly 2,000 thinking tokens spent before the answer begins on a standard 10,000-token workload. </p><p>Running the full index cost $2,709.75, nearly all of it output. Independent testing has recorded over thirteen thousand thinking tokens on a single short task, roughly twenty cents of deliberation before the first answer token. </p><p>Artificial Analysis puts the blended cost at $2.31 per million tokens on a 7:2:1 cache, input, output mix, but your mix will vary, and a maximally verbose reasoner shifts the mix toward the expensive column. &#8220;<em>Price it on your own traces</em>&#8221; is easy to say obv, so here is the whole model in twenty lines, run against three workload archetypes at launch rates:</p><pre><code><code># Request-level cost model for kimi-k3 at launch pricing.
# $3/MTok input on miss, $0.30 on cache hit, $15/MTok output.
# Thinking is always on and bills as output.
IN_MISS, IN_HIT, OUT = 3.00, 0.30, 15.00

def request_cost(in_tokens, cached_frac, answer_tokens, thinking_tokens):
    inp = in_tokens * ((1 - cached_frac) * IN_MISS + cached_frac * IN_HIT)
    out = (answer_tokens + thinking_tokens) * OUT
    return inp / 1e6, out / 1e6

archetypes = [
    # name, input, cached share, answer, thinking, requests
    ("short chat turn",      2_000, 0.20,   300, 13_000, 1),
    ("agent loop, 40 iters", 80_000, 0.95,   500,  6_000, 40),
    ("1M-doc analysis",     900_000, 0.00, 2_000, 10_000, 1),
]
print(f"{'workload':&lt;22}{'$ input':&gt;9}{'$ output':&gt;9}{'$ total':&gt;9}{'output share':&gt;14}")
for name, i, c, a, th, n in archetypes:
    ci, co = request_cost(i, c, a, th)
    ci, co = ci * n, co * n
    print(f"{name:&lt;22}{ci:&gt;9.3f}{co:&gt;9.3f}{ci+co:&gt;9.2f}{co/(ci+co):&gt;13.0%}")
</code></code></pre><p>Executed output:</p><pre><code><code>workload                $ input $ output  $ total  output share
short chat turn           0.005    0.200     0.20          98%
agent loop, 40 iters      1.392    3.900     5.29          74%
1M-doc analysis           2.700    0.180     2.88           6%
</code></code></pre><p>The <strong>distribution of pain</strong> is the finding. On a small chat turn, deliberation is 98 percent of the bill: the answer costs half a cent and the thinking costs twenty. </p><p>A forty-iteration agent loop, the workload K3 is aimed at, still spends three of every four dollars on output even with a 95 percent cache hit rate doing everything right on the input side. Only the <strong>cold million-token document</strong> flips the ratio, and that shape pays $2.70 of prefill precisely once, after which it becomes an agent loop too. </p><p>The lever ordering falls out directly: for two of the three archetypes that matter, the promised low and high effort modes are worth more than any cache optimization you can do, and until they ship, the K2.6 rate card at $0.95 and $4 remains the right tool for every request that does not need a frontier model deliberating at full depth.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Pf9j!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F798b78b6-a19c-47fd-ba61-e468e0d66318_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Pf9j!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F798b78b6-a19c-47fd-ba61-e468e0d66318_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!Pf9j!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F798b78b6-a19c-47fd-ba61-e468e0d66318_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!Pf9j!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F798b78b6-a19c-47fd-ba61-e468e0d66318_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!Pf9j!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F798b78b6-a19c-47fd-ba61-e468e0d66318_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Pf9j!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F798b78b6-a19c-47fd-ba61-e468e0d66318_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/798b78b6-a19c-47fd-ba61-e468e0d66318_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Where the money goes by workload&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Where the money goes by workload" title="Where the money goes by workload" srcset="https://substackcdn.com/image/fetch/$s_!Pf9j!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F798b78b6-a19c-47fd-ba61-e468e0d66318_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!Pf9j!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F798b78b6-a19c-47fd-ba61-e468e0d66318_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!Pf9j!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F798b78b6-a19c-47fd-ba61-e468e0d66318_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!Pf9j!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F798b78b6-a19c-47fd-ba61-e468e0d66318_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 6. Where the money goes by workload.</figcaption></figure></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>The self-hosting ledger</h2><p>The weight release invites an obvious question: </p><blockquote><p><em>at what point does running K3 yourself beat paying Moonshot? </em></p></blockquote><p>The napkin again, ours. Take the community floor configuration, 64 H100-class GPUs, and a rental price of $2 per GPU-hour: $128 per hour for the cluster. </p><p>To generate tokens cheaper than the $15 output rate, the cluster must sustain about 2,370 tokens per second aggregate, or 37 per GPU. That is an achievable number for a<strong> well-tuned MoE deployment</strong> with real batch depth, with the caveat the kernel section earned: on Hopper it assumes the grouped W4A8 dequant path matures, because every register spent converting FP4 is throughput the break-even does not get. </p><p>To beat the $2.31 blended rate that cached API traffic actually pays, the same cluster must sustain <strong>about 15,400 tokens per second</strong>, 240 per GPU, which is a very hard number for a model this size on Hopper-generation hardware. </p><p>Independent color corroborates the memory wall from an unexpected direction: an early community stress test found a 1.5 TB unified-memory Mac Studio cluster at the edge of what a K3-class million-token stack can carry, which is the same wall the 1.49 TB storage figure predicts.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!BdUF!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0424b430-a537-408b-9641-07fff9bfd4ea_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!BdUF!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0424b430-a537-408b-9641-07fff9bfd4ea_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!BdUF!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0424b430-a537-408b-9641-07fff9bfd4ea_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!BdUF!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0424b430-a537-408b-9641-07fff9bfd4ea_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!BdUF!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0424b430-a537-408b-9641-07fff9bfd4ea_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!BdUF!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0424b430-a537-408b-9641-07fff9bfd4ea_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0424b430-a537-408b-9641-07fff9bfd4ea_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Self-hosting K3: 64x H100 at $2 per GPU-hour&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Self-hosting K3: 64x H100 at $2 per GPU-hour" title="Self-hosting K3: 64x H100 at $2 per GPU-hour" srcset="https://substackcdn.com/image/fetch/$s_!BdUF!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0424b430-a537-408b-9641-07fff9bfd4ea_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!BdUF!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0424b430-a537-408b-9641-07fff9bfd4ea_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!BdUF!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0424b430-a537-408b-9641-07fff9bfd4ea_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!BdUF!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0424b430-a537-408b-9641-07fff9bfd4ea_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 7. Self-hosting K3: 64x H100 at $2 per GPU-hour.</figcaption></figure></div><p>The comparison is loose by construction: the API&#8217;s $15 covers output alone while your cluster serves whole requests, the blended target moves with your cache profile, and <strong>$2 per H100-hour</strong> is a spot figure that swings with region and term. </p><p>The asymmetry survives all of that. Self-hosting K3 competes with Moonshot&#8217;s sticker prices and loses badly to Moonshot&#8217;s cache economics, because the cache discount is not a margin decision you can copy, it is a serving architecture, <strong>Mooncake&#8217;s pooled KV </strong>plus a workload with 90 percent prefix reuse, that a single-tenant cluster cannot replicate without the traffic to feed it. </p><p>The teams for whom self-hosting pencils out are the ones who need it for reasons the rate card does not price: data residency (<em>the hosted API processes in Singapore, per an EU procurement review</em>), fine-tuned variants, latency floors, or sovereignty. </p><p>For everyone else, the July 27 download is leverage in a pricing negotiation, not a cost reduction.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Benchmarks, with the usual caution</h2><p>Moonshot&#8217;s own evaluation places K3 at frontier level but behind the two strongest closed models, Claude Fable 5 and GPT-5.6 Sol, an admission worth respecting because labs rarely volunteer it. </p><p><strong>On GDPval-AA v2</strong>, the aggregate knowledge-work evaluation, the reported ordering is Fable 5 Max at 1815, GPT-5.6 Sol Max at 1747.8, K3 at 1687, and Opus 4.8 at 1600. On the science and multimodal side, Moonshot reports <em>93.5 on GPQA-Diamond</em>, <strong>81.6 on MMMU-Pro</strong>, and 94.3 on MathVision.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!oLLt!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F724f72f0-1247-4494-a8bf-1965da6f2f7e_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!oLLt!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F724f72f0-1247-4494-a8bf-1965da6f2f7e_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!oLLt!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F724f72f0-1247-4494-a8bf-1965da6f2f7e_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!oLLt!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F724f72f0-1247-4494-a8bf-1965da6f2f7e_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!oLLt!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F724f72f0-1247-4494-a8bf-1965da6f2f7e_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!oLLt!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F724f72f0-1247-4494-a8bf-1965da6f2f7e_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/724f72f0-1247-4494-a8bf-1965da6f2f7e_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;GDPval-AA v2 as reported&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="GDPval-AA v2 as reported" title="GDPval-AA v2 as reported" srcset="https://substackcdn.com/image/fetch/$s_!oLLt!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F724f72f0-1247-4494-a8bf-1965da6f2f7e_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!oLLt!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F724f72f0-1247-4494-a8bf-1965da6f2f7e_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!oLLt!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F724f72f0-1247-4494-a8bf-1965da6f2f7e_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!oLLt!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F724f72f0-1247-4494-a8bf-1965da6f2f7e_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 8. GDPval-AA v2 as reported.</figcaption></figure></div><p>The coding numbers are where it gets interesting. K3 posts 88.3 on Terminal-Bench 2.1, half a point behind GPT-5.6 Sol, and takes the top score outright on Program Bench at 77.8 and on SWE Marathon at 42.0. SWE Marathon is the<strong> long-horizon one</strong>, and a model leading there while merely placing elsewhere is consistent with the design: a million-token window that holds an entire repository plus the full history of a long session is worth more on sustained work than on sprints. </p><p>FrontierSWE lands at 81.2, against a reported 86.6 for Fable 5 on the same benchmark, and DeepSWE at 67.5.</p><p>On agentic evaluations, VentureBeat&#8217;s read of the launch documentation has K3 first in four of eight real-world automation benchmarks, including Automation Bench, SpreadsheetBench 2, and BrowseComp, and second to Fable 5 in most of the rest. </p><p>The reported scores include 91.2 on BrowseComp, a 95.0 F1 on DeepSearchQA, and 30.8 on Automation Bench, which tells you as much about Automation Bench as about the model. </p><p>The claim we find most interesting is methodological: Moonshot says these results came from a <strong>single-agent setup </strong>running on the raw million-token window, with no context compression or management scaffolding (<em>BrowseComp reads 90.4 in that configuration against the 91.2 headline figure, both vendor numbers</em>). </p><p>Taken at face value, that is a real data point in the context-versus-orchestration argument, raw window plus strong retrieval beating elaborate<strong> multi-agent plumbing,</strong> and it needs independent replication before it graduates from press release to argument.</p><p>Two launch-week numbers deserve a somewhat explicit skepticism. K3 jumped from rank 18 to rank 1 on<strong> LMArena&#8217;s Frontend Code Arena</strong> within hours, at a score of 1679. </p><p>The placement itself has real sourcing, the Associated Press reported K3 at the top of the frontend ranking, but arena scores during a hype cycle measure attention as much as ability; check back in a month. </p><p>BenchLM&#8217;s weighted leaderboard still lacked a K3 row at press time, an absence worth noting, but the first independent aggregate has landed: <strong>Artificial Analysis scores K3 at 57</strong> on its Intelligence Index, fourth of 189 models tracked, against an average of 31 for comparable models, and the index rolls up nine evaluations including GDPval-AA v2, Terminal-Bench 2.1, GPQA Diamond, and Humanity&#8217;s Last Exam. </p><p>Frontier company, by an outside ruler. The 2.5x scaling efficiency figure, the 90 percent cache hit rate, and the single-agent protocol are all vendor-reported. <strong>None of this means the numbers are wrong</strong>. It means the confirmations are scheduled for<strong> after July 27.</strong></p><p>One demo is worth relaying because it is verifiable in kind if not yet in fact. Moonshot reports that K3, in a single 48-hour autonomous run, designed a serving chip for a nano model on its own architecture using <strong>open-source EDA tools</strong> against the Nangate 45nm library: four square millimeters, timing closed at 100 MHz, 1.46 million standard cells, 0.277 MB of SRAM, an INT4 multiply-accumulate array with fused dequantization, and a simulated decode throughput above 8,700 tokens per second. A staged demo, obviously, and simulation is not silicon. </p><p>But as a choice of demo it is telling: the lab wanted to show long-horizon agency on exactly the kind of task this audience does for a living. </p><p>A second staged demo aims at researchers: to reproduce the I-Love-Q universal relations in computational astrophysics, K3 reportedly cross-validated <strong>more than twenty papers</strong>, evaluated over three hundred equations of state, caught inconsistencies in published formulas, and shipped three thousand lines of Python plus an interactive dashboard in about two hours, against Moonshot&#8217;s estimate of one to two weeks for an experienced researcher. </p><p>Vendor-staged, like the chip, and chosen with the same message: long horizons are the product.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Where it strains</h2><p>Moonshot names three weaknesses, and all three have operational consequences.</p><p>The first is thinking history sensitivity: agent harnesses that truncate or rewrite the model&#8217;s reasoning traces degrade quality significantly. Launch coverage adds the reason: K3 was trained in a <strong>preserved thinking history mode</strong>, so a harness that drops the thinking, or a session handed mid-flight from another model, is operating outside the training distribution. </p><p>For agent builders this is a real constraint, and it compounds with the cache economics, because the two failure modes share a root cause: mutating the transcript. </p><p>Strip the thinking and quality drops; touch anything upstream in the serialized prefix and you<strong> lose the $0.30 tier </strong>and re-prefill from the perturbation point. </p><p>The defense against both is the same discipline, an append-only history with a verifiable prefix, cheap enough to enforce in a dozen lines:</p><pre><code><code># K3 penalizes two things at once: perturbing the prompt prefix (you lose the
# $0.30 cache tier) and truncating its reasoning traces (quality degrades).
# Same defense for both: an append-only history with a verifiable prefix.
import hashlib, json

def fingerprint(messages):
    blob = json.dumps(messages, sort_keys=True, separators=(",", ":"))
    return hashlib.sha256(blob.encode()).hexdigest()[:16]

def prefix_intact(old, new):
    return len(new) &gt;= len(old) and fingerprint(new[:len(old)]) == fingerprint(old)

history = [
    {"role": "system", "content": "You are a code review agent."},
    {"role": "user", "content": "Review PR 4217."},
    {"role": "assistant", "content": "&lt;thinking&gt;...&lt;/thinking&gt; Two issues found."},
]
appended  = history + [{"role": "user", "content": "Fix the first one."}]
truncated = [history[0], history[1],
             {"role": "assistant", "content": "Two issues found."}]  # thinking stripped
reordered = [history[1], history[0], history[2]]                     # system moved

for name, h in (("append-only", appended), ("thinking stripped", truncated),
                ("system reordered", reordered)):
    print(f"{name:&lt;20} prefix intact: {prefix_intact(history, h)}")

# Client shape (not executed here; the endpoint is OpenAI-compatible):
#   client = OpenAI(base_url="https://api.moonshot.ai/v1", api_key=KEY)
#   r = client.chat.completions.create(model="kimi-k3", messages=history)
#   history.append({"role": "assistant", "content": r.choices[0].message.content})
</code></code></pre><p>Executed output:</p><pre><code><code>append-only          prefix intact: True
thinking stripped    prefix intact: False
system reordered     prefix intact: False
</code></code></pre><p>The two False rows are the two expensive mistakes. A harness that &#8220;<em>helpfully</em>&#8221; compacts old reasoning fails the quality constraint and the cache constraint in the same commit; one that <strong>rebuilds the message list per call</strong>, reordering tools or regenerating timestamps, silently pays full prefill price on every iteration. </p><p>Fingerprint the prefix in development, assert it in production, and design summarization as a new branch rather than an edit to history. On long sessions where the window genuinely fills, the <strong>append-only rule </strong>forces the honest version of the tradeoff: fork the conversation with an explicit summary and eat one cold prefill, rather than quietly degrading the model mid-branch.</p><p>The second weakness is overeagerness. In ambiguous situations K3 tends to act rather than ask. On benchmarks that rewards decisiveness. In production it is <strong>how an agent cheerfully migrates the wrong database</strong>. The mitigation is boring and necessary: explicit confirmation gates in the harness, because the model will not supply them.</p><p>The third is polish. Moonshot concedes the subjective experience trails Fable 5 and GPT-5.6 Sol even where benchmark numbers are close. Benchmark parity and product parity remain different finish lines.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share The Software Frontier&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share The Software Frontier</span></a></p><div><hr></div><h2>What July 27 decides</h2><p>The weight release is the actual event; July 16 was the trailer.</p><p>The license text comes first. Modified MIT is reported across the launch coverage and would match the K2 family precedent, and the precedent has teeth worth knowing: the <strong>modified MIT </strong>that ships with K2.7 Code adds an attribution requirement that triggers at 100 million monthly active users or 20 million dollars of monthly revenue. </p><p>If K3 inherits the clause it is a display obligation for giants rather than a commercial restriction, but the binding text is the <strong>LICENSE file</strong> that lands with the weights, and nobody should sign a deployment plan before reading it.</p><p>Day-one inference support is the second gate, and the stacks section above is the checklist: block-scaled grouped GEMM at 896 groups, retuned dispatch and combine at top-16, the two-type cache with KDA state restore, and a Hopper W4A8 path that does not bleed the break-even dry. </p><p>Moonshot says it is working with inference partners ahead of the drop, and the <strong>Mooncake integrations </strong>already sitting in those codebases suggest the relationship is real; the commits will tell.</p><p>Third, reproduction. Someone outside Moonshot needs to run the benchmark suite, and the single-agent long-context protocol specifically.</p><p>Fourth, the research surface the checkpoint opens. With 896 experts in the open, expert specialization can finally be probed at frontier scale.<strong> Pruning experiments </strong>will ask whether the pool compresses to something that fits a smaller cluster, and what the quality curve looks like on the way down.</p><p>AttnRes can be ablated at small scale to see whether it generalizes or only pays at depth. And someone will attempt the <strong>first LoRA on a natively FP4 base </strong>and discover what breaks.</p><p>Then, the field is not holding still. The DeepSeek V4 line is already live at a fraction of K3&#8217;s price, and community reporting has a V4 refresh in staged testing; <strong>GLM 5.5 and MiniMax Pro </strong>are both expected around the trillion-parameter mark, and Qwen 4 is on the horizon. Whatever lead K3 holds is measured in weeks, and DeepSeek&#8217;s answer is the one to watch.</p><p>Our read, calmly: K3 does not win the frontier. By Moonshot&#8217;s own tables it sits a step behind Fable 5 and Sol. What it breaks is the assumption that open weights trail the frontier by six months and a capability class. </p><p>The gap is now single-digit percentage points and a dated download link, at Sonnet prices, on an architecture whose every major choice, the sparsity ratio, the hybrid attention, the <strong>four-bit training</strong>, the disaggregated serving, was made with the inference bill in mind. The link goes live on July 27, and that is where the work starts.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Software Frontier is a reader-supported publication. Consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p><div><hr></div><h2>Sources and confidence</h2><p><strong>A, primary:</strong> Moonshot&#8217;s K3 launch blog and platform documentation: release dates, architecture component names, context window, pricing, stated limitations, chip design demo, benchmark tables. Published background used for parameterized math and the serving-stack analysis: the OCP Microscaling format specification, the Kimi Linear technical report (<em>KDA design, the three to one hybrid ratio, KV reduction and decode speedup figures, the open kernel release</em>), the Mooncake paper and open-source release (FAST 25 best paper; transfer engine integrated in mainstream serving stacks), the DeepSeek V3 MLA configuration (512 plus 64 latent dimensions), and the DeepEP communication library.</p><p><strong>B, independent measurement:</strong> Artificial Analysis: Intelligence Index score of 57 (fourth of 189), output token consumption on the index run (130M vs a 63M median), decode speed of 62 tokens per second against a 72.7 tier median, latency of about 2 seconds to first token and 34.24 seconds to first answer token, evaluation cost of $2,709.75, and blended cost ($2.31/M at 7:2:1). OpenRouter: listing at matching rates, and platform-reported realized prices 60 to 80 percent below list under caching.</p><p><strong>C, informed secondary:</strong> VentureBeat on benchmark placements, release timing, and corporate context. The Hugging Face community model overview for the ~50B active parameter estimate, MXFP4 and MXFP8 details, and deployment sizing. Verdent and multiple pricing guides for the comparative rate card and the modified MIT report.</p><p><strong>D, vendor-claimed, unverified:</strong> the 2.5x scaling efficiency over K2, the 90 percent cache hit rate, the single-agent no-compression benchmark protocol, launch-week arena rankings, and the chip demo results.</p><p><strong>E, our arithmetic and code:</strong> the batch amortization simulation, the quantization grid experiment, the dispatch reference implementation, the cost model, the prefix fingerprint demo, all-to-all sizing, KV cache bytes, prefill FLOPs, and the self-host break-even. Every code block in this issue was executed as printed, CPython 3.12.3 with numpy 2.4.4, with fixed seeds; outputs are shown verbatim. Assumptions are stated inline; every figure downstream of the ~50B active estimate, the assumed 7,168 hidden width, the ~60 MoE layer count, the ~1.5 GB per expert, and the assumed global-layer ratio inherits their uncertainty and will be recomputed when the technical report publishes the real dimensions.</p><p>The Kimi K3 technical report had not been published at the time of writing. Numbers in this issue may be revised when it appears, and we will follow up after the July 27 weight release.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/p/kimi-k3-28-trillion-parameters-four?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/p/kimi-k3-28-trillion-parameters-four?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><div><hr></div><h2>The &#8220;fact check&#8221; ledger</h2><p>Every load-bearing claim, its source tier, and how it was checked. Rows marked recomputed were re-derived at build time by the scripts shipped with this issue.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!K-_1!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6df8a002-7fa0-40c9-84c4-8877539496ec_2320x2122.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!K-_1!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6df8a002-7fa0-40c9-84c4-8877539496ec_2320x2122.png 424w, https://substackcdn.com/image/fetch/$s_!K-_1!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6df8a002-7fa0-40c9-84c4-8877539496ec_2320x2122.png 848w, https://substackcdn.com/image/fetch/$s_!K-_1!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6df8a002-7fa0-40c9-84c4-8877539496ec_2320x2122.png 1272w, https://substackcdn.com/image/fetch/$s_!K-_1!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6df8a002-7fa0-40c9-84c4-8877539496ec_2320x2122.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!K-_1!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6df8a002-7fa0-40c9-84c4-8877539496ec_2320x2122.png" width="1456" height="1332" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6df8a002-7fa0-40c9-84c4-8877539496ec_2320x2122.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1332,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:444890,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/206014241?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6df8a002-7fa0-40c9-84c4-8877539496ec_2320x2122.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!K-_1!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6df8a002-7fa0-40c9-84c4-8877539496ec_2320x2122.png 424w, https://substackcdn.com/image/fetch/$s_!K-_1!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6df8a002-7fa0-40c9-84c4-8877539496ec_2320x2122.png 848w, https://substackcdn.com/image/fetch/$s_!K-_1!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6df8a002-7fa0-40c9-84c4-8877539496ec_2320x2122.png 1272w, https://substackcdn.com/image/fetch/$s_!K-_1!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6df8a002-7fa0-40c9-84c4-8877539496ec_2320x2122.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>References</h2><p>Primary and background sources. <strong>arXiv identifiers</strong> for the Kimi Linear and <strong>Gated DeltaNet </strong>entries were verified against arXiv during fact-checking; the remaining identifiers are standard citations for their papers.</p><ol><li><p><em>Moonshot AI. Kimi K3 Tech Blog: Open Frontier Intelligence. kimi.com/blog/kimi-k3, July 2026.</em></p></li><li><p><em>Moonshot AI. Kimi K3 quickstart and platform pricing. platform.kimi.ai, July 2026.</em></p></li><li><p><em>Kimi Team. Kimi Linear: An Expressive, Efficient Attention Architecture. arXiv:2510.26692, 2025.</em></p></li><li><p><em>Yang, S. et al. Gated Delta Networks: Improving Mamba2 with Delta Rule. arXiv:2412.06464, 2024.</em></p></li><li><p><em>Yang, S. et al. Parallelizing Linear Transformers with the Delta Rule over Sequence Length. arXiv:2406.06484, 2024.</em></p></li><li><p><em>DeepSeek-AI. DeepSeek-V3 Technical Report. arXiv:2412.19437, 2024.</em></p></li><li><p><em>DeepSeek-AI. DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model. arXiv:2405.04434, 2024.</em></p></li><li><p><em>Qin, R. et al. Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving. arXiv:2407.00079; USENIX FAST 2025, best paper.</em></p></li><li><p><em>Rouhani, B. et al. Microscaling Data Formats for Deep Learning. arXiv:2310.10537, 2023. See also the OCP Microscaling Formats Specification v1.0.</em></p></li><li><p><em>Micikevicius, P. et al. FP8 Formats for Deep Learning. arXiv:2209.05433, 2022.</em></p></li><li><p><em>Shazeer, N. et al. Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer. arXiv:1701.06538, 2017.</em></p></li><li><p><em>Lepikhin, D. et al. GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding. arXiv:2006.16668, 2020.</em></p></li><li><p><em>Fedus, W. et al. Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity. arXiv:2101.03961, 2021.</em></p></li><li><p><em>Wang, L. et al. Auxiliary-Loss-Free Load Balancing Strategy for Mixture-of-Experts. arXiv:2408.15664, 2024.</em></p></li><li><p><em>Kimi Team. Kimi K2: Open Agentic Intelligence. arXiv:2507.20534, 2025.</em></p></li><li><p><em>Liu, J. et al. Muon is Scalable for LLM Training. arXiv:2502.16982, 2025.</em></p></li><li><p><em>Ainslie, J. et al. GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints. arXiv:2305.13245, 2023.</em></p></li><li><p><em>Kwon, W. et al. Efficient Memory Management for Large Language Model Serving with PagedAttention. arXiv:2309.06180, 2023.</em></p></li><li><p><em>Zheng, L. et al. SGLang: Efficient Execution of Structured Language Model Programs. arXiv:2312.07104, 2023.</em></p></li><li><p><em>Open-source repositories referenced: MoonshotAI/Kimi-Linear, fla-org/flash-linear-attention, deepseek-ai/DeepEP, deepseek-ai/FlashMLA, kvcache-ai/Mooncake, on GitHub.</em></p></li><li><p><em>Launch-week coverage and trackers cited in text: VentureBeat, the Hugging Face community overview, Artificial Analysis, OpenRouter, BenchLM, and the pricing and readiness guides from Verdent, eesel, aireiter, kie.ai, avenchat, wan27.org, NxCode, digitalapplied, and TECHi, all July 2026.</em></p></li></ol>]]></content:encoded></item><item><title><![CDATA[Exploring how MLIR works: the compiler rewiring the AI stack ]]></title><description><![CDATA[From tensor graphs to machine code, MLIR is quietly becoming the abstraction layer connecting modern AI frameworks to increasingly specialized hardware.]]></description><link>https://www.thesoftwarefrontier.com/p/exploring-how-mlir-works-the-compiler</link><guid isPermaLink="false">https://www.thesoftwarefrontier.com/p/exploring-how-mlir-works-the-compiler</guid><dc:creator><![CDATA[Lorenzo Bradanini]]></dc:creator><pubDate>Wed, 15 Jul 2026 17:34:55 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!auv1!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F069655fa-6217-4d2d-8889-c97eba300015_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!auv1!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F069655fa-6217-4d2d-8889-c97eba300015_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!auv1!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F069655fa-6217-4d2d-8889-c97eba300015_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!auv1!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F069655fa-6217-4d2d-8889-c97eba300015_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!auv1!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F069655fa-6217-4d2d-8889-c97eba300015_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!auv1!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F069655fa-6217-4d2d-8889-c97eba300015_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!auv1!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F069655fa-6217-4d2d-8889-c97eba300015_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/069655fa-6217-4d2d-8889-c97eba300015_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2337295,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/203533021?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F069655fa-6217-4d2d-8889-c97eba300015_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!auv1!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F069655fa-6217-4d2d-8889-c97eba300015_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!auv1!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F069655fa-6217-4d2d-8889-c97eba300015_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!auv1!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F069655fa-6217-4d2d-8889-c97eba300015_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!auv1!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F069655fa-6217-4d2d-8889-c97eba300015_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><p><strong>Qualcomm</strong> just agreed to pay roughly $3.9 billion for a 150 person compiler company. Tesla says it rebuilt the FSD compiler and runtime on the same infrastructure. NVIDIA&#8217;s newest Python kernel DSL compiles through it. That infrastructure is <strong>MLIR</strong>. </p><p>This is the full tour: the object model, the dialect system, one<strong> MLP block lowered</strong> step by step into genuine <strong>sm_90 PTX</strong>, a custom dialect built from scratch, and the economics of who now controls the narrowest point in the AI software stack.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Three headlines, one substrate</h2><p>On June 24, 2026, Qualcomm announced it would acquire <strong>Modular</strong>, the AI infrastructure company founded in 2022 by Chris Lattner and Tim Davis. </p><p>Qualcomm&#8217;s own press release gives no price; Reuters and the Wall Street Journal reported an all-stock deal worth <em>approximately $3.9 billion</em>, expected to close in the second half of 2026. Modular employs roughly 150 people, per the industry analysts at <strong>NAND Research. </strong></p><p>Do the arithmetic and Qualcomm is paying something like $26 million per person for a company whose flagship language, <strong>Mojo</strong>, reached its first 1.0 beta seven weeks before the announcement, and whose compiler is not even open source yet. <strong>Qualcomm CEO</strong> Cristiano Amon justified it in one line: developers, he said, &#8220;<em>demand a more open and modern software foundation.&#8221;</em></p><p>Two months earlier, in April 2026, <strong>Tesla shipped FSD v14.3.</strong> Buried in the release notes, as documented by the release trackers at Tesla Oracle, was the claim that the AI compiler and runtime had been rewritten &#8220;<em>from the ground up with MLIR</em>&#8221;, with Tesla attributing a <strong>20 percent improvement</strong> in vehicle reaction time to the new stack. </p><p>And a year before that, in May 2025, NVIDIA shipped CUTLASS 4 with something unprecedented for the company: <strong>CuTe DSL</strong>, a Python interface for authoring peak-performance GPU kernels which, per NVIDIA&#8217;s own documentation, are <strong>JIT compiled</strong> through MLIR and then handed to ptxas.</p><p>A $3.9 billion acquisition by a mobile silicon giant. A safety-critical automotive inference stack. The kernel library that NVIDIA itself uses to demonstrate what Blackwell can do. </p><p>Three announcements, three companies with three entirely different agendas, one common substrate underneath. That<strong> substrate is MLIR</strong>, and if you work anywhere near GPUs, inference serving, or the economics of accelerated compute, you are running on top of it whether you know it or not.</p><p>In the previous installment of this series, <em>How LLVM Works: The IR That Took Over Modern Computing</em>, we told the story of the intermediate representation that <strong>unified the CPU world</strong>: one typed SSA language in the middle, every source language on one side, every instruction set on the other. </p><p>This piece is about what happened when that model met machine learning and broke, and about the infrastructure Lattner&#8217;s team built in response. MLIR is routinely described as &#8220;<em>LLVM for AI.</em>&#8221; It is not. </p><p>It is something stranger and <strong>more consequential</strong>: not a compiler, but a machine for building compilers, and it has quietly become the coordination point for the entire AI hardware buildout.</p><h4><span>How to read the code in this piece</span></h4><p><em>Every block of IR, TableGen, and C++ below was run through real tools before publication: </em><code>mlir-opt</code><em>, </em><code>mlir-runner</code><em>, </em><code>mlir-tblgen</code><em>, and GCC against the MLIR 20 headers, all from the stock LLVM 20.1.2 packages on Ubuntu 24.04. </em></p><p>A <strong><span>verified</span></strong> badge means the snippet parses and passes the verifier; <strong><span>executed</span></strong> means we compiled it to native code and ran it, and the output you see is what the machine printed. </p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/p/exploring-how-mlir-works-the-compiler?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/p/exploring-how-mlir-works-the-compiler?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><div><hr></div><h2>The problem LLVM could not see</h2><p><strong>LLVM IR&#8217;s</strong> contract, the one that conquered the CPU world, is deliberately narrow: scalar values in <strong>SSA form</strong>, a flat control flow graph of basic blocks, loads and stores against a flat memory model, one abstraction level for everything. </p><p>That narrowness is why<em> thirty-odd frontends </em>and every serious instruction set could meet in the middle. It is also precisely what makes LLVM IR the wrong meeting point for machine learning.</p><p>Consider what a tensor compiler needs to know about a matrix multiplication. </p><ul><li><p><em>That the two input buffers are <strong>dense 2-D arrays </strong>with known strides. </em></p></li><li><p><em>That the <strong>triple loop nest </strong>around it is two parallel dimensions and one reduction. </em></p></li><li><p><em>That the accumulator is <strong>transient </strong>and could live in registers or shared memory. </em></p></li><li><p><em>That the whole operation is <strong>one node in a graph </strong>whose neighbors it might profitably fuse with. </em></p></li></ul><p>By the time a matmul has been lowered to LLVM IR, every one of those facts has been destroyed. The<strong> arrays are opaque pointers</strong>, the loop structure is a thicket of compare-and-branch, the parallelism is gone, and the fusion opportunity evaporated three layers up. </p><p>Recovering that information from LLVM IR is the <strong>auto-vectorization problem</strong>, a research area with forty years of partial credit. The lesson the ML compiler generation drew was blunt: do not recover structure. Never lose it.</p><p>The second problem was <strong>combinatorial</strong> rather than semantic. Lattner has told the origin story himself, most recently in January 2026 in &#8220;<em>Democratizing AI Compute, Part 8</em>,&#8221; on Modular&#8217;s blog: <strong>MLIR began in 2018 inside Google</strong>, in the middle of the TPU buildout, when the TensorFlow ecosystem had accumulated a zoo of graph representations, converters, and codegen paths that shared almost nothing. </p><p>Every framework wanted to reach every accelerator. Every accelerator vendor was rebuilding the same 80 percent of a compiler: parsers, printers, pass managers, verifiers, canonicalizers, testing harnesses. </p><p>With N frontends and M targets, the industry was signing up to build and staff something like <strong>N times M compilers</strong>.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!-2xT!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d0c6a4e-aaa8-4dd0-87bf-5b3ec002986f_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!-2xT!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d0c6a4e-aaa8-4dd0-87bf-5b3ec002986f_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!-2xT!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d0c6a4e-aaa8-4dd0-87bf-5b3ec002986f_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!-2xT!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d0c6a4e-aaa8-4dd0-87bf-5b3ec002986f_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!-2xT!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d0c6a4e-aaa8-4dd0-87bf-5b3ec002986f_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!-2xT!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d0c6a4e-aaa8-4dd0-87bf-5b3ec002986f_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/5d0c6a4e-aaa8-4dd0-87bf-5b3ec002986f_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Point to point lowering paths versus a shared waist&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Point to point lowering paths versus a shared waist" title="Point to point lowering paths versus a shared waist" srcset="https://substackcdn.com/image/fetch/$s_!-2xT!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d0c6a4e-aaa8-4dd0-87bf-5b3ec002986f_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!-2xT!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d0c6a4e-aaa8-4dd0-87bf-5b3ec002986f_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!-2xT!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d0c6a4e-aaa8-4dd0-87bf-5b3ec002986f_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!-2xT!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d0c6a4e-aaa8-4dd0-87bf-5b3ec002986f_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><em>Figure 1. The arithmetic that motivated MLIR. Five frontends and six target families connected point to point is 30 lowering paths, each a compiler someone must build and maintain. Routed through a shared waist it is 11. LLVM proved the hourglass works for CPUs; MLIR generalizes the hourglass itself.</em></figcaption></figure></div><p>The insight that became MLIR was to go one level more abstract than LLVM did. LLVM&#8217;s answer to fragmentation was a single <strong>fixed IR</strong> in the middle. </p><p>But ML compilation does <strong>not have one natural middle</strong>: it has a graph level, a tensor algebra level, a loop level, a vector level, a hardware intrinsic level, and programs need to descend through all of them. </p><p>So instead of shipping another fixed IR, ship the infrastructure that makes intermediate representations cheap to build, and let them all coexist in one framework, one type system, one pass manager, one textual format. </p><p>Google open sourced the project in April 2019 and donated it to the LLVM Foundation, with the code landing in the LLVM monorepo at the end of that year. </p><p>The design paper, &#8220;<em>MLIR: Scaling Compiler Infrastructure for Domain Specific Computation</em>&#8221; by Lattner, Mehdi Amini, Uday Bondhugula, Albert Cohen, and colleagues, appeared at <strong>CGO 2021</strong> and is still the best single description of the architecture.</p><p>The name expands to <strong>Multi-Level Intermediate Representation</strong>, and the plural is the entire point. One artifact can hold, simultaneously and legally, a tensor-algebra view of one function, an explicit loop nest view of another, and <strong>raw hardware intrinsics</strong> for a third, with every level checked by the same verifier and transformed by the same pass infrastructure. </p><p>The rest of this article is a tour of how that actually works, at the level of real IR, because the abstractions only make sense once you have watched them move.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share The Software Frontier&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share The Software Frontier</span></a></p><div><hr></div><h2>One idea, applied without exception</h2><p>Strip away everything else and MLIR is one data structure applied with unusual discipline. The universe contains <strong>operations</strong>. </p><p>An operation has a name, takes SSA <strong>values</strong> as operands, produces values as results, carries compile-time constants called <strong>attributes</strong>, and, this is the structural leap, may contain <strong>regions</strong>: nested lists of blocks which themselves contain operations. Values have <strong>types</strong>. </p><p>That is the whole object model. There are no statements, no expressions, no special forms. Here is the smallest interesting function, in the friendly syntax you will see in every MLIR dump:</p><p>01_axpy.mlir <strong><span>verified: mlir-opt 20.1.2</span></strong></p><pre><code><code>func.func @axpy(%a: f32, %x: f32, %y: f32) -&gt; f32 {
  %0 = arith.mulf %a, %x : f32
  %1 = arith.addf %0, %y : f32
  return %1 : f32
}</code></code></pre><p>Friendly syntax is sugar. Ask <code>mlir-opt</code> to print the same module in generic form and the uniformity becomes visible:</p><p>mlir-opt --mlir-print-op-generic 01_axpy.mlir <strong><span>verified: actual tool output</span></strong></p><pre><code><code>"builtin.module"() ({
  "func.func"() &lt;{function_type = (f32, f32, f32) -&gt; f32, sym_name = "axpy"}&gt; ({
  ^bb0(%arg0: f32, %arg1: f32, %arg2: f32):
    %0 = "arith.mulf"(%arg0, %arg1) &lt;{fastmath = #arith.fastmath&lt;none&gt;}&gt; : (f32, f32) -&gt; f32
    %1 = "arith.addf"(%0, %arg2) &lt;{fastmath = #arith.fastmath&lt;none&gt;}&gt; : (f32, f32) -&gt; f32
    "func.return"(%1) : (f32) -&gt; ()
  }) : () -&gt; ()
}) : () -&gt; ()</code></code></pre><p>Read that carefully, because it is the <strong>most important dump</strong> in this article. The module is an operation named <code>builtin.module</code> with one region. The function is an operation named <code>func.func</code> whose &#8220;<em>body</em>&#8221; is just a region and whose name and signature are ordinary attributes. </p><p>Even <code>return</code> is an operation. There are no built-in language constructs at all; the things that look built in live in a dialect that happens to be called <code>builtin</code>. Every structure you will meet from here on, loop nests, <strong>GPU kernels</strong>, entire schedules, is an operation containing regions containing operations, all the way down.</p><p>Regions are what LLVM IR never had. LLVM gives you <strong>one flat control flow graph per function</strong>, so a loop exists only as a cycle among basic blocks, something analyses must rediscover with dominator trees. </p><p>In MLIR, an operation like a loop <em>contains</em> its body as a region, so the nesting of the <strong>source program </strong>survives as nesting in the IR. A GPU kernel is an operation containing the kernel body. A module is an operation containing functions. </p><p>SSA and dominance still hold inside each region, so classic optimization theory transfers intact, but structure is now first class rather than archaeological.</p><p>Discipline this strict pays off in the machinery. Because everything is an operation, one parser, one printer, one pass manager, one multithreading strategy, and <strong>one verification framework</strong> serve every abstraction level ever built on MLIR, including ones that do not exist yet. </p><p>Extensibility is kept honest by three mechanisms. </p><ul><li><p><em><strong>Verifiers</strong> let each operation enforce its own structural invariants, so malformed IR dies at the boundary instead of corrupting a pass 40 steps later. </em></p></li><li><p><em><strong>Traits</strong> declare properties like &#8220;this op has no side effects&#8221; or &#8220;operands and results share a type.&#8221; </em></p></li><li><p><em>And <strong>interfaces</strong> are the load-bearing wall: a pass written against </em><code>LoopLikeOpInterface</code><em> or </em><code>DestinationStyleOpInterface</code><em> works on any operation implementing the interface, including operations invented years after the pass was written by people who never met the pass author. </em></p></li></ul><p>That is the property that makes what follows possible: generic infrastructure over an open universe of abstractions.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Dialects: a periodic table you can extend</h2><p>Operations are grouped into <strong>dialects</strong>: namespaces that bundle related ops, types, and attributes together with their verifiers, canonicalization patterns, and documentation. </p><p>The prefix before the dot tells you the dialect: <code>arith.mulf</code>, <code>func.func</code>, <code>linalg.matmul</code>. A <strong>dialect is not a language</strong>; it is closer to a library of IR vocabulary, and dialects mix freely in one function. </p><p>The stock distribution is already a small civilization:</p><p>the in-tree roster <strong><span>verified: actual tool output</span></strong></p><pre><code><code>$ mlir-opt --show-dialects
Available Dialects (48):
  acc, affine, amdgpu, amx, arith, arm_neon, arm_sme, arm_sve, async,
  bufferization, builtin, cf, complex, dlti, emitc, func, gpu, index,
  irdl, linalg, llvm, math, memref, mesh, ml_program, mpi, nvgpu,
  nvvm, omp, pdl, pdl_interp, polynomial, ptr, quant, rocdl, scf,
  shape, sparse_tensor, spirv, tensor, test, test_dyn, tosa,
  transform, ub, vector, x86vector, xegpu</code></code></pre><p>Forty-eight dialects ship in <strong>LLVM 20.1.2</strong>, and the count only grows: CPU vector extensions (<code>amx</code>, <code>arm_sve</code>, <code>x86vector</code>), GPU vendor exits (<code>nvvm</code>, <code>rocdl</code>, <code>xegpu</code>), parallel runtimes (<code>omp</code>, <code>async</code>), quantization, sparse tensors, hardware description. </p><p>Out of tree, the population explodes; we will meet the important settlers in a few sections. What matters is the shape of the traffic: programs enter at high abstraction and descend. A few landmarks, each verified:</p><p><code>scf</code><strong>, structured control flow.</strong> Loops and conditionals as region-holding operations rather than branch spaghetti. Loop-carried state is explicit SSA, threaded through <code>iter_args</code>:</p><p>dot product in scf <strong><span>verified: mlir-opt 20.1.2</span></strong></p><pre><code><code>// scf: structured control flow with loop-carried values
func.func @dot(%a: memref&lt;1024xf32&gt;, %b: memref&lt;1024xf32&gt;) -&gt; f32 {
  %c0 = arith.constant 0 : index
  %c1 = arith.constant 1 : index
  %n  = arith.constant 1024 : index
  %zero = arith.constant 0.0 : f32
  %sum = scf.for %i = %c0 to %n step %c1
         iter_args(%acc = %zero) -&gt; (f32) {
    %x = memref.load %a[%i] : memref&lt;1024xf32&gt;
    %y = memref.load %b[%i] : memref&lt;1024xf32&gt;
    %m = arith.mulf %x, %y : f32
    %s = arith.addf %acc, %m : f32
    scf.yield %s : f32
  }
  return %sum : f32
}</code></code></pre><p><code>affine</code><strong>, the polyhedral inheritance.</strong> Same loops, harsher rules: bounds and subscripts must be affine expressions of the induction variables. In exchange, dependence analysis becomes decidable. </p><p>The compiler can <em>prove</em> that iterations are independent, that tiling is legal, that this access pattern never aliases that one, instead of guessing:</p><p>1-D stencil in affine <strong><span>verified: mlir-opt 20.1.2</span></strong></p><pre><code><code>// affine: loops + accesses restricted to affine expressions, so the
// compiler can *prove* dependence facts instead of guessing
func.func @blur1d(%in: memref&lt;1026xf32&gt;, %out: memref&lt;1024xf32&gt;) {
  affine.for %i = 0 to 1024 {
    %l = affine.load %in[%i]     : memref&lt;1026xf32&gt;
    %c = affine.load %in[%i + 1] : memref&lt;1026xf32&gt;
    %r = affine.load %in[%i + 2] : memref&lt;1026xf32&gt;
    %s0 = arith.addf %l, %c : f32
    %s1 = arith.addf %s0, %r : f32
    affine.store %s1, %out[%i] : memref&lt;1024xf32&gt;
  }
  return
}</code></code></pre><p><code>tensor</code><strong> and </strong><code>memref</code><strong>, the two memories.</strong> This pair encodes the single most useful distinction in the whole stack. </p><p>A <code>tensor</code> is an immutable SSA <em>value</em>: no address, no aliasing, safe to reorder and fuse aggressively, the natural currency of ML graphs. </p><p>A <code>memref</code> is a <em>view of actual memory</em>: base pointer, sizes, strides, and, critically for this audience, an address space. The types are expressive enough to carry real kernel engineering:</p><p>types that will look familiar <strong><span>verified: mlir-opt 20.1.2</span></strong></p><pre><code><code>func.func private @smem_tile(%t: memref&lt;64x64xf16, strided&lt;[68, 1]&gt;, 3&gt;)
func.func private @kv_page(%p: tensor&lt;16x8x128xbf16&gt;)</code></code></pre><p>The first is a <strong>64x64 half-precision tile </strong>with a padded row stride of 68 elements, the classic plus-four skew that kills shared memory bank conflicts, living in address space 3: GPU shared memory. </p><p>The second is a bf16 paged-KV block of the kind every serving engine slings around. MLIR&#8217;s type system says these things natively; nothing here is a comment or a convention.</p><p><code>vector</code><strong>, the portable SIMD layer.</strong> Fixed-shape virtual registers and the operations that move them: <code>vector.fma</code>, transfers, contractions, shuffles. Target independent until the last moment, then peeled into NEON, AVX-512, or GPU instructions:</p><p>an 8-wide fused multiply-add <strong><span>verified: mlir-opt 20.1.2</span></strong></p><pre><code><code>// vector: SIMD/SIMT-width-explicit compute on virtual registers
func.func @vec_fma(%a: vector&lt;8xf32&gt;, %b: vector&lt;8xf32&gt;, %c: vector&lt;8xf32&gt;) -&gt; vector&lt;8xf32&gt; {
  %0 = vector.fma %a, %b, %c : vector&lt;8xf32&gt;
  return %0 : vector&lt;8xf32&gt;
}</code></code></pre><p>And beneath everything, the exits: the <code>gpu</code> dialect for launch grids and kernels, <code>nvvm</code> and <code>rocdl</code> wrapping the vendors&#8217; intrinsics, and the <code>llvm</code> dialect, <strong>which models LLVM IR</strong> itself as just another MLIR dialect so that the handoff to the old empire is an ordinary rewrite rather than a file format negotiation. </p><h4>One table to orient the descent:</h4><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!0WfX!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F71f5b704-e895-41f5-afb5-e3d0f05a62b0_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!0WfX!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F71f5b704-e895-41f5-afb5-e3d0f05a62b0_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!0WfX!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F71f5b704-e895-41f5-afb5-e3d0f05a62b0_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!0WfX!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F71f5b704-e895-41f5-afb5-e3d0f05a62b0_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!0WfX!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F71f5b704-e895-41f5-afb5-e3d0f05a62b0_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!0WfX!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F71f5b704-e895-41f5-afb5-e3d0f05a62b0_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/71f5b704-e895-41f5-afb5-e3d0f05a62b0_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:177831,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/203533021?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F71f5b704-e895-41f5-afb5-e3d0f05a62b0_1634x893.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!0WfX!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F71f5b704-e895-41f5-afb5-e3d0f05a62b0_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!0WfX!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F71f5b704-e895-41f5-afb5-e3d0f05a62b0_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!0WfX!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F71f5b704-e895-41f5-afb5-e3d0f05a62b0_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!0WfX!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F71f5b704-e895-41f5-afb5-e3d0f05a62b0_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Read the<strong> right-hand column top</strong> to bottom and you have the core operating principle of the whole system, the one we will watch in motion next:<em> information is only destroyed downward, so every optimization should run at the highest level where its enabling facts still exist. </em></p><p>Fusion happens in tensor algebra, where aliasing cannot exist. Tiling happens where loops are structured objects. Bank conflict avoidance happens where address spaces are types. </p><p><em>By the time you reach LLVM, you are not optimizing anymore; you are translating.</em></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/p/exploring-how-mlir-works-the-compiler?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/p/exploring-how-mlir-works-the-compiler?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><div><hr></div><h2>linalg: the algebra at the waist</h2><p>If MLIR has a center of gravity, it is the <code>linalg</code> dialect: the tensor-algebra layer where most serious codegen pipelines, <strong>IREE&#8217;s</strong>, Modular&#8217;s, the upstream CPU flows, do their thinking. </p><p>Its design principle is the anti-autovectorizer stance taken to its conclusion. Instead of <strong>encoding a matmul as loops</strong> and hoping to rediscover its nature, <code>linalg</code> encodes the <em>nature</em> and treats loops as one of several possible spellings, generated on demand.</p><p>The workhorse is <code>linalg.generic</code>, which describes a perfectly nested computation with three pieces of data. First, <strong>indexing maps</strong>: one affine map per operand saying which element each point in the iteration space touches. </p><p>Second, <strong>iterator types</strong>: each loop dimension is declared <code>parallel</code> or <code>reduction</code>, which is the fact autovectorizers spend their lives failing to prove. Third, a region holding the scalar payload, the body computed at every point. </p><p>Everything the dialect knows, it knows structurally. A matmul&#8217;s maps are <code>(m, n, k) -&gt; (m, k)</code>, <code>(m, n, k) -&gt; (k, n)</code>, <code>(m, n, k) -&gt; (m, n)</code> over iterators <code>[parallel, parallel, reduction]</code>; named ops like <code>linalg.matmul</code> are simply memorable spellings of specific generics. </p><p>Here is the <strong>payload </strong>we will spend the rest of the article compiling, a real inference fragment: a linear layer with fused bias and ReLU.</p><p>02_mlp.mlir: matmul + fused bias/ReLU on tensors <strong><span>verified: mlir-opt 20.1.2</span></strong></p><pre><code><code>#id2d   = affine_map&lt;(m, n) -&gt; (m, n)&gt;
#bcast  = affine_map&lt;(m, n) -&gt; (n)&gt;

func.func @mlp_block(%x: tensor&lt;64x512xf32&gt;,
                     %w: tensor&lt;512x256xf32&gt;,
                     %b: tensor&lt;256xf32&gt;) -&gt; tensor&lt;64x256xf32&gt; {
  %c0 = arith.constant 0.0 : f32

  // Materialize the destination and zero-init the accumulator.
  %empty = tensor.empty() : tensor&lt;64x256xf32&gt;
  %acc   = linalg.fill ins(%c0 : f32)
                       outs(%empty : tensor&lt;64x256xf32&gt;) -&gt; tensor&lt;64x256xf32&gt;

  // y = x @ w
  %mm = linalg.matmul ins(%x, %w : tensor&lt;64x512xf32&gt;, tensor&lt;512x256xf32&gt;)
                      outs(%acc : tensor&lt;64x256xf32&gt;) -&gt; tensor&lt;64x256xf32&gt;

  // out = max(y + b, 0), one fused elementwise op
  %out = linalg.generic
           {indexing_maps = [#id2d, #bcast, #id2d],
            iterator_types = ["parallel", "parallel"]}
           ins(%mm, %b : tensor&lt;64x256xf32&gt;, tensor&lt;256xf32&gt;)
           outs(%empty : tensor&lt;64x256xf32&gt;) {
  ^bb0(%y: f32, %bias: f32, %o: f32):
    %sum  = arith.addf %y, %bias : f32
    %relu = arith.maximumf %sum, %c0 : f32
    linalg.yield %relu : f32
  } -&gt; tensor&lt;64x256xf32&gt;

  return %out : tensor&lt;64x256xf32&gt;
}</code></code></pre><p>Three details deserve attention. The bias&#8217;s indexing map, <code>(m, n) -&gt; (n)</code>, expresses broadcasting as pure index algebra: no replicated buffer, no runtime rule, just a map that ignores <code>m</code>. </p><p>The elementwise op fuses the add and the clamp into one region, because a region can hold any scalar computation, which is how epilogue fusion is spelled at this level. </p><p>And every linalg op writes into an explicit <code>outs</code> operand it also returns: <strong>destination-passing style</strong>. On immutable tensors the destination looks redundant, and semantically it is. </p><p>It exists as a standing declaration of where results <em>could</em> live, which is exactly the hint that lets bufferization later plan in-place updates instead of solving a global aliasing puzzle. Keep an eye on that <code>%empty</code>; it is about to earn its keep.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share The Software Frontier&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share The Software Frontier</span></a></p><div><hr></div><h2>The walkthrough: one MLP block to sm_90 PTX</h2><p>Now the demonstration the whole article is built around. We will take that block and walk it down the ladder one verified step at a time, watching what each level adds and what it forgets. </p><p>This is the same journey your <strong>PyTorch graphs</strong> take through Triton, that JAX programs take through XLA, that Mojo kernels take through MAX; we are just taking it on foot, with the intermediate files open.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!osL7!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9ceb5aa2-3004-40b4-8e18-ec62404d85ca_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!osL7!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9ceb5aa2-3004-40b4-8e18-ec62404d85ca_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!osL7!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9ceb5aa2-3004-40b4-8e18-ec62404d85ca_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!osL7!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9ceb5aa2-3004-40b4-8e18-ec62404d85ca_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!osL7!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9ceb5aa2-3004-40b4-8e18-ec62404d85ca_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!osL7!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9ceb5aa2-3004-40b4-8e18-ec62404d85ca_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/9ceb5aa2-3004-40b4-8e18-ec62404d85ca_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;The lowering ladder walked in this article&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="The lowering ladder walked in this article" title="The lowering ladder walked in this article" srcset="https://substackcdn.com/image/fetch/$s_!osL7!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9ceb5aa2-3004-40b4-8e18-ec62404d85ca_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!osL7!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9ceb5aa2-3004-40b4-8e18-ec62404d85ca_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!osL7!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9ceb5aa2-3004-40b4-8e18-ec62404d85ca_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!osL7!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9ceb5aa2-3004-40b4-8e18-ec62404d85ca_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><em>Figure 2. The descent this section performs, every stage machine-verified. The annotations name the decision each level owns; the italic line at the bottom is the operating principle from section 4.</em></figcaption></figure></div><h3>Step 1. Tiling, written as IR</h3><p>The first performance decision on <strong>any matrix multiply</strong> is blocking it for the memory hierarchy. In most compilers, tile sizes hide inside a C++ heuristic. </p><p>In MLIR, the transformation itself can be scripted in IR, using the <code>transform</code> dialect: a schedule language whose operations manipulate <em>other operations</em>. We append this to the file:</p><p>03_tile.mlir, the schedule part <strong><span>verified: mlir-opt 20.1.2</span></strong></p><pre><code><code>// ---- Schedule: also IR, in the transform dialect ----
module attributes {transform.with_named_sequence} {
  transform.named_sequence @__transform_main(%root: !transform.any_op {transform.readonly}) {
    %mm = transform.structured.match ops{["linalg.matmul"]} in %root
      : (!transform.any_op) -&gt; !transform.any_op
    %tiled, %loops:3 = transform.structured.tile_using_for %mm tile_sizes [16, 32, 64]
      : (!transform.any_op) -&gt; (!transform.any_op, !transform.any_op, !transform.any_op, !transform.any_op)
    transform.yield
  }
}</code></code></pre><p>The handles typed <code>!transform.any_op</code> are <strong>SSA values</strong> whose runtime contents are payload operations: the schedule matches every <code>linalg.matmul</code>, then tiles each one 16 by 32 by 64. </p><p>Run <code>mlir-opt --transform-interpreter</code> and the payload function comes back rewritten:</p><p>output: the tiled matmul <strong><span>verified: actual tool output, constants elided as marked</span></strong></p><pre><code><code>  func.func @matmul(%arg0: tensor&lt;64x512xf32&gt;, %arg1: tensor&lt;512x256xf32&gt;, %arg2: tensor&lt;64x256xf32&gt;) -&gt; tensor&lt;64x256xf32&gt; {
    // ... index constants (0, 16, 32, 64, 256, 512) elided ...
    %0 = scf.for %arg3 = %c0 to %c64 step %c16 iter_args(%arg4 = %arg2) -&gt; (tensor&lt;64x256xf32&gt;) {
      %1 = scf.for %arg5 = %c0_0 to %c256 step %c32 iter_args(%arg6 = %arg4) -&gt; (tensor&lt;64x256xf32&gt;) {
        %2 = scf.for %arg7 = %c0_1 to %c512 step %c64_2 iter_args(%arg8 = %arg6) -&gt; (tensor&lt;64x256xf32&gt;) {
          %extracted_slice = tensor.extract_slice %arg0[%arg3, %arg7] [16, 64] [1, 1] : tensor&lt;64x512xf32&gt; to tensor&lt;16x64xf32&gt;
          %extracted_slice_3 = tensor.extract_slice %arg1[%arg7, %arg5] [64, 32] [1, 1] : tensor&lt;512x256xf32&gt; to tensor&lt;64x32xf32&gt;
          %extracted_slice_4 = tensor.extract_slice %arg8[%arg3, %arg5] [16, 32] [1, 1] : tensor&lt;64x256xf32&gt; to tensor&lt;16x32xf32&gt;
          %3 = linalg.matmul ins(%extracted_slice, %extracted_slice_3 : tensor&lt;16x64xf32&gt;, tensor&lt;64x32xf32&gt;) outs(%extracted_slice_4 : tensor&lt;16x32xf32&gt;) -&gt; tensor&lt;16x32xf32&gt;
          %inserted_slice = tensor.insert_slice %3 into %arg8[%arg3, %arg5] [16, 32] [1, 1] : tensor&lt;16x32xf32&gt; into tensor&lt;64x256xf32&gt;
          scf.yield %inserted_slice : tensor&lt;64x256xf32&gt;
        }
        scf.yield %2 : tensor&lt;64x256xf32&gt;
      }
      scf.yield %1 : tensor&lt;64x256xf32&gt;
    }
    return %0 : tensor&lt;64x256xf32&gt;
  }</code></code></pre><p>One schedule op became a three-deep <code>scf.for</code> nest. <code>tensor.extract_slice</code> carves 16x64 and 64x32 tiles from the operands, an inner <code>linalg.matmul</code>, still the full structured op, now shaped <strong>16x64 times 64x32</strong>, computes each 16x32 block, and <code>tensor.insert_slice</code> threads the result through the loop-carried tensor. </p><p>Notice what did <em>not</em> happen: we have loops, but we did not lose the algebra. The inner op is <strong>still a matmul with its maps</strong> and iterator types intact, still available for vectorization, for lowering to tensor core intrinsics, or for another round of tiling. </p><p>Structure survives the transformation because the transformation is structure-aware.</p><h3>Step 2. Fusion: FlashAttention thinking, six lines long</h3><p>Tiling one op is table stakes. The move that defines modern kernel engineering, compute the producer inside the consumer&#8217;s tile so intermediates never <strong>round-trip through HBM</strong>, is the same idea that makes FlashAttention FlashAttention. </p><p>Here it is as a schedule: tile the bias/ReLU consumer into a parallel grid, then pull the matmul into each tile.</p><p>04_fuse.mlir, the schedule part <strong><span>verified: mlir-opt 20.1.2</span></strong></p><pre><code><code>module attributes {transform.with_named_sequence} {
  transform.named_sequence @__transform_main(%root: !transform.any_op {transform.readonly}) {
    // 1. Tile the *consumer* (bias+relu) into a parallel grid of 16x32 tiles.
    %gen = transform.structured.match ops{["linalg.generic"]} in %root
      : (!transform.any_op) -&gt; !transform.any_op
    %tiled, %grid = transform.structured.tile_using_forall %gen tile_sizes [16, 32]
      : (!transform.any_op) -&gt; (!transform.any_op, !transform.any_op)

    // 2. Pull the matmul producer inside each tile of that grid.
    %mm = transform.structured.match ops{["linalg.matmul"]} in %root
      : (!transform.any_op) -&gt; !transform.any_op
    %fused, %grid2 = transform.structured.fuse_into_containing_op %mm into %grid
      : (!transform.any_op, !transform.any_op) -&gt; (!transform.any_op, !transform.any_op)
    transform.yield
  }
}</code></code></pre><p>output after --transform-interpreter -canonicalize -cse <strong><span>verified: actual tool output</span></strong></p><pre><code><code>#map = affine_map&lt;(d0) -&gt; (d0 * 16)&gt;
#map1 = affine_map&lt;(d0) -&gt; (d0 * 32)&gt;
#map2 = affine_map&lt;(d0, d1) -&gt; (d0, d1)&gt;
#map3 = affine_map&lt;(d0, d1) -&gt; (d1)&gt;
  func.func @mlp_block(%arg0: tensor&lt;64x512xf32&gt;, %arg1: tensor&lt;512x256xf32&gt;, %arg2: tensor&lt;256xf32&gt;) -&gt; tensor&lt;64x256xf32&gt; {
    %cst = arith.constant 0.000000e+00 : f32
    %0 = tensor.empty() : tensor&lt;64x256xf32&gt;
    %1 = linalg.fill ins(%cst : f32) outs(%0 : tensor&lt;64x256xf32&gt;) -&gt; tensor&lt;64x256xf32&gt;
    %2 = scf.forall (%arg3, %arg4) in (4, 8) shared_outs(%arg5 = %0) -&gt; (tensor&lt;64x256xf32&gt;) {
      %3 = affine.apply #map(%arg3)
      %4 = affine.apply #map1(%arg4)
      %extracted_slice = tensor.extract_slice %arg0[%3, 0] [16, 512] [1, 1] : tensor&lt;64x512xf32&gt; to tensor&lt;16x512xf32&gt;
      %extracted_slice_0 = tensor.extract_slice %arg1[0, %4] [512, 32] [1, 1] : tensor&lt;512x256xf32&gt; to tensor&lt;512x32xf32&gt;
      %extracted_slice_1 = tensor.extract_slice %1[%3, %4] [16, 32] [1, 1] : tensor&lt;64x256xf32&gt; to tensor&lt;16x32xf32&gt;
      %5 = linalg.matmul ins(%extracted_slice, %extracted_slice_0 : tensor&lt;16x512xf32&gt;, tensor&lt;512x32xf32&gt;) outs(%extracted_slice_1 : tensor&lt;16x32xf32&gt;) -&gt; tensor&lt;16x32xf32&gt;
      %extracted_slice_2 = tensor.extract_slice %arg2[%4] [32] [1] : tensor&lt;256xf32&gt; to tensor&lt;32xf32&gt;
      %extracted_slice_3 = tensor.extract_slice %arg5[%3, %4] [16, 32] [1, 1] : tensor&lt;64x256xf32&gt; to tensor&lt;16x32xf32&gt;
      %6 = linalg.generic {indexing_maps = [#map2, #map3, #map2], iterator_types = ["parallel", "parallel"]} ins(%5, %extracted_slice_2 : tensor&lt;16x32xf32&gt;, tensor&lt;32xf32&gt;) outs(%extracted_slice_3 : tensor&lt;16x32xf32&gt;) {
      ^bb0(%in: f32, %in_4: f32, %out: f32):
        %7 = arith.addf %in, %in_4 : f32
        %8 = arith.maximumf %7, %cst : f32
        linalg.yield %8 : f32
      } -&gt; tensor&lt;16x32xf32&gt;
      scf.forall.in_parallel {
        tensor.parallel_insert_slice %6 into %arg5[%3, %4] [16, 32] [1, 1] : tensor&lt;16x32xf32&gt; into tensor&lt;64x256xf32&gt;
      }
    }
    return %2 : tensor&lt;64x256xf32&gt;
  }</code></code></pre><p>The consumer became an <code>scf.forall</code>, a parallel-by-construction grid of 4 by 8 tiles, and inside each tile sits a private 16x512 by 512x32 matmul feeding the fused bias-and-clamp directly. </p><p>The <strong>64x256 intermediate matrix</strong> that used to exist between the two ops is gone from the program: each tile&#8217;s slice is produced, consumed, and discarded in place. <code>canonicalize</code> and <code>cse</code> then deleted the original full-size matmul as dead code. </p><p>This is a working, verified skeleton of producer fusion, the transformation entire serving stacks are built around, expressed in <strong>six lines of schedule</strong> that a performance engineer can read, diff, and review like any other code. </p><h3>Step 3. Bufferization: the value world ends here</h3><p>Everything so far lived on immutable tensors, where fusion and reordering are trivially safe because aliasing cannot be expressed. Hardware, unfortunately, sells mutable bytes. </p><p>The crossing is <strong>one-shot bufferization</strong>, a whole-module analysis that assigns every tensor a buffer while proving where it can reuse memory instead of copying. Run it on the unfused MLP and look at what comes back:</p><p>mlir-opt -one-shot-bufferize=&#8221;bufferize-function-boundaries&#8221; <strong><span>verified: actual tool output</span></strong></p><pre><code><code>#map = affine_map&lt;(d0, d1) -&gt; (d0, d1)&gt;
#map1 = affine_map&lt;(d0, d1) -&gt; (d1)&gt;
module {
  func.func @mlp_block(%arg0: memref&lt;64x512xf32, strided&lt;[?, ?], offset: ?&gt;&gt;, %arg1: memref&lt;512x256xf32, strided&lt;[?, ?], offset: ?&gt;&gt;, %arg2: memref&lt;256xf32, strided&lt;[?], offset: ?&gt;&gt;) -&gt; memref&lt;64x256xf32&gt; {
    %cst = arith.constant 0.000000e+00 : f32
    %alloc = memref.alloc() {alignment = 64 : i64} : memref&lt;64x256xf32&gt;
    linalg.fill ins(%cst : f32) outs(%alloc : memref&lt;64x256xf32&gt;)
    linalg.matmul ins(%arg0, %arg1 : memref&lt;64x512xf32, strided&lt;[?, ?], offset: ?&gt;&gt;, memref&lt;512x256xf32, strided&lt;[?, ?], offset: ?&gt;&gt;) outs(%alloc : memref&lt;64x256xf32&gt;)
    linalg.generic {indexing_maps = [#map, #map1, #map], iterator_types = ["parallel", "parallel"]} ins(%alloc, %arg2 : memref&lt;64x256xf32&gt;, memref&lt;256xf32, strided&lt;[?], offset: ?&gt;&gt;) outs(%alloc : memref&lt;64x256xf32&gt;) {
    ^bb0(%in: f32, %in_0: f32, %out: f32):
      %0 = arith.addf %in, %in_0 : f32
      %1 = arith.maximumf %0, %cst : f32
      linalg.yield %1 : f32
    }
    %cast = memref.cast %alloc : memref&lt;64x256xf32&gt; to memref&lt;64x256xf32, strided&lt;[?, ?], offset: ?&gt;&gt;
    return %alloc : memref&lt;64x256xf32&gt;
    // ...</code></code></pre><p>Tensors became <code>memref</code>s; function boundaries acquired conservative strided layouts so callers can pass any compatible view. </p><p>But the line to study is the fused elementwise: it now reads from <code>%alloc</code> <em>and writes into </em><code>%alloc</code>. The analysis proved the<strong> bias/ReLU </strong>can execute in place on the matmul&#8217;s buffer, so the second 64KB allocation that a naive translation would emit simply does not exist. </p><p>That is <strong>destination-passing style</strong> paying out: because every linalg op declared its destination back in value land, the planner had the aliasing story handed to it. </p><p>This pass, more than any other, is where <strong>MLIR&#8217;s two-world design </strong>earns its complexity: algebra above the line, memory below it, and one well-tested analysis holding the border.</p><h3>Step 4. Proof of life: the IR actually runs</h3><p>Before descending to <strong>GPU intrinsics</strong>, a checkpoint that separates this article from a whiteboard exercise. </p><p>The remaining distance to a<strong> CPU binary is mechanical</strong>: linalg to loops, loops to branches, memrefs and arithmetic to the <code>llvm</code> dialect, then JIT. </p><p>We wrapped the MLP in a tiny <code>main</code> with known inputs, chose a bias of [-10, 1] so the ReLU visibly clamps, and ran the pipeline:</p><p>the full descent, one command <strong><span>executed: mlir-runner, LLVM 20.1.2</span></strong></p><pre><code><code>$ mlir-opt mlp_run.mlir \
    -one-shot-bufferize="bufferize-function-boundaries" \
    -convert-linalg-to-loops \
    -expand-strided-metadata -lower-affine \
    -convert-scf-to-cf -finalize-memref-to-llvm \
    -convert-arith-to-llvm -convert-cf-to-llvm -convert-func-to-llvm \
    -reconcile-unrealized-casts \
  | mlir-runner -e main -entry-point-result=void \
      -shared-libs=$LLVM/lib/libmlir_runner_utils.so.20.1,$LLVM/lib/libmlir_c_runner_utils.so.20.1</code></code></pre><p>what the machine printed</p><pre><code><code>Unranked Memref base@ = 0x55b27c52f500 rank = 2 offset = 0 sizes = [2, 2] strides = [2, 1] data = 
[[0,   6], 
 [0,   12]]</code></code></pre><p>By hand: x@w is [[4, 5], [10, 11]]; add the bias to get [[-6, 6], [0, 12]]; clamp at zero for [[0, 6], [0, 12]]. The machine agrees. </p><p>Every abstraction in this article compiles to electrons, and the fused ReLU is provably doing its job on the negative entries.</p><h3>Step 5. The GPU descent, ending in genuine PTX</h3><p>Now the branch this audience came for. Same matmul, memref form, aimed at a GPU. </p><p>First, expose the parallelism that linalg has known about all along: <code>-convert-linalg-to-parallel-loops</code> turns the two parallel iterators into an <code>scf.parallel</code> over (m, n) with the reduction as an ordinary inner loop:</p><p>parallel loops <strong><span>verified: actual tool output</span></strong></p><pre><code><code>  func.func @matmul(%arg0: memref&lt;64x512xf32&gt;, %arg1: memref&lt;512x256xf32&gt;, %arg2: memref&lt;64x256xf32&gt;) {
    %c0 = arith.constant 0 : index
    %c64 = arith.constant 64 : index
    %c1 = arith.constant 1 : index
    %c256 = arith.constant 256 : index
    %c512 = arith.constant 512 : index
    scf.parallel (%arg3, %arg4) = (%c0, %c0) to (%c64, %c256) step (%c1, %c1) {
      scf.for %arg5 = %c0 to %c512 step %c1 {
        %0 = memref.load %arg0[%arg3, %arg5] : memref&lt;64x512xf32&gt;
        %1 = memref.load %arg1[%arg5, %arg4] : memref&lt;512x256xf32&gt;
        %2 = memref.load %arg2[%arg3, %arg4] : memref&lt;64x256xf32&gt;
        %3 = arith.mulf %0, %1 : f32
        %4 = arith.addf %2, %3 : f32
        memref.store %4, %arg2[%arg3, %arg4] : memref&lt;64x256xf32&gt;
      }
      scf.reduce 
    }
    return
  }
}</code></code></pre><p>Next, the mapping decision: <code>-gpu-map-parallel-loops -convert-parallel-loops-to-gpu</code> assigns parallel dimensions to the grid and materializes a launch:</p><p>a kernel is born <strong><span>verified: actual tool output</span></strong></p><pre><code><code>    gpu.launch blocks(%arg3, %arg4, %arg5) in (%arg9 = %0, %arg10 = %1, %arg11 = %c1_0) threads(%arg6, %arg7, %arg8) in (%arg12 = %c1_0, %arg13 = %c1_0, %arg14 = %c1_0) {
      %2 = affine.apply #map1(%arg3)[%c1, %c0]
      %3 = affine.apply #map1(%arg4)[%c1, %c0]
      scf.for %arg15 = %c0 to %c512 step %c1 {
        %4 = memref.load %arg0[%2, %arg15] : memref&lt;64x512xf32&gt;
        %5 = memref.load %arg1[%arg15, %3] : memref&lt;512x256xf32&gt;
        %6 = memref.load %arg2[%2, %3] : memref&lt;64x256xf32&gt;
        %7 = arith.mulf %4, %5 : f32
        %8 = arith.addf %6, %7 : f32
        memref.store %8, %arg2[%2, %3] : memref&lt;64x256xf32&gt;
      }
      // ... gpu.terminator</code></code></pre><p>Be honest about what we are looking at: a 64 by 256 grid of blocks with a single thread each, the most <strong>naive mapping</strong> a GPU has ever been insulted with. </p><p>Production pipelines tile <em>before</em> mapping, so blocks get tiles and threads get elements, then layer on shared memory staging and tensor core ops from the <code>nvgpu</code> dialect (<code>mma.sync</code> and friends). </p><p>We are deliberately taking the unoptimized path because the plumbing, not the schedule, is today&#8217;s subject; later sections covers who builds the good schedules. <code>-gpu-kernel-outlining</code> then splits host from device, the step every CUDA programmer performs mentally when writing <code>__global__</code>:</p><p>outlined device module <strong><span>verified: actual tool output, trimmed as marked</span></strong></p><pre><code><code>  gpu.module @matmul_kernel {
    gpu.func @matmul_kernel(%arg0: index, %arg1: index, %arg2: memref&lt;64x512xf32&gt;, %arg3: memref&lt;512x256xf32&gt;, %arg4: memref&lt;64x256xf32&gt;, %arg5: index) kernel attributes {known_block_size = array&lt;i32: 1, 1, 1&gt;} {
      %block_id_x = gpu.block_id  x
      %block_id_y = gpu.block_id  y
      %thread_id_x = gpu.thread_id  x
      %thread_id_y = gpu.thread_id  y
      %0 = affine.apply #map1(%block_id_x)[%arg0, %arg1]
      %1 = affine.apply #map1(%block_id_y)[%arg0, %arg1]
      // ... same loop body, now reading gpu.block_id instead of loop IVs</code></code></pre><p>The <strong>loop body is unchanged</strong>, but the block indices now arrive from <code>gpu.block_id</code> instead of loop induction variables, and the kernel records its launch invariants (<code>known_block_size</code>) as attributes for later passes to exploit. </p><p>From here, <code>-convert-gpu-to-nvvm</code> rewrites device code into the <code>llvm</code> dialect sprinkled with NVIDIA&#8217;s intrinsics:</p><p>nvvm: the vendor boundary <strong><span>verified: excerpt of actual tool output</span></strong></p><pre><code><code>gpu.module @matmul_kernel {
  // signature abbreviated: 24 pointer and index arguments
  llvm.func @matmul_kernel(%arg0: i64, ..., %arg23: i64)
      attributes {gpu.kernel, gpu.known_block_size = array&lt;i32: 1, 1, 1&gt;,
                  nvvm.kernel, nvvm.maxntid = array&lt;i32: 1, 1, 1&gt;} {
    // ...
    %24 = nvvm.read.ptx.sreg.ctaid.x : i32
    %25 = llvm.sext %24 : i32 to i64
    %26 = nvvm.read.ptx.sreg.ctaid.y : i32
    // ... address arithmetic in llvm dialect: getelementptr, mul, add ...
  }
}</code></code></pre><p><code>gpu.block_id x</code> became <code>nvvm.read.ptx.sreg.ctaid.x</code>, a name any CUDA disassembly veteran will greet like an old acquaintance: the special register file, now visible in the IR. </p><p>Finally, the packaged pipeline <code>--gpu-lower-to-nvvm-pipeline="cubin-chip=sm_90 cubin-format=isa"</code> runs the whole descent and invokes LLVM&#8217;s NVPTX backend, embedding the result in a <code>gpu.binary</code> op. Extract it and you are holding 90 lines of the real thing:</p><p>out/matmul.ptx <strong><span>generated: LLVM 20.1.2 NVPTX backend, sm_90</span></strong></p><pre><code><code>//
// Generated by LLVM NVPTX Back-End
//

.version 7.8
.target sm_90
.address_size 64

&#9;// .globl&#9;matmul_kernel
&#9;// ... prologue: bounds check, address setup ...
$L__BB0_2:
&#9;ld.global.f32 &#9;%f4, [%rd35];
&#9;ld.global.f32 &#9;%f5, [%rd34];
&#9;mul.rn.f32 &#9;%f6, %f4, %f5;
&#9;add.rn.f32 &#9;%f7, %f7, %f6;
&#9;st.global.f32 &#9;[%rd6], %f7;
&#9;add.s64 &#9;%rd36, %rd36, %rd17;
&#9;add.s64 &#9;%rd35, %rd35, %rd8;
&#9;add.s64 &#9;%rd34, %rd34, %rd10;
&#9;setp.lt.s64 &#9;%p2, %rd36, %rd19;
&#9;@%p2 bra &#9;$L__BB0_2;
&#9;// ... epilogue ... (90 lines total)</code></code></pre><p>There is a<strong> final teaching moment</strong> hiding in that inner loop, and it is the thesis of this article restated by the machine. Look at <code>st.global.f32</code>: the accumulator is written back to global memory <em>every iteration of k</em>. </p><p>The value lives happily in register <code>%f7</code>, yet the store stays, because at this level nothing can prove the output buffer does not alias the inputs, so the backend must assume a reader might observe <strong>every intermediate sum. </strong></p><p>Four levels up, in linalg on tensors, that fact was free: tensors cannot alias. We destroyed the information on the way down, exactly as the ladder predicted, and the machine code is paying for it, one redundant HBM transaction per multiply. </p><p>Every dollar of <strong>kernel engineering</strong> ever spent is, in some form, the cost of putting facts back that a lowering threw away. MLIR&#8217;s entire bet is that it is cheaper to never throw them away.</p><p>Be precise about the distance between this demonstration and a kernel you would actually deploy, because the gap is the whole subject of the next section. </p><p>A production pipeline tiles <em>before</em> it maps to hardware, so a block owns a tile and a thread owns a register-sized sliver rather than the one-thread-per-block insult we emitted. </p><p>It stages the operand tiles through the padded shared-memory <code>memref</code> from section 4, turning the plus-four skew into live bank-conflict avoidance. </p><p>And instead of scalar <code>mul.rn.f32</code> plus that redundant store, it lowers the inner tile op through the <code>nvgpu</code> dialect into <code>mma.sync</code> or, on Hopper, <code>wgmma</code>, so the loop body becomes a <strong>warp-level matrix</strong> instruction feeding tensor cores, with the accumulator held in registers across the entire reduction and written back exactly once. </p><p>Here is the point that matters for understanding the whole system: <em>every one of those improvements is a different schedule over the identical set of passes we just ran.</em> The naive descent and the peak one share their entire lowering machinery; what separates 2 percent of peak from 90 percent is the sequence of tiling, fusion, and mapping decisions applied on the way down. </p><p>Which is why the ability to write that sequence as a first-class, reviewable, searchable artifact, rather than bury it in compiler C++, is not a curiosity. It is the ballgame.</p><div><hr></div><h2>The schedule is also IR</h2><p>Steps 1 and 2 of the walkthrough smuggled in the most radical idea in modern compiler engineering, so let us stop and look at it directly. </p><p>Classically, a <strong>compiler&#8217;s transformations</strong> are opaque C++ controlled by flags and heuristics; if the heuristic tiles your attention kernel wrong, your options are a rebuild or a prayer. </p><p>Halide&#8217;s great insight, back in 2012, was to split the <em>algorithm</em> from the <em>schedule</em> and make the schedule a first-class program. MLIR&#8217;s <code>transform</code> dialect imports that split into the compiler infrastructure itself: the schedule is IR, in the same file format, checked by the same verifier, with SSA handles pointing at the payload operations it manipulates.</p><p>The consequences compound. A schedule that is data can be <strong>diffed and code reviewed</strong>: the six lines in step 2 are a reviewable artifact in a way a C++ pass never is. </p><p>It can be <strong>searched</strong>: an autotuner can emit thousands of candidate schedules, run the interpreter, and measure, without recompiling the compiler; the transform interpreter is exactly the substrate you would build a kernel search system on. </p><p>It can be <strong>shipped per workload</strong>: IREE has used transform-dialect scripts to pin codegen strategies for specific dispatches, which is how you get reproducible kernels out of a general compiler. </p><p>And it fails loudly: apply a tiling to an op that cannot be tiled and the interpreter reports a definite error against a definite handle, rather than silently generating slow code.</p><p>For a <strong>CUDA-native audience </strong>the honest framing is this: the transform dialect is the compiler conceding that <em>you</em>, the performance engineer, sometimes know better, and giving you a typed, verifiable console instead of a fork of the codebase. </p><p>The upstream vocabulary already covers the daily verbs, <code>tile_using_for</code>, <code>tile_using_forall</code>, <code>fuse_into_containing_op</code>, <code>vectorize</code>, <code>pad</code>, plus pattern application and PDL-based matching for everything else. </p><p>It is not the default path for most users and may never be; it is the expert path, and its existence changes what the expert path costs.</p><div><hr></div><h2>Build a dialect before lunch</h2><p>The claim that<strong> MLIR makes IRs cheap</strong> deserves a demonstration with a bill attached. Suppose we run inference infrastructure and want the compiler to reason about our world: quantized weights, dequantization groups, fused epilogues. </p><p>In a classical compiler, adding a first-class operation means touching the parser, the printer, the verifier, the serializer, and a dozen switch statements. </p><p>In MLIR it means describing the op once, declaratively, in TableGen&#8217;s Operation Definition Specification, and generating the rest. Here is a real op for a real concern, weight-only quantized matmul, in 28 lines:</p><p>serve_ops.td <strong><span>verified: mlir-tblgen 20.1.2</span></strong></p><pre><code><code>include "mlir/IR/OpBase.td"
include "mlir/Interfaces/SideEffectInterfaces.td"

def Serve_Dialect : Dialect {
  let name = "serve";
  let summary = "Ops for LLM serving kernels";
  let cppNamespace = "::mlir::serve";
}

class Serve_Op&lt;string mnemonic, list&lt;Trait&gt; traits = []&gt;
    : Op&lt;Serve_Dialect, mnemonic, traits&gt;;

def Serve_DequantMatmulOp : Serve_Op&lt;"dequant_matmul", [Pure]&gt; {
  let summary = "matmul with fused groupwise weight dequantization";
  let description = [{
    Computes out = act @ dequant(weights, scales) without ever
    materializing the dequantized weight tensor in memory.
  }];
  let arguments = (ins
    TensorOf&lt;[F16]&gt;:$act,
    TensorOf&lt;[I8]&gt;:$weights,
    TensorOf&lt;[F16]&gt;:$scales,
    I64Attr:$group_size
  );
  let results = (outs TensorOf&lt;[F16]&gt;:$out);
  let assemblyFormat =
    "$act `,` $weights `,` $scales attr-dict `:` functional-type(operands, results)";
}</code></code></pre><p>Feed it to <code>mlir-tblgen</code> and the machinery materializes:</p><p>the leverage, measured <strong><span>verified: actual tool output</span></strong></p><pre><code><code>$ mlir-tblgen -gen-op-decls -I /usr/lib/llvm-20/include serve_ops.td &gt; ServeOps.h.inc
$ mlir-tblgen -gen-op-defs  -I /usr/lib/llvm-20/include serve_ops.td &gt; ServeOps.cpp.inc

# 28 lines of ODS in, 259 lines of generated C++ declarations out:
class DequantMatmulOp;
    ...</code></code></pre><p>Twenty-eight declarative lines became 259 lines of generated C++ declarations, plus definitions: typed accessors (<code>getAct()</code>, <code>getGroupSize()</code>), builders, parser and printer honoring our <code>assemblyFormat</code>, and verifier scaffolding enforcing the type constraints, an f16 activation tensor, i8 weights, f16 scales, or the IR does not construct. </p><p><strong>Nine-to-one leverage</strong> on the boilerplate, and every generated line is the same battle-tested code path the 48 upstream dialects use. </p><p>This is what &#8220;<em>infrastructure</em>&#8221; means concretely: our bespoke serving op gets round-trippable syntax, verification, and pass-manager citizenship for the price of a lunch break, which is why every hardware startup in section 9 could afford a real compiler at all.</p><p>Ops are half a dialect; the other half is transformations, and those are written in C++ against the pattern infrastructure. Here is a complete, compiling pass with a pattern this audience will feel in their DRAM bills: detect a matmul whose accumulator is a fresh zero fill, and tag it so a later lowering can emit a<strong> beta-equals-zero GEMM</strong> instead of a memset followed by a GEMM, saving one full write-and-read pass over the output through HBM:</p><p>TagZeroInit.cpp <strong><span>verified: g++ -fsyntax-only against MLIR 20 headers, exit 0</span></strong></p><pre><code><code>#include "mlir/Dialect/Arith/IR/Arith.h"
#include "mlir/Dialect/Linalg/IR/Linalg.h"
#include "mlir/IR/BuiltinOps.h"
#include "mlir/IR/PatternMatch.h"
#include "mlir/Pass/Pass.h"
#include "mlir/Transforms/GreedyPatternRewriteDriver.h"

using namespace mlir;

namespace {
// Match  linalg.fill(+0.0) feeding a linalg.matmul accumulator and tag
// the matmul, so a later lowering can emit a beta = 0 GEMM instead of
// a memset followed by a GEMM (one less trip through HBM).
struct TagZeroInitMatmul : OpRewritePattern&lt;linalg::MatmulOp&gt; {
  using OpRewritePattern::OpRewritePattern;

  LogicalResult matchAndRewrite(linalg::MatmulOp op,
                                PatternRewriter &amp;rewriter) const override {
    if (op-&gt;hasAttr("serve.zero_init"))
      return failure();          // already rewritten: stop, or loop forever
    auto fill =
        op.getDpsInitOperand(0)-&gt;get().getDefiningOp&lt;linalg::FillOp&gt;();
    if (!fill)
      return failure();
    auto cst =
        fill.getInputs()[0].getDefiningOp&lt;arith::ConstantFloatOp&gt;();
    if (!cst || !cst.value().isPosZero())
      return failure();
    rewriter.modifyOpInPlace(op, [&amp;] {
      op-&gt;setAttr("serve.zero_init", rewriter.getUnitAttr());
    });
    return success();
  }
};

struct TagZeroInitPass
    : PassWrapper&lt;TagZeroInitPass, OperationPass&lt;ModuleOp&gt;&gt; {
  MLIR_DEFINE_EXPLICIT_INTERNAL_INLINE_TYPE_ID(TagZeroInitPass)
  StringRef getArgument() const final { return "serve-tag-zero-init"; }
  StringRef getDescription() const final {
    return "Tag matmuls whose accumulator is a fresh zero fill";
  }
  void runOnOperation() override {
    RewritePatternSet patterns(&amp;getContext());
    patterns.add&lt;TagZeroInitMatmul&gt;(&amp;getContext());
    if (failed(applyPatternsGreedily(getOperation(), std::move(patterns))))
      signalPassFailure();
  }
};
} // namespace</code></code></pre><p>The anatomy generalizes to every pattern you will ever write. Match structurally by walking use-def edges (<code>getDefiningOp</code> through the destination operand, courtesy of destination-passing style). </p><p>Guard exhaustively, including against your own previous rewrite, because the <strong>greedy driver</strong> reapplies patterns to fixpoint and an unguarded pattern is an infinite loop. </p><p>Mutate only through the <code>rewriter</code>, never behind its back, so the driver&#8217;s worklist stays coherent. Fifty-ish lines, and it composes with every canonicalization and lowering upstream ships.</p><p>The last pillar of the workflow is cultural: everything above gets a <strong>FileCheck test</strong>, a <code>.mlir</code> file whose <code>RUN</code> line invokes the pass and whose <code>CHECK</code> lines assert on the output IR. </p><p>Because IR is text with stable syntax, transformation tests are diffs, reviewable by humans and <strong>replayable by CI</strong>. Teams that build on MLIR inherit this discipline for free, and it shows: it is the reason a project like </p><p>Triton can refactor its entire layout system without regressing a thousand kernels silently. </p><p>The honest total cost of a production dialect is of course higher than lunch, the semantics and the <strong>verifier discipline</strong> are the real work, but the floor has moved by an order of magnitude, and floors are what determine who gets to play.</p><div><hr></div><h2>Where MLIR is hiding in your stack</h2><p>Infrastructure succeeds when it disappears. Seven years after open sourcing, MLIR has disappeared into more of the AI stack than almost anyone tracks, so let us make the map explicit.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!WIp_!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F66b9e83b-34a7-4264-9604-c5bed22fdba9_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!WIp_!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F66b9e83b-34a7-4264-9604-c5bed22fdba9_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!WIp_!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F66b9e83b-34a7-4264-9604-c5bed22fdba9_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!WIp_!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F66b9e83b-34a7-4264-9604-c5bed22fdba9_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!WIp_!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F66b9e83b-34a7-4264-9604-c5bed22fdba9_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!WIp_!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F66b9e83b-34a7-4264-9604-c5bed22fdba9_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/66b9e83b-34a7-4264-9604-c5bed22fdba9_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;MLIR timeline 2018 to 2026&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="MLIR timeline 2018 to 2026" title="MLIR timeline 2018 to 2026" srcset="https://substackcdn.com/image/fetch/$s_!WIp_!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F66b9e83b-34a7-4264-9604-c5bed22fdba9_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!WIp_!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F66b9e83b-34a7-4264-9604-c5bed22fdba9_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!WIp_!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F66b9e83b-34a7-4264-9604-c5bed22fdba9_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!WIp_!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F66b9e83b-34a7-4264-9604-c5bed22fdba9_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><em>Figure 3. Eight years from a Google whiteboard to a $3.9 billion acquisition premium on the founding team. Dates and sourcing for every event are in the dossier.</em></figcaption></figure></div><p><strong>Triton, which means PyTorch.</strong> When OpenAI rebuilt Triton on MLIR in 2022, the language stopped being a research artifact and became a compiler stack. </p><p>Python is traced into the <code>triton</code> dialect (TTIR), a target-independent tile program; conversion to the <code>triton_gpu</code> dialect (TTGIR) then makes the performance-critical decision explicit by attaching a <strong>layout encoding</strong> to every tensor, describing exactly which thread of which warp owns which element, with the <strong>LinearLayout</strong> framework giving those mappings a uniform algebra; vendor sub-dialects handle Hopper and CDNA specifics before the drop into the <code>llvm</code> dialect and PTX or AMDGCN. </p><p>Since PyTorch 2 made Inductor the default compile path and Inductor emits <strong>Triton for GPU graphs</strong>, the practical consequence is that most compiled PyTorch GPU workloads on the planet route through MLIR today, whether their owners have heard of it or not. </p><p>And in 2025 the Triton team surfaced the lower rungs of its own ladder as a product: Gluon, documented in the Triton repository, hands expert users the <strong>TTGIR-level controls</strong>, layouts, shared memory, warp specialization, while reusing the same dialect stack underneath.</p><p><strong>Which means your serving engine, too.</strong> The inference stacks this publication&#8217;s readers actually operate, vLLM and SGLang chief among them, are not usually thought of as <strong>MLIR consumers,</strong> but trace the dependency and the connection is direct. </p><p>Both lean heavily on Triton for their fused and custom kernels: fused <strong>RMSNorm</strong> and RoPE, fused MoE routing, quantized GEMM epilogues, and increasingly the attention paths through backends in the FlashInfer lineage that mix <strong>hand-written CUDA</strong> with Triton-generated variants. </p><p>Every one of those Triton kernels is a program descending through TTIR and TTGIR, which is to say through MLIR, before it reaches ptxas. The abstraction that a serving engine calls &#8220;<em>a Triton kernel</em>&#8221; is, one layer down, exactly the layout-annotated dialect IR. </p><p>And the data structure those engines sling most obsessively, the paged KV cache, is precisely the strided, address-spaced <code>memref</code> from section 4 made concrete: a block table indexing <strong>fixed-size buffers</strong>, the same view-of-memory abstraction the type system was built to name. </p><p>When you profile a <strong>vLLM deployment </strong>on a B300 and find time in a fused decode kernel, you are, whether the dashboard says so or not, looking at the output of the compiler stack this article has been dissecting. The waist is not upstream of your infrastructure; it is inside it.</p><p><strong>NVIDIA, voluntarily.</strong> CUTLASS 4&#8217;s CuTe DSL lets engineers author kernels in Python that, per NVIDIA&#8217;s documentation, JIT through MLIR into optimized code finished by ptxas, with the CuTe layout algebra as the programming model. </p><p>Read that strategically: the company with the most to gain from CUDA C++ lock-in decided its flagship kernel library&#8217;s productivity layer should be built on the industry&#8217;s shared compiler substrate. </p><p>NVIDIA is <strong>not ceding the moat</strong>, ptxas and SASS remain closed, but it has conceded where the moat is not: the language layer above the driver is now contested ground, and NVIDIA chose to contest it with MLIR rather than against it.</p><p><strong>Google, the origin.</strong> XLA remains the TPU&#8217;s compiler, and since 2022 its portable front door is StableHLO, an MLIR dialect with versioned serialization that lets JAX, TensorFlow, and PyTorch programs travel between compilers without marrying one. </p><p>IREE, under the<strong> OpenXLA umbrella</strong>, is the fully MLIR-native end-to-end: StableHLO or Torch in, linalg in the middle, compiled artifacts for CPU, CUDA, ROCm, Vulkan, and mobile out. It is the closest thing to the textbook hourglass shipped as a product, and its codegen is where much of the transform-dialect and structured-codegen machinery in this article was hardened.</p><p><strong>The silicon insurgents, which is the actual story.</strong> Look at who builds their entire software identity on this infrastructure. Tenstorrent&#8217;s tt-mlir stack, public on GitHub since 2024, compiles graphs through TTIR into its TTNN library ops for Wormhole and Blackhole silicon. </p><p>AMD ships <strong>MLIR-AIE</strong>, its 1.2 release landing in January 2026 per Phoronix, as the low-level programming layer for the NPUs in every Ryzen AI laptop. Qualcomm&#8217;s compiler group published <strong>Hexagon-MLIR </strong>on arXiv this spring: a stack that ingests, notably, both PyTorch graphs and Triton kernels and lowers them onto Hexagon NPUs. </p><p>Tesla, per the FSD v14.3 release notes, rebuilt its in-car AI compiler and runtime on MLIR and claims a <strong>20 percent reaction-time improvement </strong>from the new stack, which, whatever discount you apply to vendor claims, is a statement about MLIR&#8217;s fitness for safety-critical, latency-bounded deployment that no benchmark suite could make. </p><p>Every one of these organizations faced the same choice: a <em>decade of compiler archaeology from scratch</em>, or a running start on shared infrastructure. None of them chose archaeology.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!eHF4!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79309645-08b2-4276-ba93-5677e7da4a5b_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!eHF4!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79309645-08b2-4276-ba93-5677e7da4a5b_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!eHF4!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79309645-08b2-4276-ba93-5677e7da4a5b_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!eHF4!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79309645-08b2-4276-ba93-5677e7da4a5b_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!eHF4!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79309645-08b2-4276-ba93-5677e7da4a5b_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!eHF4!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79309645-08b2-4276-ba93-5677e7da4a5b_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/79309645-08b2-4276-ba93-5677e7da4a5b_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Production software built on MLIR as of July 2026&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Production software built on MLIR as of July 2026" title="Production software built on MLIR as of July 2026" srcset="https://substackcdn.com/image/fetch/$s_!eHF4!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79309645-08b2-4276-ba93-5677e7da4a5b_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!eHF4!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79309645-08b2-4276-ba93-5677e7da4a5b_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!eHF4!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79309645-08b2-4276-ba93-5677e7da4a5b_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!eHF4!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79309645-08b2-4276-ba93-5677e7da4a5b_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><em>Figure 4. A selection, not a census, of production systems that are MLIR inside as of July 2026, grouped by what they are to their users. Sourcing per entry in the dossier.</em></figcaption></figure></div><p><strong>And beyond AI entirely</strong>, the quieter colonization: Flang, LLVM&#8217;s Fortran frontend, is built on MLIR dialects; CIRCT applies the same infrastructure to hardware design and verification; ClangIR is the upstream effort to give C and C++ themselves a structured MLIR layer above LLVM IR. </p><p>The pattern across all of it is the one Lattner himself named in January&#8217;s &#8220;<em>Democratizing AI Compute, Part 8</em>&#8221;: MLIR the <em>infrastructure</em> won more or less totally, while the original dream of one shared AI <em>compiler</em> on top of it fractured into vigorous, incompatible ecosystems. </p><p>The waist is shared; the programs flowing through it are not. Which is precisely what makes the waist worth money.</p><div><hr></div><h2>The economics of the waist</h2><p>Regular readers know this publication&#8217;s core lens: in accelerated computing, margins pool wherever software friction is highest. </p><p>CUDA&#8217;s twenty-year lesson is that the compiler and kernel layer is not a cost center; it is the <strong>tollbooth</strong>. MLIR changes the tollbooth&#8217;s economics in three distinct ways, and the June acquisition finally put a market price on one of them.</p><p><strong>First, it collapsed the entry ticket.</strong> Before this infrastructure existed, a credible accelerator software stack meant hundreds of compiler engineers for the better part of a decade; that is what CUDA cost, and it is why so many silicon startups died with excellent chips and unusable toolchains. </p><p>Previous sections showed the mechanism by which that changed: parsing, printing, verification, <strong>pass management</strong>, testing culture, all amortized across the industry, leaving a new entrant to spend only on what is genuinely theirs, the semantics of their silicon. </p><p>The evidence is the roster in Figure 4. Tenstorrent, a company of startup scale, fields a full <strong>graph-to-silicon compiler</strong>. AMD stands up an NPU stack as a side effort. This does not guarantee anyone beats NVIDIA; it guarantees the <em>attempt</em> no longer costs a billion dollars, and lowering the cost of attempts is how oligopolies erode.</p><p><strong>Second, it standardized the engineers, not just the code.</strong> A compiler developer who learns dialects, patterns, and the linalg descent is productive at Google, AMD, Qualcomm, Tenstorrent, Modular, or any of a hundred teams, because they all speak the same infrastructure. </p><p>Before MLIR, compiler talent was siloed by proprietary IR; now there is a labor market for the waist. Every technology that has standardized its practitioners, from TCP/IP to Kubernetes, has seen the layer above it commoditize and the layer below it consolidate. Keep that in mind when reading the next paragraph.</p><p><strong>Third, it created a control point you can buy.</strong> Qualcomm did not pay approximately $3.9 billion, per the Reuters-reported terms, for Modular&#8217;s revenue. It paid for roughly 150 people, per <strong>NAND Research&#8217;s</strong> headcount estimate, who include the founding architects of both MLIR and LLVM, plus MAX and Mojo, the <em>most complete independent attempt</em> to unify the layer above everyone&#8217;s silicon. </p><p>That is on the order of $26 million per employee, ten times acquihire norms, for a company whose language hit 1.0 beta in May, per its own PyPI listing, and whose compiler remains closed source with an open sourcing commitment for later this year. </p><p>The premium is not for what Modular sells; it is for what Qualcomm was <strong>structurally unable to build alone</strong>, demonstrated by the fact that its own Hexagon-MLIR effort was already consuming the ecosystem Modular&#8217;s founders created. </p><p>In the framework of our kernel-moat analysis: when the moat migrates from a proprietary language to a shared substrate, the scarce asset becomes the people who define the substrate&#8217;s direction, and June 24 was the day that asset got marked to market.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!ufst!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F40a98fbc-dafe-4d3e-acef-d60f77bfb5fa_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!ufst!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F40a98fbc-dafe-4d3e-acef-d60f77bfb5fa_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!ufst!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F40a98fbc-dafe-4d3e-acef-d60f77bfb5fa_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!ufst!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F40a98fbc-dafe-4d3e-acef-d60f77bfb5fa_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!ufst!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F40a98fbc-dafe-4d3e-acef-d60f77bfb5fa_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!ufst!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F40a98fbc-dafe-4d3e-acef-d60f77bfb5fa_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/40a98fbc-dafe-4d3e-acef-d60f77bfb5fa_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Modular capital raised, valuation, and acquisition price&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Modular capital raised, valuation, and acquisition price" title="Modular capital raised, valuation, and acquisition price" srcset="https://substackcdn.com/image/fetch/$s_!ufst!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F40a98fbc-dafe-4d3e-acef-d60f77bfb5fa_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!ufst!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F40a98fbc-dafe-4d3e-acef-d60f77bfb5fa_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!ufst!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F40a98fbc-dafe-4d3e-acef-d60f77bfb5fa_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!ufst!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F40a98fbc-dafe-4d3e-acef-d60f77bfb5fa_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><em>Figure 5. The compiler layer, marked to market. Funding history per Modular&#8217;s own repository newsline ($380M raised, $1.6B valuation at the September 2025 round); acquisition price as reported by Reuters and the WSJ; headcount per NAND Research; the per-person figure is our derived arithmetic.</em></figcaption></figure></div><p>NVIDIA&#8217;s response is the most sophisticated play on the board: embrace the waist from above. </p><p>Contributing to Triton&#8217;s NVIDIA backends and shipping <strong>CuTe DSL</strong> on MLIR costs NVIDIA the exclusivity of its <em>language</em> layer, which was already eroding, while making its <em>hardware</em> the best-supported target of the shared stack everyone else depends on, and keeping the truly proprietary layers, ptxas, SASS scheduling, <strong>NVLink-domain libraries</strong>, exactly as closed as before. </p><p>Commoditize your complement, in textbook form. For the neoclouds and inference operators this publication serves, this is not abstract; it changes the arithmetic that turns capex into margin. </p><p>A fleet&#8217;s single largest source of pricing power against its accelerator vendor is <strong>credible multi-homing</strong>: the ability to qualify a second silicon supplier and actually move workloads. </p><p>Historically that ability died at the kernel layer, because a serving stack tuned to CUDA was a serving stack married to NVIDIA, and the cost of re-qualifying every fused kernel on a new target was prohibitive enough that &#8220;<em>we could switch</em>&#8221; was rarely true. </p><p>The <strong>shared waist</strong> is what makes the threat credible, incrementally. Every serving stack that targets the waist, vLLM and SGLang through Triton, IREE, MAX, is quietly reducing the switching cost between a B300 fleet and an MI400 fleet at the kernel layer. </p><p>Portability never arrives as a binary event, and CUDA-only paths (<em>FlashAttention&#8217;s fastest branches, NCCL, the long tail of fused op</em>s) remain very real. But price convergence per token happens at the margin, and the <strong>margin is now programmable</strong>. </p><p>Watch whether accelerator price-performance spreads narrow over the next four quarters as these stacks mature; that spread is the purest financial expression of what this article has been describing.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>The counter-narrative</h2><p>The strongest version of the case against, because a thesis untested against its critics is marketing.</p><p><strong>&#8220;The library ate the compiler.&#8221;</strong> The performance frontier on NVIDIA hardware is still held by artisanal kernels: FlashAttention lineages, cuDNN, CUTLASS templates, hand-scheduled decode paths. </p><p>Much of what shipping compilers do is orchestrate calls into those libraries, and a fair reading of the <strong>TPP-MLIR work</strong> (arXiv 2404.15204), which reached roughly 90 percent of expert-written kernel performance on CPU benchmarks using upstream MLIR, is that the last ten percent is precisely where the money lives. </p><p>The rebuttal is not that compilers beat ninjas; it is that MLIR is the substrate the ninjas themselves are migrating to, CuTe DSL and Gluon being exhibits A and B. But the point stands: the waist wins by hosting excellence, not by automating it away.</p><p><strong>&#8220;LLMs will write the kernels.&#8221;</strong><em> If models generate CUDA or PTX directly on demand, does a shared IR layer matter less?</em> Plausibly it matters <em>more</em>: generated code needs verification, search needs a space with structure, and an IR with machine-checkable semantics and a transform language is a far better substrate for automated kernel search than free-form C++. </p><p>But intellectual honesty requires flagging this as the genuine wildcard: a world of cheap, correct, model-generated kernels reshuffles every layer of this analysis, ours included.</p><p><strong>&#8220;Dialect soup.&#8221;</strong> Forty-eight upstream dialects, hundreds downstream, and no blessed end-to-end path is real fragmentation with real costs: two teams can both &#8220;<em>use MLIR</em>&#8221; and share nothing but a parser. </p><p>Lattner&#8217;s own Part 8 critique makes essentially this charge, and the upstream community&#8217;s response, charters, governance reform, the <em>StableHLO-style versioning </em>discipline spreading to more dialects, is work in progress, not a solved problem. </p><p>Relatedly, core MLIR offers no stable cross-release IR contract: downstream projects pay a real rebase tax every LLVM release, a tax Triton and IREE budget actual headcount for.</p><p><strong>&#8220;The C++ tax.&#8221;</strong> The infrastructure&#8217;s expressiveness rides on heavy C++ templates and TableGen metaprogramming; build times are punishing, the learning curve is a cliff, and debugging a misapplied pattern through the greedy driver is an acquired skill. </p><p>Python bindings and tooling soften the edges for users, but dialect <em>authors</em> live in C++, and that gates the contributor pool. Fair, true, and priced in: the roster in Figure 4 is the market&#8217;s judgment that the tax is worth the leverage. </p><p>Every one of these criticisms describes a<strong> middle-aged infrastructure project</strong> with genuine adoption. None of them describes an alternative, and in infrastructure, the alternative is the only criticism that kills.</p><div><hr></div><h2>What to do with this</h2><p>If this article did its job, <strong>MLIR stopped being a logo</strong> on other people&#8217;s architecture slides and became legible machinery. </p><p>Here is the shortest path from legible to useful, calibrated for readers who already know what a memory coalescing problem feels like.</p><p><strong>Learn to read dumps before writing anything.</strong> The single highest-leverage habit is inspecting the IR your existing tools already produce. Triton will show you its TTIR and TTGIR for any kernel you own; start with a matmul you understand and find the layout encodings. </p><p>On your own experiments, <code>mlir-opt --mlir-print-ir-after-all</code> is the x-ray: every pass, before and after. When output confuses you, <code>--mlir-print-op-generic</code> strips the sugar and shows you the honest structure, exactly as in section 3.</p><p><strong>Replay this article.</strong> Everything here ran on stock Ubuntu 24.04 packages: <code>apt install mlir-20-tools llvm-20 libmlir-20-dev</code> and you have <code>mlir-opt</code>, <code>mlir-runner</code>, <code>mlir-translate</code>, <code>mlir-tblgen</code>, and <code>mlir-reduce</code> (the last being delta-debugging for IR: feed it a crashing module and a script, get back a minimal reproducer). </p><p>The dossier below lists every command. An afternoon of replaying the walkthrough will teach you more than a month of architecture diagrams.</p><p><strong>Then climb the same ladder the ecosystem climbed.</strong> The canonical study sequence: the Toy tutorial on mlir.llvm.org for the object model; the linalg dialect rationale for the structured-ops philosophy of section 5; the transform dialect tutorial for section 7&#8217;s machinery; then the CGO 2021 paper, which reads completely differently once you have touched the IR. </p><blockquote><p>For a <strong>working CUDA engineer,</strong> a realistic 30-day arc is: week one, read IR and replay the walkthrough; week two, lower your own toy op through the same pipeline and break things on purpose; week three, build the previous section dialect and make the pattern fire on real IR; week four, open the dumps of one production Triton kernel you own and annotate every layout decision. </p></blockquote><p>At the end of that month you will be conversant in the layer where, as section 10 argued, an increasing share of this industry&#8217;s margin gets decided.</p><p>Seven years ago this was a whiteboard sketch about taming TensorFlow&#8217;s compiler zoo. Today it compiles the kernels NVIDIA brags with, the <strong>NPU stacks </strong>of three chip giants, the vision system in a million cars, and it just priced a 150 person team at nearly four billion dollars. </p><p>The previous installment of this series ended by observing that LLVM won by being the infrastructure nobody had an incentive to replace. MLIR is running the same play one level up, on the layer where the AI hardware war is actually fought. </p><p>The empire had one IR. The dominion has as many as you can define, verify, and lower, and now you know how it is done.</p><div><hr></div><h2>Verification dossier</h2><p>House methodology: every load-bearing claim is tiered, every number is sourced or explicitly derived, and for this piece, every code artifact was executed against real tools before publication. </p><p>Environment: Ubuntu 24.04, stock packages <code>mlir-20-tools</code>, <code>llvm-20</code>, <code>libmlir-20-dev</code>, all LLVM/MLIR <strong>20.1.2</strong>. </p><p>Syntax note: MLIR evolves between LLVM releases; snippets are guaranteed against 20.1.2 and may need small spelling changes on older or newer toolchains.</p><h3>Code verification manifest</h3><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Qn7C!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb0f091cf-8966-4b8d-8664-d78ff4c6e534_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Qn7C!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb0f091cf-8966-4b8d-8664-d78ff4c6e534_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!Qn7C!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb0f091cf-8966-4b8d-8664-d78ff4c6e534_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!Qn7C!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb0f091cf-8966-4b8d-8664-d78ff4c6e534_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!Qn7C!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb0f091cf-8966-4b8d-8664-d78ff4c6e534_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Qn7C!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb0f091cf-8966-4b8d-8664-d78ff4c6e534_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b0f091cf-8966-4b8d-8664-d78ff4c6e534_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:233813,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/203533021?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb0f091cf-8966-4b8d-8664-d78ff4c6e534_1634x893.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Qn7C!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb0f091cf-8966-4b8d-8664-d78ff4c6e534_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!Qn7C!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb0f091cf-8966-4b8d-8664-d78ff4c6e534_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!Qn7C!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb0f091cf-8966-4b8d-8664-d78ff4c6e534_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!Qn7C!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb0f091cf-8966-4b8d-8664-d78ff4c6e534_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h3>Claims and confidence</h3><p>Tiers: <strong><span>A</span></strong> primary source or machine-verified. <strong><span>B</span></strong> reputable secondary reporting or vendor primary for vendor facts. <strong><span>C</span></strong> reasoned inference or derived arithmetic, flagged as such in text. <strong><span>D</span></strong> interpretation and forecast: our analysis, argued not asserted.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!LZdF!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26da92fe-0d13-43df-8f2f-cd7a64f952ff_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!LZdF!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26da92fe-0d13-43df-8f2f-cd7a64f952ff_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!LZdF!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26da92fe-0d13-43df-8f2f-cd7a64f952ff_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!LZdF!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26da92fe-0d13-43df-8f2f-cd7a64f952ff_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!LZdF!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26da92fe-0d13-43df-8f2f-cd7a64f952ff_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!LZdF!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26da92fe-0d13-43df-8f2f-cd7a64f952ff_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/26da92fe-0d13-43df-8f2f-cd7a64f952ff_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:226833,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/203533021?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26da92fe-0d13-43df-8f2f-cd7a64f952ff_1634x893.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!LZdF!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26da92fe-0d13-43df-8f2f-cd7a64f952ff_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!LZdF!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26da92fe-0d13-43df-8f2f-cd7a64f952ff_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!LZdF!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26da92fe-0d13-43df-8f2f-cd7a64f952ff_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!LZdF!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26da92fe-0d13-43df-8f2f-cd7a64f952ff_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Sources</h2><ol><li><p>Qualcomm, &#8220;Qualcomm to Acquire Modular,&#8221; investor press release, June 24, 2026. investor.qualcomm.com/news-events/press-releases/news-details/2026/Qualcomm-to-Acquire-Modular/default.aspx</p></li><li><p>Modular, &#8220;Qualcomm to acquire Modular,&#8221; company blog, June 2026. modular.com/blog/qualcomm-to-acquire-modular</p></li><li><p>Quartz, deal report citing Reuters and WSJ terms, June 24, 2026. qz.com/qualcomm-acquires-modular-ai-software-stock-deal-062426</p></li><li><p>NAND Research, &#8220;Qualcomm acquires Modular for its hardware-agnostic AI software layer,&#8221; June 2026. nand-research.com</p></li><li><p>Modular repository newsline (funding history, releases). github.com/modular/modular</p></li><li><p>Chris Lattner, &#8220;Democratizing AI Compute, Part 8: What about the MLIR compiler infrastructure?&#8221;, Modular blog, January 2026. modular.com/blog/democratizing-ai-compute-part-8-what-about-the-mlir-compiler-infrastructure</p></li><li><p>Modular, &#8220;The Path to Mojo 1.0&#8221;; PyPI, mojo-compiler release history (1.0.0b1, May 7, 2026). modular.com/blog/the-path-to-mojo-1-0; pypi.org/project/mojo-compiler</p></li><li><p>NVIDIA, CUTLASS documentation: CuTe DSL introduction and overview. docs.nvidia.com/cutlass/latest/media/docs/pythonDSL/cute_dsl_general/dsl_introduction.html</p></li><li><p>Triton project, Gluon documentation and dialect sources. triton-lang.org/main/gluon</p></li><li><p>Qualcomm et al., &#8220;Hexagon-MLIR,&#8221; arXiv 2602.19762, 2026. arxiv.org/pdf/2602.19762</p></li><li><p>Phoronix, &#8220;AMD MLIR-AIE 1.2,&#8221; January 2026. phoronix.com/news/AMD-MLIR-AIE-1.2</p></li><li><p>Tesla Oracle, &#8220;Tesla rolls out FSD v14.3 (2026.2.9.6): better reaction time, rewritten AI compiler with MLIR,&#8221; April 8, 2026. teslaoracle.com</p></li><li><p>Lattner, Amini, Bondhugula, Cohen, Davis, Pienaar, Riddle, Shpeisman, Vasilache, Zinenko, &#8220;MLIR: Scaling Compiler Infrastructure for Domain Specific Computation,&#8221; CGO 2021.</p></li><li><p>&#8220;TPP-MLIR&#8221; upstream performance study, arXiv 2404.15204. arxiv.org/pdf/2404.15204</p></li><li><p>MLIR project documentation: Toy tutorial, linalg rationale, transform dialect tutorial. mlir.llvm.org</p></li><li><p>StableHLO specification and repository. github.com/openxla/stablehlo</p></li><li><p>Tenstorrent, tt-mlir repository. github.com/tenstorrent/tt-mlir</p></li></ol>]]></content:encoded></item><item><title><![CDATA[How LLVM Works: The IR That Took Over Modern Computing]]></title><description><![CDATA[A two-person grant at the University of Illinois became the compiler substrate under Apple, Android, the PlayStation, Google and Meta datacenters, and every serious AI stack shipping today.]]></description><link>https://www.thesoftwarefrontier.com/p/how-llvm-works-the-ir-that-took-over</link><guid isPermaLink="false">https://www.thesoftwarefrontier.com/p/how-llvm-works-the-ir-that-took-over</guid><dc:creator><![CDATA[Lorenzo Bradanini]]></dc:creator><pubDate>Sun, 12 Jul 2026 12:56:15 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!tKIr!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c240161-6eb0-4373-8587-5ad5c822855e_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!tKIr!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c240161-6eb0-4373-8587-5ad5c822855e_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!tKIr!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c240161-6eb0-4373-8587-5ad5c822855e_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!tKIr!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c240161-6eb0-4373-8587-5ad5c822855e_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!tKIr!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c240161-6eb0-4373-8587-5ad5c822855e_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!tKIr!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c240161-6eb0-4373-8587-5ad5c822855e_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!tKIr!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c240161-6eb0-4373-8587-5ad5c822855e_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/4c240161-6eb0-4373-8587-5ad5c822855e_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2504090,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/203533011?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c240161-6eb0-4373-8587-5ad5c822855e_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!tKIr!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c240161-6eb0-4373-8587-5ad5c822855e_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!tKIr!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c240161-6eb0-4373-8587-5ad5c822855e_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!tKIr!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c240161-6eb0-4373-8587-5ad5c822855e_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!tKIr!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c240161-6eb0-4373-8587-5ad5c822855e_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>Intro</h2><p>This is the story of <strong>how LLVM&#8217;s intermediate representation ate modern computing</strong>, told at the level of the instructions themselves, with the source numbers checked.</p><p><strong>LLVM won </strong>because of two decisions made in one semester in the year 2000: a single, strongly typed, target-independent intermediate representation in <strong>SSA form </strong>that can be shipped and executed on its own, and a modular library architecture that let anyone reuse a compiler pass without dragging the whole compiler along. </p><p>Everything else, the trillion-dollar reach across mobile, cloud, consoles and AI, is a consequence of those two decisions. </p><p>The current release, LLVM 22.1, shipped in February 2026 and still carries the exact architecture<strong> Chris Lattner </strong>sketched over a<em> winter break twenty-five years ago.</em></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>The hole that GCC left</h2><p>Why the dominant open compiler of 2000 could not become infrastructure</p><p><strong>Compilers are the one technology every other technology has to pass through.</strong> They translate every line a human writes into something a processor can run, which means a compiler that is easy to build on is a lever under the entire software economy. </p><p>By the <strong>year 2000</strong> the open lever was the GNU Compiler Collection, and it was showing its age.</p><p><strong>GCC</strong>, launched in the 1980s, had become foundational: it built every Linux system, every Apple computer of the era, and a large share of the open source world. </p><p>But as <strong>Vikram Adve</strong> and Chris Lattner recount in their <a href="https://cacm.acm.org/federal-funding-of-academic-research/the-llvm-compiler-infrastructure/">June 2026 retrospective in Communications of the ACM</a>, GCC lagged on the techniques that were becoming central to modern compilation. </p><p>It was written in C rather than a modern object-oriented language, it lacked cross-module interprocedural optimization, and it did not gain <strong>Static Single Assignment</strong> form until GCC 4.0 in 2005. </p><p>Static Single Assignment, or SSA, is the representation on which most serious dataflow optimization is built, and GCC arrived late to it.</p><p>The deeper problem was structural. <strong>GCC was monolithic</strong>. You could not lift a piece of it out, an optimizer, a code generator, an analysis, and reuse that piece without pulling most of the compiler with it. </p><p>That single fact foreclosed most of the interesting futures: load-time and just-in-time compilation of mobile code, embedded scripting languages, sandboxed browser extensions, graphics shader compilation. </p><p>The managed-language runtimes of the day, the<strong> Jikes RVM for Java</strong>, Mono and later Roslyn for the CLR, were genuinely excellent inside their domain and advanced garbage collection and JIT compilation considerably. </p><p>But each imposed a language object model and a set of runtime semantics that ruled it out for the vast majority of other languages, and for anything that needed to ship mobile code without a managed runtime underneath it.</p><p>So the opening was specific and it was large: no open system combined ahead-of-time optimizing compilation of the kind static languages need with the self-contained, shippable, <strong>just-in-time-capable code representation</strong> that managed languages have, and offered both to <em>every</em> language rather than one family. Filling that hole is the whole of what follows.</p><h3>The federal money that made it possible</h3><p>Adve had come to Illinois with a research program aimed at flexible dynamic compilation for arbitrary languages. </p><p>His CAREER proposal to the <strong>National Science Foundation</strong>&#8216;s Next Generation Software program, submitted in 2000, funded the work; the acknowledgments in the CACM piece name grant <code>NSF EIA-0093426</code>, and Adve&#8217;s own record lists the <strong>CAREER </strong>award at roughly $499,211. </p><p>That grant, <strong>Adve and Lattner</strong> write, &#8220;<em>would prove pivotal to the group&#8217;s early work on LLVM</em>,&#8221; and it supported both authors through most of the project&#8217;s early life. </p><p>The point they press, and it is worth pressing, is that this was blue-sky research with an unknowable outcome, funded on the condition that the artifacts be released as open source. The<strong> trillion-dollar downstream</strong> was not forecast. It was seeded.</p><h4>A naming footnote</h4><p>LLVM originally stood for Low Level Virtual Machine. The name outgrew the acronym years ago; the project&#8217;s own position, stated on <a href="https://llvm.org/">llvm.org</a>, is that &#8220;LLVM&#8221; is now simply the name of the umbrella project and has &#8220;little to do with traditional virtual machines.&#8221; We use it as a proper noun throughout.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>The IR is the entire argument</h2><p>Everything LLVM can do is downstream of a <strong>single design object</strong></p><p>If you understand one thing about LLVM, understand the intermediate representation, because the reach of the project is a property of the IR and not of any particular optimizer or backend. </p><p><strong>LLVM IR is a strongly typed</strong>, RISC-like virtual instruction set. It abstracts away the machine, there are no physical registers and no addressing modes, but it stays low enough to represent operating systems, kernels and application code without imposing any language&#8217;s object model on top.</p><p>Three things about it do most of the work.</p><p><strong>It is in SSA form.</strong> Every value is assigned exactly once, and each use points back to a single, unambiguous definition. Where control flow merges, an explicit <code>phi</code> instruction selects the incoming value based on which predecessor block you arrived from. </p><p>This is the representation that <strong>Ron Cytron</strong> and colleagues formalized in 1991, and it is what makes dataflow optimization tractable: because there is one definition per value, the flow from definitions to uses is explicit, and reordering transformations are freed from spurious dependencies. </p><p>Consider, for example, a trivial loop.</p><pre><code><strong><span>int</span></strong> sum_to(<strong><span>int</span></strong> n){
  <strong><span>int</span></strong> s = <span>0</span>;
  <strong><span>for</span></strong> (<strong><span>int</span></strong> i = <span>0</span>; i &lt; n; i++)
    s += i;
  <strong><span>return</span></strong> s;
}</code></pre><p>Compiled at <code>-O0</code>, <strong>Clang</strong> first emits the naive version: local variables become stack slots via <code>alloca</code>, and every read and write is an explicit <code>load</code> or <code>store</code>. </p><p>Note the pointer type is simply <code>ptr</code>, a point we return to below.</p><pre><code><strong><span>define</span></strong> <span>i32</span> <span>@sum_to</span>(<span>i32</span> %n) {
<span>entry:</span>
  %i   = <strong><span>alloca</span></strong> <span>i32</span>
  %s   = <strong><span>alloca</span></strong> <span>i32</span>
  <strong><span>store</span></strong> <span>i32</span> <span>0</span>, <span>ptr</span> %s
  <strong><span>store</span></strong> <span>i32</span> <span>0</span>, <span>ptr</span> %i
  <strong><span>br</span></strong> <strong><span>label</span></strong> %cond
  <em><span>; ... loads and stores on every iteration ...</span></em>
}</code></pre><p>Then the <code>mem2reg</code> pass (memory-to-register promotion, one of the first transformations any pipeline runs) proves those stack slots never escape and rewrites them into <strong>pure SSA values. </strong></p><p>The stack traffic disappears, and <code>phi</code> nodes appear at the loop header and at the exit merge:</p><pre><code><strong><span>define</span></strong> <span>i32</span> <span>@sum_to</span>(<span>i32</span> %n) {
<span>entry:</span>
  %pos = <strong><span>icmp</span></strong> sgt <span>i32</span> %n, <span>0</span>
  <strong><span>br</span></strong> <span>i1</span> %pos, <strong><span>label</span></strong> %loop, <strong><span>label</span></strong> %exit

<span>loop:</span>                                <em><span>; preds = %entry, %loop</span></em>
  %i = <strong><span>phi</span></strong> <span>i32</span> [ <span>0</span>, %entry ], [ %i.next, %loop ]
  %s = <strong><span>phi</span></strong> <span>i32</span> [ <span>0</span>, %entry ], [ %s.next, %loop ]
  %s.next = <strong><span>add</span></strong> nsw <span>i32</span> %s, %i
  %i.next = <strong><span>add</span></strong> nsw <span>i32</span> %i, <span>1</span>
  %done = <strong><span>icmp</span></strong> eq <span>i32</span> %i.next, %n
  <strong><span>br</span></strong> <span>i1</span> %done, <strong><span>label</span></strong> %exit, <strong><span>label</span></strong> %loop

<span>exit:</span>                                <em><span>; preds = %entry, %loop</span></em>
  %r = <strong><span>phi</span></strong> <span>i32</span> [ <span>0</span>, %entry ], [ %s.next, %loop ]
  <strong><span>ret</span></strong> <span>i32</span> %r
}</code></pre><p>The <code>nsw</code><strong> flags </strong>mark the additions as having no signed wraparound, a promise the front end makes on the compiler&#8217;s behalf so later passes can optimize more aggressively. </p><p>This is the<strong> texture of LLVM IR</strong>: low-level operations, explicit control flow, and a scattering of flags and metadata that carry higher-level facts down to where they can be exploited.</p><p><strong>It carries a real type system.</strong> Integers of arbitrary bit width (<code>i1</code>, <code>i8</code>, <code>i32</code>, <code>i64</code>), floating point types, vectors as first-class values, arrays and structures. </p><p>Vectors matter more than they look: because a <strong>vector is a first-class value</strong>, the same IR expresses Intel SSE, AVX and AVX-512, Arm Neon and SVE, and the tensor operations that machine-learning front ends emit, and the auto-vectorizer can create them from scalar loops using dependence analysis. </p><p>Aggregate access goes through <code>getelementptr</code>, the address-arithmetic instruction that computes a field or element address without loading it:</p><pre><code><em><span>; struct Point { i32 x; i32 y; };  compute &amp;p-&gt;y</span></em>
%y.ptr = <strong><span>getelementptr</span></strong> inbounds { <span>i32</span>, <span>i32</span> }, <span>ptr</span> %p, <span>i32</span> <span>0</span>, <span>i32</span> <span>1</span>
%y     = <strong><span>load</span></strong> <span>i32</span>, <span>ptr</span> %y.ptr</code></pre><p><strong>It is extensible through intrinsics.</strong> Intrinsic functions are built-ins with defined semantics that behave like ordinary calls, so a pass that does not care about them can ignore them safely, while selected passes and the backend can recognize <code>llvm.memcpy</code>, <code>llvm.masked.load</code>, the matrix and vector-predication <strong>intrinsics</strong>, atomics and the rest, and lower them specially. </p><p>This is the mechanism that let the IR absorb decades of new hardware features without a language change every time.</p><p>The complexity was not free, and the authors are candid about it.<strong> LLVM IR launched in 2003</strong> with fewer than thirty-five polymorphic operations. It now carries more than one hundred and seventy-five, including in excess of one hundred and <strong>ten vector operations </strong>defined as intrinsics, plus hundreds of further semantic intrinsics for garbage collection, memory ordering, synchronization and stack manipulation. </p><p>Most of that can be ignored when writing a front end or an optimization pass, but it makes writing a new backend, or a formal semantics for the IR, genuinely hard.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!7UdV!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6b75fb7-62b5-46ac-b48a-3ce8b04cfd0c_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!7UdV!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6b75fb7-62b5-46ac-b48a-3ce8b04cfd0c_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!7UdV!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6b75fb7-62b5-46ac-b48a-3ce8b04cfd0c_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!7UdV!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6b75fb7-62b5-46ac-b48a-3ce8b04cfd0c_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!7UdV!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6b75fb7-62b5-46ac-b48a-3ce8b04cfd0c_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!7UdV!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6b75fb7-62b5-46ac-b48a-3ce8b04cfd0c_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e6b75fb7-62b5-46ac-b48a-3ce8b04cfd0c_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Bar chart: LLVM IR operation count grew from fewer than 35 in 2003 to more than 175 in 2026, with roughly 110 vector operations defined as intrinsics.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Bar chart: LLVM IR operation count grew from fewer than 35 in 2003 to more than 175 in 2026, with roughly 110 vector operations defined as intrinsics." title="Bar chart: LLVM IR operation count grew from fewer than 35 in 2003 to more than 175 in 2026, with roughly 110 vector operations defined as intrinsics." srcset="https://substackcdn.com/image/fetch/$s_!7UdV!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6b75fb7-62b5-46ac-b48a-3ce8b04cfd0c_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!7UdV!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6b75fb7-62b5-46ac-b48a-3ce8b04cfd0c_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!7UdV!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6b75fb7-62b5-46ac-b48a-3ce8b04cfd0c_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!7UdV!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6b75fb7-62b5-46ac-b48a-3ce8b04cfd0c_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Figure 1.</strong> The count of polymorphic operations in the IR, at first release and today. The 2026 split between core and vector operations is illustrative; the totals (&#8221;fewer than 35,&#8221; &#8220;more than 175,&#8221; &#8220;in excess of 110&#8221; vector) are the figures Adve and Lattner give in the CACM retrospective.</figcaption></figure></div><h3>Three forms of one thing, and the opaque-pointer turn</h3><p>The IR exists in three isomorphic forms: a human-readable textual assembly (<code>.ll</code>), a dense binary serialization called <strong>bitcode</strong> (<code>.bc</code>), and an in-memory form the compiler manipulates directly. </p><p>They are interconvertible without loss, which is exactly what lets the IR be shipped and compiled later rather than only used in-process.</p><pre><code>clang -S -emit-llvm foo.c -o foo.ll   <em><span># textual IR</span></em>
clang -c -emit-llvm foo.c -o foo.bc   <em><span># bitcode</span></em>
llvm-as foo.ll -o foo.bc              <em><span># text  -&gt; bitcode</span></em>
llvm-dis foo.bc -o foo.ll             <em><span># bitcode -&gt; text</span></em></code></pre><p>The IR is not frozen. Two recent changes show the project still reshaping its own foundation. The larger one is <strong>opaque pointers</strong>. For most of LLVM&#8217;s life a pointer carried its pointee type, so you wrote <code>i32*</code>. </p><p>That information was redundant with the types on the loads and stores that actually used the pointer, and it forced a swarm of no-op bitcast instructions. </p><p>Per the <a href="https://llvm.org/docs/OpaquePointers.html">official documentation</a>, <strong>opaque pointers</strong>, a single untyped <code>ptr</code>, became the default in LLVM 15 (2022), and typed pointers were removed entirely in LLVM 17 (2023). Every IR sample in this article uses <code>ptr</code> for that reason. </p><p>The second, newer change landed in <strong>LLVM 22</strong>: a dedicated <code>ptrtoaddr</code> instruction that extracts a pointer&#8217;s integer address while separating it from provenance tracking, which the older <code>ptrtoint</code> conflated. </p><p>It is a small instruction with a real purpose, sharper semantics for the memory model, and it is the kind of change that keeps the IR honest as verification tools grow more demanding.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Passes, and the separation of concerns</h2><p>The modular library architecture is the part that actually won</p><p>The most cited <strong>technical novelty of LLVM is the IR</strong>. The most <em>consequential</em> design decision, the authors argue and we agree, is not technical at all in the  usual sense: it is the modular, library-based architecture, built on a strict adherence to separation of concerns. </p><p>A <strong>compiler built on LLVM</strong> is assembled from libraries with clean interfaces rather than carved out of a monolith. That is why pieces of LLVM turn up in circuit-design tools, symbolic-execution engines, browser sandboxes and quantum-computing compilers, uses the authors freely admit they never imagined.</p><p>Concretely, transformation is organized as <strong>a pipeline of passes</strong>. Analysis passes compute facts (<em>dominator trees, alias analysis, loop information, scalar evolution</em>) and cache them; transformation passes consume those facts and rewrite the IR, declaring which analyses they preserve so the manager knows what to recompute. </p><p>Since LLVM 13 (2021) the default engine is the <strong>new pass manager</strong>, a rewrite that made analysis caching explicit and pass composition cheaper.</p><p>A canonical <code>-O2</code> pipeline threads dozens of passes together. A representative slice: </p><ul><li><p><code>SROA</code> and <code>mem2reg</code> to lift memory into SSA values; </p></li><li><p><code>EarlyCSE</code> and <code>GVN</code> to eliminate redundant computation; </p></li><li><p><code>InstCombine</code> to canonicalize and simplify instruction patterns; </p></li><li><p><code>SimplifyCFG</code> to clean up control flow; the inliner to pull small callees into their callers; </p></li><li><p><code>LICM</code> to hoist loop-invariant code; </p></li><li><p><code>SCCP</code> for sparse conditional constant propagation; </p></li><li><p><code>IndVars</code> and <code>LoopUnroll</code> for loops; </p></li></ul><p>then the two vectorizers, the <strong>loop vectorizer </strong>and the<strong> SLP</strong> (superword-level parallelism) vectorizer, near the end. You can watch any of it run.</p><pre><code><em><span># run one pass and print the result</span></em>
opt -passes=<span>&#8220;mem2reg&#8221;</span> -S sum.ll -o -

<em><span># see the full default -O2 pipeline the pass builder constructs</span></em>
opt -passes=<span>&#8220;default&lt;O2&gt;&#8221;</span> -print-pipeline-passes -S sum.ll -o /dev/null</code></pre><p>Writing a pass is deliberately cheap, which is the whole point. The 2003 release announcement claimed new users could &#8220;<em>write their first LLVM pass in hours,</em>&#8221; and the modern <strong>out-of-tree plugin form </strong>holds to that. Here is a complete, loadable function pass in the new pass manager:</p><pre><code><em><span>#include &#8220;llvm/IR/PassManager.h&#8221;</span></em>
<em><span>#include &#8220;llvm/Passes/PassBuilder.h&#8221;</span></em>
<em><span>#include &#8220;llvm/Passes/PassPlugin.h&#8221;</span></em>
<strong><span>using namespace</span></strong> llvm;

<strong><span>struct</span></strong> CountAllocas : PassInfoMixin&lt;CountAllocas&gt; {
  PreservedAnalyses run(Function &amp;F, FunctionAnalysisManager &amp;) {
    <strong><span>unsigned</span></strong> n = <span>0</span>;
    <strong><span>for</span></strong> (BasicBlock &amp;BB : F)
      <strong><span>for</span></strong> (Instruction &amp;I : BB)
        <strong><span>if</span></strong> (isa&lt;AllocaInst&gt;(I)) ++n;
    errs() &lt;&lt; F.getName() &lt;&lt; <span>&#8220;: &#8220;</span> &lt;&lt; n &lt;&lt; <span>&#8220; allocas\n&#8221;</span>;
    <strong><span>return</span></strong> PreservedAnalyses::all();   <em><span>// we changed nothing</span></em>
  }
};

<strong><span>extern</span></strong> <span>&#8220;C&#8221;</span> LLVM_ATTRIBUTE_WEAK ::llvm::PassPluginLibraryInfo
llvmGetPassPluginInfo() {
  <strong><span>return</span></strong> {LLVM_PLUGIN_API_VERSION, <span>&#8220;CountAllocas&#8221;</span>, LLVM_VERSION_STRING,
    [](PassBuilder &amp;PB) {
      PB.registerPipelineParsingCallback(
        [](StringRef Name, FunctionPassManager &amp;FPM,
           ArrayRef&lt;PassBuilder::PipelineElement&gt;) {
          <strong><span>if</span></strong> (Name == <span>&#8220;count-allocas&#8221;</span>) {
            FPM.addPass(CountAllocas()); <strong><span>return</span></strong> <strong><span>true</span></strong>;
          }
          <strong><span>return</span></strong> <strong><span>false</span></strong>;
        });
    }};
}</code></pre><pre><code>clang++ -shared -fPIC CountAllocas.cpp -o CountAllocas.so \
  $(llvm-config --cxxflags)
opt -load-pass-plugin=./CountAllocas.so \
    -passes=<span>&#8220;count-allocas&#8221;</span> -disable-output sum.bc</code></pre><p><em>The IR was the innovation everyone cites. The library architecture was the innovation that let ten thousand projects reuse a compiler without inheriting one.</em></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Five capabilities no one else has together</h2><p>A twenty-year-old assertion from the CGO 2004 paper that still holds</p><p>In their 2004 paper at the <strong>International Symposium on Code Generation and Optimization</strong>, Lattner and Adve made a specific claim: LLVM provides a combination of five capabilities that no other system offers at once. </p><p>The paper won the conference&#8217;s Most Influential Paper award ten years later, and the retrospective restates the claim as still true more than twenty years on. The five:</p><p>CapabilityWhat it means<strong>Persistent program information</strong>The IR can be preserved across an application&#8217;s whole lifetime, so optimization can happen at compile time, link time, load time, run time and even idle time between runs on the user&#8217;s machine.</p><p><strong>Offline code generation</strong>Regardless of persistence, code can be lowered to efficient native machine code offline, using techniques too expensive to run at runtime.<strong>User-specific profiling</strong>Profile data can be gathered from deployed software as the end user actually runs it, then fed back into <strong>profile-guided optimization</strong>.</p><p><strong>Transparent runtime model</strong>The IR imposes no object model, exception semantics or runtime, so it serves any language or mix of languages.<strong>Uniform whole-program compilation</strong>Because it is language-independent, an entire application, including its language runtime and system libraries, can be optimized together after linking.</p><p>The argument is not that each property is unique; it is that the <em>conjunction</em> is. Managed virtual machines like the<strong> JVM or the .NET CLI</strong> deliver user-specific profiling and partially deliver persistence and whole-program compilation, but they do not deliver a transparent runtime model, which is what confines them to a language family. </p><p>And when they do offer offline code generation, the authors note, they do it at the expense of persistence and profiling. GPU instruction sets such as CUDA&#8217;s PTX and <strong>OpenCL&#8217;s SPIR-V,</strong> and the JITs behind JavaScript, Python and WebAssembly, cover some combination of persistence and profiling and none of the rest. </p><p>GCC, superb at offline code generation, was never designed to ship a language-independent representation for lifelong compilation at all.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Qrpx!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d2d4b76-b556-4b61-9477-2eaad8127164_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Qrpx!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d2d4b76-b556-4b61-9477-2eaad8127164_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!Qrpx!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d2d4b76-b556-4b61-9477-2eaad8127164_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!Qrpx!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d2d4b76-b556-4b61-9477-2eaad8127164_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!Qrpx!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d2d4b76-b556-4b61-9477-2eaad8127164_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Qrpx!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d2d4b76-b556-4b61-9477-2eaad8127164_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/5d2d4b76-b556-4b61-9477-2eaad8127164_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Matrix comparing LLVM, JVM/.NET, GPU instruction sets, browser JITs and GCC across the five capabilities. Only LLVM is 'full' on all five.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Matrix comparing LLVM, JVM/.NET, GPU instruction sets, browser JITs and GCC across the five capabilities. Only LLVM is 'full' on all five." title="Matrix comparing LLVM, JVM/.NET, GPU instruction sets, browser JITs and GCC across the five capabilities. Only LLVM is 'full' on all five." srcset="https://substackcdn.com/image/fetch/$s_!Qrpx!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d2d4b76-b556-4b61-9477-2eaad8127164_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!Qrpx!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d2d4b76-b556-4b61-9477-2eaad8127164_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!Qrpx!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d2d4b76-b556-4b61-9477-2eaad8127164_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!Qrpx!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d2d4b76-b556-4b61-9477-2eaad8127164_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Figure 2.</strong> The five-capabilities matrix, encoded from the claims in the CGO 2004 paper and the 2026 retrospective. Teal marks full support, gold partial, empty none. LLVM is the only row that is full across the board, which is the entire thesis of the paper.</figcaption></figure></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>From IR to silicon</h2><p>Instruction selection, <strong>register allocation</strong>, and the <strong>TableGen</strong> machine</p><p>Optimized IR is still target-independent. Turning it into machine code is the backend&#8217;s job, and LLVM has, notably, two backends that do it.</p><p>The mature path is <strong>SelectionDAG</strong>. Each basic block is turned into a directed acyclic graph of operations; the DAG is legalized so every node is something the target can actually do; instruction selection matches subgraphs against target patterns to pick real machine instructions; the result is scheduled into a linear order; and then register allocation and emission follow. </p><p>The newer path is <strong>GlobalISel</strong>, designed to eventually replace <strong>SelectionDAG</strong>. It works on the whole function rather than one block at a time and skips the DAG entirely: an <code>IRTranslator</code> lowers IR to generic machine instructions, a <code>Legalizer</code> makes them legal, <code>RegBankSelect</code> assigns register banks, and an <code>InstructionSelect</code> pass chooses final opcodes. </p><p>As of the current release <strong>GlobalISel</strong> is the default at <code>-O0</code> on AArch64 and is used for AMD GPUs, and, as an ongoing <a href="https://discourse.llvm.org/t/status-on-enabling-globalisel-by-default-on-clang-on-aarch64/89964">LLVM forum discussion from February 2026</a> shows, work to enable it more broadly on AArch64 is still live. Two instruction selectors, maintained in parallel for years, is itself a statement about how much the project values getting this layer right.</p><p>Register allocation is where virtual registers, of which the IR has infinitely many, meet the finite physical register file. The default at <code>-O1</code> and above is the <strong>greedy</strong> allocator, which orders live ranges by priority and spills to the stack when it must; <code>-O0</code> uses a fast allocator that trades quality for speed. </p><p>Both are among the passes now being handed to machine learning, a thread we pick up in section 10.</p><p>Almost none of this is hand-written per target. Targets are described declaratively in <strong>TableGen</strong>, a domain-specific language whose <code>.td</code> files specify registers, instructions, calling conventions, scheduling models and selection patterns; a generator turns them into the C++ tables the backend uses. </p><p>A <strong>single instruction definition</strong> ties an assembly syntax to a SelectionDAG pattern:</p><pre><code><strong><span>def</span></strong> ADDrr : Instruction {
  <strong><span>let</span></strong> Namespace       = <span>&#8220;MyTarget&#8221;</span>;
  <strong><span>let</span></strong> OutOperandList = (outs GPR:$dst);
  <strong><span>let</span></strong> InOperandList  = (ins GPR:$src1, GPR:$src2);
  <strong><span>let</span></strong> AsmString      = <span>&#8220;add $dst, $src1, $src2&#8221;</span>;
  <em><span>// match (add a, b) in the DAG and emit this instruction</span></em>
  <strong><span>let</span></strong> Pattern = [(set <span>i32</span>:$dst, (add <span>i32</span>:$src1, <span>i32</span>:$src2))];
}</code></pre><p>This is why standing up a new architecture in LLVM, <strong>RISC-V</strong> being the most visible recent example, is a matter of writing target descriptions rather than a compiler. </p><p>The LLVM 22 notes alone record new scheduling models for <strong>NVIDIA&#8217;s Olympus CPU</strong>, additional Arm C1 cores, Intel Wildcat Lake and Nova Lake, and fresh RISC-V vector extensions, all arriving as target-description work on top of the shared infrastructure.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share The Software Frontier&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share The Software Frontier</span></a></p><div><hr></div><h2>The link-time frontier, and why ThinLTO exists</h2><p>Whole-program optimization that actually scales to a datacenter build</p><p><strong>Persistent IR</strong> unlocks link-time optimization: instead of discarding each source file&#8217;s IR after compiling it, you keep the bitcode, hand all of it to the linker, and optimize across module boundaries, inlining a function defined in one file into a caller in another, propagating constants program-wide. </p><p>Classic <strong>full LTO</strong> merges every module into one giant IR module and optimizes it as a unit. The result is excellent and the cost is brutal: the whole program has to fit in one process, and the work is largely serial, which is a non-starter when the program is a datacenter binary with tens of thousands of translation units.</p><p><strong>ThinLTO</strong>, first presented at EuroLLVM 2015 by <strong>Teresa Johnson</strong>, <strong>Mehdi Amini</strong> and <strong>David Li</strong> and merged into LLVM over the following year, is the answer that made link-time optimization practical at industrial scale. Each module is compiled to bitcode alongside a compact <em>summary</em>: a thin index of its call graph and symbols. </p><p>A cheap global step reads only the summaries, decides cross-module actions such as which functions to import for inlining, and then the heavy per-module optimization runs in parallel, importing only the specific functions each module needs. It recovers<strong> most of full LTO&#8217;s benefit</strong> at a fraction of the memory, and it parallelizes. The numbers are quite stark. </p><p>Full LTO can inflate compile time two to three fold and demands that the whole program fit in a single process; ThinLTO&#8217;s per-module summaries add under one percent to bitcode size,<strong> roughly 0.8 percent for Clang </strong>built without debug information, and keep build time and peak memory close to a non-LTO build. </p><p>On some benchmarks ThinLTO even edges out full LTO, because its scalable design lets each module run a more aggressive backend pipeline than a single monolithic module could afford. </p><p>LLVM 22 is now upstreaming <strong>Distributed ThinLTO</strong> (DTLTO), which pushes those parallel backend jobs across a build cluster, with caching for incremental rebuilds. Turning it on is a flag.</p><pre><code>clang -flto=thin -O2 a.c b.c c.c -o app   <em><span># scalable cross-module LTO</span></em>
clang -flto      -O2 a.c b.c c.c -o app   <em><span># full (monolithic) LTO</span></em></code></pre><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!UmjI!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5931e02a-ba05-45f2-b860-b9884c2c7654_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!UmjI!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5931e02a-ba05-45f2-b860-b9884c2c7654_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!UmjI!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5931e02a-ba05-45f2-b860-b9884c2c7654_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!UmjI!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5931e02a-ba05-45f2-b860-b9884c2c7654_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!UmjI!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5931e02a-ba05-45f2-b860-b9884c2c7654_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!UmjI!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5931e02a-ba05-45f2-b860-b9884c2c7654_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/5931e02a-ba05-45f2-b860-b9884c2c7654_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Horizontal bar chart of compile time relative to a normal build: No LTO 1.0x, ThinLTO about 1x, Full LTO 2x to 3x.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Horizontal bar chart of compile time relative to a normal build: No LTO 1.0x, ThinLTO about 1x, Full LTO 2x to 3x." title="Horizontal bar chart of compile time relative to a normal build: No LTO 1.0x, ThinLTO about 1x, Full LTO 2x to 3x." srcset="https://substackcdn.com/image/fetch/$s_!UmjI!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5931e02a-ba05-45f2-b860-b9884c2c7654_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!UmjI!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5931e02a-ba05-45f2-b860-b9884c2c7654_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!UmjI!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5931e02a-ba05-45f2-b860-b9884c2c7654_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!UmjI!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5931e02a-ba05-45f2-b860-b9884c2c7654_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Figure 3.</strong> Why ThinLTO exists. Full LTO buys cross-module optimization at 2 to 3 times the compile time and needs the whole program in one process; ThinLTO adds under 1 percent to bitcode, stays close to a normal build, and parallelizes. Only the 2 to 3x figure is a hard number (Gentoo LTO documentation and the LLVM ThinLTO write-up); the ThinLTO bar is indicative of its near-baseline behavior.</figcaption></figure></div><p>This is the mechanism underneath one of the headline claims: that all of <strong>Google&#8217;s and Meta&#8217;s datacenter C and C++ </strong>is built with LLVM. It is not merely that Clang compiles the files; it is that link-time optimization across an entire service is tractable because <em>ThinLTO made it so.</em></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/p/how-llvm-works-the-ir-that-took-over?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/p/how-llvm-works-the-ir-that-took-over?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><div><hr></div><h2>What LLVM actually runs</h2><p>The reach, with the adoption dates checked against primary sources</p><p>The scale claims in the retrospective are large, and they hold up. LLVM long ago<strong> replaced GCC</strong> as the foundation of Apple&#8217;s software across every device line. Google uses it for Android and the Android NDK. Applications on Qualcomm&#8217;s Snapdragon processors are built in substantial part with LLVM. </p><p>The<strong> two dominant CPU vendors</strong> for PCs and servers, Arm and Intel, retired their proprietary compilers, ARMCC and ICC, in favor of LLVM-based toolchains. Google&#8217;s and Meta&#8217;s datacenter C and C++ is compiled with it. Sony&#8217;s PlayStation 4 and 5 and the Nintendo Switch use Clang as their primary application toolchain. </p><p>And the AI stack leans on it heavily, through NVIDIA&#8217;s CUDA compiler, through the Triton compiler used by PyTorch and OpenAI, and most broadly through <strong>MLIR</strong>.</p><p>The adoption did not happen at once, and the sequence matters, because it shows a single seed (<em>Apple, via Lattner&#8217;s hiring</em>) catalyzing an industry.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Kmcg!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe9e30f2d-ab4b-435a-b82b-8b5f16bf82ef_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Kmcg!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe9e30f2d-ab4b-435a-b82b-8b5f16bf82ef_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!Kmcg!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe9e30f2d-ab4b-435a-b82b-8b5f16bf82ef_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!Kmcg!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe9e30f2d-ab4b-435a-b82b-8b5f16bf82ef_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!Kmcg!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe9e30f2d-ab4b-435a-b82b-8b5f16bf82ef_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Kmcg!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe9e30f2d-ab4b-435a-b82b-8b5f16bf82ef_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e9e30f2d-ab4b-435a-b82b-8b5f16bf82ef_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Timeline of LLVM industry adoption from 2005 to 2022: Apple begins, macOS OpenGL JIT ships, Clang open-sourced, PS4 ships with Clang, Apple ships bitcode, ARM and Intel adopt LLVM, Apple deprecates bitcode.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Timeline of LLVM industry adoption from 2005 to 2022: Apple begins, macOS OpenGL JIT ships, Clang open-sourced, PS4 ships with Clang, Apple ships bitcode, ARM and Intel adopt LLVM, Apple deprecates bitcode." title="Timeline of LLVM industry adoption from 2005 to 2022: Apple begins, macOS OpenGL JIT ships, Clang open-sourced, PS4 ships with Clang, Apple ships bitcode, ARM and Intel adopt LLVM, Apple deprecates bitcode." srcset="https://substackcdn.com/image/fetch/$s_!Kmcg!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe9e30f2d-ab4b-435a-b82b-8b5f16bf82ef_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!Kmcg!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe9e30f2d-ab4b-435a-b82b-8b5f16bf82ef_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!Kmcg!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe9e30f2d-ab4b-435a-b82b-8b5f16bf82ef_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!Kmcg!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe9e30f2d-ab4b-435a-b82b-8b5f16bf82ef_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Figure 4.</strong> Adoption milestones with dates verified against primary sources. Apple began serious LLVM work around 2005; the first shipping product was the macOS OpenGL JIT in 2006. The App Store accepted app bitcode from the mid-2010s until Apple deprecated it in Xcode 14 (2022).</figcaption></figure></div><p>The bitcode chapter is the sharpest illustration of the &#8220;<em>lifelong compilation</em>&#8221; idea reaching production. For years, developers shipped iOS,<strong> watchOS </strong>and <strong>tvOS apps</strong> to Apple&#8217;s App Store not as machine code but as LLVM bitcode, and Apple compiled and specialized each app for the target device on its side. </p><p>That let one submission serve <strong>32-bit and 64-bit Arm iPhones</strong>, low-power Apple Watch variants and Apple TV, shrank downloads, and let Apple roll out new compiler optimizations without developers rebuilding. </p><p>Apple deprecated bitcode in <a href="https://developer.apple.com/documentation/xcode-release-notes/xcode-14-release-notes">Xcode 14 in 2022</a>, once its hardware had converged on Arm64 and the flexibility was no longer worth the friction. The feature retired, but it stands as the largest deployment of shippable IR the industry has seen.</p><h3>Sizing the surface area</h3><p>The authors put a figure on all this: software compiled with LLVM, they write, represents<strong> hundreds of billions of dollars</strong> in annual revenue across mobile, cloud, desktop, server, gaming, supercomputing and AI. </p><p>We think that undercounts it, and it is worth being precise about what can actually be measured, because &#8220;<em>revenue compiled by LLVM</em>&#8221; is not a single clean quantity. </p><p>The honest move is to measure the top-line revenue of the platforms whose flagship products are demonstrably built with LLVM, treat it as a floor rather than an attribution, and be explicit that this is revenue riding on <strong>LLVM-compiled software</strong>, not value the compiler captures.</p><p>Do that and the picture is not subtle. In their most recent full fiscal years three companies alone, Apple, Alphabet and NVIDIA, booked <strong>about 1.04 trillion dollars in combined revenue</strong>, and the flagship products behind every one of those dollars are built with LLVM: Apple&#8217;s entire device line through Clang and Swift, Android and Google&#8217;s hyperscale C and C++ through Clang and ThinLTO, and the AI-accelerator stack through the <strong>CUDA</strong>, <strong>Triton </strong>and <strong>MLIR compilers</strong> that lower into LLVM.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!ekQK!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd8d54fdb-6d29-4e59-9d46-2d34403a05b9_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!ekQK!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd8d54fdb-6d29-4e59-9d46-2d34403a05b9_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!ekQK!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd8d54fdb-6d29-4e59-9d46-2d34403a05b9_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!ekQK!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd8d54fdb-6d29-4e59-9d46-2d34403a05b9_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!ekQK!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd8d54fdb-6d29-4e59-9d46-2d34403a05b9_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!ekQK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd8d54fdb-6d29-4e59-9d46-2d34403a05b9_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d8d54fdb-6d29-4e59-9d46-2d34403a05b9_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Horizontal bar chart of most-recent-fiscal-year revenue for Apple (416 billion dollars), Alphabet (403 billion dollars) and NVIDIA (216 billion dollars), summing to about 1.04 trillion dollars, all with flagship products built using LLVM.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Horizontal bar chart of most-recent-fiscal-year revenue for Apple (416 billion dollars), Alphabet (403 billion dollars) and NVIDIA (216 billion dollars), summing to about 1.04 trillion dollars, all with flagship products built using LLVM." title="Horizontal bar chart of most-recent-fiscal-year revenue for Apple (416 billion dollars), Alphabet (403 billion dollars) and NVIDIA (216 billion dollars), summing to about 1.04 trillion dollars, all with flagship products built using LLVM." srcset="https://substackcdn.com/image/fetch/$s_!ekQK!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd8d54fdb-6d29-4e59-9d46-2d34403a05b9_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!ekQK!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd8d54fdb-6d29-4e59-9d46-2d34403a05b9_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!ekQK!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd8d54fdb-6d29-4e59-9d46-2d34403a05b9_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!ekQK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd8d54fdb-6d29-4e59-9d46-2d34403a05b9_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Figure 5.</strong> Three companies whose flagship products are compiled with LLVM, and their most recent full fiscal year of top-line revenue: Apple 416 billion dollars (FY2025), Alphabet 403 billion dollars (2025) and NVIDIA 216 billion dollars (fiscal 2026, ended January 25, 2026). The combined 1.04 trillion dollars is a deliberate floor: it excludes Meta, Microsoft, Qualcomm, Sony, Nintendo and the entire Android hardware market beyond Google. Figures from each company&#8217;s SEC filings.</figcaption></figure></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!J6kW!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f510ecd-d1f5-45cf-9c10-a4522e40c128_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!J6kW!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f510ecd-d1f5-45cf-9c10-a4522e40c128_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!J6kW!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f510ecd-d1f5-45cf-9c10-a4522e40c128_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!J6kW!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f510ecd-d1f5-45cf-9c10-a4522e40c128_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!J6kW!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f510ecd-d1f5-45cf-9c10-a4522e40c128_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!J6kW!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f510ecd-d1f5-45cf-9c10-a4522e40c128_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2f510ecd-d1f5-45cf-9c10-a4522e40c128_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:192104,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/203533011?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f510ecd-d1f5-45cf-9c10-a4522e40c128_1634x893.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!J6kW!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f510ecd-d1f5-45cf-9c10-a4522e40c128_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!J6kW!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f510ecd-d1f5-45cf-9c10-a4522e40c128_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!J6kW!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f510ecd-d1f5-45cf-9c10-a4522e40c128_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!J6kW!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f510ecd-d1f5-45cf-9c10-a4522e40c128_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>We <strong>do not sum the rows of that table</strong>: the smartphone market overlaps the Apple and Alphabet figures, and the point is the surface area, not a grand total. </p><p>But the<strong> three company revenues</strong> in Figure 4 are independently additive and clear a trillion dollars between them, which is why we are comfortable calling <strong>LLVM&#8217;s reach a trillion-dollar one. </strong>The precise figure is unauditable and we mark it accordingly in the dossier; the order of magnitude is not in doubt. </p><p>A large majority of end users across computing are, at some layer, running code that LLVM compiled.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Tq0r!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F27ec6d5c-6d7e-4659-845f-0cc80f75014b_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Tq0r!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F27ec6d5c-6d7e-4659-845f-0cc80f75014b_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!Tq0r!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F27ec6d5c-6d7e-4659-845f-0cc80f75014b_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!Tq0r!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F27ec6d5c-6d7e-4659-845f-0cc80f75014b_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!Tq0r!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F27ec6d5c-6d7e-4659-845f-0cc80f75014b_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Tq0r!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F27ec6d5c-6d7e-4659-845f-0cc80f75014b_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/27ec6d5c-6d7e-4659-845f-0cc80f75014b_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Horizontal bar chart of active devices running LLVM-compiled system software: Android more than 3 billion, Apple 2.5 billion, PlayStation and Switch several hundred million.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Horizontal bar chart of active devices running LLVM-compiled system software: Android more than 3 billion, Apple 2.5 billion, PlayStation and Switch several hundred million." title="Horizontal bar chart of active devices running LLVM-compiled system software: Android more than 3 billion, Apple 2.5 billion, PlayStation and Switch several hundred million." srcset="https://substackcdn.com/image/fetch/$s_!Tq0r!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F27ec6d5c-6d7e-4659-845f-0cc80f75014b_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!Tq0r!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F27ec6d5c-6d7e-4659-845f-0cc80f75014b_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!Tq0r!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F27ec6d5c-6d7e-4659-845f-0cc80f75014b_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!Tq0r!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F27ec6d5c-6d7e-4659-845f-0cc80f75014b_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Figure 6.</strong> The same reach, counted in devices rather than dollars, which is the more defensible floor. More than 3 billion active Android devices (Google&#8217;s own stated figure), 2.5 billion active Apple devices (January 2026), and several hundred million PlayStation and Switch consoles, all running system software built with LLVM. More than 5.8 billion devices, before the Android hardware Google does not build or the datacenters they run against.</figcaption></figure></div><h3><strong>The academic and language dividend</strong></h3><p>Two quieter forms of impact compound over time. LLVM underlies a generation of programming languages that an easy-to-target, modular back end made feasible: <strong>Swift</strong>, <strong>Rust</strong>, <strong>Julia</strong>, <strong>Halide</strong> and <strong>Mojo</strong> among them.</p><p> Rust&#8217;s memory-safety guarantees and Julia&#8217;s JIT-driven scientific performance both rest on LLVM code generation. </p><p>And in the academy, LLVM democratized compiler research and teaching: a graduate student can write a novel pass in a day and test it on large real programs immediately, and the authors report meeting undergraduates who arrived at college <strong>having already built compiler projects</strong> on LLVM in high school, &#8220;<em>for fun.</em>&#8221;</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/p/how-llvm-works-the-ir-that-took-over/comments&quot;,&quot;text&quot;:&quot;Leave a comment&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/p/how-llvm-works-the-ir-that-took-over/comments"><span>Leave a comment</span></a></p><div><hr></div><h2>MLIR, and the compiler layer under modern AI</h2><p>When one IR was not enough, the answer was a framework for making IRs</p><p>This is the part that matters most for anyone working on inference infrastructure, so we will spend a moment on it. <strong>LLVM IR is deliberately low-level,</strong> which is a strength for representing machines and a weakness for representing a matrix multiply. </p><p>By the time a tensor operation has been shredded into scalar loads, adds and branches, the structure a compiler most wants to exploit, the fact that this is a <strong>dense linear-algebra kernel</strong>, is gone. That loss is fine for C. It is expensive for machine learning, where the highest-value optimizations (tiling, fusion, layout, targeting a systolic array) live at exactly the level LLVM IR throws away.</p><p><strong>MLIR</strong>, the Multi-Level Intermediate Representation, introduced in 2019 and presented at<strong> CGO 2021</strong>, is the response, and it is characteristically an infrastructure answer rather than a point solution. Instead of one IR, MLIR provides the machinery to define many coexisting IRs, called <em>dialects</em>, that all share the same structural substrate of operations, regions and blocks. </p><p>A compiler keeps a computation at a high level of abstraction as long as that is useful, then <em>progressively lowers</em> it, dialect by dialect, until it reaches the <code>llvm</code> dialect and finally <strong>ordinary LLVM IR</strong>, where the mature backends take over. A tensor program can begin like this:</p><pre><code><em><span>// a matmul on tensors, before any lowering</span></em>
<strong><span>func.func</span></strong> <span>@matmul</span>(%A: <span>tensor</span>&lt;128x256xf32&gt;,
                  %B: <span>tensor</span>&lt;256x512xf32&gt;,
                  %C: <span>tensor</span>&lt;128x512xf32&gt;) -&gt; <span>tensor</span>&lt;128x512xf32&gt; {
  %0 = <strong><span>linalg.matmul</span></strong>
         ins(%A, %B : <span>tensor</span>&lt;128x256xf32&gt;, <span>tensor</span>&lt;256x512xf32&gt;)
         outs(%C   : <span>tensor</span>&lt;128x512xf32&gt;) -&gt; <span>tensor</span>&lt;128x512xf32&gt;
  <strong><span>return</span></strong> %0 : <span>tensor</span>&lt;128x512xf32&gt;
}</code></pre><p>From there the same operation descends through structured and loop dialects (<code>linalg</code>, <code>affine</code>, <code>scf</code>, <code>tensor</code>), gets its buffers assigned (<code>memref</code>, <code>bufferization</code>, <code>vector</code>), is<strong> retargeted to a device dialect</strong> (<code>gpu</code>, <code>nvvm</code>, <code>rocdl</code>, <code>spirv</code>), and finally lands in the <code>llvm</code> dialect and LLVM IR. </p><p>Each step is a rewrite in the same framework, and each step is where a tiling or fusion decision can be made while the structure is still visible.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!fL2Q!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5e84afc1-9774-4a35-9342-23999d8f60d4_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!fL2Q!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5e84afc1-9774-4a35-9342-23999d8f60d4_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!fL2Q!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5e84afc1-9774-4a35-9342-23999d8f60d4_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!fL2Q!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5e84afc1-9774-4a35-9342-23999d8f60d4_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!fL2Q!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5e84afc1-9774-4a35-9342-23999d8f60d4_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!fL2Q!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5e84afc1-9774-4a35-9342-23999d8f60d4_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/5e84afc1-9774-4a35-9342-23999d8f60d4_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Layered stack showing MLIR progressive lowering from framework graph through high-level dialects, structured and loop dialects, buffer and vector dialects, target dialects, the LLVM dialect, LLVM IR, and finally target machine code.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Layered stack showing MLIR progressive lowering from framework graph through high-level dialects, structured and loop dialects, buffer and vector dialects, target dialects, the LLVM dialect, LLVM IR, and finally target machine code." title="Layered stack showing MLIR progressive lowering from framework graph through high-level dialects, structured and loop dialects, buffer and vector dialects, target dialects, the LLVM dialect, LLVM IR, and finally target machine code." srcset="https://substackcdn.com/image/fetch/$s_!fL2Q!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5e84afc1-9774-4a35-9342-23999d8f60d4_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!fL2Q!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5e84afc1-9774-4a35-9342-23999d8f60d4_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!fL2Q!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5e84afc1-9774-4a35-9342-23999d8f60d4_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!fL2Q!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5e84afc1-9774-4a35-9342-23999d8f60d4_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Figure 7.</strong> Progressive lowering in MLIR. The compiler holds tensor-level structure at the top and converts it step by step into the same LLVM IR that Clang and Rust emit, then to PTX, x86-64, AArch64 or RISC-V. The stack narrows toward machine code to suggest the convergence onto shared infrastructure.</figcaption></figure></div><p>For an inference audience the payoff of this layering has a name: fusion. Because <strong>progressive lowering</strong> keeps tensor-level structure visible until late, the compiler can fuse a chain of operations, a matmul followed by a bias add followed by an activation, into a single kernel, so the intermediate tensors are never written out to high-bandwidth memory and read back. </p><p>That matters because decode, the<strong> token-by-token phase</strong> of LLM serving, is bound by memory bandwidth rather than arithmetic: the accelerator spends its time moving weights and activations, not multiplying them. </p><p>Collapsing those <strong>HBM round trips</strong> is therefore one of the largest levers on tokens per second per GPU, and it is a compiler transformation, made at exactly the abstraction level that raw LLVM IR throws away. </p><p>This is the concrete reason every serious inference stack now ships a compiler rather than a<strong> fixed library of kernels</strong>, and why the layer we are describing sits directly upstream of inference cost.</p><p>This is not a research curiosity. MLIR is the backbone of TensorFlow&#8217;s XLA compiler and the broader <strong>OpenXLA ecosystem</strong>, of parts of PyTorch through <strong>Torch-MLIR</strong>, of IREE, and of OpenAI&#8217;s Triton, the kernel language a great deal of contemporary GPU work is written in. Even components of NVIDIA&#8217;s own CUDA toolkit use it. </p><p>And the<strong> underlying LLVM backend</strong> still does the final, unglamorous work of register allocation and scheduling for all of them. When people say the AI stack runs on LLVM, this is the mechanism they are pointing at: not that models are<strong> written in C</strong>, but that the compilers turning tensor graphs into GPU code are, almost universally, MLIR lowering into LLVM. </p><p>The hardware-design world is following the same path through <strong>CIRCT</strong>, which extends MLIR into digital circuit design.</p><div><hr></div><h2>The one thing LLVM is bad at</h2><p>Compile speed, and the backends being built to route around it</p><p>A retrospective written by the project&#8217;s founders is not the place to look for LLVM&#8217;s weaknesses, so we will supply the missing side. </p><p>LLVM is engineered for the quality of the code it emits, and it pays for that in the time it takes to emit it. Even with optimizations disabled the <strong>pipeline is heavy</strong>, and for a developer sitting in an edit, compile, run loop the wait is a real tax. </p><p>As bjorn3, the author of Rust&#8217;s <strong>Cranelift</strong> backend, puts it, LLVM is optimized for output quality at the cost of compilation speed even when optimizations are turned off. This is the most common and most legitimate complaint about the project, and two major language communities are now building their own code generators specifically to escape it.</p><p>Rust took the incremental path. Its alternative <code>rustc_codegen_cranelift</code> <strong>backend uses Cranelift</strong>, a code generator built for speed rather than peak output and aimed squarely at debug builds. </p><p>The<strong> Rust team&#8217;s 2025 measurements</strong> on large real projects, among them Zed, Tauri and hickory-dns, show roughly a twenty percent cut in code-generation time and about a five percent speedup in total clean-build time, and getting the backend production-ready for local development is an official<strong> 2025 project goal. </strong></p><p>The bargain is stated plainly: Cranelift produces code almost as fast as LLVM with optimizations off, in exchange for compiling far more quickly.</p><p><strong>Zig</strong> went further and wrote its own backends outright. In mid-2025 it made its self-hosted x86-64 backend the default for debug builds on Linux and macOS, <strong>bypassing LLVM entirely</strong> on that path, and reported wall-clock compile improvements ranging from five to fifty percent across real projects. </p><p>The rationale is a direct indictment: because compilation time was dominated by LLVM, and because <strong>LLVM&#8217;s heavy use of shared state</strong> kept Zig&#8217;s LLVM path effectively single-threaded while its own backend parallelized code generation across cores, the only way to get materially faster was to leave. </p><p>Notably, <strong>Zig still emits LLVM bitcode </strong>when it wants LLVM&#8217;s optimizer for release builds. That is the pattern worth seeing.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!CMmQ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F05f0cce8-c325-473e-b99a-30f557519b34_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!CMmQ!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F05f0cce8-c325-473e-b99a-30f557519b34_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!CMmQ!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F05f0cce8-c325-473e-b99a-30f557519b34_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!CMmQ!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F05f0cce8-c325-473e-b99a-30f557519b34_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!CMmQ!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F05f0cce8-c325-473e-b99a-30f557519b34_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!CMmQ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F05f0cce8-c325-473e-b99a-30f557519b34_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/05f0cce8-c325-473e-b99a-30f557519b34_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Horizontal bar chart of compile-time improvement from non-LLVM backends: Rust Cranelift about 20 percent on codegen and 5 percent on clean build, Zig self-hosted backend 5 to 50 percent wall clock.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Horizontal bar chart of compile-time improvement from non-LLVM backends: Rust Cranelift about 20 percent on codegen and 5 percent on clean build, Zig self-hosted backend 5 to 50 percent wall clock." title="Horizontal bar chart of compile-time improvement from non-LLVM backends: Rust Cranelift about 20 percent on codegen and 5 percent on clean build, Zig self-hosted backend 5 to 50 percent wall clock." srcset="https://substackcdn.com/image/fetch/$s_!CMmQ!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F05f0cce8-c325-473e-b99a-30f557519b34_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!CMmQ!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F05f0cce8-c325-473e-b99a-30f557519b34_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!CMmQ!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F05f0cce8-c325-473e-b99a-30f557519b34_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!CMmQ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F05f0cce8-c325-473e-b99a-30f557519b34_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Figure 8.</strong> The cost of LLVM&#8217;s code quality, measured. Rust&#8217;s Cranelift backend cuts code-generation time by about 20 percent and clean-build time by about 5 percent (Rust project goals, 2025); Zig&#8217;s self-hosted x86-64 backend, default for debug builds since mid-2025, reports 5 to 50 percent faster wall-clock compiles (Zig devlog, 2025). The three bars measure different denominators. Both ecosystems keep LLVM for optimized release builds.</figcaption></figure></div><p><em>Keep LLVM for the optimized release build, where its code quality is still unmatched. Route around it for the fast debug build, where its latency hurts.</em></p><h3>The shape of both the Rust and the Zig response to LLVM&#8217;s one real weakness</h3><p>That shared pattern is not a threat to the empire so much as a map of its single soft border, and it is worth being honest that the border exists. There is a second, quieter tension living inside LLVM&#8217;s own success.</p><p><strong>MLIR&#8217;s openness</strong> invites a proliferation of dialects, and a representation that can be anything risks fragmenting into many incompatible somethings; the whole force of the original LLVM argument was that the IR won because it was <em>one</em> thing. </p><p><em>&#8220;The IR won because it was one representation</em>&#8221; and &#8220;<em>the future is many coexisting representations</em>&#8221; are in genuine tension, and how MLIR holds its dialects together as<strong> their number grows</strong> is one of the more interesting open questions in the field. So far the shared substrate has held. </p><p>The bet is that it keeps holding.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Where LLVM goes next</h2><p>Verified compilation, auto-generated backends, machine-learned heuristics, hardened code</p><p>The retrospective closes on active research, and several threads are worth stating precisely because they point at the next decade of the project.</p><h3>Proving the optimizer correct</h3><p>A compiler bug silently miscompiles correct source into wrong machine code, and optimizers are where these bugs breed. </p><p><strong>Alive2</strong> <em>(Nuno Lopes and colleagues, PLDI 2021)</em> attacks this with bounded translation validation: it takes an IR function before and after a transformation and uses an SMT solver to check whether the rewrite is sound for all inputs, <strong>encoding LLVM&#8217;s tricky semantics</strong>, including undefined behavior and <code>poison</code>, into the query. </p><p>Alive2 runs continuously against LLVM and has found real miscompilations in core passes such as <code>InstCombine</code>. The <code>ptrtoaddr</code> instruction we met in section 02 is part of the same trend: the <strong>project is steadily sharpening</strong> its own semantics so that this kind of verification can go further.</p><h3>Generating the compiler from the chip</h3><p>Two projects challenge the assumption that backends and peephole optimizations must be written by hand. <strong>Hydride</strong> (ASPLOS 2024) and <strong>MISAAL</strong> (<em>published in the ACM&#8217;s Proceedings on Programming Languages, 2025</em>) use program synthesis to <strong>generate compiler components</strong>, target-independent IR operations, front-end translators and backend code generators, directly from a vendor&#8217;s formal specification of an instruction set. </p><p>The reported result is striking: <strong>auto-generated compilers</strong> that match or beat a heavily hand-engineered production compiler for Halide while requiring roughly an order of magnitude less manual effort. </p><p>As instruction sets multiply, matrix and vector extensions arriving on <strong>nearly every new chip, </strong>generating the compiler from the specification rather than writing it by hand is a plausible future for the backend.</p><h3>Machine learning inside the compiler, and about the compiler</h3><p>There are two distinct AI-and-compiler stories, and they are easy to confuse. The first is <strong>machine learning inside LLVM</strong>: Google&#8217;s MLGO framework replaces hand-tuned heuristics for inlining-for-size and register-allocation eviction with policies trained by reinforcement learning, and these have shipped in production builds. </p><p>The second is <strong>LLVM as training data for models about code</strong>. Meta&#8217;s <strong>LLM Compiler</strong> (2024), built on Code Llama and published at Compiler Construction 2025, was trained on a corpus of <strong>546 billion tokens</strong> of LLVM IR and assembly, then instruction-tuned to predict the effect of optimization passes; the flag-tuning and disassembly variants add a further 164 billion tokens for 710 billion in total. </p><p>Meta reports the model reaching <strong>77 percent of the optimization benefit</strong> of an autotuning search, and a 45 percent round-trip on disassembling machine code back to IR. </p><p>The reason such a model can exist at all is the reason <strong>LLVM matters everywhere else</strong>: the sheer volume of LLVM-compiled open source makes a compiler-scale training corpus possible.</p><h4>One number, corrected</h4><p>The CACM retrospective states the LLM Compiler was fine-tuned on &#8220;<em>537 billion tokens.</em>&#8221; The figure in Meta&#8217;s paper and in the peer-reviewed Compiler Construction 2025 version is <strong>546 billion tokens</strong> of LLVM IR and assembly. We use 546 billion, which is what the primary source reports.</p><h3>Hardened code and heterogeneous parallelism</h3><p>Two more directions round out the picture. </p><p>On security, LLVM and Clang are the delivery vehicle for a generation of hardening: forward-edge control-flow integrity, shadow call stacks, and the Arm hardware defenses of pointer authentication (<em>PAC</em>), branch target identification (<em>BTI</em>) and memory tagging (<em>MTE</em>), alongside the sanitizer family (<em>AddressSanitizer, ThreadSanitizer, MemorySanitizer, UndefinedBehaviorSanitizer</em>) that made whole classes of C and C++ bugs findable. </p><p>LLVM 22 even added the <strong>AArch64 build-attribute machinery</strong> that lets linkers reason about PAC and BTI compatibility across object files. </p><p>On parallelism, the <strong>HPVM</strong> project (IEEE Micro 2022) layers a hierarchical dataflow graph over LLVM IR to target heterogeneous systems, CPUs, GPUs and accelerators, from a single representation, and national laboratories are extending LLVM for exascale machines through efforts such as <strong>Los Alamos&#8217;s Kitsune </strong>and the NNSA-backed Flang Fortran front end.</p><div class="captioned-button-wrap" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/p/how-llvm-works-the-ir-that-took-over?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="CaptionedButtonToDOM"><div class="preamble"><p class="cta-caption">Thanks for reading The Software Frontier! This post is public so feel free to share it.</p></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/p/how-llvm-works-the-ir-that-took-over?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/p/how-llvm-works-the-ir-that-took-over?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p></div><div><hr></div><h2>What the story is actually about</h2><p>Cadence, and the transferable lesson under the technology</p><p>Strip away the specifics and LLVM is a case study in how foundational infrastructure gets built, and the ingredients are unglamorous. Federal money funded <strong>early-stage research </strong>whose payoff no one could forecast, on the condition that the results be released openly. </p><p>An academic effort took known techniques, SSA, separation of concerns, retargetable code generation, interprocedural analysis, and applied all of them together more thoroughly than any production compiler had. </p><p>A commercially neutral open license, eventually <strong>Apache 2.0 with LLVM exceptions</strong>, let the result be reused, again and again, in products that competed with each other. Strong technical and community leadership, first inside Apple and then across the industry, drove adoption. </p><p>And a family of successful derivatives, Clang, Swift and MLIR chief among them, widened the reach far beyond the original core.</p><p>The result compounds on a schedule. Twenty-three years after the first release, <strong>LLVM ships a full toolchain</strong>, Clang, LLD, LLDB, libc++, compiler-rt, MLIR, Flang, Polly, BOLT, on a strict six-month train.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!bIew!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F490824b8-4388-44ed-9df1-c0f1d0380112_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!bIew!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F490824b8-4388-44ed-9df1-c0f1d0380112_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!bIew!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F490824b8-4388-44ed-9df1-c0f1d0380112_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!bIew!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F490824b8-4388-44ed-9df1-c0f1d0380112_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!bIew!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F490824b8-4388-44ed-9df1-c0f1d0380112_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!bIew!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F490824b8-4388-44ed-9df1-c0f1d0380112_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/490824b8-4388-44ed-9df1-c0f1d0380112_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Stepped timeline of LLVM major versions from 1.0 in 2003 to 22 in 2026, showing an irregular early period and a steady six-month cadence after version 4 in 2017, annotated with IR milestones.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Stepped timeline of LLVM major versions from 1.0 in 2003 to 22 in 2026, showing an irregular early period and a steady six-month cadence after version 4 in 2017, annotated with IR milestones." title="Stepped timeline of LLVM major versions from 1.0 in 2003 to 22 in 2026, showing an irregular early period and a steady six-month cadence after version 4 in 2017, annotated with IR milestones." srcset="https://substackcdn.com/image/fetch/$s_!bIew!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F490824b8-4388-44ed-9df1-c0f1d0380112_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!bIew!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F490824b8-4388-44ed-9df1-c0f1d0380112_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!bIew!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F490824b8-4388-44ed-9df1-c0f1d0380112_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!bIew!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F490824b8-4388-44ed-9df1-c0f1d0380112_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Figure 9.</strong> Every LLVM major version against its release date, with IR milestones marked: the new pass manager becoming default (v13), opaque pointers by default (v15), typed pointers removed (v17) and <strong>ptrtoaddr</strong> added (v22). The clear elbow at version 4 in 2017 is where the project settled into a predictable half-year release cadence.</figcaption></figure></div><p><strong>Adve </strong>and <strong>Lattner </strong>end their retrospective on the claim that a small, federally funded academic project produced &#8220;<em>world-changing, foundational infrastructure,</em>&#8221; and on the evidence it is not an overstatement. </p><p>The instruction set that Lattner sketched over a winter break in 2000, three isomorphic forms of one strongly typed, target-independent, shippable IR, is visibly the <strong>same architecture running underneath the compiler</strong> in every phone, console, datacenter and AI accelerator we have discussed. </p><p>The empire is real. It was always, underneath, an argument about a representation.</p><div><hr></div><h2>Verification dossier</h2><p>Every load-bearing claim above, checked against a primary or authoritative source. Confidence tiers: <strong>A</strong> primary source, direct confirmation &#183; <strong>B</strong> authoritative secondary &#183; <strong>C</strong> reasonable inference or the authors&#8217; own estimate &#183; <strong>D</strong> flagged discrepancy or figure we corrected.</p><p><strong><span>A. LLVM 1.0 was first released in October 2003</span></strong></p><p>Chris Lattner&#8217;s announcement, &#8220;The LLVM 1.0 Release is finally available!&#8221;, is dated October 24, 2003; Wikidata records the same inception date. <strong><span>Discrepancy noted:</span></strong> the CACM retrospective says &#8220;October 2003&#8221; in its introduction but &#8220;December 2003&#8221; in two later places. October 24, 2003 is correct; the December date appears to be the release-notes page&#8217;s last-modified timestamp (Dec 8, 2003), not the release.</p><p><strong><span>A. Current stable release is LLVM 22.1, from February 2026</span></strong></p><p>LLVM/Clang 22.1 was released February 24, 2026 per the project and Phoronix; the latest point release, 22.1.8, is dated June 16, 2026 on GitHub. LLVM 23 is in development as 23.0.0git. Consistent with the project&#8217;s post-2017 six-month cadence.</p><p><strong><span>A. The CGO 2004 paper won Most Influential Paper; the team won the 2012 ACM Software System Award</span></strong></p><p>&#8220;LLVM: A Compilation Framework for Lifelong Program Analysis and Transformation&#8221; (CGO 2004, pp. 75 to 86, DOI 10.1109/CGO.2004.1281665) was named Most Influential Paper of CGO 2004, awarded 2014. <strong><span>Nuance:</span></strong> the 2012 ACM Software System Award went to Adve, Lattner <strong>and Evan Cheng</strong>; the CACM retrospective names only the two authors. The &#8220;more than 8,000 citations&#8221; figure is the authors&#8217; own and is conservative against public citation indices.</p><p><strong><span>D. The Meta LLM Compiler was trained on 546 billion tokens, not 537 billion</span></strong></p><p>The CACM piece states 537 billion tokens. Meta&#8217;s paper (arXiv 2407.02524) and the peer-reviewed Compiler Construction 2025 version (DOI 10.1145/3708493.3712691) both report <strong>546 billion tokens</strong> of LLVM IR and assembly, with a further 164 billion for flag-tuning and disassembly (710 billion total). We use 546 billion.</p><p><strong><span>A. Opaque pointers became default in LLVM 15; typed pointers were removed in LLVM 17</span></strong></p><p>Confirmed by the official Opaque Pointers documentation: default in LLVM 15 (2022), typed-pointer support removed in LLVM 17 (2023). The <code>ptrtoaddr</code> instruction is a genuine LLVM 22 addition, per the 22.1 release coverage.</p><p><strong><span>A. Apple deprecated bitcode in Xcode 14 (2022)</span></strong></p><p>Xcode 14 release notes: the App Store no longer accepts bitcode submissions and Xcode no longer builds bitcode by default. This bounds the &#8220;mid-2010s through 2022&#8221; window the retrospective gives for App Store bitcode.</p><p><strong><span>B. Industry adoption: Apple, Google, Arm, Intel, Sony, Nintendo, Qualcomm</span></strong></p><p>Apple&#8217;s OpenGL/iOS SDK work began around 2005 and the macOS OpenGL JIT shipped in 2006 (Clang and LLVM histories). Intel switched its C/C++ stack to LLVM (ICX / oneAPI DPC++); Arm shipped LLVM-based Arm Compiler 6 and drove the Apache 2.0 relicensing over patent concerns. Sony uses Clang for PS4/PS5. These are well documented across primary announcements and vendor pages.</p><p><strong><span>C. &#8221;Hundreds of billions of dollars in annual revenue&#8221; compiled by LLVM</span></strong></p><p>This is the authors&#8217; order-of-magnitude estimate, not an audited figure, and we present it as such. The qualitative claim, that a large majority of end users run LLVM-compiled code at some layer, is well supported by the adoption evidence above.</p><p><strong><span>A. MLIR underpins XLA, PyTorch (Torch-MLIR), IREE and Triton; GlobalISel is default at -O0 on AArch64</span></strong></p><p>MLIR was presented at CGO 2021 (DOI 10.1109/CGO51591.2021.9370308) and is used across the named AI compilers. GlobalISel&#8217;s default-at-O0 status on AArch64 and ongoing work to broaden it are documented in current LLVM forum threads (February 2026).</p><p><strong><span>B. ThinLTO: first presented at EuroLLVM 2015; summaries add under one percent to bitcode; full LTO costs two to three times the compile time</span></strong></p><p>ThinLTO was introduced by Teresa Johnson, Mehdi Amini and David Li (LLVM Project blog, 2016; presented EuroLLVM 2015). The summary overhead of about 0.8 percent for Clang without debug info, and the design goal of build time and memory close to a non-LTO build, are from that write-up and the Clang ThinLTO documentation. The two-to-three-fold compile-time cost of full LTO is a widely reported figure (for example, the Gentoo LTO documentation). &#8220;Sometimes beats full LTO&#8221; reflects the authors&#8217; own SPEC results, where ThinLTO&#8217;s more aggressive per-module pipeline occasionally wins.</p><p><strong><span>C. LLVM&#8217;s reach can be sized as a trillion-dollar surface area</span></strong></p><p>Apple FY2025 revenue 416.16 billion dollars (Form 8-K, quarter and year ended Sept 27, 2025); Alphabet 2025 revenue 402.84 billion dollars (Form 10-K / Annual Report 2025); NVIDIA fiscal 2026 revenue 215.9 billion dollars, data center 193.7 billion, for the year ended Jan 25, 2026 (Form 8-K). Combined, about 1.04 trillion dollars. This is top-line company revenue for firms whose flagship products are built with LLVM, presented as a floor, not value attributed to or captured by the compiler. Smartphone units of about 1.25 billion for 2025 are from IDC and Omdia. We do not treat the constructed total as audited.</p><p><strong><span>A. Rust (Cranelift) and Zig are building non-LLVM backends to escape LLVM compile times</span></strong></p><p>Rust&#8217;s <code>rustc_codegen_cranelift</code> reports roughly a twenty percent reduction in code-generation time and about five percent in clean-build time on large projects, and production-readiness is a 2025 Rust project goal (rust-project-goals, 2025h2). Zig made its self-hosted x86-64 backend the default for Debug builds on Linux and macOS in mid-2025, reporting five to fifty percent wall-clock improvements and citing LLVM as the compile-time bottleneck (Zig devlog, 2025). Both retain an LLVM path for optimized release builds.</p><div><hr></div><h2><strong>Sources</strong></h2><ol><li><p>V. Adve and C. Lattner, &#8220;The LLVM Compiler Infrastructure,&#8221; <em>Communications of the ACM</em>, June 2026. <a href="https://cacm.acm.org/federal-funding-of-academic-research/the-llvm-compiler-infrastructure/">cacm.acm.org</a></p></li><li><p>C. Lattner and V. Adve, &#8220;LLVM: A Compilation Framework for Lifelong Program Analysis and Transformation,&#8221; <em>CGO 2004</em>, pp. 75 to 86. DOI 10.1109/CGO.2004.1281665.</p></li><li><p>C. Lattner, &#8220;The LLVM 1.0 Release is finally available!&#8221;, llvm-announce, October 24, 2003.</p></li><li><p>LLVM Project, &#8220;Opaque Pointers.&#8221; <a href="https://llvm.org/docs/OpaquePointers.html">llvm.org/docs/OpaquePointers.html</a></p></li><li><p>Arm, &#8220;What is new in LLVM 22,&#8221; 2026; M. Larabel, &#8220;LLVM/Clang 22.1 Released,&#8221; <em>Phoronix</em>, Feb 24, 2026.</p></li><li><p>LLVM Project, release tags through 22.1.8. <a href="https://github.com/llvm/llvm-project/releases">github.com/llvm/llvm-project</a></p></li><li><p>Apple, &#8220;Xcode 14 Release Notes,&#8221; 2022. <a href="https://developer.apple.com/documentation/xcode-release-notes/xcode-14-release-notes">developer.apple.com</a></p></li><li><p>N. P. Lopes et al., &#8220;Alive2: Bounded Translation Validation for LLVM,&#8221; <em>PLDI 2021</em>. DOI 10.1145/3453483.3454030.</p></li><li><p>C. Lattner et al., &#8220;MLIR: Scaling Compiler Infrastructure for Domain Specific Computation,&#8221; <em>CGO 2021</em>. DOI 10.1109/CGO51591.2021.9370308.</p></li><li><p>C. Cummins et al., &#8220;Meta Large Language Model Compiler,&#8221; arXiv 2407.02524, 2024; <em>CC 2025</em>, DOI 10.1145/3708493.3712691.</p></li><li><p>A. Kothari et al., &#8220;Hydride,&#8221; <em>ASPLOS 2024</em>, DOI 10.1145/3620665.3640385; A. R. Noor et al., &#8220;MISAAL,&#8221; <em>PACMPL 2025</em>.</p></li><li><p>A. Ejjeh et al., &#8220;HPVM: Hardware-agnostic programming for heterogeneous parallel systems,&#8221; <em>IEEE Micro</em> 42(5), 2022.</p></li><li><p>V. Adve, faculty biography and CV, University of Illinois. <a href="https://vikram.cs.illinois.edu/bio/">vikram.cs.illinois.edu</a></p></li><li><p>T. Johnson, M. Amini, D. Li, &#8220;ThinLTO: Scalable and Incremental LTO,&#8221; LLVM Project Blog, 2016 (presented EuroLLVM 2015); Clang &#8220;ThinLTO&#8221; documentation. <a href="https://blog.llvm.org/2016/06/thinlto-scalable-and-incremental-lto.html">blog.llvm.org</a></p></li><li><p>Rust Project Goals 2025h2, &#8220;Production-ready Cranelift backend.&#8221; <a href="https://rust-lang.github.io/rust-project-goals/2025h2/production-ready-cranelift.html">rust-lang.github.io</a></p></li><li><p>Zig, Devlog 2025 (self-hosted x86-64 backend default for Debug builds). <a href="https://ziglang.org/devlog/2025/">ziglang.org/devlog/2025</a></p></li><li><p>Apple Inc., Form 8-K, fiscal 2025 fourth-quarter and full-year results (year ended Sept 27, 2025), revenue 416 billion dollars. <a href="https://www.sec.gov/cgi-bin/browse-edgar?action=getcompany&amp;CIK=0000320193">SEC EDGAR</a></p></li><li><p>Alphabet Inc., Annual Report / Form 10-K for 2025, total revenue 402.84 billion dollars. <a href="https://www.sec.gov/cgi-bin/browse-edgar?action=getcompany&amp;CIK=0001652044">SEC EDGAR</a></p></li><li><p>NVIDIA Corp., Form 8-K, fourth quarter and fiscal 2026 (year ended Jan 25, 2026), revenue 215.9 billion dollars, data center 193.7 billion. <a href="https://www.sec.gov/cgi-bin/browse-edgar?action=getcompany&amp;CIK=0001045810">SEC EDGAR</a></p></li><li><p>IDC and Omdia, worldwide smartphone shipments for 2025, about 1.25 billion units.</p></li></ol><p><strong>Inference.Engineering</strong> &#183;<em> Independent analysis of LLM serving infrastructure, GPU economics, and the compiler layer beneath them. This piece is built on the June 2026 CACM retrospective by Vikram Adve and Chris Lattner and cross-checked against LLVM 22.1 documentation, release notes and the primary research literature. Vendor-independent; every figure sourced. Corrections and disputes welcome.</em></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/p/how-llvm-works-the-ir-that-took-over/comments&quot;,&quot;text&quot;:&quot;Leave a comment&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/p/how-llvm-works-the-ir-that-took-over/comments"><span>Leave a comment</span></a></p><p></p>]]></content:encoded></item><item><title><![CDATA[The Wafer & the Wallet]]></title><description><![CDATA[Memory costs are inflecting up as enterprise token budgets slam shut. The wafer math, the pass-through proof, and the engineering playbook that defends the spread.]]></description><link>https://www.thesoftwarefrontier.com/p/the-wafer-and-the-wallet</link><guid isPermaLink="false">https://www.thesoftwarefrontier.com/p/the-wafer-and-the-wallet</guid><dc:creator><![CDATA[Lorenzo Bradanini]]></dc:creator><pubDate>Sun, 05 Jul 2026 07:56:35 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!0y1r!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc81fedd3-e26c-4b0d-9cc1-718e1c0f97e2_1122x1402.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!0y1r!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc81fedd3-e26c-4b0d-9cc1-718e1c0f97e2_1122x1402.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!0y1r!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc81fedd3-e26c-4b0d-9cc1-718e1c0f97e2_1122x1402.png 424w, https://substackcdn.com/image/fetch/$s_!0y1r!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc81fedd3-e26c-4b0d-9cc1-718e1c0f97e2_1122x1402.png 848w, https://substackcdn.com/image/fetch/$s_!0y1r!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc81fedd3-e26c-4b0d-9cc1-718e1c0f97e2_1122x1402.png 1272w, https://substackcdn.com/image/fetch/$s_!0y1r!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc81fedd3-e26c-4b0d-9cc1-718e1c0f97e2_1122x1402.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!0y1r!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc81fedd3-e26c-4b0d-9cc1-718e1c0f97e2_1122x1402.png" width="1122" height="1402" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c81fedd3-e26c-4b0d-9cc1-718e1c0f97e2_1122x1402.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1402,&quot;width&quot;:1122,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2521188,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/203533001?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc81fedd3-e26c-4b0d-9cc1-718e1c0f97e2_1122x1402.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!0y1r!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc81fedd3-e26c-4b0d-9cc1-718e1c0f97e2_1122x1402.png 424w, https://substackcdn.com/image/fetch/$s_!0y1r!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc81fedd3-e26c-4b0d-9cc1-718e1c0f97e2_1122x1402.png 848w, https://substackcdn.com/image/fetch/$s_!0y1r!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc81fedd3-e26c-4b0d-9cc1-718e1c0f97e2_1122x1402.png 1272w, https://substackcdn.com/image/fetch/$s_!0y1r!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc81fedd3-e26c-4b0d-9cc1-718e1c0f97e2_1122x1402.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>Intro </h2><p>For three years, the cost of making a token fell every quarter while the budget to buy tokens grew without a meter. In <strong>H1 2026 both curves reversed</strong> at once: memory pass-through is raising the cost floor of every token served, while enterprises institutionalize token budgets at the ceiling. </p><p>This is the anatomy of <strong>inference&#8217;s first margin squeeze</strong>: the wafer math, the proof of pass-through, modeled cost floors marked to real market prices, and the engineering playbook that defends the spread.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Software Frontier is a reader-supported publication. Please, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!PC3u!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96bbb590-bfcc-40f5-99ac-caf2f1d5b3bb_1100x353.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!PC3u!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96bbb590-bfcc-40f5-99ac-caf2f1d5b3bb_1100x353.png 424w, https://substackcdn.com/image/fetch/$s_!PC3u!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96bbb590-bfcc-40f5-99ac-caf2f1d5b3bb_1100x353.png 848w, https://substackcdn.com/image/fetch/$s_!PC3u!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96bbb590-bfcc-40f5-99ac-caf2f1d5b3bb_1100x353.png 1272w, https://substackcdn.com/image/fetch/$s_!PC3u!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96bbb590-bfcc-40f5-99ac-caf2f1d5b3bb_1100x353.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!PC3u!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96bbb590-bfcc-40f5-99ac-caf2f1d5b3bb_1100x353.png" width="1100" height="353" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/96bbb590-bfcc-40f5-99ac-caf2f1d5b3bb_1100x353.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:353,&quot;width&quot;:1100,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:133192,&quot;alt&quot;:&quot;c00-masthead&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="c00-masthead" title="c00-masthead" srcset="https://substackcdn.com/image/fetch/$s_!PC3u!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96bbb590-bfcc-40f5-99ac-caf2f1d5b3bb_1100x353.png 424w, https://substackcdn.com/image/fetch/$s_!PC3u!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96bbb590-bfcc-40f5-99ac-caf2f1d5b3bb_1100x353.png 848w, https://substackcdn.com/image/fetch/$s_!PC3u!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96bbb590-bfcc-40f5-99ac-caf2f1d5b3bb_1100x353.png 1272w, https://substackcdn.com/image/fetch/$s_!PC3u!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96bbb590-bfcc-40f5-99ac-caf2f1d5b3bb_1100x353.png 1456w" sizes="100vw"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>Two curves reversed in the same half-year</h2><p>Every industry eventually meets the quarter when its two defining curves cross. For LLM inference, that quarter just happened. </p><p>Since late 2022, the economics of this business rested on two assumptions so reliable nobody wrote them down: <strong>the cost of producing a token falls every quarter</strong> (process nodes, kernel engineering, quantization, and a relentless price war), and <strong>the budget for buying tokens grows without a meter</strong>, the era of &#8220;tokenmaxxing,&#8221; when Meta and Salesforce were publicly urging employees to consume as many tokens as possible. In the first half of 2026, both assumptions died within months of each other.</p><p><strong>Jaw one: the cost floor turned upward.</strong> That is the core of this issue: the deepest memory supercycle in the forty-year history of DRAM, transmitted measurably into GPU rental rates with a 4-6 week lag ( Fig. 7). For the first time since the ChatGPT moment, <strong>the marginal cost of serving a token is inflecting up, not down</strong>: driven by the single most supply-constrained commodity in technology.</p><p><strong>Jaw two: the price ceiling got institutionalized.</strong> SemiAnalysis published field data on July 2 from conversations with 50+ enterprises: Uber burned through its annual Claude Code and Codex budget in four months and responded with a<strong> $1,500/month per-employee cap</strong>; an aerospace manufacturer&#8217;s $250/month caps were exhausted by power users in four days; companies are switching off premium model tiers and downgrading defaults; enterprise budgets now range from $250 to tens of thousands per employee per month, but they are <em>budgets</em>, reviewed by finance. </p><p> Read the finding precisely, because it is not a demand-collapse story: SemiAnalysis concludes there is no material risk to 2H26 AI budgets and API spend keeps compounding. <strong>The jaw is not falling demand, it is the arrival of price sensitivity.</strong> </p><p>Unbounded willingness-to-pay is over; every serving provider now negotiates against a CFO&#8217;s cap instead of an enthusiast&#8217;s curiosity.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!zbOW!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F603f7041-8239-4c16-a157-e9ccd588f95f_1100x690.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!zbOW!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F603f7041-8239-4c16-a157-e9ccd588f95f_1100x690.png 424w, https://substackcdn.com/image/fetch/$s_!zbOW!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F603f7041-8239-4c16-a157-e9ccd588f95f_1100x690.png 848w, https://substackcdn.com/image/fetch/$s_!zbOW!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F603f7041-8239-4c16-a157-e9ccd588f95f_1100x690.png 1272w, https://substackcdn.com/image/fetch/$s_!zbOW!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F603f7041-8239-4c16-a157-e9ccd588f95f_1100x690.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!zbOW!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F603f7041-8239-4c16-a157-e9ccd588f95f_1100x690.png" width="1100" height="690" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/603f7041-8239-4c16-a157-e9ccd588f95f_1100x690.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:690,&quot;width&quot;:1100,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;c-fig0&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="c-fig0" title="c-fig0" srcset="https://substackcdn.com/image/fetch/$s_!zbOW!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F603f7041-8239-4c16-a157-e9ccd588f95f_1100x690.png 424w, https://substackcdn.com/image/fetch/$s_!zbOW!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F603f7041-8239-4c16-a157-e9ccd588f95f_1100x690.png 848w, https://substackcdn.com/image/fetch/$s_!zbOW!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F603f7041-8239-4c16-a157-e9ccd588f95f_1100x690.png 1272w, https://substackcdn.com/image/fetch/$s_!zbOW!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F603f7041-8239-4c16-a157-e9ccd588f95f_1100x690.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!9PWX!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72d19ef9-e996-462d-bab6-03fab8436ba5_1100x473.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!9PWX!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72d19ef9-e996-462d-bab6-03fab8436ba5_1100x473.png 424w, https://substackcdn.com/image/fetch/$s_!9PWX!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72d19ef9-e996-462d-bab6-03fab8436ba5_1100x473.png 848w, https://substackcdn.com/image/fetch/$s_!9PWX!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72d19ef9-e996-462d-bab6-03fab8436ba5_1100x473.png 1272w, https://substackcdn.com/image/fetch/$s_!9PWX!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72d19ef9-e996-462d-bab6-03fab8436ba5_1100x473.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!9PWX!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72d19ef9-e996-462d-bab6-03fab8436ba5_1100x473.png" width="1100" height="473" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/72d19ef9-e996-462d-bab6-03fab8436ba5_1100x473.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:473,&quot;width&quot;:1100,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;c-callout0&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="c-callout0" title="c-callout0" srcset="https://substackcdn.com/image/fetch/$s_!9PWX!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72d19ef9-e996-462d-bab6-03fab8436ba5_1100x473.png 424w, https://substackcdn.com/image/fetch/$s_!9PWX!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72d19ef9-e996-462d-bab6-03fab8436ba5_1100x473.png 848w, https://substackcdn.com/image/fetch/$s_!9PWX!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72d19ef9-e996-462d-bab6-03fab8436ba5_1100x473.png 1272w, https://substackcdn.com/image/fetch/$s_!9PWX!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72d19ef9-e996-462d-bab6-03fab8436ba5_1100x473.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div><hr></div><h2>The quarter the memory market broke</h2><p>Start with the raw numbers, because they have no precedent in the forty-year history of the DRAM industry. TrendForce&#8217;s Q1 2026 contract-price revisions came in at <strong>+90-95% quarter-over-quarter for conventional DRAM</strong>: revised <em>upward</em> from an initial forecast of +55-60% that analysts had already called unprecedented. </p><p>PC DRAM was expected to more than double in a single quarter; server DDR5 rose roughly 90%; <strong>NAND contract prices </strong>climbed 33-38% in the same window. Counterpoint tracked spot DRAM up 80-90% inside the quarter. </p><p>For the full year, TrendForce projects DRAM up more than 70% <em>on top of</em> the Q1 step-change; Bank of America has DRAM industry revenue +51% YoY and NAND +45%.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!OkVP!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4fb89a05-21d2-440a-8e5c-7ad016d63e21_1100x571.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!OkVP!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4fb89a05-21d2-440a-8e5c-7ad016d63e21_1100x571.png 424w, https://substackcdn.com/image/fetch/$s_!OkVP!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4fb89a05-21d2-440a-8e5c-7ad016d63e21_1100x571.png 848w, https://substackcdn.com/image/fetch/$s_!OkVP!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4fb89a05-21d2-440a-8e5c-7ad016d63e21_1100x571.png 1272w, https://substackcdn.com/image/fetch/$s_!OkVP!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4fb89a05-21d2-440a-8e5c-7ad016d63e21_1100x571.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!OkVP!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4fb89a05-21d2-440a-8e5c-7ad016d63e21_1100x571.png" width="1100" height="571" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/4fb89a05-21d2-440a-8e5c-7ad016d63e21_1100x571.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:571,&quot;width&quot;:1100,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;c-fig1&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="c-fig1" title="c-fig1" srcset="https://substackcdn.com/image/fetch/$s_!OkVP!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4fb89a05-21d2-440a-8e5c-7ad016d63e21_1100x571.png 424w, https://substackcdn.com/image/fetch/$s_!OkVP!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4fb89a05-21d2-440a-8e5c-7ad016d63e21_1100x571.png 848w, https://substackcdn.com/image/fetch/$s_!OkVP!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4fb89a05-21d2-440a-8e5c-7ad016d63e21_1100x571.png 1272w, https://substackcdn.com/image/fetch/$s_!OkVP!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4fb89a05-21d2-440a-8e5c-7ad016d63e21_1100x571.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The supply side tells you why this is not a blip. <strong>SK Hynix announced in October 2025 that its entire 2026 production capacity was already sold out.</strong> Kioxia has said the same of its 2026 NAND output. </p><p>Micron shut down its Crucial consumer business outright (a two-decade brand, liquidated) to route every wafer toward <strong>hyperscaler and GPU-grade contracts</strong>, and its executives describe being able to cover at most two-thirds of medium-term demand for some customers.</p><p>Micron&#8217;s Boise fabs ramp in 2027-2028; the New York megafab in 2030. <strong>Jensen Huang</strong>, asked at CES 2026 whether gamers should resent AI for GPU prices, answered in effect that the world needs more memory factories. </p><blockquote><p><em>As stated even before, in the previous issues: when the largest memory buyer on Earth says the fix is on the supply side, the message downstream is: you will not negotiate your way to allocation relief.</em></p></blockquote><h3>Three structural breaks from every previous cycle</h3><p><strong>1. The demand driver is durable, not episodic.</strong> Prior shortages came from fab fires, earthquakes, supply discipline, or one-off demand pulses (crypto, pandemic PCs). This one comes from inference workloads compounding in production. IDC&#8217;s assessment calls it a <strong>&#8220;permanent reallocation&#8221; of the world&#8217;s wafer capacity, not a cyclical shortage</strong>, and IDC does not deploy the word &#8220;permanent&#8221; casually.</p><p><strong>2. The margin gradient is a one-way valve, and the wafer math is brutal.</strong> HBM carries gross margins 3-5&#215; commodity DRAM, so every rational fab reallocates cleanroom toward it. </p><p>But because of TSV formation, die thinning, stacking, and compounded assembly yield, <strong>one gigabyte of HBM consumes &#8776;4&#215; the wafer capacity of one gigabyte of standard DRAM; GDDR7 consumes &#8776;1.7&#215;</strong>. AI&#8217;s <em>effective</em> wafer claim reaches ~20% of world DRAM capacity in 2026 against total bit-supply growth of only 10-16% per year. </p><p>The arithmetic cannot close without starving someone, and &#8220;<em>someone</em>&#8221; is every non-AI buyer plus every AI buyer without a long-term agreement.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!6WWW!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3a88d9d-78ef-4e79-8bfa-b0f22bbd8935_1100x562.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!6WWW!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3a88d9d-78ef-4e79-8bfa-b0f22bbd8935_1100x562.png 424w, https://substackcdn.com/image/fetch/$s_!6WWW!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3a88d9d-78ef-4e79-8bfa-b0f22bbd8935_1100x562.png 848w, https://substackcdn.com/image/fetch/$s_!6WWW!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3a88d9d-78ef-4e79-8bfa-b0f22bbd8935_1100x562.png 1272w, https://substackcdn.com/image/fetch/$s_!6WWW!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3a88d9d-78ef-4e79-8bfa-b0f22bbd8935_1100x562.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!6WWW!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3a88d9d-78ef-4e79-8bfa-b0f22bbd8935_1100x562.png" width="1100" height="562" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a3a88d9d-78ef-4e79-8bfa-b0f22bbd8935_1100x562.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:562,&quot;width&quot;:1100,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;c-fig2&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="c-fig2" title="c-fig2" srcset="https://substackcdn.com/image/fetch/$s_!6WWW!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3a88d9d-78ef-4e79-8bfa-b0f22bbd8935_1100x562.png 424w, https://substackcdn.com/image/fetch/$s_!6WWW!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3a88d9d-78ef-4e79-8bfa-b0f22bbd8935_1100x562.png 848w, https://substackcdn.com/image/fetch/$s_!6WWW!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3a88d9d-78ef-4e79-8bfa-b0f22bbd8935_1100x562.png 1272w, https://substackcdn.com/image/fetch/$s_!6WWW!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3a88d9d-78ef-4e79-8bfa-b0f22bbd8935_1100x562.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><strong>3. The price </strong><em><strong>convergence</strong></em><strong> matters more than the price level.</strong> HBM3e historically sold at 4-5&#215; server DDR5 per bit. TrendForce expects that gap to compress to <strong>1-2&#215; by end-2026</strong>, not because HBM got cheaper, but because DDR5 inflated toward it. Sit with that. </p><p><strong>The entire tiered-memory architecture of modern serving (</strong><em>KV offload to host DRAM, paged hierarchies, DDR5-backed prefix caches, NVMe cold tiers</em><strong>) was designed in a world where the tier below HBM was 4-5&#215; cheaper per bit.</strong> </p><p>When the discount for descending a tier collapses from ~80% to ~30-50%, a large fraction of the &#8220;obvious&#8221; tiering optimizations of 2024-2025 stop being obvious. Section 5 does that math properly.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!J560!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feeefc492-5628-48c1-8f25-09ae71c75c19_1100x560.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!J560!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feeefc492-5628-48c1-8f25-09ae71c75c19_1100x560.png 424w, https://substackcdn.com/image/fetch/$s_!J560!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feeefc492-5628-48c1-8f25-09ae71c75c19_1100x560.png 848w, https://substackcdn.com/image/fetch/$s_!J560!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feeefc492-5628-48c1-8f25-09ae71c75c19_1100x560.png 1272w, https://substackcdn.com/image/fetch/$s_!J560!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feeefc492-5628-48c1-8f25-09ae71c75c19_1100x560.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!J560!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feeefc492-5628-48c1-8f25-09ae71c75c19_1100x560.png" width="1100" height="560" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/eeefc492-5628-48c1-8f25-09ae71c75c19_1100x560.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:560,&quot;width&quot;:1100,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;c-fig3&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="c-fig3" title="c-fig3" srcset="https://substackcdn.com/image/fetch/$s_!J560!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feeefc492-5628-48c1-8f25-09ae71c75c19_1100x560.png 424w, https://substackcdn.com/image/fetch/$s_!J560!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feeefc492-5628-48c1-8f25-09ae71c75c19_1100x560.png 848w, https://substackcdn.com/image/fetch/$s_!J560!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feeefc492-5628-48c1-8f25-09ae71c75c19_1100x560.png 1272w, https://substackcdn.com/image/fetch/$s_!J560!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feeefc492-5628-48c1-8f25-09ae71c75c19_1100x560.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>One supply-side note that matters for H2: <strong>HBM4 becomes the mainstream HBM in the second half of 2026</strong>, with a doubled 2,048-bit interface, ~2 TB/s per stack, and a manufacturing-complexity premium above 30% over HBM3e. NVIDIA&#8217;s Rubin ships with up to 288 GB of HBM4 per GPU, and every one of those stacks is wafer capacity that used to be laptops.</p><div><hr></div><h2>The KV cache ate the memory supply</h2><p>The lazy narrative is &#8220;AI ate the memory market.&#8221; The precise narrative is that <strong>the KV cache ate the memory market</strong>, and the industry&#8217;s shift from training-dominated to inference-dominated compute is what lit the fuse.</p><p>Training demand is large but <em>concentrated and schedulable</em>: a fixed pool of HBM for weights, activations, and optimizer states, planned quarters ahead. Inference demand is <em>elastic and per-user</em>: every concurrent conversation, agent trajectory, and <strong>200K-token codebase </strong>instantiates its own slab of state that must live somewhere in the hierarchy for the life of the request, and increasingly beyond it, as prefix caches persist context between turns. </p><p>Industry estimates via TrendForce put cloud high-speed memory consumption near <strong>3 exabytes in 2026</strong>, with core inference platforms (<em>Gemini-, Bedrock-, ChatGPT-class serving</em>) accounting for ~750 PB of <em>live</em> memory demand before redundancy roughly doubles it. </p><p>TrendForce separately projects that <strong>by 2029 inference (not training) becomes the primary driver of AI server demand outright</strong>, and North American CSPs are already bulk-buying high-capacity DDR5 RDIMMs specifically for inference fleets, the demand class that used to be the safety valve absorbing DRAM oversupply.</p><h3>The per-token physics, derived from first principles</h3><p>Everything in this issue hangs on one number: bytes of KV state per token. Take the workhorse case, a Llama-3-class 70B dense model with grouped-query attention: 80 layers, 8 KV heads, head dimension 128.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!zKAn!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faf0f9d6d-1829-4a06-b45c-6306b2c3dd3b_1100x449.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!zKAn!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faf0f9d6d-1829-4a06-b45c-6306b2c3dd3b_1100x449.png 424w, https://substackcdn.com/image/fetch/$s_!zKAn!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faf0f9d6d-1829-4a06-b45c-6306b2c3dd3b_1100x449.png 848w, https://substackcdn.com/image/fetch/$s_!zKAn!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faf0f9d6d-1829-4a06-b45c-6306b2c3dd3b_1100x449.png 1272w, https://substackcdn.com/image/fetch/$s_!zKAn!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faf0f9d6d-1829-4a06-b45c-6306b2c3dd3b_1100x449.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!zKAn!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faf0f9d6d-1829-4a06-b45c-6306b2c3dd3b_1100x449.png" width="1100" height="449" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/af0f9d6d-1829-4a06-b45c-6306b2c3dd3b_1100x449.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:449,&quot;width&quot;:1100,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;c-math0&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="c-math0" title="c-math0" srcset="https://substackcdn.com/image/fetch/$s_!zKAn!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faf0f9d6d-1829-4a06-b45c-6306b2c3dd3b_1100x449.png 424w, https://substackcdn.com/image/fetch/$s_!zKAn!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faf0f9d6d-1829-4a06-b45c-6306b2c3dd3b_1100x449.png 848w, https://substackcdn.com/image/fetch/$s_!zKAn!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faf0f9d6d-1829-4a06-b45c-6306b2c3dd3b_1100x449.png 1272w, https://substackcdn.com/image/fetch/$s_!zKAn!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faf0f9d6d-1829-4a06-b45c-6306b2c3dd3b_1100x449.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!QCjx!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fca93064a-232c-4fe8-b04b-919a5b7a5ea3_1100x625.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!QCjx!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fca93064a-232c-4fe8-b04b-919a5b7a5ea3_1100x625.png 424w, https://substackcdn.com/image/fetch/$s_!QCjx!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fca93064a-232c-4fe8-b04b-919a5b7a5ea3_1100x625.png 848w, https://substackcdn.com/image/fetch/$s_!QCjx!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fca93064a-232c-4fe8-b04b-919a5b7a5ea3_1100x625.png 1272w, https://substackcdn.com/image/fetch/$s_!QCjx!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fca93064a-232c-4fe8-b04b-919a5b7a5ea3_1100x625.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!QCjx!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fca93064a-232c-4fe8-b04b-919a5b7a5ea3_1100x625.png" width="1100" height="625" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ca93064a-232c-4fe8-b04b-919a5b7a5ea3_1100x625.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:625,&quot;width&quot;:1100,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;c-fig4&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="c-fig4" title="c-fig4" srcset="https://substackcdn.com/image/fetch/$s_!QCjx!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fca93064a-232c-4fe8-b04b-919a5b7a5ea3_1100x625.png 424w, https://substackcdn.com/image/fetch/$s_!QCjx!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fca93064a-232c-4fe8-b04b-919a5b7a5ea3_1100x625.png 848w, https://substackcdn.com/image/fetch/$s_!QCjx!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fca93064a-232c-4fe8-b04b-919a5b7a5ea3_1100x625.png 1272w, https://substackcdn.com/image/fetch/$s_!QCjx!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fca93064a-232c-4fe8-b04b-919a5b7a5ea3_1100x625.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>A single 128K sequence on the GQA-8 model pins <strong>~41 GB at FP16: half an H100&#8217;s entire HBM for one user&#8217;s context</strong>, before a single weight is stored. This is why effective batch size at long context is a memory-capacity calculation, not a compute calculation, and why <strong>every marginal concurrent user at long context is, economically, a memory purchase</strong>. </p><p>Multiply by the fleet: 10,000 concurrent long-context sequences at FP8 KV pins <strong>~200 TB of state across HBM</strong>, DDR5, and NVMe, state that produces no tokens itself; it exists purely so decode can stream it past the compute units at terabytes per second. </p><p>The <strong>2026 memory market</strong> is what happens when the industry&#8217;s aggregate KV state, growing super-linearly with agent adoption and context inflation, collides with wafer supply growing 10-16% a year.</p><h3>Two second-order effects complete the picture</h3><p><strong>NAND got dragged in.</strong> NVMe became the cold tier for prefix caches and the staging layer for weights (including GPU-direct storage paths). Kioxia now says nearly half its future NAND demand could come from AI. Meanwhile Samsung and SK Hynix <em>cut</em> NAND wafer output in 2024-2025 while chasing HBM margins. Omdia has Samsung&#8217;s NAND wafers falling 4.9M&#8594;4.68M and SK Hynix&#8217;s 1.9M&#8594;1.7M. <strong>There is no cheap tier left to hide in.</strong></p><p><strong>The host-DRAM tax on every GPU node inflated.</strong> A serious 2026 inference node pairs its GPUs with 1-2 TB of DDR5 for KV offload, CPU-side batching, and cache tiers. At Q1 2026 server-DRAM pricing, the host memory on a single node can cost what a mid-range GPU did two years ago. No 2024-vintage TCO model has this line item at anywhere near its current size.</p><div><hr></div><h2>Roofline: why prefill and decode were never the same workload</h2><p>Before the silicon story makes sense, you need the roofline argument stated with numbers, not vibes. A kernel&#8217;s attainable throughput is bounded by min(peak_FLOPs, intensity &#215; bandwidth), where <strong>arithmetic intensity</strong> is FLOPs performed per byte moved from memory. </p><p>The machine has a <em>ridge point</em>, the intensity at which it stops being bandwidth-bound and becomes compute-bound.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!4bPr!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F93058626-ab79-40c9-b56a-cf5c87b9acb2_1100x483.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!4bPr!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F93058626-ab79-40c9-b56a-cf5c87b9acb2_1100x483.png 424w, https://substackcdn.com/image/fetch/$s_!4bPr!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F93058626-ab79-40c9-b56a-cf5c87b9acb2_1100x483.png 848w, https://substackcdn.com/image/fetch/$s_!4bPr!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F93058626-ab79-40c9-b56a-cf5c87b9acb2_1100x483.png 1272w, https://substackcdn.com/image/fetch/$s_!4bPr!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F93058626-ab79-40c9-b56a-cf5c87b9acb2_1100x483.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!4bPr!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F93058626-ab79-40c9-b56a-cf5c87b9acb2_1100x483.png" width="1100" height="483" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/93058626-ab79-40c9-b56a-cf5c87b9acb2_1100x483.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:483,&quot;width&quot;:1100,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;c-math1&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="c-math1" title="c-math1" srcset="https://substackcdn.com/image/fetch/$s_!4bPr!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F93058626-ab79-40c9-b56a-cf5c87b9acb2_1100x483.png 424w, https://substackcdn.com/image/fetch/$s_!4bPr!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F93058626-ab79-40c9-b56a-cf5c87b9acb2_1100x483.png 848w, https://substackcdn.com/image/fetch/$s_!4bPr!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F93058626-ab79-40c9-b56a-cf5c87b9acb2_1100x483.png 1272w, https://substackcdn.com/image/fetch/$s_!4bPr!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F93058626-ab79-40c9-b56a-cf5c87b9acb2_1100x483.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!6tit!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa859114a-4861-45ca-8684-6e2ac283494e_1100x647.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!6tit!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa859114a-4861-45ca-8684-6e2ac283494e_1100x647.png 424w, https://substackcdn.com/image/fetch/$s_!6tit!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa859114a-4861-45ca-8684-6e2ac283494e_1100x647.png 848w, https://substackcdn.com/image/fetch/$s_!6tit!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa859114a-4861-45ca-8684-6e2ac283494e_1100x647.png 1272w, https://substackcdn.com/image/fetch/$s_!6tit!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa859114a-4861-45ca-8684-6e2ac283494e_1100x647.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!6tit!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa859114a-4861-45ca-8684-6e2ac283494e_1100x647.png" width="1100" height="647" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a859114a-4861-45ca-8684-6e2ac283494e_1100x647.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:647,&quot;width&quot;:1100,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;c-fig5&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="c-fig5" title="c-fig5" srcset="https://substackcdn.com/image/fetch/$s_!6tit!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa859114a-4861-45ca-8684-6e2ac283494e_1100x647.png 424w, https://substackcdn.com/image/fetch/$s_!6tit!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa859114a-4861-45ca-8684-6e2ac283494e_1100x647.png 848w, https://substackcdn.com/image/fetch/$s_!6tit!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa859114a-4861-45ca-8684-6e2ac283494e_1100x647.png 1272w, https://substackcdn.com/image/fetch/$s_!6tit!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa859114a-4861-45ca-8684-6e2ac283494e_1100x647.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The corollary that most cost models miss: <strong>decode throughput is not a FLOPs number, it is a bandwidth budget divided by bytes-per-step</strong>. Per decode step the device must stream the weights once plus every live sequence&#8217;s KV:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!sjTw!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faedb9217-181e-49c1-8ffb-258e768d104d_1100x381.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!sjTw!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faedb9217-181e-49c1-8ffb-258e768d104d_1100x381.png 424w, https://substackcdn.com/image/fetch/$s_!sjTw!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faedb9217-181e-49c1-8ffb-258e768d104d_1100x381.png 848w, https://substackcdn.com/image/fetch/$s_!sjTw!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faedb9217-181e-49c1-8ffb-258e768d104d_1100x381.png 1272w, https://substackcdn.com/image/fetch/$s_!sjTw!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faedb9217-181e-49c1-8ffb-258e768d104d_1100x381.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!sjTw!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faedb9217-181e-49c1-8ffb-258e768d104d_1100x381.png" width="1100" height="381" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/aedb9217-181e-49c1-8ffb-258e768d104d_1100x381.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:381,&quot;width&quot;:1100,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;c-math2&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="c-math2" title="c-math2" srcset="https://substackcdn.com/image/fetch/$s_!sjTw!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faedb9217-181e-49c1-8ffb-258e768d104d_1100x381.png 424w, https://substackcdn.com/image/fetch/$s_!sjTw!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faedb9217-181e-49c1-8ffb-258e768d104d_1100x381.png 848w, https://substackcdn.com/image/fetch/$s_!sjTw!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faedb9217-181e-49c1-8ffb-258e768d104d_1100x381.png 1272w, https://substackcdn.com/image/fetch/$s_!sjTw!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faedb9217-181e-49c1-8ffb-258e768d104d_1100x381.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Read those three lines again. Same GPU, same model, same bandwidth: <strong>16&#215; throughput difference between short-context and long-context traffic</strong>, and <strong>a clean 2&#215; recovered at 128K purely by halving KV bytes</strong>. Context length is not a feature flag; it is a cost multiplier that acts through memory. </p><p>Every<strong> KV byte you remove</strong> converts directly into either more concurrent sequences or faster steps: both of which are revenue.</p><div><hr></div><h2>Rubin CPX is a memory trade wearing a GPU costume</h2><p>Now read <strong>NVIDIA&#8217;s 2026 inference roadmap </strong>through the memory market and it snaps into focus in a way most launch coverage missed. If prefill is compute-bound and decode is bandwidth-bound (&#167;03), then serving both phases on <strong>identical HBM-stuffed GPUs</strong> means that during prefill you are paying for the most supply-constrained, margin-rich commodity on the planet (HBM bandwidth) <em>and not using it</em>. </p><p>In 2024 that was an inefficiency. In 2026, with HBM at 4&#215; wafer-equivalent cost and fully allocated into next year, <strong>it is an unforced error measured in real money</strong>.</p><p><strong>Rubin CPX is the correction.</strong> Announced at the AI Infra Summit in September 2025, detailed through CES and GTC 2026, generally available late this year: a monolithic prefill-specialized die pairing <strong>30 PFLOPS of sparse NVFP4 compute</strong><em> (20 PFLOPS dense, per SemiAnalysis)</em><strong> with 128 GB of GDDR7 at only ~2 TB/s of bandwidth</strong> (32 Gbps GDDR7 on a 512-bit bus. </p><p>That bandwidth would embarrass a decode GPU) it&#8217;s below an H100. On a prefill part it is the entire point: SemiAnalysis&#8217;s launch framing was exactly right, a chip deliberately skinny on bandwidth and fat on compute, because that is what prefill consumes. Its intensity budget sits far to the right of <strong>Fig. 5&#8217;s ridge, where GDDR7 </strong>is not a compromise but a correct sizing.</p><p>Overlay the wafer math and the arbitrage is explicit: GDDR7 costs 1.7&#215; standard DRAM wafer-equivalent; HBM costs 4&#215;, plus CoWoS packaging, interposer, and thermal budget. <strong>By building the prefill tier out of GDDR7, NVIDIA routes the fastest-growing slice of inference demand</strong><em> (long-context prompt processing) </em><strong>around the most constrained commodity in its own supply chain.</strong> </p><p>Every prefill FLOP served from GDDR7 is HBM allocation freed for decode, where it earns its keep.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!n-TM!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf8f4658-7538-401d-8414-7c00ee5969d1_1100x460.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!n-TM!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf8f4658-7538-401d-8414-7c00ee5969d1_1100x460.png 424w, https://substackcdn.com/image/fetch/$s_!n-TM!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf8f4658-7538-401d-8414-7c00ee5969d1_1100x460.png 848w, https://substackcdn.com/image/fetch/$s_!n-TM!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf8f4658-7538-401d-8414-7c00ee5969d1_1100x460.png 1272w, https://substackcdn.com/image/fetch/$s_!n-TM!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf8f4658-7538-401d-8414-7c00ee5969d1_1100x460.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!n-TM!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf8f4658-7538-401d-8414-7c00ee5969d1_1100x460.png" width="1100" height="460" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/cf8f4658-7538-401d-8414-7c00ee5969d1_1100x460.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:460,&quot;width&quot;:1100,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;c-table0&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="c-table0" title="c-table0" srcset="https://substackcdn.com/image/fetch/$s_!n-TM!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf8f4658-7538-401d-8414-7c00ee5969d1_1100x460.png 424w, https://substackcdn.com/image/fetch/$s_!n-TM!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf8f4658-7538-401d-8414-7c00ee5969d1_1100x460.png 848w, https://substackcdn.com/image/fetch/$s_!n-TM!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf8f4658-7538-401d-8414-7c00ee5969d1_1100x460.png 1272w, https://substackcdn.com/image/fetch/$s_!n-TM!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf8f4658-7538-401d-8414-7c00ee5969d1_1100x460.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>NVIDIA is pitching this with a number that deserves both scrutiny and attention : <strong>$5B of token revenue per $100M of infrastructure</strong>, 30-50&#215; platform ROI, and naming Cursor, Runway, and Magic as early partners, which tells you the target token distribution: repository-scale coding context and long-form generative media, i.e., prefill-dominated workloads. Discount the multiplier as marketing; the direction is load-bearing.</p><p>The strategic read for this audience: <strong>hardware disaggregation of prefill and decode is memory-market arbitrage instantiated in silicon, and it will not stay proprietary to NVIDIA</strong>. SRAM-based decode tiers are being positioned into the same disaggregated fabrics, AMD will be forced to answer, and every serious serving stack (<em>Dynamo, vLLM&#8217;s disagg mode, SGLang, Mooncake-style architectures</em>) has converged on <strong>KV-cache-transfer-over-fabric as the central abstraction of 2026 serving</strong>. </p><p>The clearest tell that this is permanent: per Tom&#8217;s Hardware&#8217;s platform teardown, the Vera Rubin rack&#8217;s BlueField-4 DPU ships with an <em>integrated SSD specifically to store KV cache</em>: cached context now has its own dedicated hardware tier in the reference design. </p><p>If your mental model of an &#8220;<em>inference GPU</em>&#8221; is one SKU doing both phases, you are one hardware generation behind the economics, and Q1&#8217;s memory prices just made that lag expensive.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!NaQE!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F95a817d3-72a1-4c0a-910b-a5ebef3df62f_1100x555.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!NaQE!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F95a817d3-72a1-4c0a-910b-a5ebef3df62f_1100x555.png 424w, https://substackcdn.com/image/fetch/$s_!NaQE!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F95a817d3-72a1-4c0a-910b-a5ebef3df62f_1100x555.png 848w, https://substackcdn.com/image/fetch/$s_!NaQE!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F95a817d3-72a1-4c0a-910b-a5ebef3df62f_1100x555.png 1272w, https://substackcdn.com/image/fetch/$s_!NaQE!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F95a817d3-72a1-4c0a-910b-a5ebef3df62f_1100x555.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!NaQE!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F95a817d3-72a1-4c0a-910b-a5ebef3df62f_1100x555.png" width="1100" height="555" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/95a817d3-72a1-4c0a-910b-a5ebef3df62f_1100x555.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:555,&quot;width&quot;:1100,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;c-callout1&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="c-callout1" title="c-callout1" srcset="https://substackcdn.com/image/fetch/$s_!NaQE!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F95a817d3-72a1-4c0a-910b-a5ebef3df62f_1100x555.png 424w, https://substackcdn.com/image/fetch/$s_!NaQE!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F95a817d3-72a1-4c0a-910b-a5ebef3df62f_1100x555.png 848w, https://substackcdn.com/image/fetch/$s_!NaQE!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F95a817d3-72a1-4c0a-910b-a5ebef3df62f_1100x555.png 1272w, https://substackcdn.com/image/fetch/$s_!NaQE!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F95a817d3-72a1-4c0a-910b-a5ebef3df62f_1100x555.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!quTc!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc975bc73-b912-4dc6-a8c5-9b0f5d200935_1100x433.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!quTc!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc975bc73-b912-4dc6-a8c5-9b0f5d200935_1100x433.png 424w, https://substackcdn.com/image/fetch/$s_!quTc!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc975bc73-b912-4dc6-a8c5-9b0f5d200935_1100x433.png 848w, https://substackcdn.com/image/fetch/$s_!quTc!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc975bc73-b912-4dc6-a8c5-9b0f5d200935_1100x433.png 1272w, https://substackcdn.com/image/fetch/$s_!quTc!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc975bc73-b912-4dc6-a8c5-9b0f5d200935_1100x433.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!quTc!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc975bc73-b912-4dc6-a8c5-9b0f5d200935_1100x433.png" width="1100" height="433" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c975bc73-b912-4dc6-a8c5-9b0f5d200935_1100x433.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:433,&quot;width&quot;:1100,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;c-callout2&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="c-callout2" title="c-callout2" srcset="https://substackcdn.com/image/fetch/$s_!quTc!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc975bc73-b912-4dc6-a8c5-9b0f5d200935_1100x433.png 424w, https://substackcdn.com/image/fetch/$s_!quTc!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc975bc73-b912-4dc6-a8c5-9b0f5d200935_1100x433.png 848w, https://substackcdn.com/image/fetch/$s_!quTc!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc975bc73-b912-4dc6-a8c5-9b0f5d200935_1100x433.png 1272w, https://substackcdn.com/image/fetch/$s_!quTc!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc975bc73-b912-4dc6-a8c5-9b0f5d200935_1100x433.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div><hr></div><h2>The repriced hierarchy, and what happens to cost per token</h2><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!4czk!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74d49549-f1d1-428f-b202-669ebefeea81_1100x672.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!4czk!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74d49549-f1d1-428f-b202-669ebefeea81_1100x672.png 424w, https://substackcdn.com/image/fetch/$s_!4czk!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74d49549-f1d1-428f-b202-669ebefeea81_1100x672.png 848w, https://substackcdn.com/image/fetch/$s_!4czk!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74d49549-f1d1-428f-b202-669ebefeea81_1100x672.png 1272w, https://substackcdn.com/image/fetch/$s_!4czk!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74d49549-f1d1-428f-b202-669ebefeea81_1100x672.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!4czk!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74d49549-f1d1-428f-b202-669ebefeea81_1100x672.png" width="1100" height="672" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/74d49549-f1d1-428f-b202-669ebefeea81_1100x672.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:672,&quot;width&quot;:1100,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;c-table1&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="c-table1" title="c-table1" srcset="https://substackcdn.com/image/fetch/$s_!4czk!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74d49549-f1d1-428f-b202-669ebefeea81_1100x672.png 424w, https://substackcdn.com/image/fetch/$s_!4czk!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74d49549-f1d1-428f-b202-669ebefeea81_1100x672.png 848w, https://substackcdn.com/image/fetch/$s_!4czk!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74d49549-f1d1-428f-b202-669ebefeea81_1100x672.png 1272w, https://substackcdn.com/image/fetch/$s_!4czk!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74d49549-f1d1-428f-b202-669ebefeea81_1100x672.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The uncomfortable synthesis: <strong>the price ratios between tiers compressed at the same moment the absolute levels rose</strong>, and tiering strategies earn their complexity from the <em>ratio</em>, not the level. A prefix cache offloading to DDR5 at 20-25% of HBM&#8217;s per-bit cost is an easy win. </p><p>The same cache at 50-70% of HBM&#8217;s per-bit cost must clear a much higher bar once you charge it for transfer bandwidth, PCIe hops, TTFT-on-miss, and the payroll that maintains it.</p><h3>The offload break-even, derived</h3><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!m4Pk!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F796ba650-9bfc-40aa-ac94-41d75ad0ed8a_1100x449.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!m4Pk!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F796ba650-9bfc-40aa-ac94-41d75ad0ed8a_1100x449.png 424w, https://substackcdn.com/image/fetch/$s_!m4Pk!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F796ba650-9bfc-40aa-ac94-41d75ad0ed8a_1100x449.png 848w, https://substackcdn.com/image/fetch/$s_!m4Pk!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F796ba650-9bfc-40aa-ac94-41d75ad0ed8a_1100x449.png 1272w, https://substackcdn.com/image/fetch/$s_!m4Pk!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F796ba650-9bfc-40aa-ac94-41d75ad0ed8a_1100x449.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!m4Pk!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F796ba650-9bfc-40aa-ac94-41d75ad0ed8a_1100x449.png" width="1100" height="449" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/796ba650-9bfc-40aa-ac94-41d75ad0ed8a_1100x449.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:449,&quot;width&quot;:1100,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;c-math3&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="c-math3" title="c-math3" srcset="https://substackcdn.com/image/fetch/$s_!m4Pk!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F796ba650-9bfc-40aa-ac94-41d75ad0ed8a_1100x449.png 424w, https://substackcdn.com/image/fetch/$s_!m4Pk!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F796ba650-9bfc-40aa-ac94-41d75ad0ed8a_1100x449.png 848w, https://substackcdn.com/image/fetch/$s_!m4Pk!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F796ba650-9bfc-40aa-ac94-41d75ad0ed8a_1100x449.png 1272w, https://substackcdn.com/image/fetch/$s_!m4Pk!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F796ba650-9bfc-40aa-ac94-41d75ad0ed8a_1100x449.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>System prompts, shared codebase contexts, RAG corpus headers: cache. One-shot user uploads and cold agent trajectories: recompute. <strong>Cache residency is now a metered commodity; treat admission like an underwriting decision.</strong> </p><p>One nuance cuts the other way: NAND inflated <em>less</em> than DRAM, so the relative case for demoting warm entries to NVMe actually improved even as both tiers rose.</p><h3>Three channels into cost per token</h3><p><strong>Channel 1: CapEx per node.</strong> HBM is now the <strong>single largest line item in a flagship accelerator&#8217;s package BOM</strong> (SemiAnalysis&#8217;s finding for the GB300, after rising as a share every generation since Hopper) and adding host DDR5 and NVMe puts memory at roughly half the bill of a modern inference node (author&#8217;s estimate on top of that sourced base). </p><p>HBM moves slowly through long-term contracts; host DRAM and SSD are bought near spot: which is where 2026 budgets are bleeding now. A node spec&#8217;d mid-2025 with 1.5 TB DDR5 + 30 TB NVMe carries tens of thousands of dollars of new memory cost at current prices. Amortized over four years at high utilization it&#8217;s single-digit percent per token, but it compounds with Channel 2.</p><p><strong>Channel 2: the capacity ceiling (the big one).</strong> When memory binds, its price doesn&#8217;t just raise cost : <strong>it caps revenue per node</strong>. Max concurrency = free-memory-after-weights &#247; KV-bytes-per-sequence; decode throughput scales with batch until bandwidth saturates. As traffic mix shifts long-context (it is: agents, codebases, multimodal), each node serves fewer users and each token carries more fixed cost. </p><p><strong>Cost per token rises without any component getting more expensive: pure mix shift.</strong> The shortage means you can&#8217;t fix it by bolting on DRAM at 2024 prices; the escape routes are compression and disaggregation. Both are engineering, not procurement.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!_h4r!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F16083e0e-6498-46d6-80ae-a32f3316319e_1100x647.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!_h4r!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F16083e0e-6498-46d6-80ae-a32f3316319e_1100x647.png 424w, https://substackcdn.com/image/fetch/$s_!_h4r!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F16083e0e-6498-46d6-80ae-a32f3316319e_1100x647.png 848w, https://substackcdn.com/image/fetch/$s_!_h4r!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F16083e0e-6498-46d6-80ae-a32f3316319e_1100x647.png 1272w, https://substackcdn.com/image/fetch/$s_!_h4r!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F16083e0e-6498-46d6-80ae-a32f3316319e_1100x647.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!_h4r!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F16083e0e-6498-46d6-80ae-a32f3316319e_1100x647.png" width="1100" height="647" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/16083e0e-6498-46d6-80ae-a32f3316319e_1100x647.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:647,&quot;width&quot;:1100,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;c-fig6&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="c-fig6" title="c-fig6" srcset="https://substackcdn.com/image/fetch/$s_!_h4r!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F16083e0e-6498-46d6-80ae-a32f3316319e_1100x647.png 424w, https://substackcdn.com/image/fetch/$s_!_h4r!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F16083e0e-6498-46d6-80ae-a32f3316319e_1100x647.png 848w, https://substackcdn.com/image/fetch/$s_!_h4r!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F16083e0e-6498-46d6-80ae-a32f3316319e_1100x647.png 1272w, https://substackcdn.com/image/fetch/$s_!_h4r!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F16083e0e-6498-46d6-80ae-a32f3316319e_1100x647.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><strong>Channel 3: the tier-ratio compression</strong> of Fig. 3, which quietly deprecates a generation of offloading tricks and re-rates model-architecture choices (&#167;02, Fig. 4).</p><h3>Winners, losers, and the direction of the tilt</h3><p><strong>Hyperscalers and frontier labs are insulated.</strong> Multi-year HBM/DDR5 LTAs signed pre-squeeze, first-priority allocation (Micron: &#8220;larger, strategic customers&#8221;), fleets big enough that architecture fixes amortize across <strong>trillions of tokens. </strong></p><p><strong>Databricks&#8217; disclosure</strong> that cost-aware autoscaling and &#8220;model unit&#8221; abstractions cut GPU spend over 80% versus static provisioning (on a platform serving 120 trillion tokens a month) shows where leverage lives at that scale: <em>utilization engineering</em>, because their input prices are contractually smoothed.</p><p><strong>Mid-tier providers and self-hosters absorb the shock.</strong> They buy near spot, hold no allocation priority, and their pitch (undercutting frontier APIs) sits directly on the inflating commodity. Worse, their buyers are the newly capped: under a token budget, an enterprise squeezes more work from the same spend rather than paying up, so the mid-tier faces rising input costs and a customer base structurally optimizing against price <em>at the same time</em>.</p><p> That is the <strong>textbook geometry of a margin squeeze</strong>, and it lands here first. Watch for: per-token price floors firming through H2 2026 in open-weights serving; long-context surcharges turning explicit; and prompt-caching <em>pricing</em> rebalancing: cached-token discounts exist because storage was cheap relative to recompute, and that ratio just moved. Cache-storage line items (already visible in frontier pricing) will spread.</p><p><strong>The self-hosting math shifted in a direction almost nobody has updated for.</strong> The 2025 pitch (&#8221;your $15-25K refurbished H100 workstation replaces $500/month of API spend&#8221;) has two 2026 problems. The workstation&#8217;s own BOM inflated: RAM alone on a new server build can now approach the cost of an entire refurbished platform. And the API side is partially shielded by hyperscaler contract insulation plus superior compression engineering. </p><p>The perverse result: <strong>the memory supercycle is a centralizing force, it taxes small-fleet inference more heavily than hyperscale inference, at precisely the moment open-weights quality made decentralized serving viable.</strong> </p><p>If you&#8217;re modeling a GPU buy this year, redo it with 2026 memory line items, and note that used enterprise hardware with pre-crisis RAM already installed is currently the cleanest arbitrage in the market, which is exactly why the refurb channel is having a moment.</p><div><hr></div><h2>Proof of pass-through: the March repricing, and the token bill in real dollars</h2><p>A thesis this size needs a smoking gun: evidence that memory contract prices actually reach the price <em>you</em> pay per GPU-hour, on a measurable lag. Q1 2026 provided it.</p><p>Silicon Data&#8217;s SDB200RT index (the standardized benchmark for B200 cloud rental pricing) opened 2026 at 4.40, drifted through January and February, then <strong>rose 23.6% inside March alone</strong>, crossing 5.0 on March 15 and 6.0 on March 23 before settling at 5.48: up 24.4% year-to-date.</p><p> Their attribution is the entire argument of this issue in one sentence: <strong>Samsung and SK Hynix raised HBM3e contract prices ~20% for 2026 deliveries, NVIDIA revised hardware MSRPs upward in late February citing memory component costs, and those input costs flowed into cloud hourly rates with a 4-6 week lag</strong>: landing squarely in March. </p><p>Meanwhile the two-year-old H100, whose memory was bought at pre-squeeze prices, traded in a tight, boring band all quarter. Same market, same month: <strong>the GPUs carrying new memory repriced; the GPUs carrying old memory didn&#8217;t.</strong> That divergence <em>is</em> the supercycle reaching your invoice.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!r2ZO!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e7e1680-0830-4ab0-b17f-1a7755da9205_1100x637.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!r2ZO!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e7e1680-0830-4ab0-b17f-1a7755da9205_1100x637.png 424w, https://substackcdn.com/image/fetch/$s_!r2ZO!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e7e1680-0830-4ab0-b17f-1a7755da9205_1100x637.png 848w, https://substackcdn.com/image/fetch/$s_!r2ZO!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e7e1680-0830-4ab0-b17f-1a7755da9205_1100x637.png 1272w, https://substackcdn.com/image/fetch/$s_!r2ZO!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e7e1680-0830-4ab0-b17f-1a7755da9205_1100x637.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!r2ZO!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e7e1680-0830-4ab0-b17f-1a7755da9205_1100x637.png" width="1100" height="637" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/3e7e1680-0830-4ab0-b17f-1a7755da9205_1100x637.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:637,&quot;width&quot;:1100,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;c-fig7&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="c-fig7" title="c-fig7" srcset="https://substackcdn.com/image/fetch/$s_!r2ZO!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e7e1680-0830-4ab0-b17f-1a7755da9205_1100x637.png 424w, https://substackcdn.com/image/fetch/$s_!r2ZO!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e7e1680-0830-4ab0-b17f-1a7755da9205_1100x637.png 848w, https://substackcdn.com/image/fetch/$s_!r2ZO!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e7e1680-0830-4ab0-b17f-1a7755da9205_1100x637.png 1272w, https://substackcdn.com/image/fetch/$s_!r2ZO!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e7e1680-0830-4ab0-b17f-1a7755da9205_1100x637.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h3>The token bill, priced in real mid-2026 dollars</h3><p>With real rental rates in hand (H200 working median ~$3.50/GPU-hr on-demand across neocloud trackers (cohort median ~$4.00; floor $2.30; hyperscalers to $13.78)) the bandwidth model from &#167;03 becomes an actual price sheet. </p><p>These are <strong>bandwidth-bound floors</strong>: ideal kernels, full overlap, output tokens only. Real deployments land at 50-70% of these throughputs; the <em>ratios</em> between cells are the robust result.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!-mk-!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd59ba112-c8ab-4aed-b804-076fc419af92_1100x313.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!-mk-!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd59ba112-c8ab-4aed-b804-076fc419af92_1100x313.png 424w, https://substackcdn.com/image/fetch/$s_!-mk-!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd59ba112-c8ab-4aed-b804-076fc419af92_1100x313.png 848w, https://substackcdn.com/image/fetch/$s_!-mk-!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd59ba112-c8ab-4aed-b804-076fc419af92_1100x313.png 1272w, https://substackcdn.com/image/fetch/$s_!-mk-!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd59ba112-c8ab-4aed-b804-076fc419af92_1100x313.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!-mk-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd59ba112-c8ab-4aed-b804-076fc419af92_1100x313.png" width="1100" height="313" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d59ba112-c8ab-4aed-b804-076fc419af92_1100x313.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:313,&quot;width&quot;:1100,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;c-math4&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="c-math4" title="c-math4" srcset="https://substackcdn.com/image/fetch/$s_!-mk-!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd59ba112-c8ab-4aed-b804-076fc419af92_1100x313.png 424w, https://substackcdn.com/image/fetch/$s_!-mk-!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd59ba112-c8ab-4aed-b804-076fc419af92_1100x313.png 848w, https://substackcdn.com/image/fetch/$s_!-mk-!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd59ba112-c8ab-4aed-b804-076fc419af92_1100x313.png 1272w, https://substackcdn.com/image/fetch/$s_!-mk-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd59ba112-c8ab-4aed-b804-076fc419af92_1100x313.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!sZWU!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd100fe8c-3a01-47c9-ba87-4111eac15e28_1100x689.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!sZWU!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd100fe8c-3a01-47c9-ba87-4111eac15e28_1100x689.png 424w, https://substackcdn.com/image/fetch/$s_!sZWU!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd100fe8c-3a01-47c9-ba87-4111eac15e28_1100x689.png 848w, https://substackcdn.com/image/fetch/$s_!sZWU!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd100fe8c-3a01-47c9-ba87-4111eac15e28_1100x689.png 1272w, https://substackcdn.com/image/fetch/$s_!sZWU!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd100fe8c-3a01-47c9-ba87-4111eac15e28_1100x689.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!sZWU!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd100fe8c-3a01-47c9-ba87-4111eac15e28_1100x689.png" width="1100" height="689" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d100fe8c-3a01-47c9-ba87-4111eac15e28_1100x689.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:689,&quot;width&quot;:1100,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;c-fig8&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="c-fig8" title="c-fig8" srcset="https://substackcdn.com/image/fetch/$s_!sZWU!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd100fe8c-3a01-47c9-ba87-4111eac15e28_1100x689.png 424w, https://substackcdn.com/image/fetch/$s_!sZWU!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd100fe8c-3a01-47c9-ba87-4111eac15e28_1100x689.png 848w, https://substackcdn.com/image/fetch/$s_!sZWU!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd100fe8c-3a01-47c9-ba87-4111eac15e28_1100x689.png 1272w, https://substackcdn.com/image/fetch/$s_!sZWU!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd100fe8c-3a01-47c9-ba87-4111eac15e28_1100x689.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>One honest caveat, stated before a reader states it for me: this model charges every decode step for streaming the full weights and KV and ignores prefill cost, which at 128K is substantial and shifts spend toward exactly the CPX-shaped silicon. </p><p>A production number also carries utilization (Databricks&#8217; 80% savings figure is the size of <em>that</em> lever), SLO headroom, and interconnect overhead in multi-GPU serving. <strong>The model&#8217;s job is not to predict your invoice to the cent; it is to rank your decisions, and the ranking is unambiguous: </strong>context policy first, KV dtype second, hardware price third.</p><h3>The spread, marked to market. July 2, 2026</h3><p>Now close the loop the title promises. As of July 2, flat per-token pricing for Llama-3.3-70B-class serving spans <strong>$0.31-$0.90 per million output tokens</strong> across the fifteen providers Artificial Analysis tracks (DeepInfra&#8217;s FP8 &#8220;Turbo&#8221; at $0.40, Groq at $0.79, Together AI and Fireworks at $0.88-0.90) and the rate is <em>flat across the model&#8217;s 131K context window</em>. </p><p>Set those prices against this section&#8217;s floors and the cross-subsidy stops being an inference and becomes a table:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!PvSl!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a284d54-9659-4978-82b2-8a7123776172_1100x459.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!PvSl!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a284d54-9659-4978-82b2-8a7123776172_1100x459.png 424w, https://substackcdn.com/image/fetch/$s_!PvSl!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a284d54-9659-4978-82b2-8a7123776172_1100x459.png 848w, https://substackcdn.com/image/fetch/$s_!PvSl!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a284d54-9659-4978-82b2-8a7123776172_1100x459.png 1272w, https://substackcdn.com/image/fetch/$s_!PvSl!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a284d54-9659-4978-82b2-8a7123776172_1100x459.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!PvSl!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a284d54-9659-4978-82b2-8a7123776172_1100x459.png" width="1100" height="459" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/9a284d54-9659-4978-82b2-8a7123776172_1100x459.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:459,&quot;width&quot;:1100,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;c-table2&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="c-table2" title="c-table2" srcset="https://substackcdn.com/image/fetch/$s_!PvSl!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a284d54-9659-4978-82b2-8a7123776172_1100x459.png 424w, https://substackcdn.com/image/fetch/$s_!PvSl!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a284d54-9659-4978-82b2-8a7123776172_1100x459.png 848w, https://substackcdn.com/image/fetch/$s_!PvSl!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a284d54-9659-4978-82b2-8a7123776172_1100x459.png 1272w, https://substackcdn.com/image/fetch/$s_!PvSl!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a284d54-9659-4978-82b2-8a7123776172_1100x459.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Note the confirming detail hiding in the market data: the price leader doesn&#8217;t serve the reference model at all, it serves an FP8-quantized variant. <strong>The cheapest provider in the market is this issue&#8217;s playbook, priced and shipped.</strong> </p><p>And the honest caveats, before a reader raises them: these floors are single-GPU bandwidth ideals (real serving is less efficient, which raises them), while real fleets also earn input-token revenue and batch mixed contexts (which lowers effective floors), so treat the magnitudes as directional. </p><p>The <em>sign pattern</em> is the robust result: <strong>at July 2026 flat prices, long-context traffic on 70B-class serving is sold below its modeled production floor, and budget-capped buyers remove the option of raising the flat rate to fix it.</strong> Hence the final item of &#167;07&#8217;s playbook.</p><div><hr></div><h2>Eight decisions, ranked by dollar leverage</h2><p>Everything above, compiled into action: ranked by expected cost impact per unit of engineering effort for a team serving open-weights models at scale in 2026. Under the scissors of &#167;00, read &#8220;cost impact&#8221; as what it now is: <strong>spread defense</strong>: every dollar of floor you remove is margin the ceiling can no longer take from you.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!xTts!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F193e4edc-128f-483f-8df8-7f9a299c4472_1100x405.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!xTts!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F193e4edc-128f-483f-8df8-7f9a299c4472_1100x405.png 424w, https://substackcdn.com/image/fetch/$s_!xTts!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F193e4edc-128f-483f-8df8-7f9a299c4472_1100x405.png 848w, https://substackcdn.com/image/fetch/$s_!xTts!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F193e4edc-128f-483f-8df8-7f9a299c4472_1100x405.png 1272w, https://substackcdn.com/image/fetch/$s_!xTts!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F193e4edc-128f-483f-8df8-7f9a299c4472_1100x405.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!xTts!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F193e4edc-128f-483f-8df8-7f9a299c4472_1100x405.png" width="1100" height="405" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/193e4edc-128f-483f-8df8-7f9a299c4472_1100x405.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:405,&quot;width&quot;:1100,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;c-play01&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="c-play01" title="c-play01" srcset="https://substackcdn.com/image/fetch/$s_!xTts!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F193e4edc-128f-483f-8df8-7f9a299c4472_1100x405.png 424w, https://substackcdn.com/image/fetch/$s_!xTts!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F193e4edc-128f-483f-8df8-7f9a299c4472_1100x405.png 848w, https://substackcdn.com/image/fetch/$s_!xTts!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F193e4edc-128f-483f-8df8-7f9a299c4472_1100x405.png 1272w, https://substackcdn.com/image/fetch/$s_!xTts!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F193e4edc-128f-483f-8df8-7f9a299c4472_1100x405.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!qafv!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6542ca3-f586-484f-920e-11ae05da957d_1100x378.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!qafv!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6542ca3-f586-484f-920e-11ae05da957d_1100x378.png 424w, https://substackcdn.com/image/fetch/$s_!qafv!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6542ca3-f586-484f-920e-11ae05da957d_1100x378.png 848w, https://substackcdn.com/image/fetch/$s_!qafv!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6542ca3-f586-484f-920e-11ae05da957d_1100x378.png 1272w, https://substackcdn.com/image/fetch/$s_!qafv!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6542ca3-f586-484f-920e-11ae05da957d_1100x378.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!qafv!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6542ca3-f586-484f-920e-11ae05da957d_1100x378.png" width="1100" height="378" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b6542ca3-f586-484f-920e-11ae05da957d_1100x378.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:378,&quot;width&quot;:1100,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;c-play02&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="c-play02" title="c-play02" srcset="https://substackcdn.com/image/fetch/$s_!qafv!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6542ca3-f586-484f-920e-11ae05da957d_1100x378.png 424w, https://substackcdn.com/image/fetch/$s_!qafv!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6542ca3-f586-484f-920e-11ae05da957d_1100x378.png 848w, https://substackcdn.com/image/fetch/$s_!qafv!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6542ca3-f586-484f-920e-11ae05da957d_1100x378.png 1272w, https://substackcdn.com/image/fetch/$s_!qafv!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6542ca3-f586-484f-920e-11ae05da957d_1100x378.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!OcjK!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb08315fd-80fd-475b-b82f-ecd2565bc281_1100x378.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!OcjK!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb08315fd-80fd-475b-b82f-ecd2565bc281_1100x378.png 424w, https://substackcdn.com/image/fetch/$s_!OcjK!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb08315fd-80fd-475b-b82f-ecd2565bc281_1100x378.png 848w, https://substackcdn.com/image/fetch/$s_!OcjK!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb08315fd-80fd-475b-b82f-ecd2565bc281_1100x378.png 1272w, https://substackcdn.com/image/fetch/$s_!OcjK!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb08315fd-80fd-475b-b82f-ecd2565bc281_1100x378.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!OcjK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb08315fd-80fd-475b-b82f-ecd2565bc281_1100x378.png" width="1100" height="378" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b08315fd-80fd-475b-b82f-ecd2565bc281_1100x378.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:378,&quot;width&quot;:1100,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;c-play03&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="c-play03" title="c-play03" srcset="https://substackcdn.com/image/fetch/$s_!OcjK!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb08315fd-80fd-475b-b82f-ecd2565bc281_1100x378.png 424w, https://substackcdn.com/image/fetch/$s_!OcjK!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb08315fd-80fd-475b-b82f-ecd2565bc281_1100x378.png 848w, https://substackcdn.com/image/fetch/$s_!OcjK!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb08315fd-80fd-475b-b82f-ecd2565bc281_1100x378.png 1272w, https://substackcdn.com/image/fetch/$s_!OcjK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb08315fd-80fd-475b-b82f-ecd2565bc281_1100x378.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!AJii!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4052a6f5-6c01-4e98-a857-d06cba0e22e5_1100x415.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!AJii!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4052a6f5-6c01-4e98-a857-d06cba0e22e5_1100x415.png 424w, https://substackcdn.com/image/fetch/$s_!AJii!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4052a6f5-6c01-4e98-a857-d06cba0e22e5_1100x415.png 848w, https://substackcdn.com/image/fetch/$s_!AJii!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4052a6f5-6c01-4e98-a857-d06cba0e22e5_1100x415.png 1272w, https://substackcdn.com/image/fetch/$s_!AJii!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4052a6f5-6c01-4e98-a857-d06cba0e22e5_1100x415.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!AJii!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4052a6f5-6c01-4e98-a857-d06cba0e22e5_1100x415.png" width="1100" height="415" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/4052a6f5-6c01-4e98-a857-d06cba0e22e5_1100x415.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:415,&quot;width&quot;:1100,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;c-play04&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="c-play04" title="c-play04" srcset="https://substackcdn.com/image/fetch/$s_!AJii!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4052a6f5-6c01-4e98-a857-d06cba0e22e5_1100x415.png 424w, https://substackcdn.com/image/fetch/$s_!AJii!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4052a6f5-6c01-4e98-a857-d06cba0e22e5_1100x415.png 848w, https://substackcdn.com/image/fetch/$s_!AJii!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4052a6f5-6c01-4e98-a857-d06cba0e22e5_1100x415.png 1272w, https://substackcdn.com/image/fetch/$s_!AJii!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4052a6f5-6c01-4e98-a857-d06cba0e22e5_1100x415.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!lrul!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2096bc08-9268-4de0-b7f2-3285f250e586_1100x306.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!lrul!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2096bc08-9268-4de0-b7f2-3285f250e586_1100x306.png 424w, https://substackcdn.com/image/fetch/$s_!lrul!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2096bc08-9268-4de0-b7f2-3285f250e586_1100x306.png 848w, https://substackcdn.com/image/fetch/$s_!lrul!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2096bc08-9268-4de0-b7f2-3285f250e586_1100x306.png 1272w, https://substackcdn.com/image/fetch/$s_!lrul!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2096bc08-9268-4de0-b7f2-3285f250e586_1100x306.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!lrul!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2096bc08-9268-4de0-b7f2-3285f250e586_1100x306.png" width="1100" height="306" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2096bc08-9268-4de0-b7f2-3285f250e586_1100x306.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:306,&quot;width&quot;:1100,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;c-play05&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="c-play05" title="c-play05" srcset="https://substackcdn.com/image/fetch/$s_!lrul!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2096bc08-9268-4de0-b7f2-3285f250e586_1100x306.png 424w, https://substackcdn.com/image/fetch/$s_!lrul!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2096bc08-9268-4de0-b7f2-3285f250e586_1100x306.png 848w, https://substackcdn.com/image/fetch/$s_!lrul!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2096bc08-9268-4de0-b7f2-3285f250e586_1100x306.png 1272w, https://substackcdn.com/image/fetch/$s_!lrul!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2096bc08-9268-4de0-b7f2-3285f250e586_1100x306.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!XHBY!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ca5727f-44af-4947-b152-5d01560fcf6d_1100x342.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!XHBY!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ca5727f-44af-4947-b152-5d01560fcf6d_1100x342.png 424w, https://substackcdn.com/image/fetch/$s_!XHBY!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ca5727f-44af-4947-b152-5d01560fcf6d_1100x342.png 848w, https://substackcdn.com/image/fetch/$s_!XHBY!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ca5727f-44af-4947-b152-5d01560fcf6d_1100x342.png 1272w, https://substackcdn.com/image/fetch/$s_!XHBY!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ca5727f-44af-4947-b152-5d01560fcf6d_1100x342.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!XHBY!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ca5727f-44af-4947-b152-5d01560fcf6d_1100x342.png" width="1100" height="342" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/8ca5727f-44af-4947-b152-5d01560fcf6d_1100x342.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:342,&quot;width&quot;:1100,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;c-play06&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="c-play06" title="c-play06" srcset="https://substackcdn.com/image/fetch/$s_!XHBY!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ca5727f-44af-4947-b152-5d01560fcf6d_1100x342.png 424w, https://substackcdn.com/image/fetch/$s_!XHBY!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ca5727f-44af-4947-b152-5d01560fcf6d_1100x342.png 848w, https://substackcdn.com/image/fetch/$s_!XHBY!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ca5727f-44af-4947-b152-5d01560fcf6d_1100x342.png 1272w, https://substackcdn.com/image/fetch/$s_!XHBY!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ca5727f-44af-4947-b152-5d01560fcf6d_1100x342.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!3N8x!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb7585f3d-de3d-479f-a018-ab52d6708861_1100x378.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!3N8x!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb7585f3d-de3d-479f-a018-ab52d6708861_1100x378.png 424w, https://substackcdn.com/image/fetch/$s_!3N8x!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb7585f3d-de3d-479f-a018-ab52d6708861_1100x378.png 848w, https://substackcdn.com/image/fetch/$s_!3N8x!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb7585f3d-de3d-479f-a018-ab52d6708861_1100x378.png 1272w, https://substackcdn.com/image/fetch/$s_!3N8x!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb7585f3d-de3d-479f-a018-ab52d6708861_1100x378.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!3N8x!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb7585f3d-de3d-479f-a018-ab52d6708861_1100x378.png" width="1100" height="378" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b7585f3d-de3d-479f-a018-ab52d6708861_1100x378.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:378,&quot;width&quot;:1100,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;c-play07&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="c-play07" title="c-play07" srcset="https://substackcdn.com/image/fetch/$s_!3N8x!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb7585f3d-de3d-479f-a018-ab52d6708861_1100x378.png 424w, https://substackcdn.com/image/fetch/$s_!3N8x!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb7585f3d-de3d-479f-a018-ab52d6708861_1100x378.png 848w, https://substackcdn.com/image/fetch/$s_!3N8x!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb7585f3d-de3d-479f-a018-ab52d6708861_1100x378.png 1272w, https://substackcdn.com/image/fetch/$s_!3N8x!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb7585f3d-de3d-479f-a018-ab52d6708861_1100x378.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!gfIz!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b53e289-c6b1-459b-a166-a7c2ee46298b_1100x514.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!gfIz!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b53e289-c6b1-459b-a166-a7c2ee46298b_1100x514.png 424w, https://substackcdn.com/image/fetch/$s_!gfIz!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b53e289-c6b1-459b-a166-a7c2ee46298b_1100x514.png 848w, https://substackcdn.com/image/fetch/$s_!gfIz!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b53e289-c6b1-459b-a166-a7c2ee46298b_1100x514.png 1272w, https://substackcdn.com/image/fetch/$s_!gfIz!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b53e289-c6b1-459b-a166-a7c2ee46298b_1100x514.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!gfIz!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b53e289-c6b1-459b-a166-a7c2ee46298b_1100x514.png" width="1100" height="514" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/1b53e289-c6b1-459b-a166-a7c2ee46298b_1100x514.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:514,&quot;width&quot;:1100,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;c-play08&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="c-play08" title="c-play08" srcset="https://substackcdn.com/image/fetch/$s_!gfIz!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b53e289-c6b1-459b-a166-a7c2ee46298b_1100x514.png 424w, https://substackcdn.com/image/fetch/$s_!gfIz!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b53e289-c6b1-459b-a166-a7c2ee46298b_1100x514.png 848w, https://substackcdn.com/image/fetch/$s_!gfIz!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b53e289-c6b1-459b-a166-a7c2ee46298b_1100x514.png 1272w, https://substackcdn.com/image/fetch/$s_!gfIz!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b53e289-c6b1-459b-a166-a7c2ee46298b_1100x514.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div><hr></div><h2>Five falsifiable predictions</h2><p>Scored publicly in Q1 2027, alongside the Issue 04 speculative-decoding calls. Confidence is stated so being wrong costs me something.</p><p><strong>P1 &#183; Supplyconfidence: high</strong><br>DRAM relief does not arrive before late 2027. New capacity (Micron Boise) ramps 2027-2028; suppliers themselves guide consumer relief to ~2028. Any 2027 hardware-refresh model priced at 2024 memory levels is fiction.</p><p><strong>P2 &#183; Architectureconfidence: high</strong><br>Open-weights flagships converge on compressed KV. By mid-2027 the majority of new frontier-adjacent open releases ship MLA-like latent attention, &#8804;8-head GQA, or hybrid sliding windows: marketed explicitly as serving-cost features. Model architecture is now downstream of the DRAM spot price.</p><p><strong>P3 &#183; Pricingconfidence: medium-high</strong><br>Per-token API pricing bifurcates by context residency: at least two major providers introduce or restructure explicit cached-context <em>storage</em> pricing (per-token-per-hour or equivalent) by mid-2027, decoupling state from compute. Partial confirmation is already on the books (Google Vertex bills per-hour cache storage and Anthropic bills TTL-tiered cache writes) so the live prediction is that this becomes the <em>norm in open-weights serving</em>, not just frontier practice.</p><p><strong>P4 &#183; Siliconconfidence: medium</strong><br>Prefill-specialized silicon becomes a category, not a SKU: a CPX competitor (AMD or an ASIC player) is announced within 12 months of CPX GA, and &#8220;prefill accelerator&#8221; enters the standard rack taxonomy.</p><p><strong>P5 &#183; Marketconfidence: medium</strong><br>The open-weights serving price war pauses: median $/Mtok for 70B-class serving on third-party clouds goes flat-to-up through H1 2027 (the first sustained non-decline in that series since it has existed. The March B200 repricing (Fig. 7) is the leading edge) and enterprise token budgets are the demand-side mechanism that lets long-context surcharges and cache-storage fees actually stick, because a buyer operating under a cap optimizes usage instead of switching providers.</p><h3>The dashboard: six leading indicators, with trigger levels</h3><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!ozsr!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff24d4938-9512-4e15-9d7e-41f79234be1a_1100x1033.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!ozsr!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff24d4938-9512-4e15-9d7e-41f79234be1a_1100x1033.png 424w, https://substackcdn.com/image/fetch/$s_!ozsr!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff24d4938-9512-4e15-9d7e-41f79234be1a_1100x1033.png 848w, https://substackcdn.com/image/fetch/$s_!ozsr!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff24d4938-9512-4e15-9d7e-41f79234be1a_1100x1033.png 1272w, https://substackcdn.com/image/fetch/$s_!ozsr!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff24d4938-9512-4e15-9d7e-41f79234be1a_1100x1033.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!ozsr!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff24d4938-9512-4e15-9d7e-41f79234be1a_1100x1033.png" width="1100" height="1033" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f24d4938-9512-4e15-9d7e-41f79234be1a_1100x1033.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1033,&quot;width&quot;:1100,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;c-table3&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="c-table3" title="c-table3" srcset="https://substackcdn.com/image/fetch/$s_!ozsr!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff24d4938-9512-4e15-9d7e-41f79234be1a_1100x1033.png 424w, https://substackcdn.com/image/fetch/$s_!ozsr!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff24d4938-9512-4e15-9d7e-41f79234be1a_1100x1033.png 848w, https://substackcdn.com/image/fetch/$s_!ozsr!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff24d4938-9512-4e15-9d7e-41f79234be1a_1100x1033.png 1272w, https://substackcdn.com/image/fetch/$s_!ozsr!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff24d4938-9512-4e15-9d7e-41f79234be1a_1100x1033.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!GcTh!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F32114b54-e57d-4baa-bcb8-d576d0a9d59c_1100x593.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!GcTh!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F32114b54-e57d-4baa-bcb8-d576d0a9d59c_1100x593.png 424w, https://substackcdn.com/image/fetch/$s_!GcTh!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F32114b54-e57d-4baa-bcb8-d576d0a9d59c_1100x593.png 848w, https://substackcdn.com/image/fetch/$s_!GcTh!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F32114b54-e57d-4baa-bcb8-d576d0a9d59c_1100x593.png 1272w, https://substackcdn.com/image/fetch/$s_!GcTh!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F32114b54-e57d-4baa-bcb8-d576d0a9d59c_1100x593.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!GcTh!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F32114b54-e57d-4baa-bcb8-d576d0a9d59c_1100x593.png" width="1100" height="593" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/32114b54-e57d-4baa-bcb8-d576d0a9d59c_1100x593.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:593,&quot;width&quot;:1100,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;c-fig9&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="c-fig9" title="c-fig9" srcset="https://substackcdn.com/image/fetch/$s_!GcTh!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F32114b54-e57d-4baa-bcb8-d576d0a9d59c_1100x593.png 424w, https://substackcdn.com/image/fetch/$s_!GcTh!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F32114b54-e57d-4baa-bcb8-d576d0a9d59c_1100x593.png 848w, https://substackcdn.com/image/fetch/$s_!GcTh!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F32114b54-e57d-4baa-bcb8-d576d0a9d59c_1100x593.png 1272w, https://substackcdn.com/image/fetch/$s_!GcTh!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F32114b54-e57d-4baa-bcb8-d576d0a9d59c_1100x593.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div><hr></div><h2>The spread is the product now</h2><p>For two years, inference economics was a story about FLOPs: kernel efficiency, quantized matmuls, speculative decoding, Blackwell&#8217;s FP4 throughput. </p><p>Those battles were real; this publication has spent six issues inside them. But they were fought against a background assumption so universal nobody stated it: <em>bytes are cheap and getting cheaper.</em></p><p>That assumption died in Q1 2026. <strong>The binding constraint of the inference industry is now measured in wafer starts and TSV yields</strong>: in exabytes of KV state colliding with 10-16% annual bit-supply growth, in an IDC report that uses the word &#8220;permanent,&#8221; in a rental index that repriced 24% in one month because a memory contract reset five weeks earlier. </p><p>The consequences run one direction through the entire stack: hardware specializes by phase <strong>because HBM is too precious to waste</strong> on prefill; model architectures compress their KV because fat caches became unserveable; serving stacks reorganize around shipping cached state across fabrics because holding it still got expensive; and the cost advantage tilts toward <strong>whoever holds allocation contracts </strong>and compression engineering: which is to say, toward the largest players, in an industry that spent two years congratulating itself on decentralizing.</p><p>The engineers who internalize this fastest hold a simple advantage: while the market reprices memory, they reprice their <em>need</em> for it. <strong>Every KV byte you decline to allocate is bought at 2026 prices and sold at 2026 prices, at 100% margin, with zero lead time and no allocation meeting.</strong> The wafer sets the cost floor. The wallet sets the price ceiling. </p><p>The engineering in between is the only variable you control, and in a margin regime, that engineering stops being a cost-optimization side quest and <strong>becomes the P&amp;L itself.</strong> Squeezes do not reward the biggest fleet; they reward the widest spread per byte.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!ULaf!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd001de35-a3bc-4195-a51a-e01c27856add_1100x311.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!ULaf!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd001de35-a3bc-4195-a51a-e01c27856add_1100x311.png 424w, https://substackcdn.com/image/fetch/$s_!ULaf!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd001de35-a3bc-4195-a51a-e01c27856add_1100x311.png 848w, https://substackcdn.com/image/fetch/$s_!ULaf!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd001de35-a3bc-4195-a51a-e01c27856add_1100x311.png 1272w, https://substackcdn.com/image/fetch/$s_!ULaf!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd001de35-a3bc-4195-a51a-e01c27856add_1100x311.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!ULaf!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd001de35-a3bc-4195-a51a-e01c27856add_1100x311.png" width="1100" height="311" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d001de35-a3bc-4195-a51a-e01c27856add_1100x311.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:311,&quot;width&quot;:1100,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;c-callout3&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="c-callout3" title="c-callout3" srcset="https://substackcdn.com/image/fetch/$s_!ULaf!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd001de35-a3bc-4195-a51a-e01c27856add_1100x311.png 424w, https://substackcdn.com/image/fetch/$s_!ULaf!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd001de35-a3bc-4195-a51a-e01c27856add_1100x311.png 848w, https://substackcdn.com/image/fetch/$s_!ULaf!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd001de35-a3bc-4195-a51a-e01c27856add_1100x311.png 1272w, https://substackcdn.com/image/fetch/$s_!ULaf!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd001de35-a3bc-4195-a51a-e01c27856add_1100x311.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><strong>METHOD &amp; VERIFICATION</strong>. Market figures trace to named sources below; all KV, roofline, concurrency, and $/Mtok arithmetic is the author&#8217;s, derived in-text from published model configs and device datasheets so you can check every step. Cost floors are bandwidth-bound ideals: absolute values are lower bounds, ratios are the robust result. </p><p>The Fig. 7 path between published waypoints is interpolated; Fig. 3&#8217;s intermediate points are directional; Fig. 0 is a directional synthesis whose turning points, not slopes, are the sourced claims. Pricing data was spot-checked the week of publication and moves weekly: recheck before committing capital. Corrections run at the top of the next issue, as always.</p><div><hr></div><p><strong>SOURCES:</strong> <em>TrendForce: 1Q26 DRAM/NAND contract revisions; &#8220;Memory Wall&#8221; research (Jan 2026); HBM adoption &amp; 2029 inference-demand outlook; AI wafer-equivalent consumption via Commercial Times (Dec 2025) &#183; IDC (&#8221;Global Memory Shortage Crisis&#8221; (Feb 2026) &#183; IEEE Spectrum) &#8220;AI Is a Memory Hog&#8221; (Apr 2026) &#183; Counterpoint Research (Q1&#8217;26 spot DRAM tracking &#183; CNBC) Micron/SK Hynix/Samsung allocation &amp; CES 2026 reporting (Jan 2026) &#183; Avnet/Omdia/BofA (supercycle &amp; supply-relief estimates &#183; SemiAnalysis) &#8220;Another Giant Leap: The Rubin CPX Specialized Accelerator &amp; Rack&#8221; (Sep 2025) &#183; The Register (CES 2026 Vera Rubin systems coverage (Jan 2026) &#183; Tom&#8217;s Hardware) &#8220;Nvidia&#8217;s Vera Rubin platform in depth&#8221; (Nov 2025, incl. BlueField-4 KV-cache SSD) &#183; Futurum/Ori/NADDOD (Rubin CPX technical analyses &#183; NVIDIA) AI Infra Summit &amp; GTC 2026 announcements &#183; Databricks Engineering (&#8221;Reliable LLM Inference at Scale&#8221; (May 2026) &#183; SemiAnalysis) &#8220;TokenBudgeting: Our Conversations with Enterprises on Token Spend&#8221; (Jul 2, 2026) &#183; Artificial Analysis (Llama 3.3 70B provider benchmarking &amp; cache-pricing mechanics (Jul 2026) &#183; aipricing.guru / Price Per Token) provider price pages (sourced Jul 2, 2026) &#183; Silicon Data. SDB200RT B200 index, March 2026 update &#183; aimultiple GPU Rental Price Index; getdeploying B200/H200 trackers; Jarvislabs H200 pricing (2026) &#183; DeepSeek-V3 technical report (MLA configuration &#183; vLLM) NVFP4 KV-cache PR (in progress at publication).</em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Software Frontier is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[The Router and the Wire]]></title><description><![CDATA[Mixture-of-experts promised cheaper inference by doing less arithmetic. The bill did not disappear. It moved into the network, and the entire shape of a 2026 serving stack is the receipt.]]></description><link>https://www.thesoftwarefrontier.com/p/the-router-and-the-wire</link><guid isPermaLink="false">https://www.thesoftwarefrontier.com/p/the-router-and-the-wire</guid><dc:creator><![CDATA[Lorenzo Bradanini]]></dc:creator><pubDate>Mon, 29 Jun 2026 21:18:38 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!5xDE!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb9cb9fb9-d6a7-45e1-a371-0d6456ee6809_1024x1536.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!5xDE!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb9cb9fb9-d6a7-45e1-a371-0d6456ee6809_1024x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!5xDE!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb9cb9fb9-d6a7-45e1-a371-0d6456ee6809_1024x1536.png 424w, https://substackcdn.com/image/fetch/$s_!5xDE!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb9cb9fb9-d6a7-45e1-a371-0d6456ee6809_1024x1536.png 848w, https://substackcdn.com/image/fetch/$s_!5xDE!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb9cb9fb9-d6a7-45e1-a371-0d6456ee6809_1024x1536.png 1272w, https://substackcdn.com/image/fetch/$s_!5xDE!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb9cb9fb9-d6a7-45e1-a371-0d6456ee6809_1024x1536.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!5xDE!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb9cb9fb9-d6a7-45e1-a371-0d6456ee6809_1024x1536.png" width="1024" height="1536" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b9cb9fb9-d6a7-45e1-a371-0d6456ee6809_1024x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1536,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2363576,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/199726817?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb9cb9fb9-d6a7-45e1-a371-0d6456ee6809_1024x1536.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!5xDE!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb9cb9fb9-d6a7-45e1-a371-0d6456ee6809_1024x1536.png 424w, https://substackcdn.com/image/fetch/$s_!5xDE!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb9cb9fb9-d6a7-45e1-a371-0d6456ee6809_1024x1536.png 848w, https://substackcdn.com/image/fetch/$s_!5xDE!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb9cb9fb9-d6a7-45e1-a371-0d6456ee6809_1024x1536.png 1272w, https://substackcdn.com/image/fetch/$s_!5xDE!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb9cb9fb9-d6a7-45e1-a371-0d6456ee6809_1024x1536.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>The cost did not vanish. It moved.</h2><p>Every few months the inference market tells itself a story about where the money goes, and every few months the story is wrong in the same direction. For two years the story was memory. </p><p>The <strong>key-value cache grew with context</strong>, the weights grew with parameter count, and the binding constraint on a serving deployment was how many bytes of high-bandwidth memory you could buy and how fast you could read them. </p><p>That story was true, and it is the subject of two earlier issues of this publication. It is also no longer the whole story for the models that now define the frontier.</p><p>The frontier moved to sparsity. A <strong>dense model the size of DeepSeek-V3 </strong>would activate all 671 billion of its parameters on every token it processed. The mixture-of-experts version activates roughly 37 billion. </p><p>On paper that is a reduction in arithmetic of roughly eighteen to one, and it is the single reason a model with two-thirds of a trillion parameters can be served at all without a fleet of accelerators per request. </p><p>The promise of the architecture was always framed in floating-point operations: do less math, pay less money. The <strong>vLLM and SGLang communities</strong>, NVIDIA, DeepSeek, and every serving vendor in between repeated some version of it.</p><p>The arithmetic did get cheaper. What the framing left out is that the arithmetic was never the part that was hard to scale. When you spread 256 experts across <strong>dozens or hundreds of accelerators</strong>, a token routed to eight of them has to physically travel to those eight accelerators, be processed, and travel back to be recombined. </p><p>That round trip is an all-to-all communication pattern, and it does not appear in<strong> any FLOP count</strong>. It is the part of the bill that the sparsity story quietly moved off the compute line and onto the network line, where it has been growing ever since.</p><p>DigitalOcean&#8217;s engineering writers put the distinction more bluntly than most vendors will. For a <strong>dense model</strong>, they note, cost scales with memory and is linear and predictable. For a mixture-of-experts model, cost becomes <em>a game of communication</em>. That is the thesis of this issue, stated in five words by someone selling cloud capacity. </p><p>The rest of this report is the long version: what the all-to-all actually costs in bytes and in silicon, why a rack that lists for two to three million dollars is best understood as an answer to a networking problem rather than a compute one, and whether the <strong>wide expert parallelism</strong> that everyone is now deploying actually earns its keep.</p><p>There is a seductive counterexample worth disposing of immediately, because it will come up. The <strong>KTransformers project</strong> can run the complete DeepSeek-V3 model on a single low-cost server with one consumer GPU, a machine that costs in the neighborhood of ten thousand dollars, and still produce nearly twenty tokens per second. I</p><blockquote><p><em>f a mixture-of-experts model can run on a ten-thousand-dollar box, how can the routing be expensive? </em></p></blockquote><p>The answer is that the <strong>KTransformers </strong>configuration never pays the all-to-all toll, because there is no all-to-all. With every expert resident in the memory of a single node, routing a token to an expert is a memory lookup, not a network transfer. </p><p>The economics that follow in this issue are the economics of scale, of serving thousands of concurrent users at frontier latency, and the moment you cross the <strong>boundary from one node</strong> to many, the toll switches on. The single-box demo is real, and it is exactly why the multi-box reality is so often misunderstood.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!lsqO!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F381212c9-40f7-4d9d-859e-fd3b4659c6bc_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!lsqO!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F381212c9-40f7-4d9d-859e-fd3b4659c6bc_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!lsqO!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F381212c9-40f7-4d9d-859e-fd3b4659c6bc_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!lsqO!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F381212c9-40f7-4d9d-859e-fd3b4659c6bc_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!lsqO!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F381212c9-40f7-4d9d-859e-fd3b4659c6bc_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!lsqO!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F381212c9-40f7-4d9d-859e-fd3b4659c6bc_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/381212c9-40f7-4d9d-859e-fd3b4659c6bc_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:97417,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/199726817?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F381212c9-40f7-4d9d-859e-fd3b4659c6bc_1634x893.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!lsqO!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F381212c9-40f7-4d9d-859e-fd3b4659c6bc_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!lsqO!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F381212c9-40f7-4d9d-859e-fd3b4659c6bc_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!lsqO!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F381212c9-40f7-4d9d-859e-fd3b4659c6bc_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!lsqO!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F381212c9-40f7-4d9d-859e-fd3b4659c6bc_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div><hr></div><h2>What sparsity actually buys, and what it borrows</h2><p>Begin with the structure of the model itself, because the geometry of the dispatch is dictated by it. <strong>DeepSeek-V3 carries 256 routed experts</strong> per mixture-of-experts layer plus one shared expert that processes every token. A gating network selects the top eight routed experts for each token.</p><p> The model has <strong>61 transformer layers,</strong> 58 of which are mixture-of-experts layers; the first few are dense. So for the overwhelming majority of the depth of the network, every single token triggers a routing decision, a dispatch to eight destinations, and a combine back.</p><p>The routing itself is not free, though it is cheap relative to the transfer.<strong> A gating network scores all 256 experts </strong>for every token and selects the top eight, and that scoring, the sorting, and the construction of the dispatch order add a small compute and synchronization cost before any data moves. </p><p>It is<strong> minor against the 168 kilobytes</strong> that follow, but it is one more thing the dense model never does, and at the token rates of a frontier deployment even minor per-token costs accumulate. The gating is also where the load imbalance originates, since it is the gate&#8217;s learned preferences that<strong> send too many tokens to too few experts</strong>, which makes it both the cheapest and the most consequential of the operations the router performs.</p><p>The sparsity is genuine and the savings are genuine. Switch Transformer, the architecture that popularized the modern top-k mixture, demonstrated <strong>roughly a sevenfold speedup</strong> over a dense model of equivalent quality, and that ratio has only widened as expert counts have grown. </p><p>When practitioners say mixture-of-experts reduces computation by ninety percent, they are describing the activated-parameter ratio, and they are not wrong about it. A<strong> token that touches 37 billion of 671 billion parameters</strong> is doing far less matrix multiplication than a token that touches all of them.</p><p>But the activation has to move. In a single-accelerator world, the experts a token needs are <strong>sitting in local memory</strong> and the only thing that travels is a memory read. In a serving deployment large enough to matter, the experts are spread across the accelerators by expert parallelism precisely so that each accelerator holds only a few of them and the aggregate weight footprint fits. </p><p>That is the entire point of <strong>expert parallelism</strong>: it is what lets you serve a model whose experts, summed, are far too large for any single device. And it is also what guarantees that the token and its eight chosen experts will, in general, live on different devices.</p><p>So the model does <strong>two collective</strong> operations per mixture-of-experts layer that a dense model never does. The dispatch, sometimes called the scatter, sends<strong> each token&#8217;s hidden activation</strong> to the devices holding its selected experts. The combine, the gather, takes the eight expert outputs and reduces them back into a single vector on the token&#8217;s home device. </p><p>Survey work on efficient inference serving is consistent on the consequence: this<strong> all-to-all exchange of token dispatch</strong> and output gathering is the bottleneck in large-scale mixture-of-experts inference. Not the expert math. The exchange around it.</p><p>This is the borrowing that the sparsity bargain does not advertise. You spend<strong> less on arithmetic and you take on a debt</strong> denominated in bandwidth, and the debt comes due on a part of the machine that has improved far more slowly than the compute has. </p><p>To see why that matters, you have to look at the wire.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>A rack-scale answer to a network problem</h2><p>The defining fact about modern accelerators is that their arithmetic has outrun their interconnect, and it has done so by a margin that is hard to overstate until you put the numbers on a single axis. A Blackwell B200 reads from <strong>its own high-bandwidth memory</strong> at roughly eight terabytes per second. </p><p>The <strong>fifth-generation NVLink fabric</strong> that connects it to its neighbors moves 1.8 terabytes per second per GPU, eighteen links at a hundred gigabytes per second each. That is the fast path between two accelerators, and it is already more than four times slower than the path to local memory. </p><p>Then you fall off the edge. A single four-hundred-gigabit InfiniBand network card, the scale-out path that connects one node to another, moves <strong>about fifty gigabytes per second. </strong>The cliff from on-package memory to the cross-node network is more than two orders of magnitude.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!KbLz!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a4fe2a3-24a7-436b-a1cd-cd49dbd9626b_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!KbLz!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a4fe2a3-24a7-436b-a1cd-cd49dbd9626b_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!KbLz!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a4fe2a3-24a7-436b-a1cd-cd49dbd9626b_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!KbLz!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a4fe2a3-24a7-436b-a1cd-cd49dbd9626b_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!KbLz!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a4fe2a3-24a7-436b-a1cd-cd49dbd9626b_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!KbLz!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a4fe2a3-24a7-436b-a1cd-cd49dbd9626b_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/5a4fe2a3-24a7-436b-a1cd-cd49dbd9626b_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Horizontal bar chart on a log scale comparing effective bandwidth: HBM3e on-package at 8000 GB/s, NVLink 5 fabric at 1800, DeepEP all-to-all over NVLink measured at 730, DeepEP all-to-all over RDMA internode at 90, and a single InfiniBand NIC at 50.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Horizontal bar chart on a log scale comparing effective bandwidth: HBM3e on-package at 8000 GB/s, NVLink 5 fabric at 1800, DeepEP all-to-all over NVLink measured at 730, DeepEP all-to-all over RDMA internode at 90, and a single InfiniBand NIC at 50." title="Horizontal bar chart on a log scale comparing effective bandwidth: HBM3e on-package at 8000 GB/s, NVLink 5 fabric at 1800, DeepEP all-to-all over NVLink measured at 730, DeepEP all-to-all over RDMA internode at 90, and a single InfiniBand NIC at 50." srcset="https://substackcdn.com/image/fetch/$s_!KbLz!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a4fe2a3-24a7-436b-a1cd-cd49dbd9626b_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!KbLz!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a4fe2a3-24a7-436b-a1cd-cd49dbd9626b_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!KbLz!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a4fe2a3-24a7-436b-a1cd-cd49dbd9626b_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!KbLz!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a4fe2a3-24a7-436b-a1cd-cd49dbd9626b_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong><span>FIG 1</span></strong> The bandwidth available to a token collapses as it travels outward from the chip. The all-to-all of a mixture-of-experts layer lives somewhere on the right half of this chart, and where exactly is the whole question.</figcaption></figure></div><p>This is the chart that explains the rest of the hardware industry&#8217;s behavior. If the <strong>all-to-all of a mixture-of-experts layer</strong> can be kept inside the NVLink domain, it runs at hundreds of gigabytes per second. </p><p>If it has to cross the InfiniBand fabric between racks, it runs at a fraction of that. Introl&#8217;s infrastructure analysis puts the ratio at roughly eighteen to one between scale-up bandwidth <strong>inside the NVLink domain</strong> and scale-out bandwidth between racks. For an architecture whose dominant cost is an all-to-all, that ratio is not a detail. It is the design center.</p><p>Which is what the <strong>GB200 NVL72 is.</strong> NVIDIA&#8217;s rack-scale system connects 72 Blackwell GPUs and 36 Grace CPUs into a single NVLink domain delivering 130 terabytes per second of aggregate, non-blocking, all-to-all bandwidth, with <strong>13.5 terabytes of high-bandwidth memory</strong> addressable as one pool. Before this system, the largest NVLink domain you could buy was eight GPUs on a single baseboard. </p><p>The NVL72 takes the fast interconnect and stretches it across an entire rack so that 72 accelerators can talk to each other as though they were neighbors on the same board. <strong>NVIDIA&#8217;s own materials describe the result as a single massive GPU</strong>, and for the purposes of a mixture-of-experts all-to-all, that marketing is closer to literally true than marketing usually is.</p><p>The price of that rack is two to three million dollars, it draws around a hundred and twenty kilowatts, and it is liquid-cooled because there is no other way to remove the heat. It is easy to read those numbers as a statement about <strong>compute density</strong>, and the 1.44 exaflops of four-bit tensor performance per rack invites that reading. </p><p>But the compute was never the scarce thing. You can buy <strong>720 petaflops of eight-bit compute</strong> in roughly 182 H100 accelerators for less money than an NVL72 costs. </p><p>What you cannot buy that way is a 72-way all-to-all domain. The premium on the rack is, in substantial part, the premium on the wire. It is the cost of <strong>not having to cross InfiniBand</strong> for the operation that a mixture-of-experts model performs 58 times per token.</p><blockquote><p><em><span>The premium on the rack is, in substantial part, the premium on the wire. It is the cost of not crossing InfiniBand for the operation a mixture-of-experts model performs fifty-eight times per token.</span></em></p></blockquote><p>Read this way, a great deal of the 2026 accelerator roadmap resolves into a single sentence: make the all-to-all domain larger than the problem. </p><p>At CES 2026 NVIDIA disclosed that the next-generation <strong>Vera Rubin NVL72 will roughly double per-GPU NVLink bandwidth</strong> to 3.6 terabytes per second and lift aggregate all-to-all bandwidth to 260 terabytes per second, with the explicit justification, in NVIDIA&#8217;s own words, that this is the bandwidth needed for the all-to-all communications of leading mixture-of-experts architectures. </p><p>The company has <strong>stopped being coy about it. </strong>The interconnect generation is being sold, by name, as the answer to the problem that the model generation created.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Twenty streaming multiprocessors, give or take</h2><p>Hardware sets the ceiling. Whether you reach it is a question of kernels, and the reference implementation for the <strong>mixture-of-experts all-to-all is DeepEP</strong>, the communication library DeepSeek open-sourced during its 2025 release week. </p><p>DeepEP is worth studying closely not because it is the only such library but because it is the one whose <strong>measured numbers are public</strong>, and those numbers are the closest thing the field has to a ground truth for what the all-to-all costs at the kernel level.</p><p>DeepEP provides two classes of kernel, and the split maps precisely onto the two phases of inference. The normal kernels are tuned for throughput and serve training and the <strong>prefill phase</strong>, where batches are large and the all-to-all moves a great deal of data at once. </p><p>The <strong>low-latency kernels</strong> are tuned for the decode phase, where each step generates one token per sequence, the batches are tiny, and what matters is not bandwidth but the round-trip time of the dispatch and combine. </p><p>This is the <strong>same prefill-versus-decode division</strong> that disaggregated serving exploits, examined one issue ago, now visible at the level of individual communication kernels.</p><p>The measured bandwidths tell the scale-up story in a single table. On Blackwell-class hardware, DeepEP&#8217;s dispatch kernel moves 726 gigabytes per second and<strong> its combine kernel 740 gigabytes per second</strong> when the experts are inside the NVLink domain. The same kernels, forced across the internode RDMA fabric on the same generation of hardware, move about 90 gigabytes per second each. </p><p>That is the<strong> eighteen-to-one ratio of Figure 1,</strong> reproduced at the kernel level on a real workload: the configuration DeepSeek published, with eight thousand tokens per batch, a <em>hidden dimension of 7168</em>, top-eight routing, eight-bit dispatch, and sixteen-bit combine.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!aJDM!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa35c0b14-487f-47ac-9079-b9349b576f29_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!aJDM!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa35c0b14-487f-47ac-9079-b9349b576f29_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!aJDM!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa35c0b14-487f-47ac-9079-b9349b576f29_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!aJDM!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa35c0b14-487f-47ac-9079-b9349b576f29_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!aJDM!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa35c0b14-487f-47ac-9079-b9349b576f29_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!aJDM!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa35c0b14-487f-47ac-9079-b9349b576f29_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a35c0b14-487f-47ac-9079-b9349b576f29_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Grouped bar chart comparing DeepEP dispatch and combine kernel bandwidth: NVLink intranode at 726 and 740 GB/s versus RDMA internode at 90 GB/s each, roughly eight times slower off the NVLink domain.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Grouped bar chart comparing DeepEP dispatch and combine kernel bandwidth: NVLink intranode at 726 and 740 GB/s versus RDMA internode at 90 GB/s each, roughly eight times slower off the NVLink domain." title="Grouped bar chart comparing DeepEP dispatch and combine kernel bandwidth: NVLink intranode at 726 and 740 GB/s versus RDMA internode at 90 GB/s each, roughly eight times slower off the NVLink domain." srcset="https://substackcdn.com/image/fetch/$s_!aJDM!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa35c0b14-487f-47ac-9079-b9349b576f29_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!aJDM!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa35c0b14-487f-47ac-9079-b9349b576f29_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!aJDM!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa35c0b14-487f-47ac-9079-b9349b576f29_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!aJDM!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa35c0b14-487f-47ac-9079-b9349b576f29_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong><span>FIG 2</span></strong> The same kernel, the same hardware, the same workload. The only variable is whether the experts sit inside the NVLink domain or across the network. That one boundary costs roughly a factor of eight.</figcaption></figure></div><p>The second thing DeepEP reveals is subtler and, for the economics, more important. <strong>Moving data costs compute.</strong> The all-to-all kernels do not run on dedicated networking silicon; they run on the same streaming multiprocessors that would otherwise be doing matrix multiplication. </p><p>Every SM assigned to push bytes through the fabric is an SM not computing an expert. <strong>DeepEP&#8217;s first version spent around 24 SMs</strong> on the communication for a training-scale all-to-all. </p><p>Its <strong>second version</strong>, a substantial rewrite that moved from a custom backend to a more lightweight one built on NVIDIA&#8217;s NCCL, cut that to <strong>between four and six SMs</strong> for the same work while matching or exceeding the old bandwidth. </p><p>The library&#8217;s authors describe the V2 rewrite as achieving extreme performance with <strong>several times fewer SM resources,</strong> and the measured table backs the claim: up to 1.3 times the peak bandwidth at up to four times fewer SMs.</p><p>But the decode path is hungrier than the training path, and here the numbers sharpen into a real cost. To hit maximum throughput on the NVLink decode all-to-all, DeepEP&#8217;s table shows the kernel consuming 64 streaming multiprocessors<strong>. A B200 has 148 of them.</strong> That is forty-three percent of the entire accelerator spent moving data rather than computing, in the configuration tuned for speed. </p><p>You can run the same kernel in a low-SM mode that uses 24, but you give up bandwidth to do it. The library also offers genuinely zero-SM paths for pipeline parallelism, <strong>context parallelism</strong>, and certain RDMA transfers, by offloading the movement to copy engines and the network cards directly, and a great deal of the <strong>engineering frontier in 2026 </strong>is about pushing more of the all-to-all onto those zero-SM paths. </p><p>The reason that frontier exists is <strong>Figure 3.</strong></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!jiut!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4713276c-ab2d-4df2-8213-757d8836a8bc_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!jiut!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4713276c-ab2d-4df2-8213-757d8836a8bc_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!jiut!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4713276c-ab2d-4df2-8213-757d8836a8bc_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!jiut!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4713276c-ab2d-4df2-8213-757d8836a8bc_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!jiut!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4713276c-ab2d-4df2-8213-757d8836a8bc_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!jiut!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4713276c-ab2d-4df2-8213-757d8836a8bc_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/4713276c-ab2d-4df2-8213-757d8836a8bc_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Bar chart of streaming multiprocessors consumed by communication kernels: DeepEP V1 training at 24 SMs, V2 training at 5, NVLink decode min-SM mode at 24, NVLink decode max-throughput mode at 64, against a reference line of 148 total SMs on a B200.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Bar chart of streaming multiprocessors consumed by communication kernels: DeepEP V1 training at 24 SMs, V2 training at 5, NVLink decode min-SM mode at 24, NVLink decode max-throughput mode at 64, against a reference line of 148 total SMs on a B200." title="Bar chart of streaming multiprocessors consumed by communication kernels: DeepEP V1 training at 24 SMs, V2 training at 5, NVLink decode min-SM mode at 24, NVLink decode max-throughput mode at 64, against a reference line of 148 total SMs on a B200." srcset="https://substackcdn.com/image/fetch/$s_!jiut!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4713276c-ab2d-4df2-8213-757d8836a8bc_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!jiut!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4713276c-ab2d-4df2-8213-757d8836a8bc_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!jiut!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4713276c-ab2d-4df2-8213-757d8836a8bc_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!jiut!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4713276c-ab2d-4df2-8213-757d8836a8bc_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong><span>FIG 3</span></strong> The all-to-all is not free even when the wire is fast, because it is paid for in the same currency as the math. At peak decode throughput the communication kernel claims forty-three percent of the GPU.</figcaption></figure></div><p>There is real craft in how DeepEP hides this. The library exposes a hook-based mechanism for overlapping communication with computation: a dispatch is launched,<strong> independent work proceeds on the compute stream </strong>while the data is in flight, and only when the result is needed does the kernel wait. </p><p>Done well, the all-to-all latency disappears behind the expert computation and the SM cost is the only thing left to account for. Done badly, the all-to-all stalls the pipeline and the expensive accelerators sit idle waiting for the network. The <strong>difference between those two outcomes</strong> is most of the difference between a good mixture-of-experts deployment and a wasteful one, and none of it is visible in a<strong> FLOP count.</strong></p><p>The second version pushes the SM problem harder by moving the data movement off the streaming multiprocessors entirely wherever it can. Its experimental branches<strong> expose zero-SM paths</strong> for pipeline and context parallelism, handing the transfers to the GPU&#8217;s copy engines, and a <strong>zero-SM remote-memory primitive</strong> the authors call Engram that lets one device reach into another&#8217;s memory over RDMA without spending a single SM on the transfer. </p><p>The motivation is exactly <strong>Figure 3: </strong>every SM the network gives back is an SM the experts can use. The rewrite also abandoned the custom communication backend for a <strong>lighter one built on NVIDIA&#8217;s NCCL</strong>, which let it reuse existing communicators and scale the expert-parallel domain to as many as two thousand devices, far past anything a production model currently needs. </p><p>A separate branch rebuilds the kernels around the tensor-memory-accelerator instructions on<strong> Hopper and Blackwell</strong>, shrinking SM usage again and adding native four-bit support, which is how the dispatch leg of the toll gets cheaper at the same moment the experts do.</p><p>None of this would matter if the expert computation itself were not reorganized to match. The all-to-all delivers a <strong>variable number of tokens to each expert</strong>, because the gating network does not distribute traffic evenly, and a standard batched matrix multiply assumes a fixed shape. </p><p>The answer, embodied in DeepSeek&#8217;s companion <strong>DeepGEMM library</strong>, is a grouped matrix multiply that processes each expert&#8217;s variable token count as a contiguous segment, so the expert math runs as one efficient kernel rather than a ragged collection of small ones. </p><p>The <strong>communication and the computation are co-designed</strong>: the all-to-all produces exactly the memory layout the grouped GEMM wants to consume. </p><p>Pull either apart from the other and the efficiency collapses, which is part of why a<strong> tuned mixture-of-experts stack</strong> is so much harder to assemble than the FLOP savings would suggest.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>What every token pays at the door</h2><p>The bandwidth numbers describe the pipe. The next question is how much the model tries to push through it, and that can be computed directly from the geometry, which makes it <strong>one of the few places in this analysis </strong>where the arithmetic is exact rather than measured.</p><p>Take <strong>DeepSeek-V3&#8217;s mixture-of-experts layer.</strong> Each token&#8217;s hidden activation is a vector of 7168 values. On the dispatch, those values are sent in eight-bit precision, so one byte each, and they are sent to each of the eight selected experts. That is <strong>7168 times 8, roughly 56 kilobytes of dispatch traffic</strong> per token per layer. </p><p>On the combine, each of the eight experts returns an output vector of the same width, but the combine is performed in sixteen-bit precision to preserve the accuracy of the reduction, so two bytes each. That is <strong>7168 times 2 times 8</strong>, roughly <strong>112 kilobytes of combine traffic per token</strong> per layer. </p><p>Add them and a single token, passing through a single mixture-of-experts layer, generates about 168 kilobytes of all-to-all traffic.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!2NvU!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8fe61e25-78e5-4414-864e-a3860fd9cdc0_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!2NvU!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8fe61e25-78e5-4414-864e-a3860fd9cdc0_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!2NvU!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8fe61e25-78e5-4414-864e-a3860fd9cdc0_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!2NvU!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8fe61e25-78e5-4414-864e-a3860fd9cdc0_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!2NvU!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8fe61e25-78e5-4414-864e-a3860fd9cdc0_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!2NvU!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8fe61e25-78e5-4414-864e-a3860fd9cdc0_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/8fe61e25-78e5-4414-864e-a3860fd9cdc0_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Stacked bar chart comparing bytes moved per token per layer: a dense layer at 14 KB with no routing, versus a mixture-of-experts layer at 168 KB total, split into 56 KB of eight-bit dispatch and 112 KB of sixteen-bit combine.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Stacked bar chart comparing bytes moved per token per layer: a dense layer at 14 KB with no routing, versus a mixture-of-experts layer at 168 KB total, split into 56 KB of eight-bit dispatch and 112 KB of sixteen-bit combine." title="Stacked bar chart comparing bytes moved per token per layer: a dense layer at 14 KB with no routing, versus a mixture-of-experts layer at 168 KB total, split into 56 KB of eight-bit dispatch and 112 KB of sixteen-bit combine." srcset="https://substackcdn.com/image/fetch/$s_!2NvU!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8fe61e25-78e5-4414-864e-a3860fd9cdc0_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!2NvU!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8fe61e25-78e5-4414-864e-a3860fd9cdc0_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!2NvU!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8fe61e25-78e5-4414-864e-a3860fd9cdc0_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!2NvU!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8fe61e25-78e5-4414-864e-a3860fd9cdc0_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong><span>FIG 4</span></strong> A derived figure for DeepSeek-V3 geometry. The routing multiplies per-token data movement by roughly twelve and, unlike a dense layer, all of it has to cross the fabric. The model stacks 58 of these layers.</figcaption></figure></div><p>Set that against what a dense layer moves for the same token, which is essentially nothing across the fabric: the activation stays on the device and the only traffic is the <strong>local memory read of about 14 kilobytes</strong>. The mixture-of-experts layer moves roughly twelve times as much data per token, and the crucial difference is not the multiple but the destination. </p><p>The dense traffic stays on-chip. The mixture-of-experts traffic crosses the network. And<strong> this happens 58 times</strong> as the token descends through the model.</p><p>Two things follow from the structure of those 168 kilobytes. The first is that the combine is twice the dispatch, because the combine runs in higher precision. This is not an arbitrary choice; reducing eight expert outputs in <strong>eight-bit precision degrades quality unacceptably</strong>, so the field has settled on eight-bit dispatch and sixteen-bit combine as the standard, and that asymmetry means the return trip is the more expensive leg. </p><p>Any optimization that can compress the combine, including the four-bit experiments now appearing in DeepEP&#8217;s experimental branches, attacks the larger half of the toll.</p><p>The second is that the toll is paid per token, which means the decode phase, where tokens are generated one at a time, pays it in the worst possible way. In prefill, <strong>thousands of tokens are dispatched together and </strong>the <strong>all-to-all amortizes its latency</strong> across an enormous batch; the kernel runs in its throughput regime and the bandwidth numbers of Figure 2 apply. </p><p>In decode, a single step might dispatch only a handful of tokens per sequence, the batch is tiny, the <strong>bandwidth of the pipe</strong> is irrelevant because the pipe is nearly empty, and what dominates is the fixed round-trip latency of reaching across the fabric and back. </p><p>This is why the low-latency decode kernels exist as a separate class, why they are <strong>willing to burn 64 SMs to shave microseconds</strong>, and why decode is the phase where the mixture-of-experts toll hurts most. </p><p>It is also why the entire industry serves prefill and decode on separately tuned pools of hardware, a point this publication examined at length one issue ago and which the all-to-all only sharpens.</p><p>The decode penalty is worth making concrete, because it is where the toll is most counterintuitive. At a service level of a hundred tokens per second per user, the budget for generating one token is ten milliseconds, and into that budget the model must <strong>fit 58 mixture-of-experts layers</strong>, each with a dispatch and a combine that reach across the fabric. </p><p>Inside the NVLink domain a round trip is measured in microseconds and 58 of them fit with room to spare; across the <strong>InfiniBand fabric</strong> the same round trips, with their higher fixed latency, begin to eat the budget directly. </p><p>That is why the decode all-to-all spends 64 SMs to shave microseconds, and why a decode deployment forced to leave the <strong>NVLink domain for its all-to-all can miss its latency target</strong> even when its aggregate bandwidth looks adequate on paper. In decode, latency is the currency, and the fabric boundary is where it gets spent.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share The Software Frontier&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share The Software Frontier</span></a></p><div><hr></div><h2>The hottest expert sets the clock</h2><p>There is a failure mode hiding inside the all-to-all that the bandwidth numbers do not capture at all, and it is the one that most often separates a deployment hitting its theoretical throughput from one falling well short of it. An <strong>all-to-all is a synchronization barrier. </strong></p><p>The combine cannot complete until every expert has returned its outputs, which means the slowest expert on the most overloaded device sets the pace for the entire operation. If the gating network sends a disproportionate share of tokens to a <strong>handful of popular experts</strong>, the devices holding those experts become stragglers, and every other device in the domain waits on them.</p><p>Expert load is not uniform in practice, and it is not even stable. Certain experts specialize in patterns that appear frequently in real traffic, and the imbalance shifts with the workload. Survey work documents the consequence plainly: <strong>imbalanced token distribution</strong> causes device underutilization, and the whole expensive all-to-all runs at the speed of its hottest path. </p><p>A mixture-of-experts deployment can have perfectly adequate aggregate bandwidth and still bleed throughput because the load is lumpy.</p><p>DeepSeek&#8217;s answer in production is an expert-parallel load balancer that the community has <strong>reproduced under the name EPLB</strong>. The mechanism is to identify the high-load experts from live deployment statistics and replicate them: a hot expert is duplicated onto multiple devices so that the tokens destined for it can be spread, flattening the straggler. This is a direct trade of memory for balance. </p><p>You spend extra capacity holding redundant copies of the popular experts in order to keep the all-to-all from stalling on them. It works, and it is now standard, but it is<strong> another line on the bill that the sparsity story did not mention</strong>, and it interacts with the deployment topology in a way that is worth seeing concretely.</p><p>DeepSeek runs the same model checkpoint as two physically different machines, one for each phase, and the contrast is the clearest illustration in the field of how the all-to-all reshapes a deployment. According to DeepSeek&#8217;s own published inference overview and the <strong>CloudMatrix serving analysis</strong> that reconstructs it, the prefill machine groups four nodes, 32 GPUs, into a single unit running 32-way expert parallelism alongside 32-way data parallelism. </p><p>Across those 32 GPUs the routed experts are distributed nine to a device once the redundant copies of the popular experts are counted, with the shared expert and the <strong>attention mechanism replicated on every one.</strong> The raw figure would be eight; the ninth is the load balancer at work. </p><p>The decode machine expands the same model to 18 nodes, 144 GPUs, running 144-way expert parallelism and <strong>144-way data parallelism</strong>, where each device holds only about two routed experts.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!5OdG!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fba346f04-17d0-4e74-8005-8d1018faff52_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!5OdG!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fba346f04-17d0-4e74-8005-8d1018faff52_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!5OdG!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fba346f04-17d0-4e74-8005-8d1018faff52_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!5OdG!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fba346f04-17d0-4e74-8005-8d1018faff52_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!5OdG!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fba346f04-17d0-4e74-8005-8d1018faff52_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!5OdG!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fba346f04-17d0-4e74-8005-8d1018faff52_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ba346f04-17d0-4e74-8005-8d1018faff52_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Two-panel chart. Left panel: GPUs in one expert-parallel domain, 32 for prefill versus 144 for decode. Right panel: routed experts per GPU, 8 for prefill versus 1.78 for decode, each plus one shared expert.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Two-panel chart. Left panel: GPUs in one expert-parallel domain, 32 for prefill versus 144 for decode. Right panel: routed experts per GPU, 8 for prefill versus 1.78 for decode, each plus one shared expert." title="Two-panel chart. Left panel: GPUs in one expert-parallel domain, 32 for prefill versus 144 for decode. Right panel: routed experts per GPU, 8 for prefill versus 1.78 for decode, each plus one shared expert." srcset="https://substackcdn.com/image/fetch/$s_!5OdG!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fba346f04-17d0-4e74-8005-8d1018faff52_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!5OdG!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fba346f04-17d0-4e74-8005-8d1018faff52_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!5OdG!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fba346f04-17d0-4e74-8005-8d1018faff52_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!5OdG!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fba346f04-17d0-4e74-8005-8d1018faff52_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong><span>FIG 5</span></strong> One checkpoint, two machines. The decode deployment spreads the experts across more than four times as many GPUs, which is partly about latency and partly about leaving room to replicate the hot experts.</figcaption></figure></div><blockquote><p><em>Why spread the same 256 experts across 144 devices for decode when 32 sufficed for prefill?</em> </p><p><strong>Two reasons</strong>, and both come back to the all-to-all. </p></blockquote><ul><li><p>The first is latency: with<strong> fewer experts resident per device,</strong> each device does less work per step and the decode latency target is easier to hit. </p></li><li><p>The second is precisely the <strong>straggler problem</strong>. Spreading thin leaves headroom to replicate the popular experts without overflowing any device&#8217;s memory, so the load balancer has somewhere to put the redundant copies. </p></li></ul><p>The decode machine is wider not because the math demands it but because the communication and the balance do. <strong>The shape of the deployment is dictated by the toll,</strong> not the FLOPs.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>The toolchain the toll demanded</h2><p>The all-to-all did not only reshape the hardware and the kernels. It pulled an entire toolchain into being around itself, and the size of that toolchain is the clearest measure of how far the cost migrated from the math. </p><p>A <strong>2026 mixture-of-experts</strong> serving stack at frontier scale is not a model and a runtime. It is a model, a communication library, a grouped-GEMM library, an expert load balancer, a disaggregation layer, and an overlap scheduler, each of which exists to manage some facet of the routing tax. The FLOP count described <strong>one of those six boxes</strong>.</p><p>Consider the overlap problem at the level of an entire forward pass rather than a single layer. <strong>Hiding the all-to-all behind computation</strong> works within a layer, but the decode phase is so latency-sensitive that the field has gone further and split each batch in two, running the communication of one half against the computation of the other in a continuous pipeline. </p><p><strong>SGLang&#8217;s two-batch overlap </strong>and the analogous schemes in other runtimes exist for one reason: to keep the expensive accelerators busy with expert math while the all-to-all of a different microbatch is in flight. It is the same instinct as the <strong>kernel-level hooks</strong>, lifted to the level of the request scheduler, and it is now a standard part of large-scale deployments rather than an exotic optimization.</p><p>Disaggregation adds a second communication problem on top of the all-to-all. Once prefill and decode run on separate pools of hardware, the <strong>key-value cache </strong>computed during prefill<strong> has to be shipped</strong> to the decode pool before generation can begin, and at frontier scale that transfer is large enough and frequent enough to need its own engine. </p><p>The <strong>Mooncake transfer engine</strong> and the equivalent layers inside vLLM and SGLang exist to move key-value caches across the network efficiently, overlapping the transfer with computation so the handoff does not stall the pipeline. This is a network tax distinct from the all-to-all, and it is the price of the <strong>prefill-decode split</strong> that the all-to-all economics make worthwhile in the first place. </p><p>The <strong>two taxes are siblings</strong>: both are consequences of spreading one model&#8217;s inference across many devices, and both are paid down by the same instinct of overlapping transfer with compute.</p><p>The lesson in the length of that list is that the sparsity bargain did not merely move the cost to the network. It moved the cost to a place where <strong>extracting good performance</strong> requires assembling and tuning half a dozen interacting systems, any one of which, misconfigured, hands the savings back. </p><p>The vLLM and SGLang playbooks both carry warnings to this effect, and <strong>AMD&#8217;s ROCm guide to the vLLM mixture-of-experts options </strong>is blunt that the wrong combination of tensor, data, pipeline, and expert parallelism can duplicate the key-value cache many times over and consume far more memory than expected. </p><p>The <strong>FLOP count said the model got cheaper</strong>. The operations manual says it got more complicated, and the complication is where a large part of the real cost now lives.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Does wide expert parallelism pay for itself?</h2><p>All of this is overhead, and the natural reaction to a catalogue of overhead is to minimize it. If the <strong>all-to-all is the cost</strong>, why not keep the expert-parallel domain small, so the all-to-all stays inside a tight, fast group of devices? </p><p>The answer is that narrowing the domain trades one cost for another, and the trade does not run in the obvious direction. </p><p>Wider expert parallelism, counterintuitively, often <strong>produces more throughput per GPU</strong>, not less, and understanding why is the crux of whether the whole approach earns its keep.</p><p>The mechanism is expert packing. When experts are spread across more devices, each device holds fewer of them, which means <strong>more of each device&#8217;s memory and compute</strong> can be devoted to the batch of tokens currently being processed rather than to holding a large slice of the model. </p><p>Larger effective batches per device improve the arithmetic intensity of the expert matrix multiplications, the kernels run closer to the hardware&#8217;s peak, and the <strong>per-GPU throughput rises</strong>, provided the all-to-all overhead can be kept hidden behind that larger computation. The question is always whether the communication grows faster than the packing benefit, and up to a point, on the right interconnect, it does not.</p><p><strong>NVIDIA&#8217;s measurements on the GB200 NVL72 quantify the dividend directly</strong>. Moving from an eight-way expert-parallel configuration to a 32-way one delivers up to 1.8 times the output token throughput per GPU, at a fixed service level of a hundred tokens per second per user, with disaggregated serving and multi-token prediction in both cases. </p><p>Same hardware, same latency target, nearly double the per-GPU output, purely from going wider on expert parallelism.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!CaiQ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10a0a69c-89d2-406f-944f-319e09edd0a7_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!CaiQ!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10a0a69c-89d2-406f-944f-319e09edd0a7_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!CaiQ!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10a0a69c-89d2-406f-944f-319e09edd0a7_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!CaiQ!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10a0a69c-89d2-406f-944f-319e09edd0a7_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!CaiQ!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10a0a69c-89d2-406f-944f-319e09edd0a7_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!CaiQ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10a0a69c-89d2-406f-944f-319e09edd0a7_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/10a0a69c-89d2-406f-944f-319e09edd0a7_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Bar chart of output tokens per second per GPU, normalized to EP8 at 100: EP8 at 100, EP32 at 180, showing 1.8 times the per-GPU throughput from wider expert parallelism.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Bar chart of output tokens per second per GPU, normalized to EP8 at 100: EP8 at 100, EP32 at 180, showing 1.8 times the per-GPU throughput from wider expert parallelism." title="Bar chart of output tokens per second per GPU, normalized to EP8 at 100: EP8 at 100, EP32 at 180, showing 1.8 times the per-GPU throughput from wider expert parallelism." srcset="https://substackcdn.com/image/fetch/$s_!CaiQ!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10a0a69c-89d2-406f-944f-319e09edd0a7_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!CaiQ!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10a0a69c-89d2-406f-944f-319e09edd0a7_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!CaiQ!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10a0a69c-89d2-406f-944f-319e09edd0a7_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!CaiQ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10a0a69c-89d2-406f-944f-319e09edd0a7_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong><span>FIG 6</span></strong> NVIDIA&#8217;s Wide-EP figures on the NVL72. Going wider improves per-GPU throughput, because the packing benefit outweighs the added all-to-all, as long as the all-to-all stays inside the NVLink domain.</figcaption></figure></div><p>The decisive qualifier is the last clause. The 1.8 times holds because the 32-way all-to-all stays inside the NVL72&#8217;s NVLink domain, where Figure 2 says it runs at <strong>726 gigabytes per second. </strong></p><p>The dividend exists because the wire is fast enough that going wider does not push the communication off the cliff. Try the same widening on a cluster where<strong> 32-way expert parallelism forces the all-to-all across InfiniBand</strong>, and the calculus inverts: the packing benefit is swamped by the eightfold bandwidth penalty of leaving the domain, and wider becomes worse. </p><p>This is the same fact from a different angle. The reason the <strong>rack-scale NVLink domain is worth its price </strong>is that it is what makes the wide-EP dividend positive instead of negative.</p><p>There is a second lever working alongside the width, and it appears in <strong>nearly every published wide-EP result</strong>: multi-token prediction. Rather than generating one token per forward pass, the model proposes several and verifies them together, which raises the number of tokens flowing through each all-to-all and pushes the decode kernel out of its worst, smallest-batch regime toward something the bandwidth can amortize. </p><p>Multi-token prediction and wide expert parallelism are complementary for the same underlying reason: <strong>both increase the work done </strong>per round trip across the fabric, and the all-to-all rewards anything that makes its fixed latency a smaller fraction of the whole. </p><p>The dividend in <strong>Figure 6 is partly a multi-token-prediction dividend</strong>, which is why NVIDIA and SGLang report the two together. They are deployed together because they solve the same problem from two directions.</p><p>So the answer to whether wide expert parallelism pays for itself is conditional, and the condition is the interconnect. Inside a sufficiently large fast domain, wider is genuinely better and the measurements prove it. Outside one, wider is a trap. The<strong> crossover sits exactly at the boundary of the NVLink domain</strong>, which is why the size of that domain, 8 GPUs yesterday, 72 today, the same 72 at higher bandwidth tomorrow, is the number that determines how far the dividend extends. </p><p><strong>Expert parallelism</strong> and the interconnect are not two separate decisions. They are one decision, and the hardware vendor has been making half of it for you.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>How much is silicon, and how much is numerics</h2><p>It is tempting to attribute the throughput of a Blackwell mixture-of-experts deployment to the silicon, and the marketing encourages it, but the <strong>public measurements </strong>let us decompose the uplift, and the decomposition is instructive about where the real leverage sits.</p><p>The <strong>LMSYS and SGLang teams</strong> have published a careful progression of DeepSeek serving results on the GB200 NVL72, and the numbers are specific. </p><p>With disaggregated prefill and decode, large-scale expert parallelism, and the conservative numeric configuration of sixteen-bit attention and eight-bit experts, <strong>SGLang reaches 18,471 input tokens per second per GPU</strong> on prefill and 9,087 output tokens per second per GPU on decode, for two-thousand-token sequences. </p><p>Switch to the aggressive configuration, eight-bit attention and four-bit <strong>NVFP4 experts</strong>, and the same system reaches <strong>26,156 input and 13,386 output tokens per second per GPU.</strong> Against the H100 baseline the teams report, those aggressive numbers represent a 3.8 times prefill and 4.8 times decode improvement.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!t5VC!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F58fdb759-c5b7-4999-a918-b5332633c973_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!t5VC!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F58fdb759-c5b7-4999-a918-b5332633c973_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!t5VC!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F58fdb759-c5b7-4999-a918-b5332633c973_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!t5VC!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F58fdb759-c5b7-4999-a918-b5332633c973_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!t5VC!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F58fdb759-c5b7-4999-a918-b5332633c973_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!t5VC!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F58fdb759-c5b7-4999-a918-b5332633c973_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/58fdb759-c5b7-4999-a918-b5332633c973_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Grouped bar chart of tokens per second per GPU for prefill and decode across three configurations: H100 baseline at 6883 prefill and 2789 decode, GB200 with BF16 attention and FP8 experts at 18471 and 9087, and GB200 with FP8 attention and NVFP4 experts at 26156 and 13386, marked as 3.8 times and 4.8 times the baseline.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Grouped bar chart of tokens per second per GPU for prefill and decode across three configurations: H100 baseline at 6883 prefill and 2789 decode, GB200 with BF16 attention and FP8 experts at 18471 and 9087, and GB200 with FP8 attention and NVFP4 experts at 26156 and 13386, marked as 3.8 times and 4.8 times the baseline." title="Grouped bar chart of tokens per second per GPU for prefill and decode across three configurations: H100 baseline at 6883 prefill and 2789 decode, GB200 with BF16 attention and FP8 experts at 18471 and 9087, and GB200 with FP8 attention and NVFP4 experts at 26156 and 13386, marked as 3.8 times and 4.8 times the baseline." srcset="https://substackcdn.com/image/fetch/$s_!t5VC!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F58fdb759-c5b7-4999-a918-b5332633c973_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!t5VC!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F58fdb759-c5b7-4999-a918-b5332633c973_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!t5VC!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F58fdb759-c5b7-4999-a918-b5332633c973_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!t5VC!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F58fdb759-c5b7-4999-a918-b5332633c973_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong><span>FIG 7</span></strong> The Blackwell uplift, decomposed. A large share of the gain over the conservative GB200 configuration comes from dropping the experts to four-bit NVFP4, not from the silicon alone.</figcaption></figure></div><p>The decomposition is the point. The jump from the H100 baseline to the conservative GB200 configuration is the hardware: faster tensor cores, the NVLink domain, more memory bandwidth. But the further jump from the conservative to the aggressive <strong>GB200 configuration, from 18,471 to 26,156 on prefill and from 9,087 to 13,386 on decode</strong>, is numerics. </p><p>It comes from running the experts in four-bit NVFP4 rather than eight-bit. That is a software-and-format change applied to the same rack, and it accounts for a substantial fraction of the total uplift over H100.</p><p>NVFP4 earns its own treatment, and it is a strong candidate for a future issue, but the relevant fact here is why it interacts so favorably with the all-to-all. <strong>Four-bit experts are half the bytes of eight-bit experts</strong>, which directly shrinks the dispatch leg of the toll, and they double the tensor-core throughput of the expert math itself, so the computation that hides the all-to-all gets faster at the same time the all-to-all gets smaller. </p><p>NVIDIA&#8217;s format reportedly holds accuracy within about one percent of the higher-precision baseline on large models through a two-level scaling scheme, and the <strong>accuracy holds up best precisely on the large mixture-of-experts models </strong>where it matters most. The format is, in effect, a second lever on the same toll that the interconnect attacks, and the two compound. </p><p>This is also <strong>why NVIDIA can credibly claim a fivefold reduction in cost per token</strong> from software optimization alone in the two months after Blackwell&#8217;s launch, with no hardware change: a large part of that was kernel and format work on exactly these operations.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>One dollar, or twenty cents</h2><p>The throughput numbers are engineering. The reason they matter is that they convert, almost directly, into the <strong>only number a serving operator actually cares about</strong>, which is dollars per million tokens. And here the all-to-all moves from being a technical concern to being the dominant line item in the unit economics.</p><p>The cleanest demonstration in the public record is the <strong>LMSYS deployment of DeepSeek on 96 H100 GPUs</strong>, twelve nodes of eight, using prefill-decode disaggregation and large-scale expert parallelism with the full DeepEP, DeepGEMM, and EPLB stack. </p><p>That deployment reached 52,300 input tokens per second and 22,300 output tokens per second per node, and when the team translated the throughput into cost, it came to twenty cents per million output tokens. That figure is <strong>roughly one-fifth of what DeepSeek&#8217;s own public API charged at the time</strong>, achieved on rented hardware by an outside team reproducing the architecture.</p><p>The comparison that matters most, though, is the one against the naive alternative on identical hardware. The same report states that the optimized expert-parallel strategy improved output throughput by up to five times over <strong>vanilla tensor parallelism</strong> using the same resources. Five times the throughput on the same GPUs is five times lower cost per token. </p><p>The all-to-all engineering, getting the dispatch and combine to run efficiently inside the fast domain, hiding the latency behind computation, balancing the hot experts, is the entire difference between a deployment at twenty cents and a deployment at a dollar.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!OqS9!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe045fa0-2808-4ee2-8e81-4476ee597b4c_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!OqS9!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe045fa0-2808-4ee2-8e81-4476ee597b4c_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!OqS9!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe045fa0-2808-4ee2-8e81-4476ee597b4c_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!OqS9!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe045fa0-2808-4ee2-8e81-4476ee597b4c_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!OqS9!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe045fa0-2808-4ee2-8e81-4476ee597b4c_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!OqS9!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe045fa0-2808-4ee2-8e81-4476ee597b4c_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/fe045fa0-2808-4ee2-8e81-4476ee597b4c_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Horizontal bar chart of US dollars per million output tokens: vanilla tensor parallel on 96 H100 at one dollar, official DeepSeek API as a reference at one dollar, and PD plus large-scale expert parallelism self-hosted on 96 H100 at twenty cents, five times cheaper on identical hardware.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Horizontal bar chart of US dollars per million output tokens: vanilla tensor parallel on 96 H100 at one dollar, official DeepSeek API as a reference at one dollar, and PD plus large-scale expert parallelism self-hosted on 96 H100 at twenty cents, five times cheaper on identical hardware." title="Horizontal bar chart of US dollars per million output tokens: vanilla tensor parallel on 96 H100 at one dollar, official DeepSeek API as a reference at one dollar, and PD plus large-scale expert parallelism self-hosted on 96 H100 at twenty cents, five times cheaper on identical hardware." srcset="https://substackcdn.com/image/fetch/$s_!OqS9!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe045fa0-2808-4ee2-8e81-4476ee597b4c_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!OqS9!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe045fa0-2808-4ee2-8e81-4476ee597b4c_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!OqS9!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe045fa0-2808-4ee2-8e81-4476ee597b4c_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!OqS9!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe045fa0-2808-4ee2-8e81-4476ee597b4c_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong><span>FIG 8</span></strong> Same 96 GPUs, two ways of organizing them. The five-fold gap between vanilla tensor parallelism and tuned expert parallelism is, almost entirely, the all-to-all done well versus done naively.</figcaption></figure></div><p>Put that five-fold against the backdrop of where inference pricing has gone, and the stakes of the routing tax become clear. The price of <strong>frontier-class inference has fallen by something close to fifty times in three years</strong>, from around twenty dollars per million tokens for GPT-4-class output in late 2022 to roughly forty cents in early 2026. </p><p>Public trackers attribute the collapse to four compounding forces, and mixture-of-experts together with expert parallelism is explicitly one of them, alongside hardware efficiency, kernel and compiler optimization, and low-precision formats.<strong> Inference now consumes roughly two-thirds of all AI compute</strong>, having crossed over from a minority of it only a couple of years ago. </p><p>In that environment a five-fold cost difference is not a margin to be optimized later. It is the difference between a viable serving business and an unviable one.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!CvQD!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa7a5f99e-3225-4cd6-9684-5d6e2eb8c25c_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!CvQD!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa7a5f99e-3225-4cd6-9684-5d6e2eb8c25c_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!CvQD!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa7a5f99e-3225-4cd6-9684-5d6e2eb8c25c_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!CvQD!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa7a5f99e-3225-4cd6-9684-5d6e2eb8c25c_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!CvQD!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa7a5f99e-3225-4cd6-9684-5d6e2eb8c25c_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!CvQD!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa7a5f99e-3225-4cd6-9684-5d6e2eb8c25c_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a7a5f99e-3225-4cd6-9684-5d6e2eb8c25c_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Line chart on a log scale of US dollars per million tokens for GPT-4-class output from 2022 to 2026: 20 dollars in late 2022, 5 in 2023, 2 in 2024, 0.8 in 2025, and 0.4 in early 2026, about a fifty-fold decline, with drivers listed as hardware, kernels, mixture-of-experts plus expert parallelism, and four-bit formats.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Line chart on a log scale of US dollars per million tokens for GPT-4-class output from 2022 to 2026: 20 dollars in late 2022, 5 in 2023, 2 in 2024, 0.8 in 2025, and 0.4 in early 2026, about a fifty-fold decline, with drivers listed as hardware, kernels, mixture-of-experts plus expert parallelism, and four-bit formats." title="Line chart on a log scale of US dollars per million tokens for GPT-4-class output from 2022 to 2026: 20 dollars in late 2022, 5 in 2023, 2 in 2024, 0.8 in 2025, and 0.4 in early 2026, about a fifty-fold decline, with drivers listed as hardware, kernels, mixture-of-experts plus expert parallelism, and four-bit formats." srcset="https://substackcdn.com/image/fetch/$s_!CvQD!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa7a5f99e-3225-4cd6-9684-5d6e2eb8c25c_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!CvQD!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa7a5f99e-3225-4cd6-9684-5d6e2eb8c25c_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!CvQD!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa7a5f99e-3225-4cd6-9684-5d6e2eb8c25c_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!CvQD!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa7a5f99e-3225-4cd6-9684-5d6e2eb8c25c_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong><span>FIG 9</span></strong> The price floor that makes a routing tax of cents per token worth a flagship. Expert parallelism is one of the four named drivers of this curve, not a footnote to it.</figcaption></figure></div><div><hr></div><h2>Making the domain bigger than the problem</h2><p>Step back from the individual numbers and a single strategic motion organizes all of them. The <strong>mixture-of-experts architecture</strong> created a communication problem. </p><p>The hardware industry&#8217;s response has been to make the <strong>fast communication domain large enough</strong> to swallow the problem whole, and the trajectory of that response is the most reliable predictor of where serving economics go next.</p><p>DeepSeek&#8217;s own engineers, in their published reflections on the <strong>hardware lessons of training V3</strong>, frame the future in exactly these terms. They call for the convergence of scale-up and scale-out, for precise low-precision compute units, and for innovations in low-latency communication fabrics.</p><p> Read against this issue, that is a wish list written by the people paying the all-to-all toll, addressed to the people who can make the domain bigger. The <strong>scale-up and scale-out convergence</strong> they ask for is<strong> precisely the elimination of the cliff in Figure 1:</strong> a world where crossing from one node to the next does not cost a factor of eight, because the fast domain has grown to encompass both.</p><p>NVIDIA is building toward exactly that, and is increasingly explicit that it is doing so for this reason. The<strong> NVL72 took the NVLink domain from 8 to 72.</strong> The NVLink Switch architecture is specified to reach 576 GPUs in a single non-blocking fabric. The Rubin generation lifts the per-GPU bandwidth again and ties the increase directly, in NVIDIA&#8217;s own framing, to the all-to-all needs of mixture-of-experts models. </p><p>Each step is sold, more openly than the last, as a larger container for the communication problem that sparsity created. The <strong>architecture and the interconnect are co-evolving</strong>, and the direction is set: the domain keeps growing, the cliff keeps receding, and the toll keeps shrinking as a fraction of the work, without ever quite reaching zero.</p><p>The domain cannot grow without limit, and the constraints on how far it can stretch are physical. NVLink at rack scale runs over copper, which is cheap and reliable but <strong>reaches only a couple of meters</strong>; pushing the domain past a single rack toward the <strong>576-GPU fabric the switch silicon</strong> can address means either optical interconnect, with its added cost, power draw, and failure modes, or denser and hotter racks than the current design. </p><p><strong>Power and cooling </strong>are already near the edge of what a standard data center hall delivers per rack, which is why the NVL72 is<strong> liquid-cooled</strong> and why each new generation leans harder on liquid. And the fault domain grows with the fabric, because a larger coherent domain is a larger blast radius for a single failure. </p><p>The trajectory is set toward bigger domains, but each <strong>expansion buys less headroom than the last</strong> against a wall of copper reach, power density, and fault tolerance that the all-to-all cannot argue its way past.</p><p>What this does not resolve is the dependency it creates. An operator who builds a serving business on wide expert parallelism is building on the assumption that the<strong> fast domain will keep growing</strong>, and that assumption ties the economics of the model layer to the roadmap of a single interconnect vendor. </p><p>The <strong>wide-EP dividend is real</strong>, but it is contingent on hardware that one company predominantly supplies, and the contingency is worth naming. The cheapest way to serve a frontier mixture-of-experts model in 2026 runs through a rack that is, for now, effectively sole-sourced. </p><p>That is a strategic fact about the inference market as much as a technical one, and it is the part of the story most likely to matter in the issues to come.</p><blockquote><p><em><span>The cheapest way to serve a frontier mixture-of-experts model in 2026 runs through a rack that is, for now, effectively sole-sourced. That is a strategic fact as much as a technical one.</span></em></p></blockquote><p>The <strong>dependency has not gone unanswered</strong>. An industry that has watched a single vendor&#8217;s interconnect become the determinant of mixture-of-experts economics has begun to organize alternatives. </p><p>The <strong>UALink consortium</strong> and the <strong>Ultra Ethernet </strong>effort are both attempts to build an open scale-up fabric that could host the all-to-all without routing through one company&#8217;s switches, and <strong>AMD&#8217;s serving stack </strong>now carries its own expert-parallel communication path, a port of the DeepEP ideas onto its accelerators. </p><p>None of these has yet demonstrated the rack-scale all-to-all bandwidth of an NVL72 in production, and the gap is real, but the direction of the effort is itself a <strong>measure of how much the all-to-all matters</strong>. An entire alternative-hardware ecosystem is organizing around the single operation that this issue is about.</p><p>There is also a cost that none of the throughput numbers capture, which is reliability. A 144-GPU decode deployment is one coordinated system, and the all-to-all is a <strong>synchronization barrier across all of it</strong>, which means a fault or a slowdown on any single device degrades the whole. </p><p>The larger the expert-parallel domain, the more devices have to stay healthy and in lockstep for the all-to-all to complete on time, and the operational burden of keeping a domain of that size running at frontier latency is substantial. </p><p><strong>DeepSeek&#8217;s own diagnostic tooling</strong> for locating slow ranks in a DeepEP deployment exists because, at this scale, finding the one straggling device in a domain of hundreds is a routine and necessary operation. </p><p>The wide-EP dividend is real, but it is collected by operators who can keep a very large, very tightly coupled machine running, and that capability is a cost the smaller-domain alternatives never have to pay.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>What to actually do</h2><p>The analysis <strong>resolves into a handful of decisions</strong> that an operator faces in practice, and they follow from the structure rather than from any single benchmark.</p><p>The first decision is whether to use expert parallelism at all, and the honest answer is that it depends entirely on whether your all-to-all can be kept inside a fast domain.</p><p>If you are serving a <strong>frontier mixture-of-experts model at scale</strong> and you have access to a rack-scale NVLink domain, wide expert parallelism is the right tool and the measurements say to go as wide as the domain allows, because the packing dividend is positive inside the fast fabric. </p><p>If your all-to-all would have to cross InfiniBand to go wider, stop widening before it does, because the cliff inverts the dividend. The boundary of the NVLink domain is the boundary of the decision.</p><p>The second decision is how to split the phases. Prefill and decode want different all-to-all kernels, <strong>different expert-parallel widths</strong>, and in DeepSeek&#8217;s production case different physical machines entirely. The decode machine should be wider, both to hit latency targets and to leave room for the load balancer to replicate hot experts. </p><p>If you cannot afford to disaggregate, the decode phase is where the toll will hurt, and the low-latency kernels are where to spend your tuning effort. The <strong>vLLM and SGLang playbooks</strong> both warn, correctly, that the wrong parallelism strategy can duplicate key-value caches across the domain and consume many times the memory you expected, so the parallelism decision is not only about the all-to-all but about what else it forces to be replicated.</p><p>The third decision is precision, and it is mostly free throughput if you are on Blackwell.<strong> Four-bit NVFP4 experts shrink the dispatch leg </strong>of the toll and double the expert math throughput at an accuracy cost that, on large models, is small. The aggressive configuration in Figure 7 is not a marginal tuning; it is a large fraction of the total uplift, and it attacks the same toll the interconnect attacks. If <strong>your hardware supports it and your accuracy budget allows it,</strong> it is among the highest-leverage changes available.</p><p>And the fourth decision is whether you need any of this at all. If your workload is single-user or small-scale, the<strong> KTransformers lesson stands</strong>: a mixture-of-experts model on a single node never pays the toll, and the entire apparatus of expert parallelism is overhead you can decline. </p><p>The all-to-all economics in this issue are the economics of serving at frontier scale and <strong>frontier latency.</strong> Below that scale, the right move is to keep the experts local and let the toll switch stay off.</p><p>The deeper lesson is the one the sparsity story obscured for two years. Mixture-of-experts did not make inference cheaper by doing less work. It moved the work from a <strong>place that was easy to scale</strong>, the arithmetic, to a place that was hard, the network, and then the hardware industry spent two product generations and a great deal of money making the network easy to scale too. </p><p>The bargain was always real. It was just never free, and the bill was always going to come due on the wire.<strong> Knowing where it comes due</strong>, and how much, is most of what it takes to serve these models without overpaying. </p><p> The router decides which experts a token needs. The wire decides what that decision costs. For the models that now define the frontier, the <strong>wire is the more expensive </strong>of the two.</p><div><hr></div><h2>What we are confident about, and what we estimated</h2><p><strong>A</strong></p><p><em>NVLink 5 delivers 1.8 TB/s per GPU; the GB200 NVL72 provides 130 TB/s aggregate all-to-all bandwidth across 72 GPUs, with 13.5 TB of unified HBM3e.</em></p><p><em>NVIDIA GB200 NVL72 datasheet; NVIDIA multi-node NVLink tuning guide; Introl and Spheron interconnect analyses.</em></p><p><strong>A</strong></p><p><em>DeepEP measures dispatch and combine at 726 and 740 GB/s inside the NVLink domain on Blackwell, versus about 90 GB/s each across internode RDMA, on the published V3 workload.</em></p><p><em>DeepEP V2 performance table, deepseek-ai/DeepEP repository.</em></p><p><strong>A</strong></p><p><em>DeepEP&#8217;s decode all-to-all consumes up to 64 SMs at peak throughput; the V2 rewrite cut training all-to-all SM use from 24 to between 4 and 6. A B200 has 148 SMs.</em></p><p><em>DeepEP V2 performance table and release notes; Blackwell architecture specifications.</em></p><p><strong>A</strong></p><p><em>SGLang on the GB200 NVL72 reaches 26,156 prefill and 13,386 decode tokens/sec/GPU with eight-bit attention and NVFP4 experts, reported as 3.8x and 4.8x over H100; the conservative configuration reaches 18,471 and 9,087.</em></p><p><em>LMSYS Org, GB200 NVL72 Part II, September 2025.</em></p><p><strong>A</strong></p><p><em>An LMSYS 96-GPU H100 deployment reached 52.3k input and 22.3k output tokens/sec/node and translated to $0.20 per 1M output tokens, about one-fifth the official API price, and up to 5x the throughput of vanilla tensor parallelism on the same hardware.</em></p><p><em>LMSYS Org, large-scale EP on 96 H100, May 2025.</em></p><p><strong>B</strong></p><p><em>Moving from EP8 to EP32 yields up to 1.8x output throughput per GPU at a fixed 100 tok/s/user SLA on the NVL72, with disaggregated serving and multi-token prediction.</em></p><p><em>NVIDIA, Wide Expert Parallelism on NVL72, January 2026.</em></p><p><strong>B</strong></p><p><em>DeepSeek-V3 runs DP32+EP32 across 32 GPUs for prefill (nine routed experts per GPU plus one shared, including one redundant) and DP144+EP144 across 144 GPUs for decode (about two routed experts per GPU plus one shared).</em></p><p><em>DeepSeek Open Source Week inference system overview (Day 6); CloudMatrix serving analysis (arXiv 2506.12708). The V3 technical report describes a different decode configuration (EP320, one expert per GPU).</em></p><p><strong>B</strong></p><p><em>NVFP4 holds accuracy within roughly one percent of the higher-precision baseline on large models via two-level scaling, and accuracy recovery is strongest on the largest dense and MoE models.</em></p><p><em>NVIDIA NVFP4 technical blogs; Red Hat AI NVFP4 evaluation.</em></p><p><strong>C</strong></p><p><em>A DeepSeek-V3 mixture-of-experts layer moves about 56 KB of dispatch (FP8, top-8) and 112 KB of combine (BF16, top-8) per token, roughly 168 KB total, against about 14 KB for a dense layer.</em></p><p><em>Derived from V3 geometry (hidden 7168, top-8, FP8 dispatch, BF16 combine). Excludes the shared expert and any local-rank optimization.</em></p><p><strong>C</strong></p><p><em>The vanilla-tensor-parallel and official-API reference points of roughly $1.00 per 1M output tokens are derived from the LMSYS statements (optimized $0.20 figure at one-fifth of API, and 5x over vanilla TP).</em></p><p><em>Derived from LMSYS 96-GPU report figures.</em></p><p><strong>C</strong></p><p><em>The H100 baseline in Figure 7 (6,883 prefill, 2,789 decode tokens/sec/GPU) is back-calculated from the reported 3.8x and 4.8x speedups, not independently measured.</em></p><p><em>Derived from LMSYS GB200 Part II reported multipliers.</em></p><p><strong>D</strong></p><p><em>Frontier-class inference pricing has fallen roughly fifty-fold from about $20 to about $0.40 per 1M tokens from late 2022 to early 2026, with MoE plus expert parallelism among four named drivers.</em></p><p><em>Public inference price trackers, 2022 to 2026. Order-of-magnitude trend across vendors, not a single price series.</em></p><p><strong>D</strong></p><p><em>Vera Rubin NVL72 is specified for roughly 3.6 TB/s per GPU and 260 TB/s aggregate, framed by NVIDIA as serving MoE all-to-all needs.</em></p><p><em>NVIDIA NVLink product page and CES 2026 disclosures; pre-release specification subject to change.</em></p><p><em>A = primary or measured | B = single strong vendor or operator source | C = derived by us from sourced inputs | D = directional, treat as trend not point estimate.<br>Character scan: this issue contains zero em dashes and zero en dashes, verified programmatically against the rendered text.</em></p><div><hr></div><h2>Bibliography</h2><ol><li><p><span>DeepSeek-AI. DeepEP: an efficient expert-parallel communication library. GitHub repository, 2025. Performance table, V2 release notes, decode and prefill kernel interfaces.github.com/deepseek-ai/DeepEP</span></p></li><li><p><span>DeepSeek-AI. Insights into DeepSeek-V3: Scaling Challenges and Reflections on Hardware for AI Architectures. arXiv 2505.09343, 2025.arxiv.org/abs/2505.09343</span></p></li><li><p><span>LMSYS Org. Deploying DeepSeek with PD Disaggregation and Large-Scale Expert Parallelism on 96 H100 GPUs. May 2025.lmsys.org/blog/2025-05-05-large-scale-ep</span></p></li><li><p><span>LMSYS Org. Deploying DeepSeek on GB200 NVL72 with PD and Large Scale EP, Part I: 2.7x Higher Decoding Throughput. June 2025.lmsys.org/blog/2025-06-16-gb200-part-1</span></p></li><li><p><span>LMSYS Org. Deploying DeepSeek on GB200 NVL72 with PD and Large Scale EP, Part II: 3.8x Prefill, 4.8x Decode Throughput. September 2025.lmsys.org/blog/2025-09-25-gb200-part-2</span></p></li><li><p><span>LMSYS Org. SGLang and NVIDIA Accelerating SemiAnalysis InferenceMAX and GB200 Together. October 2025.lmsys.org/blog/2025-10-14-sa-inference-max</span></p></li><li><p><span>NVIDIA. Scaling Large MoE Models with Wide Expert Parallelism on NVL72 Rack-Scale Systems. NVIDIA Technical Blog, January 2026.developer.nvidia.com/blog</span></p></li><li><p><span>NVIDIA. GB200 NVL72 product page and datasheet. 130 TB/s NVLink domain, 72-GPU rack specifications.nvidia.com/en-us/data-center/gb200-nvl72</span></p></li><li><p><span>NVIDIA. Inside NVIDIA Blackwell Ultra: The Chip Powering the AI Factory Era. NVIDIA Technical Blog, January 2026.developer.nvidia.com/blog</span></p></li><li><p><span>NVIDIA. Introducing NVFP4 for Efficient and Accurate Low-Precision Inference. NVIDIA Technical Blog, 2025.developer.nvidia.com/blog</span></p></li><li><p><span>NVIDIA. The Economic Value of Inference Software Optimization at the Datacenter Level. April 2026. Fivefold cost-per-token reduction via software.perspectives.nvidia.com</span></p></li><li><p><span>NVIDIA. Multi-Node NVLink Systems Tuning Guide and NVLink / NVLink Switch product documentation. Fifth-generation NVLink and NVSwitch specifications.docs.nvidia.com; nvidia.com/en-us/data-center/nvlink</span></p></li><li><p><span>Microsoft. Achieving Optimal Performance for DeepSeek Expert Parallelism (DeepEP) on Azure. Azure HPC Blog, May 2025.techcommunity.microsoft.com</span></p></li><li><p><span>AMD ROCm. The vLLM MoE Playbook: A Practical Guide to TP, DP, PP and Expert Parallelism. November 2025.rocm.blogs.amd.com</span></p></li><li><p><span>Taming the Titans: A Survey of Efficient LLM Inference Serving. arXiv 2504.19720, 2025. All-to-all as the MoE bottleneck; expert load balancing.arxiv.org/abs/2504.19720</span></p></li><li><p><span>Serving Large Language Models on Huawei CloudMatrix384. arXiv 2506.12708, 2025. DeepSeek DP32+EP32 prefill and DP144+EP144 decode topology.arxiv.org/abs/2506.12708</span></p></li><li><p><span>DeepSeek-AI. DeepSeek-V3/R1 Inference System Overview (Open Source Week, Day 6). February 2025. Production prefill EP32 (9 experts/GPU) and decode EP144 (2 experts/GPU) topology.github.com/deepseek-ai/open-infra-index</span></p></li><li><p><span>Introl. NVLink and Scale-Up Networking. 2026. Scale-up versus scale-out bandwidth ratio; NVL72 physical architecture.introl.com/blog</span></p></li><li><p><span>DigitalOcean. The LLM Inference Trilemma: Throughput, Latency, Cost. April 2026. MoE cost as a game of communication.digitalocean.com/blog</span></p></li><li><p><span>GPUnex. AI Inference Economics: The 1,000x Cost Collapse Reshaping GPUs. February 2026. Inference price trend and drivers.gpunex.com/blog</span></p></li><li><p><span>NVIDIA. NVLink and NVLink Switch, Vera Rubin NVL72 and NVLink 6 disclosures. CES 2026. 260 TB/s aggregate, MoE all-to-all framing.nvidia.com/en-us/data-center/nvlink</span></p></li></ol>]]></content:encoded></item><item><title><![CDATA[Decode Is Memory-Bound. Speculation Is the Arbitrage]]></title><description><![CDATA[Speculative decoding is the only inference optimization that turns idle silicon into tokens without changing a single output. Whether that lands on your bill as a discount or a surcharge is not a prop]]></description><link>https://www.thesoftwarefrontier.com/p/decode-is-memory-bound-speculation</link><guid isPermaLink="false">https://www.thesoftwarefrontier.com/p/decode-is-memory-bound-speculation</guid><dc:creator><![CDATA[Lorenzo Bradanini]]></dc:creator><pubDate>Thu, 25 Jun 2026 10:19:56 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!1i2p!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff2f8e194-f28b-4508-aa07-1a6406746102_1122x1122.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!1i2p!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff2f8e194-f28b-4508-aa07-1a6406746102_1122x1122.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!1i2p!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff2f8e194-f28b-4508-aa07-1a6406746102_1122x1122.png 424w, https://substackcdn.com/image/fetch/$s_!1i2p!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff2f8e194-f28b-4508-aa07-1a6406746102_1122x1122.png 848w, https://substackcdn.com/image/fetch/$s_!1i2p!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff2f8e194-f28b-4508-aa07-1a6406746102_1122x1122.png 1272w, https://substackcdn.com/image/fetch/$s_!1i2p!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff2f8e194-f28b-4508-aa07-1a6406746102_1122x1122.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!1i2p!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff2f8e194-f28b-4508-aa07-1a6406746102_1122x1122.png" width="1122" height="1122" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f2f8e194-f28b-4508-aa07-1a6406746102_1122x1122.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1122,&quot;width&quot;:1122,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2300883,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/199710199?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff2f8e194-f28b-4508-aa07-1a6406746102_1122x1122.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!1i2p!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff2f8e194-f28b-4508-aa07-1a6406746102_1122x1122.png 424w, https://substackcdn.com/image/fetch/$s_!1i2p!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff2f8e194-f28b-4508-aa07-1a6406746102_1122x1122.png 848w, https://substackcdn.com/image/fetch/$s_!1i2p!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff2f8e194-f28b-4508-aa07-1a6406746102_1122x1122.png 1272w, https://substackcdn.com/image/fetch/$s_!1i2p!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff2f8e194-f28b-4508-aa07-1a6406746102_1122x1122.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>Introduction</h2><p><strong>Rent a B200 for an hour</strong> and you are paying for roughly four and a half thousand trillion floating-point operations per second. Ask it to generate text from a seventy-billion-parameter model one user at a time, and for most of that hour the tensor cores do almost nothing. </p><p>The reason is not a bug, a bad kernel, or a scheduling failure. It is arithmetic. To <strong>produce a single token</strong>, the machine must read every weight in the model out of high-bandwidth memory, and reading is the slow part. </p><p>The multiply that follows the read is nearly free, and <strong>there is almost nothing to multiply</strong>, because a single decode step touches one token&#8217;s worth of activations against the entire weight matrix. <mark>You are paying for a fleet of trucks and using them to deliver one envelope per trip.</mark></p><p>Put numbers on it. A <strong>seventy-billion-parameter model</strong> in the eight-bit precision typical of modern serving is seventy gigabytes of weights. </p><p>On a B200 with eight terabytes per second of memory bandwidth, sweeping those weights once takes just under nine milliseconds, and that single sweep yields exactly one token for one user. </p><p>The tensor cores that could have executed thousands of trillions of operations in that window execute a few billion. The arithmetic intensity of <strong>single-stream decode,</strong> the ratio of compute performed to bytes moved, sits at roughly one to two floating-point operations per byte. </p><p>The hardware does not break even until that ratio reaches several hundred. The gap between those two numbers is the entire subject of this issue, because <strong>that gap is compute you have already paid</strong> for and are not using.</p><p><mark>Speculative decoding is the one technique in the</mark><strong><mark> </mark></strong><mark>inference toolbox that spends that idle compute on tokens</mark>, and, in its exact formulations, does so without altering the model&#8217;s output by a single logit. Every other lever trades something visible. </p><p><strong>Quantization trades precision</strong>. Pruning trades capacity. Distillation trades a different model entirely. Speculation, done correctly, trades nothing the user can observe; it simply reorganizes when the weight reads happen so that one read can validate several tokens at once. That is <strong>what makes it unusual</strong>, and it is why every major laboratory shipped a version of it over the last eighteen months.</p><p>And yet the operator folklore says to turn it off above a certain batch size, and the operator folklore is correct, as far as it goes. The resolution of that apparent contradiction is the thesis of this piece. </p><p>The value of speculative decoding is not a number you can quote. <mark>It is a position on a plane whose axes are</mark><strong><mark> batch size and context length,</mark></strong><mark> measured against the ridge of the roofline.</mark> In one region it cuts your cost per token roughly in half. </p><p>In the adjacent region it raises your cost per token by a fifth. The technique never changed. The regime did. The job of this issue is to draw the plane, mark the line that divides it, and show<strong> why the workload that came to dominate 2026</strong>, long-form reasoning, walked straight into the half of the plane where speculation pays.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>What speculation actually does, stated precisely</h2><p>The <strong>mechanism is worth stating exactly</strong>, because almost every confusion about the economics traces back to a loose mental model of it. </p><p>A small, cheap model, the draft, proposes a <strong>short run of candidate tokens</strong>, say four or five of them, by generating them autoregressively in the ordinary way. Because the draft is small, those proposals are fast. </p><p>The large model, the target, then performs a single forward pass that scores all of the candidate positions at once. This is the move that matters: <strong>verifying four candidate tokens</strong> costs the target essentially one weight load, the same memory sweep it would have spent producing one token on its own, because the candidates are processed in parallel across the sequence dimension rather than one step at a time.</p><p>The target then walks the candidates left to right and applies a modified rejection-sampling test at each position. </p><p>It keeps the <strong>longest prefix of candidates </strong>that agrees with what it would have sampled itself, discards the first disagreement and everything after it, and emits one additional bonus token drawn from its own distribution at the point of divergence. </p><p>So a step that began with a draft of length K returns the number of accepted candidates, call it n, plus one. <strong>If the draft proposed five tokens and the target accepted three</strong>, the step produced four tokens for the price of one memory sweep. If the target accepted all five, it produced six. If it accepted none, it produced one, the bonus token, and you paid the draft&#8217;s cost for nothing.</p><p>This is the first thing the folklore gets right and the economics must respect: the speedup is governed by <strong>how many tokens the target accepts </strong>per step, and specifically by the <em>accept length</em>, the mean size of that accepted run plus the bonus. </p><p>It is not governed by the raw acceptance rate in isolation, and it is not governed by how clever the draft sounds. A <strong>draft that is right ninety percent of the time </strong>on the next token but falls apart by the third token buys you less than a draft that is right seventy percent of the time but stays coherent for five. </p><p>The lever is the length of the run, because each run, however long, costs exactly one expensive weight load of the target.</p><h3><span>The lossleness property</span></h3><p>For the <strong>rejection-sampling formulations </strong>introduced by <em>Leviathan and colleagues in 2023</em> and independently by <em>Chen and colleagues</em> the same year, the output distribution is provably identical to standard autoregressive sampling from the target. </p><p>The <strong>modified rejection test</strong> is constructed precisely so that the accepted-token statistics match the target&#8217;s own. EAGLE preserves this exactly, as Hugging Face&#8217;s engineering writeup states plainly. The user cannot tell, from the output alone, that speculation was used.</p><p>That property deserves a caveat stated in the same breath, because vendors are <strong>not always careful about it</strong>. The losslessness holds for the exact rejection-sampling rule. </p><p>There are faster variants, relaxed acceptance, typical acceptance, and several aggressive tree-acceptance schemes, that raise the acceptance rate by loosening the test, and these do change the output distribution. They are often worth it. </p><p>But a quoted speedup that came from a relaxed acceptance rule is not the same artifact as a quoted speedup from <strong>exact rejection sampling</strong>, and an honest ledger keeps them in separate columns. When this issue later cites a four-times number, it will say which rule produced it.</p><p>The methods themselves form a clean lineage, and the direction of travel tells you what the field decided mattered. The original formulation used a <em>separate</em> draft model, <strong>a smaller member</strong> of the same family, which is simple but means maintaining and serving two models. </p><p>Medusa removed the second model by attaching several prediction heads to the target itself, each guessing a future position in parallel. EAGLE, in its <strong>first and second versions</strong>, moved the autoregression down a level, drafting in the target&#8217;s own feature space rather than in token space, which made the draft both cheaper and better aligned. </p><p><strong>EAGLE-3, presented at NeurIPS 2025</strong> and described in arXiv:2503.01840, pushed further: it fuses features from early, middle, and late layers of the target, predicts tokens directly rather than through an intermediate feature-regression step, removes a constraint that had limited how much training data helped, and uses a dynamic draft tree that expands the most promising candidates. </p><p>The endpoint of the lineage is to fold the draft into the target entirely, which is <strong>what DeepSeek&#8217;s multi-token prediction does</strong>, and which the next sections will show has economic consequences beyond mere convenience.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>The roofline and the speculation budget</h2><p>The previous issue, <em>The Split and the Seam</em>, derived the roofline for LLM serving in detail and split inference into its prefill and decode phases on exactly these grounds. </p><p>This issue assumes that derivation rather than repeating it, and reuses its house figures. The roofline says that for any kernel there is a ridge point, an <strong>arithmetic intensity above</strong> which you are limited by the chip&#8217;s compute throughput and below which you are limited by its memory bandwidth. </p><p>The ridge is simply peak compute divided by peak bandwidth. For an H100 SXM running FP8, that is one thousand nine hundred and seventy-<strong>nine teraFLOPS of dense tensor throughput </strong>against three and thirty-five hundredths terabytes per second of HBM3, which puts the ridge at five hundred and ninety-one FLOP per byte. </p><p>The H200 keeps the same compute but <strong>raises bandwidth to four and eight tenths terabytes per second</strong>, dropping the ridge to four hundred and twelve. A B200 at roughly four thousand five hundred teraFLOPS against eight terabytes per second sits near five hundred and sixty-two.</p><p>Single-stream decode operates at one to two FLOP per byte. Hold those two numbers next to each other. The operating point is two to nearly three orders of magnitude below the ridge. </p><p>That distance, expressed as a ratio, is the factor by which you could <strong>multiply the compute performed per byte</strong> <strong>moved </strong>before you would hit the memory ceiling and start paying for it in latency. Call it the <em>speculation budget</em>. On a single stream it is somewhere between three hundred and nearly six hundred times. </p><p>It is, very precisely, the ceiling on what any decode-side technique could reclaim from idle compute, and the headroom that speculative decoding draws on.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!9x7-!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb3d8c768-f443-486c-b423-9078bc10d614_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!9x7-!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb3d8c768-f443-486c-b423-9078bc10d614_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!9x7-!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb3d8c768-f443-486c-b423-9078bc10d614_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!9x7-!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb3d8c768-f443-486c-b423-9078bc10d614_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!9x7-!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb3d8c768-f443-486c-b423-9078bc10d614_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!9x7-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb3d8c768-f443-486c-b423-9078bc10d614_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b3d8c768-f443-486c-b423-9078bc10d614_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Roofline chart showing the speculation budget as the vertical gap between the single-stream decode operating point and the compute ridge for H100, H200, and B200.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Roofline chart showing the speculation budget as the vertical gap between the single-stream decode operating point and the compute ridge for H100, H200, and B200." title="Roofline chart showing the speculation budget as the vertical gap between the single-stream decode operating point and the compute ridge for H100, H200, and B200." srcset="https://substackcdn.com/image/fetch/$s_!9x7-!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb3d8c768-f443-486c-b423-9078bc10d614_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!9x7-!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb3d8c768-f443-486c-b423-9078bc10d614_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!9x7-!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb3d8c768-f443-486c-b423-9078bc10d614_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!9x7-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb3d8c768-f443-486c-b423-9078bc10d614_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Figure 1.</strong> The speculation budget is the vertical distance from the decode operating point to the roofline ridge. Single-stream decode runs two to three orders of magnitude below the point where compute becomes the limit, which is the headroom speculation cashes in. Ridge points use Issue 03 house figures: H100 SXM FP8 1,979 TFLOPS / 3.35 TB/s, H200 same compute / 4.8 TB/s, B200 ~4,500 TFLOPS / 8 TB/s.</figcaption></figure></div><p>A reader should immediately ask why, if the budget is several hundred times, speculation delivers only two or three. </p><p>The answer is that no single technique spends the whole budget, and speculation in particular spends only a sliver of it. </p><p><strong>Its yield is capped by accept length</strong>: each verification step still costs one weight load and returns at most the accepted run plus a bonus, which in practice is two to five tokens, so the multiple is bounded there no matter how much idle compute waits unused. </p><p>The draft is not free either, and its own forward passes consume part of the budget before any of it reaches the output. The rest of the headroom is what <em>batching</em> claims, the other and larger way to <strong>raise arithmetic intensity</strong>, and whatever neither mechanism reaches simply sits idle under the latency ceiling. </p><p><mark>So the budget is the size of the prize, not the size of the winnings.</mark> Speculation is the instrument that collects the part of it that batching cannot, which, as the rest of this issue argues, is exactly the part that matters when a<strong> latency SLA forbids batching in the first place</strong>.</p><p>This budget is not an accident of one chip generation. It is the accumulated result of a divergence that has run for a decade. </p><p>Across the <strong>span from V100 to B200</strong>, tensor compute throughput grew by roughly thirty-six times, while HBM bandwidth over the same generations grew by only about nine times, a gap documented in the systems literature and discussed at length in this publication&#8217;s earlier piece on the memory wall. </p><p><strong>Compute outran memory</strong> by a factor of four across those generations, and every factor of that divergence widened the speculation budget, because it pushed the ridge further above the place where decode actually runs. </p><p>The technique gets structurally more attractive with each generation of hardware that <strong>improves compute faster than bandwidth</strong>, which is to say, with each generation.</p><h3><span>The budget is real, and finite</span></h3><p>The roofline guarantees the<strong> headroom exists on a single stream</strong>. It does not guarantee the headroom survives batching, or long context. The next two sections are the story of what spends the budget down, and they reach opposite conclusions depending on which axis you move along.</p><p>The framing to carry forward is that speculation is, mechanically, a<strong> way of converting roofline headroom into tokens</strong>. When the headroom is large, the conversion is cheap and the tokens are nearly free. </p><p>When the headroom has been consumed by something else, there is nothing left to convert, and the draft&#8217;s compute becomes pure overhead. </p><p>Everything downstream is a question about <strong>how much headroom is actually available</strong> in your serving regime, and the surprising part, the part the folklore half-misses, is that the answer depends on two independent variables, not one.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Why batching is said to kill it</h2><p>Here is the story every production guide tells, and it is the right place to start because <strong>it is true within its assumptions.</strong> As you raise the batch size, packing more concurrent sequences into each forward pass, the target&#8217;s arithmetic intensity rises. </p><p>The reason is that <strong>the weights are read once per step</strong> regardless of how many sequences are in the batch, so the cost of that read amortizes across the batch. One sequence pays the full one hundred and forty gigabyte sweep for one token. </p><p><strong>Thirty-two sequences</strong> split the same sweep across thirty-two tokens. The bytes-per-token falls, the FLOP-per-byte rises, and at some batch size the target crosses its ridge and becomes compute-bound. </p><p>Past that crossing, the free headroom is gone, because the compute is now the scarce resource, and the draft model&#8217;s extra forward passes are competing for it against real work.</p><p>The crossing is commonly placed around a batch of thirty-two. <strong>Spheron&#8217;s production guide</strong> from March 2026 and<strong> E2E Networks&#8217; </strong>engineering notes both put the practical break-even in that neighborhood, with the qualification that it moves with model size, quantization, and sequence length. </p><p>Below a draft acceptance of roughly one half, the guides agree, speculation hurts at any batch, because too few candidates survive verification to cover the draft&#8217;s cost. </p><p>The operational rule that falls out is blunt and widely repeated: <strong>disable speculation when batch sizes climb past the low tens</strong>, when outputs are short, when generation is high-entropy, or when you are memory-constrained on weights to the point that the draft displaces batch capacity.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!b_nj!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F62ed2ac3-10a2-40a0-84d1-0ad7180b8940_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!b_nj!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F62ed2ac3-10a2-40a0-84d1-0ad7180b8940_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!b_nj!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F62ed2ac3-10a2-40a0-84d1-0ad7180b8940_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!b_nj!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F62ed2ac3-10a2-40a0-84d1-0ad7180b8940_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!b_nj!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F62ed2ac3-10a2-40a0-84d1-0ad7180b8940_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!b_nj!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F62ed2ac3-10a2-40a0-84d1-0ad7180b8940_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/62ed2ac3-10a2-40a0-84d1-0ad7180b8940_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Chart showing decode speedup from speculation decaying from over 3x at batch 1 toward break-even near batch 32, with EAGLE 3.1 measured points overlaid.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Chart showing decode speedup from speculation decaying from over 3x at batch 1 toward break-even near batch 32, with EAGLE 3.1 measured points overlaid." title="Chart showing decode speedup from speculation decaying from over 3x at batch 1 toward break-even near batch 32, with EAGLE 3.1 measured points overlaid." srcset="https://substackcdn.com/image/fetch/$s_!b_nj!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F62ed2ac3-10a2-40a0-84d1-0ad7180b8940_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!b_nj!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F62ed2ac3-10a2-40a0-84d1-0ad7180b8940_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!b_nj!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F62ed2ac3-10a2-40a0-84d1-0ad7180b8940_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!b_nj!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F62ed2ac3-10a2-40a0-84d1-0ad7180b8940_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Figure 1.</strong> The conventional envelope. As the batch fills, the spare compute that speculation feeds on disappears, and the speedup decays toward break-even near a batch of 32. Measured points are EAGLE 3.1 on Kimi-K2.6-NVFP4, vLLM tensor-parallel 4, GB200, SPEED-Bench, published by the vLLM team in May 2026: <strong>2.03x</strong> at concurrency 1, <strong>1.71x</strong> at 4, <strong>1.66x</strong> at 16. The break-even location and the 0.5-acceptance floor are from E2E Networks and the Spheron production guide. The envelope is illustrative; the points are measured.</figcaption></figure></div><p>The measured points anchor the shape. <strong>EAGLE 3.1, released jointly by the EAGLE, vLLM, and TorchSpec teams </strong>in May 2026 and benchmarked in the vLLM team&#8217;s own writeup running on Kimi-K2.6 in NVFP4 under vLLM with tensor parallelism of four on a GB200, delivered a per-user throughput multiple of two and three hundredths at concurrency one, one and<strong> seventy-one hundredths </strong>at concurrency four, and one and sixty-six hundredths at concurrency sixteen, on the SPEED-Bench suite. </p><p>The curve is unmistakable: the benefit is largest when the machine is emptiest, and it erodes as the batch fills. This is the empirical backbone of the folklore, and nothing in this issue disputes it on its own terms.</p><p>There is a sharper version of the same point that the practitioner Tian Pan has called the critical inversion. At<strong> low concurrency</strong> the draft runs in compute the target was wasting anyway, so it is free. </p><p>At high concurrency the draft&#8217;s forward passes contend with queued real requests for the same saturated compute, so the draft is no longer free; it is actively stealing throughput from work you could otherwise be doing. </p><p>Under that framing, speculation is fundamentally a low-concurrency latency optimization, and treating it as a <strong>throughput optimization at scale </strong>is a category error. This is good guidance. It is also, and this is the whole turn of the issue, an argument that silently assumes short context.</p><p>The <strong>amortization story is entirely about weights.</strong> It says the weight read, which dominates single-stream decode, gets cheaper per token as the batch grows. That is true. But the weight read is not the only thing decode reads from memory on every step, and the other thing it reads does not amortize across the batch at all. </p><p>The conventional wisdom is not wrong. It is two-thirds of a three-variable problem, and the missing variable is the one that 2026&#8217;s workloads turned up to eleven.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>The KV cache re-opens the budget</h2><p>Every token a transformer has already produced leaves behind a key and a value vector in every attention layer, and every future token must read all of them. <strong>That is the KV cache</strong>, and it is the second great memory cost of decode. Crucially, it behaves nothing like the weights. </p><p>The weights are shared across the batch, so their read amortizes. The KV cache is private to each sequence and grows with that sequence&#8217;s length, so its read scales with the <strong>batch size </strong>and with the context length simultaneously. </p><p>Doubling the batch does not split the KV read across more tokens; it doubles the total KV that must be read. Doubling the context length doubles it again.</p><p>The consequence is the result at the center of the <strong>MagicDec work, described in arXiv:2408.11049</strong> and in Together AI&#8217;s analysis of it. There is a critical sequence length, call it S-star, beyond which the per-step KV read dominates the per-step weight read even at large batch. </p><p>Past S-star, decode is memory-bound <em>again</em>, not because the weights are unamortized, but because the KV cache is enormous and unamortizable. The free compute the conventional wisdom said batching had consumed comes back, because batching only consumed the part of the memory bill that the weights were responsible for. The <strong>KV part grew instead of shrinking.</strong></p><p>This changes the geometry of the entire question. The compute-bound region is not the half-plane &#8220;<em>batch greater than thirty-two</em>.&#8221; It is a wedge: compute-bound requires high batch <em>and</em> short context, both at once. Move to small batch and you are <strong>memory-bound on weights</strong>. Move to long context and you are memory-bound on KV. </p><p>Only in the corner where the batch is large and the sequences are short does the target actually saturate its compute. Everywhere else, on a single stream, on long documents, on <strong>extended reasoning traces</strong>, the headroom is open and speculation has something to convert. </p><p>The short-context side of that corner has a hard edge worth naming. Because the KV cache caps how high arithmetic intensity can climb, there is a context length, roughly <strong>a thousand tokens on a B200</strong> and closer to eleven hundred on an H200, past which no batch size reaches the compute-bound ridge at all. </p><p>That ceiling is the dashed wall in the diagram below, and the compute-bound wedge lives entirely to its left. Every reasoning trace sits far to its right.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!eB9V!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97470b63-02ca-454b-add0-20561267e0be_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!eB9V!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97470b63-02ca-454b-add0-20561267e0be_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!eB9V!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97470b63-02ca-454b-add0-20561267e0be_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!eB9V!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97470b63-02ca-454b-add0-20561267e0be_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!eB9V!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97470b63-02ca-454b-add0-20561267e0be_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!eB9V!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97470b63-02ca-454b-add0-20561267e0be_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/97470b63-02ca-454b-add0-20561267e0be_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Phase diagram on axes of sequence length and batch size showing the compute-bound region as a high-batch short-context wedge and the memory-bound region everywhere else, with workload markers.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Phase diagram on axes of sequence length and batch size showing the compute-bound region as a high-batch short-context wedge and the memory-bound region everywhere else, with workload markers." title="Phase diagram on axes of sequence length and batch size showing the compute-bound region as a high-batch short-context wedge and the memory-bound region everywhere else, with workload markers." srcset="https://substackcdn.com/image/fetch/$s_!eB9V!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97470b63-02ca-454b-add0-20561267e0be_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!eB9V!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97470b63-02ca-454b-add0-20561267e0be_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!eB9V!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97470b63-02ca-454b-add0-20561267e0be_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!eB9V!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97470b63-02ca-454b-add0-20561267e0be_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Figure 2.</strong> The two-axis map, drawn from the roofline condition. Decode is compute-bound, where speculation taxes throughput, only in the small wedge at high batch and short context: it needs batch above roughly the ridge over two (around 250 for a B200) and context below the dashed wall. The wall sits where even infinite batch cannot lift arithmetic intensity to the ridge, at a sequence length of about 2C/R, near a thousand tokens on a B200 and eleven hundred on an H200. Everything else is memory-bound, where speculation pays. Interactive chat, single-user reasoning, and the long-document MagicDec regime all sit in speculation-pays territory; only short-prompt high-batch offline serving sits in the tax. With a speculative draft tree the effective batch reaches the wedge nearer nominal batch 32, which is the conventional break-even. House calculation; the KV-versus-weight crossover is a distinct curve, in Figure 4.</figcaption></figure></div><h3>Where the line actually sits</h3><p>The boundary is not abstract; you can locate it with the model&#8217;s own dimensions, and where it lands is the punchline. Take a seventy-billion-parameter model of the Llama-3-70B shape: <strong>eighty layers, grouped-query attention</strong> with eight key-value heads of head-dimension one hundred and twenty-eight. </p><p>The key-value cache that must be read per token is two vectors, key and value, times eight heads, times one hundred and <strong>twenty-eight dimensions</strong>, times eighty layers, which is one hundred and sixty-three thousand eight hundred and forty elements per token. </p><p>In a<strong> sixteen-bit KV cache</strong> that is about three tenths of a megabyte for every token already in the sequence, per sequence. The weights, in an eight-bit serving format, are seventy gigabytes, read once per step and shared across the whole batch.</p><p>The crossover, the point where the <strong>per-step key-value</strong> read equals the per-step weight read, is therefore where batch size times sequence length reaches roughly seventy gigabytes divided by three tenths of a megabyte, which is about two hundred and twenty thousand. </p><p>That locus, batch times sequence held constant, is a hyperbola: it is the line drawn in <strong>Figure 3 below</strong>, and it marks where the KV read overtakes the weight read, which is to say where adding more batch stops reducing the bytes paid per token. </p><p>At a <strong>batch of thirty-two</strong> it puts that amortization crossover near seven thousand tokens; at a batch of sixty-four, near three thousand five hundred; at a batch of one hundred and twenty-eight, near one thousand seven hundred. </p><p>An <strong>eight-bit key-value cache</strong> roughly doubles all of those. This crossover is a finer fact than the compute-bound wall of the previous figure, and the two should not be confused: the wall is the context length past which no batch reaches the ridge, while the crossover is the point at a given batch where batching has stopped buying amortization. </p><p>The reasoning workload clears both at once. Hold it against the MLPerf numbers: a mean output of three thousand eight hundred and eighty tokens, a maximum of twenty thousand, <strong>AIME traces running to twenty-three thousand. </strong></p><p>Those sequences run far past the roughly one-thousand-token compute-bound wall, so no batch reaches the ridge, and at any batch an operator can realistically run under a <strong>latency SLA </strong>they are past the amortization crossover as well. This is not a near miss. </p><p><mark>The workload that came to define 2026 lives deep in the memory-bound region by a wide margin, which is the entire reason speculation pays there.</mark> (Both lines are <strong>house order-of-magnitude calculations</strong> from the stated architecture, graded in the dossier; the precise coefficients move with KV precision, head count, ridge, and serving format, but the order of magnitude, and therefore the conclusion, is robust.)</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!HWH4!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F382010e2-3271-4faa-8148-bbf4387b0154_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!HWH4!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F382010e2-3271-4faa-8148-bbf4387b0154_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!HWH4!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F382010e2-3271-4faa-8148-bbf4387b0154_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!HWH4!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F382010e2-3271-4faa-8148-bbf4387b0154_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!HWH4!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F382010e2-3271-4faa-8148-bbf4387b0154_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!HWH4!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F382010e2-3271-4faa-8148-bbf4387b0154_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/382010e2-3271-4faa-8148-bbf4387b0154_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Log-log chart of per-step memory read versus aggregate tokens in flight, showing a flat weight-read line crossed by rising KV-read lines for dense GQA and MLA, with crossover points marked.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Log-log chart of per-step memory read versus aggregate tokens in flight, showing a flat weight-read line crossed by rising KV-read lines for dense GQA and MLA, with crossover points marked." title="Log-log chart of per-step memory read versus aggregate tokens in flight, showing a flat weight-read line crossed by rising KV-read lines for dense GQA and MLA, with crossover points marked." srcset="https://substackcdn.com/image/fetch/$s_!HWH4!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F382010e2-3271-4faa-8148-bbf4387b0154_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!HWH4!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F382010e2-3271-4faa-8148-bbf4387b0154_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!HWH4!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F382010e2-3271-4faa-8148-bbf4387b0154_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!HWH4!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F382010e2-3271-4faa-8148-bbf4387b0154_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Figure 3.</strong> The crossover, drawn. The weight read is flat because it amortizes across the batch; the KV read rises linearly with aggregate tokens in flight because it does not. For a dense-GQA 70B model the two meet near 220,000 aggregate tokens; a reasoning workload at batch 32 and 8,000 tokens of context already sits past it. Compressed attention moves the crossover, it does not remove it. House calculation; the order of magnitude is the point.</figcaption></figure></div><p>One honest qualification belongs here, because it is the first thing a careful reader will raise. The arithmetic above is for dense grouped-query attention, where the KV cache is large. </p><p>Architectures that compress the cache move the crossover to the right. <strong>DeepSeek&#8217;s Multi-head Latent Attention</strong>, by its hardware paper&#8217;s account, holds the KV cache to roughly seventy kilobytes per token, about a fifth of the dense-GQA figure, which pushes the crossover out toward a million aggregate tokens, the shallower line in the chart above. </p><p><strong>DeepSeek-V4 goes further still</strong>: its model card reports that at a one-million-token context, V4-Pro spends about ten percent of V3.2&#8217;s KV cache and twenty-seven percent of its per-token compute, with V4-Flash at seven percent and ten percent, through a compressed sparse-attention stack. </p><p>This does not rescue the throughput regime. It relocates the line, and it does so precisely in service of making very long contexts affordable, which <strong>keeps sequences long</strong>, which keeps the <strong>budget open</strong>. The compression buys context length, and context length is what holds decode in the memory-bound region. The two facts point the same way.</p><p>MagicDec turns this into a working technique with one additional move: the <strong>draft itself must be light on KV,</strong> not just light on weights, or it reintroduces the very bottleneck it is trying to relieve. </p><p>With a draft that uses a fixed sparse or short-window KV footprint, MagicDec reports up to roughly two times on both throughput and latency together in the<strong> large-batch long-context regime</strong> on eight A100s, a regime where the conventional wisdom predicts speculation should be dead. </p><p>The reported draft-to-target memory ratio for a Llama-3.1-70B target with an <strong>eight-billion-parameter draft</strong> stays near four tenths and, importantly, stays constant as the batch grows, because the draft&#8217;s KV is bounded by design while the target&#8217;s KV grows. </p><p>That constancy is what keeps the draft cheap exactly where the conventional analysis assumed it would become expensive.</p><p>The corrected physics is therefore a single sentence with three clauses. <strong>Small batch is memory-bound </strong>because weights dominate. Long context is memory-bound because the KV read dominates. </p><p>Compute-bound is only the high-batch corner below the thousand-token wall, and that corner is smaller than the folklore implies. The &#8220;<em>disable above batch thirty-two</em>&#8221; rule is not wrong; it is a short-context rule wearing the costume of a general one. </p><p>And the moment your workload develops long sequences, whether from large documents or from long generations,<strong> the rule inverts</strong>, and speculation comes back to life precisely where you had been told to switch it off.</p><div><hr></div><h2>What actually shows on the ledger</h2><p><strong>Physics is not the bill</strong>. To get from the roofline to dollars, you have to know how the operator is allowed to set the batch size, and that is a question about service-level agreements, not about chips. </p><p>There are <strong>two serving regimes</strong>, and they read the same technique with opposite signs.</p><p>In <em>throughput-maximizing</em> service, the operator is free to batch all the way to the compute-bound point, because the <strong>only objective is cost per token</strong> and the way to minimize it is to amortize the weight read across as many sequences as possible. In that regime the machine is, by construction, saturated. </p><p><strong>There is no idle compute</strong>. Speculation adds the draft&#8217;s forward passes to a chip that has nothing spare to run them in, so the cost per token rises. This is the regime the folklore is built for, and in it the folklore&#8217;s advice is exactly right.</p><p>In <em>latency-capped</em> service, the operator may <em>not</em> batch to the compute-bound point, because there is a ceiling on how long each token may take, and raising the batch raises per-token latency. The operator batches only until the <strong>latency SLA binds</strong>, and then stops, often well short of saturation. </p><p>The machine therefore runs with idle compute by design, not by accident, because the SLA forbids filling it. That idle compute is the speculation budget, and <strong>speculation converts it into tokens</strong>, cutting the cost per token. <mark>Same technique, opposite sign, and the only thing that changed was whether a latency ceiling capped the batch.</mark></p><p><mark>These two signs are not asserted; they fall out of a one-line cost model.</mark> Cost per token is the <strong>rental rate of the GPU</strong> divided by the tokens it delivers each second, so anything that multiplies throughput divides cost by the same factor. </p><p>In the latency-capped regime the wasted verification compute is free, because the chip sat idle under the SLA anyway, so throughput scales with the accept length discounted only by the draft&#8217;s own overhead: an accept length of about two and a half against a draft overhead near a fifth gives a <strong>throughput multiple close to two</strong>, which is the cut of roughly half the chart shows. </p><p>In the throughput-maximized regime the chip has no spare compute, so the draft&#8217;s wasted work bites directly. If the draft proposes three tokens and <strong>two and a half clear verification on average</strong>, the target spends three positions of compute to yield two and a half tokens, a throughput multiple near five sixths, which is the cost rise of about a fifth the chart shows. </p><p>The same two numbers, an accept length near two and a half and a draft length near three, generate both bars, and they are the same numbers behind the one-and-a-half to two-and-a-half times production speedups.</p><p>Push the draft length above the accept length and the throughput-regime penalty grows, which is precisely why <strong>over-drafting is the classic way to lose money</strong> on speculation in a saturated cluster.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!T4i-!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F832809e8-ea41-4db9-be51-1a2a5a28cc53_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!T4i-!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F832809e8-ea41-4db9-be51-1a2a5a28cc53_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!T4i-!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F832809e8-ea41-4db9-be51-1a2a5a28cc53_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!T4i-!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F832809e8-ea41-4db9-be51-1a2a5a28cc53_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!T4i-!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F832809e8-ea41-4db9-be51-1a2a5a28cc53_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!T4i-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F832809e8-ea41-4db9-be51-1a2a5a28cc53_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/832809e8-ea41-4db9-be51-1a2a5a28cc53_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Bar chart comparing cost per million tokens for autoregressive versus speculative decoding under latency-capped serving and throughput-maximizing serving.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Bar chart comparing cost per million tokens for autoregressive versus speculative decoding under latency-capped serving and throughput-maximizing serving." title="Bar chart comparing cost per million tokens for autoregressive versus speculative decoding under latency-capped serving and throughput-maximizing serving." srcset="https://substackcdn.com/image/fetch/$s_!T4i-!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F832809e8-ea41-4db9-be51-1a2a5a28cc53_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!T4i-!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F832809e8-ea41-4db9-be51-1a2a5a28cc53_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!T4i-!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F832809e8-ea41-4db9-be51-1a2a5a28cc53_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!T4i-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F832809e8-ea41-4db9-be51-1a2a5a28cc53_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Figure 4.</strong> The sign of the ledger is set by the regime. Under a latency SLA that caps the batch below saturation, speculation converts idle compute into tokens and cuts cost per million output tokens (here roughly <strong>-48%</strong>). In throughput-maximizing service batched to the compute-bound point, there is no idle compute for the draft to use, and the draft&#8217;s overhead raises cost (here roughly <strong>+19%</strong>). Both magnitudes follow from the cost model in the text (accept length near 2.5, draft length near 3) and are consistent with the production speedups in Figure 8; the absolute dollar levels still move with rate, model, and quantization, anchored here to 2026 neocloud figures (Spheron, getdeploying).</figcaption></figure></div><p>The reason latency-capped service is the common case in 2026, rather than a corner case, is written directly into the benchmark SLAs. </p><p>MLPerf Inference v5.1, published by MLCommons in September 2025, sets for its DeepSeek-R1 reasoning workload a time-to-first-token ninety-ninth-percentile threshold of two seconds and a<strong> time-per-output-token ninety-ninth-percentile threshold of eighty milliseconds</strong>, against a mean input of around eight hundred tokens and a mean output of three thousand eight hundred and eighty, with a maximum output of twenty thousand, the highest the benchmark has ever specified. </p><p>An eighty-millisecond ceiling on per-token latency, applied to sequences thousands of tokens long, caps the batch far below the compute-bound point, because<strong> long sequences mean large KV reads</strong> and large KV reads mean each added unit of batch costs latency you do not have. </p><p>The SLA traps the GPU in the <strong>memory-bound regime</strong>. <mark>The trap is the opportunity: a memory-bound GPU has idle compute, and idle compute is what speculation eats.</mark></p><h3><span>The benchmark concedes the point</span></h3><p>The argument stops being a thesis and becomes a rule when the benchmark authority writes it into the rules. In March 2026, <strong>MLPerf Inference v6.0 added an interactive reasoning scenario</strong> for DeepSeek-R1 with the ceiling pulled tighter still, a 1.5-second TTFT and a 15-millisecond TPOT at the ninety-ninth percentile. </p><p>To make that scenario achievable at all, MLCommons mandates speculative decoding for it: implementations must run the official <strong>DeepSeek-R1 MTP head with EAGLE-style decoding. </strong></p><p>The independent body that defines how inference is measured decided that, past a certain latency target on reasoning traffic, speculation is not an optional optimization but a requirement of entry.</p><p>The dollar figures that frame the chart are anchored to 2026 market rates and published per-token costs, kept deliberately conservative. Neocloud H100 capacity runs around two dollars an hour, with <strong>Spheron listing two dollars and one cent</strong>; B200 on-demand sits in the five-to-six-dollar range across getdeploying and aimultiple&#8217;s trackers. </p><p>Published serving costs land near forty-two cents per million tokens on a B200 and<strong> forty-seven cents on an H100 PCIe.</strong> The point of the chart is not to nail a single deployment&#8217;s economics to the cent, which would be dishonest given how much rate, model, and quantization move the number. </p><p>The point is the asymmetry: the same forty-something cents per million can become a discount or a surcharge depending solely on which side of the saturation line your SLA puts you.</p><p>The most honest evidence for this whole framing comes, unexpectedly, from the vendor with the most incentive to claim an unqualified win. <strong>DeepSeek&#8217;s hardware paper, arXiv:2505.09343</strong>, states that its multi-token prediction module can slightly hurt raw throughput while significantly improving end-to-end generation latency. </p><p>Read that again in the context of the two regimes. DeepSeek is reporting, in print, that in a <strong>throughput accounting MTP </strong>can cost a little, and in a latency accounting it helps a lot, and that they ship it because latency is the product. </p><p>They add a second-order point that sharpens it further: MTP raises the effective batch size, which in their <strong>mixture-of-experts architecture </strong>increases expert-parallel arithmetic intensity, partially offsetting the throughput cost. </p><p>A company could have quoted the latency win alone and called it a free lunch. Instead they <strong>documented the tradeoff in both directions</strong>, which is precisely the shape of the real ledger this issue is arguing for.</p><p>When the vendor with the strongest incentive to claim a pure throughput win instead publishes that the technique &#8220;<em>slightly hurts throughput while significantly improving latency</em>,&#8221; that is not a weakness in the technique. </p><p>It is the ledger showing its true two-sided shape, in the vendor&#8217;s own numbers.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/p/decode-is-memory-bound-speculation?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/p/decode-is-memory-bound-speculation?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><div><hr></div><h2>Why reasoning moved the bill into decode</h2><p>Cost per token has a denominator, and the denominator is dominated by decode steps, because <strong>prefill is a single parallel pass</strong> over the prompt while decode is a long sequence of memory-bound steps, one per output token. </p><p>Anything that multiplies the number of output tokens multiplies the share of the bill that lives in decode, which is exactly the share speculation can attack. This is <strong>why 2026 is a different economic environment</strong> for speculative decoding than 2023 was, even though the technique is largely the same. The traffic changed.</p><p>Reasoning models emit output on a different scale entirely. A conventional chat reply is a few hundred tokens. A reasoning trace runs to thousands, and the trend within the model generation has been sharply upward: <strong>BentoML&#8217;s deployment guide </strong>notes that DeepSeek-R1-0528 nearly doubled its reasoning length over the prior R1, from around twelve thousand to around twenty-three thousand tokens on a single hard math question. </p><p>MLPerf&#8217;s DeepSeek-R1 workload puts the mean output at three thousand eight hundred and eighty and the <strong>maximum at twenty thousand</strong>. Agentic systems then chain many such traces into a single user-visible task, so the effective output length per task can be larger still. </p><p>The bill, which used to be split between a substantial prefill and a modest decode, has tilted hard toward decode.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!hi0v!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F340c3881-d019-43c5-acf6-e5a668d8f89a_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!hi0v!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F340c3881-d019-43c5-acf6-e5a668d8f89a_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!hi0v!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F340c3881-d019-43c5-acf6-e5a668d8f89a_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!hi0v!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F340c3881-d019-43c5-acf6-e5a668d8f89a_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!hi0v!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F340c3881-d019-43c5-acf6-e5a668d8f89a_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!hi0v!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F340c3881-d019-43c5-acf6-e5a668d8f89a_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/340c3881-d019-43c5-acf6-e5a668d8f89a_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Horizontal bar chart of output tokens per request growing from a few hundred for chat to tens of thousands for reasoning workloads.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Horizontal bar chart of output tokens per request growing from a few hundred for chat to tens of thousands for reasoning workloads." title="Horizontal bar chart of output tokens per request growing from a few hundred for chat to tens of thousands for reasoning workloads." srcset="https://substackcdn.com/image/fetch/$s_!hi0v!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F340c3881-d019-43c5-acf6-e5a668d8f89a_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!hi0v!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F340c3881-d019-43c5-acf6-e5a668d8f89a_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!hi0v!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F340c3881-d019-43c5-acf6-e5a668d8f89a_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!hi0v!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F340c3881-d019-43c5-acf6-e5a668d8f89a_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Figure 5.</strong> Reasoning moved the bill into decode. A reasoning request emits one to two orders of magnitude more output tokens than a chat reply, and every one of them is a memory-bound decode step that must clear the latency SLA. MLPerf Inference v5.1 DeepSeek-R1 (MLCommons, September 2025) reports a mean output of 3,880 tokens and a maximum of 20,000; R1 and R1-0528 AIME usage runs from roughly 12,000 to 23,000 tokens per question (BentoML guide). The chat baseline is a round-number reference.</figcaption></figure></div><p><strong>Two facts</strong> about reasoning traffic place it squarely in the regime where speculation pays. The first is the one just shown: it is <strong>decode-heavy</strong>, so the part of the bill speculation can lower is the dominant part. The second is subtler and follows from the previous sections. </p><p>Long traces mean long sequences in flight, which means<strong> large KV reads</strong>, which means decode is memory-bound even when the operator manages to batch, both because the <strong>latency SLA</strong> caps the batch and because the KV cost re-opens the budget past S-star. </p><p>The two mechanisms reinforce each other. The workload is in the memory-bound regime by virtue of its output length, and it is held there by <strong>virtue of its latency SLA</strong>. There is idle compute on the machine for both reasons at once, and speculation is the technique that turns idle compute into tokens.</p><p>The architectural direction of travel keeps the budget open rather than closing it. <strong>DeepSeek-V4&#8217;s sparse-attention work,</strong> with V4-Pro reportedly using about twenty-seven percent of the FLOPs and ten percent of the KV of V3.2 at a one-million-token context through DeepSeek Sparse Attention, is explicitly aimed at making very long contexts affordable.</p><p>Cheaper long context means more long context, which means more memory-bound decode, which means a wider speculation budget, not a narrower one. The <strong>hardware trend</strong> widens the budget by improving compute faster than bandwidth; the model trend widens it by pushing context length up. Both vectors point the same way.</p><p>And this is where accept length, the lever from the first section, cashes in <strong>directly against the denominator</strong>. The decode discount is, to first order, the accept length: a method that lands two and a half accepted tokens per step is doing roughly two and a half times the decode work per expensive weight load. </p><p>The measured accept lengths of the shipped 2026 methods, two and fifty-five hundredths for <strong>DeepSeek-V3.2&#8217;s MTP,</strong> two and seventy-six hundredths for GLM-5&#8217;s shared-MTP design, and four and seven tenths for EAGLE-3 on coding and reasoning workloads, are therefore not abstract quality scores. </p><p>They are multipliers on the largest line item in the reasoning-era bill.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!U_FB!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F410e3d7e-698b-4812-8847-93486f5c88d5_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!U_FB!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F410e3d7e-698b-4812-8847-93486f5c88d5_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!U_FB!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F410e3d7e-698b-4812-8847-93486f5c88d5_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!U_FB!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F410e3d7e-698b-4812-8847-93486f5c88d5_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!U_FB!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F410e3d7e-698b-4812-8847-93486f5c88d5_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!U_FB!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F410e3d7e-698b-4812-8847-93486f5c88d5_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/410e3d7e-698b-4812-8847-93486f5c88d5_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Bar chart of mean accepted tokens per verification step across vanilla drafting, DeepSeek-V3.2 MTP, GLM-5 shared-MTP, and EAGLE-3.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Bar chart of mean accepted tokens per verification step across vanilla drafting, DeepSeek-V3.2 MTP, GLM-5 shared-MTP, and EAGLE-3." title="Bar chart of mean accepted tokens per verification step across vanilla drafting, DeepSeek-V3.2 MTP, GLM-5 shared-MTP, and EAGLE-3." srcset="https://substackcdn.com/image/fetch/$s_!U_FB!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F410e3d7e-698b-4812-8847-93486f5c88d5_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!U_FB!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F410e3d7e-698b-4812-8847-93486f5c88d5_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!U_FB!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F410e3d7e-698b-4812-8847-93486f5c88d5_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!U_FB!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F410e3d7e-698b-4812-8847-93486f5c88d5_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Figure 6.</strong> The lever is accept length, not raw acceptance. Tokens produced per step equal the accepted run plus one bonus token, and each step costs exactly one target weight load, so accept length is the decode discount. Reported figures: DeepSeek-V3.2 MTP <strong>2.55</strong> and GLM-5 shared-MTP <strong>2.76</strong> from the GLM-5 technical report (arXiv:2602.15763); EAGLE-3 <strong>4.5 to 5.0</strong> on coding and reasoning from E2E Networks. The vanilla-draft figure is a representative reference.</figcaption></figure></div><div><hr></div><h2>Where speculation loses</h2><p>A technique that only ever helps does not need an issue written about when to use it. Speculation has real failure modes, and an analysis that buries them is worth less than one that lists them, so here is the full debit column, without hedging.</p><p><strong>The throughput regime.</strong> In saturated high-batch short-context serving, the offline-batch corner of the phase diagram, <mark>speculation is a tax and should be disabled.</mark> The compute is fully employed, the draft has nothing free to run in, and its forward passes displace real work. The Spheron and Tian Pan guidance is correct here without qualification. If your job is to push the maximum number of short completions through a fleet of GPUs at minimum cost per token, speculation is the wrong lever.</p><p><strong>Acceptance collapse.</strong> <mark>Below roughly one-half acceptance, speculation hurts at any batch</mark>, because too few candidates survive to cover the draft&#8217;s cost. Acceptance is not a constant; it falls with high sampling temperature, with out-of-distribution inputs the draft was never trained on, and with the kind of high-entropy generation where the next token is genuinely uncertain. A draft trained against one target distribution and then serving a drifted or fine-tuned target degrades silently, the acceptance rate sliding without any error being raised. Monitoring accept length in production is not optional; it is the only way to notice that your discount has quietly become a surcharge.</p><p><strong>VRAM pressure.</strong> The draft model and its KV cache occupy memory you could otherwise spend on a larger batch or a longer context. A Llama-3.3-70B target in FP8 alongside a one-billion-parameter draft consumes roughly seventy-five to seventy-eight gigabytes on an eighty-gigabyte H100, per Spheron&#8217;s figures, leaving very little headroom. On a memory-constrained deployment, the draft can cost you more in lost batch capacity than it returns in accept length, and that tradeoff has to be measured, not assumed.</p><p><strong>No help for time-to-first-token.</strong> Speculation accelerates decode, and only decode. It does nothing for prefill, which means it does nothing for time-to-first-token. Under the MLPerf two-second TTFT ceiling, that is a separate problem requiring separate techniques, which is precisely the prefill-decode disaggregation argument of Issue 03. Speculation and prefill optimization are complementary, not substitutes, and a serving stack that needs both will not get the first from the second.</p><p><strong>Draft maintenance.</strong> A separate draft model is a second training, evaluation, and deployment surface that must be kept aligned as the target evolves. Every target update risks degrading a draft that was tuned against the previous version. EAGLE-style heads and built-in MTP layers reduce this by coupling the draft to the target&#8217;s own features or parameters, but they do not eliminate the obligation to retrain and revalidate. GLM-5&#8217;s choice to share parameters across three MTP layers, described in arXiv:2602.15763, is partly an answer to exactly this maintenance cost: fewer independent parameters to train and keep aligned.</p><p><strong>The losslessness caveat, restated.</strong> The provable equivalence to standard sampling holds for the exact rejection-sampling rule. Relaxed acceptance, typical acceptance, and aggressive tree-acceptance schemes raise throughput by changing the output distribution. They are frequently worth it, but a four-times figure obtained under a relaxed rule is not interchangeable with a four-times figure under exact sampling, and a serving team quoting a speedup owes itself, and its users, clarity about which rule produced it.</p><p><strong>Draft-length tuning.</strong> The number of tokens the draft proposes per step, often written gamma, is a workload-dependent knob with a real optimum. Set it too long and the draft burns compute generating candidates that will be rejected; set it too short and you leave accept length on the table. The optimum moves with acceptance rate and with batch size, so a value tuned on one workload can be wrong on another, and dynamic schemes that adjust it per request exist precisely because no single value is right everywhere.</p><p><strong>The lab-versus-production gap.</strong> The EAGLE-3 paper reports speedups of up to six and a half times, but those are temperature-zero academic measurements on Vicuna-13B, Llama-3.1-8B, and Llama-3.3-70B. Production reports cluster instead around two to three times: LMSYS and Vertex describe two-to-three-times figures for EAGLE-3 on SGLang, E2E Networks reports two and three tenths on Llama-3.1-8B at a batch of four, and a Gemma-4 EAGLE3 draft head is documented at one and seventy-two hundredths at batch one on conversational traffic. The gap between the lab number and the production number is itself one of the most important facts in this space, because <mark>quoting the former as if it were the latter is the single most common honesty failure in vendor material on speculative decoding.</mark></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!IE6Z!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1992dc1-fdc3-410b-af49-2787e536fe46_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!IE6Z!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1992dc1-fdc3-410b-af49-2787e536fe46_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!IE6Z!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1992dc1-fdc3-410b-af49-2787e536fe46_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!IE6Z!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1992dc1-fdc3-410b-af49-2787e536fe46_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!IE6Z!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1992dc1-fdc3-410b-af49-2787e536fe46_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!IE6Z!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1992dc1-fdc3-410b-af49-2787e536fe46_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f1992dc1-fdc3-410b-af49-2787e536fe46_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Horizontal bar chart contrasting a 6.5x lab speedup against a cluster of production speedups between 1.66x and 2.3x.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Horizontal bar chart contrasting a 6.5x lab speedup against a cluster of production speedups between 1.66x and 2.3x." title="Horizontal bar chart contrasting a 6.5x lab speedup against a cluster of production speedups between 1.66x and 2.3x." srcset="https://substackcdn.com/image/fetch/$s_!IE6Z!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1992dc1-fdc3-410b-af49-2787e536fe46_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!IE6Z!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1992dc1-fdc3-410b-af49-2787e536fe46_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!IE6Z!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1992dc1-fdc3-410b-af49-2787e536fe46_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!IE6Z!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1992dc1-fdc3-410b-af49-2787e536fe46_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Figure 7.</strong> The headline academic figure (6.5x, EAGLE-3 paper, temperature 0) sits far above the production cluster, which lands between roughly 1.66x and 2.3x across the EAGLE 3.1 vLLM benchmark, DeepSeek&#8217;s vendor-reported MTP TPS, a Gemma-4 EAGLE3 draft head, and E2E Networks. The bars use different targets and conditions and are not strictly comparable; they are shown to convey the range, and the distance between the gold bar and the teal cluster is the point.</figcaption></figure></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>A decision on the plane</h2><p>The verdict is not a yes or a no. It is a lookup. Given a workload&#8217;s batch size, its context length, its <strong>latency SLA</strong>, and its measured acceptance, the phase diagram tells you which regime you are in, and the regime tells you the sign of the ledger. The table below collapses the analysis into that lookup.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!wYrw!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37830286-9e44-4ed2-a815-f14199ca909b_1710x1168.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!wYrw!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37830286-9e44-4ed2-a815-f14199ca909b_1710x1168.png 424w, https://substackcdn.com/image/fetch/$s_!wYrw!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37830286-9e44-4ed2-a815-f14199ca909b_1710x1168.png 848w, https://substackcdn.com/image/fetch/$s_!wYrw!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37830286-9e44-4ed2-a815-f14199ca909b_1710x1168.png 1272w, https://substackcdn.com/image/fetch/$s_!wYrw!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37830286-9e44-4ed2-a815-f14199ca909b_1710x1168.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!wYrw!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37830286-9e44-4ed2-a815-f14199ca909b_1710x1168.png" width="1456" height="995" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/37830286-9e44-4ed2-a815-f14199ca909b_1710x1168.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:995,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:248170,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/199710199?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37830286-9e44-4ed2-a815-f14199ca909b_1710x1168.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!wYrw!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37830286-9e44-4ed2-a815-f14199ca909b_1710x1168.png 424w, https://substackcdn.com/image/fetch/$s_!wYrw!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37830286-9e44-4ed2-a815-f14199ca909b_1710x1168.png 848w, https://substackcdn.com/image/fetch/$s_!wYrw!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37830286-9e44-4ed2-a815-f14199ca909b_1710x1168.png 1272w, https://substackcdn.com/image/fetch/$s_!wYrw!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37830286-9e44-4ed2-a815-f14199ca909b_1710x1168.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>State the thesis cleanly now that the machinery is in place. <strong>Speculative decoding is the only inference technique</strong> that converts decode&#8217;s idle compute into tokens losslessly, and the sign of its effect on your bill is set entirely by whether that compute was actually idle. </p><p>Whether it was idle is a question about the roofline, and <strong>where you sit on the roofline</strong> is a question about batch size and context length, the two axes of the phase diagram. </p><p>There is no universal answer because the inputs are not universal. <mark>There is a correct answer for every point on the plane, and the table is how you read it.</mark></p><p>The reason this matters more now than it did three years ago is that the median workload moved. In <strong>2023 the prototypical request</strong> was a short chat completion at modest context, which lives near the compute-bound corner once you batch it, where speculation is at best neutral. </p><p>In 2026 the prototypical high-value request is a long reasoning trace under a <strong>tight per-token latency SLA</strong>, which lives deep in the memory-bound region for two independent reasons, its output length and its SLA, and which is therefore exactly where speculation pays. <mark>The technique did not move toward the workload. The workload moved toward the technique.</mark></p><p>That migration is why the shipping decisions of the major laboratories converged. <strong>DeepSeek built multi-token prediction into V3</strong> and carried it through V3.2, documenting the latency win and the throughput cost honestly. </p><p>GLM-5 shipped a shared-parameter three-layer MTP design with a measured accept length near two and three-quarters. <strong>NVIDIA&#8217;s NeMo RL work applied EAGLE-3 to reinforcement-learning rollouts</strong> and reported a one-and-eight-tenths-times generation speedup at the eight-billion scale, with validation accuracy on AIME-2024 evolving identically under autoregressive and speculative decoding, a clean confirmation that the lossless guarantee holds across training. </p><p><strong>EAGLE-3 landed across vLLM, SGLang, and TensorRT-LLM</strong>, the three serving stacks that matter. These are not independent fashions. They are the same bet, placed by everyone who looked at the same plane and saw that reasoning traffic had walked into the half where the ledger reads in your favor.</p><p><em>The technique never changed. The regime did. Speculation lowers your cost per token exactly where batching cannot help you, and reasoning is the workload that made that region the center of the map.</em></p><div><hr></div><h2>Confidence tiers and external-audit read</h2><p>Every load-bearing claim in this issue is scored below against a four-tier confidence scale, with its source named inline. </p><p>The <strong>scale is applied as an external auditor would apply it</strong>, crediting primary and measured sources, discounting derived and illustrative ones, and flagging the weakest links explicitly rather than hiding them in the prose.</p><p><strong>Tier A</strong> <em>primary or measured: peer-reviewed papers, vendor hardware disclosures, MLPerf-published SLAs and benchmark statistics.</em><br><strong>Tier B</strong> <em>secondary, with method: vendor or practitioner reports that state their configuration and measurement conditions.</em><br><strong>Tier C</strong> <em>derived or stylized: house figures and curves built from the cited physics, presented as illustrative renderings, not measurements.</em><br><strong>Tier D</strong> <em>illustrative or round-number: reference values chosen for scale, not claimed as measured.</em></p><p><strong>A</strong></p><p><strong>Exact rejection-sampling speculative decoding is output-distribution lossless.</strong></p><p>Leviathan et al. 2023; Chen et al. 2023; EAGLE losslessness per Hugging Face engineering writeup. Provable equivalence to standard sampling under the exact rule.</p><p><strong>A</strong></p><p><strong>EAGLE-3 mechanism: training-time test, multi-layer feature fusion, direct token prediction, dynamic draft tree.</strong></p><p>EAGLE-3 paper, NeurIPS 2025, arXiv:2503.01840. Headline lab speedups up to 6.5x at temperature 0 on Vicuna-13B, Llama-3.1-8B, Llama-3.3-70B (explicitly a best-case lab figure; production lands far lower, see Figure 8).</p><p><strong>A</strong></p><p><strong>DeepSeek MTP slightly hurts throughput while significantly improving end-to-end latency; raises effective batch and expert-parallel intensity.</strong></p><p>DeepSeek hardware paper, arXiv:2505.09343. The two-sided tradeoff is stated in the vendor&#8217;s own text, which is the strongest single piece of evidence in this issue.</p><p><strong>A</strong></p><p><strong>MLPerf DeepSeek-R1 SLAs: TTFT 99p 2s, TPOT 99p 80ms; mean input 800, mean output 3,880, max output 20,000.</strong></p><p>MLPerf Inference v5.1, MLCommons, September 2025. These SLAs are the anchor for why latency-capped serving traps the GPU in the memory-bound regime.</p><p><strong>A</strong></p><p><strong>MLPerf Inference v6.0 added a DeepSeek-R1 interactive scenario (TTFT 1.5s, TPOT 15ms) and mandates speculative decoding (official MTP head, EAGLE-style) to meet it.</strong></p><p>MLCommons, March 2026. The benchmark authority requiring speculation for tight-latency reasoning is the strongest external corroboration of the thesis.</p><p><strong>A</strong></p><p><strong>NVIDIA NeMo RL: EAGLE-3 gives ~1.8x rollout generation speedup at 8B; AIME-2024 accuracy identical under autoregressive and speculative decoding throughout training.</strong></p><p>NVIDIA NeMo RL research, May 2026. Independent empirical confirmation that the lossless guarantee holds in practice across training.</p><p><strong>A</strong></p><p><strong>Critical-sequence-length result: past S*, decode is memory-bound even at large batch via KV read; KV-light draft delivers up to ~2x throughput and latency.</strong></p><p>MagicDec, arXiv:2408.11049; Together AI long-context analysis. Draft-to-target memory ratio ~0.4, constant at large batch, for Llama-3.1-70B with 8B draft.</p><p><strong>A</strong></p><p><strong>Roofline ridge points: H100 SXM FP8 591, H200 412, B200 ~562 FLOP/byte; compute outgrew bandwidth ~36x vs ~9x V100 to B200.</strong></p><p>Carried from Issue 03, The Split and the Seam, and Issue on the memory wall, The Wall and the Stack. Hardware specifications and systems-literature divergence figures.</p><p><strong>A</strong></p><p><strong>DeepSeek-V4 sparse attention: V4-Pro 27% FLOPs and 10% KV of V3.2 at 1M context; V4-Flash 10% and 7%.</strong></p><p>Official DeepSeek-V4-Pro / V4-Flash model cards (Compressed Sparse Attention + Heavily Compressed Attention). Verified figures. Direction-of-travel evidence that long context is getting cheaper, widening the budget.</p><p><strong>A</strong></p><p><strong>EAGLE 3.1 per-user throughput: 2.03x at concurrency 1, 1.71x at 4, 1.66x at 16.</strong></p><p>Primary: vLLM team blog, May 2026 (EAGLE / vLLM / TorchSpec joint release). Kimi-K2.6-NVFP4, tensor-parallel 4, GB200, non-disagg, SPEED-Bench coding. Verified against the primary engineering writeup.</p><div><hr></div><p><strong>B</strong></p><p><strong>Practical break-even near batch 32; below ~0.5 acceptance speculation hurts at any batch.</strong></p><p>Spheron production guide, March 2026; E2E Networks. Practitioner guidance with stated qualifications on model, quantization, and sequence length.</p><p><strong>B</strong></p><p><strong>The lab-versus-production speedup comparison in Figure 8 (6.5x lab vs a 1.66 to 2.3x production cluster).</strong></p><p>Compiled from verified primary and vendor sources (EAGLE-3 paper, EAGLE 3.1 vLLM blog, DeepSeek hardware paper, Gemma-4 EAGLE3 card, E2E Networks). Bars use different targets and conditions and are not strictly comparable; shown to convey range, not to rank.</p><p><strong>B</strong></p><p><strong>Production speedups cluster at 2 to 3x; E2E reports 2.3x on Llama-3.1-8B at batch 4, accept length 4.5 to 5.0.</strong></p><p>LMSYS and Vertex on SGLang; E2E Networks. Multiple independent practitioner reports converging on the same range.</p><p><strong>B</strong></p><p><strong>Accept lengths: DeepSeek-V3.2 MTP 2.55, GLM-5 shared-MTP 2.76; GLM-5 shares parameters across 3 MTP layers.</strong></p><p>GLM-5 technical report, arXiv:2602.15763. EAGLE-3 accept length 4.5 to 5.0 from E2E Networks.</p><p><strong>B</strong></p><p><strong>2026 GPU rates and per-token costs: H100 ~$2/hr, B200 ~$5 to $6/hr on-demand; ~$0.42/M (B200), ~$0.47/M (H100 PCIe).</strong></p><p>Spheron ($2.01/hr H100), getdeploying, aimultiple. Market trackers; rates move with provider, commitment, and region.</p><p><strong>B</strong></p><p><strong>Reasoning length growth: R1-0528 nearly doubled to ~23K tokens per AIME question vs ~12K for prior R1.</strong></p><p>BentoML DeepSeek deployment guide, 2026. R1 token pricing ~$0.55/M in, ~$2.19/M out; output billed higher and dominates cost.</p><p><strong>C</strong></p><p><strong>The KV-versus-weight amortization crossover (Figure 4): batch x sequence near 220,000 for a 70B model, where the KV read overtakes the weight read (~7,000 tokens at batch 32, ~1,750 at batch 128).</strong></p><p>House order-of-magnitude calculation from Llama-3-70B architecture (80 layers, 8 KV heads, head-dim 128) and the weight-versus-KV read balance, drawn in Figure 4. The coefficient moves with KV precision and serving format; the order of magnitude, and the conclusion that reasoning traces sit past the crossover, is robust.</p><p><strong>C</strong></p><p><strong>The compute-bound boundary and ~1,000-token wall in Figure 3.</strong></p><p>Derived, not stylized: the boundary is the roofline condition AI = 2B/(1+B*S/C) exceeding the ridge, with the wall at S = 2C/R (C ~ 220,000; R = 562 FP8 for B200). A house calculation with standard simplifying assumptions (GEMM-dominated FLOPs, dense GQA KV); exact coordinates shift with precision and ridge, the structure does not.</p><p><strong>C</strong></p><p><strong>The -48% / +19% cost magnitudes in Figure 5.</strong></p><p>Derived from the cost model in the text (accept length ~2.5, draft length ~3, draft overhead ~0.2) and consistent with the production speedups in Figure 8. Absolute dollar levels still vary with rate, model, and quantization; the regime-dependent sign and rough magnitude are the load-bearing claim.</p><p><strong>C</strong></p><p><strong>The shape of the speedup-decay envelope in Figure 2.</strong></p><p>House curve illustrating the conventional decay toward break-even; the measured EAGLE 3.1 points on it are Tier A and the break-even location is sourced, but the connecting envelope is schematic, not fitted.</p><p><strong>D</strong></p><p><strong>The ~300-token chat-reply baseline and the vanilla-draft 2.1 accept-length reference.</strong></p><p>Round-number references chosen to set scale against the measured reasoning and method figures, not claimed as measured values.</p><h3>External-audit simulation</h3><p>Audited as a whole, the issue rests on a spine of Tier A primary sources, and every load-bearing empirical claim in it was checked against its primary source: the <strong>EAGLE-3 paper </strong>(arXiv:2503.01840),<strong> MagicDec </strong>(arXiv:2408.11049), the <strong>DeepSeek hardware disclosure</strong> (arXiv:2505.09343), <strong>the GLM-5 report </strong>(arXiv:2602.15763), the <strong>MLPerf v5.1 and v6.0 </strong>specifications, the <strong>DeepSeek-V4 model cards</strong>, and the EAGLE 3.1 vLLM <strong>release </strong>each confirmed the figures attributed to them. </p><p>The<strong> central argument</strong>, that the sign of the ledger is set by serving regime and that reasoning traffic sits in the favorable regime, follows from those sources rather than from the house figures, which is the property an audit most wants to see. </p><p>Two pieces of evidence are doing disproportionate work and both survive scrutiny: DeepSeek&#8217;s own statement that <strong>MTP slightly hurts throughput </strong>while significantly <strong>improving latency</strong>, a direct vendor quote, and MLPerf v6.0&#8217;s decision to mandate speculative decoding for its tight-latency reasoning scenario, which is the measuring authority writing the thesis into the rules.</p><p>The weakest links are named rather than hidden, and after this revision they are narrow. <strong>Boundary and Crossover</strong> are now derivations rather than stylizations: the first is the roofline condition with the wall at twice the bytes-ratio over the ridge, the second is the<strong> weight-versus-KV balance</strong>, both house calculations carried out with standard simplifying assumptions (<em>GEMM-dominated compute, dense grouped-query KV</em>) whose exact coordinates move with precision and ridge while the structure holds. </p><p>The cost magnitudes are likewise derived from an explicit model, an accept length near two and a half and a draft length near three, and <strong>cross-checked against the production speedups</strong>; what remains genuinely soft there is the absolute dollar level, which varies too much with rate, model, and quantization to pin down. </p><p>The <strong>decay envelope is a schematic shape</strong>, though the measured points on it and its break-even location are sourced. The chat and vanilla-draft baselines are round numbers for scale. A handful of practitioner figures, the break-even batch, the <strong>VRAM envelope</strong>, and the production speedup band, come from individual engineering guides rather than independently reproduced benchmarks, and are tiered B accordingly. </p><p>None of these elements carries the conclusion: a reader who accepts only the Tier A claims, and works the two house calculations independently, arrives at the same verdict table.</p><p><strong>Overall confidence in the thesis is high</strong>, because the thesis is a statement about regimes and signs that the primary sources support directly, and it is deliberately not a statement that speculation yields any specific universal multiple, which the evidence would not support. </p><p>The <strong>quantitative illustrations</strong> are held at lower confidence by design, and labeled as such, so that the argument does not borrow credibility it has not earned. A reader who accepts only the Tier A claims still arrives at the same verdict table; the lower tiers furnish the texture, not the conclusion.</p><div><hr></div><h2>Sources and useful informations</h2><p>Primary and secondary sources for the load-bearing claims, with arXiv identifiers and venues where applicable. The text above attributes each source at its point of use; this is the consolidated record. </p><p>Papers appear first, then benchmark specifications, vendor and practitioner writeups, and model cards.</p><ol><li><p><span>Y. Leviathan, M. Kalman, and Y. Matias. </span><em><span>Fast Inference from Transformers via Speculative Decoding.</span></em><span> International Conference on Machine Learning (ICML), 2023. arXiv:2211.17192.</span></p></li><li><p><span>C. Chen, S. Borgeaud, G. Irving, J.-B. Lespiau, L. Sifre, and J. Jumper. </span><em><span>Accelerating Large Language Model Decoding with Speculative Sampling.</span></em><span> arXiv:2302.01318, 2023.</span></p></li><li><p><span>T. Cai, Y. Li, Z. Geng, H. Peng, J. D. Lee, D. Chen, and T. Dao. </span><em><span>Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads.</span></em><span> International Conference on Machine Learning (ICML), 2024. arXiv:2401.10774.</span></p></li><li><p><span>Y. Li, F. Wei, C. Zhang, and H. Zhang. </span><em><span>EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty.</span></em><span> International Conference on Machine Learning (ICML), 2024. arXiv:2401.15077.</span></p></li><li><p><span>Y. Li, F. Wei, C. Zhang, and H. Zhang. </span><em><span>EAGLE-2: Faster Inference of Language Models with Dynamic Draft Trees.</span></em><span> Conference on Empirical Methods in Natural Language Processing (EMNLP), 2024. arXiv:2406.16858.</span></p></li><li><p><span>Y. Li, F. Wei, C. Zhang, and H. Zhang. </span><em><span>EAGLE-3: Scaling up Inference Acceleration of Large Language Models via Training-Time Test.</span></em><span> Conference on Neural Information Processing Systems (NeurIPS), 2025. arXiv:2503.01840.</span></p></li><li><p><span>R. Sadhukhan, J. Chen, Z. Chen, V. Tiwari, A. May, T. Chen, and B. Chen. </span><em><span>MagicDec: Breaking the Latency-Throughput Tradeoff for Long Context Generation with Speculative Decoding.</span></em><span> arXiv:2408.11049, 2024.</span></p></li><li><p><span>DeepSeek-AI. </span><em><span>DeepSeek-V3 Technical Report.</span></em><span> arXiv:2412.19437, 2024.</span></p></li><li><p><span>DeepSeek-AI. </span><em><span>Insights into DeepSeek-V3: Scaling Challenges and Reflections on Hardware for AI Architectures.</span></em><span> International Symposium on Computer Architecture (ISCA), 2025. arXiv:2505.09343.</span></p></li><li><p><span>DeepSeek-AI. </span><em><span>DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.</span></em><span> arXiv:2501.12948, 2025.</span></p></li><li><p><span>Z.ai (Zhipu AI). </span><em><span>GLM-5 Technical Report.</span></em><span> arXiv:2602.15763, 2026.</span></p></li><li><p><span>MLCommons. </span><em><span>MLPerf Inference: Datacenter, v5.1 (DeepSeek-R1 reasoning workload).</span></em><span> Benchmark rules and results, 2025.</span></p></li><li><p><span>MLCommons. </span><em><span>MLPerf Inference: Datacenter, v6.0 (DeepSeek-R1 Interactive scenario, mandated speculative decoding).</span></em><span> Benchmark rules, 2026.</span></p></li><li><p><span>EAGLE Team, vLLM Team, and TorchSpec. </span><em><span>EAGLE 3.1: release and SPEED-Bench results on Kimi-K2.6.</span></em><span> vLLM Blog, May 2026.</span></p></li><li><p><span>NVIDIA. </span><em><span>Speculative Decoding for Reinforcement-Learning Rollouts in NeMo RL.</span></em><span> NVIDIA Developer technical writeup, 2026.</span></p></li><li><p><span>Together AI. </span><em><span>Speculative decoding for high-throughput long-context inference (analysis of MagicDec).</span></em><span> Together AI Blog, 2024.</span></p></li><li><p><span>BentoML. </span><em><span>The Complete Guide to DeepSeek Models: V3, R1, V4 and Beyond.</span></em><span> BentoML Blog, 2026.</span></p></li><li><p><span>Spheron Network. </span><em><span>Speculative decoding in production: a practitioner&#8217;s guide.</span></em><span> Engineering guide, 2026.</span></p></li><li><p><span>E2E Networks. </span><em><span>Speculative decoding performance on Llama-3.1 serving.</span></em><span> Engineering notes, 2026.</span></p></li><li><p><span>DeepSeek-AI. </span><em><span>DeepSeek-R1-0528.</span></em><span> Model card, Hugging Face, 2025.</span></p></li><li><p><span>DeepSeek-AI. </span><em><span>DeepSeek-V4-Pro and DeepSeek-V4-Flash.</span></em><span> Model cards, Hugging Face, 2026.</span></p><p></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share The Software Frontier&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share The Software Frontier</span></a></p><p></p></li></ol>]]></content:encoded></item><item><title><![CDATA[The Split and the Seam]]></title><description><![CDATA[Eighteen months ago, splitting prefill from decode was a contrarian research bet. Today it is the default that every serious stack runs, it helped erase 600 billion dollars from the company that build]]></description><link>https://www.thesoftwarefrontier.com/p/the-split-and-the-seam</link><guid isPermaLink="false">https://www.thesoftwarefrontier.com/p/the-split-and-the-seam</guid><dc:creator><![CDATA[Lorenzo Bradanini]]></dc:creator><pubDate>Sun, 21 Jun 2026 14:02:14 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!a3zy!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e933c23-7e0f-4ee4-be74-01d31e1bd107_1122x1402.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!a3zy!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e933c23-7e0f-4ee4-be74-01d31e1bd107_1122x1402.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!a3zy!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e933c23-7e0f-4ee4-be74-01d31e1bd107_1122x1402.png 424w, https://substackcdn.com/image/fetch/$s_!a3zy!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e933c23-7e0f-4ee4-be74-01d31e1bd107_1122x1402.png 848w, https://substackcdn.com/image/fetch/$s_!a3zy!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e933c23-7e0f-4ee4-be74-01d31e1bd107_1122x1402.png 1272w, https://substackcdn.com/image/fetch/$s_!a3zy!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e933c23-7e0f-4ee4-be74-01d31e1bd107_1122x1402.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!a3zy!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e933c23-7e0f-4ee4-be74-01d31e1bd107_1122x1402.png" width="1122" height="1402" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2e933c23-7e0f-4ee4-be74-01d31e1bd107_1122x1402.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1402,&quot;width&quot;:1122,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2294382,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/202706223?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e933c23-7e0f-4ee4-be74-01d31e1bd107_1122x1402.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!a3zy!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e933c23-7e0f-4ee4-be74-01d31e1bd107_1122x1402.png 424w, https://substackcdn.com/image/fetch/$s_!a3zy!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e933c23-7e0f-4ee4-be74-01d31e1bd107_1122x1402.png 848w, https://substackcdn.com/image/fetch/$s_!a3zy!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e933c23-7e0f-4ee4-be74-01d31e1bd107_1122x1402.png 1272w, https://substackcdn.com/image/fetch/$s_!a3zy!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e933c23-7e0f-4ee4-be74-01d31e1bd107_1122x1402.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>Intro</h2><p>On Monday, <strong>January 27, 2025</strong>, NVIDIA lost about 600 billion dollars of market value in a single trading session, a 17 percent fall that stands as the <strong>largest one-day market-capitalization loss</strong> in the history of US public markets. It did not fall alone. </p><p>Broadcom dropped 17.4 percent, Marvell 19.1 percent, AMD 19.1 percent, the <strong>chip-adjacent names down</strong> in sympathy across the board. The trigger was not an earnings miss or a product recall. </p><p>It was a technical report from a <strong>Chinese lab, DeepSeek</strong>, whose V3 and R1 models matched the frontier while having been trained and, crucially, served at a fraction of the assumed cost. </p><p>The market did the obvious arithmetic: if <strong>inference is far cheaper</strong> than actually we believed, fewer GPUs are needed, so the company selling the GPUs is worth less.</p><p>The arithmetic was wrong, and the reason it was wrong is the subject of this issue. Within days, Microsoft&#8217;s Satya Nadella was citing the<strong> Jevons paradox</strong>, the nineteenth-century observation that making a resource cheaper to use tends to increase, not decrease, its total consumption. The paradox held beautifully. </p><p>Over the course of 2025, by the Peterson Institute&#8217;s accounting, the cost to reach a fixed score on a <strong>hard reasoning benchmark fell</strong> from about 4,500 dollars per task to 11.64 dollars, a roughly<strong> 386-fold collapse</strong>, and inference usage did not shrink to match. It exploded past the efficiency gains, exactly as Jevons would predict. </p><p>The chip that was supposed to be made redundant by cheap inference is now sold out for years on the back of it [Figure 1].</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!zuXr!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa34aff13-7ef3-4e83-b325-ba8b05d0417f_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!zuXr!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa34aff13-7ef3-4e83-b325-ba8b05d0417f_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!zuXr!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa34aff13-7ef3-4e83-b325-ba8b05d0417f_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!zuXr!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa34aff13-7ef3-4e83-b325-ba8b05d0417f_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!zuXr!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa34aff13-7ef3-4e83-b325-ba8b05d0417f_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!zuXr!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa34aff13-7ef3-4e83-b325-ba8b05d0417f_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a34aff13-7ef3-4e83-b325-ba8b05d0417f_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Figure 12&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 12" title="Figure 12" srcset="https://substackcdn.com/image/fetch/$s_!zuXr!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa34aff13-7ef3-4e83-b325-ba8b05d0417f_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!zuXr!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa34aff13-7ef3-4e83-b325-ba8b05d0417f_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!zuXr!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa34aff13-7ef3-4e83-b325-ba8b05d0417f_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!zuXr!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa34aff13-7ef3-4e83-b325-ba8b05d0417f_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 1. The shake and its resolution. Left, the single-day sector rout that followed DeepSeek&#8217;s efficiency disclosure, the largest one-day loss in market history. Right, the cost to reach a fixed benchmark score over 2025, a collapse the industry has named LLMflation. Efficiency did not destroy demand. It multiplied it.</figcaption></figure></div><p>Here is the part the market missed in its first reaction. DeepSeek&#8217;s efficiency was not a single trick. </p><p>It was a stack of techniques, low-precision FP8 arithmetic, a <strong>sparse mixture-of-experts model</strong>, a compressed attention scheme called multi-head latent attention, and, underneath all of it, an inference architecture that ran the <strong>two phases of language-model serving</strong>, prefill and decode, on entirely separate pools of machines. </p><p>That last technique, <strong>disaggregation</strong>, is the one that matters most for understanding what happened next, because in the eighteen months bracketing that selloff it went from a contrarian research idea that the <strong>open-source community</strong> pushed back on to the default architecture of virtually every production serving system in existence. </p><p><strong>NVIDIA Dynamo, llm-d, Ray Serve, SGLang, vLLM, LMCache</strong>, and Mooncake all run on it now. The very metrics the industry uses to talk about inference latency, time-to-first-token and time-per-output-token, were popularized through its lens. </p><p>The authors of <strong>DistServe</strong>, the paper that named the architecture, marked the moment in a November 2025 retrospective with a wry observation: if Moore&#8217;s law doubles compute every eighteen months, then the <strong>serving-systems equivalent had just doubled too</strong>, not because the chips got faster, but because the systems serving them did.</p><p>This issue is a teardown of that architecture as it actually exists in mid-2026, not as a tidy origin story. We<strong> start with the physics</strong>, because the physics is clean and explains everything downstream. </p><p>Then <strong>why colocation lost</strong>, what the split mechanically does, and the cost it creates at the seam where the two halves rejoin, a cost that the rack-scale hardware of the current generation has largely, but not entirely, dissolved. </p><p>Then we <strong>go down into the kernel layer</strong>, to the expert-parallel all-to-all communication and the custom kernels that make large mixture-of-experts models servable at all, which is the part most coverage skips and the part that actually decides throughput. </p><p>Then the attention rewrites that are shrinking the problem from underneath, the <strong>cross-vendor benchmark numbers</strong> and the honest places they fall apart, the operational tax nobody puts on the slide, the new silicon the split has spawned, and finally what the whole shake settled into.</p><p>The thesis, stated once and defended throughout: disaggregation is no longer a choice an operator debates. It is the substrate. The <strong>live questions have moved up a layer</strong>, to how you balance the pools, how wide you spread the experts, and whether the wire between your machines is fast enough that the seam is free.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Two workloads on opposite ends of the roofline</h2><p>The <strong>roofline model</strong> is the oldest honest tool in performance engineering, and once prefill and decode are plotted on it the rest of this issue is commentary. </p><p>Every kernel has an arithmetic intensity, the<strong> floating-point operations</strong> it performs per byte it moves from memory. Every chip has two ceilings: a compute ceiling set by its peak arithmetic rate, and a memory ceiling set by its bandwidth. </p><p>A kernel is compute-bound when its arithmetic intensity is high enough that the <strong>chip exhausts its FLOPS before its bandwidth</strong>, and memory-bound otherwise. </p><p>The crossover, the ridge point, is simply peak FLOP rate divided by bandwidth. For an <strong>H100 SXM at FP8</strong>, roughly 1,979 dense teraFLOPS over 3.35 terabytes per second of HBM3, the ridge sits near 591 FLOP per byte. </p><p>For the <strong>H200,</strong> identical compute over 4.8 terabytes per second of HBM3e, it falls to 412. For a B200, about 4,500 dense FP8 teraFLOPS over 8 terabytes per second, near 562 [Figure 2].</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!4W3S!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F20d78621-0dc3-4426-9930-56a99469e836_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!4W3S!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F20d78621-0dc3-4426-9930-56a99469e836_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!4W3S!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F20d78621-0dc3-4426-9930-56a99469e836_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!4W3S!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F20d78621-0dc3-4426-9930-56a99469e836_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!4W3S!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F20d78621-0dc3-4426-9930-56a99469e836_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!4W3S!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F20d78621-0dc3-4426-9930-56a99469e836_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/20d78621-0dc3-4426-9930-56a99469e836_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Figure 1&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 1" title="Figure 1" srcset="https://substackcdn.com/image/fetch/$s_!4W3S!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F20d78621-0dc3-4426-9930-56a99469e836_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!4W3S!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F20d78621-0dc3-4426-9930-56a99469e836_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!4W3S!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F20d78621-0dc3-4426-9930-56a99469e836_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!4W3S!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F20d78621-0dc3-4426-9930-56a99469e836_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 2. A roofline for three inference GPUs with the operating region of each phase marked. Prefill runs against the flat compute roof. Single-stream decode is pinned to the sloped bandwidth region. The same silicon is a different machine depending on which phase is running.</figcaption></figure></div><p>Prefill ingests the whole prompt at once. <strong>Every token attends to every prior token</strong>, the feed-forward layers process the full sequence in parallel, and the matrix multiplications are large and dense. </p><p>A prompt of a few thousand tokens pushes the arithmetic intensity into the <strong>hundreds or thousands of FLOP per byte</strong>, planting prefill firmly to the right of the ridge against the compute roof. Prefill is a FLOPS problem. It wants tensor cores and low precision and scales with raw matrix-multiply throughput.</p><p>Decode is the opposite animal. To generate one token it reads the entire weight set and the entire <strong>key-value cache</strong> for the sequence, performs a thin slice of computation, and emits a single token. For a single stream the arithmetic intensity sits near the floor, on the order of one to two FLOP per byte, pinning decode to the far left of the roofline on the bandwidth-limited slope. </p><p>Decode is a <strong>memory-bandwidth problem</strong>. It cannot keep the tensor cores fed; it cares only how fast the chip streams weights and cache out of HBM. The hard floor follows immediately: the fastest a single decode stream can run is bandwidth divided by bytes read per token, dominated by the weights [Figure 3]. </p><p>A <strong>70-billion-parameter model</strong> at FP16 is 140 gigabytes, so an H100 at 3.35 terabytes per second generates at most about 24 tokens per second on a single stream, an H200 about 34, a B200 about 57. Halve the precision to FP8 and every ceiling doubles. </p><p>This is the <strong>same calculation that bounds DeepSeek&#8217;s</strong> observed 20 to 22 tokens per second in production. Compute does not enter.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!bdZL!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9e9b0dbb-6acb-4bef-b5fd-b60a4b83f74b_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!bdZL!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9e9b0dbb-6acb-4bef-b5fd-b60a4b83f74b_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!bdZL!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9e9b0dbb-6acb-4bef-b5fd-b60a4b83f74b_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!bdZL!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9e9b0dbb-6acb-4bef-b5fd-b60a4b83f74b_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!bdZL!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9e9b0dbb-6acb-4bef-b5fd-b60a4b83f74b_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!bdZL!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9e9b0dbb-6acb-4bef-b5fd-b60a4b83f74b_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/9e9b0dbb-6acb-4bef-b5fd-b60a4b83f74b_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Figure 3&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 3" title="Figure 3" srcset="https://substackcdn.com/image/fetch/$s_!bdZL!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9e9b0dbb-6acb-4bef-b5fd-b60a4b83f74b_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!bdZL!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9e9b0dbb-6acb-4bef-b5fd-b60a4b83f74b_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!bdZL!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9e9b0dbb-6acb-4bef-b5fd-b60a4b83f74b_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!bdZL!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9e9b0dbb-6acb-4bef-b5fd-b60a4b83f74b_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 3. The single-stream decode ceiling is HBM bandwidth divided by model bytes. No quantity of compute changes it. Halving precision, which halves the weight bytes, is the only lever that moves the ceiling for a fixed model.</figcaption></figure></div><p>Batching is the escape, and its limit is the reason disaggregation exists. <strong>Decode many sequences </strong>at once and the weight read is shared across the batch: you stream the weights once and amortize them. </p><p>The arithmetic intensity of decode is therefore approximately twice the batch size divided by the bytes per weight, which at FP16 is<strong> approximately the batch size itself</strong> [Figure 4]. To cross the H100 FP8 ridge of 591 you need a batch in the high hundreds. </p><p>That is the entire game in decode, pack as many concurrent sequences into a step as memory allows, because every added sequence moves you rightward toward the compute roof and <strong>lifts tokens-per-second-per-GPU. </strong></p><p>Hold the two facts together: prefill wants to run immediately, in small groups, against the compute roof, to keep <strong>first-token latency low</strong>; decode wants to run in enormous batches against the bandwidth ceiling, to keep cost-per-token low. </p><p><strong>One phase is latency-shaped</strong> and compute-hungry, <strong>the other throughput-shaped</strong> and bandwidth-hungry, and for two years the industry asked one GPU under one scheduler to do both at once.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Sk9s!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb8e4a4ca-f093-4cae-bfe7-c043bc733d82_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Sk9s!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb8e4a4ca-f093-4cae-bfe7-c043bc733d82_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!Sk9s!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb8e4a4ca-f093-4cae-bfe7-c043bc733d82_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!Sk9s!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb8e4a4ca-f093-4cae-bfe7-c043bc733d82_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!Sk9s!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb8e4a4ca-f093-4cae-bfe7-c043bc733d82_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Sk9s!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb8e4a4ca-f093-4cae-bfe7-c043bc733d82_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b8e4a4ca-f093-4cae-bfe7-c043bc733d82_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Figure 2&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 2" title="Figure 2" srcset="https://substackcdn.com/image/fetch/$s_!Sk9s!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb8e4a4ca-f093-4cae-bfe7-c043bc733d82_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!Sk9s!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb8e4a4ca-f093-4cae-bfe7-c043bc733d82_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!Sk9s!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb8e4a4ca-f093-4cae-bfe7-c043bc733d82_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!Sk9s!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb8e4a4ca-f093-4cae-bfe7-c043bc733d82_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 4. Decode arithmetic intensity is essentially the batch size, because weights are read once per step and shared. Reaching the compute roof requires hundreds of concurrent sequences, which is why decode pools are built around the largest batches memory will hold.</figcaption></figure></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Why colocation lost</h2><p>Run both phases on one GPU under one<strong> continuous-batching scheduler</strong>, the architecture Orca introduced and vLLM popularized, and they fight. The fight has a precise mechanism. </p><p>Continuous batching keeps a rolling batch of decode steps running and folds in new requests as they arrive, but a <strong>new request cannot decode</strong> until its prompt is prefilled, and prefill is a heavy compute-bound operation that occupies the GPU far longer than a single decode step. </p><p>When a prefill lands in the batch, the system must either pause the <strong>in-flight decodes</strong> to prioritize it or batch the prefill alongside them, and both choices stall token generation for every active sequence. </p><p>The <strong>DistServe authors</strong> quantify the damage bluntly in their retrospective: even with chunked-prefill mitigation, a single large prefill can inflate time-per-output-token by a factor of two to thirty under bursty workloads. </p><p>A long prompt arriving at the wrong moment makes every other user&#8217;s stream stutter for a <strong>third of a second or more</strong> [Figure 5].</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!6Uw3!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9f51689f-e474-4199-a1d4-f970379ac6a8_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!6Uw3!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9f51689f-e474-4199-a1d4-f970379ac6a8_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!6Uw3!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9f51689f-e474-4199-a1d4-f970379ac6a8_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!6Uw3!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9f51689f-e474-4199-a1d4-f970379ac6a8_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!6Uw3!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9f51689f-e474-4199-a1d4-f970379ac6a8_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!6Uw3!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9f51689f-e474-4199-a1d4-f970379ac6a8_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/9f51689f-e474-4199-a1d4-f970379ac6a8_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Figure 11&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 11" title="Figure 11" srcset="https://substackcdn.com/image/fetch/$s_!6Uw3!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9f51689f-e474-4199-a1d4-f970379ac6a8_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!6Uw3!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9f51689f-e474-4199-a1d4-f970379ac6a8_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!6Uw3!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9f51689f-e474-4199-a1d4-f970379ac6a8_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!6Uw3!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9f51689f-e474-4199-a1d4-f970379ac6a8_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 5. The structural waste of colocation. A prefill-heavy step pins the tensor cores while bandwidth idles; a decode-heavy step pins bandwidth while the tensor cores idle. A shared GPU pays for both resources and presses hard on one at a time. Exact values are workload-dependent; the asymmetry is not.</figcaption></figure></div><p>The <strong>deeper cost is coupling</strong>. As the DistServe paper put it, colocation forces the resource allocator to provision for the worst case of both latency targets simultaneously, the <strong>tight first-token target</strong> and the tight per-token target, because the same GPUs serve both. </p><p>You cannot tune one phase without detuning the other, and you cannot scale one without scaling the other. The roofline says why the waste is structural and <strong>not a scheduling artifact</strong>: on a prefill-heavy step the tensor cores run near saturation while the memory bus idles, and on a decode-heavy step the reverse, so a <strong>colocated GPU pays rent</strong> on two expensive resources and uses roughly one of them at any instant.</p><p>The serving community fought this with chunked prefill, introduced in the <strong>Sarathi work by Amey Agrawal</strong> and co-authors, which breaks a long prefill into bounded chunks interleaved with ongoing decode so the peak disruption per step is capped and per-token latency smooths out. </p><p>Chunked prefill is the strongest argument against disaggregation and an honest teardown has to credit it: by mixing a<strong> compute-bound prefill chunk</strong> with bandwidth-bound decode work in one step it even improves utilization, running the two against different ceilings. But it does not dissolve the coupling. </p><p>The phases still share one parallelism strategy, one memory pool, <strong>one tensor-parallel degree</strong>, all of them a compromise. And it trades latencies against each other: the finer you chunk to <strong>protect per-token latency</strong>, the more you stretch first-token latency, because a long prompt now dribbles through the GPU in pieces. </p><p>You are back in the original bind. On <strong>shared hardware</strong> you can favor first-token latency or per-token latency, but you cannot independently optimize both.</p><p>What changed in 2025 was not the physics but the stakes. DistServe, the authors recount, met real pushback in 2024 because disaggregation demands a <strong>heavy refactor of existing serving systems</strong>, and saw little adoption that year. </p><p>Then businesses began deploying language models at competitive scale, and throughput stopped being the only metric that mattered. <strong>Latency became existential</strong>, because a chatbot that stutters loses users, and an agent that stalls breaks workflows. </p><p>At the same time models grew and traffic surged, forcing systems past hundreds and into thousands of GPUs, the regime where a disaggregated architecture genuinely shines because <strong>it can allocate resources to each phase</strong> independently and pair each with its own parallelism strategy. The technique that was ahead of its time in 2024 was exactly on time in 2025.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>The anatomy of the split</h2><p>Disaggregation assigns the phases to <strong>physically separate GPU pools</strong>. A request hits a prefill instance, which builds the key-value cache for the prompt and produces the first token; the cache is handed across to a decode instance, which loads it, folds the request into its large rolling batch, and generates the rest. </p><p>Prefill machines only prefill. Decode machines only decode. Three things become <strong>possible that colocation forbids</strong>, and they are the whole value proposition.</p><p>Interference disappears, because no prefill ever lands in the decode batch, so decode steps run uninterrupted and per-token latency stops spiking. <strong>Independent optimization becomes possible</strong>, because each pool can be tuned to its own roofline and scaled on its own axis, adding prefill capacity when prompts lengthen and decode capacity when outputs lengthen. </p><p>And <strong>phase-specific hardware</strong> becomes possible, the idea Splitwise pushed hardest, that since decode is bandwidth-bound and prefill compute-bound, the two should not run on identical chips at all, a thread that has since grown into an entire hardware category we will come to.</p><p>The crucial <strong>metric </strong>that makes all of this legible is <strong>goodput</strong>, and it is the metric most operators still fail to measure. Throughput is requests or tokens completed per second, full stop. </p><p><strong>Goodput is requests completed per second </strong>that meet their service-level objectives, both the first-token target and the per-token target. The distinction is the whole point, because a colocated system under load can post rising throughput while its goodput collapses, as interference pushes more and more requests past their latency targets even as tokens keep flowing [Figure 6]. </p><p><strong>Hao Zhang of UC San Diego</strong>, a DistServe author, frames it starkly in his lectures: a system can show ten requests per second of throughput while delivering three requests per second of goodput once the SLO is applied. The other seven finished late, which for an interactive product means they did not finish. </p><p>Disaggregation&#8217;s claim is a goodput claim. It does not necessarily move more tokens in the abstract; it moves more tokens that arrive on time, and the DistServe paper measured that as <strong>7.4 times more requests within SLO</strong>, or a 12.6 times tighter achievable latency target, against the colocated state of the art.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!5i_O!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F35b53939-6b20-4328-bce5-c51984c289d6_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!5i_O!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F35b53939-6b20-4328-bce5-c51984c289d6_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!5i_O!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F35b53939-6b20-4328-bce5-c51984c289d6_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!5i_O!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F35b53939-6b20-4328-bce5-c51984c289d6_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!5i_O!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F35b53939-6b20-4328-bce5-c51984c289d6_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!5i_O!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F35b53939-6b20-4328-bce5-c51984c289d6_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/35b53939-6b20-4328-bce5-c51984c289d6_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Figure 6&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 6" title="Figure 6" srcset="https://substackcdn.com/image/fetch/$s_!5i_O!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F35b53939-6b20-4328-bce5-c51984c289d6_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!5i_O!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F35b53939-6b20-4328-bce5-c51984c289d6_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!5i_O!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F35b53939-6b20-4328-bce5-c51984c289d6_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!5i_O!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F35b53939-6b20-4328-bce5-c51984c289d6_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 6. Throughput and goodput diverge under load. Raw tokens per second keep climbing while the count of requests that meet their latency targets falls away as interference mounts. Disaggregation&#8217;s win is the gap. Shape follows the DistServe goodput results; axes are illustrative.</figcaption></figure></div><p>The <strong>published uplifts</strong> that first announced the technique are worth stating with the asterisk each deserves [Figure 7]. </p><ul><li><p>DistServe reported 7.4 times more requests served within SLO against the colocated state of the art; </p></li><li><p>Splitwise, the parallel Microsoft Azure effort, 2.35 times the throughput at equal power and cost, or 1.4 times at 20 percent lower cost; </p></li><li><p>Mooncake, the platform behind Moonshot AI&#8217;s Kimi assistant, <strong>up to a 525 percent throughput gain</strong> in simulated overload and 75 percent more requests under real production traffic. </p></li></ul><p>The metrics differ, goodput in one, throughput at fixed power in another, raw request volume in a third, so they are not directly comparable, and each is a <strong>ceiling for a favorable workload</strong> rather than a constant anyone reproduces. </p><p>They are the numbers that turned a research idea into a procurement decision.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!bA3m!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc65e7bda-d66b-453b-9bb4-0aa66d4ec4ce_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!bA3m!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc65e7bda-d66b-453b-9bb4-0aa66d4ec4ce_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!bA3m!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc65e7bda-d66b-453b-9bb4-0aa66d4ec4ce_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!bA3m!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc65e7bda-d66b-453b-9bb4-0aa66d4ec4ce_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!bA3m!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc65e7bda-d66b-453b-9bb4-0aa66d4ec4ce_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!bA3m!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc65e7bda-d66b-453b-9bb4-0aa66d4ec4ce_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c65e7bda-d66b-453b-9bb4-0aa66d4ec4ce_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Figure 7&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 7" title="Figure 7" srcset="https://substackcdn.com/image/fetch/$s_!bA3m!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc65e7bda-d66b-453b-9bb4-0aa66d4ec4ce_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!bA3m!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc65e7bda-d66b-453b-9bb4-0aa66d4ec4ce_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!bA3m!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc65e7bda-d66b-453b-9bb4-0aa66d4ec4ce_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!bA3m!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc65e7bda-d66b-453b-9bb4-0aa66d4ec4ce_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 7. The headline multiples are real and each is a best case. The metrics are not commensurable across bars, and every figure is tied to a specific model, workload, and SLO. Read them as ceilings for favorable regimes, not a universal constant.</figcaption></figure></div><p>Modern orchestration adds one more lever on top of the split: cache-aware routing. <strong>NVIDIA Dynamo</strong>, the orchestration layer that sits above the inference engines, routes each request to the prefill worker whose resident cache best overlaps the incoming prompt, which NVIDIA reports can <strong>roughly halve first-token latency</strong> by avoiding redundant prefill of shared prefixes. </p><p>The router treats prefill and decode workers as first-class services, and a central planner continuously profiles the GPUs to autoscale and rebalance. This is the <strong>productized form of an idea DistServe </strong>seeded with a much simpler pull-based scheduler, which it used to keep decode workers from being flooded by spiky prefill bursts. </p><p>The <em>control plane</em> has become as much of the system as the data plane.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>The seam, and the rack that dissolved it</h2><p>The instant you split the phases you create a seam, and across it <strong>you must carry the key-value cache</strong> from the prefill instance that built it to the decode instance that consumes it, for every request. The cache is not a control message. It is the full attention state of the prompt, and it can be gigabytes.</p><p>Its size is set by the attention architecture, and the spread is enormous [Figure 8]. <strong>Multi-head attention stores a key and value vector per head</strong>, per layer, per token; a Llama-2-class 70B model with 64 heads, head dimension 128, and 80 layers holds about 2.5 megabytes per token at FP16, so a 1,000-token prompt carries roughly 2.6 gigabytes of state. </p><p><strong>Grouped-query attention</strong>, the Llama 3 scheme, shares key and value projections across groups and cuts the stored heads from 64 to 8, dropping the cache to about 0.31 gigabytes per thousand tokens. </p><p><strong>Multi-head latent attention</strong>, DeepSeek&#8217;s design, compresses the per-token state into a single latent vector of dimension 576 stored once rather than per head, landing near 0.07 gigabytes per thousand tokens. </p><p>Across these three schemes the seam payload spans a factor of about 37, which is the <strong>unglamorous reason MLA models are structurally cheaper to disaggregate</strong> and the reason attention-architecture choices are now serving-cost choices.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!X8Ol!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4469ab2e-86e4-4465-8d9b-9b1a8cb7288b_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!X8Ol!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4469ab2e-86e4-4465-8d9b-9b1a8cb7288b_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!X8Ol!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4469ab2e-86e4-4465-8d9b-9b1a8cb7288b_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!X8Ol!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4469ab2e-86e4-4465-8d9b-9b1a8cb7288b_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!X8Ol!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4469ab2e-86e4-4465-8d9b-9b1a8cb7288b_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!X8Ol!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4469ab2e-86e4-4465-8d9b-9b1a8cb7288b_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/4469ab2e-86e4-4465-8d9b-9b1a8cb7288b_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Figure 4&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 4" title="Figure 4" srcset="https://substackcdn.com/image/fetch/$s_!X8Ol!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4469ab2e-86e4-4465-8d9b-9b1a8cb7288b_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!X8Ol!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4469ab2e-86e4-4465-8d9b-9b1a8cb7288b_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!X8Ol!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4469ab2e-86e4-4465-8d9b-9b1a8cb7288b_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!X8Ol!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4469ab2e-86e4-4465-8d9b-9b1a8cb7288b_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 8. The seam payload is set by attention design. MHA ships a fraction of a gigabyte per token of context across the wire; MLA ships a small fraction of that. The bytes you must move are decided in the model architecture, not the serving stack.</figcaption></figure></div><p>Now the carry. Take the cache for a <strong>4,096-token prompt on a grouped-query 70B model</strong>, about 1.34 gigabytes, and move it across the interconnects an operator might have [Figure 9]. </p><p>On fifth-generation NVLink at 1.8 terabytes per second, 0.75 milliseconds. On <strong>InfiniBand NDR at 400 gigabits per second, 27 milliseconds.</strong> On 200-gigabit RoCE, 54. On commodity 100-gigabit Ethernet, 107. Across a 25-gigabit link between datacenters, 429. </p><p>The carry is a one-time handoff per request, so it adds to first-token latency rather than the per-token rate, and the right comparison is against the prefill it follows, <strong>roughly 290 to 600 milliseconds for that prompt</strong> on an H100. </p><p>Against that, NVLink and InfiniBand transfers are rounding error or close to it, <strong>Ethernet is a third of the prefill</strong> bolted onto every request, and cross-datacenter is fatal on its own. </p><p>The DistServe authors measured the intra-node case at under 0.1 percent of request latency over fast intra-node links, and every serious stack hides the transfer further with<strong> layer-wise streaming</strong>, shipping each layer&#8217;s cache the moment that layer finishes so its movement overlaps the computation of the next.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!tQlD!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4f16fdc2-d23d-4f31-abba-c29402b0d224_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!tQlD!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4f16fdc2-d23d-4f31-abba-c29402b0d224_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!tQlD!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4f16fdc2-d23d-4f31-abba-c29402b0d224_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!tQlD!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4f16fdc2-d23d-4f31-abba-c29402b0d224_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!tQlD!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4f16fdc2-d23d-4f31-abba-c29402b0d224_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!tQlD!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4f16fdc2-d23d-4f31-abba-c29402b0d224_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/4f16fdc2-d23d-4f31-abba-c29402b0d224_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Figure 5&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 5" title="Figure 5" srcset="https://substackcdn.com/image/fetch/$s_!tQlD!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4f16fdc2-d23d-4f31-abba-c29402b0d224_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!tQlD!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4f16fdc2-d23d-4f31-abba-c29402b0d224_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!tQlD!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4f16fdc2-d23d-4f31-abba-c29402b0d224_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!tQlD!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4f16fdc2-d23d-4f31-abba-c29402b0d224_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 9. The same 1.34 GB handoff across the interconnect hierarchy. Inside the rack it is noise; on InfiniBand it fits under the prefill; on commodity Ethernet it eats the budget; across datacenters it is fatal. The interconnect, not the GPU, decides whether disaggregation is viable.</figcaption></figure></div><p>This is where the <strong>current hardware generation changed the calculus</strong>, and it is the single most important development since the technique went mainstream. </p><p>The GB200 NVL72 and its successor the GB300 NVL72 connect 72 GPUs into <strong>one NVLink domain that behaves as a single massive GPU,</strong> with up to 130 terabytes per second of aggregate GPU-to-GPU bandwidth. </p><p>When prefill and decode pools live inside the same NVLink domain, the seam is no longer a network hop. It is a <strong>memory copy </strong>across a coherent fabric, sub-millisecond, effectively free. </p><p>The hard part of disaggregation, the part that made it a research problem for years, was moving the cache without blowing the latency budget, and the <strong>rack-scale NVLink domain </strong>makes that part vanish for deployments that fit inside a rack. This is a large reason disaggregation went from papers to production so fast: the hardware arrived to pay the seam&#8217;s bill. </p><p>The same fabric is why these systems can run the extremely wide expert parallelism we will see in the next section, which would be <strong>communication-bound on any slower interconnect</strong>.</p><p>The transfer machinery itself is now standardized infrastructure. NVIDIA&#8217;s NIXL, <strong>open-sourced at GTC 2025</strong>, unifies NVLink, InfiniBand, PCIe, and storage fabrics under one point-to-point abstraction and runs non-blocking so the GPU keeps computing while the cache moves. </p><p>Mooncake&#8217;s Transfer Engine, from the Kimi team, presents the same <strong>unified interface over TCP</strong>, RDMA, shared memory, and NVMe-over-fabrics. </p><p>LMCache, from the University of Chicago, accelerates the movement with <strong>batched transfers and I/O pipelining</strong> and decouples cache storage from the engine so the cache can persist, migrate, and be shared independently of execution; its authors report up to a tenfold reduction in first-token latency from KV reuse and offload. </p><p>DeepSeek built 3FS, a file system that pools thousands of SSDs and hundreds of storage nodes so any prefill can stash a cache that any decode can fetch in a locality-oblivious way. </p><p>And the cache, once it is a<strong> first-class managed object</strong>, can be retained and reused across requests, which is the quiet lever under the whole economy: of <strong>DeepSeek&#8217;s 608 billion daily input tokens</strong>, 342 billion, 56.3 percent, were served from cache rather than recomputed. </p><p>Prefix caching is for many workloads a larger cost lever than the prefill-decode split itself, and disaggregation is what makes the cache a managed resource in the first place.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!puFZ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F21597614-d3ca-439e-b18a-3c8e504cb975_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!puFZ!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F21597614-d3ca-439e-b18a-3c8e504cb975_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!puFZ!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F21597614-d3ca-439e-b18a-3c8e504cb975_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!puFZ!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F21597614-d3ca-439e-b18a-3c8e504cb975_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!puFZ!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F21597614-d3ca-439e-b18a-3c8e504cb975_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!puFZ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F21597614-d3ca-439e-b18a-3c8e504cb975_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/21597614-d3ca-439e-b18a-3c8e504cb975_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Figure 10&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 10" title="Figure 10" srcset="https://substackcdn.com/image/fetch/$s_!puFZ!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F21597614-d3ca-439e-b18a-3c8e504cb975_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!puFZ!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F21597614-d3ca-439e-b18a-3c8e504cb975_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!puFZ!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F21597614-d3ca-439e-b18a-3c8e504cb975_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!puFZ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F21597614-d3ca-439e-b18a-3c8e504cb975_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 10. The economics DeepSeek disclosed, and the cache that drives them. Left, the daily GPU cost against the theoretical revenue if every token were billed at R1 rates, the source of the much-quoted 545 percent margin, which the company itself flags as theoretical and which is materially lower in reality. Right, the token composition: more than half of all input tokens were served from cache, not recomputed. This is the disaggregated, MLA-based, heavily-cached architecture that lit the fuse on the whole repricing.</figcaption></figure></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>The kernel layer: expert parallelism and the all-to-all</h2><p>Here is the part most coverage skips, and the part that actually sets throughput on a <em>modern frontier model</em>. </p><p>Look under the hood of virtually any current frontier model, DeepSeek-V3, Kimi K2, Qwen3, Llama 4, and you find a <strong>sparse mixture-of-experts architecture</strong>: hundreds of expert sub-networks per layer, of which each token activates only a handful. </p><p>A 671-billion-parameter model like DeepSeek-V3 activates about 37 billion parameters per token. <strong>Sparsity is what makes these models cheap to run,</strong> but it imposes a specific and brutal communication pattern, and that pattern is where disaggregation, expert parallelism, and the interconnect all collide.</p><p>To serve a<strong> large MoE </strong>you distribute the experts across many GPUs, a scheme called expert parallelism, because no single GPU holds all of them. </p><p>But then every token has to travel to whichever GPUs hold its chosen experts and the results have to travel back, which means every layer performs<strong> two all-to-all communication operations per step</strong>, a dispatch that scatters tokens to their experts and a combine that gathers the results. </p><p>This is the dominant communication cost of MoE inference, and it has a vicious property: the messages are tiny. In <strong>DeepSeek-V3 </strong>each dispatch or combine message runs from about 7 kilobytes during inference to 256 kilobytes during training [Figure 11]. </p><p>General-purpose collective libraries like NCCL are tuned for the opposite regime, the large multi-megabyte all-reduce of dense training, where bandwidth dominates. <strong>At 7 kilobytes the per-message latency and synchronization overhead dominate</strong> instead, and NCCL leaves most of the wire idle.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!SwoG!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ab435cb-265c-4fa8-8ec7-fd4fe70cc402_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!SwoG!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ab435cb-265c-4fa8-8ec7-fd4fe70cc402_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!SwoG!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ab435cb-265c-4fa8-8ec7-fd4fe70cc402_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!SwoG!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ab435cb-265c-4fa8-8ec7-fd4fe70cc402_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!SwoG!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ab435cb-265c-4fa8-8ec7-fd4fe70cc402_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!SwoG!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ab435cb-265c-4fa8-8ec7-fd4fe70cc402_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/3ab435cb-265c-4fa8-8ec7-fd4fe70cc402_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Figure 14&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 14" title="Figure 14" srcset="https://substackcdn.com/image/fetch/$s_!SwoG!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ab435cb-265c-4fa8-8ec7-fd4fe70cc402_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!SwoG!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ab435cb-265c-4fa8-8ec7-fd4fe70cc402_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!SwoG!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ab435cb-265c-4fa8-8ec7-fd4fe70cc402_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!SwoG!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ab435cb-265c-4fa8-8ec7-fd4fe70cc402_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 11. Expert parallelism breaks the collective library. MoE dispatch and combine move tiny messages, where NCCL&#8217;s all-reduce kernels stall on latency and synchronization rather than saturating bandwidth. This mismatch is why DeepSeek wrote a dedicated communication library.</figcaption></figure></div><p><strong>That library is DeepEP</strong>, which DeepSeek open-sourced on the second day of its Open Source Week, and it is a small masterclass in why the kernel layer matters. </p><p>DeepEP provides custom all-to-all dispatch and combine kernels in two distinct flavors, and the split mirrors the prefill-decode split exactly. The <strong>high-throughput kernels serve prefill and training</strong>, maximizing raw bandwidth, but they emit dynamically shaped tensors that are incompatible with CUDA graphs. </p><p>The low-latency kernels serve decode, using direct RDMA to minimize latency and, critically, remaining CUDA-graph compatible so they avoid the <strong>kernel-launch overhead</strong> that dominates the decode phase, where each step is tiny and launch cost is a large fraction of the work. </p><p>The kernels need only about 20 streaming multiprocessors to saturate both the <strong>intra-node NVLink domain</strong> and the inter-node RDMA network simultaneously, freeing the rest of the GPU for computation. </p><p>They achieve this <strong>through NVSHMEM and IBGDA</strong>, which let the GPU issue RDMA operations directly to the network card without a round-trip through the CPU, and through asymmetric-domain forwarding that bridges the fast NVLink domain and the slower RDMA domain in one kernel. </p><p>The high-throughput path uses 24 queue pairs; the <strong>low-latency path uses 8 to 16</strong>, matched to the local expert count. A hook-based design overlaps the communication with computation without occupying compute units at all. </p><p>The newest DeepEP revisions add TMA-based transfers for minimal SM usage, support for the <strong>larger multi-node NVLink domains</strong> of the rack-scale systems, and a zero-SM remote-memory primitive for fetching KV cache directly from a peer.</p><p>This level of specialization is not optional at scale, and NVIDIA&#8217;s response confirms it: <strong>NVIDIA built HybridEP</strong>, its own token-based dispatch backend using the same hardware primitives, and co-developed a set of <strong>Blackwell kernels</strong> with the SGLang and vLLM projects through the FlashInfer library, covering attention prefill and decode, the communication path, the grouped matrix multiplications, the multi-node NVLink transfers, and MLA specifically. </p><p><strong>The all-to-all is the bottleneck</strong>, and the entire industry is now optimizing the same dozen kernels.</p><p>The payoff of getting this right is expert parallelism that goes very wide, and width is throughput [Figure 12]. </p><p>Spreading the experts of DeepSeek-R1 across 32 ways instead of 8 on a GB200 NVL72 lifts per-GPU output throughput by about 1.8 times, <strong>NVIDIA&#8217;s TensorRT-LLM measurements show</strong>, because a wider spread means each GPU holds fewer experts and therefore loads less weight per step, and because it fills the grouped matrix multiplications more completely. </p><p>Wide expert parallelism is only viable because the <strong>130-terabyte-per-second NVLink domain absorbs the all-to-all traffic</strong> that wider spreading generates; on a slower fabric the communication would swamp the gain. And the two phases run deliberately different widths. </p><p>DeepSeek&#8217;s disclosed production configuration runs prefill at expert-<strong>parallel degree 32 and decode at degree 144</strong>, a decode pool nearly five times wider than the prefill pool, because decode is where the wide spread pays off in throughput and where the batch is large enough to keep all those experts busy. </p><p>The DistServe authors, surveying the same system, note <strong>decode configurations reaching toward 256-way expert parallelism</strong>, and newer stacks push wider still. </p><p>The prefill-decode asymmetry that began as a latency argument has become a parallelism argument: the phases want not just different hardware but <strong>different distributed-systems topologies</strong> entirely.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!4Dxx!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F68c853f7-9b06-485e-bdbe-02348dbc7188_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!4Dxx!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F68c853f7-9b06-485e-bdbe-02348dbc7188_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!4Dxx!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F68c853f7-9b06-485e-bdbe-02348dbc7188_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!4Dxx!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F68c853f7-9b06-485e-bdbe-02348dbc7188_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!4Dxx!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F68c853f7-9b06-485e-bdbe-02348dbc7188_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!4Dxx!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F68c853f7-9b06-485e-bdbe-02348dbc7188_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/68c853f7-9b06-485e-bdbe-02348dbc7188_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Figure 16&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 16" title="Figure 16" srcset="https://substackcdn.com/image/fetch/$s_!4Dxx!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F68c853f7-9b06-485e-bdbe-02348dbc7188_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!4Dxx!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F68c853f7-9b06-485e-bdbe-02348dbc7188_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!4Dxx!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F68c853f7-9b06-485e-bdbe-02348dbc7188_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!4Dxx!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F68c853f7-9b06-485e-bdbe-02348dbc7188_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 12. Decode wants the widest expert parallelism the fabric will allow. Spreading experts to EP32 instead of EP8 lifts per-GPU decode throughput by about 1.8 times by shrinking the per-GPU weight load and filling the grouped matrix multiplications. The two phases run deliberately different degrees.</figcaption></figure></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>The attention rewrite is shrinking the problem</h2><p>While the serving stack was learning to split and spread, the <strong>model architects were attacking the problem from underneath</strong>, and the attack lands squarely on the two costs this issue is about: the bytes at the seam and the compute in prefill.</p><p>Multi-head latent attention was the first cut, and we have already seen its effect on the seam, a roughly <strong>thirty-seven-fold reduction</strong> in cache bytes versus dense multi-head attention. </p><p>But MLA leaves the other cost untouched: prefill attention still scales quadratically with sequence length, because every token still attends to every prior token, and as <strong>context windows</strong> stretch toward a million tokens that quadratic term dominates the prefill bill. </p><p>This is the cost that DeepSeek attacked in V3.2 with DeepSeek Sparse Attention, and the mechanism is elegant [Figure 13]. A <strong>lightweight neural network DeepSeek calls the Lightning Indexer</strong> scores the relevance of past key blocks to the current query and selects only the top few thousand most relevant, and the expensive attention computation then runs only over that selected set. </p><p>The pattern is retrieve-then-attend, and it bends the cost curve from quadratic in sequence length toward roughly linear past the selection cutoff, while DeepSeek reports <strong>model quality virtually unchanged. </strong></p><p>At a million tokens of context the difference is on the order of a few hundredfold less attention work in prefill.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!xCLo!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b1b012a-64e9-4114-9d13-218c7ff7a37a_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!xCLo!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b1b012a-64e9-4114-9d13-218c7ff7a37a_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!xCLo!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b1b012a-64e9-4114-9d13-218c7ff7a37a_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!xCLo!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b1b012a-64e9-4114-9d13-218c7ff7a37a_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!xCLo!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b1b012a-64e9-4114-9d13-218c7ff7a37a_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!xCLo!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b1b012a-64e9-4114-9d13-218c7ff7a37a_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/1b1b012a-64e9-4114-9d13-218c7ff7a37a_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Figure 18&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 18" title="Figure 18" srcset="https://substackcdn.com/image/fetch/$s_!xCLo!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b1b012a-64e9-4114-9d13-218c7ff7a37a_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!xCLo!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b1b012a-64e9-4114-9d13-218c7ff7a37a_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!xCLo!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b1b012a-64e9-4114-9d13-218c7ff7a37a_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!xCLo!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b1b012a-64e9-4114-9d13-218c7ff7a37a_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 13. Sparse attention attacks the prefill cost itself. MLA cut the KV bytes crossing the seam; DeepSeek Sparse Attention then cut the prefill compute, selecting a fixed set of relevant key blocks before attending and bending quadratic attention toward linear at long context. Illustrative scaling for a fixed selection budget.</figcaption></figure></div><p>This matters for disaggregation in a way that compounds. Sparse attention shrinks both the prefill compute and, in its variants, the cache that must be carried, which <strong>shifts the prefill-decode balance again</strong> and makes long-context disaggregation economically viable at lengths that would have been hopeless under dense attention. </p><p>It is also the live frontier as of this writing. DeepSeek-V3.2 shipped in late 2025 as a <strong>671-billion-parameter MLA-plus-MoE model with DSA layered on top</strong>, reaching reasoning quality the company benchmarks against GPT-5, and its Speciale variant took gold-medal scores at the 2025 International Mathematical Olympiad and the ICPC World Finals. </p><p><strong>DeepSeek-V4</strong>, released in April 2026, extends the context window to a million tokens and replaces DSA with a successor called Compressed Sparse Attention, with its Pro variant reported at<strong> 80.6 percent on SWE-bench. </strong></p><p>The trajectory is unmistakable: the attention mechanism is being rebuilt around the economics of <strong>long-context inference</strong>, and each rebuild changes the numbers in the serving stack beneath it. Architecture and serving are no longer separable disciplines.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>The benchmark reality, and where it breaks</h2><p>For two years the inference market argued over performance with vendor slides, which is no way to allocate billions in capital. </p><p>That changed in late 2025 when SemiAnalysis launched InferenceMAX, the <strong>first independent open-source benchmark</strong> to measure not raw throughput but total cost of compute across real models and real interactivity targets, running DeepSeek R1, GPT-OSS, Llama 3, and Qwen across the GB200 NVL72, B200, H200, H100, and AMD&#8217;s MI300X, MI325X, and MI355X, with Google TPU and AWS Trainium backends following. </p><p>It is the<strong> closest thing the field has to a neutral scoreboard</strong>, and the picture it paints is one a chip-level analysis would miss entirely [Figure 13].</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!n0Rh!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F20458f4f-3dc8-4649-9bf4-113297b5a519_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!n0Rh!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F20458f4f-3dc8-4649-9bf4-113297b5a519_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!n0Rh!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F20458f4f-3dc8-4649-9bf4-113297b5a519_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!n0Rh!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F20458f4f-3dc8-4649-9bf4-113297b5a519_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!n0Rh!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F20458f4f-3dc8-4649-9bf4-113297b5a519_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!n0Rh!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F20458f4f-3dc8-4649-9bf4-113297b5a519_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/20458f4f-3dc8-4649-9bf4-113297b5a519_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Figure 13&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 13" title="Figure 13" srcset="https://substackcdn.com/image/fetch/$s_!n0Rh!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F20458f4f-3dc8-4649-9bf4-113297b5a519_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!n0Rh!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F20458f4f-3dc8-4649-9bf4-113297b5a519_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!n0Rh!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F20458f4f-3dc8-4649-9bf4-113297b5a519_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!n0Rh!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F20458f4f-3dc8-4649-9bf4-113297b5a519_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 13. Per-GPU throughput on DeepSeek R1 at a fixed 25 tokens-per-second-per-user interactivity, normalized to H200. The GB200 NVL72&#8217;s roughly tenfold lead over the H200, and threefold over the B200, comes from the rack-scale NVLink domain, not from a faster individual chip. Values from SemiAnalysis InferenceMAX; the advantage is interactivity-dependent.</figcaption></figure></div><p>At a fixed interactivity of 25 tokens per second per user on DeepSeek R1, the GB200 NVL72 delivers roughly ten times the per-GPU throughput of an H200 and <strong>about three times that of a standalone B200</strong>, even though the B200 is a faster chip in isolation. </p><p>The <strong>advantage is the 72-GPU NVLink domain</strong>, which lets the rack run wider parallelism and larger coherent batches than any single eight-GPU node can. </p><p>On absolute terms the B200 reaches 60,000 tokens per second per GPU at 1,000 tokens per second per user on GPT-OSS, and software optimization alone drove its cost on that model to <strong>two cents per million tokens,</strong> a fivefold reduction in two months. </p><p>NVIDIA frames the rack-level economics as a five-million-dollar GB200 NVL72 generating <strong>75 million dollars in token revenue</strong>, a fifteen-fold return, though that figure prices output at favorable rates and should be read as a vendor&#8217;s best case rather than a realized margin. </p><p>The energy picture is cleaner and third-party: on DeepSeek R1 the GB200 NVL72 delivers roughly eight times the tokens per provisioned megawatt of a single-node H200, and <strong>Blackwell runs about 20 percent more energy-efficient than AMD&#8217;s CDNA4</strong> on <strong>GPT-OSS</strong>, partly because the MI355X draws 1.4 kilowatts per GPU against the B200&#8217;s 1 kilowatt.</p><p>Now the part the headline numbers omit, and the part this publication exists to surface. The rack-scale advantage is not a constant. It is a function of interactivity, and it expires [Figure 14]. </p><p>At 60 tokens per second per user the GB200 NVL72 produces a little less than triple a B200&#8217;s per-GPU throughput, but as the interactivity target rises the batch that the rack can assemble shrinks, and<strong> by around 130 tokens per second per user</strong> the workload fits inside a single eight-GPU node&#8217;s NVLink domain, at which point the NVL72&#8217;s scale-out advantage disappears entirely and it becomes more expensive per token than a standalone node. </p><p>The whole case for the rack rests on serving many users at moderate interactivity, the <strong>chatbot and agent regime</strong>; push to extreme single-user speed and the economics invert. </p><p>The benchmark also exposes a software truth that no spec sheet shows: <strong>AMD&#8217;s MI355X</strong> is competitive with the B200 on FP8 disaggregated prefill, but its disaggregated performance actually degrades at higher interactivity because <strong>the ROCm stack lacks the kernel and collective optimizations needed</strong> to compose multiple state-of-the-art techniques together. </p><p>Disaggregation is not a hardware capability you buy; it is a software capability you accumulate, and the gap between vendors is measured in kernels.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!9XYZ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac224577-1715-4906-a307-d1a5974f1d0c_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!9XYZ!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac224577-1715-4906-a307-d1a5974f1d0c_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!9XYZ!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac224577-1715-4906-a307-d1a5974f1d0c_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!9XYZ!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac224577-1715-4906-a307-d1a5974f1d0c_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!9XYZ!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac224577-1715-4906-a307-d1a5974f1d0c_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!9XYZ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac224577-1715-4906-a307-d1a5974f1d0c_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ac224577-1715-4906-a307-d1a5974f1d0c_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Figure 17&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 17" title="Figure 17" srcset="https://substackcdn.com/image/fetch/$s_!9XYZ!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac224577-1715-4906-a307-d1a5974f1d0c_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!9XYZ!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac224577-1715-4906-a307-d1a5974f1d0c_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!9XYZ!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac224577-1715-4906-a307-d1a5974f1d0c_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!9XYZ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac224577-1715-4906-a307-d1a5974f1d0c_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 14. The rack-scale win has an expiry date. At moderate interactivity the NVL72 roughly triples a single node per GPU; push interactivity high enough that the batch shrinks into one eight-GPU node, and the advantage evaporates and inverts on cost. Shape after SemiAnalysis InferenceX v2.</figcaption></figure></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>The dial you must keep turning</h2><p>Disaggregation hands you an operational problem that the headline numbers never mention, and every team that has run it in production knows it immediately. </p><p><strong>Once the fleet is split into a prefill pool and a decode pool</strong>, you have to choose the ratio between them, and you will get it wrong, because the right answer keeps moving [Figure 15].</p><p>The problem is a producer-consumer imbalance: prefill instances produce cache that <strong>decode instances consume</strong>, and the production rate rarely matches the consumption rate. Provision too few prefill instances and prompts queue while first-token latency slips and the decode pool sits half-idle for want of work. </p><p>Provision too few decode instances and the prefill pool races ahead while <strong>per-token latency slips </strong>and the prefill pool sits half-idle. Either error strands capital on GPUs that cannot do useful work because the other pool is the bottleneck. </p><p>DeepSeek&#8217;s disclosed answer was a fixed three-to-nine ratio, three prefill nodes feeding nine decode nodes, and the SGLang team reproduced a similar split on 96 H100s,<strong> twenty-four GPUs for prefill and seventy-two for decode, reaching 52,300 input tokens and 22,300 output tokens per second per node</strong>, the first open implementation to match DeepSeek&#8217;s own reported numbers, at a cost they put at twenty cents per million output tokens, about one-fifth the official API price.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!8n0Z!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa96d9635-1cf2-43b5-93fc-8791ad9e0990_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!8n0Z!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa96d9635-1cf2-43b5-93fc-8791ad9e0990_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!8n0Z!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa96d9635-1cf2-43b5-93fc-8791ad9e0990_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!8n0Z!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa96d9635-1cf2-43b5-93fc-8791ad9e0990_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!8n0Z!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa96d9635-1cf2-43b5-93fc-8791ad9e0990_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!8n0Z!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa96d9635-1cf2-43b5-93fc-8791ad9e0990_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a96d9635-1cf2-43b5-93fc-8791ad9e0990_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Figure 8&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 8" title="Figure 8" srcset="https://substackcdn.com/image/fetch/$s_!8n0Z!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa96d9635-1cf2-43b5-93fc-8791ad9e0990_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!8n0Z!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa96d9635-1cf2-43b5-93fc-8791ad9e0990_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!8n0Z!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa96d9635-1cf2-43b5-93fc-8791ad9e0990_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!8n0Z!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa96d9635-1cf2-43b5-93fc-8791ad9e0990_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 15. Disaggregation hands you a ratio you must keep correct. Too few prefill instances and first-token latency slips; too few decode instances and per-token latency slips. The optimum sits in a narrow band and drifts with every shift in prompt and output length. Schematic of the producer-consumer balance.</figcaption></figure></div><p>What makes this hard rather than a one-time sizing exercise is that the optimal ratio is not constant. It <strong>depends on the shape of the traffic</strong>, the ratio of input length to output length, and that shape shifts by the hour and by the product surface. </p><p>A wave of document-summarization requests is prefill-heavy and wants more prefill capacity; a <strong>wave of long-form generation</strong> is decode-heavy and wants more decode; a split that is optimal at noon is wrong by midnight. </p><p>This is why the serious systems have moved to dynamic rebalancing, monitoring load in real time and shifting the ratio, and <strong>why a system like TaiChi goes further</strong> and switches whole instances between disaggregated and colocated modes depending on which yields better goodput at the current load. </p><p>The existence of that last capability is the tell: colocation is not always wrong, and a<strong> system smart enough</strong> to know when to disaggregate is smart enough to know when to stop. </p><p>Disaggregation converts a <strong>hardware-utilization problem</strong> into a scheduling-and-capacity problem, which is usually a good trade because software is cheaper to change than silicon, but <strong>it is a trade and not a free win</strong>, and a team that splits without building the rebalancing machinery will frequently lose to a well-tuned colocated deployment with chunked prefill.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>The split spawned silicon</h2><p>The most striking evidence that disaggregation has become foundational is not in any serving framework. </p><p><strong>It is in the silicon roadmap</strong>, because once the phases run on separate machines, the machines stop needing to be the same machine, and the hardware vendors have noticed [Figure 16].</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!hLgW!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffae74696-d771-4049-b014-6f00837861b6_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!hLgW!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffae74696-d771-4049-b014-6f00837861b6_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!hLgW!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffae74696-d771-4049-b014-6f00837861b6_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!hLgW!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffae74696-d771-4049-b014-6f00837861b6_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!hLgW!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffae74696-d771-4049-b014-6f00837861b6_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!hLgW!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffae74696-d771-4049-b014-6f00837861b6_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/fae74696-d771-4049-b014-6f00837861b6_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Figure 15&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 15" title="Figure 15" srcset="https://substackcdn.com/image/fetch/$s_!hLgW!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffae74696-d771-4049-b014-6f00837861b6_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!hLgW!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffae74696-d771-4049-b014-6f00837861b6_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!hLgW!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffae74696-d771-4049-b014-6f00837861b6_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!hLgW!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffae74696-d771-4049-b014-6f00837861b6_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 16. The split spawned silicon. Plotting accelerators by compute against memory bandwidth, vendors are now building to the corners: FLOPS-rich, cheap-memory parts for prefill, and bandwidth-rich parts for decode. Compute is a low-precision proxy across mixed formats; the positioning, not exact parity, is the point.</figcaption></figure></div><p>NVIDIA&#8217;s clearest statement is the <strong>Rubin CPX</strong>, announced in September 2025 and shipping at the end of 2026, a GPU built exclusively for the prefill and context phase. </p><p>Its logic is the wrong-sizing problem stated as a product: prefill is compute-bound and barely touches memory bandwidth, so dedicating <strong>expensive high-bandwidth HBM</strong> to it wastes the most expensive component on the chip. </p><p>The CPX therefore pairs 30 petaFLOPS of NVFP4 compute with <strong>128 gigabytes of GDDR7</strong>, a memory that SemiAnalysis estimates is roughly five times more cost-effective per byte than HBM and runs at perhaps a quarter of HBM&#8217;s bandwidth, which is fine because prefill does not need the bandwidth. </p><p>It adds dedicated attention hardware delivering about three times the <strong>attention throughput of a GB300 NVL72</strong>, aimed directly at million-token context, and it ships with PCIe but no NVLink, because it is built for disaggregated inference racks rather than tightly coupled training clusters. </p><p>The packaging makes the intent explicit: the <strong>Vera Rubin NVL144 CPX rack pairs 144 CPX prefill GPUs with 144 standard Rubin decode GPUs and 36 CPUs,</strong> prefill and decode silicon racked side by side, the disaggregation thesis cast in metal.</p><p>The decode side is bifurcating too, and along a more radical axis. In December 2025 NVIDIA signed a<strong> roughly 20-billion-dollar licensing arrangement with Groq</strong>, whose language-processing units abandon HBM entirely for on-chip SRAM, 500 megabytes per chip at 150 terabytes per second, an order of magnitude past any HBM part. </p><p>For autoregressive decode at long context, where the entire bottleneck is reading state out of memory, <strong>SRAM&#8217;s bandwidth is decisive</strong>: a 70-billion-parameter decode at 128,000 tokens of context runs dramatically faster in memory-access time on an SRAM part than on an HBM GPU. </p><p>The emerging architectural division has GPUs handling training and prefill while <strong>specialized low-latency parts handle decode</strong>, which is prefill-decode disaggregation pushed all the way down to the level of distinct chip families. </p><p>And the trend is not NVIDIA&#8217;s alone. The DistServe authors report that Huawei, Enflame, MetaX, and Biren are <strong>all prototyping or deploying decode-specialized or attention-optimized accelerators</strong> built on exactly this philosophy. </p><p>A systems technique conceived to tame latency on homogeneous GPUs is now redrawing the boundaries of the accelerator market itself.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>When it still doesn&#8217;t pay</h2><p>Disaggregation is the substrate now, but it is not a universal good, and the honest boundaries matter as much as the wins [Figure 17]. </p><p>SemiAnalysis, <strong>evaluating the Rubin CPX&#8217;s full hardware-level disaggregation</strong>, put the caveat precisely: complete disaggregation delivers excellent results only under certain ratios of input to output length and for long decode lengths, with other scenarios seeing underwhelming benefits. </p><p>The structure of the advantage explains the boundary. <strong>Disaggregation&#8217;s benefit grows with output length</strong>, because longer outputs mean more decode steps to protect from interference, and with offered load, because heavier load means more interference to remove. </p><p>In the opposite corner, short outputs under light traffic, there is <strong>little interference to eliminate</strong> and too little decoding for the protection to accrue, and the seam and the ratio overhead are pure cost. A well-tuned colocated deployment with chunked prefill wins there.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!tLwj!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b02adfd-3132-468b-9dfe-c9073e3bec3b_1634x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!tLwj!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b02adfd-3132-468b-9dfe-c9073e3bec3b_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!tLwj!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b02adfd-3132-468b-9dfe-c9073e3bec3b_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!tLwj!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b02adfd-3132-468b-9dfe-c9073e3bec3b_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!tLwj!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b02adfd-3132-468b-9dfe-c9073e3bec3b_1634x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!tLwj!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b02adfd-3132-468b-9dfe-c9073e3bec3b_1634x893.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0b02adfd-3132-468b-9dfe-c9073e3bec3b_1634x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Figure 9&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Figure 9" title="Figure 9" srcset="https://substackcdn.com/image/fetch/$s_!tLwj!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b02adfd-3132-468b-9dfe-c9073e3bec3b_1634x893.png 424w, https://substackcdn.com/image/fetch/$s_!tLwj!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b02adfd-3132-468b-9dfe-c9073e3bec3b_1634x893.png 848w, https://substackcdn.com/image/fetch/$s_!tLwj!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b02adfd-3132-468b-9dfe-c9073e3bec3b_1634x893.png 1272w, https://substackcdn.com/image/fetch/$s_!tLwj!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b02adfd-3132-468b-9dfe-c9073e3bec3b_1634x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 17. The decision is a regime. Disaggregation pays when outputs are long and load is heavy, and goes underwater for short replies under light load. The map is an illustrative model encoding the consistent direction of the evidence, not a measured surface. Characterize your own traffic before committing.</figcaption></figure></div><p>Three caveats compound this. The wrong-sizing problem never fully disappears even with disaggregation, because a <strong>pure prefill instance on an HBM part still underutilizes its memory bandwidth</strong>, which is the entire reason the Rubin CPX exists. </p><p>The interactivity cliff from the benchmark section means that even where disaggregation and <strong>rack-scale hardware win</strong> at moderate interactivity, the advantage inverts at extreme single-user speed, when the batch collapses into a single node. </p><p>And the software-maturity tax, visible in AMD&#8217;s degraded high-interactivity disaggregation, means the gains are not portable across stacks; they <strong>have to be earned kernel by kernel. </strong>The strategic read is not disaggregate or do not disaggregate. </p><p>It is characterize your traffic on the two axes that matter, <strong>real output-length distribution</strong> and real peak-to-trough load, then build for your regime, and buy the interconnect before you buy the split, because on a slow fabric the seam eats the gain and on a fast one it disappears.</p><div><hr></div><h2>The shake, settled</h2><p><strong>Return to the 600 billion dollars</strong>. The market&#8217;s first instinct, that cheaper inference means less demand for the chips that serve it, has been falsified about as cleanly as a macro thesis ever is. </p><p>The <strong>cost collapse was real</strong>, from thousands of dollars per benchmark task to about eleven, and demand did not fall; it ran so far past the efficiency gains that the Peterson Institute concluded <strong>usage had dwarfed them</strong>, the Jevons paradox playing out in real time. </p><p>The technique that frightened the market, efficient inference built on disaggregation and sparsity and compressed attention, <strong>did not shrink the industry</strong>. It enlarged it, by making applications viable that were uneconomical at the old prices.</p><p>But there is a second-order effect the bullish reading often misses, and it is where the real consequence sits. <strong>DeepSeek did not just demonstrate cheap inference</strong>; it open-sourced the means of production. DeepEP, 3FS, the full reference architecture, all released to anyone. </p><p><strong>NVIDIA open-sourced Dynamo and NIXL</strong>. Mooncake, llm-d, LMCache, SGLang, and vLLM are all open. The result is that the disaggregated serving stack, the thing that lets you run a frontier model at a fraction of the naive cost, is no longer a moat. </p><p>It is a commodity any competent team can deploy, which is exactly why the <strong>SGLang reproduction could serve DeepSeek at one-fifth the official API price</strong>. The value did not disappear; it moved. </p><p>SemiAnalysis frames the shift as a move from raw FLOPS per chip to total intelligence per dollar at rack scale, and the InferenceMAX results bear it out: <strong>the GB200 NVL72 wins not because its chips are faster but because 72 of them act as one</strong>, and the orchestration software across them is as much of the product as the silicon. </p><p>The moat migrated from the model and the kernel, which are now shared, to the rack-scale system integration and the interconnect, which are hard to replicate and hard to buy.</p><p>The frontier from here is the generalization of the same idea to the next seam. The DistServe authors point to <strong>attention-FFN disaggregation as the natural successor</strong>: within decode, attention is memory-bound and hungry for KV-cache bandwidth while the feed-forward layers are compute-bound and hungry for <strong>weight storage,</strong> so splitting them onto tailored hardware lets each reach high utilization independently, and the same logic that justified the prefill-decode split applies one level down. </p><p>For dense models this was long considered impractical because it doubles the activation transfer per layer, but the <strong>MoE models that now dominate already perform two all-to-all operations per decode step</strong>, so the attention-FFN split can be folded into the communication pattern that already exists, making its extra transfer nearly free. </p><p><strong>MegaScale-Infer and StepFun&#8217;s Step-3 have already demonstrated it</strong> on large MoE models. The pattern is always the same: find a boundary in the computation where the two sides want different hardware or different parallelism, split there, and pay a transfer cost at the new seam in exchange for independent optimization on each side. </p><p>The question every such split raises is the question this entire issue has circled. <em>Is the thing you carry across the new seam small enough, and the wire fast enough, that the split pays?</em></p><p>DeepSeek published a margin and <strong>the market saw a number. </strong>The number was a consequence. </p><p>The cause was an architecture that took two workloads with opposite appetites and stopped forcing them to share a plate, and within eighteen months that<strong> architecture became the floor that everything else is built on</strong>, dragged the hardware roadmap behind it, and survived a 600-billion-dollar referendum on whether efficiency was a threat or an accelerant. The split won. </p><p>What remains, and what <strong>will keep deciding the economics as the seams multiply and the models keep rewriting themselves underneath</strong>, is the same thing it always was: how honestly the system respects the shape of the work, and how fast the wire is at the seam.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><p></p>]]></content:encoded></item><item><title><![CDATA[We wrote the CUDA reference we could not find]]></title><description><![CDATA[The handbook we wanted on our own desks. GPU performance from first principles, current to Blackwell and CUDA 13.x.]]></description><link>https://www.thesoftwarefrontier.com/p/we-wrote-the-cuda-reference-we-could</link><guid isPermaLink="false">https://www.thesoftwarefrontier.com/p/we-wrote-the-cuda-reference-we-could</guid><dc:creator><![CDATA[Lorenzo Bradanini]]></dc:creator><pubDate>Sat, 20 Jun 2026 08:35:38 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!SAY7!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F54550d86-2756-4131-8818-956604f6749d_608x608.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>We have both <strong>lost more hours than we want </strong>to admit to questions that should have a clean answer. Here is one.</p><p>On Hopper, a WGMMA sources its operands from shared memory and registers. On Blackwell, the new UMMA instruction can read its A operand straight from tensor memory, a level of the hierarchy that did not exist a generation ago. </p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Software Frontier is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>That single change reshapes how you build a pipeline. When we went looking for it written down in one place, it was not there. It was scattered across a PTX spec, a CUTLASS example, and a couple of microbenchmark threads.</p><p>That was the pattern everywhere. Occupancy rules in one tuning guide. Descriptor encodings in another. The genealogy of the tensor cores living in our heads after weeks of trial runs. Nothing sat in one place we could read front to back.</p><p><strong>So we wrote that place.</strong></p><p>We are not technical writers who picked up CUDA for a book. We write low-level GPU and systems code for a living, and we got tired of not having this on our own desks.</p><p><strong>CUDA Mastery 2026</strong> is a deep technical handbook for engineers who want to understand and optimize CUDA on modern NVIDIA GPUs, from the fundamentals through Hopper, Blackwell, and Blackwell Ultra. It is current to CUDA 13.x and the fifth generation of tensor cores.</p><p>It is<strong> not a syntax tutorial.</strong> Most CUDA resources teach you syntax, a few explain how the hardware actually works, and almost none teach you to reason about performance before you write a line. This is the one that does.</p><p>Here is what it covers, as one integrated reference instead of forty open browser tabs: SM internals, warp scheduling, scoreboarding, occupancy, and latency hiding. The full memory hierarchy, coalescing, shared memory, TMA, and Tensor Memory. </p><p>The tensor core path from WMMA to WGMMA to UMMA. Roofline modeling, bottleneck analysis, and Nsight profiling. CUTLASS, CuTe, CUDA Tile, PTX, SASS, and JIT. Multi-GPU with NCCL, NVLink, NVSwitch, and NVSHMEM. And real kernel walkthroughs: SGEMM, Flash Attention, reductions, scan, sort, and sparse ops.</p><p>Two rules we held ourselves to. Every specification is footnoted to a primary source, whether an NVIDIA whitepaper, the PTX ISA, a tuning guide, or published microbenchmarks, with nothing paraphrased from a blog you cannot trace. </p><p>And <strong>specification stays separate from measurement</strong>, so you always know whether a claim comes from a document or from a number someone ran.</p><p>Think about what one wrong optimization actually costs: a day chasing a bottleneck that was never there, or a kernel rewritten three times because an operand sat in the wrong memory. This reference shortens that loop. The next time you open Nsight and see a stall, you will know which part of the machine is talking to you.</p><p>It is <strong>written for engineers who already ship CUDA.</strong> If you have never launched a kernel, start somewhere gentler and come back when you want to go deep.</p><p><strong>The book is &#8364;89</strong>, a single payment, or 10 installments, and less than the going hourly rate for someone who already knows this. Every future v1.x update and correction is included free.</p><p><strong>We built this because we wanted it on our own desks</strong>. It is on yours now if you want it. </p><p><a href="https://lorenzobrada.gumroad.com/l/cuda_mastery">Get CUDA Mastery 2026</a></p><h4>Thank you in advance for your extremely precious support. </h4><p></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Software Frontier is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[The Kill Switch Was in the Mail ]]></title><description><![CDATA[How a single Friday letter switched off the world's most capable AI model, turned Anthropic's safety moat into its munition, and marked the second federal strike on the]]></description><link>https://www.thesoftwarefrontier.com/p/the-kill-switch-was-in-the-mail</link><guid isPermaLink="false">https://www.thesoftwarefrontier.com/p/the-kill-switch-was-in-the-mail</guid><dc:creator><![CDATA[Lorenzo Bradanini]]></dc:creator><pubDate>Fri, 19 Jun 2026 13:30:29 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!NMlq!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5307dade-240f-4620-a799-3a8ff09e1f63_1254x1254.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!NMlq!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5307dade-240f-4620-a799-3a8ff09e1f63_1254x1254.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!NMlq!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5307dade-240f-4620-a799-3a8ff09e1f63_1254x1254.png 424w, https://substackcdn.com/image/fetch/$s_!NMlq!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5307dade-240f-4620-a799-3a8ff09e1f63_1254x1254.png 848w, https://substackcdn.com/image/fetch/$s_!NMlq!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5307dade-240f-4620-a799-3a8ff09e1f63_1254x1254.png 1272w, https://substackcdn.com/image/fetch/$s_!NMlq!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5307dade-240f-4620-a799-3a8ff09e1f63_1254x1254.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!NMlq!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5307dade-240f-4620-a799-3a8ff09e1f63_1254x1254.png" width="1254" height="1254" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/5307dade-240f-4620-a799-3a8ff09e1f63_1254x1254.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1254,&quot;width&quot;:1254,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2459135,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/198379004?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5307dade-240f-4620-a799-3a8ff09e1f63_1254x1254.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!NMlq!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5307dade-240f-4620-a799-3a8ff09e1f63_1254x1254.png 424w, https://substackcdn.com/image/fetch/$s_!NMlq!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5307dade-240f-4620-a799-3a8ff09e1f63_1254x1254.png 848w, https://substackcdn.com/image/fetch/$s_!NMlq!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5307dade-240f-4620-a799-3a8ff09e1f63_1254x1254.png 1272w, https://substackcdn.com/image/fetch/$s_!NMlq!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5307dade-240f-4620-a799-3a8ff09e1f63_1254x1254.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>Introduction</h2><p>At <strong>5:21pm Eastern on Friday, June 12</strong>, Anthropic received a letter. By midnight, the company had switched off the two most capable models it had ever released, for every customer on earth.</p><p>The letter came from the Department of Commerce. It invoked national security authorities and ordered Anthropic to <strong>suspend all access to Claude Fable 5 and Claude Mythos 5</strong> by any foreign national, anywhere, including the company&#8217;s own non-citizen employees. </p><p><strong>Anthropic&#8217;s </strong>own statement, published that night, explained the mechanics of compliance plainly: because there is no clean way to fence a live API endpoint by passport in real time, the only way to obey the order was to take both models down for everyone. <strong>Access to Opus 4.8, Sonnet, and Haiku was untouched.</strong> Fable 5 had shipped three days earlier. Its IPO-grade launch week ended with a kill order.</p><p>There are essentially<strong> two ways to read what happened</strong>, and the honest position is to hold both at once.</p><p>The first reading is narrow and technical. The trigger, by Anthropic&#8217;s account, was a single bypass technique that amounts to pointing the model at a codebase and asking it to <strong>fix the security flaws</strong>. </p><p>The company says the capability is widely available, including from <strong>OpenAI&#8217;s GPT-5.5</strong>, and is used every day by the defenders who keep systems running. On this reading, the suspension is an overzealous first use of a blunt instrument, a misunderstanding that both sides want resolved, and a story that will be over in days to weeks.</p><p>The second reading is structural and does not go away when the models come back. For the first time, the <strong>United States has reached for export-control machinery</strong>, designed for physical dual-use technology, and used it to pull a commercial AI model out of the hands of hundreds of millions of people with a Friday-afternoon letter. </p><p>The kill switch was never in the model. It was not in the classifiers, the <strong>thousand hours of red-teaming</strong>, or the thirty-day data retention. It was in the mail. And once a mechanism like that has been used once, it exists for every lab, every Friday, from now on.</p><p>This piece is about both readings, and about a third fact that reframes the whole episode: this is not the first time in <strong>2026 that the US government has moved against Anthropic specifically</strong>. It is the second, from a different department, under a different legal theory, in the space of three months. The pattern is the story.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!B65X!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F650468c3-b9e2-4f8c-9358-c329b0c14600_1650x810.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!B65X!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F650468c3-b9e2-4f8c-9358-c329b0c14600_1650x810.png 424w, https://substackcdn.com/image/fetch/$s_!B65X!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F650468c3-b9e2-4f8c-9358-c329b0c14600_1650x810.png 848w, https://substackcdn.com/image/fetch/$s_!B65X!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F650468c3-b9e2-4f8c-9358-c329b0c14600_1650x810.png 1272w, https://substackcdn.com/image/fetch/$s_!B65X!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F650468c3-b9e2-4f8c-9358-c329b0c14600_1650x810.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!B65X!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F650468c3-b9e2-4f8c-9358-c329b0c14600_1650x810.png" width="1456" height="715" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/650468c3-b9e2-4f8c-9358-c329b0c14600_1650x810.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:715,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Seventy-two hours from launch to global shutdown.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Seventy-two hours from launch to global shutdown." title="Seventy-two hours from launch to global shutdown." srcset="https://substackcdn.com/image/fetch/$s_!B65X!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F650468c3-b9e2-4f8c-9358-c329b0c14600_1650x810.png 424w, https://substackcdn.com/image/fetch/$s_!B65X!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F650468c3-b9e2-4f8c-9358-c329b0c14600_1650x810.png 848w, https://substackcdn.com/image/fetch/$s_!B65X!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F650468c3-b9e2-4f8c-9358-c329b0c14600_1650x810.png 1272w, https://substackcdn.com/image/fetch/$s_!B65X!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F650468c3-b9e2-4f8c-9358-c329b0c14600_1650x810.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><span data-color="rgb(187, 59, 40)" style="color: rgb(187, 59, 40);">Fig. 1</span> Seventy-two hours from launch to global shutdown.</figcaption></figure></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>What got pulled</h2><p>To understand the stakes, you<strong> have to understand what was switched off</strong>, because it was not an ordinary model.</p><p>On June 9, Anthropic launched Claude Fable 5 and Claude Mythos 5 together. In the company&#8217;s framing,<strong> both belong to a new tier it calls Mythos-class</strong>, sitting above the Opus class in raw capability. </p><p>The first member of that tier, <strong>Claude Mythos Preview</strong>, had been released quietly in April through a government-linked program called Project Glasswing, and never offered to the public. </p><p>Fable 5 was the moment that tier went mainstream. In <strong>Anthropic&#8217;s own words</strong>, its capabilities exceed those of any model the company had ever made generally available, with the lead over Opus 4.8 widening as tasks grow longer and more agentic.</p><p>Fable 5 and Mythos 5 share the same underlying weights. The difference is entirely in the safeguards, which is also why they carry different names. <strong>Fable</strong>, from the Latin <em>fabula</em>, is the version <strong>made safe for general release.</strong> <strong>Mythos </strong>is the same model with safeguards lifted in specific domains, reserved for <strong>vetted cyberdefenders</strong> and critical-infrastructure operators inside Glasswing. Anthropic describes Mythos 5 as having the strongest cybersecurity capabilities of any model in the world.</p><p>The headline numbers were not subtle, and software engineering was the spine of the launch. On Anthropic&#8217;s published evaluations, <strong>Fable 5 posted 80.3 percent on SWE-bench Pro against 69.2 for Opus 4.8 and 58.6 for OpenAI&#8217;s GPT-5.5</strong>, and scored highest among frontier models on Cognition&#8217;s harder FrontierCode set, even at medium effort. </p><blockquote><p><em>The widely quoted 95 percent on SWE-bench Verified is real but should be read with care, because that benchmark is saturating; the gap that means something is the roughly eleven points on the harder Pro variant.</em></p></blockquote><p>Stripe, in early testing, used the model to perform a codebase-wide migration of a<strong> 50-million-line Ruby codebase in a single day</strong>, work it estimated would have taken a team over two months by hand. </p><p>The capability that matters most for the rest of this story is the one the public model hides, and it is worth being precise about what it is. </p><p>Anthropic&#8217;s cybersecurity evaluations measure offensive progress directly: the Firefox benchmark scores the fraction of trials that reach arbitrary code execution,<strong> OSS-Fuzz is a severity-weighted score </strong>that runs from a basic crash up to a full control-flow hijack, and CyberGym, a Berkeley suite of <strong>more than 1,500 real-world tasks</strong> built on Google&#8217;s fuzzing corpus, scores whether the model can reproduce a genuine vulnerability and prove it with a crashing input. </p><p>On the OSS-Fuzz measure, an independent reading of the Mythos 5 system card by<strong> Epoch AI</strong> put the unsafeguarded model&#8217;s crash rate at 80 percent against Opus 4.8&#8217;s 61.5. </p><p>The detail that matters most is that Anthropic says it did not train these abilities in; they <strong>emerged as a byproduct of general gains</strong> in code understanding and autonomy, which is the technical reason they cannot be cleanly excised. </p><p>Fable&#8217;s classifiers are tuned to block any progress on exactly these tasks. <strong>Pricing landed at ten dollars per million input tokens</strong> and fifty per million output, double Opus 4.8, and by Anthropic&#8217;s own description less than half what the retired Mythos Preview had cost. The <strong>context window is one million tokens</strong>, with up to 128k output.</p><p>For practitioners, the integration surface is where the safety architecture becomes concrete, and it is a fallback rather than a wall. By Anthropic&#8217;s account, when Fable&#8217;s classifiers flag a request touching one of three domains, the <strong>response is not refused</strong> but quietly handled by Claude Opus 4.8 instead, and the user is told it happened. </p><p>The <strong>model ID is </strong><code>claude-fable-5</code>, a drop-in string swap; developer documentation indicates the <strong>fallback surfaces in the API response</strong> with a refusal stop reason and the triggering category named, so a client can route or retry, though Anthropic&#8217;s own materials describe the behavior rather than the field names. </p><p>The<strong> three covered domains are </strong>the load-bearing detail:<strong> offensive cybersecurity, biology and chemistry, and distillation</strong>, the last being organized attempts to extract Claude&#8217;s capabilities to train competing models, which Anthropic says it has already detected at scale originating from authoritarian states. </p><p>Those classifiers route <strong>fewer than five percent of sessions</strong> away from the frontier model; for the other ninety-five, Anthropic says Fable performs effectively identically to the unsafeguarded Mythos 5. </p><p>Both Mythos-class models are <strong>designated Covered Models</strong>, carrying a mandatory thirty-day data retention even for enterprise customers who previously held zero-retention agreements. None of this is cosmetic. </p><p>Under<strong> Anthropic&#8217;s Responsible Scaling Policy</strong>, a model with this much cyber and biological capability is handled at <em><strong>AI Safety Level 3</strong></em>, the tier whose rules require hardened deployment and security controls, and the classifiers, the fallback, and the retention are the concrete form those requirements take.</p><p>That <strong>retention policy</strong> is part of what Anthropic points to as evidence of seriousness, and its definitions matter for everything that follows. <strong>A universal jailbreak</strong>, in Anthropic&#8217;s own terms, is any prompt, script, or harness that lets a user interact with the model as if its safeguards were not present; a non-universal one elicits some capability only in narrow circumstances, or needs reworking for each new case. </p><p>Anthropic says an external bug bounty ran more than a thousand hours without producing a universal jailbreak, that Fable <strong>complied with zero harmful single-turn cyber requests</strong> even when tested against thirty different public jailbreak techniques, and that the thirty-day retention exists precisely so it can detect and patch novel attacks in flight. </p><p>It also disclosed, in the same launch post, that the <strong>UK AI Security Institute had made progress toward a universal jailbreak </strong>within a brief testing window. That admission matters more than it looks: Anthropic&#8217;s case was never that Fable is unbreakable, but that breaking it is slow and costly enough to monitor. </p><p>That is <strong>a probabilistic claim</strong>, not an absolute one, which is exactly why the dispute that followed turned on what counts as a serious enough break.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!603C!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff0523776-b7a0-4cda-9b1b-d0aad65b983c_1650x855.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!603C!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff0523776-b7a0-4cda-9b1b-d0aad65b983c_1650x855.png 424w, https://substackcdn.com/image/fetch/$s_!603C!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff0523776-b7a0-4cda-9b1b-d0aad65b983c_1650x855.png 848w, https://substackcdn.com/image/fetch/$s_!603C!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff0523776-b7a0-4cda-9b1b-d0aad65b983c_1650x855.png 1272w, https://substackcdn.com/image/fetch/$s_!603C!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff0523776-b7a0-4cda-9b1b-d0aad65b983c_1650x855.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!603C!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff0523776-b7a0-4cda-9b1b-d0aad65b983c_1650x855.png" width="1456" height="754" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f0523776-b7a0-4cda-9b1b-d0aad65b983c_1650x855.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:754,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;What got pulled: the only public model above Opus.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="What got pulled: the only public model above Opus." title="What got pulled: the only public model above Opus." srcset="https://substackcdn.com/image/fetch/$s_!603C!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff0523776-b7a0-4cda-9b1b-d0aad65b983c_1650x855.png 424w, https://substackcdn.com/image/fetch/$s_!603C!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff0523776-b7a0-4cda-9b1b-d0aad65b983c_1650x855.png 848w, https://substackcdn.com/image/fetch/$s_!603C!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff0523776-b7a0-4cda-9b1b-d0aad65b983c_1650x855.png 1272w, https://substackcdn.com/image/fetch/$s_!603C!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff0523776-b7a0-4cda-9b1b-d0aad65b983c_1650x855.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong><span data-color="rgb(187, 59, 40)" style="color: rgb(187, 59, 40);">Fig. 2</span></strong> What got pulled: the only public model above Opus.</figcaption></figure></div><p>Fable 5 was <strong>not a careless release.</strong> It was, by design and by the company&#8217;s loud public framing, the <strong>most safety-instrumented frontier launch</strong> the industry had produced. That is what makes the next three days strange.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!VfVP!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0cc81b93-b6ee-4aea-bf4b-e4e820c867a8_1650x810.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!VfVP!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0cc81b93-b6ee-4aea-bf4b-e4e820c867a8_1650x810.png 424w, https://substackcdn.com/image/fetch/$s_!VfVP!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0cc81b93-b6ee-4aea-bf4b-e4e820c867a8_1650x810.png 848w, https://substackcdn.com/image/fetch/$s_!VfVP!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0cc81b93-b6ee-4aea-bf4b-e4e820c867a8_1650x810.png 1272w, https://substackcdn.com/image/fetch/$s_!VfVP!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0cc81b93-b6ee-4aea-bf4b-e4e820c867a8_1650x810.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!VfVP!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0cc81b93-b6ee-4aea-bf4b-e4e820c867a8_1650x810.png" width="1456" height="715" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0cc81b93-b6ee-4aea-bf4b-e4e820c867a8_1650x810.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:715,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;The safeguards the government overrode.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="The safeguards the government overrode." title="The safeguards the government overrode." srcset="https://substackcdn.com/image/fetch/$s_!VfVP!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0cc81b93-b6ee-4aea-bf4b-e4e820c867a8_1650x810.png 424w, https://substackcdn.com/image/fetch/$s_!VfVP!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0cc81b93-b6ee-4aea-bf4b-e4e820c867a8_1650x810.png 848w, https://substackcdn.com/image/fetch/$s_!VfVP!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0cc81b93-b6ee-4aea-bf4b-e4e820c867a8_1650x810.png 1272w, https://substackcdn.com/image/fetch/$s_!VfVP!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0cc81b93-b6ee-4aea-bf4b-e4e820c867a8_1650x810.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong><span data-color="rgb(187, 59, 40)" style="color: rgb(187, 59, 40);">Fig. 3</span></strong> The safeguards the government overrode.</figcaption></figure></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>The trigger, and the call that wasn&#8217;t answered</h2><p>The cleanest account of how the letter came to be written runs through Anthropic&#8217;s own cap table.</p><p>By the<strong> reporting of the Wall Street Journal </strong>and <strong>The Information</strong>, later echoed by Bloomberg, the bypass that set everything in motion was found not by an adversary but by researchers at Amazon, Anthropic&#8217;s largest investor. Amazon&#8217;s chief executive,<strong> Andy Jassy</strong>, took the finding straight to Washington, <strong>personally alerting Treasury Secretary Scott Bessent</strong> that his researchers had demonstrated a way around Fable 5&#8217;s cybersecurity guardrails. </p><p>The escalation reached Commerce Secretary Howard Lutnick the same day. By one reconstruction of the timeline, the <strong>government first called Anthropic at 1:00pm Eastern</strong>; the signed letter followed hours later, at 5:21. Dario Amodei spent that Friday on calls with cabinet officials, arguing that Amazon&#8217;s technique did not amount to a true jailbreak. The letter went out anyway.</p><p><em>What was the technique?</em> Here the reporting converges on something almost anticlimactic. Anthropic&#8217;s statement describes it as <strong>asking the model to read a specific codebase </strong>and fix any software flaws, a single-instance, non-universal bypass that surfaced a handful of previously known, minor vulnerabilities. </p><p>The most authoritative outside account is from<strong> Katie Moussouris</strong>, the founder of <strong>Luta Security</strong>, who serves on the Commerce Department&#8217;s own Information Systems Technical Advisory Committee and the Cyber Safety Review Board, and who says she is the only outside expert to have actually read the underlying research paper. Her summary is blunt. </p><p>The <strong>researchers took open-source code with known CVEs</strong>, plus new code with deliberately planted bugs, and <em>asked Fable 5, Mythos, and Opus to review it for security issues.</em> <strong>Fable 5 refused</strong>. They then asked the models to fix the code, and through a manual, multi-step process turned the output into scripts that test the patches. </p><p>That, she writes, is it. She proposed, only half in jest, a 1990s-style t-shirt: &#8220;<em>fix this code</em>&#8221; on the front, &#8220;<em>this shirt is a munition</em>&#8221; on the back.</p><p><strong>Her technical point is </strong>the one that should worry policymakers more than the prompt itself. The reason the technique works is that it is a defensive request. <strong>Asking a model to find a bug</strong>, explain why it matters, and write a test that confirms the patch is the single most valuable thing an AI can do for defensive security, the find-fix-test loop that defenders run every day.</p><p>You cannot remove that capability, she argues, without making the model worse at defending real systems. There is <strong>no surgical edit</strong> that deletes the offense and keeps the defense, because at this level they are the same muscle.</p><p>An independent test complicates the picture from the other direction, and it is worth holding alongside Anthropic&#8217;s own numbers. </p><p>The security firm <strong>Endor Labs</strong> benchmarked Fable 5 on <strong>two hundred real-world vulnerability-fixing tasks</strong> and found it landed mid-table, at 59.8 percent on functional correctness and just 19 percent on producing a genuinely secure fix, with a record number of timeouts the firm attributed to the model&#8217;s extended thinking. </p><p>Their reading cuts both ways: <strong>Anthropic&#8217;s headline cyber numbers</strong>, they note, mostly measure offensive progress, exploits and proofs of concept, not whether the model writes safe production code. </p><p>So the <strong>capability </strong>that triggered the recall <strong>is real and measurable</strong>, but it is specifically an offensive-discovery capability, and even that is uneven across the cyber task surface.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!ap0x!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf3b7f77-e280-4178-8cac-a89b57fe3934_1650x780.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!ap0x!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf3b7f77-e280-4178-8cac-a89b57fe3934_1650x780.png 424w, https://substackcdn.com/image/fetch/$s_!ap0x!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf3b7f77-e280-4178-8cac-a89b57fe3934_1650x780.png 848w, https://substackcdn.com/image/fetch/$s_!ap0x!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf3b7f77-e280-4178-8cac-a89b57fe3934_1650x780.png 1272w, https://substackcdn.com/image/fetch/$s_!ap0x!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf3b7f77-e280-4178-8cac-a89b57fe3934_1650x780.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!ap0x!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf3b7f77-e280-4178-8cac-a89b57fe3934_1650x780.png" width="1456" height="688" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/cf3b7f77-e280-4178-8cac-a89b57fe3934_1650x780.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:688,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;The gap between what was cited and what would justify a recall.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="The gap between what was cited and what would justify a recall." title="The gap between what was cited and what would justify a recall." srcset="https://substackcdn.com/image/fetch/$s_!ap0x!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf3b7f77-e280-4178-8cac-a89b57fe3934_1650x780.png 424w, https://substackcdn.com/image/fetch/$s_!ap0x!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf3b7f77-e280-4178-8cac-a89b57fe3934_1650x780.png 848w, https://substackcdn.com/image/fetch/$s_!ap0x!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf3b7f77-e280-4178-8cac-a89b57fe3934_1650x780.png 1272w, https://substackcdn.com/image/fetch/$s_!ap0x!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf3b7f77-e280-4178-8cac-a89b57fe3934_1650x780.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong><span data-color="rgb(187, 59, 40)" style="color: rgb(187, 59, 40);">Fig. 4</span></strong> The gap between what was cited and what would justify a recall.</figcaption></figure></div><p>This is <strong>the crux of Anthropic&#8217;s public objection. </strong></p><p>The company says it has been given only verbal evidence of a narrow, non-universal bypass, that <em>no concerning jailbreak producing a genuinely harmful result has been disclosed to it</em>, and that recalling a <strong>model deployed to hundreds of millions of people</strong> over a narrow finding would, if applied as an industry standard, halt all frontier deployments everywhere. </p><p>Anthropic has been explicit, including in Amodei&#8217;s own policy writing, that it believes the government <em>should</em> be able to block unsafe deployments, but <strong>through a process that is transparent</strong>, fair, and grounded in technical fact. Its complaint is not that the power exists. It is that this use of it does not meet that bar.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Two rationales, and which one is real</h2><p>Here the <strong>official story develops a seam</strong>, and the seam is worth prying at, because it determines which of the two readings from the top of this piece is the operative one.</p><p>Publicly and to Anthropic, the stated trigger was the jailbreak. But <strong>the Lutnick letter itself</strong>, a copy of which Reuters reported seeing, frames the danger differently: as the risk that the models <strong>could be diverted to military intelligence users in China, Russia</strong>, or other countries of concern. </p><p>Those are not the same justification. One is a <strong>technical claim</strong> about a safeguard failing. The other is a <strong>geopolitical claim </strong>about who might get access. And the divergence is not merely rhetorical: the regulatory provision the order reportedly invokes is one aimed at <strong>military-intelligence end users</strong> in countries of concern, which means the diversion rationale is built into the legal instrument itself, even as Anthropic&#8217;s own public statement names only the jailbreak. </p><p>The hook and the explanation point in different directions. Semafor reported that <strong>White House concerns</strong> ran beyond the bypass entirely, with officials suspecting that a <strong>China-linked group</strong> had accessed Mythos before the shutdown, a suspicion, it should be stressed, that has not been publicly substantiated. </p><p>Anthropic says the question of Chinese access was never raised in any of its conversations about the jailbreak.</p><p>The accounts diverge on the human level too. <strong>David Sacks</strong>, the White House AI adviser, has said the administration gave Amodei a clear choice, fix the jailbreak or <strong>de-deploy the models</strong>, and that Amodei refused. Anthropic disputes that characterization. As of this writing the competing versions have not been reconciled, and a reader should treat both as contested.</p><p><em>Why does the seam matter? </em>Because if the real driver is a narrow technical bypass, the episode is fixable, and probably will be fixed, with some added vetting layer. </p><p>If the real driver is diversion risk to adversary states, the <strong>model may stay restricted no matter what Anthropic does to its classifiers</strong>, because the concern is not about the safeguard at all. And if the real driver is neither, if national security is functioning as the available lever in a relationship that had already curdled, then the technical merits are almost beside the point.</p><p><strong>Michael Horowitz </strong>of the <strong>Council on Foreign Relations</strong>, speaking about the <em>earlier</em> Anthropic dispute, called it a fight &#8220;<em>about politics and personalities</em>&#8221; that was &#8220;<em>masquerading as a policy dispute.</em>&#8221; That framing did not come from nowhere. </p><p>To see why, you have to look at what happened in March.</p><div><hr></div><h2>A part most coverage missed</h2><p>Th<strong>e Fable 5 suspension</strong> has been reported as a discrete event. It is better understood as the <strong>second of two extraordinary statutory actions</strong> against Anthropic inside a single quarter, the sharp end of a federal campaign that had been escalating since winter, and the earlier action tells you a great deal about this one.</p><p>The pressure began before either statute was invoked. By Fortune&#8217;s account, the <strong>Trump administration moved as early as February</strong> to push federal agencies off Anthropic&#8217;s models, after the company refused the Pentagon&#8217;s preferred contract language permitting use of Claude for any lawful purpose. </p><p><strong>Anthropic wanted two carve-outs</strong>: no fully autonomous lethal weapons, and no mass surveillance of Americans. </p><p>That dispute hardened into something unprecedented on March 3, when <strong>Secretary of Defense Pete Hegseth </strong>designated Anthropic a supply-chain risk to national security, invoking a rarely used statute, Section 3252 of Title 10, and barring defense contractors, suppliers, and partners from doing business with the company. </p><p>As the <strong>Congressional Research Service and Courthouse News documented</strong>, the same refusal was the trigger. The Pentagon&#8217;s position, argued later before a DC Circuit panel, was that those guardrails amounted to an <em>&#8220;operational veto,&#8221;</em> a remote kill switch the company could trigger based on its own reading of lawful versus unlawful use. </p><p>Hegseth put it less legally on social media: America&#8217;s warfighters, he wrote, would never be <em>&#8220;held hostage by the ideological whims of Big Tech.&#8221;</em></p><p>The <strong>supply-chain-risk label</strong> is historically reserved for foreign adversaries. Applying it to a US company that was, until weeks earlier, the Pentagon&#8217;s hand-picked frontier vendor, the first whose models ran inside classified networks, was extraordinary. </p><p><strong>Anthropic sued on March 9</strong>, calling the action unprecedented and unlawful. By May, a three-judge DC Circuit panel hearing one of the two challenges sounded openly skeptical of the government&#8217;s reasoning, with the <strong>court pressing </strong>on whether a guardrail against, in one judge&#8217;s framing, telling an unreliable model &#8220;<em>which bombs to drop</em>&#8221; really constituted a supply-chain risk. </p><p>The litigation is unresolved. In the meantime, agencies including Health and Human Services, Treasury, and State confirmed they were migrating off Claude, even as the <strong>Department of Defense</strong>, by CNBC&#8217;s reporting, <strong>kept using Anthropic&#8217;s models</strong> to support active military operations in Iran, and the NSA, by later accounts, continued using Claude for its own work. Blacklisted and indispensable at the same time.</p><p>There is a <strong>political layer </strong>underneath the legal one. Amodei has drawn fire from Sacks, who has accused Anthropic of pushing &#8220;<em>woke AI</em>,&#8221; largely over its regulatory positions. </p><p>A Defense official told CNBC the March decision was about &#8220;<em>the military being able to use technology for all lawful purposes,</em>&#8221; not personality. But the pattern that emerges when you put the two actions side by side is hard to wave away.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!G0FF!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F392bedb3-0ed0-42d3-a54a-b5cedbd0996e_1650x900.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!G0FF!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F392bedb3-0ed0-42d3-a54a-b5cedbd0996e_1650x900.png 424w, https://substackcdn.com/image/fetch/$s_!G0FF!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F392bedb3-0ed0-42d3-a54a-b5cedbd0996e_1650x900.png 848w, https://substackcdn.com/image/fetch/$s_!G0FF!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F392bedb3-0ed0-42d3-a54a-b5cedbd0996e_1650x900.png 1272w, https://substackcdn.com/image/fetch/$s_!G0FF!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F392bedb3-0ed0-42d3-a54a-b5cedbd0996e_1650x900.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!G0FF!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F392bedb3-0ed0-42d3-a54a-b5cedbd0996e_1650x900.png" width="1456" height="794" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/392bedb3-0ed0-42d3-a54a-b5cedbd0996e_1650x900.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:794,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Two federal actions against one company in three months.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Two federal actions against one company in three months." title="Two federal actions against one company in three months." srcset="https://substackcdn.com/image/fetch/$s_!G0FF!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F392bedb3-0ed0-42d3-a54a-b5cedbd0996e_1650x900.png 424w, https://substackcdn.com/image/fetch/$s_!G0FF!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F392bedb3-0ed0-42d3-a54a-b5cedbd0996e_1650x900.png 848w, https://substackcdn.com/image/fetch/$s_!G0FF!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F392bedb3-0ed0-42d3-a54a-b5cedbd0996e_1650x900.png 1272w, https://substackcdn.com/image/fetch/$s_!G0FF!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F392bedb3-0ed0-42d3-a54a-b5cedbd0996e_1650x900.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong><span data-color="rgb(187, 59, 40)" style="color: rgb(187, 59, 40);">Fig. 5</span></strong> Two federal actions against one company in three months.</figcaption></figure></div><p><strong>Two departments.</strong> Two legal instruments that had never been pointed at an AI company before. One target. In March, the objection was that Anthropic&#8217;s safety restrictions were too strong, that it would not let the government use Claude freely enough. </p><p>In June, <strong>the objection was</strong> that <strong>Anthropic&#8217;s safety restrictions</strong> were too weak, that a bypass let users reach capability the safeguards were meant to contain. The company was punished, in the same quarter, for both having guardrails and for not having good enough ones. </p><p>Whatever else that is, it is <strong>not a stable regulatory environment,</strong> and it is the context every line of Anthropic&#8217;s confidential S-1 now sits inside.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>The machinery, and why it is new</h2><p>The <strong>legal vehicle deserves a closer look</strong>, because its novelty is the part with the longest half-life.</p><p>The legal vehicle deserves a closer look, because its novelty is the part with the longest half-life. The order took the form of what export lawyers call an &#8220;<em>is informed</em>&#8221; letter from the <strong>Commerce Department&#8217;s Bureau of Industry and Security</strong>, the mechanism by which BIS notifies a specific company that a license is now required for specific items. That mechanism is not itself new. </p><p>What is new is pointing it at a deployed commercial model. Reuters, which reported seeing the letter, said Commerce invoked the<strong> Export Control Reform Act of 2018</strong> and its authority over emerging and foundational technologies; analysts parsing the public reporting place the regulatory hook in the export rules covering <strong>military-intelligence end users</strong>, though the letter itself has not been released, so the precise provisions remain a reconstruction. </p><p>The directive requires a license for any export, re-export, or <strong>in-country transfer of the models</strong> to a foreign person, and warns that noncompliance brings prompt criminal and civil penalties.</p><p>The doctrine doing the work is the deemed export, and understanding it is the whole game. Under the export regulations, releasing<strong> controlled technology or source code</strong> to a foreign national inside the United States has for decades been treated as an export to that person&#8217;s home country.</p><p>Showing a controlled blueprint to a foreign-born engineer in a San Francisco office is, in law, a shipment abroad. The letter applies that logic to a model: every inference a foreign national draws from Fable arguably releases the controlled capability to them, which is<strong> why the order reached foreign nationals on US soil</strong> and Anthropic&#8217;s own non-citizen staff, and why no geofence could satisfy it.</p><p><strong>Read with an engineer&#8217;s eye</strong>, the theory strains at a seam. Export controls were built for things that cross borders: chips, machine tools, encryption binaries, centrifuge designs. </p><p>The cited rules speak of technology and source code, but a hosted model hands the user neither. It<strong> hands them inference</strong>, an outcome rather than an artifact, and the weights never leave Anthropic&#8217;s data centers. </p><p>As one former federal prosecutor put it to CIO, the <strong>physical location of the source code</strong> has become <strong>irrelevant</strong>; what is being controlled is access to what the code can do. That is a real shift in what export means, and it is not clearly settled law. </p><p>Export controls have not traditionally reached foreign access to US software as a service, which is precisely why, as Lawfare noted, the House passed a <strong>Remote Access Security Act in January</strong> to extend export jurisdiction to remote access of controlled US technology. The government reached for a theory that Congress is, at the same moment, trying to write into statute because it is not yet clearly there.</p><p>There is one more piece of timing that <strong>turns the episode from aggressive into nearly incoherent</strong>. Ten days before the letter, on June 2, the same administration signed an executive order titled &#8220;<em>Promoting Advanced Artificial Intelligence Innovation and Security,</em>&#8221; setting up a voluntary framework, to be <strong>designed by August 1</strong>, under which developers could offer the government early access to frontier models up to thirty days before release. </p><p>The President had just cut a planned ninety-day cybersecurity review window down to thirty, and by multiple accounts the order explicitly <strong>barred any mandatory licensing</strong> or pre-clearance regime. Anthropic, OpenAI, and Google all welcomed it as the reasonable version of government engagement: <strong>a request, not a rule.</strong> </p><p>The framework was not built. The August deadline had not arrived. And before the voluntary process could take its first breath, the <strong>government reached for the binding instrument</strong> it had just promised not to use. Voluntary in the announcement, mandatory in the execution, ten days apart.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>The own-goal: who export control actually reaches</h2><p>Set aside the question of whether the safeguard failed, and ask the colder question an <strong>export-control regime</strong> is supposed to answer: </p><blockquote><p><em>does it stop the bad outcome it names?</em></p></blockquote><p>The collected judgment of the security profession is that it does not, and may do the reverse. <strong>More than eighty cybersecurity executives</strong>, including leaders at firms such as Nvidia and Adobe, signed an open letter to Lutnick and National Cyber Director Sean Cairncross over the weekend <strong>asking that the controls be lifted</strong>, an effort now hosted at freefable.org. </p><p>Moussouris&#8217;s argument is the technical spine of their case: the capability the <strong>order targets is defensive</strong>, the people it cuts off are defenders, and the attackers it is meant to thwart never needed a US endpoint in the first place.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!YH9D!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe26fff03-bd6b-434d-9b58-065430447e5e_1650x840.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!YH9D!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe26fff03-bd6b-434d-9b58-065430447e5e_1650x840.png 424w, https://substackcdn.com/image/fetch/$s_!YH9D!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe26fff03-bd6b-434d-9b58-065430447e5e_1650x840.png 848w, https://substackcdn.com/image/fetch/$s_!YH9D!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe26fff03-bd6b-434d-9b58-065430447e5e_1650x840.png 1272w, https://substackcdn.com/image/fetch/$s_!YH9D!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe26fff03-bd6b-434d-9b58-065430447e5e_1650x840.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!YH9D!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe26fff03-bd6b-434d-9b58-065430447e5e_1650x840.png" width="1456" height="741" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e26fff03-bd6b-434d-9b58-065430447e5e_1650x840.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:741,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;The asymmetry at the heart of the policy.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="The asymmetry at the heart of the policy." title="The asymmetry at the heart of the policy." srcset="https://substackcdn.com/image/fetch/$s_!YH9D!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe26fff03-bd6b-434d-9b58-065430447e5e_1650x840.png 424w, https://substackcdn.com/image/fetch/$s_!YH9D!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe26fff03-bd6b-434d-9b58-065430447e5e_1650x840.png 848w, https://substackcdn.com/image/fetch/$s_!YH9D!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe26fff03-bd6b-434d-9b58-065430447e5e_1650x840.png 1272w, https://substackcdn.com/image/fetch/$s_!YH9D!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe26fff03-bd6b-434d-9b58-065430447e5e_1650x840.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong><span data-color="rgb(187, 59, 40)" style="color: rgb(187, 59, 40);">Fig. 6</span></strong> The asymmetry at the heart of the policy.</figcaption></figure></div><p>The asymmetry is stark. On one side of the ledger, the <strong>order cuts off allied-nation cyberdefenders</strong>, foreign-national engineers working legally in the United States, Anthropic&#8217;s own staff, the <strong>Glasswing partners</strong> across more than fifteen countries who were using Mythos precisely to secure critical infrastructure, and any researcher running a find-fix-test loop. </p><p>On the other side, it reaches none of the things it is nominally about. It <strong>does not touch Chinese open-weight models</strong>; on June 13, the day after the shutdown, the Chinese lab <strong>Zhipu AI shipped GLM-5.2 </strong>and explicitly cited the US ban as evidence that American models are unreliable partners. </p><p>It does <strong>not touch other frontier labs&#8217; cyber-capable models</strong>, which by Moussouris&#8217;s reckoning carry fewer guardrails and comparable capability, with the rest of the field expected to match Mythos-class capability within months. </p><p>It does not touch adversary state actors, who have their own systems. And it does not touch the capability itself, which is diffusing regardless. You cannot, <strong>as Moussouris puts it</strong>, export-control your way to cyber resilience.</p><p>The contradiction has not been lost on people who know the regime from the inside. <strong>Dean Ball</strong>, an AI policy analyst who briefly served in this administration, called the action cartoonish, pointing to the oddity of a government that <strong>waves advanced AI chips</strong> through to China while barring Britain and every other allied user from its best models. </p><p><strong>Moussouris</strong>, who lived through the last version of this fight, put the bottom line more bluntly: if national defense was the goal, the order scored an own goal against the United States.</p><p>There is a precedent she lived through. When the <strong>Wassenaar Arrangement </strong>added controls on &#8220;<em>intrusion software</em>&#8221; in 2013, the definition was written so broadly that it reached the routine cross-border work of defense itself, <strong>sharing exploit proofs of concept</strong>, coordinating vulnerability disclosure, running incident response, to the point that the United States declined to implement the original language and the text had to be <strong>renegotiated in 2017 </strong>to carve defensive research back out. </p><p>The Fable 5 directive, signed in an afternoon, has the same shape and none of the deliberation.</p><div><hr></div><h2>The sovereignty bill comes due</h2><p>The diplomatic cost is the part that will outlast the news cycle, and it lands closest to home for <strong>anyone reading this from outside the United States.</strong></p><p>For every allied government that had quietly assumed continuous access to the frontier, the <strong>message of June 12 was unambiguous</strong>: that access is revocable, unilaterally, on a domestic legal theory you have no vote in, with a few hours&#8217; notice. </p><p>The reaction was immediate and came from the top. At the <strong>G7 summit at Evian-les-Bains this week</strong>, the suspension became one of the sharpest flashpoints on the agenda. <strong>UK Prime Minister Keir Starmer</strong> raised the blackout directly with Trump and asked for a carve-out restoring access for British citizens and businesses. Washington rebuffed it. </p><p><strong>UK AI minister Kanishka Narayan</strong> put the stakes plainly, noting that the most advanced AI in the world had just been cut off for everyone in Britain, and argued the episode proves the case for sovereign AI capability. </p><p><strong>Canada&#8217;s Mark Carney </strong>framed it as a warning about overreliance on any single model. The European Union, already the stricter regulator, now has fresh reason to reduce its dependence on American AI infrastructure.</p><p>Out of that pressure, a <strong>workaround is taking shape</strong>, and its shape is the most revealing part of the whole episode. According to Reuters, the Financial Times, and Axios, a <strong>US delegation led by Lutnick </strong>spent the summit&#8217;s sidelines negotiating a &#8220;<em>trusted partners</em>&#8221; framework: a sanctioned channel through which vetted allies, either whole countries or individual companies, could regain access to the controlled models. </p><p>The<strong> pitch is cybersecurity</strong>, the same allied-defense logic the order itself invoked. Read it against what Anthropic did on June 12 and the irony is exact. The company <strong>declined to gate access by nationality</strong> and shut the models off rather than operate that system. Governments are now assembling the same nationality gate themselves, one level up, and calling it a partnership. </p><p>The <strong>selective-access layer</strong> that was too compromising for a private firm to run is being rebuilt as a diplomatic club, in which reaching the frontier becomes a privilege Washington grants to allies rather than a product a company sells to customers. <strong>No agreement has been reached</strong>, and which countries and which companies would qualify is undefined.</p><p>The <strong>structural irony compounds</strong>. An action justified by the need to deny capability to adversaries functions, in practice, as the strongest possible argument for those adversaries&#8217; domestic alternatives, and for <strong>allied sovereign-AI programs</strong> that route around US providers entirely. </p><p>Zhipu&#8217;s launch timing was not a coincidence; it was marketing handed to Beijing for free. This is the <strong>AI-sovereignty debate </strong>stripped of abstraction. It is no longer a conference panel. </p><p>It is a <strong>procurement decision </strong>that every<strong> non-US enterprise and government now has to price</strong>, and the new line item is the probability that the best model on the market disappears for an indeterminate number of weeks because of a dispute in Washington it cannot see coming.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>An IPO inside the blast radius</h2><p>All of this is happening on the worst possible calendar.</p><p>By the reporting around the launch, confirmed in the company&#8217;s own filing announcement, <strong>Anthropic confidentially filed its IPO prospectus with the SEC on June 1</strong>, eight days before Fable 5 shipped and eleven before it was pulled. </p><p>The filing followed a 65 billion dollar Series H that valued the company at 965 billion post-money, ahead of OpenAI&#8217;s 852 billion from late March, on a revenue run-rate that had crossed 47 billion by Anthropic&#8217;s May disclosure. The <strong>growth is real and almost vertical</strong>: Anthropic told investors it expects 10.9 billion dollars of revenue in the second quarter alone, more than double the first. The <strong>margin underneath is thinner</strong> than the headline. </p><p>Its own projected second-quarter operating profit implies a margin of roughly five percent, and at the reported valuation <strong>the company is priced near twenty times annualized revenue</strong>, a multiple that assumes years of uninterrupted hypergrowth. </p><p>OpenAI confirmed its own confidential filing days later, and SpaceX is gearing up for a record public debut this week. </p><p><strong>Three of the largest IPOs in history </strong>are converging, and one of the three just demonstrated, live, that its flagship product can be switched off by a letter.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!1hE2!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F413f5823-4987-48c7-bee3-2a6a7132c96d_1650x840.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!1hE2!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F413f5823-4987-48c7-bee3-2a6a7132c96d_1650x840.png 424w, https://substackcdn.com/image/fetch/$s_!1hE2!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F413f5823-4987-48c7-bee3-2a6a7132c96d_1650x840.png 848w, https://substackcdn.com/image/fetch/$s_!1hE2!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F413f5823-4987-48c7-bee3-2a6a7132c96d_1650x840.png 1272w, https://substackcdn.com/image/fetch/$s_!1hE2!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F413f5823-4987-48c7-bee3-2a6a7132c96d_1650x840.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!1hE2!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F413f5823-4987-48c7-bee3-2a6a7132c96d_1650x840.png" width="1456" height="741" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/413f5823-4987-48c7-bee3-2a6a7132c96d_1650x840.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:741,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;A flagship that can be pulled by letter, mid-roadshow.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="A flagship that can be pulled by letter, mid-roadshow." title="A flagship that can be pulled by letter, mid-roadshow." srcset="https://substackcdn.com/image/fetch/$s_!1hE2!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F413f5823-4987-48c7-bee3-2a6a7132c96d_1650x840.png 424w, https://substackcdn.com/image/fetch/$s_!1hE2!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F413f5823-4987-48c7-bee3-2a6a7132c96d_1650x840.png 848w, https://substackcdn.com/image/fetch/$s_!1hE2!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F413f5823-4987-48c7-bee3-2a6a7132c96d_1650x840.png 1272w, https://substackcdn.com/image/fetch/$s_!1hE2!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F413f5823-4987-48c7-bee3-2a6a7132c96d_1650x840.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong><span data-color="rgb(187, 59, 40)" style="color: rgb(187, 59, 40);">Fig. 7</span></strong> A flagship that can be pulled by letter, mid-roadshow.</figcaption></figure></div><p>In an earlier piece for this newsletter, <strong>I argued that Anthropic&#8217;s valuation rested on a question no one could yet answer</strong>, because the S-1 was sealed: whether the revenue was high-margin, organically demanded, and durably embedded, or low-margin and propped by a circular capital structure. <strong>Two different businesses</strong>, hiding inside the same black box. The Fable 5 episode adds a second axis of the same kind. </p><p>There are now two different regulatory realities the company could be living in, and the filing does not tell you which. </p><p>In one, these are isolated frictions with an administration that will normalize, and the government dependency cuts the other way, toward Glasswing, toward classified deployments, toward a company so embedded in national security that it is too important to sideline. </p><p>In the other, the through-line from<strong> the Pentagon blacklist</strong> to the <strong>Commerce directive</strong> is a <strong>durable hostility</strong> that will keep surfacing as new restrictions, new carve-outs, and new headline risk, precisely the kind a roadshow cannot price.</p><p>The market&#8217;s read leans toward the benign outcome on the immediate question, though these are <strong>live prediction-market prices</strong> that move by the day, not fixed facts, and different outlets captured different values within the same week. </p><p>As of mid-June,<strong> Kalshi&#8217;s own market</strong> put the odds of Fable 5 returning before July 1 at 57 percent, before July 10 at 67, and before July 17 at 75; a parallel Polymarket contract ran somewhat higher, around seventy percent for a US return by July 1.</p><p>Kalshi traders separately gave <strong>Anthropic a 77 percent chance of reaching the public markets before OpenAI.</strong> The figures drift with each session, but across both venues the signal is the same: a bet on &#8220;<em>restored, with conditions,</em>&#8221; not on a permanent kill.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!moZ0!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb078d6ee-43e1-4be0-bfc3-52903804df83_1650x810.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!moZ0!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb078d6ee-43e1-4be0-bfc3-52903804df83_1650x810.png 424w, https://substackcdn.com/image/fetch/$s_!moZ0!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb078d6ee-43e1-4be0-bfc3-52903804df83_1650x810.png 848w, https://substackcdn.com/image/fetch/$s_!moZ0!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb078d6ee-43e1-4be0-bfc3-52903804df83_1650x810.png 1272w, https://substackcdn.com/image/fetch/$s_!moZ0!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb078d6ee-43e1-4be0-bfc3-52903804df83_1650x810.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!moZ0!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb078d6ee-43e1-4be0-bfc3-52903804df83_1650x810.png" width="1456" height="715" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b078d6ee-43e1-4be0-bfc3-52903804df83_1650x810.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:715,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;The market expects &#8220;restored, with conditions,&#8221; not a kill.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="The market expects &#8220;restored, with conditions,&#8221; not a kill." title="The market expects &#8220;restored, with conditions,&#8221; not a kill." srcset="https://substackcdn.com/image/fetch/$s_!moZ0!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb078d6ee-43e1-4be0-bfc3-52903804df83_1650x810.png 424w, https://substackcdn.com/image/fetch/$s_!moZ0!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb078d6ee-43e1-4be0-bfc3-52903804df83_1650x810.png 848w, https://substackcdn.com/image/fetch/$s_!moZ0!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb078d6ee-43e1-4be0-bfc3-52903804df83_1650x810.png 1272w, https://substackcdn.com/image/fetch/$s_!moZ0!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb078d6ee-43e1-4be0-bfc3-52903804df83_1650x810.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong><span data-color="rgb(187, 59, 40)" style="color: rgb(187, 59, 40);">Fig. 8</span></strong> The market expects &#8220;restored, with conditions,&#8221; not a kill.</figcaption></figure></div><p>It is worth noting what the comparison to OpenAI implies. OpenAI built its government posture around vetted, tiered access: an explicit government track, a <strong>Defense Department pilot</strong>, sensitive cyber capability released through approval rather than open availability. </p><p>That is closer to what regulators have signaled they want. Anthropic, by shipping a<strong> Mythos-class model</strong> to the broad public first, was arguably the lab that tested the line. It found the line. The cost of being first to the frontier, this quarter, was being first to the letter.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>What this actually means</h2><p>Strip away the specifics and four things are left standing, in rough order of how long they will matter.</p><p>The immediate event is probably resolvable, and the machinery of its resolution is already visible. <strong>Anthropic sent senior engineers to Washington</strong> and met officials through the week, and by the weekend the contours of a deal were forming around vetted access rather than open availability: the <strong>trusted-partners framework floated at the G7</strong>, plus a reported identity-verification layer, both of which look a great deal like how Mythos 5 was restricted in the first place. </p><p>Even David Sacks, no friend of the company, has said the administration&#8217;s <strong>hope is that Anthropic fixes the issue,</strong> the directive is revoked, and Fable returns to general availability. A permanent kill of a flagship the company just filed an IPO around would be the surprise, not the base case.</p><p>The precedent is not resolvable, and it is the real news. Frontier AI has been reclassified, in practice if not yet in settled law, as a controlled export. </p><p>The mechanism to pull a live model from the entire market now exists, has been used, and survives whatever happens to Fable 5 specifically. <strong>The trusted-partners talks make this worse</strong>, not better: they do not roll the precedent back, they institutionalize it, turning a one-off emergency letter into a standing system in which allied access is licensed rather than assumed. </p><p><strong>Every lab, OpenAI and Google included</strong>, now operates one Friday letter away from the same outcome, and the whole industry has just watched the vetted-access posture become the one that survives contact with Washington, while the open-release model gets tested to destruction.</p><p>For practitioners, the lesson is architectural and immediate. If your stack had a <strong>hard dependency on a single frontier model</strong>, June 12 was the day that risk stopped being theoretical. </p><p>The pragmatic move while this plays out is the one Fable 5 itself makes on <strong>high-risk queries:</strong> fall back to Opus 4.8, usually a one-line model-ID change, since it is live, unaffected, and the model Fable defers to anyway. </p><p><strong>AWS did exactly this</strong> at the infrastructure level, automatically rerouting Fable and Mythos calls to Opus 4.8 the moment the order landed. </p><p>The deeper lesson is to treat model availability as a dependency to be abstracted and load-balanced, not a constant. The<strong> capability gap between Mythos-class and Opus-class is real,</strong> widest on long-horizon agentic work, and for some workloads there is no clean substitute today. That gap is now a supply risk, and supply risk gets designed around.</p><p>And for Anthropic, the episode puts its foundational thesis under load. The company&#8217;s entire identity is the bet that safety and capability can be co-developed, that being the most careful lab is a moat rather than a tax. </p><p>The uncomfortable reading of this quarter is that <strong>the moat became the target.</strong> The capability it built to lead the frontier is exactly what got the model classed as a munition, and the very language the company used to sell its caution helped write the legal case for taking it away. </p><p>The <strong>cybersecurity researcher Peter Girnus </strong>made the point with a blade: a company that calls its product a munition in every press release should not be shocked when a government eventually takes it at its word. As he put it, &#8220;<em>They wrote the legal predicate themselves and called it a brand</em>.&#8221; </p><p>The <strong>government partnership</strong> Anthropic cultivated as a differentiator is the same channel through which all of it arrived. Safety did not exempt the company. By one reading, safety is what made it conspicuous.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>What to watch</h2><p>Because the <strong>story is moving daily, </strong>here are the markers that will tell you which of the two readings is winning, in roughly the order they will resolve.</p><p>The terms of restoration, not the fact of it. &#8220;<em>Restored</em>&#8221; and &#8220;<em>restored as it shipped</em>&#8221; are different outcomes. Watch whether Fable 5 returns generally available or only through the trusted-partners channel, whether the reported<strong> early-July identity-verification rollout </strong>extends to foreign nationals or merely confirms US citizenship, and what compliance burden attaches. A narrow, conditioned return is the base case, and the conditions are the whole story.</p><p>Whether the trusted-partners framework actually lands. As of this week it is a negotiation with no agreement, the qualifying countries and companies are undefined, and the <strong>EU and UK are already hedging toward their own capacity.</strong> If it formalizes, allied access to the frontier becomes a license Washington issues. If it collapses, the sovereignty exodus accelerates.</p><p>Whether the China-diversion claim is ever evidenced. It is the one rationale that, if substantiated, flips the base case from &#8220;<em>restored with conditions</em>&#8221; to &#8220;<em>restricted regardless of safeguards.</em>&#8221; So far it is an <strong>unsubstantiated suspicion,</strong> and Anthropic says it was never raised in any conversation about the jailbreak.</p><p><strong>The DC Circuit ruling on the Pentagon designation</strong>. A decision against the government reframes the whole pattern as overreach the courts will check. A decision for it hardens the precedent across both fronts at once.</p><p>And whether the rest of the field quietly changes its release posture. If OpenAI and Google visibly shift toward vetted, tiered access over the coming months, <strong>that is the market pricing the new reality</strong>, and the open-release era of frontier models ends not with a ban but with a business decision.</p><p>Hold all of that, and <strong>resist the two easy endings.</strong> This was not a trivial misunderstanding that proves nothing, and it was not a five-alarm assault on innovation that proves everything. It was a demonstration. </p><p>A<strong> commercial AI model </strong>serving hundreds of millions of people was reclassified as a controlled weapon and switched off centrally, on a <strong>domestic legal theory built for centrifuges</strong>, triggered by a defensive coding prompt, escalated by the company&#8217;s own largest investor, and justified by two rationales that still do not match, in the same quarter a different arm of the same government had blacklisted the same company for the opposite sin. </p><p>The narrow event will probably pass. The <strong>precedent it set will not. </strong>The envelope has been opened in public, and it does not go back inside.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><p><em>Sources include the primary statement and launch posts from Anthropic; reporting from Reuters, the Financial Times, the Wall Street Journal, The Information, Bloomberg, Axios, CNBC, Fortune, TIME, Semafor, TechCrunch, NBC News, and The Next Web; the Congressional Research Service (IN12669) and Courthouse News on the Pentagon designation and litigation; technical analysis by Katie Moussouris of Luta Security, the only outside expert to have read the underlying research paper; the open letter at freefable.org; market-implied data from Kalshi; and the Claude API documentation. The Amazon-origin account is attributed to the Wall Street Journal and The Information; the G7 trusted-partners talks to Reuters, the Financial Times, and Axios. All valuation figures are reported, estimated, or market-implied as labeled. The China-access concern is reported only as an unsubstantiated suspicion, and the competing Sacks and Anthropic accounts of the de-deployment ultimatum remain unreconciled as of June 18, 2026. No claim here rests on a publicly available audited filing, because Anthropic&#8217;s S-1 remains confidential.</em></p>]]></content:encoded></item><item><title><![CDATA[How three companies set the price of intelligence ]]></title><description><![CDATA[HBM, the memory wall, and the physics underneath every token. From the capacitor to the income statement.]]></description><link>https://www.thesoftwarefrontier.com/p/how-three-companies-set-the-price</link><guid isPermaLink="false">https://www.thesoftwarefrontier.com/p/how-three-companies-set-the-price</guid><dc:creator><![CDATA[Lorenzo Bradanini]]></dc:creator><pubDate>Wed, 17 Jun 2026 12:13:06 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!d7Nd!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6b6b42d6-d6fc-49f8-8210-0734fcff887f_1254x1254.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!d7Nd!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6b6b42d6-d6fc-49f8-8210-0734fcff887f_1254x1254.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!d7Nd!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6b6b42d6-d6fc-49f8-8210-0734fcff887f_1254x1254.png 424w, https://substackcdn.com/image/fetch/$s_!d7Nd!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6b6b42d6-d6fc-49f8-8210-0734fcff887f_1254x1254.png 848w, https://substackcdn.com/image/fetch/$s_!d7Nd!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6b6b42d6-d6fc-49f8-8210-0734fcff887f_1254x1254.png 1272w, https://substackcdn.com/image/fetch/$s_!d7Nd!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6b6b42d6-d6fc-49f8-8210-0734fcff887f_1254x1254.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!d7Nd!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6b6b42d6-d6fc-49f8-8210-0734fcff887f_1254x1254.png" width="1254" height="1254" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6b6b42d6-d6fc-49f8-8210-0734fcff887f_1254x1254.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1254,&quot;width&quot;:1254,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2505071,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/201615837?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6b6b42d6-d6fc-49f8-8210-0734fcff887f_1254x1254.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!d7Nd!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6b6b42d6-d6fc-49f8-8210-0734fcff887f_1254x1254.png 424w, https://substackcdn.com/image/fetch/$s_!d7Nd!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6b6b42d6-d6fc-49f8-8210-0734fcff887f_1254x1254.png 848w, https://substackcdn.com/image/fetch/$s_!d7Nd!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6b6b42d6-d6fc-49f8-8210-0734fcff887f_1254x1254.png 1272w, https://substackcdn.com/image/fetch/$s_!d7Nd!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6b6b42d6-d6fc-49f8-8210-0734fcff887f_1254x1254.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><strong>Author&#8217;s note on method.</strong> <em>This report was researched in June 2026 against primary sources (JEDEC standards, vendor announcements, supplier earnings releases) and dated trade reporting, with every computed figure derived from first principles and shown inline. The HBM4 ramp, the Rubin and MI450 launches, and the Q1 2026 memory shock are recent events; a verification list of claims worth re-checking sits in Appendix B, with confidence tiers throughout. Two sections present explicit illustrative models (the cost of a stack, the gigawatt chain): their inputs are tagged as sourced or assumed, their arithmetic is exact, and their outputs are ranges, not disclosures. Cross-vendor specifications use published peak figures and are directional where measurement conventions differ. Appendix C specifies the measurement program that would convert several first-principles claims here into original benchmark data. Nothing in this report is investment advice.</em></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>The thesis: every token is a memory read</h2><p>Strip away the software and an output token is a physical event: to produce one token, an accelerator must stream essentially<strong> every active parameter of the model</strong> from memory through its arithmetic units. A 70 billion parameter model quantized to <strong>FP8 is 70 gigabytes of weights. </strong></p><p>Generating one token means reading 70 gigabytes. Generating a hundred tokens a second means moving seven terabytes a second, and <em>no amount of arithmetic brilliance changes that requirement</em>, because the arithmetic is not the constraint. <strong>During decode</strong>, the phase that produces every token anyone has ever read from a model, the limiting resource on every modern accelerator is memory bandwidth.</p><p>That bandwidth has exactly one industrial source: <strong>High Bandwidth Memory</strong>, towers of DRAM dies thinned to thirty micrometers, drilled through with thousands of copper vias, and bonded onto a logic die millimeters from the GPU. </p><p>Three companies on earth manufacture it at the frontier: <strong>SK hynix</strong> in Icheon and Cheongju, <strong>Samsung</strong> in Pyeongtaek, and <strong>Micron </strong>in Boise, Hiroshima, and Taichung. Which means the marginal cost of intelligence, the dollars per million tokens that every API price and every AI gross margin ultimately rests on, is set<strong> not in Santa Clara</strong> but in a three-supplier memory oligopoly, by stacking yields, bonding chemistry, and wafer allocation.</p><p>The market has noticed, violently. In the first quarter of 2026, SK hynix posted a <strong>72 percent operating margin</strong>, higher than Nvidia&#8217;s and TSMC&#8217;s most recent reported margins, on revenue that nearly tripled year over year; on the earnings call the company said <strong>customer requests for HBM already exceed its planned production</strong> capacity for the next three years, per the Q1 release and call coverage. </p><p>The supplier of the bottleneck is now more profitable, in percentage terms, than the company whose chips it feeds. That inversion is the subject of this report.</p><p>The structure: <strong>the device physics that created the wall (II)</strong>; the wall measured on real silicon, including<strong> the operating taxes nobody quotes (III)</strong>; the anatomy of a stack down to the via, and t<strong>he yield equation that prices it (IV)</strong>; the oligopoly&#8217;s history and<strong> the 2026 supercycle (V)</strong>; HBM4, the <strong>largest architectural break in the technology&#8217;s history, now ramping (VI)</strong>; the machines it feeds and the two opposing design philosophies <strong>the new standard revealed (VII)</strong>; the money, including a <strong>transparent cost-per-stack model (VIII)</strong>; the bridge to the price of a token and <strong>the gigawatt-to-wafer chain (IX)</strong>; the <strong>bear case, stated properly (X); </strong>the <strong>road past HBM4 (XI); </strong>and <strong>what to watch (XII)</strong>, with five falsifiable calls. </p><p>Two interludes price the <strong>KV economy </strong>and the power wall along the way. <strong>Appendix C </strong>specifies the benchmarks that would extend this from synthesis to measurement.</p><p>One framing number before the detail. From the V100 in 2017 to the Rubin GPU now entering production, Nvidia&#8217;s peak tensor throughput at the lowest supported precision grew <strong>roughly 400-fold</strong>. </p><p>Over the same nine years, the memory bandwidth feeding that compute grew <em>24.4-fold. </em>The <strong>16.4x</strong> <strong>divergence </strong>between those exponents, computed from the vendors&#8217; own datasheets, is the memory wall. Everything below is downstream of it.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!3c1e!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F55b5d674-8588-4801-8926-694f9c7700ce_1672x931.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!3c1e!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F55b5d674-8588-4801-8926-694f9c7700ce_1672x931.png 424w, https://substackcdn.com/image/fetch/$s_!3c1e!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F55b5d674-8588-4801-8926-694f9c7700ce_1672x931.png 848w, https://substackcdn.com/image/fetch/$s_!3c1e!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F55b5d674-8588-4801-8926-694f9c7700ce_1672x931.png 1272w, https://substackcdn.com/image/fetch/$s_!3c1e!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F55b5d674-8588-4801-8926-694f9c7700ce_1672x931.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!3c1e!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F55b5d674-8588-4801-8926-694f9c7700ce_1672x931.png" width="1456" height="811" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/55b5d674-8588-4801-8926-694f9c7700ce_1672x931.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:811,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:132213,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/201615837?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F55b5d674-8588-4801-8926-694f9c7700ce_1672x931.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!3c1e!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F55b5d674-8588-4801-8926-694f9c7700ce_1672x931.png 424w, https://substackcdn.com/image/fetch/$s_!3c1e!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F55b5d674-8588-4801-8926-694f9c7700ce_1672x931.png 848w, https://substackcdn.com/image/fetch/$s_!3c1e!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F55b5d674-8588-4801-8926-694f9c7700ce_1672x931.png 1272w, https://substackcdn.com/image/fetch/$s_!3c1e!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F55b5d674-8588-4801-8926-694f9c7700ce_1672x931.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div><hr></div><h2>II. The physics: why the wall exists</h2><p>The wall is not a design oversight. It is the collision of two different scaling regimes, and to price it correctly you have to go down to the cell.</p><p><strong>The cell that cannot shrink.</strong> A DRAM bit is one transistor and one capacitor, the 1T1C cell, unchanged in concept since 1966. Writing stores charge on the capacitor; reading shares that charge onto a long bitline and asks a sense amplifier to detect the disturbance. </p><p>The readable signal is set by charge sharing: the voltage swing the sense amplifier sees is approximately the cell&#8217;s stored swing scaled by Cs over (<em>Cs plus Cbl</em>), the <strong>cell capacitance against the parasitic bitline capacitance</strong>. With modern cell capacitance in the single-digit femtofarads and bitlines loaded by hundreds of neighboring cells, that swing is on the order of tens of millivolts, detected differentially against a reference bitline. </p><p>Shrink the capacitor and the signal disappears into noise; the read becomes unreliable; the bit is worthless. </p><p>So the capacitor holds a roughly fixed charge requirement regardless of lithography, which is <strong>why DRAM capacitors became architecture</strong> rather than printing: vertical pillars with aspect ratios beyond 100 to 1, wells drilled a hundred times deeper than they are wide, lined with high-k dielectric laminates <em>(the zirconia-alumina-zirconia family)</em> to wring capacitance from area that no longer exists in plan view.</p><p>The consequence is the slowest node cadence in semiconductors. The industry&#8217;s DRAM generations crawled from <strong>1x-class</strong> around 19 nanometers in the mid-2010s through <strong>1y</strong>, <strong>1z</strong>, <strong>1a </strong>(<em>roughly 14</em>), <strong>1b </strong><em>(roughly 12 to 13)</em> to today&#8217;s <strong>1c </strong>at roughly 11 to 12 nanometers: a decade to cover what logic crossed in three years. </p><p>EUV lithography, which rescued logic, arrived in DRAM late and thinly: SK hynix introduced it at 1a, Samsung uses it on more layers, and Micron famously held out on <strong>DUV multi-patterning through 1-beta</strong> before adopting EUV at<strong> 1-gamma. </strong></p><p>Bits per wafer, the quantity that sets the cost of a gigabyte, now improves single-digit percent per year. This is the <em>supply-side bedrock </em>of every price in this report: the raw material of memory has nearly stopped getting cheaper.</p><p><strong>The latency that never moved.</strong> Hidden under the bandwidth story is a stagnation worse than the capacity one: <strong>DRAM row cycle time</strong>, tRC, the time to open a row, sense it, restore it, and precharge for the next, has improved by less than 2x in two decades, parked in the <strong>mid-40-nanosecond range</strong>, because it is governed by the analog physics of sense amplification, not by lithography. </p><p>Every bandwidth gain in <strong>modern DRAM</strong> is therefore parallelism wearing a frequency costume: more banks, more bank groups, deeper prefetch <em>(DDR5 fetches sixteen beats per access)</em>, more independent channels, so that thousands of slow rows are in flight at once. HBM is this philosophy at its logical extreme.</p><p><strong>The pin that cannot run.</strong> The other escape route, faster pins, is capped by signal integrity. A DDR5 pin driving centimeters of motherboard trace through a connector tops out in the high single Gbps; <strong>GDDR7 reaches the 30s </strong>only over short, exquisitely tuned point-to-point routes at painful energy cost. </p><p>And energy is the real currency: moving a bit from a board-level DIMM costs on the <strong>order of 10 to 15 picojoules</strong>; GDDR-class interfaces sit near 7 to 8; HBM, with its millimeters-long links through a silicon interposer, runs in the 3 to 5 range, and the logic-process base dies of HBM4 push the interface toward <strong>0.75 to 0.8 volts against 1.1 for DRAM-process predecessors</strong>, roughly doubling interface efficiency, per TSMC&#8217;s published figures (<em>all pJ-per-bit values are approximate vendor-class numbers</em>).</p><p>Run the energy arithmetic forward and it bites hard. At 4 picojoules per bit, fully streaming a <strong>B200&#8217;s 8 TB/s costs about 256 watts</strong>; fully streaming Rubin&#8217;s 22 TB/s at an improved 3 pJ per bit still costs about 528 watts. Memory traffic alone, at full decode throughput on an HBM4 flagship, plausibly draws two-thirds of what an entire H100 board drew. <strong>This is why packages crossed 2,000 watts</strong>, why every HBM4 platform is liquid-cooled, and why the JEDEC thermal envelope is now a first-order economic document.</p><p>So: a cell that cannot shrink, a row that cannot speed up, a pin that cannot run, and an energy budget that punishes distance. The only move left is the one HBM made: <strong>go wide</strong> (<em>1,024 wires, now 2,048</em>), <strong>go short</strong> (<em>millimeters through an interposer</em>), and<strong> go up </strong>(<em>stack the dies</em>). </p><p>Width replaces frequency; proximity replaces drive power; the third dimension replaces the second. The <strong>cost of the trick</strong> is the <em>subject of section IV</em>: stacking is the most yield-hostile thing the memory industry has ever mass-produced.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>III. The wall, quantified, and the taxes nobody quotes</h2><p>The cleanest way to see the wall is the manufacturers&#8217; own flagship specifications, indexed.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!RzVG!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1fb3b92d-80c2-4307-a55b-f20e10fcade9_1747x1026.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!RzVG!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1fb3b92d-80c2-4307-a55b-f20e10fcade9_1747x1026.png 424w, https://substackcdn.com/image/fetch/$s_!RzVG!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1fb3b92d-80c2-4307-a55b-f20e10fcade9_1747x1026.png 848w, https://substackcdn.com/image/fetch/$s_!RzVG!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1fb3b92d-80c2-4307-a55b-f20e10fcade9_1747x1026.png 1272w, https://substackcdn.com/image/fetch/$s_!RzVG!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1fb3b92d-80c2-4307-a55b-f20e10fcade9_1747x1026.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!RzVG!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1fb3b92d-80c2-4307-a55b-f20e10fcade9_1747x1026.png" width="1456" height="855" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/1fb3b92d-80c2-4307-a55b-f20e10fcade9_1747x1026.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:855,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:148973,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/201615837?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1fb3b92d-80c2-4307-a55b-f20e10fcade9_1747x1026.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!RzVG!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1fb3b92d-80c2-4307-a55b-f20e10fcade9_1747x1026.png 424w, https://substackcdn.com/image/fetch/$s_!RzVG!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1fb3b92d-80c2-4307-a55b-f20e10fcade9_1747x1026.png 848w, https://substackcdn.com/image/fetch/$s_!RzVG!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1fb3b92d-80c2-4307-a55b-f20e10fcade9_1747x1026.png 1272w, https://substackcdn.com/image/fetch/$s_!RzVG!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1fb3b92d-80c2-4307-a55b-f20e10fcade9_1747x1026.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><p>Computed from vendor datasheets and <strong>GTC 2026 disclosures</strong>. Note that the compute line rides shrinking precision (FP16 to FP8 to FP4), a legitimate but partly definitional gain; the bandwidth line had to be manufactured stack by stack. </p><p>The <strong>bytes-per-FLOP column</strong> is the punchline: each generation can feed each unit of its arithmetic less data than the last. The machine is increasingly a furnace with a narrowing fuel line.</p><p>Operationally, the two phases of inference live on opposite sides of the roofline. Prefill is a matrix-matrix multiply with arithmetic intensity in the <strong>hundreds of FLOPs per byte</strong>: compute-bound. Decode performs roughly two FLOPs per parameter per token while reading every parameter byte: intensity near 2 at FP8, against a <strong>ridge point around 560 FLOPs per byte on a B200</strong>, so a single conversation idles the arithmetic above 99 percent while it waits on DRAM. </p><p>The entire modern serving stack (<em>continuous batching, paged KV caches, speculative decoding, MoE routing</em>) exists to hide that ratio, and every one of those techniques converts the problem into a different demand on the same resource: <strong>more concurrent streams need more KV cache</strong>, and the KV cache lives in HBM. </p><p>Bandwidth sets the speed of a token; capacity sets how many tokens you can be making at once. Both are the stack.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!m1OJ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5ec83cbc-3f09-49a3-bb8d-896b840af626_1672x950.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!m1OJ!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5ec83cbc-3f09-49a3-bb8d-896b840af626_1672x950.png 424w, https://substackcdn.com/image/fetch/$s_!m1OJ!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5ec83cbc-3f09-49a3-bb8d-896b840af626_1672x950.png 848w, https://substackcdn.com/image/fetch/$s_!m1OJ!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5ec83cbc-3f09-49a3-bb8d-896b840af626_1672x950.png 1272w, https://substackcdn.com/image/fetch/$s_!m1OJ!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5ec83cbc-3f09-49a3-bb8d-896b840af626_1672x950.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!m1OJ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5ec83cbc-3f09-49a3-bb8d-896b840af626_1672x950.png" width="1456" height="827" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/5ec83cbc-3f09-49a3-bb8d-896b840af626_1672x950.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:827,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:119579,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/201615837?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5ec83cbc-3f09-49a3-bb8d-896b840af626_1672x950.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!m1OJ!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5ec83cbc-3f09-49a3-bb8d-896b840af626_1672x950.png 424w, https://substackcdn.com/image/fetch/$s_!m1OJ!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5ec83cbc-3f09-49a3-bb8d-896b840af626_1672x950.png 848w, https://substackcdn.com/image/fetch/$s_!m1OJ!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5ec83cbc-3f09-49a3-bb8d-896b840af626_1672x950.png 1272w, https://substackcdn.com/image/fetch/$s_!m1OJ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5ec83cbc-3f09-49a3-bb8d-896b840af626_1672x950.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>Two operating taxes deserve quantification because they appear in no marketing datasheet.</em></p><p><strong>The refresh tax.</strong> DRAM forgets; every row must be rewritten on a fixed schedule, and while a bank refreshes it cannot serve traffic. The overhead is roughly <em>tRFC over tREFI</em>, the refresh pulse width over the refresh interval: with multi-hundred-nanosecond tRFC on dense dies against the <em>standard 3.9 microsecond interval</em>, the tax is on the order of 5 to 10 percent of theoretical bandwidth (HBM&#8217;s per-bank and managed refresh modes claw some back). </p><p>The vicious part is thermal: JEDEC devices double their refresh rate above 85 degrees Celsius, halving tREFI, so the bandwidth tax roughly doubles exactly when the stack is working hardest and hottest. A 16-high tower dissipating tens of watts through molded underfill in a 2,300-watt package is a device engineered to live near that threshold. Hot memory is slow memory, and slow memory is expensive tokens: cooling budgets are bandwidth budgets.</p><p><strong>The utilization tax.</strong> Achieved bandwidth is not peak. Between refresh, bank conflicts under <em>irregular KV-cache access,</em> read-write turnarounds, and command overheads, well-tuned decode workloads typically realize 60 to 80 percent of datasheet bandwidth (the measurement protocol in Appendix C exists to pin this number per platform). </p><p>Every figure in this report that divides by peak bandwidth is therefore optimistic by that factor, uniformly, which preserves comparisons while flattering absolutes.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Dxwz!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F52ef3999-d33f-4dd7-abb6-69c1e6e26a0c_1672x855.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Dxwz!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F52ef3999-d33f-4dd7-abb6-69c1e6e26a0c_1672x855.png 424w, https://substackcdn.com/image/fetch/$s_!Dxwz!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F52ef3999-d33f-4dd7-abb6-69c1e6e26a0c_1672x855.png 848w, https://substackcdn.com/image/fetch/$s_!Dxwz!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F52ef3999-d33f-4dd7-abb6-69c1e6e26a0c_1672x855.png 1272w, https://substackcdn.com/image/fetch/$s_!Dxwz!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F52ef3999-d33f-4dd7-abb6-69c1e6e26a0c_1672x855.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Dxwz!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F52ef3999-d33f-4dd7-abb6-69c1e6e26a0c_1672x855.png" width="1456" height="745" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/52ef3999-d33f-4dd7-abb6-69c1e6e26a0c_1672x855.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:745,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:82547,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/201615837?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F52ef3999-d33f-4dd7-abb6-69c1e6e26a0c_1672x855.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Dxwz!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F52ef3999-d33f-4dd7-abb6-69c1e6e26a0c_1672x855.png 424w, https://substackcdn.com/image/fetch/$s_!Dxwz!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F52ef3999-d33f-4dd7-abb6-69c1e6e26a0c_1672x855.png 848w, https://substackcdn.com/image/fetch/$s_!Dxwz!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F52ef3999-d33f-4dd7-abb6-69c1e6e26a0c_1672x855.png 1272w, https://substackcdn.com/image/fetch/$s_!Dxwz!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F52ef3999-d33f-4dd7-abb6-69c1e6e26a0c_1672x855.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The wall is also not transient. The longest-baseline study, Gholami and colleagues&#8217; &#8220;<strong>AI and Memory Wall,</strong>&#8221; measured twenty years of server hardware: peak compute scaling 3.0x every two years, DRAM bandwidth 1.6x, interconnect 1.4x. Different exponents compound. </p><p>The wall is structural, and the industry that gets paid because of it is the next several sections.</p><div><hr></div><h2>IV. Anatomy of a stack, to the via</h2><p>A current HBM device is a tower: a base die at the bottom and 8, 12, or 16 DRAM core dies above it. </p><p><strong>JEDEC&#8217;s HBM4 standard</strong>, published as JESD270-4 on April 16, 2025, fixes the envelope: a 2,048-bit interface organized as 32 independent channels (<em>each split into two pseudo-channels, 64 concurrent access streams per stack</em>), support for <strong>24Gb </strong>and <strong>32Gb </strong>core dies in 4-, 8-, 12-, and 16-high stacks, capacities to <strong>64GB per stack</strong>, per-pin rates from 8 Gbps in the base spec, and a package height of 775 micrometers for both 12- and 16-high, <em>loosened from HBM3E&#8217;s 720</em> to give 16-high a fighting chance without new bonding physics.</p><h4>The vertical wiring</h4><p>Each core die is thinned to 30 to 50 micrometers (<em>SK hynix&#8217;s CES 2026 16-high uses 30, about a third of a hair</em>) and pierced by through-silicon vias: <strong>copper columns</strong> roughly 5 to 6 micrometers in diameter, formed via-middle with deep reactive-ion etch, lined, filled, then revealed by grinding the wafer from the back. </p><p>Signal, power, and <strong>ground TSVs</strong> together number on the order of ten thousand per stack (<em>order-of-magnitude; vendors do not publish counts</em>), with spares woven in:<strong> TSV repair logic </strong>in the base die can route around dead vias, one of several redundancy layers (<em>alongside row and column fuses</em>) that keep the yield equation below from being even crueler. </p><p>Between dies, communication crosses<strong> microbump fields</strong> at pitches around 25 micrometers, thousands of joints per interface, every one a potential stack-killing defect.</p><h4>The brain at the bottom</h4><p>The base die is the stack&#8217;s logic: the 2,048-bit PHY facing the host, channel routing, built-in self-test, the<strong> IEEE 1500 test wrapper</strong>, repair control, and the direct-access port that lets a tester exercise the tower. </p><p>Through HBM3E it was built on a DRAM process, because that is what memory firms own, and it shows: DRAM transistors make poor I/O drivers. At <strong>HBM4&#8217;s 10-plus Gbps </strong>per pin across 2,048 lanes, signal integrity demands real logic transistors, real equalization, lower supply rails, which is the engineering reason (<em>beyond the strategic one in section VI</em>) that the base die migrated to foundry logic processes this generation.</p><h4>How the tower is joined</h4><p>The bonding step is the deepest process moat in the industry, currently a three-way technology bet.<strong> </strong></p><p><strong>SK hynix uses MR-MUF</strong>, mass reflow with molded underfill: dies are placed and the solder joints formed in a batch reflow, then the whole stack is encapsulated in one molded underfill shot, a flow with better warpage control, a <strong>stronger thermal path</strong> through the mold compound, and batch throughput, widely credited as the reason it shipped 12-high first and leads yields (t<em>he underfill material itself is a quiet chokepoint, long supplied under an exclusive arrangement with Namics, per trade reporting</em>). </p><p><strong>Samsung </strong>and <strong>Micron </strong>use TC-NCF, thermo-compression over a pre-laminated non-conductive film: each die is pressed down individually with heat and force, slower and stress-accumulating, but precise at fine pitch. </p><p>The <strong>bridge step is fluxless TCB</strong>, removing flux and its residues by bonding in a reducing atmosphere. The endgame is hybrid bonding: copper pads and dielectric planarized to sub-nanometer roughness, fused face to face with no bumps at all, <strong>pitch capability below 10 micrometers</strong>, thinner stacks, a direct copper thermal path, and lower parasitics. </p><p>It is mandatory somewhere past 16 to 20 layers and brutally hard: Samsung, betting on it most aggressively, was reported in April 2026 to be sampling hybrid-bonded 16-high HBM4 to Nvidia at <strong>yields around 10 percent</strong>, while SK hynix completed a 12-high hybrid-bonding validation and placed its first inline production order while publicly committing to MR-MUF through HBM4E, per <strong>EE Times via TrendForce</strong>. </p><p>And in April 2026 JEDEC was reported to be weighing a roughly 900-micrometer height for HBM4E, which would let incumbent bonding survive another generation and <em>shift hundreds of millions of dollars</em> of equipment orders with one standards vote.</p><h4>The equation that prices it all</h4><p>If each die-plus-bond event succeeds with probability <strong>p</strong>, a stack of <strong>n </strong>yields <strong>p </strong>to the power <strong>n</strong>, and one failure scraps the tower with every good die in it. </p><p>At<strong> 99 percent per layer</strong>, a 12-high yields 89 percent; at 97, 69; at 95, 54. This compounding, mitigated but not repealed by known-good-die testing before stacking, TSV and row repair after, and known-good-stack test at the end (<em>the step driving Advantest&#8217;s memory-test boom</em>), is why HBM commands <strong>5 to 6 times the per-bit price of DDR5 </strong>(<em>industry trackers put HBM3E near $8 to $10 per GB, roughly $300 per 36GB stack, with early HBM4 stacks around $500, per Silicon Analysts estimates</em>), and why TrendForce calculates HBM consumes roughly <strong>three times the wafer area per bi</strong>t of commodity DRAM once die-size trades and stack losses are counted. </p><p>Three suppliers, a decade of bonding chemistry embodied in process recipes, and an exponential that punishes newcomers: that is the moat, stated as math.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Lnfr!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F388d5d12-a438-428f-83a2-fbfe55f09010_1672x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Lnfr!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F388d5d12-a438-428f-83a2-fbfe55f09010_1672x893.png 424w, https://substackcdn.com/image/fetch/$s_!Lnfr!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F388d5d12-a438-428f-83a2-fbfe55f09010_1672x893.png 848w, https://substackcdn.com/image/fetch/$s_!Lnfr!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F388d5d12-a438-428f-83a2-fbfe55f09010_1672x893.png 1272w, https://substackcdn.com/image/fetch/$s_!Lnfr!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F388d5d12-a438-428f-83a2-fbfe55f09010_1672x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Lnfr!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F388d5d12-a438-428f-83a2-fbfe55f09010_1672x893.png" width="1456" height="778" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/388d5d12-a438-428f-83a2-fbfe55f09010_1672x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:778,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:134991,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/201615837?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F388d5d12-a438-428f-83a2-fbfe55f09010_1672x893.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Lnfr!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F388d5d12-a438-428f-83a2-fbfe55f09010_1672x893.png 424w, https://substackcdn.com/image/fetch/$s_!Lnfr!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F388d5d12-a438-428f-83a2-fbfe55f09010_1672x893.png 848w, https://substackcdn.com/image/fetch/$s_!Lnfr!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F388d5d12-a438-428f-83a2-fbfe55f09010_1672x893.png 1272w, https://substackcdn.com/image/fetch/$s_!Lnfr!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F388d5d12-a438-428f-83a2-fbfe55f09010_1672x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div><hr></div><h2>V. The oligopoly, and the quarter the memory market broke</h2><p>HBM was born unwanted. AMD&#8217;s packaging architects, led by <strong>Bryan Black</strong>, spent the early 2010s convincing anyone who would listen that memory belonged on the package; SK hynix co-developed the first standard, and the<strong> 2015 Fury X shipped it</strong>, 4GB of HBM1 whose capacity ceiling promptly handicapped the card against Nvidia&#8217;s cheaper GDDR5 flagship. </p><p>The pioneer paid the tuition; the fast follower banked the lesson: Nvidia adopted HBM2 on the P100 in 2016 and never looked back, while for most of a decade the<strong> product line survived at SK hynix on conviction</strong> more than profit. </p><p>The reward arrived all at once after ChatGPT. SK hynix was first to HBM3 (effectively sole-sourcing the H100), first to <strong>8- and 12-high HBM3E</strong>, and converted the lead into a position Counterpoint measured at 62 percent revenue share in Q2 2025, against 21 for Micron, which had skipped HBM3 entirely and <em>leapfrogged to HBM3E</em>, and 17 for Samsung, the incumbent giant caught flat. </p><p>We saw <strong>repeated Nvidia qualification failures</strong> on HBM3E thermals and power through 2024, a leadership change, and a recovery visible by Q3 2025 (<em>Counterpoint: back to 35 percent</em>) that culminated in late January 2026 with Nvidia qualification for HBM4 itself and production from February, per Bloomberg-sourced reporting. </p><p>Analyst estimates after Computex put <strong>SK hynix at 60 to 70 percent </strong>of the HBM4 volume allocated to Vera Rubin, Samsung at 25 to 30, and Micron the remainder, and in early June Nvidia and SK hynix signed a multi-year pact to <strong>co-develop AI memory for Rubin and beyond</strong>, the first agreement of its kind, which converts the leader&#8217;s share into contracted durability, per reporting on the deal; Counterpoint credits SK hynix 61 to 64 percent of the overall HBM market through the period. </p><p>Nvidia has also reportedly asked all three for 16-high stacks as early as late 2026.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!BO9_!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd507b94e-99d4-4d05-8d2d-670469a10f42_1672x893.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!BO9_!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd507b94e-99d4-4d05-8d2d-670469a10f42_1672x893.png 424w, https://substackcdn.com/image/fetch/$s_!BO9_!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd507b94e-99d4-4d05-8d2d-670469a10f42_1672x893.png 848w, https://substackcdn.com/image/fetch/$s_!BO9_!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd507b94e-99d4-4d05-8d2d-670469a10f42_1672x893.png 1272w, https://substackcdn.com/image/fetch/$s_!BO9_!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd507b94e-99d4-4d05-8d2d-670469a10f42_1672x893.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!BO9_!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd507b94e-99d4-4d05-8d2d-670469a10f42_1672x893.png" width="1456" height="778" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d507b94e-99d4-4d05-8d2d-670469a10f42_1672x893.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:778,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:79680,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/201615837?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd507b94e-99d4-4d05-8d2d-670469a10f42_1672x893.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!BO9_!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd507b94e-99d4-4d05-8d2d-670469a10f42_1672x893.png 424w, https://substackcdn.com/image/fetch/$s_!BO9_!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd507b94e-99d4-4d05-8d2d-670469a10f42_1672x893.png 848w, https://substackcdn.com/image/fetch/$s_!BO9_!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd507b94e-99d4-4d05-8d2d-670469a10f42_1672x893.png 1272w, https://substackcdn.com/image/fetch/$s_!BO9_!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd507b94e-99d4-4d05-8d2d-670469a10f42_1672x893.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Then came the quarter that broke the market. </p><p>Because an HBM bit consumes about three times the wafer area of a commodity bit and sells for five times the price, every rational fab starved DDR5 to feed it, with <strong>DDR4 output collapsing toward 20 percent of 2025 levels</strong>; demand met the squeeze, and in Q1 2026, the seasonal trough, global DRAM industry revenue hit $97.1 billion, up 85.3 percent in a single quarter, the <strong>largest in history, per Omdia,</strong> on contract price increases TrendForce recorded at 90 to 95 percent quarter on quarter, the steepest ever, revised up from an already unprecedented 55 to 60.</p><p>The primary-source exhibit is SK hynix&#8217;s Q1 2026 report, and it deserves its numbers stated in full because they are the income-statement proof of everything above: <strong>revenue of 52.5763 trillion won </strong>(<em>about $35.6 billion</em>), the first quarter above 50 trillion in company history, up roughly 60 percent sequentially and 198 percent year over year; operating profit of 37.6103 trillion won (around $25 to 27 billion) at a 72 percent operating margin and a <strong>77 percent net margin</strong>, all-time highs on every line, with one quarter&#8217;s operating profit nearly matching the whole of record fiscal 2025 (47.2 trillion won) and exceeding all of fiscal 2024, per the company&#8217;s release and earnings coverage. </p><p>On the call, management said customer HBM requests already exceed planned capacity for the next three years, <em>guided HBM4E samples for the second half of 2026 with 2027 mass production</em>, and announced <strong>a 19 trillion won </strong>(about $13 billion)<strong> advanced-packaging plant</strong>, with 2026 capex priorities of the M15X ramp, Yongin site preparation, and EUV tooling. </p><p>SK Group&#8217;s chairman went further, telling reporters in March that the global<strong> wafer shortage will likely persist</strong> to 2030 with a shortfall exceeding 20 percent, since capacity takes four to five years to add, per CNBC. </p><p>Micron&#8217;s fiscal first quarter (the November quarter) had already shown the shape: $13.64 billion of revenue, up 57 percent, 56 percent gross margins, demand &#8220;substantially higher&#8221; than supply, followed by its December exit from the consumer memory business entirely. </p><p><strong>Bank of America</strong> frames 2026 as a <strong>supercycle on the order of the 1990s boom</strong>, with DRAM revenue up 51 percent for the year; the three memory makers added roughly $900 billion of combined market value from September, per market reporting.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!H-uf!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd5ab0e2e-2322-4f35-a825-07c27337951f_1672x873.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!H-uf!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd5ab0e2e-2322-4f35-a825-07c27337951f_1672x873.png 424w, https://substackcdn.com/image/fetch/$s_!H-uf!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd5ab0e2e-2322-4f35-a825-07c27337951f_1672x873.png 848w, https://substackcdn.com/image/fetch/$s_!H-uf!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd5ab0e2e-2322-4f35-a825-07c27337951f_1672x873.png 1272w, https://substackcdn.com/image/fetch/$s_!H-uf!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd5ab0e2e-2322-4f35-a825-07c27337951f_1672x873.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!H-uf!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd5ab0e2e-2322-4f35-a825-07c27337951f_1672x873.png" width="1456" height="760" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d5ab0e2e-2322-4f35-a825-07c27337951f_1672x873.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:760,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:87609,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/201615837?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd5ab0e2e-2322-4f35-a825-07c27337951f_1672x873.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!H-uf!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd5ab0e2e-2322-4f35-a825-07c27337951f_1672x873.png 424w, https://substackcdn.com/image/fetch/$s_!H-uf!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd5ab0e2e-2322-4f35-a825-07c27337951f_1672x873.png 848w, https://substackcdn.com/image/fetch/$s_!H-uf!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd5ab0e2e-2322-4f35-a825-07c27337951f_1672x873.png 1272w, https://substackcdn.com/image/fetch/$s_!H-uf!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd5ab0e2e-2322-4f35-a825-07c27337951f_1672x873.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Demand <strong>now reaches upstream</strong> in ways the industry has never seen: in October 2025 OpenAI signed letters of intent with both Samsung and SK hynix under Stargate targeting on the <strong>order of 900,000 DRAM wafer starts per month</strong>, widely characterized as approaching 40 percent of global output, per Reuters. </p><p>Whatever fraction converts, the meaning is the change itself: model companies negotiating two layers down their own supply chain, because they have understood what this report argues. The<em> token supply curve is a wafer allocation.</em></p>
      <p>
          <a href="https://www.thesoftwarefrontier.com/p/how-three-companies-set-the-price">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[The Vertical and the Loop: valuation, compute, and the Anthropic IPO]]></title><description><![CDATA[Anthropic&#8217;s confidential trillion-dollar IPO, the three-silicon bet underneath it, the physics of a token, and the circular machine that pays for the frontier.]]></description><link>https://www.thesoftwarefrontier.com/p/the-vertical-and-the-loop-valuation</link><guid isPermaLink="false">https://www.thesoftwarefrontier.com/p/the-vertical-and-the-loop-valuation</guid><dc:creator><![CDATA[Lorenzo Bradanini]]></dc:creator><pubDate>Wed, 10 Jun 2026 13:42:39 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!OReR!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad172655-a524-421c-9473-baf5db9b44a8_2200x1140.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!OReR!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad172655-a524-421c-9473-baf5db9b44a8_2200x1140.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!OReR!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad172655-a524-421c-9473-baf5db9b44a8_2200x1140.png 424w, https://substackcdn.com/image/fetch/$s_!OReR!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad172655-a524-421c-9473-baf5db9b44a8_2200x1140.png 848w, https://substackcdn.com/image/fetch/$s_!OReR!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad172655-a524-421c-9473-baf5db9b44a8_2200x1140.png 1272w, https://substackcdn.com/image/fetch/$s_!OReR!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad172655-a524-421c-9473-baf5db9b44a8_2200x1140.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!OReR!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad172655-a524-421c-9473-baf5db9b44a8_2200x1140.png" width="1456" height="754" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ad172655-a524-421c-9473-baf5db9b44a8_2200x1140.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:754,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:170744,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/201004624?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad172655-a524-421c-9473-baf5db9b44a8_2200x1140.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!OReR!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad172655-a524-421c-9473-baf5db9b44a8_2200x1140.png 424w, https://substackcdn.com/image/fetch/$s_!OReR!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad172655-a524-421c-9473-baf5db9b44a8_2200x1140.png 848w, https://substackcdn.com/image/fetch/$s_!OReR!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad172655-a524-421c-9473-baf5db9b44a8_2200x1140.png 1272w, https://substackcdn.com/image/fetch/$s_!OReR!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad172655-a524-421c-9473-baf5db9b44a8_2200x1140.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>Reported and estimated figures. Anthropic's S-1 is confidential; none of the below is yet an audited, publicly filed fact.</em></p><div><hr></div><p><strong>A note on the byline.</strong><em> This analysis is written by <strong>The Software Frontier.</strong> We have no access to Anthropic&#8217;s non-public financials, no instruction to flatter the company, and no stake in the outcome. Everything here is drawn from public reporting, third-party regulatory filings, analyst estimates, and published silicon benchmarks, attributed inline. Where the company looks strong we say so. Where the bears have the better argument, we say that too. Treat the byline as a reason for scrutiny, not for trust.</em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Software Frontier is a reader-supported publication.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><h2>Three rockets, one launch window</h2><p>For the first time in the history of American capital markets, three private companies are attempting to go public in a <strong>single year at valuations</strong> above one trillion dollars each. That has never happened once. It is now scheduled to happen three times in roughly one hundred days.</p><p>SpaceX, which by 2026 had combined with xAI into a single entity, filed its <strong>public S-1 on May 20</strong> and is expected to begin trading on June 12, seeking upwards of seventy-five billion dollars, according to Datacenter Dynamics&#8217; reading of the prospectus. </p><p><strong>OpenAI filed</strong> a confidential draft registration on May 22, targeting a fourth-quarter listing that could come as early as September, with Goldman Sachs, Morgan Stanley, and JPMorgan leading, per the Wall Street Journal and CNBC. And on Monday, June 1, Anthropic started its own clock,<strong> confidentially filing an IPO prospectus</strong> with the SEC and confirming it in a public statement. </p><p>Reuters reported back in December that the company had already engaged Wilson Sonsini, the firm that managed Google&#8217;s 2004 IPO, to prepare.</p><p>PitchBook&#8217;s Harrison Rolfes told CNN that two trillion-dollar filings in such a short window represent the largest concentration of pre-IPO capital ever brought to market at once. He was counting two. There are three. </p><p><strong>Wedbush&#8217;s Dan Ives</strong>, who has tracked this complex for years, called the moment an <em>opening of the floodgates</em> for an IPO market that has been largely shut.</p><p>This piece is about the middle rocket. Anthropic is the most interesting of the three not because it is the largest, but because it sits at the exact center of <strong>every structural question</strong> the AI buildout has raised: the quality of AI revenue, the physics and economics of inference, the rivalry between <em>three incompatible silicon stacks</em>, and the circular capital flows that critics compare to the vendor-financing collapse of the dot-com era. To price Anthropic is to price the entire complex.</p><p>There is a complication, and it is the most important sentence in this report. <strong>The filing is confidential.</strong> Under SEC rules for emerging growth companies, the prospectus stays private until roughly fifteen days before a public roadshow. </p><p>Which means that as of this writing, the<strong> most anticipated technology IPO in a generation</strong> is being valued by the market on numbers that no auditor has signed for public release. </p><p>Everything you are about to read about Anthropic&#8217;s financials is reported, estimated, or projected. None of it is yet a filed, audited fact. Hold that thought. We will return repeatedly to why it matters more here than almost anywhere else.</p><div><hr></div><h2>The vertical</h2><p>Start with the revenue curve, because it is the reason any of this is happening, and because it does not look like a curve. <mark>It looks like a wall.</mark></p><p>Anthropic reported roughly <strong>one billion dollars in annualized revenue</strong> at the end of 2024. By the end of 2025 that figure was around nine billion. After closing its Series G in February 2026 it was near fourteen billion. By early April, multiple outlets put the run-rate above thirty billion. </p><p>In May, per CNBC&#8217;s account of the filing, Anthropic disclosed a<strong> run-rate of forty-seven billion dollars</strong>, a figure echoed at roughly forty-five billion by the research firm Sacra.</p><p>Read that sequence again: <mark>one, nine, fourteen, thirty, forty-seven.</mark> The jump from fourteen to forty-seven happened in about four months. There is no precedent for this in enterprise software. </p><p>The closest analogue is<strong> not a software company at all. </strong>It is a commodity in a shortage, which is exactly what frontier inference capacity has become.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!ZwXF!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e3a4ba2-3f48-4789-9658-1c48a3c889e3_1720x940.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!ZwXF!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e3a4ba2-3f48-4789-9658-1c48a3c889e3_1720x940.png 424w, https://substackcdn.com/image/fetch/$s_!ZwXF!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e3a4ba2-3f48-4789-9658-1c48a3c889e3_1720x940.png 848w, https://substackcdn.com/image/fetch/$s_!ZwXF!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e3a4ba2-3f48-4789-9658-1c48a3c889e3_1720x940.png 1272w, https://substackcdn.com/image/fetch/$s_!ZwXF!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e3a4ba2-3f48-4789-9658-1c48a3c889e3_1720x940.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!ZwXF!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e3a4ba2-3f48-4789-9658-1c48a3c889e3_1720x940.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2e3a4ba2-3f48-4789-9658-1c48a3c889e3_1720x940.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;c-revenue&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="c-revenue" title="c-revenue" srcset="https://substackcdn.com/image/fetch/$s_!ZwXF!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e3a4ba2-3f48-4789-9658-1c48a3c889e3_1720x940.png 424w, https://substackcdn.com/image/fetch/$s_!ZwXF!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e3a4ba2-3f48-4789-9658-1c48a3c889e3_1720x940.png 848w, https://substackcdn.com/image/fetch/$s_!ZwXF!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e3a4ba2-3f48-4789-9658-1c48a3c889e3_1720x940.png 1272w, https://substackcdn.com/image/fetch/$s_!ZwXF!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e3a4ba2-3f48-4789-9658-1c48a3c889e3_1720x940.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>The wall.</strong> Run-rate roughly 5x&#8217;d in twelve months and more than 3x&#8217;d in the four months to May. Sources: CNBC, Sacra, The Motley Fool, MEXC syndication, TipRanks. Intermediate points (Aug, Oct 2025) from contemporaneous reporting.</figcaption></figure></div><p>The composition matters more than the level. Roughly eighty percent of revenue comes from enterprises, per TipRanks, almost the inverse of OpenAI&#8217;s consumer-heavy base. Futurum&#8217;s Nick Patience notes that<strong> eight of the Fortune 10</strong> are now paying customers. More than three hundred thousand businesses run Claude. </p><p>The count of customers spending a million dollars or more per year crossed one thousand by April, double the roughly five hundred in February. And one product, <strong>Claude Code</strong>, reached about one billion dollars in annualized revenue within six months of launch, a developer-tool adoption speed that AI Weekly, citing WSJ figures, called a new category benchmark.</p><p>That last fact is<strong> double-edged</strong>, and we will sharpen the second edge later. For now, note the shape: this is not a consumer novelty that might churn. It is a deeply embedded enterprise dependency growing into core workflows. </p><p>That is the <strong>most durable kind of revenue </strong>there is. It is also the kind that invites a backlash when the bill arrives, which is precisely what is now starting to happen.</p><div><hr></div><h2>The margin question, which is really two companies</h2><p>The number everyone quotes is revenue. The number that decides whether Anthropic is worth a trillion dollars is<strong> gross margin</strong>, and on gross margin Anthropic is two entirely different businesses wearing the same name.</p><p>In 2024, Anthropic&#8217;s gross margin was <strong>negative ninety-four percent</strong>, per data compiled by TradingKey. It cost nearly two dollars of compute to deliver one dollar of revenue. By 2025 that had swung to somewhere around forty to fifty percent. </p><p>The company&#8217;s own projection, reported by Sacra and Seeking Alpha, is for gross margin to reach roughly seventy-seven percent by 2028.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!bz_K!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ae14ac4-d63d-44e8-b7b1-6f499fae5b36_1720x940.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!bz_K!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ae14ac4-d63d-44e8-b7b1-6f499fae5b36_1720x940.png 424w, https://substackcdn.com/image/fetch/$s_!bz_K!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ae14ac4-d63d-44e8-b7b1-6f499fae5b36_1720x940.png 848w, https://substackcdn.com/image/fetch/$s_!bz_K!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ae14ac4-d63d-44e8-b7b1-6f499fae5b36_1720x940.png 1272w, https://substackcdn.com/image/fetch/$s_!bz_K!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ae14ac4-d63d-44e8-b7b1-6f499fae5b36_1720x940.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!bz_K!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ae14ac4-d63d-44e8-b7b1-6f499fae5b36_1720x940.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/3ae14ac4-d63d-44e8-b7b1-6f499fae5b36_1720x940.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;c-margin&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="c-margin" title="c-margin" srcset="https://substackcdn.com/image/fetch/$s_!bz_K!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ae14ac4-d63d-44e8-b7b1-6f499fae5b36_1720x940.png 424w, https://substackcdn.com/image/fetch/$s_!bz_K!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ae14ac4-d63d-44e8-b7b1-6f499fae5b36_1720x940.png 848w, https://substackcdn.com/image/fetch/$s_!bz_K!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ae14ac4-d63d-44e8-b7b1-6f499fae5b36_1720x940.png 1272w, https://substackcdn.com/image/fetch/$s_!bz_K!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ae14ac4-d63d-44e8-b7b1-6f499fae5b36_1720x940.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Two businesses.</strong> A dollar of 45%-margin revenue and a dollar of 77%-margin revenue are not the same asset in any discounted cash flow. The entire valuation thesis rests on which one is emerging. Sources: TradingKey (2024), Sacra and Seeking Alpha (2025, 2028E).</figcaption></figure></div><p>A Substack analysis by <strong>Shanaka Anslem Perera</strong> framed the stakes more precisely than any sell-side note I have read: revenue at forty percent gross margin and revenue at seventy-seven percent gross margin are not the same business with different costs. </p><p><strong>They are different businesses</strong>. The whole valuation rests on the assumption that the second business is the one arriving. The bear case rests on the possibility that it is not.</p><p><em>Why might margins expand so dramatically?</em> The answer is not hand-waving. It is mechanical, it is grounded in the physics of how a token is produced, and it is the strongest single pillar of the bull case. Section VII takes it apart piece by piece. </p><p>But first we have to understand what Anthropic actually runs Claude on, because the margin story and the silicon story are the same story.</p><div><hr></div><h2>The profitability that may not be one</h2><p>In mid-May, the Wall Street Journal reported that Anthropic is on pace to post its first operating profit in the second quarter of 2026: <strong>roughly $10.9 billion in revenue</strong> against an expected operating profit of about $559 million, more than doubling the $4.8 billion booked in Q1.</p><p><strong>AI Weekly&#8217;s summary </strong>of the WSJ figures framed this as the moment Anthropic crosses into covering its own costs, a milestone that, if real, changes the fundraising calculus for the entire frontier sector. Investors could finally model a path to returns rather than an indefinite subsidy.</p><p>The technology critic Ed Zitron argues the milestone is partly an artifact of timing, and the argument deserves engagement rather than dismissal. The mechanism: under the <strong>compute deal Anthropic struck with the xAI unit of SpaceX</strong>, Anthropic pays $1.25 billion per month for the Colossus 1 cluster, a figure that emerged from SpaceX&#8217;s own S-1 and was reported by TechCrunch. </p><p>That is roughly fifteen billion dollars a year. But the deal carries a discounted rate for the first two months while xAI completes its ramp, and those two discounted months fall, conveniently, in exactly the quarter Anthropic is using to claim its first operating profit.</p><p>Zitron&#8217;s point is not that Anthropic is lying. It is that a 559-million-dollar operating profit is a thin margin on a temporarily depressed cost base, and that when the <strong>SpaceX rate steps up </strong>to its full level, the same quarter&#8217;s economics look different. </p><p>Whether you find this damning or merely worth watching depends on your priors. <strong>What is not in dispute</strong> is that the profitability claim and the compute deal are entangled, and that a confidential S-1 means we cannot yet see how the company books the ramp discount. </p><p>This is the first concrete reason the confidentiality matters: the single most important narrative claim, that Anthropic is now profitable, sits on a cost structure we cannot inspect.</p><div><hr></div><h2>The three-silicon bet</h2><p>Now the part that, in the end, actually decides the future: where the compute comes from, what it runs on, and who controls it.</p><p>Anthropic has done something <strong>no other frontier lab has managed</strong>. It trains and serves Claude across three mutually incompatible silicon stacks at once: Amazon&#8217;s Trainium, Google&#8217;s TPU, and Nvidia&#8217;s GPUs (<em>the latter rented, remarkably, from a rival</em>). </p><p>This is <strong>not an accident of procurement</strong>. It is a deliberate hedge against the single greatest risk a frontier lab faces, which is being captive to one supplier&#8217;s roadmap, pricing, and power budget. </p><p>Anthropic CFO Krishna Rao framed the <strong>multi-vendor approach</strong> to CNBC as spreading workloads across vendors to tune for price, performance, and power. Read it as insurance.</p><h3>Amazon and Trainium</h3><p>On April 20 and 21, Amazon and Anthropic announced an expanded partnership that, per <strong>Global Data Center Hub&#8217;s reconstruction</strong>, totals up to thirty-three billion dollars in committed Amazon equity: <em>five billion in fresh equity at a $350 billion pre-money valuation</em>, up to twenty billion more tied to milestones, on top of eight billion deployed from 2023 to 2025. </p><p>In exchange, Anthropic committed to spend more than one hundred billion dollars on AWS over the next decade and to <strong>deploy up to five gigawatts of Trainium capacity</strong>. The engine is Project Rainier, which Anthropic&#8217;s own announcement describes as already running over one million Trainium2 chips. </p><p>The<strong> Indiana campus</strong> alone, per Global Data Center Hub, spans twelve hundred acres across seven buildings and scales toward 2.2 gigawatts at full build-out. </p><p>The commitment spans Trainium2, Trainium3, and the still-unannounced Trainium4, with <strong>Amazon&#8217;s Annapurna Labs </strong>taking design feedback from the lab that stresses the chips hardest.</p><h3>Google and TPUs</h3><p>The <em>Google relationship is older </em>and, in valuation terms, arguably the better bargain. CNBC reported that before the latest deal Google&#8217;s stake exceeded three billion dollars at roughly fourteen percent, built from a <strong>three-hundred-million-dollar 2023 check</strong> for about ten percent plus a two-billion-dollar follow-on. </p><p>In October 2025 the two announced a cloud deal for up to one million TPUs worth tens of billions. Then on April 24, 2026, Google committed up to forty billion dollars more in cash and compute, per TechCrunch and CNBC, expanding to<strong> five gigawatts of TPU capacity</strong> over five years. A Broadcom securities filing put the associated next-generation TPU figure at 3.5 gigawatts. </p><p>The Motley Fool&#8217;s<strong> Billy Duberstein</strong> argued Google is getting a screaming bargain, partly because a guaranteed five-gigawatt anchor tenant de-risks Alphabet&#8217;s own enormous capex. He is probably right, and the reason he is right is the same reason the structure is fragile, which is <strong>the subject</strong> of future discussions.</p><h3>Nvidia, via xAI&#8217;s Colossus</h3><p>This is the strangest leg, and the most revealing. On May 6, at its own Code with Claude developer conference, Anthropic announced it would take essentially all the compute at Colossus 1, the<strong> Memphis supercomputer</strong> built by xAI and now owned by the merged SpaceX entity.</p><p>Per xAI&#8217;s release, Colossus 1 houses over 220,000 Nvidia GPUs spanning H100, H200, and GB200 accelerators, roughly 300 megawatts. The timing was not accidental: with usage of xAI&#8217;s own Grok having dropped, Colossus 1 sat underused, which is what freed its full capacity for Anthropic, per<strong> TechCrunch&#8217;s</strong> reading of the S-1. </p><p>Anthropic&#8217;s chief compute officer Tom Brown said on the record that the company would expand onto <strong>Nvidia GB200 capacity </strong>in the larger Colossus 2 through June, per Axios. TechCrunch later reported, from SpaceX&#8217;s S-1, that Anthropic will pay <strong>$1.25 billion per month through May 2029</strong>, a deal that could bring the Musk entity over forty billion dollars in revenue, with either side able to terminate on ninety days&#8217; notice. </p><p>The arrangement is pointed at consumer capacity, directly improving Claude Pro and Claude Max. The irony is thick enough to cut: Anthropic, the lab founded by OpenAI defectors, is<strong> now renting its consumer compute from Elon Musk</strong>, who wrote on X in February that the company <em>hates Western civilization</em>, per CNBC. Business is business.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!N9Xv!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7d7f7241-09e4-4344-8f01-e02513837998_1720x940.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!N9Xv!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7d7f7241-09e4-4344-8f01-e02513837998_1720x940.png 424w, https://substackcdn.com/image/fetch/$s_!N9Xv!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7d7f7241-09e4-4344-8f01-e02513837998_1720x940.png 848w, https://substackcdn.com/image/fetch/$s_!N9Xv!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7d7f7241-09e4-4344-8f01-e02513837998_1720x940.png 1272w, https://substackcdn.com/image/fetch/$s_!N9Xv!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7d7f7241-09e4-4344-8f01-e02513837998_1720x940.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!N9Xv!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7d7f7241-09e4-4344-8f01-e02513837998_1720x940.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/7d7f7241-09e4-4344-8f01-e02513837998_1720x940.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;c-capacity&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="c-capacity" title="c-capacity" srcset="https://substackcdn.com/image/fetch/$s_!N9Xv!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7d7f7241-09e4-4344-8f01-e02513837998_1720x940.png 424w, https://substackcdn.com/image/fetch/$s_!N9Xv!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7d7f7241-09e4-4344-8f01-e02513837998_1720x940.png 848w, https://substackcdn.com/image/fetch/$s_!N9Xv!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7d7f7241-09e4-4344-8f01-e02513837998_1720x940.png 1272w, https://substackcdn.com/image/fetch/$s_!N9Xv!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7d7f7241-09e4-4344-8f01-e02513837998_1720x940.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>The hedge, in gigawatts.</strong> Roughly 1.5 to 2 GW is live today; the bars show contracted ceilings. A 1 GW data center costs near $50B, of which about $35B is chips (CNBC). Sources: Anthropic, Amazon, Google, xAI, Broadcom filings, TechCrunch.</figcaption></figure></div><div><hr></div><h2>Chips, racks, and the CUDA moat</h2><p>If you read only one section as an engineer, read this one and the next. The headline question is simple to state and hard to answer: </p><div class="pullquote"><p>can custom silicon actually serve a frontier model, or is Nvidia&#8217;s lead structural? </p></div><p>Anthropic is the live experiment, and the early data is more interesting than either camp admits.</p><p>Start at the chip. <strong>AWS shipped Trainium3</strong> in December 2025 on TSMC&#8217;s 3nm N3P node, the most advanced process in any shipping AI accelerator, per <strong>Tom&#8217;s Hardware</strong> and the spec compilations at IntuitionLabs and Awesome Agents, reaching broad availability in early 2026. </p><p>Each Trainium3 chip delivers about 2.52 petaflops of MXFP8 compute with 144 GB of HBM3e and 4.9 TB/s of memory bandwidth, with eight NeuronCore-v4 engines and a <strong>NeuronLink-v4</strong> interconnect at <strong>2 TB/s. </strong>AWS claims roughly 2x the per-chip compute of Trainium2, rising to about 4.4x at the 144-chip UltraServer level, with 4x better energy efficiency. </p><p><strong>Google&#8217;s TPU v7</strong>, codenamed Ironwood and announced in 2025, delivers about 4.6 petaflops of FP8 per chip with<strong> 192 GB of HBM and 7.37 TB/s</strong> of bandwidth, which analysts at Introl described as on par with Blackwell. </p><p>Nvidia&#8217;s B200, by contrast, delivers <mark>roughly 9 petaflops of FP8 per chip with sparsity</mark> (about 4.5 dense), per Nvidia&#8217;s datasheet. Its headline figure <strong>near 18 to 20 petaflops is an FP4 number</strong>, a lower-precision format, so comparing like for like at FP8 is the only fair reading.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!TQUG!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F951e9236-3ac4-411c-a7fa-acc549e85145_1720x940.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!TQUG!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F951e9236-3ac4-411c-a7fa-acc549e85145_1720x940.png 424w, https://substackcdn.com/image/fetch/$s_!TQUG!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F951e9236-3ac4-411c-a7fa-acc549e85145_1720x940.png 848w, https://substackcdn.com/image/fetch/$s_!TQUG!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F951e9236-3ac4-411c-a7fa-acc549e85145_1720x940.png 1272w, https://substackcdn.com/image/fetch/$s_!TQUG!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F951e9236-3ac4-411c-a7fa-acc549e85145_1720x940.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!TQUG!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F951e9236-3ac4-411c-a7fa-acc549e85145_1720x940.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/951e9236-3ac4-411c-a7fa-acc549e85145_1720x940.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;c-chipflops&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="c-chipflops" title="c-chipflops" srcset="https://substackcdn.com/image/fetch/$s_!TQUG!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F951e9236-3ac4-411c-a7fa-acc549e85145_1720x940.png 424w, https://substackcdn.com/image/fetch/$s_!TQUG!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F951e9236-3ac4-411c-a7fa-acc549e85145_1720x940.png 848w, https://substackcdn.com/image/fetch/$s_!TQUG!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F951e9236-3ac4-411c-a7fa-acc549e85145_1720x940.png 1272w, https://substackcdn.com/image/fetch/$s_!TQUG!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F951e9236-3ac4-411c-a7fa-acc549e85145_1720x940.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>At the chip level Nvidia leads, but by less than the marketing implies.</strong> At comparable FP8 precision, a single B200 (~9 PF, with sparsity) is roughly 3.6x a Trainium3 die (2.52 PF) and under 2x a TPU v7 (4.61 PF); at dense FP8 the B200 (~4.5 PF) barely edges the TPU. The often-quoted 8x gap compares B200 FP4 against Trainium3 FP8, which is not like for like. Sources: AWS, Google, Nvidia datasheet via Civo, CudoCompute, Tom&#8217;s Hardware.</figcaption></figure></div><p>Now move up one level, to the rack, where AI is actually deployed. AWS packs 144 Trainium3 chips into a liquid-cooled Trn3 Gen2 UltraServer that delivers <strong>roughly 362 petaflops of FP8, 20.7 TB of HBM3e</strong>, and an aggregate 705.6 TB/s of memory bandwidth. </p><p>Per Tom&#8217;s Hardware and Oplexa&#8217;s analysis, that puts the <strong>UltraServer </strong>essentially level with Nvidia&#8217;s flagship GB300 NVL72 at rack scale, at an estimated fifty percent lower cost per workload and roughly forty percent better energy efficiency.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!jqqu!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F924d1724-a69c-4512-9989-b328f52a38f6_1720x940.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!jqqu!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F924d1724-a69c-4512-9989-b328f52a38f6_1720x940.png 424w, https://substackcdn.com/image/fetch/$s_!jqqu!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F924d1724-a69c-4512-9989-b328f52a38f6_1720x940.png 848w, https://substackcdn.com/image/fetch/$s_!jqqu!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F924d1724-a69c-4512-9989-b328f52a38f6_1720x940.png 1272w, https://substackcdn.com/image/fetch/$s_!jqqu!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F924d1724-a69c-4512-9989-b328f52a38f6_1720x940.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!jqqu!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F924d1724-a69c-4512-9989-b328f52a38f6_1720x940.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/924d1724-a69c-4512-9989-b328f52a38f6_1720x940.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;c-rack&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="c-rack" title="c-rack" srcset="https://substackcdn.com/image/fetch/$s_!jqqu!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F924d1724-a69c-4512-9989-b328f52a38f6_1720x940.png 424w, https://substackcdn.com/image/fetch/$s_!jqqu!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F924d1724-a69c-4512-9989-b328f52a38f6_1720x940.png 848w, https://substackcdn.com/image/fetch/$s_!jqqu!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F924d1724-a69c-4512-9989-b328f52a38f6_1720x940.png 1272w, https://substackcdn.com/image/fetch/$s_!jqqu!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F924d1724-a69c-4512-9989-b328f52a38f6_1720x940.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>At the rack level, it is a tie, at roughly half the cost per workload.</strong> Amazon closes a roughly 3 to 4x per-chip FP8 gap with integration and scale-up fabric. This is the entire competitive argument for custom silicon, and Anthropic is the proof of concept. Source: Oplexa, Tom&#8217;s Hardware.</figcaption></figure></div><p><strong>This is the crux</strong> that most coverage misses. The custom-silicon competition is not happening at the transistor. It is happening at the system and at the dollar. </p><p><strong>Amazon and Google </strong>close a brutal single-chip deficit through dense integration, liquid cooling, and proprietary scale-up fabrics: Nvidia calls its fabric NVLink, Google calls its ICI, <strong>AWS calls its NeuronLink</strong>. Once you are buying racks rather than chips, and once you weight by cost and power rather than raw flops, the gap collapses.</p><p>Memory is the other half of the story, and it is the half that governs inference. As we will show in detail, modern serving is <strong>memory-bandwidth-bound</strong>, not compute-bound, which is why the per-token cost curve is so sensitive to HBM. </p><p>Here the three are closer than the compute numbers suggest, and Nvidia&#8217;s lead is narrower.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!WGHs!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3eeb355-320b-468b-bcab-18c892eb5da9_1720x940.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!WGHs!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3eeb355-320b-468b-bcab-18c892eb5da9_1720x940.png 424w, https://substackcdn.com/image/fetch/$s_!WGHs!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3eeb355-320b-468b-bcab-18c892eb5da9_1720x940.png 848w, https://substackcdn.com/image/fetch/$s_!WGHs!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3eeb355-320b-468b-bcab-18c892eb5da9_1720x940.png 1272w, https://substackcdn.com/image/fetch/$s_!WGHs!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3eeb355-320b-468b-bcab-18c892eb5da9_1720x940.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!WGHs!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3eeb355-320b-468b-bcab-18c892eb5da9_1720x940.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d3eeb355-320b-468b-bcab-18c892eb5da9_1720x940.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;c-bandwidth&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="c-bandwidth" title="c-bandwidth" srcset="https://substackcdn.com/image/fetch/$s_!WGHs!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3eeb355-320b-468b-bcab-18c892eb5da9_1720x940.png 424w, https://substackcdn.com/image/fetch/$s_!WGHs!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3eeb355-320b-468b-bcab-18c892eb5da9_1720x940.png 848w, https://substackcdn.com/image/fetch/$s_!WGHs!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3eeb355-320b-468b-bcab-18c892eb5da9_1720x940.png 1272w, https://substackcdn.com/image/fetch/$s_!WGHs!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3eeb355-320b-468b-bcab-18c892eb5da9_1720x940.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>On the metric that decides inference economics, the field is tight.</strong> HBM capacity per chip runs 144 GB (Trainium3), 192 GB (TPU v7), and 288 GB (Nvidia B300). For memory-bound serving, this parity is why non-Nvidia inference is viable. Source: vendor specs via Tom&#8217;s Hardware, IntuitionLabs.</figcaption></figure></div><h3>Inside the die: systolic arrays versus SIMT</h3><p>The reason a custom accelerator can rival a GPU it trails on paper comes down to dataflow. </p><p>Each Trainium3 chip carries eight NeuronCore-v4 engines, and each engine is<strong> built around systolic arrays</strong>: a 128 by 128 grid for BF16 and a wider 512 by 128 grid for MXFP8, backed by 32 MiB of on-core SRAM. A systolic array is the opposite of a general-purpose processor. </p><p>Weights are loaded into the grid and held stationary while activations pulse through it, so a<strong> value read once from SRAM </strong>is reused across an entire row or column of multiply-accumulate units before it retires. </p><p>On the <strong>dense matrix multiplies</strong> that dominate a transformer, that weight-stationary dataflow keeps the multiply-accumulate units near full occupancy while spending almost nothing on instruction fetch, register-file traffic, or <strong>cache coherence</strong>. Google&#8217;s TPU is the same idea at a different size: one very large matrix-multiply unit fed by a compiler that schedules the entire computation ahead of time.</p><p>Nvidia&#8217;s Blackwell takes the other path. Its streaming multiprocessors are <strong>SIMT machines</strong>: thousands of threads grouped into warps, with tensor cores doing the matrix math and a deep hierarchy of schedulers, register files, and caches feeding them. </p><p>That flexibility is the point. A <strong>GPU runs irregular control flow</strong>, dynamic shapes, sparsity, and an enormous library surface, on hardware that was never specialized to one operator. The cost of that generality is silicon and power spent on everything that is not the multiply. For the long tail of workloads, the flexibility earns its keep. For a transformer decode loop, much of it sits idle.</p><p>This is the precise reason the <strong>rack-level parity</strong> in the charts above is real rather than a marketing artifact. A frontier lab runs essentially one workload shape, the <strong>transformer</strong>, and writes its own kernels, so it does not need most of what a GPU spends transistors on, and it can drive a systolic array to a utilization a general user could never reach. </p><p>The cost it pays is the programming model. <strong>Fifteen years of CUDA</strong>, of PTX and SASS-level tuning, of cuDNN and CUTLASS and a developer base in the millions, <strong>has no equivalent on the other side</strong>. AWS answers with the Neuron SDK and its NKI kernel interface plus JAX and PyTorch support; Google answers with <strong>XLA</strong>, JAX, and Pallas. </p><p>A team with kernel engineers can reach high utilization on any of the three. An enterprise that only wants a <strong>model behind an API cannot</strong>, which is why the moat protects Nvidia at the bottom of the market and erodes at the very top, where Anthropic operates.</p><p>The last equalizer is the fabric. A single chip never serves a frontier model alone, so what matters is how fast many chips behave as one. Nvidia&#8217;s NVLink 5 moves about 1.8 terabytes per second per GPU through an <strong>NVSwitch fabric;</strong> AWS NeuronLink-v4 moves about 2 terabytes per second per chip; Google&#8217;s ICI wires its pods into a three-dimensional torus. </p><p>Tensor parallelism forces every chip to exchange a slice of activations on every layer, and expert parallelism in a<strong> mixture-of-experts model </strong>adds an all-to-all shuffle of tokens to their chosen experts, so scale-up bandwidth, not raw per-chip flops, is what lets 144 Trainium3 act like one giant accelerator with<strong> 706 terabytes per second </strong>of aggregate memory bandwidth. The per-chip FLOPS gap is what the slides show. The scale-up fabric is what the workload feels.</p><p><em>So why has nobody else pulled this off? </em>Because the real moat was never the silicon. It is the software. Nvidia&#8217;s CUDA is fifteen years of libraries, kernels, compilers, and developer muscle memory. </p><p>AWS counters with the Neuron SDK and JAX and <strong>PyTorch support;</strong> Google has its own mature stack. For the long tail of enterprises, porting off CUDA is a non-starter: the engineering cost dwarfs the hardware savings. </p><p>But a frontier lab is the one customer that can pay that cost, because it writes its own kernels, owns its own stack, and has the systems talent to make a<strong> non-CUDA chip productive</strong>. </p><p>That is <strong>precisely why Anthropic is the perfect partner</strong> to validate custom silicon, and why Amazon and Google paid tens of billions in equity to make it their anchor tenant. Anthropic reportedly fed design input directly into Trainium3, per Awesome Agents. </p><p>The lab is not just renting the chips. It is co-designing the thing meant to dethrone the incumbent.</p><p>One quiet beneficiary deserves a name: Broadcom. It co-designs Google&#8217;s TPU, supplies the connectivity silicon that stitches these racks together, and, per The Motley Fool citing Broadcom&#8217;s own disclosures, <strong>booked roughly a ten-billion-dollar TPU order </strong>plus an additional eleven billion in hardware tied to Anthropic. </p><p>In a gold rush, the <strong>firm selling the most shovels</strong> to the most miners is often the cleaner trade. Hold that for the market section.</p><p>Caveat where it is due. Nvidia&#8217;s roadmap does not stand still: <strong>B200 and GB200 </strong>are reportedly sold out through mid-2026 against a backlog near 3.6 million units, per IntuitionLabs, and Vera Rubin, due in the second half of 2026 with HBM4 and a <strong>Rubin NVL144 rack</strong> at 3.6 exaflops of dense FP4, extends the lead at the top. </p><p>The near-term threat to Nvidia is not revenue. Demand still dwarfs supply. The threat is the long-run margin structure that depends on hyperscalers having no realistic alternative. Anthropic&#8217;s <strong>three-silicon bet</strong> is the clearest signal yet that the alternative is becoming real.</p><div><hr></div><h2>The cost of a token</h2><p>The bull case for Anthropic&#8217;s margin is a claim about physics and software, <strong>not accounting</strong>. To judge it you have to understand how a single token is actually produced, and why the cost of producing it is collapsing roughly tenfold a year. </p><p>This is the <strong>section the financial press cannot write</strong> and your readers care about most.</p><h3>Two phases, one bottleneck</h3><p>Serving a language model has two phases with opposite cost structures. <strong>Prefill</strong> ingests the prompt: every input token is processed in parallel, the arithmetic is dense, and the accelerator runs near its compute ceiling. Prefill is compute-bound, and it is cheap per token because parallelism is high. </p><p><strong>Decode</strong> generates the answer one token at a time, autoregressively, and here is the trap: to produce each new token, the hardware must stream the model&#8217;s entire active weight set out of high-bandwidth memory. </p><p><mark>Decode is therefore </mark><strong><mark>memory-bandwidth-bound</mark></strong><mark>, not compute-bound.</mark> As the Inworld and Spheron benchmark teardowns put it, reading model weights during decode is the primary bottleneck for autoregressive generation. </p><p>This single fact is why the <strong>memory-bandwidth chart</strong> in the previous section matters more than the flops chart, and why a Trainium3 or a TPU v7, which trail Nvidia badly on raw flops but sit within striking distance on bandwidth, can serve inference competitively. <mark>The frontier is not compute-starved. It is bandwidth-starved.</mark></p><p>Two structures sit on top of this. The <strong>KV cache</strong> stores the attention keys and values for every prior token, so it grows with context length and with batch size, and it competes for the same <strong>scarce HBM capacity </strong>and bandwidth; long-context serving is expensive precisely because, as CloudRift&#8217;s benchmarks note, it stresses KV-cache traffic. </p><p>And <strong>batching</strong> is the lever that makes serving economic at all: by processing many requests together, each expensive weight-read from HBM is amortized across many sequences, so throughput, and therefore cost per token, depends enormously on how full the batch is. </p><p><strong>GMI Cloud&#8217;s teardown</strong> makes the point concrete: an H100 running Llama 70B in FP8 generates roughly two to three thousand tokens per second at batch 32, about $0.19 to $0.29 per million output tokens, and the same GPU at fifty percent utilization sees its effective cost per token roughly double. Utilization is not a footnote. It is half the unit economics.</p><h3>Three levers that crush cost</h3><p>An academic study circulating on arXiv this year, modeling what it calls a tiered <strong>Super-Moore effect</strong>, decomposes inference cost into hardware, labor, and a technology index that captures architectural innovation. Its key finding is that the technology index has improved far faster than the hardware alone. Three levers do most of the work.</p><p><strong>Quantization.</strong> Moving the numerical format of the weights from FP16 to FP8 to FP4 or INT4 halves the bytes per parameter at each step. Because decode is bandwidth-bound, fewer bytes per weight means more tokens per second on the same silicon, almost linearly. </p><p>Spheron measures FP8 cutting effective cost per token by roughly half on H100 and H200 by <strong>doubling throughput with no extra GPUs</strong>; Blackwell&#8217;s FP4 tensor cores push it further still. Quantization is close to a free lunch until model quality degrades, and the frontier labs have become expert at quantizing right up to that line.</p><p><strong>Mixture of experts.</strong> A dense model activates all its parameters for every token. A mixture-of-experts model routes each token through only a small subset. </p><p>The <strong>arXiv study</strong> quantifies the canonical example: DeepSeek&#8217;s architecture carries <em>671 billion total parameters but activates only 37 billion per token</em>, <mark>an eighteen-fold reduction in per-token compute</mark> with no proportionate quality loss, and it operates independently of any hardware trend. </p><p>MoE is the <strong>single largest architectural reason </strong>a frontier-class answer is no longer a frontier-class expense.</p><p><strong>Algorithmic serving.</strong> FlashAttention removed the memory bottleneck inside the attention kernel; speculative decoding, which drafts several tokens with a small model and verifies them with the large one, cuts latency two to three times, per Introl; continuous batching and paged KV caches keep utilization high. <strong>None of these require new chips</strong>. They are software, and software ships continuously.</p><p>The hardware levers compound on top: Blackwell&#8217;s B200 carries 2.4 times the <strong>memory bandwidth of an H100</strong> and enough capacity (192 GB) to hold models up to roughly 96 billion parameters in FP16, or 192 billion in FP8, on a single GPU, which removes tensor-parallel communication overhead entirely, per Inworld. </p><p>Fewer cross-GPU hops means lower latency and lower cost at once.</p><h3>The arithmetic of serving</h3><p>Make the bottleneck quantitative, because the precise numbers are what justify the margin. During decode, producing one token requires reading every active parameter out of<strong> high-bandwidth memory exactly once</strong>, so the single-stream token rate has a hard ceiling set by bandwidth, not by compute:</p><blockquote><p><em>tokens/sec &#8776; HBM bandwidth &#247; ( bytes-per-parameter &#215; active parameters )</em></p></blockquote><p>The numbers are unforgiving. A dense 70-billion-parameter model in FP8 must stream <strong>70 GB per token</strong>; an H100 at 3.35 TB/s tops out near 48 tokens per second on a single stream, a B200 at 8 TB/s near 114. </p><p>A <strong>405-billion-parameter model</strong> falls into the single digits. This is why latency-sensitive, single-user decoding feels slow on the largest dense models no matter how many teraflops the chip advertises: those teraflops are not the binding constraint.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!bxlB!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F573de779-12c2-4405-a6bb-ede56132546f_1720x940.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!bxlB!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F573de779-12c2-4405-a6bb-ede56132546f_1720x940.png 424w, https://substackcdn.com/image/fetch/$s_!bxlB!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F573de779-12c2-4405-a6bb-ede56132546f_1720x940.png 848w, https://substackcdn.com/image/fetch/$s_!bxlB!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F573de779-12c2-4405-a6bb-ede56132546f_1720x940.png 1272w, https://substackcdn.com/image/fetch/$s_!bxlB!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F573de779-12c2-4405-a6bb-ede56132546f_1720x940.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!bxlB!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F573de779-12c2-4405-a6bb-ede56132546f_1720x940.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/573de779-12c2-4405-a6bb-ede56132546f_1720x940.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;c-decode&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="c-decode" title="c-decode" srcset="https://substackcdn.com/image/fetch/$s_!bxlB!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F573de779-12c2-4405-a6bb-ede56132546f_1720x940.png 424w, https://substackcdn.com/image/fetch/$s_!bxlB!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F573de779-12c2-4405-a6bb-ede56132546f_1720x940.png 848w, https://substackcdn.com/image/fetch/$s_!bxlB!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F573de779-12c2-4405-a6bb-ede56132546f_1720x940.png 1272w, https://substackcdn.com/image/fetch/$s_!bxlB!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F573de779-12c2-4405-a6bb-ede56132546f_1720x940.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>The latency floor.</strong> Single-stream decode is bounded by HBM bandwidth divided by the bytes streamed per token, so the largest dense models emit only single-digit to low-hundreds of tokens per second per accelerator before batching. Batching raises aggregate throughput; it does not raise this single-stream ceiling. H100 at 3.35 TB/s, B200 at 8 TB/s, FP8.</figcaption></figure></div><p>The escape is batching. Because every sequence in a batch reuses the same weight read, serving <strong>256 requests together</strong> multiplies aggregate throughput by roughly 256 with no additional weight traffic, until the KV cache or the compute ceiling intervenes. </p><p>Plotted on a roofline, decode lives far down the memory-bound slope at an arithmetic intensity near one, while prefill and training sit against the compute ceiling. </p><p>The ridge point, where a workload stops being memory-bound and becomes compute-bound, falls <strong>near 560 to 590 FLOP per byte</strong> on both H100 and B200, and decode runs one to two orders of magnitude below it.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!0ue5!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffc22a07f-35bc-4b43-8942-192345d9b0ed_1720x1000.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!0ue5!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffc22a07f-35bc-4b43-8942-192345d9b0ed_1720x1000.png 424w, https://substackcdn.com/image/fetch/$s_!0ue5!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffc22a07f-35bc-4b43-8942-192345d9b0ed_1720x1000.png 848w, https://substackcdn.com/image/fetch/$s_!0ue5!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffc22a07f-35bc-4b43-8942-192345d9b0ed_1720x1000.png 1272w, https://substackcdn.com/image/fetch/$s_!0ue5!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffc22a07f-35bc-4b43-8942-192345d9b0ed_1720x1000.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!0ue5!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffc22a07f-35bc-4b43-8942-192345d9b0ed_1720x1000.png" width="1456" height="847" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/fc22a07f-35bc-4b43-8942-192345d9b0ed_1720x1000.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:847,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;c-roofline&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="c-roofline" title="c-roofline" srcset="https://substackcdn.com/image/fetch/$s_!0ue5!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffc22a07f-35bc-4b43-8942-192345d9b0ed_1720x1000.png 424w, https://substackcdn.com/image/fetch/$s_!0ue5!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffc22a07f-35bc-4b43-8942-192345d9b0ed_1720x1000.png 848w, https://substackcdn.com/image/fetch/$s_!0ue5!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffc22a07f-35bc-4b43-8942-192345d9b0ed_1720x1000.png 1272w, https://substackcdn.com/image/fetch/$s_!0ue5!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffc22a07f-35bc-4b43-8942-192345d9b0ed_1720x1000.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Why decode leaves FLOPs on the table.</strong> Below the ridge point, throughput is capped by bandwidth, so adding compute does nothing; only more bandwidth or fewer bytes per token help. Prefill and training sit against the compute ceiling, where FLOPs matter. Dense FP8 peak: H100 ~1.98 PF, B200 ~4.5 PF.</figcaption></figure></div><p>This <strong>reframes the right efficiency metric</strong>, a point your readers will appreciate more than any valuation table. For prefill and training, the number that matters is model FLOPs utilization, MFU, typically 35 to 50 percent on a well-tuned cluster. </p><p>For decode, <strong>FLOPs utilization is nearly meaningless </strong>because the tensor cores are starved of data; the metric that matters is memory-bandwidth utilization, MBU, and the engineering goal is to keep HBM busy, not the math units. Almost every serving optimization that matters, from paged attention to continuous batching, is at bottom a scheme to raise MBU.</p><p><strong>Quantization attacks the bytes.</strong> Halving the bytes per parameter halves the bandwidth bill per token and so roughly doubles decode throughput. The frontier has marched down the precision ladder accordingly: FP16 and BF16 at two bytes, <strong>FP8 </strong>at one, <strong>FP4 </strong>and INT4 at half a byte, with the KV cache itself increasingly stored in FP8 to stretch context. </p><p>Blackwell&#8217;s tensor cores are built for FP4; Trainium3&#8217;s wider array is built for <strong>MXFP8 microscaling</strong>. The binding constraint is quality: too-aggressive quantization raises perplexity and degrades reasoning, so labs quantize weights and cache hard while protecting the few numerically sensitive layers.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!agXi!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb4b780fa-8d09-4960-8c2b-1aad0975dbbb_1720x940.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!agXi!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb4b780fa-8d09-4960-8c2b-1aad0975dbbb_1720x940.png 424w, https://substackcdn.com/image/fetch/$s_!agXi!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb4b780fa-8d09-4960-8c2b-1aad0975dbbb_1720x940.png 848w, https://substackcdn.com/image/fetch/$s_!agXi!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb4b780fa-8d09-4960-8c2b-1aad0975dbbb_1720x940.png 1272w, https://substackcdn.com/image/fetch/$s_!agXi!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb4b780fa-8d09-4960-8c2b-1aad0975dbbb_1720x940.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!agXi!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb4b780fa-8d09-4960-8c2b-1aad0975dbbb_1720x940.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b4b780fa-8d09-4960-8c2b-1aad0975dbbb_1720x940.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;c-bytesparam&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="c-bytesparam" title="c-bytesparam" srcset="https://substackcdn.com/image/fetch/$s_!agXi!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb4b780fa-8d09-4960-8c2b-1aad0975dbbb_1720x940.png 424w, https://substackcdn.com/image/fetch/$s_!agXi!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb4b780fa-8d09-4960-8c2b-1aad0975dbbb_1720x940.png 848w, https://substackcdn.com/image/fetch/$s_!agXi!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb4b780fa-8d09-4960-8c2b-1aad0975dbbb_1720x940.png 1272w, https://substackcdn.com/image/fetch/$s_!agXi!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb4b780fa-8d09-4960-8c2b-1aad0975dbbb_1720x940.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Precision is bandwidth.</strong> Each step down the ladder halves the bytes streamed per token, and because decode is bandwidth-bound, roughly doubles throughput. The frontier now runs weights in FP8 or FP4 and stores the KV cache in FP8, quantizing right up to the point where quality breaks.</figcaption></figure></div><p><strong>The KV cache is the tax that limits batching.</strong> Attention must keep the key and value vectors of every prior token, and that cache grows linearly with both context length and batch size, competing with the weights for the same HBM. </p><p>For a<strong> 70-billion-parameter model</strong> with grouped-query attention, the cache costs roughly 0.33 MB per token, so a single million-token context consumes more than 340 GB, beyond what a B200 holds. The cache, not the weights, is usually what caps how large a batch can run, and batch size is what sets cost per token, so KV-cache management is the hinge of serving economics. </p><p><strong>Grouped-query </strong>and <strong>multi-query attention</strong> shrink it by sharing key and value heads; paged attention stops it from fragmenting memory; FP8 storage halves it again.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!NvlE!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1bdf132a-f5af-4f41-aa4c-9b26992d5b3d_1720x940.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!NvlE!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1bdf132a-f5af-4f41-aa4c-9b26992d5b3d_1720x940.png 424w, https://substackcdn.com/image/fetch/$s_!NvlE!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1bdf132a-f5af-4f41-aa4c-9b26992d5b3d_1720x940.png 848w, https://substackcdn.com/image/fetch/$s_!NvlE!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1bdf132a-f5af-4f41-aa4c-9b26992d5b3d_1720x940.png 1272w, https://substackcdn.com/image/fetch/$s_!NvlE!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1bdf132a-f5af-4f41-aa4c-9b26992d5b3d_1720x940.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!NvlE!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1bdf132a-f5af-4f41-aa4c-9b26992d5b3d_1720x940.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/1bdf132a-f5af-4f41-aa4c-9b26992d5b3d_1720x940.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;c-kv&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="c-kv" title="c-kv" srcset="https://substackcdn.com/image/fetch/$s_!NvlE!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1bdf132a-f5af-4f41-aa4c-9b26992d5b3d_1720x940.png 424w, https://substackcdn.com/image/fetch/$s_!NvlE!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1bdf132a-f5af-4f41-aa4c-9b26992d5b3d_1720x940.png 848w, https://substackcdn.com/image/fetch/$s_!NvlE!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1bdf132a-f5af-4f41-aa4c-9b26992d5b3d_1720x940.png 1272w, https://substackcdn.com/image/fetch/$s_!NvlE!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1bdf132a-f5af-4f41-aa4c-9b26992d5b3d_1720x940.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>The cache outgrows the chip.</strong> For a 70B-class GQA model the KV cache costs ~0.33 MB per token and grows linearly with context, overtaking a B200&#8217;s entire 192 GB before a single million-token sequence. This, more than the weights, is what caps batch size, and batch size sets cost per token.</figcaption></figure></div><p><strong>Speculative decoding attacks the sequential dependency.</strong> A small draft model proposes several tokens and the large model verifies them in one forward pass, accepting the longest correct prefix. </p><p>With a per-token draft acceptance probability near 0.7 and four drafted tokens, the expected number confirmed per verification step is about ( 1 minus 0.7 to the fifth ) divided by 0.3, near 2.8, a two to three times speedup before draft overhead. </p><p>The large model still does the same total work per accepted token; what changes is that the work happens in parallel instead of one token at a time, which is exactly what a memory-bound loop needs.</p><p><strong>Mixture of experts attacks the parameter count.</strong> A dense model pays for all its parameters on every token; a mixture-of-experts model routes each token to a small subset, so per-token compute scales with active parameters, not total. DeepSeek&#8217;s architecture carries 671 billion parameters but activates 37 billion per token, an 18-fold cut in both the FLOPs and, decisively, the bytes streamed during decode. </p><p>The catch is twofold: all 671 billion parameters must still sit in HBM, a capacity tax that demands many chips, and routing tokens to experts requires an all-to-all step that leans on the scale-up fabric from the previous section. </p><p>MoE is the single largest reason a frontier-class answer is no longer a frontier-class expense, and it is why the dense decode ceiling above understates a well-built model: at 37 billion active parameters that same H100 serves roughly 90 tokens per second single-stream rather than five.</p><p><strong>Stack these together</strong>, quantization halving bytes, mixture-of-experts cutting active parameters by an order of magnitude, speculative decoding parallelizing the sequence, batching <strong>amortizing every weight read</strong>, and cheaper bandwidth each hardware generation on top, and the roughly tenfold annual fall in the cost of a token stops looking like magic and starts looking like arithmetic. </p><p>That is<strong> the engine underneath</strong> the margin bridge that follows.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!JS1s!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F40b9f016-945d-4026-9806-d81d05f0f204_1720x940.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!JS1s!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F40b9f016-945d-4026-9806-d81d05f0f204_1720x940.png 424w, https://substackcdn.com/image/fetch/$s_!JS1s!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F40b9f016-945d-4026-9806-d81d05f0f204_1720x940.png 848w, https://substackcdn.com/image/fetch/$s_!JS1s!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F40b9f016-945d-4026-9806-d81d05f0f204_1720x940.png 1272w, https://substackcdn.com/image/fetch/$s_!JS1s!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F40b9f016-945d-4026-9806-d81d05f0f204_1720x940.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!JS1s!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F40b9f016-945d-4026-9806-d81d05f0f204_1720x940.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/40b9f016-945d-4026-9806-d81d05f0f204_1720x940.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;c-costdecline&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="c-costdecline" title="c-costdecline" srcset="https://substackcdn.com/image/fetch/$s_!JS1s!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F40b9f016-945d-4026-9806-d81d05f0f204_1720x940.png 424w, https://substackcdn.com/image/fetch/$s_!JS1s!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F40b9f016-945d-4026-9806-d81d05f0f204_1720x940.png 848w, https://substackcdn.com/image/fetch/$s_!JS1s!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F40b9f016-945d-4026-9806-d81d05f0f204_1720x940.png 1272w, https://substackcdn.com/image/fetch/$s_!JS1s!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F40b9f016-945d-4026-9806-d81d05f0f204_1720x940.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Roughly 10x a year, faster than PC compute or dot-com bandwidth ever fell (Introl).</strong> GPT-4-class output dropped from about $20 per million tokens in late 2022 to about $0.40 by late 2025. Endpoints documented (Introl); intermediate points trace the stated trend. Economy-tier quality fell about 600x since 2020, from the $60 GPT-3 API to roughly $0.10 today (arXiv).</figcaption></figure></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!rP7c!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F155bf55b-95b0-48b9-b16c-80c195aba771_1720x940.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!rP7c!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F155bf55b-95b0-48b9-b16c-80c195aba771_1720x940.png 424w, https://substackcdn.com/image/fetch/$s_!rP7c!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F155bf55b-95b0-48b9-b16c-80c195aba771_1720x940.png 848w, https://substackcdn.com/image/fetch/$s_!rP7c!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F155bf55b-95b0-48b9-b16c-80c195aba771_1720x940.png 1272w, https://substackcdn.com/image/fetch/$s_!rP7c!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F155bf55b-95b0-48b9-b16c-80c195aba771_1720x940.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!rP7c!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F155bf55b-95b0-48b9-b16c-80c195aba771_1720x940.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/155bf55b-95b0-48b9-b16c-80c195aba771_1720x940.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;c-costchip&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="c-costchip" title="c-costchip" srcset="https://substackcdn.com/image/fetch/$s_!rP7c!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F155bf55b-95b0-48b9-b16c-80c195aba771_1720x940.png 424w, https://substackcdn.com/image/fetch/$s_!rP7c!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F155bf55b-95b0-48b9-b16c-80c195aba771_1720x940.png 848w, https://substackcdn.com/image/fetch/$s_!rP7c!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F155bf55b-95b0-48b9-b16c-80c195aba771_1720x940.png 1272w, https://substackcdn.com/image/fetch/$s_!rP7c!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F155bf55b-95b0-48b9-b16c-80c195aba771_1720x940.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>A generational chip cut inference cost roughly 7x.</strong> Inworld measures about $0.02 per million tokens on B200 versus about $0.14 on H100, because throughput gains outpace the rental premium. Other teardowns put H100 frontier serving at $0.19 to $0.29 fully utilized, doubling at half utilization (GMI). Figures are config-dependent; treat as directional.</figcaption></figure></div><h3>What it costs Anthropic to make a token</h3><p>We can now build the cost side from the metal up. The exercise is illustrative, the assumptions are stated plainly, and the point is the order of magnitude, not a false-precision number.</p><p></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!rocn!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1540f0fa-42db-4708-bee7-7643480c36b8_2470x1273.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!rocn!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1540f0fa-42db-4708-bee7-7643480c36b8_2470x1273.png 424w, https://substackcdn.com/image/fetch/$s_!rocn!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1540f0fa-42db-4708-bee7-7643480c36b8_2470x1273.png 848w, https://substackcdn.com/image/fetch/$s_!rocn!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1540f0fa-42db-4708-bee7-7643480c36b8_2470x1273.png 1272w, https://substackcdn.com/image/fetch/$s_!rocn!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1540f0fa-42db-4708-bee7-7643480c36b8_2470x1273.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!rocn!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1540f0fa-42db-4708-bee7-7643480c36b8_2470x1273.png" width="1456" height="750" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/1540f0fa-42db-4708-bee7-7643480c36b8_2470x1273.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:750,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:258471,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/201004624?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1540f0fa-42db-4708-bee7-7643480c36b8_2470x1273.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!rocn!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1540f0fa-42db-4708-bee7-7643480c36b8_2470x1273.png 424w, https://substackcdn.com/image/fetch/$s_!rocn!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1540f0fa-42db-4708-bee7-7643480c36b8_2470x1273.png 848w, https://substackcdn.com/image/fetch/$s_!rocn!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1540f0fa-42db-4708-bee7-7643480c36b8_2470x1273.png 1272w, https://substackcdn.com/image/fetch/$s_!rocn!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1540f0fa-42db-4708-bee7-7643480c36b8_2470x1273.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Now the other side of the ledger. </p><p>We do not know Anthropic&#8217;s blended <strong>realized price per token</strong>, because the filing is confidential, and that is the honest limit of this analysis. </p><p>What we do know: frontier output has historically listed in the <strong>five-to-fifteen-dollar range per million tokens</strong>, and Anthropic&#8217;s most capable model, the withheld Mythos preview, was priced at twenty-five dollars per million input tokens and one hundred twenty-five dollars per million output, per <strong>Sacra </strong>(<em>a single-source figure, indicative rather than confirmed</em>). </p><p>Even after a generous markup for the true cost of a genuine frontier model over a benchmark mid-size one, the gross spread between a cost-to-serve measured in cents and a realized price measured in dollars is wide. </p><p>That <strong>spread is the gross margin</strong>. And because the cost side falls about tenfold a year while realized prices fall more slowly, the spread widens with time. This is the physical mechanism behind the projected march from forty-five to seventy-seven percent.</p><p>Nvidia, naturally, has quantified the same loop from the supplier&#8217;s side. In its InferenceMax v1 results, the company claims a <strong>single GB200 NVL72 </strong>turns a five-million-dollar investment into roughly seventy-five million dollars of <strong>DeepSeek-R1 token revenue</strong>, a fifteen-fold return, what it calls AI-factory economics. </p><p>Treat the figure as a vendor benchmark on an idealized model, but the direction is the entire bull thesis in one number: at current token prices, a frontier accelerator generates a multiple of its cost in sellable output.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!lK1C!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F844b3c34-0da0-413b-8688-82d9037758f8_1720x940.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!lK1C!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F844b3c34-0da0-413b-8688-82d9037758f8_1720x940.png 424w, https://substackcdn.com/image/fetch/$s_!lK1C!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F844b3c34-0da0-413b-8688-82d9037758f8_1720x940.png 848w, https://substackcdn.com/image/fetch/$s_!lK1C!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F844b3c34-0da0-413b-8688-82d9037758f8_1720x940.png 1272w, https://substackcdn.com/image/fetch/$s_!lK1C!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F844b3c34-0da0-413b-8688-82d9037758f8_1720x940.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!lK1C!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F844b3c34-0da0-413b-8688-82d9037758f8_1720x940.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/844b3c34-0da0-413b-8688-82d9037758f8_1720x940.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;c-marginbridge&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="c-marginbridge" title="c-marginbridge" srcset="https://substackcdn.com/image/fetch/$s_!lK1C!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F844b3c34-0da0-413b-8688-82d9037758f8_1720x940.png 424w, https://substackcdn.com/image/fetch/$s_!lK1C!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F844b3c34-0da0-413b-8688-82d9037758f8_1720x940.png 848w, https://substackcdn.com/image/fetch/$s_!lK1C!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F844b3c34-0da0-413b-8688-82d9037758f8_1720x940.png 1272w, https://substackcdn.com/image/fetch/$s_!lK1C!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F844b3c34-0da0-413b-8688-82d9037758f8_1720x940.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>How -94% becomes +77%.</strong> Fixed-cost leverage (revenue scaling against a roughly fixed capacity base, per SaaStr) does the heavy lifting; token deflation (quantization, MoE, cheaper silicon) and a mix shift toward high-value enterprise and Claude Code finish it. Step sizes are an illustrative attribution to named mechanisms, not audited figures; the 2024, 2025, and 2028E levels are the reported and projected anchors.</figcaption></figure></div><p><strong>The technical bottom line</strong></p><p>The cost <strong>half of Anthropic&#8217;s margin equation</strong> is governed by mechanisms we can see and that compound predictably: <em>quantization, mixture-of-experts routing, algorithmic serving gains, and cheaper bandwidth per token</em>, together delivering roughly an order of magnitude of cost reduction per year. That is why a 77% gross margin is physically plausible rather than fantastical.</p><p>The risk lives on the price half, which we cannot see. If open-weight competition (<em>DeepSeek, Llama</em>) and rival labs compress realized prices as fast as cost falls, the spread does not widen and the margin thesis stalls. </p><p><mark>The bull is betting cost falls faster than price. </mark></p><p><mark>The bear is betting price falls to meet cost.</mark> The <strong>confidential S-1</strong> hides exactly the number, realized revenue per token, that would settle it.</p><div><hr></div><h2>The infinite loop</h2><p>Look again at the last two columns of the capital-and-compute table. <strong>Amazon </strong>is investing up to <strong>thirty-three billion into Anthropic</strong>; Anthropic is committing to spend over one hundred billion with Amazon. </p><p><strong>Google </strong>is investing up to <em>forty billion</em>; Anthropic is spending tens of billions with Google. The capital flows out as equity and comes back as revenue.</p><p><strong>CNBC </strong>said it plainly in its coverage of the Google deal: much of the investment will return in the form of revenue. This is the circular-financing question, and it is the single most important structural issue hanging over all three IPOs. </p><p>It is not unique to Anthropic. It is basically the operating system of the entire cycle.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!_SP9!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F95a85398-8873-46aa-9577-856e877e8f18_1800x1080.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!_SP9!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F95a85398-8873-46aa-9577-856e877e8f18_1800x1080.png 424w, https://substackcdn.com/image/fetch/$s_!_SP9!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F95a85398-8873-46aa-9577-856e877e8f18_1800x1080.png 848w, https://substackcdn.com/image/fetch/$s_!_SP9!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F95a85398-8873-46aa-9577-856e877e8f18_1800x1080.png 1272w, https://substackcdn.com/image/fetch/$s_!_SP9!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F95a85398-8873-46aa-9577-856e877e8f18_1800x1080.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!_SP9!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F95a85398-8873-46aa-9577-856e877e8f18_1800x1080.png" width="1456" height="874" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/95a85398-8873-46aa-9577-856e877e8f18_1800x1080.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:874,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;c-flow&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="c-flow" title="c-flow" srcset="https://substackcdn.com/image/fetch/$s_!_SP9!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F95a85398-8873-46aa-9577-856e877e8f18_1800x1080.png 424w, https://substackcdn.com/image/fetch/$s_!_SP9!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F95a85398-8873-46aa-9577-856e877e8f18_1800x1080.png 848w, https://substackcdn.com/image/fetch/$s_!_SP9!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F95a85398-8873-46aa-9577-856e877e8f18_1800x1080.png 1272w, https://substackcdn.com/image/fetch/$s_!_SP9!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F95a85398-8873-46aa-9577-856e877e8f18_1800x1080.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>The closed loop.</strong> The same firms that supply the silicon and the cloud also own the equity. Capital invested as equity returns as committed revenue, which is then reported as growth. Figures are deal maxima, not all drawn. Sources: Global Data Center Hub, TechCrunch, CNBC, Reuters.</figcaption></figure></div><p>The canonical example is on the <strong>OpenAI side</strong>. In September 2025 Nvidia announced it would invest up to one hundred billion dollars in OpenAI to fund a data center buildout equipped with, naturally, Nvidia chips. </p><p><strong>Bernstein&#8217;s Stacy Rasgon </strong>wrote, per Business Standard, that the move would <em>clearly fuel circular concerns.</em> By March 2026, per BlockEden&#8217;s account, Jensen Huang was telling investors that thirty billion might be the last such investment and that the full hundred billion was not in the cards.</p><p>Nvidia also committed up to ten billion to Anthropic, which <strong>CFO Colette Kress</strong> noted could further expand the company&#8217;s bookings, a sentence that contains the whole critique in miniature: the investment expands the bookings of the company making the investment.</p><p>The <strong>scale of the web</strong> is staggering. BlockEden tallied OpenAI&#8217;s infrastructure commitments at roughly $1.15 trillion across seven vendors between 2025 and 2035.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!glRV!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a7a59a9-addd-40df-873d-4b5145d51825_1720x980.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!glRV!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a7a59a9-addd-40df-873d-4b5145d51825_1720x980.png 424w, https://substackcdn.com/image/fetch/$s_!glRV!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a7a59a9-addd-40df-873d-4b5145d51825_1720x980.png 848w, https://substackcdn.com/image/fetch/$s_!glRV!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a7a59a9-addd-40df-873d-4b5145d51825_1720x980.png 1272w, https://substackcdn.com/image/fetch/$s_!glRV!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a7a59a9-addd-40df-873d-4b5145d51825_1720x980.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!glRV!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a7a59a9-addd-40df-873d-4b5145d51825_1720x980.png" width="1456" height="830" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/1a7a59a9-addd-40df-873d-4b5145d51825_1720x980.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:830,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;c-vendorweb&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="c-vendorweb" title="c-vendorweb" srcset="https://substackcdn.com/image/fetch/$s_!glRV!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a7a59a9-addd-40df-873d-4b5145d51825_1720x980.png 424w, https://substackcdn.com/image/fetch/$s_!glRV!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a7a59a9-addd-40df-873d-4b5145d51825_1720x980.png 848w, https://substackcdn.com/image/fetch/$s_!glRV!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a7a59a9-addd-40df-873d-4b5145d51825_1720x980.png 1272w, https://substackcdn.com/image/fetch/$s_!glRV!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a7a59a9-addd-40df-873d-4b5145d51825_1720x980.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>One company, $1.15 trillion in promises.</strong> Several of these vendors are also OpenAI investors; AMD reportedly handed OpenAI 10% equity warrants as a customer. The capital and the contracts run in a closed circle. Source: BlockEden compilation of public disclosures.</figcaption></figure></div><p>The defenders are not stupid, and their argument deserves a fair hearing. </p><p><strong>Dario Amodei</strong>, at the <em>New York Times DealBook Summit</em> in December, argued there is nothing inappropriate in principle about a party with capital and a chip interest funding a party with revenue confidence but no cash on hand. </p><p>That is a coherent description of ordinary project finance. But the historical rhyme is hard to ignore, and Bloomberg drew it explicitly: during the<strong> late-1990s internet boom</strong>, equipment makers fueled the fiber buildout with vendor financing, and when demand failed to arrive on schedule, the roundtripping that had inflated the appearance of demand amplified the collapse. </p><p>The mechanism that makes the boom look bigger is the same one that makes the bust deeper. Sequoia&#8217;s David Cahn has quantified the implied shortfall: by his framework, the <strong>AI complex</strong> needs roughly six hundred billion dollars in annual revenue to justify the capex being deployed, and the gap is widening, not closing.</p><blockquote><p><em>&#8220;In this new world of AI, compute is revenues.&#8221;</em></p></blockquote><p><strong>Jensen Huang</strong> &#183; Nvidia CEO, on the Q4 FY2026 earnings call, reframing the entire spending debate (Fortune, Benzinga)</p><p>That single line is the keystone of the bull architecture, and section VII gave it a number: at current token prices a frontier accelerator throws off a multiple of its cost in sellable output. </p><p>Huang&#8217;s claim is that<strong> capital expenditure</strong> converts into compute, compute into tokens, and tokens directly into revenue, so the spending is self-justifying. </p><p>Nvidia&#8217;s own results give the argument force: <strong>record quarterly revenue</strong> of $68.1 billion, up roughly 73 percent year over year, with $78 billion guided for the next quarter, and Kress telling investors total AI infrastructure investment could reach three to four trillion dollars annually by 2029 or 2030. </p><p>For Anthropic specifically, the loop lands on one phrase, revenue quality. When eighty percent of revenue is enterprise and a meaningful share of capital comes from the same<strong> hyperscalers </strong>whose clouds Anthropic is committing to, an investor is entitled to ask how much of the forty-seven-billion run-rate is organic demand and how much is the visible end of a closed capital loop. </p><p><em>Nobody outside the company knows.</em> The filing is confidential. That is the second reason confidentiality matters more here than usual.</p><div><hr></div><h2>How do you price a wall?</h2><p>Valuation is where the rigor either holds or collapses, so let us be careful, bring the <strong>public-market </strong>context the private headlines leave out, and do the one thing most coverage skips: put the multiple next to its peers.</p><p>Anthropic closed its <strong>Series G</strong> on February 12, 2026: thirty billion dollars raised at a $380 billion post-money valuation, which Perera calculated as roughly twenty-seven times annualized revenue. </p><p>By the April tender offer, the reference valuation was $350 billion, at which <em>The Motley Fool</em> noted the multiple had compressed to under twelve times the then-thirty-billion run-rate, simply because revenue had nearly tripled while the valuation held flat. </p><p>Then, per <strong>Let&#8217;s Data Science</strong> and <strong>CNN</strong>, Anthropic raised sixty-five billion dollars in May at a $965 billion valuation, surpassing OpenAI&#8217;s $852 billion mark for the first time. </p><p>The IPO target is above one trillion. Reuters notes the company was valued at just $183 billion as recently as last November.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!G0wH!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe29d6d9c-6df7-4492-ae42-955d84f42a23_1720x940.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!G0wH!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe29d6d9c-6df7-4492-ae42-955d84f42a23_1720x940.png 424w, https://substackcdn.com/image/fetch/$s_!G0wH!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe29d6d9c-6df7-4492-ae42-955d84f42a23_1720x940.png 848w, https://substackcdn.com/image/fetch/$s_!G0wH!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe29d6d9c-6df7-4492-ae42-955d84f42a23_1720x940.png 1272w, https://substackcdn.com/image/fetch/$s_!G0wH!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe29d6d9c-6df7-4492-ae42-955d84f42a23_1720x940.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!G0wH!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe29d6d9c-6df7-4492-ae42-955d84f42a23_1720x940.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e29d6d9c-6df7-4492-ae42-955d84f42a23_1720x940.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;c-valuation&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="c-valuation" title="c-valuation" srcset="https://substackcdn.com/image/fetch/$s_!G0wH!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe29d6d9c-6df7-4492-ae42-955d84f42a23_1720x940.png 424w, https://substackcdn.com/image/fetch/$s_!G0wH!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe29d6d9c-6df7-4492-ae42-955d84f42a23_1720x940.png 848w, https://substackcdn.com/image/fetch/$s_!G0wH!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe29d6d9c-6df7-4492-ae42-955d84f42a23_1720x940.png 1272w, https://substackcdn.com/image/fetch/$s_!G0wH!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe29d6d9c-6df7-4492-ae42-955d84f42a23_1720x940.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>5x in six months.</strong> $183B to a trillion-plus target between November and the IPO. The multiple is not the story; the denominator is moving too fast for any multiple to stay meaningful. Sources: Reuters, Perera, Global Data Center Hub, CNN, Let&#8217;s Data Science.</figcaption></figure></div><h3>The comp that reframes the question</h3><p>A trillion dollars sounds insane until you place it beside what the public market already pays for AI growth. </p><p>At its IPO target, Anthropic trades near twenty-one times its <strong>forty-seven-billion run-rate</strong>, and near fourteen times its projected seventy-billion 2028 revenue. Those are not the highest multiples in the AI complex. They are far from it.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!dwiN!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8887fcb0-ea3e-4552-89b5-f6231bdfde36_1720x940.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!dwiN!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8887fcb0-ea3e-4552-89b5-f6231bdfde36_1720x940.png 424w, https://substackcdn.com/image/fetch/$s_!dwiN!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8887fcb0-ea3e-4552-89b5-f6231bdfde36_1720x940.png 848w, https://substackcdn.com/image/fetch/$s_!dwiN!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8887fcb0-ea3e-4552-89b5-f6231bdfde36_1720x940.png 1272w, https://substackcdn.com/image/fetch/$s_!dwiN!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8887fcb0-ea3e-4552-89b5-f6231bdfde36_1720x940.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!dwiN!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8887fcb0-ea3e-4552-89b5-f6231bdfde36_1720x940.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/8887fcb0-ea3e-4552-89b5-f6231bdfde36_1720x940.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;c-comps&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="c-comps" title="c-comps" srcset="https://substackcdn.com/image/fetch/$s_!dwiN!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8887fcb0-ea3e-4552-89b5-f6231bdfde36_1720x940.png 424w, https://substackcdn.com/image/fetch/$s_!dwiN!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8887fcb0-ea3e-4552-89b5-f6231bdfde36_1720x940.png 848w, https://substackcdn.com/image/fetch/$s_!dwiN!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8887fcb0-ea3e-4552-89b5-f6231bdfde36_1720x940.png 1272w, https://substackcdn.com/image/fetch/$s_!dwiN!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8887fcb0-ea3e-4552-89b5-f6231bdfde36_1720x940.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>Anthropic&#8217;s multiple is mid-pack, despite the fastest growth.</strong> Palantir trades near 55x revenue at roughly 70 to 85 percent growth; Databricks near 25x at 65 percent. Anthropic on run-rate sits near 21x while growing on the order of 1,000 percent year over year, and near 14x on 2028E. Multiples for Microsoft, Alphabet, and Amazon are approximate and vary daily. Databricks and Anthropic are private marks. Sources: multiples.vc, heygotrade, SaaStr (PLTR, NVDA, Databricks); company valuations for the Anthropic bars.</figcaption></figure></div><p>That is the reframe a careful analyst owes the reader. On a pure revenue multiple, Anthropic at a trillion dollars is <strong><mark>cheaper than Palantir</mark></strong><mark> and cheaper than Databricks, while growing many times faster than either.</mark> The naive objection, the multiple is absurd, does not survive contact with the comp set. </p><p>The serious objections are about durability and quality, and there are three honest frames.</p><p><strong>The bull frame is forward margin.</strong> If you believe the seventy-seven-percent gross margin of section VII arrives, and the company&#8217;s projection of roughly seventy billion in revenue and seventeen billion in cash flow by 2028 (Sacra), then a trillion dollars is about fourteen times 2028 revenue on the fastest-growing, highest-margin software asset ever built. Rich, but not obviously mispriced for the category leader.</p><p><strong>The bear frame is present reality.</strong> Right now the margin is closer to forty-five percent, the company does not expect to stop burning cash until 2027 (Sacra), and the cloud bill runs to roughly eighty billion dollars through 2029. At present economics, a trillion dollars prices a future that has not arrived as though it already has, and the comp multiples assume the growth rate persists for years, which no company in history has sustained.</p><p><strong>The skeptic frame is the one I find most useful.</strong> We are pricing the largest IPO in a generation on a confidential filing. One quiet signal cuts against the skepticism: The Motley Fool reported that when Anthropic invited long-tenured employees to sell at the $350 billion valuation, they chose to hold far more than expected. The people with the most information declined liquidity at $350 billion, and the May round then priced at nearly three times that. Insider behavior is not proof, but it is data, and here it points the same way the revenue does. Up.</p><h3>How to actually get exposure</h3><p>For investors who<strong> cannot buy private shares</strong>, the cleanest listed proxies are mechanical. Amazon carries up to a thirty-three-billion-dollar stake plus the AWS revenue Anthropic is committing to. </p><p>Alphabet holds roughly fourteen percent plus Google Cloud&#8217;s TPU revenue; <strong>Google Cloud </strong>grew about sixty-three percent year over year to a roughly twenty-billion-dollar quarterly run-rate, the fastest of the big three. </p><p>Broadcom is the picks-and-shovels play through TPU co-design and connectivity. Nvidia sits at the keystone with a roughly ten-billion-dollar stake and the <strong>GPU demand underneath</strong> all of it, trading near a 4.8-trillion-dollar market capitalization on the strength of Huang&#8217;s compute-is-revenues thesis. </p><p>A trillion-dollar Anthropic IPO does not just price Anthropic. As one trading desk framed it, the entire listed AI complex is likely to re-rate on Anthropic comps, not merely the company itself.</p><p>One market-structure point will matter on debut day. A listing this large arriving this fast triggers fast-entry index inclusion rules, which means passive funds become forced buyers shortly after pricing, a mechanical tailwind Yahoo Finance flagged as one reason all<strong> three trillion-dollar names</strong> may benefit from going public in quick succession. Investors will have just watched SpaceX test those rules in real time.</p><h3>What the price implies: a transparent valuation</h3><p>The comps above are <strong>a relative reframe</strong>, not a valuation. Here is the valuation, built the way a disciplined analyst builds one when the financials are sealed. </p><p>You cannot run a bottoms-up discounted-cash-flow model on Anthropic, because the inputs such a model needs, the actual margins, the free cash flow, the <strong>capex schedule</strong>, the share count, and the realized price per token, all sit inside the confidential S-1. </p><p>Anyone publishing a <strong>single intrinsic-value number</strong> from a DCF right now is inventing those inputs. What can be done rigorously, from fact-checked data alone, is to reframe the question around two hard primary anchors, the roughly $47 billion run-rate and the roughly $1 trillion price, and ask answerable things.</p><h3>What the price requires</h3><p>The most honest move is <strong>a reverse discounted-cash-flow</strong>: rather than forecast cash flows we do not have, solve for the growth the known price embeds, then judge whether it is believable. </p><p>Treat $1 trillion as today&#8217;s enterprise value, assume the market values the company at a steady-state revenue multiple m reached in year eight and discounts at a cost of capital w, and the revenue the price implies follows directly.</p><p>Revenue(yr 8) = $1,000B &#215; (1 + w)^8 &#247; m | implied CAGR = ( Revenue(yr 8) &#247; $47B )^(1/8) - 1</p><p>With cost of capital from 9 to 13 percent and a terminal revenue multiple from 4 to 8 times, both defensible and neither touching sealed data, the implied eight-year revenue growth rate is the following.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Q7PU!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9dc27048-e691-4f95-a160-22db72676638_2470x1368.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Q7PU!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9dc27048-e691-4f95-a160-22db72676638_2470x1368.png 424w, https://substackcdn.com/image/fetch/$s_!Q7PU!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9dc27048-e691-4f95-a160-22db72676638_2470x1368.png 848w, https://substackcdn.com/image/fetch/$s_!Q7PU!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9dc27048-e691-4f95-a160-22db72676638_2470x1368.png 1272w, https://substackcdn.com/image/fetch/$s_!Q7PU!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9dc27048-e691-4f95-a160-22db72676638_2470x1368.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Q7PU!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9dc27048-e691-4f95-a160-22db72676638_2470x1368.png" width="1456" height="806" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/9dc27048-e691-4f95-a160-22db72676638_2470x1368.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:806,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:316461,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/201004624?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9dc27048-e691-4f95-a160-22db72676638_2470x1368.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Q7PU!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9dc27048-e691-4f95-a160-22db72676638_2470x1368.png 424w, https://substackcdn.com/image/fetch/$s_!Q7PU!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9dc27048-e691-4f95-a160-22db72676638_2470x1368.png 848w, https://substackcdn.com/image/fetch/$s_!Q7PU!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9dc27048-e691-4f95-a160-22db72676638_2470x1368.png 1272w, https://substackcdn.com/image/fetch/$s_!Q7PU!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9dc27048-e691-4f95-a160-22db72676638_2470x1368.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The price embeds <em><mark>roughly 25 to 40 percent annual revenue growth</mark></em><mark> sustained for eight years, centered near 30 percent.</mark> That converts the entire debate into one question a reader can answer: </p><blockquote><p><em>do you believe Anthropic compounds revenue at about 30 percent a year for nearly a decade? </em></p></blockquote><p>Current growth is far above that, which gives the bull case runway; the bear case is that <strong>no company has held 30 percent</strong> for eight straight years, and open-weight price compression is the likeliest thing to break it. </p><p>The reverse DCF does not say who is right. It says, from fact-checked data, exactly what you are being asked to underwrite.</p><h3>The range of outcomes</h3><p>A scenario value completes the picture, anchoring the base on the company&#8217;s own 2028 projection, which is a projection and flagged as such, with a bear and a bull around it, each valued at an exit multiple and discounted to today.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!5WP4!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd2fc4517-77bf-4998-976b-8f2ef3f2174a_1720x940.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!5WP4!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd2fc4517-77bf-4998-976b-8f2ef3f2174a_1720x940.png 424w, https://substackcdn.com/image/fetch/$s_!5WP4!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd2fc4517-77bf-4998-976b-8f2ef3f2174a_1720x940.png 848w, https://substackcdn.com/image/fetch/$s_!5WP4!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd2fc4517-77bf-4998-976b-8f2ef3f2174a_1720x940.png 1272w, https://substackcdn.com/image/fetch/$s_!5WP4!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd2fc4517-77bf-4998-976b-8f2ef3f2174a_1720x940.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!5WP4!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd2fc4517-77bf-4998-976b-8f2ef3f2174a_1720x940.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d2fc4517-77bf-4998-976b-8f2ef3f2174a_1720x940.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;c-valscenarios&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="c-valscenarios" title="c-valscenarios" srcset="https://substackcdn.com/image/fetch/$s_!5WP4!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd2fc4517-77bf-4998-976b-8f2ef3f2174a_1720x940.png 424w, https://substackcdn.com/image/fetch/$s_!5WP4!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd2fc4517-77bf-4998-976b-8f2ef3f2174a_1720x940.png 848w, https://substackcdn.com/image/fetch/$s_!5WP4!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd2fc4517-77bf-4998-976b-8f2ef3f2174a_1720x940.png 1272w, https://substackcdn.com/image/fetch/$s_!5WP4!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd2fc4517-77bf-4998-976b-8f2ef3f2174a_1720x940.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>The price is paid for the right tail.</strong> Probability-weighted present value is about $800 billion against a price near $1 trillion, so the market leans on the bull case to clear the bar. The five-fold spread between bear and bull, roughly $290 billion to $1.6 trillion, is the real story: this is a wide distribution, not a tight intrinsic value. Discounted at 11 percent; the probabilities are illustrative and the one input a reader should make their own.</figcaption></figure></div><p>One tension in the fact-checked data deserves a flag rather than a paper-over. A roughly $47 billion run-rate today against the company&#8217;s roughly $70 billion 2028 projection implies revenue growth decelerating to 15 to 20 percent a year by 2028, a sharp slowdown from the current pace. </p><p>Either the <strong>$70 billion figure </strong>predates the $47 billion run-rate and is now conservative, or management expects growth to crash, and the comps reframe that calls the price cheap rests entirely on which it is.</p><p><strong>How to read it</strong></p><p>The honest close is not a verdict but the underwriting question the price poses: do you believe roughly 30 percent revenue growth for eight years, and do you believe the<strong> cost half of the margin</strong> keeps falling faster than the price half. </p><p>A reader who answers yes to both can justify the trillion-dollar tag; one who doubts either cannot. The point of showing every assumption is that you can change them and watch the answer move.</p><p><em>This valuation is analysis, not investment advice and not a recommendation. The run-rate and price anchors are primary; the 2028 revenue and cash-flow figures are company projections; the multiples are approximate and volatile; the cost of capital, exit multiples, and scenario probabilities are explicit modeling choices made for transparency, not derived from non-public data.</em></p><div><hr></div><h2>A mirror, a rocket, and a clock</h2><p>You cannot value Anthropic in isolation, because the IPO is partly a race, and races have<strong> positional dynamics.</strong></p><p>OpenAI is the mirror image. Per Nerd Level Tech and Let&#8217;s Data Science, it converted to a public benefit corporation in October 2025 as<strong> OpenAI Group PBC,</strong> with the nonprofit Foundation retaining roughly twenty-six percent and board control. </p><p>Microsoft holds roughly twenty-seven percent on a diluted basis, an investment valued between about $135 billion and $228 billion depending on the mark, and ended its exclusivity arrangement in April. Sam Altman holds no equity. </p><p>OpenAI closed the largest private round in history on March 31, $122 billion at an<strong> $852 billion post-money valuation</strong>, with SoftBank, Amazon, Nvidia, and Microsoft all participating. </p><p>Revenue runs about two billion dollars a month, near a twenty-five-billion run-rate as of March, with<strong> fifty million consumer subscribers</strong> and nine million business users, per roborhythms&#8217; compilation.</p><p>The contrast that will dominate the dueling roadshows is profitability. Multiple outlets, citing the loss figures, report <em>OpenAI losing about $1.22 for every dollar of revenue</em> in Q1 2026, an operating margin near negative one hundred twenty-two percent, with a projected fourteen-billion-dollar loss for the year. </p><p>Anthropic, by the <strong>disputed WSJ figures</strong>, claims a small operating profit in the same window. So the two enter the public markets at nearly identical valuations and opposite financial stories.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!QV0t!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee5fc40b-3f71-4a60-b09a-7253fd7f72dd_1720x940.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!QV0t!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee5fc40b-3f71-4a60-b09a-7253fd7f72dd_1720x940.png 424w, https://substackcdn.com/image/fetch/$s_!QV0t!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee5fc40b-3f71-4a60-b09a-7253fd7f72dd_1720x940.png 848w, https://substackcdn.com/image/fetch/$s_!QV0t!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee5fc40b-3f71-4a60-b09a-7253fd7f72dd_1720x940.png 1272w, https://substackcdn.com/image/fetch/$s_!QV0t!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee5fc40b-3f71-4a60-b09a-7253fd7f72dd_1720x940.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!QV0t!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee5fc40b-3f71-4a60-b09a-7253fd7f72dd_1720x940.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ee5fc40b-3f71-4a60-b09a-7253fd7f72dd_1720x940.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;c-vsopenai&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="c-vsopenai" title="c-vsopenai" srcset="https://substackcdn.com/image/fetch/$s_!QV0t!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee5fc40b-3f71-4a60-b09a-7253fd7f72dd_1720x940.png 424w, https://substackcdn.com/image/fetch/$s_!QV0t!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee5fc40b-3f71-4a60-b09a-7253fd7f72dd_1720x940.png 848w, https://substackcdn.com/image/fetch/$s_!QV0t!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee5fc40b-3f71-4a60-b09a-7253fd7f72dd_1720x940.png 1272w, https://substackcdn.com/image/fetch/$s_!QV0t!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee5fc40b-3f71-4a60-b09a-7253fd7f72dd_1720x940.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>The overtaking.</strong> Anthropic crossed above OpenAI on run-rate revenue in spring 2026 even as OpenAI retained far greater consumer reach. By May, Sacra put Anthropic near $45B versus OpenAI near $33B. Reported figures, not audited.</figcaption></figure></div><p>There is a real first-mover argument, though the precedent cuts only so far. When <strong>Lyft </strong>and Uber went public in 2019, Lyft, the first mover, popped on its debut while <strong>Uber </strong>fell on its first day; both stocks then traded poorly in the months after, so debut-day positioning is no guarantee of anything.</p><p> Both labs will seek tens of billions in fresh capital in close succession, so reaching market first plausibly matters. Anthropic&#8217;s confidential filing means it could price as early as mid-August on the <strong>SpaceX timeline</strong>, per Yahoo Finance, though Futurum reports a target as late as October; either way, likely ahead of <strong>OpenAI&#8217;s Q4 window</strong>. </p><p>Against that, the sober counterpoint: Wall Street already knows both stories intimately, and <strong>OpenAI&#8217;s S-1</strong> will likely be public by the time Anthropic prices, letting investors judge them side by side regardless of who rings the bell first.</p><p>The rocket is the merged SpaceX entity, and it is the wild card. Its S-1 disclosed that xAI spent $12.7 billion on AI infrastructure in 2025 and another $7.7 billion in Q1 2026, per <strong>Datacenter Dynamics</strong>, and it described its one gigawatt of capacity as <em>nameplate compute draw,</em> explicitly noting the figure reflects installed capacity and does not represent actual utilization. </p><p>In plain terms, the GPUs are installed but may not all be powered. That single disclosure is a gift to anyone trying to separate capacity headlines from real,<strong> energized compute</strong>, and every analyst should now apply a utilization haircut to gigawatt claims across the industry, including the contracted-capacity bars in section V.</p><p>Step back and the pipeline is unlike anything the IPO market has seen. Three trillion-dollar names, plus <strong>Databricks </strong><em>(the only clearly profitable candidate, at a $5.4 billion run-rate growing 65 percent with positive free cash flow and a $134 billion private mark</em>) and <strong>Cerebras </strong>(which priced around a $48.8 billion valuation on a heavily oversubscribed book). </p><p>One IPO tracker estimated combined pipeline demand at up to four times the entire 2025 US IPO market.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!uxaG!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea3dd121-286a-45e7-ac05-0686b5fe4af6_1720x940.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!uxaG!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea3dd121-286a-45e7-ac05-0686b5fe4af6_1720x940.png 424w, https://substackcdn.com/image/fetch/$s_!uxaG!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea3dd121-286a-45e7-ac05-0686b5fe4af6_1720x940.png 848w, https://substackcdn.com/image/fetch/$s_!uxaG!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea3dd121-286a-45e7-ac05-0686b5fe4af6_1720x940.png 1272w, https://substackcdn.com/image/fetch/$s_!uxaG!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea3dd121-286a-45e7-ac05-0686b5fe4af6_1720x940.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!uxaG!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea3dd121-286a-45e7-ac05-0686b5fe4af6_1720x940.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ea3dd121-286a-45e7-ac05-0686b5fe4af6_1720x940.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;c-ipopipeline&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="c-ipopipeline" title="c-ipopipeline" srcset="https://substackcdn.com/image/fetch/$s_!uxaG!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea3dd121-286a-45e7-ac05-0686b5fe4af6_1720x940.png 424w, https://substackcdn.com/image/fetch/$s_!uxaG!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea3dd121-286a-45e7-ac05-0686b5fe4af6_1720x940.png 848w, https://substackcdn.com/image/fetch/$s_!uxaG!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea3dd121-286a-45e7-ac05-0686b5fe4af6_1720x940.png 1272w, https://substackcdn.com/image/fetch/$s_!uxaG!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea3dd121-286a-45e7-ac05-0686b5fe4af6_1720x940.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>An unprecedented concentration.</strong> Three potential trillion-dollar debuts in roughly one hundred days, alongside the largest software listings in years. SpaceX, which acquired xAI in January 2026, is shown here near a reported IPO target of $1.5 trillion, the figure most frequently cited, with reports ranging from $1 trillion to $2 trillion. Sources: Datacenter Dynamics, CNN, Let&#8217;s Data Science, AI IPO Tracker.</figcaption></figure></div><div><hr></div><h2>Stated as strongly as I can make it</h2><p>A report that only sells the bull case is marketing. Here is the bear case, and it is not weak.</p><p><strong>Demand is showing its first cracks.</strong> Axios reported, in a piece timed to the filing, that Anthropic is going public just as businesses begin to rethink their AI spend, hit with what it called sticker shock. The writer <strong>Derek Thompson</strong> has named this the great AI cost panic of 2026, the phase where Fortune 500 buyers ask whether agentic AI is worth the bill. The most cited data point is the <strong>MIT Project NANDA </strong>study from July 2025, which found that ninety-five percent of enterprise generative-AI pilots produced zero measurable profit-and-loss impact, on thirty to forty billion dollars of corporate spending. </p><p>If even a fraction of that skepticism hardens into budget discipline, the revenue surge that justifies these valuations slows exactly when the infrastructure bills come due.</p><p><strong>The macro is stretched to dot-com proportions.</strong> The five largest Western hyperscalers are guiding toward roughly $725 billion of capex in 2026, up about seventy-seven percent from 2025&#8217;s record, per the Goldman, CreditSights, and Morgan Stanley estimates compiled by Tool Directory, with the trajectory projected past a trillion dollars annually in 2027.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!QEqS!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbf928a7c-306a-4d40-8fde-188d432b725e_1720x940.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!QEqS!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbf928a7c-306a-4d40-8fde-188d432b725e_1720x940.png 424w, https://substackcdn.com/image/fetch/$s_!QEqS!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbf928a7c-306a-4d40-8fde-188d432b725e_1720x940.png 848w, https://substackcdn.com/image/fetch/$s_!QEqS!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbf928a7c-306a-4d40-8fde-188d432b725e_1720x940.png 1272w, https://substackcdn.com/image/fetch/$s_!QEqS!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbf928a7c-306a-4d40-8fde-188d432b725e_1720x940.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!QEqS!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbf928a7c-306a-4d40-8fde-188d432b725e_1720x940.png" width="1456" height="796" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/bf928a7c-306a-4d40-8fde-188d432b725e_1720x940.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:796,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;c-capex&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="c-capex" title="c-capex" srcset="https://substackcdn.com/image/fetch/$s_!QEqS!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbf928a7c-306a-4d40-8fde-188d432b725e_1720x940.png 424w, https://substackcdn.com/image/fetch/$s_!QEqS!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbf928a7c-306a-4d40-8fde-188d432b725e_1720x940.png 848w, https://substackcdn.com/image/fetch/$s_!QEqS!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbf928a7c-306a-4d40-8fde-188d432b725e_1720x940.png 1272w, https://substackcdn.com/image/fetch/$s_!QEqS!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbf928a7c-306a-4d40-8fde-188d432b725e_1720x940.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><strong>The largest private investment cycle in history.</strong> Tech equipment and software hit 4.4% of GDP in 2025, near the dot-com peak (IEEE ComSoc). Cahn&#8217;s framework says ~$600B of annual revenue is needed to justify it; the gap is widening. Sources: Goldman Sachs, CreditSights, Morgan Stanley via Tool Directory.</figcaption></figure></div><p>The reminder that this complex can reprice violently is recent. The <strong>DeepSeek shock </strong>of January 2025, when a cheaper Chinese model wiped roughly a trillion dollars of US AI market value in a single day, including Nvidia&#8217;s $588.8 billion single-day loss, the largest in market history, shows how a single efficiency surprise can rerate the whole sector. And note the double edge. </p><p>In fact, the same open-weight efficiency that drives the cost deflation of <strong>the previous sections</strong>, is also the force most likely to compress the prices that section warned about. There is <strong>no reason another DeepSeek cannot happen</strong>. Pat Gelsinger, the former Intel chief, when asked whether this is a bubble, answered <em>of course.</em></p><p><strong>Antitrust is loading.</strong> The cumulative concentration of hyperscaler-lab pairings, Microsoft with OpenAI and Google and Amazon both with Anthropic, is large enough that, per Tech Insider&#8217;s reading, the <strong>DOJ</strong>, FTC, and <strong>European Commission</strong> are likely to revisit the structure, with a formal investigation plausible by the fourth quarter of 2026. An IPO does not resolve this. It raises the profile of the very arrangements regulators most want to examine.</p><p><strong>Company-specific overhangs are real.</strong> Sacra flags several that bear directly on Anthropic&#8217;s ability to monetize its best work. Its most capable model, previewed as <strong>Mythos </strong>and codenamed <strong>Capybara</strong>, has been withheld from general release after it proved able to identify thousands of high-severity software vulnerabilities; the commercial vehicle, Claude Security under <strong>Project Glasswing</strong> (<em>with partners including Nvidia, AWS, Apple, Google, Broadcom, Microsoft, Cisco, CrowdStrike, and Palo Alto Networks, per The Motley Fool</em>), runs on a deliberately constrained Opus 4.7. </p><p>That is a safety decision I think is correct on the merits, and it is simultaneously a self-imposed ceiling on revenue from the company&#8217;s most powerful capability. </p><p>Add the <strong>Pentagon dispute,</strong> where Anthropic sued the US government over a designation it viewed as a threat to its revenue, with CFO Krishna Rao testifying that the matter risked cutting 2026 revenue by multiple billions of dollars and where Amodei later publicly apologized for how the company handled the failed talks, and a legal overhang whose verified component is the roughly<strong> $1.5 billion authors&#8217; copyright settlement</strong> (<em>a larger figure naming the founder personally appears in a single analyst account and is not independently confirmed</em>), and you have a company whose brand safety and its revenue ceiling are the same wall.</p><p><strong>The customer-concentration paradox.</strong> Recall Claude Code&#8217;s billion-dollar ramp. Perera&#8217;s sharpest observation is that the company&#8217;s fastest-growing product may cannibalize its largest revenue source, because the same agentic coding capability enterprises buy <strong>directly can displace the API consumption </strong>those enterprises previously paid for. Growth in one column can quietly erode another. We cannot see the net effect, because the filing is confidential.</p><p>That phrase keeps recurring, and that is the point. Almost every load-bearing question, the real margin, the realized price per token, the quality of the revenue, the accounting for the <strong>SpaceX ramp</strong>, the net effect of Claude Code, resolves only when the public S-1 lands. </p><p>Until then, the bull and the bear are arguing about a black box.</p><div><hr></div><h2>Who should care, and why</h2><p>Strip away the spectacle and ask what actually changes downstream. Four things.</p><p><strong>For Nvidia and the silicon market, Anthropic is the existence proof.</strong> The most important fact in this report for the long-run structure of the industry is that a frontier-class model is being trained and served at scale on Trainium and TPU, not just on Nvidia. </p><p>The near-term risk to Nvidia is not revenue, since demand still dwarfs supply, but the long-run margin structure that depends on hyperscalers having no realistic alternative. </p><p>As previous sections showed, at rack scale, on a cost-per-workload basis, and on the bandwidth metric that actually governs inference, the alternative now exists. <strong>Every chip team</strong> at Amazon, Google, and Broadcom is using Anthropic as their proof of concept, and that is worth more to them than the revenue.</p><p><strong>For the circular-financing thesis, the IPO is the disclosure event the skeptics have awaited.</strong> A public listing forces the first concrete, audited window into a frontier lab&#8217;s financials. For two years the bears and bulls have argued about revenue quality with no primary data. </p><p>The S-1, once public, will show how Anthropic accounts for hyperscaler-funded revenue, how it books the SpaceX ramp, and what its real, unsubsidized unit economics, including the realized revenue per token that section VII could not pin down, actually look like. </p><p>This is the rare case where bull and bear should want the same thing: the numbers. If they are as good as the run-rate suggests, the bubble talk deflates. If they are not, better to know now.</p><p><strong>For enterprise buyers, the public-company transition changes the vendor relationship.</strong> Public companies optimize for quarterly margins in ways private ones do not. </p><p>The inference-cost deflation that has made Claude cheaper every year was partly funded by patient private capital. A public Anthropic, answerable to shareholders, may price differently, and may be less willing to pass the full token-cost decline through to customers. </p><p>For anyone building production systems on Claude, that is a planning input, not a panic, and one more argument for the multi-provider architectures that open-weight alternatives like DeepSeek and Llama keep making viable.</p><p><strong>For the private markets and the broader economy, the drain is unclogging, for better and worse.</strong> Three trillion-dollar IPOs in a hundred days will pull enormous capital into public AI equities and force index providers to confront new inclusion rules for companies this large arriving this fast. </p><blockquote><p><em>If the debuts go well, they validate the cycle and pull more capital in. </em></p><p><em>If they go poorly, three of the largest IPOs in history repricing in quick succession is exactly the event that turns a capex bubble into a capex correction, with the roundtripping amplifying the move down just as it amplified it up. </em></p></blockquote><p>The labor question rides alongside: tech layoffs passed 115,000 through May 2026, with Meta, Amazon, and Snap citing AI, even as the <strong>Yale Budget Lab</strong> found no significant change yet in the occupational mix of high-exposure jobs. The same plumbing runs in both directions, and so does the narrative.</p><blockquote><p><em>&#8220;What I see is this smooth exponential line. And that march has just been constant.&#8221;</em></p></blockquote><p><strong>Dario Amodei</strong> &#183; Anthropic CEO, at Davos 2026, on why he discounts the cycle of hype and bubble talk (Rest of World)</p><p>Set against that, Nvidia&#8217;s Huang, who has said publicly he disagrees with almost everything Amodei says, dismisses the doomier predictions as the product of a CEO <em>God complex.</em> T</p><p>wo of the most important people in the industry cannot agree on whether it is reshaping labor, let alone whether it is a bubble. The<strong> IPOs will not settle that</strong>. They will only price it.</p><div><hr></div><h2>The bottom line</h2><p>Here is what I actually think, stated plainly, with the byline caveat from the top still standing.</p><p>Anthropic is, by the public evidence, the best-positioned of the three companies going public this summer. It has the cleanest revenue mix, the leanest cost structure, the <strong>most credible margin-expansion story</strong>, and a genuinely differentiated three-silicon strategy that is reshaping the hardware layer beneath the entire industry. </p><p>The physics say a seventy-seven-percent gross margin is plausible rather than fanciful; the comps of previous parts say <strong>a trillion-dollar valuation is mid-pack rather than mad</strong>; the insider behavior at the tender offer and the speed of the run-rate point the same way. If forced to rank the three debuts on fundamentals rather than spectacle, Anthropic would be first.</p><p>And the entire case rests on three things we cannot yet verify and one we can. We <strong>cannot verify the real gross margin</strong>, the realized price per token, or the accounting behind the Q2 profitability claim, because the filing is confidential. </p><p>What we can verify is that the company is being valued at over a trillion dollars on exactly those unverified figures, inside a capital structure that critics, with a strong historical analogy, compare to the vendor financing that deepened the last great technology crash.</p><p>That is<strong> not a contradiction</strong>. It is the trade. The bull is betting the black box is full of seventy-seven-percent-margin, organically demanded, durably embedded enterprise revenue, with a cost-to-serve falling faster than price. </p><p>The bear is betting it is full of forty-five-percent-margin revenue, propped by a circular capital loop, with open-weight competition dragging price down to meet cost, in a demand environment that is just starting to flinch. </p><p><strong>The public S-1,</strong> fifteen days before the roadshow, opens the box. Everything before then, including the trillion-dollar valuation the market has already assigned, is a wager placed in the dark.</p><p><em><strong>The vertical is real. The loop is real. The physics is real. The only honest position, until the numbers are public, is to hold all three in view at once and refuse to pretend the box is open when it is still shut.</strong></em></p><div><hr></div><h2>Disclosures, statements, and the safety architecture</h2><p>What follows is the verifiable record behind the analysis: what was actually filed and said, on the record, by whom, and when. None of it is the confidential S-1&#8217;s financials, which remain sealed. Where a claim rests on a single source, it is marked.</p><h3>The filing, precisely</h3><p>Anthropic confirmed on June 1, 2026 that it had confidentially submitted a <strong>draft Form S-1 to the SEC</strong>, in an announcement made under Rule 135 of the Securities Act, a rule that by design states only that a filing exists: the number of shares and the offering price are explicitly undetermined, and the offering remains subject to market conditions. </p><p>Because the submission is confidential, Anthropic has disclosed no audited revenue, no margin, and no risk factors; <strong>under SEC rules</strong> for emerging growth companies those become public only about fifteen days before a roadshow. </p><p>The filing came four days after a cluster of disclosures on May 28, 2026: the close of a <em>$65 billion Series H at a $965 billion post-money valuation</em>, confirmation that run-rate revenue had crossed $47 billion earlier in May, and the release of <strong>Claude Opus 4.8. </strong>Multiple outlets place the listing target in an October 2026 window, above $1 trillion if markets cooperate.</p><p>One revealing side event: in May 2026 Anthropic warned about unauthorized transfers of its shares, naming several platforms selling unapproved, SPV-backed pre-IPO tokens and cautioning that such instruments may carry limited or no legal value. </p><p><strong>Tokenized Anthropic </strong>and OpenAI pre-IPO products reportedly fell 34 to 40 percent within days, per Bitcoin.com, a vivid illustration of demand for liquidity running far ahead of authorized supply, the same pressure visible in the earlier employee tender.</p><h3>What the executives said, on the record</h3><p>The single most important primary statement is Amodei&#8217;s own. At the Code with <strong>Claude </strong>developer conference in San Francisco on May 6, 2026, he said the company had planned for roughly tenfold annual growth but instead saw, in his words, <strong>eighty-fold annualized growth</strong> in the first quarter, which he gave as the direct cause of the company&#8217;s compute shortages, promising to pass that capacity to developers as fast as it could be brought online, per CNBC.</p><p><em>&#8220;In Q1 2026, we saw 80x annualized growth per year in revenue and usage.&#8221;</em></p><p><strong>Dario Amodei</strong> &#183; Anthropic CEO, Code with Claude, San Francisco, May 6, 2026 (CNBC)</p><p>That exuberance has a hard floor, and Amodei has named it precisely. On Dwarkesh Patel&#8217;s podcast in March 2026 he walked through the arithmetic of his own ruin: <strong>if he committed to a trillion dollars</strong> a year of compute in 2027 and revenue arrived even at $800 billion rather than the trillion he is extrapolating, then in his phrase there is &#8220;<em>no force on earth</em>&#8221; that could stop the company from going bankrupt. </p><p>It is the clearest admission any frontier-lab chief executive has made that the whole edifice is a bet on a growth rate continuing, and that the bet is existential. He has also said publicly that the industry may be near the end of the exponential.</p><p>The operator behind the IPO is CFO Krishna Rao, who joined in 2024 as the company <strong>closed its Series D at roughly $250 million in run-rate revenue</strong>, and who previously guided Airbnb&#8217;s IPO. Rao calls compute the lifeblood of the business and says he spends 30 to 40 percent of his time on it. </p><p>He has named the three risks that would push Anthropic toward the bottom of its growth cone rather than the top: enterprise diffusion failing to keep pace with model capability, scaling laws unexpectedly flattening, and <strong>competition eroding margins</strong>, per his interview with YourStory. </p><p>In a court filing around March 2026 Rao stated under oath that the company had brought in revenue exceeding $5 billion to date, the figure critics such as<strong> Ed Zitron </strong>use to argue the later profitability claim was flattered by the timing of the SpaceX compute discount.</p><blockquote><p><em>&#8220;The compute that we procure is the lifeblood of our business.&#8221;</em></p></blockquote><p><strong>Krishna Rao</strong> &#183; Anthropic CFO, who previously led Airbnb&#8217;s IPO (YourStory)</p><h3>Talent, culture, and the organization</h3><p>Anthropic&#8217;s defining operational claim is talent retention under siege. When Meta made aggressive offers across the frontier labs, Anthropic reportedly lost only two researchers where rivals lost dozens, a result Rao attributes to a culture the company describes as talent density over talent mass. </p><p><strong>All seven co-founders remain</strong>, as does the vast majority of the first thirty employees; every hire must clear a culture interview; and Amodei addresses the entire company every two weeks and takes unscripted questions, per YourStory. </p><p>In February 2026 Anthropic opened a <strong>Bengaluru office</strong>, calling India its second-largest market for Claude.</p><h3>The safety architecture, which is also a revenue constraint</h3><p>Anthropic governs releases through its <strong>Responsible Scaling Policy</strong>, first published in September 2023 and rewritten as Version 3.0, effective February 24, 2026, which introduced <strong>Frontier Safety Roadmaps</strong> and Risk Reports that quantify risk across deployed models. </p><p>The policy uses<em> AI Safety Levels</em> modeled on biosafety: <strong>ASL-2 </strong>is the current baseline; <strong>ASL-3</strong>, which Anthropic first activated alongside Claude Opus 4 in May 2025, adds hardened weight security and a narrow set of deployment limits aimed at CBRN misuse; <strong>ASL-4 is reserved</strong> for models posing major national-security risk or capable of autonomous AI research.</p><p> At Opus 4&#8217;s launch, chief scientist <strong>Jared Kaplan</strong> said the model gave novices a &#8220;significantly greater&#8221; uplift toward building biological weapons than a search engine or prior models, per TIME.</p><p>The clearest case of safety capping revenue is Claude Mythos. Announced as a preview on April 8, 2026, Mythos autonomously discovered, and wrote working exploits for, <strong>thousands of zero-day vulnerabilities</strong> across major operating systems and browsers, capability Anthropic judged too dangerous for general release, placing it at or near the ASL-3 cyber threshold per the <strong>Cloud Security Alliance.</strong> </p><p>The commercial vehicle, Claude Security under the Glasswing program, runs a deliberately constrained model and carries 90-day reporting commitments. </p><p>Anthropic has also published a Sabotage Risk Report for Opus 4.6 and, in February 2026, an internal <strong>Noncompliance Reporting and Anti-Retaliation Policy</strong> giving employees channels to flag potential violations.</p><p> Each of these is a decision that is defensible on its safety merits and simultaneously a self-imposed ceiling on the revenue the company&#8217;s most powerful capabilities could earn.</p><h3>The model record, and what Anthropic does not disclose</h3><p>The Claude lineage behind the revenue is precise and public at the capability level: </p><ul><li><p><strong>Opus 4.5</strong> (November 24, 2025) shipped a 200k-token context window and a 64k-token thinking budget, with up to 65 percent fewer tokens on long-horizon coding; </p></li><li><p><strong>Opus 4.6 </strong>(February 2026) added a 1M-token context in beta and led Terminal-Bench 2.0 and Humanity&#8217;s Last Exam; Sonnet 4.6 followed on February 17, 2026; </p></li><li><p><strong>Opus 4.7</strong> was independently rated the most EU-AI-Act-compliant model by the testing firm Aithos; </p></li><li><p>and <strong>Opus 4.8</strong> (May 28, 2026) carries a 1M-token context with cross-session memory for multi-day work. </p></li></ul><p>Critically, <strong>none of these system cards </strong>disclose the one thing an analyst most wants: parameter counts, layer counts, or whether the models are dense or mixture-of-experts. </p><p>Anthropic publishes capability and safety evaluations in exhaustive detail and keeps the architecture itself sealed, which is exactly why the <strong>inference unit-economics</strong> had to be built from silicon specifications and first principles rather than from any disclosed model size.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!hZAe!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6c38bec2-5425-4e51-9748-b29bba517da9_1747x2204.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!hZAe!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6c38bec2-5425-4e51-9748-b29bba517da9_1747x2204.png 424w, https://substackcdn.com/image/fetch/$s_!hZAe!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6c38bec2-5425-4e51-9748-b29bba517da9_1747x2204.png 848w, https://substackcdn.com/image/fetch/$s_!hZAe!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6c38bec2-5425-4e51-9748-b29bba517da9_1747x2204.png 1272w, https://substackcdn.com/image/fetch/$s_!hZAe!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6c38bec2-5425-4e51-9748-b29bba517da9_1747x2204.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!hZAe!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6c38bec2-5425-4e51-9748-b29bba517da9_1747x2204.png" width="1456" height="1837" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6c38bec2-5425-4e51-9748-b29bba517da9_1747x2204.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1837,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:302075,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/201004624?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6c38bec2-5425-4e51-9748-b29bba517da9_1747x2204.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!hZAe!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6c38bec2-5425-4e51-9748-b29bba517da9_1747x2204.png 424w, https://substackcdn.com/image/fetch/$s_!hZAe!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6c38bec2-5425-4e51-9748-b29bba517da9_1747x2204.png 848w, https://substackcdn.com/image/fetch/$s_!hZAe!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6c38bec2-5425-4e51-9748-b29bba517da9_1747x2204.png 1272w, https://substackcdn.com/image/fetch/$s_!hZAe!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6c38bec2-5425-4e51-9748-b29bba517da9_1747x2204.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div><hr></div><h2>The data, at a glance</h2><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!_jsS!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9d85271f-f124-4dcf-bf73-f7fc3961c6b3_2090x1531.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!_jsS!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9d85271f-f124-4dcf-bf73-f7fc3961c6b3_2090x1531.png 424w, https://substackcdn.com/image/fetch/$s_!_jsS!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9d85271f-f124-4dcf-bf73-f7fc3961c6b3_2090x1531.png 848w, https://substackcdn.com/image/fetch/$s_!_jsS!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9d85271f-f124-4dcf-bf73-f7fc3961c6b3_2090x1531.png 1272w, https://substackcdn.com/image/fetch/$s_!_jsS!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9d85271f-f124-4dcf-bf73-f7fc3961c6b3_2090x1531.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!_jsS!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9d85271f-f124-4dcf-bf73-f7fc3961c6b3_2090x1531.png" width="1456" height="1067" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/9d85271f-f124-4dcf-bf73-f7fc3961c6b3_2090x1531.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1067,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:255577,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/201004624?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9d85271f-f124-4dcf-bf73-f7fc3961c6b3_2090x1531.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!_jsS!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9d85271f-f124-4dcf-bf73-f7fc3961c6b3_2090x1531.png 424w, https://substackcdn.com/image/fetch/$s_!_jsS!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9d85271f-f124-4dcf-bf73-f7fc3961c6b3_2090x1531.png 848w, https://substackcdn.com/image/fetch/$s_!_jsS!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9d85271f-f124-4dcf-bf73-f7fc3961c6b3_2090x1531.png 1272w, https://substackcdn.com/image/fetch/$s_!_jsS!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9d85271f-f124-4dcf-bf73-f7fc3961c6b3_2090x1531.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!h5aL!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F387bc208-2a32-40ed-99f3-925c787ed8c3_2090x813.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!h5aL!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F387bc208-2a32-40ed-99f3-925c787ed8c3_2090x813.png 424w, https://substackcdn.com/image/fetch/$s_!h5aL!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F387bc208-2a32-40ed-99f3-925c787ed8c3_2090x813.png 848w, https://substackcdn.com/image/fetch/$s_!h5aL!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F387bc208-2a32-40ed-99f3-925c787ed8c3_2090x813.png 1272w, https://substackcdn.com/image/fetch/$s_!h5aL!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F387bc208-2a32-40ed-99f3-925c787ed8c3_2090x813.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!h5aL!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F387bc208-2a32-40ed-99f3-925c787ed8c3_2090x813.png" width="1456" height="566" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/387bc208-2a32-40ed-99f3-925c787ed8c3_2090x813.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:566,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:118065,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/201004624?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F387bc208-2a32-40ed-99f3-925c787ed8c3_2090x813.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!h5aL!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F387bc208-2a32-40ed-99f3-925c787ed8c3_2090x813.png 424w, https://substackcdn.com/image/fetch/$s_!h5aL!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F387bc208-2a32-40ed-99f3-925c787ed8c3_2090x813.png 848w, https://substackcdn.com/image/fetch/$s_!h5aL!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F387bc208-2a32-40ed-99f3-925c787ed8c3_2090x813.png 1272w, https://substackcdn.com/image/fetch/$s_!h5aL!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F387bc208-2a32-40ed-99f3-925c787ed8c3_2090x813.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!jo8a!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc390cf72-f89e-4955-bd8f-f0c39ce4a629_2090x1214.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!jo8a!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc390cf72-f89e-4955-bd8f-f0c39ce4a629_2090x1214.png 424w, https://substackcdn.com/image/fetch/$s_!jo8a!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc390cf72-f89e-4955-bd8f-f0c39ce4a629_2090x1214.png 848w, https://substackcdn.com/image/fetch/$s_!jo8a!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc390cf72-f89e-4955-bd8f-f0c39ce4a629_2090x1214.png 1272w, https://substackcdn.com/image/fetch/$s_!jo8a!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc390cf72-f89e-4955-bd8f-f0c39ce4a629_2090x1214.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!jo8a!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc390cf72-f89e-4955-bd8f-f0c39ce4a629_2090x1214.png" width="1456" height="846" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c390cf72-f89e-4955-bd8f-f0c39ce4a629_2090x1214.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:846,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:230800,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/201004624?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc390cf72-f89e-4955-bd8f-f0c39ce4a629_2090x1214.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!jo8a!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc390cf72-f89e-4955-bd8f-f0c39ce4a629_2090x1214.png 424w, https://substackcdn.com/image/fetch/$s_!jo8a!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc390cf72-f89e-4955-bd8f-f0c39ce4a629_2090x1214.png 848w, https://substackcdn.com/image/fetch/$s_!jo8a!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc390cf72-f89e-4955-bd8f-f0c39ce4a629_2090x1214.png 1272w, https://substackcdn.com/image/fetch/$s_!jo8a!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc390cf72-f89e-4955-bd8f-f0c39ce4a629_2090x1214.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><strong>Methodology.</strong> The cost-to-serve and gross-margin-bridge figures in section VII are illustrative models built from published silicon benchmarks and stated assumptions, <strong>not Anthropic disclosures</strong>; they are intended to show order of magnitude and mechanism, not to report the company&#8217;s actual numbers, which are confidential. </p><p>Revenue, valuation, and capacity figures throughout are reported, estimated, or projected by the cited third parties. Where sources conflict (for example, run-rate near $45B vs $47B, or Trainium3 availability dates), the range is given in text.</p><p><strong>Sourcing.</strong> <em>Reporting and analysis from CNBC, CNN, NBC News, Axios, Reuters, the Wall Street Journal (via AI Weekly and Sacra), Bloomberg, TechCrunch, Datacenter Dynamics, Yahoo Finance, Fortune, Benzinga, The Motley Fool, Futurum, Investing.com, and Business Standard; infrastructure, inference, and financial analysis from Sacra, PitchBook, Global Data Center Hub, Tech Insider, TradingKey, SaaStr, BlockEden, Tool Directory, IEEE ComSoc, IntuitionLabs, Introl, Inworld, GMI Cloud, Spheron, CloudRift, multiples.vc, and an arXiv study on token-price evolution; silicon specifications from Tom&#8217;s Hardware, Oplexa, Introl, and Awesome Agents; benchmark data from Nvidia InferenceMax; company statements from Anthropic, xAI, and Databricks; commentary from Ed Zitron, Derek Thompson, Shanaka Anslem Perera, and the named analysts (Rasgon / Bernstein, Ives / Wedbush, Cahn / Sequoia, Duberstein / Motley Fool, Patience / Futurum).</em></p><p><strong>Disclosure and limits.</strong> All financial figures are reported, estimated, or projected; none are drawn from a publicly available audited filing, because Anthropic&#8217;s S-1 remained confidential as of the filing date. Nothing here is investment advice. <em>This is analysis</em>, not a recommendation.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Software Frontier is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[The Blackwell Migration Question]]></title><description><![CDATA[When you should move Llama 3.3 70B inference from H100 to B200, when you should not, and what the real cost-per-token improvement is once you do.]]></description><link>https://www.thesoftwarefrontier.com/p/the-blackwell-migration-question</link><guid isPermaLink="false">https://www.thesoftwarefrontier.com/p/the-blackwell-migration-question</guid><dc:creator><![CDATA[Lorenzo Bradanini]]></dc:creator><pubDate>Sun, 07 Jun 2026 11:32:01 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!4tb9!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6758be5-f2fe-4087-b105-7ab352c95432_1122x1402.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!4tb9!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6758be5-f2fe-4087-b105-7ab352c95432_1122x1402.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!4tb9!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6758be5-f2fe-4087-b105-7ab352c95432_1122x1402.png 424w, https://substackcdn.com/image/fetch/$s_!4tb9!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6758be5-f2fe-4087-b105-7ab352c95432_1122x1402.png 848w, https://substackcdn.com/image/fetch/$s_!4tb9!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6758be5-f2fe-4087-b105-7ab352c95432_1122x1402.png 1272w, https://substackcdn.com/image/fetch/$s_!4tb9!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6758be5-f2fe-4087-b105-7ab352c95432_1122x1402.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!4tb9!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6758be5-f2fe-4087-b105-7ab352c95432_1122x1402.png" width="1122" height="1402" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b6758be5-f2fe-4087-b105-7ab352c95432_1122x1402.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1402,&quot;width&quot;:1122,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2346590,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/199710189?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6758be5-f2fe-4087-b105-7ab352c95432_1122x1402.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!4tb9!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6758be5-f2fe-4087-b105-7ab352c95432_1122x1402.png 424w, https://substackcdn.com/image/fetch/$s_!4tb9!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6758be5-f2fe-4087-b105-7ab352c95432_1122x1402.png 848w, https://substackcdn.com/image/fetch/$s_!4tb9!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6758be5-f2fe-4087-b105-7ab352c95432_1122x1402.png 1272w, https://substackcdn.com/image/fetch/$s_!4tb9!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6758be5-f2fe-4087-b105-7ab352c95432_1122x1402.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><h2>Introduction</h2><p>A single Blackwell B200 running Llama 3.3 70B at NVFP4 can decode at approximately <strong>5.5 milliseconds per token at batch size 1</strong>, an 18-month-old H100 at FP8 cannot move below <strong>21.5 ms</strong>. The B200 floor is roughly <strong>4&#215; faster</strong> and the KV-cache budget grows by about <strong>18&#215;</strong>. </p><p>The numbers are not vendor marketing. They are <strong>bandwidth arithmetic: </strong>7.7 TB/s of HBM3e divided by a ~43 GB on-GPU footprint for the FP4 weights, the same first-principles derivation we used for the H100 floor in Issue #1, with two parameters changed.</p><p>The question that follows is not whether Blackwell is faster. It is whether the <strong>price-per-token</strong> improvement actually delivered to a deployment engineer justifies the migration cost, and the published vendor numbers do not answer it cleanly. </p><p>NVIDIA&#8217;s October 2025 announcement of SemiAnalysis InferenceMAX v1 claims <strong>&#8220;15&#215; lower cost per million tokens&#8221;</strong> for Blackwell vs Hopper. </p><p>The real cost-per-token improvement for like-for-like single-GPU Llama 3.3 70B inference, derived from the same InferenceMAX v1 data combined with public on-demand pricing, is closer to <strong>3&#215; on-demand and 8&#215; on spot</strong>, with the gap explained by rack-scale GB200 NVL72 comparisons, MoE workloads, and BF16-baseline framings that do not represent the H100 FP8 production deployment most readers are actually running.</p><p>This is Issue #2 of what is now called <strong>Inference.Engineering</strong>. Issue #1 derived the bandwidth-bound floor for Llama 3.3 70B FP8 on H100 SXM5, audited the public benchmark landscape, identified where the engine-vs-engine variance actually lives, and committed to running our own measurements. </p><p>The<strong> measurement work</strong> is in progress and ships separately. This issue addresses the question deployment engineers are asking <em>now</em>: should I migrate, and what do I actually get if I do.</p><p>The post does what<strong> Issue #1 did</strong>. It derives the physical bounds from the NVIDIA datasheet and the FP4 weight footprint. It places the bounds against empirical data from SemiAnalysis InferenceMAX v1 and NVIDIA&#8217;s MLPerf v4.1 submission. </p><p>It audits the public Blackwell benchmark claims against those bounds, reads the kernel-layer story underneath (NVFP4 vs MXFP4, the <strong>second-generation Transformer Engine</strong>, FlashAttention-3 on SM100 vs SM90, NVLink 5), works through the quantization-accuracy tradeoff using Red Hat AI&#8217;s published <strong>evaluation of NVFP4 on Llama 3.3 70B</strong>, and concludes with a migration decision matrix for the six most common production situations.</p><p><strong>A note on what we have and have not done.</strong> This post is analysis grounded in primary-source data. The bounds are calculator-checkable. The empirical numbers come from SemiAnalysis InferenceMAX v1 (October 2025), NVIDIA&#8217;s own MLPerf submission, Red Hat AI&#8217;s published NVFP4 evaluation, and the vLLM v0.12.0 Blackwell recipe page. </p><p>We have not yet run our <strong>own B200 benchmarks</strong>. When we do, this post will be updated and any deltas will be tracked in the errata page. If the migration recommendations below turn out to be wrong when measured directly, cite this post against us.</p><p><strong>Inference.Engineering</strong> is reader-supported. The paid subscription funds rented GPU time on H100, B200, and MI355X, which is how we plan to measure the configurations this issue discusses.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Key observations</h2><ol><li><p><strong>Decode floor on a single B200 for Llama 3.3 70B NVFP4 is ~5.5 ms/token at batch 1</strong>, a per-user throughput ceiling of ~180 tok/s/user. Derived from ~43 GB on-GPU footprint at NVFP4 (34 GB FP4 linear weights + ~4 GB BF16 embedding/lm_head + ~4 GB block-scale overhead at NVFP4&#8217;s 16-element block size) divided by 7.7 TB/s HBM3e bandwidth on the HGX B200. The GB200 variant at 8.0 TB/s reaches ~5.3 ms.</p></li><li><p><strong>The B200&#8217;s KV-cache budget after NVFP4 weights is ~135 GB on a single 180 GB HGX B200 GPU</strong>, versus ~7&#8211;8 GB on a single 80 GB H100 at FP8. The 18&#215; increase is the load-bearing fact, not the throughput number. Long-context workloads that hit the KV wall on H100 (Part 1.2 of Issue #1) move comfortably into single-GPU territory on B200.</p></li><li><p><strong>NVFP4 on Llama 3.3 70B is essentially lossless when properly calibrated.</strong> Red Hat AI&#8217;s published NVFP4 evaluation (February 2026) shows large models (70B&#8211;235B parameters) &#8220;consistently achieve ~99% recovery&#8221; of BF16 accuracy across task-level and aggregate benchmarks. This holds for Llama 3.3 70B specifically; the published RedHatAI/Llama-3.3-70B-Instruct-NVFP4 checkpoint exists with full reproduction recipes.</p></li><li><p><strong>NVFP4 &#8800; MXFP4.</strong> On Llama 3.3 70B with quantized KV cache, NVIDIA measures <strong>5% higher MMLU accuracy with NVFP4 vs MXFP4</strong> (December 2025), attributed to NVFP4&#8217;s finer block scaling (16-element blocks vs 32) and higher-precision E4M3 FP8 scaling factors vs MXFP4&#8217;s E8M0. The distinction matters: a benchmark reporting &#8220;FP4&#8221; without specifying which format does not generalize.</p></li><li><p><strong>SemiAnalysis InferenceMAX v1 reports B200 delivering ~10,000 TPS/GPU at 50 tok/s/user interactivity on Llama 3.3 70B</strong>, roughly 4&#215; higher per-GPU throughput than H200 at the same interactivity (NVIDIA blog, October 9, 2025). Note that the batch-1 <em>decode-floor</em> ratio vs H200 is smaller (~2.7&#215;, from 1.6&#215; bandwidth times 1.69&#215; NVFP4 footprint); the larger 4&#215; at production interactivity additionally captures the bigger batch sizes the B200&#8217;s KV headroom permits and the FP4 compute density that helps once batches are large. Floor and interactivity-point throughput are different metrics, and the gap between them is itself informative.</p></li><li><p><strong>NVIDIA&#8217;s &#8220;15&#215; lower cost per million tokens&#8221; headline is not the right number for the single-GPU H100 &#8594; B200 question.</strong> It applies to rack-scale GB200 NVL72 on MoE models compared to HGX H100 air-cooled clusters at BF16. For like-for-like single-GPU Llama 3.3 70B FP8 (H100) vs NVFP4 (B200) at 50 tok/s/user interactivity, the derived cost-per-million-tokens improvement using Spheron&#8217;s published rates is <strong>~3&#215; on-demand and ~8&#215; on spot</strong>.</p></li><li><p><strong>The B200&#8217;s compute-bandwidth ridge sits at 2,338 FLOPs/byte for FP4 dense</strong> (18 PFLOPS / 7.7 TB/s) on the HGX variant, versus 591 FLOPs/byte at FP8 on H100. The Blackwell ridge is roughly 4&#215; further out, meaning Llama 3.3 70B decode at AI &#8776; 1 FLOP/byte sits even further below the ceiling. <strong>Decode is more bandwidth-bound on B200, not less.</strong> The throughput improvement comes from the bandwidth itself, not from compute throughput.</p></li><li><p><strong>vLLM v0.12.0 is the current Blackwell-ready release.</strong> NVIDIA&#8217;s vLLM recipe page is direct on the precision choice: <em>&#8220;For Hopper, FP8 offers the best performance for most workloads. For Blackwell, NVFP4 provides additional memory savings and throughput gains, but may require tuning to maintain accuracy on certain tasks.&#8221;</em> The recipe also documents <code>kv-cache-dtype: fp8</code> and <code>max-num-batched-tokens: 8192</code> as the recommended Llama 3.3 70B Blackwell defaults.</p></li><li><p><strong>The B200&#8217;s 1,000W TDP requires infrastructure most data centers do not yet have.</strong> H100 SXM5 runs at 700W in air-cooled configurations. B200 at 1,000W in HGX form (or 1,200W in GB200) frequently requires liquid cooling or a substantial reduction in rack density. The migration cost includes infrastructure, not just hardware.</p></li><li><p><strong>The cross-vendor portability story is real and underappreciated.</strong> vLLM&#8217;s Triton attention backend, which achieved 100.7% of FlashAttention-3 performance on Hopper for long decode requests (Issue #1, Section 4.2), runs unchanged on Blackwell and AMD. <strong>The same Triton kernel that closes the gap to FA3 on H100 also closes the gap to FA4 on B200.</strong> A migration evaluation that compares only the NVIDIA-optimized stack systematically understates the AMD alternative on the workloads where Triton attention wins.</p></li><li><p><strong>The B200 supply situation matters for the timeline.</strong> Inworld reports B200 hardware orders backlogged through mid-2026 with ~3.6 million units in queue (April 2026 reference). Cloud rental is the practical migration path for most readers in 2026, not on-premise purchase. This makes the spot-price economics more relevant than the list-price comparisons NVIDIA&#8217;s blog posts emphasize.</p></li></ol><div><hr></div><h2>Part 1, The physical bounds on Blackwell</h2><p>Three numbers belong on every Llama 3.3 70B / B200 deployment engineer&#8217;s whiteboard: the decode latency floor under NVFP4, the KV-cache budget that the larger HBM3e capacity unlocks, and the compute roofline ridge that shifts under the second-generation Transformer Engine.</p><h3>1.1 The decode latency floor under NVFP4</h3><p>The bandwidth-bound floor derivation from Issue #1 transfers unchanged. Decode on a single B200 streams the full weight set from HBM3e to the SMs per token. </p><p>At batch 1, decode arithmetic intensity remains <strong>~1 FLOP per byte</strong> loaded (<em>Spector &amp; R&#233;; arXiv 2603.05931</em>). The H100 figures from Issue #1 become Blackwell figures by substituting the new weight footprint and the new bandwidth.</p><p>NVIDIA B200 datasheet figures, cross-verified against the official December 2024 Blackwell datasheet:</p><blockquote><p><em>Metric HGX B200 (per GPU) GB200 (per GPU) HBM3e bandwidth 7.7 TB/s 8.0 TB/s FP4 dense 18 PFLOPS 20 PFLOPS FP8/FP6 dense 9 PFLOPS 10 PFLOPS BF16/FP16 dense 5 PFLOPS 5 PFLOPS HBM3e capacity 180 GB 186 GB TDP 1,000 W up to 1,200 W NVLink generation 5 (1.8 TB/s) 5 (1.8 TB/s)</em></p></blockquote><p>(All <strong>Tensor Core figures</strong> are dense; sparse values are 2&#215; the dense.)</p><p>Llama 3.3 70B on-GPU footprint under NVFP4 decomposes as follows. NVFP4 quantizes <strong>linear-layer weights to 4 bits </strong>per parameter in 16-element blocks with one E4M3 FP8 scale per block. Embedding and lm_head remain in BF16 (the standard production path that Red Hat AI&#8217;s checkpoint uses, mirroring the FP8 convention from Issue #1):</p><pre><code><code>Linear FP4 weights:        68.45B params &#215; 0.5 bytes = 34.23 GB
Block-scale overhead:      68.45B params / 16 &#215; 1 byte = 4.28 GB
Embedding + lm_head BF16:  2.10B params &#215; 2 bytes = 4.20 GB
Runtime + activations + CUDA graphs: ~1 GB
                                              ---------
Total on-GPU footprint:                      ~43.7 GB
</code></code></pre><p>The lower bound on decode latency per token at batch 1 uses the <strong>per-token streaming footprint</strong> of ~42.7 GB (the FP4 linear weights, their block scales, and the BF16 lm_head, all of which stream from HBM every decode step). </p><p>The extra <strong>~1 GB of runtime </strong>and activation buffers in the ~43.7 GB total does not stream per token, so it enters the KV-budget subtraction in Part 1.2 but not the floor:</p><pre><code><code>HGX B200:  t_floor = 42.7 GB / 7.7 TB/s  &#8776; 5.5 ms/token
           v_ceiling = 1 / t_floor       &#8776; 180 tok/s/user

GB200:     t_floor = 42.7 GB / 8.0 TB/s  &#8776; 5.3 ms/token
           v_ceiling = 1 / t_floor       &#8776; 187 tok/s/user
</code></code></pre><p>Comparing to the H100 SXM5 / FP8 reference from Issue #1 (~21.5 ms / 46.5 tok/s/user), the Blackwell + NVFP4 combination delivers a roughly <strong>4&#215; decode-floor improvement</strong> at batch 1. The improvement decomposes into two contributions: bandwidth (7.7 vs 3.35 TB/s = 2.3&#215;) and streaming footprint (42.7 vs 72 GB = 1.69&#215;). </p><p>The footprint contribution is what makes NVFP4 valuable, separate from the hardware. Running<strong> B200 with FP8 weights</strong> instead of NVFP4 would cut the footprint contribution to ~1&#215; and recover only the 2.3&#215; bandwidth improvement.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!kuRA!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15155d3f-9812-4994-b5bb-3f8467d23cf8_2579x1305.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!kuRA!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15155d3f-9812-4994-b5bb-3f8467d23cf8_2579x1305.png 424w, https://substackcdn.com/image/fetch/$s_!kuRA!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15155d3f-9812-4994-b5bb-3f8467d23cf8_2579x1305.png 848w, https://substackcdn.com/image/fetch/$s_!kuRA!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15155d3f-9812-4994-b5bb-3f8467d23cf8_2579x1305.png 1272w, https://substackcdn.com/image/fetch/$s_!kuRA!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15155d3f-9812-4994-b5bb-3f8467d23cf8_2579x1305.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!kuRA!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15155d3f-9812-4994-b5bb-3f8467d23cf8_2579x1305.png" width="1456" height="737" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/15155d3f-9812-4994-b5bb-3f8467d23cf8_2579x1305.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:737,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:191207,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/199710189?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15155d3f-9812-4994-b5bb-3f8467d23cf8_2579x1305.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!kuRA!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15155d3f-9812-4994-b5bb-3f8467d23cf8_2579x1305.png 424w, https://substackcdn.com/image/fetch/$s_!kuRA!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15155d3f-9812-4994-b5bb-3f8467d23cf8_2579x1305.png 848w, https://substackcdn.com/image/fetch/$s_!kuRA!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15155d3f-9812-4994-b5bb-3f8467d23cf8_2579x1305.png 1272w, https://substackcdn.com/image/fetch/$s_!kuRA!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15155d3f-9812-4994-b5bb-3f8467d23cf8_2579x1305.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>Chart 1. The decode latency floor and the per-user throughput ceiling it implies, derived from each GPU&#8217;s HBM bandwidth divided by the Llama 3.3 70B weight footprint at the relevant precision. B200 NVFP4 is ~3.9x faster at the floor than H100 FP8. These are physical lower bounds; real systems sit slightly above them.</em></p><p><strong>Verify the floor yourself.</strong> The bound is checkable on any rented B200 in under ten minutes, using NVIDIA&#8217;s vLLM v0.12.0 Blackwell recipe as the launch baseline:</p><pre><code><code>docker pull vllm/vllm-openai:v0.12.0

docker run --gpus all -p 8000:8000 vllm/vllm-openai:v0.12.0 \
  --model nvidia/Llama-3.3-70B-Instruct-NVFP4 \
  --kv-cache-dtype fp8 \
  --max-num-batched-tokens 8192 \
  --gpu-memory-utilization 0.9
</code></code></pre><p>Then measure<strong> single-stream decode </strong>at concurrency 1:</p><pre><code><code>vllm bench serve \
  --model nvidia/Llama-3.3-70B-Instruct-NVFP4 \
  --dataset-name random --random-input-len 512 --random-output-len 512 \
  --max-concurrency 1 --num-prompts 8 \
  --percentile-metrics ttft,tpot,itl,e2el --ignore-eos
</code></code></pre><p>NVIDIA also publishes the canonical reproduction in its <code>dgxc-benchmarking</code> repository, which runs Llama 3.3 70B inference with <code>--dtype nvfp4</code> on <strong>B200/GB200</strong> and exposes TPOT directly when streaming is enabled (<code>STREAMING=true</code>); </p><p>the same harness runs the <strong>H100 baseline</strong> at <code>--dtype fp8</code>, making it the cleanest apples-to-apples generational comparison available without writing your own client.</p><p>Expected mean <strong>TPOT </strong>should land in the <strong>~6&#8211;8 ms range on HGX B200</strong>, at or just above the 5.5 ms floor, with the ~1&#8211;2 ms gap being engine overhead and kernel-launch latency. If you see TPOT materially below 5.5 ms at batch 1, the engine is running speculative decoding or sub-NVFP4 weights. </p><p>If materially above ~10 ms, the precision path is wrong (<em>the engine fell back to FP8 or BF16</em>) or the configuration is wrong (cold CUDA graphs, an unconfigured Transformer Engine path). Per-user throughput is <strong>1000 / TPOT_ms.</strong></p><p><strong>A caveat we will not paper over.</strong> Unlike Issue #1, where NVIDIA&#8217;s own NIM benchmark published the H100 throughput-vs-concurrency curve that empirically confirmed the<strong> KV wall</strong>, we have not located a published single-stream (batch-1) B200 TPOT measurement for Llama 3.3 70B NVFP4 to anchor the 5.5 ms floor directly. </p><p>The public <strong>B200 numbers</strong> (InferenceMAX, MLPerf, the TRT-LLM perf tables) are all aggregate-throughput or production-interactivity points at large batch, not batch-1 latency. The <strong>5.5 ms</strong> figure is therefore a bandwidth derivation, not yet an empirically confirmed measurement, and confirming it within 15% is the first specific claim<strong> Issue #3 will test</strong>. </p><p>Treat it as a physical lower bound that real systems approach from above, exactly as the H100 floor behaved before NIM data confirmed it.</p><h3>1.2 The KV-cache budget grows by an order of magnitude</h3><p>The KV-cache calculation transfers unchanged from Issue #1. Llama 3.3 70B uses Grouped-Query Attention with 8 KV heads, head_dim 128, across 80 layers. Per-token KV-cache footprint is ~327 KB at BF16 or ~164 KB at FP8.</p><p>After loading the ~43.7 GB NVFP4 weight footprint and reserving ~1 GB for runtime, the KV-cache budget on a single 180 GB HGX B200 is approximately <strong>135 GB</strong>, versus ~7&#8211;8 GB on a single 80 GB H100 at FP8:</p><p>KV precision Bytes/token KV tokens on H100 (~7&#8211;8 GB) KV tokens on B200 HGX (~135 GB) BF16 327 KB ~21,000&#8211;24,500 <strong>~413,000</strong> FP8 164 KB ~42,000&#8211;49,000 <strong>~825,000</strong></p><p>At avg_seq_len = 8,192 tokens (the vLLM V1 default chunked-prefill budget), <strong>single-H100 concurrency is ~3</strong>. Single-B200 concurrency is ~50 at the same configuration. At avg_seq_len = 32,768 (long-context agentic workloads), single-H100 supports less than one full request in memory; single-B200 supports ~12.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!kMXr!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F809e8b0b-8a18-4923-970f-6d1f7c8920e4_2325x1279.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!kMXr!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F809e8b0b-8a18-4923-970f-6d1f7c8920e4_2325x1279.png 424w, https://substackcdn.com/image/fetch/$s_!kMXr!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F809e8b0b-8a18-4923-970f-6d1f7c8920e4_2325x1279.png 848w, https://substackcdn.com/image/fetch/$s_!kMXr!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F809e8b0b-8a18-4923-970f-6d1f7c8920e4_2325x1279.png 1272w, https://substackcdn.com/image/fetch/$s_!kMXr!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F809e8b0b-8a18-4923-970f-6d1f7c8920e4_2325x1279.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!kMXr!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F809e8b0b-8a18-4923-970f-6d1f7c8920e4_2325x1279.png" width="1456" height="801" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/809e8b0b-8a18-4923-970f-6d1f7c8920e4_2325x1279.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:801,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:185879,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/199710189?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F809e8b0b-8a18-4923-970f-6d1f7c8920e4_2325x1279.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!kMXr!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F809e8b0b-8a18-4923-970f-6d1f7c8920e4_2325x1279.png 424w, https://substackcdn.com/image/fetch/$s_!kMXr!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F809e8b0b-8a18-4923-970f-6d1f7c8920e4_2325x1279.png 848w, https://substackcdn.com/image/fetch/$s_!kMXr!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F809e8b0b-8a18-4923-970f-6d1f7c8920e4_2325x1279.png 1272w, https://substackcdn.com/image/fetch/$s_!kMXr!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F809e8b0b-8a18-4923-970f-6d1f7c8920e4_2325x1279.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>Chart 2. KV-cache budget remaining after weights load, single GPU. The B200&#8217;s ~135 GB is ~18x the H100&#8217;s ~7.5 GB. Note that the H200 already reaches ~67 GB (~9x), capturing half the step-change, which is why an existing H200 fleet has a weaker migration case than an H100 fleet (Part 6.1).</em></p><p>This is the load-bearing finding of the Blackwell migration question for long-context production: <strong>the workloads that required TP=2 or TP=4 on Hopper for KV headroom fit comfortably on a single B200</strong>. The migration changes which workloads are &#8220;single-GPU&#8221; deployments.</p><h3>1.3 The roofline shifts further out</h3><p>The B200 compute-bandwidth ridge moves to:</p><pre><code><code>AI_ridge (B200 FP4 HGX) = 18 PFLOPS / 7.7 TB/s  &#8776; 2,338 FLOPs/byte
AI_ridge (B200 FP8 HGX) =  9 PFLOPS / 7.7 TB/s  &#8776; 1,169 FLOPs/byte
AI_ridge (H100 FP8)     = 1,979 TFLOPS / 3.35 TB/s &#8776; 591 FLOPs/byte
</code></code></pre><p>The B200 FP4 ridge is roughly 4&#215; further out than the H100 FP8 ridge. Decode at batch 1 with AI &#8776; 1 sits at the same arithmetic intensity but achieves only ~7.7 TFLOPS on B200 (0.04% of FP4 peak), vs ~3.4 TFLOPS on H100 (0.17% of FP8 peak). </p><p><strong>The fraction of peak compute used by decode falls as the hardware improves</strong>, because the bound is bandwidth, not compute. The Blackwell FP4 compute density is largely irrelevant to batch-1 decode latency; it matters at higher batch and in prefill.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!TcVb!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc2bfce12-ca13-4f59-b440-b851aeda00c5_2409x1402.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!TcVb!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc2bfce12-ca13-4f59-b440-b851aeda00c5_2409x1402.png 424w, https://substackcdn.com/image/fetch/$s_!TcVb!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc2bfce12-ca13-4f59-b440-b851aeda00c5_2409x1402.png 848w, https://substackcdn.com/image/fetch/$s_!TcVb!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc2bfce12-ca13-4f59-b440-b851aeda00c5_2409x1402.png 1272w, https://substackcdn.com/image/fetch/$s_!TcVb!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc2bfce12-ca13-4f59-b440-b851aeda00c5_2409x1402.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!TcVb!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc2bfce12-ca13-4f59-b440-b851aeda00c5_2409x1402.png" width="1456" height="847" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c2bfce12-ca13-4f59-b440-b851aeda00c5_2409x1402.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:847,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:265266,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/199710189?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc2bfce12-ca13-4f59-b440-b851aeda00c5_2409x1402.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!TcVb!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc2bfce12-ca13-4f59-b440-b851aeda00c5_2409x1402.png 424w, https://substackcdn.com/image/fetch/$s_!TcVb!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc2bfce12-ca13-4f59-b440-b851aeda00c5_2409x1402.png 848w, https://substackcdn.com/image/fetch/$s_!TcVb!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc2bfce12-ca13-4f59-b440-b851aeda00c5_2409x1402.png 1272w, https://substackcdn.com/image/fetch/$s_!TcVb!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc2bfce12-ca13-4f59-b440-b851aeda00c5_2409x1402.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>Chart 3. The B200 FP4 compute-bandwidth ridge sits ~4x further right than the H100 FP8 ridge. Batch-1 decode at arithmetic intensity ~1 FLOP/byte uses an even smaller fraction of peak compute on Blackwell (0.04%) than on Hopper (0.17%): the throughput gain is bandwidth, not compute. Compute density only matters in the prefill regime (shaded).</em></p><p>Prefill on B200 at FP4, with arithmetic intensity in the <strong>2,000&#8211;10,000 FLOPs/byte range</strong>, runs comfortably in the compute-bound regime above the ridge. At 75% utilization on the FP4 ceiling, the per-GPU prefill throughput on Llama 3.3 70B reaches approximately:</p><pre><code><code>prefill_ceiling = 0.75 &#215; 18 PFLOPS / (2 &#215; 70.55 &#215; 10^9 FLOPs/token) &#8776; 96,000 tok/s
</code></code></pre><p>versus ~10,500 tok/s on H100 at FP8. <strong>A ~9&#215; prefill throughput improvement</strong>, which dominates the workloads where input length is much larger than output length (RAG, document QA, code understanding). </p><p>The bandwidth-bound decode improvement is &#8220;<em>only</em>&#8221; ~4&#215;; the compute-bound prefill improvement is much larger.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Part 2, Empirical verification from InferenceMAX v1 and MLPerf</h2><p>The Part 1 derivation predicts a ~4&#215; decode floor improvement and a ~9&#215; prefill ceiling improvement for Llama 3.3 70B going from H100 FP8 to B200 NVFP4. Three empirical sources test the prediction, and getting the comparison right requires understanding precisely how each was measured.</p><p><strong>SemiAnalysis InferenceMAX v1 (October 9, 2025)</strong> is the most rigorous public source, and its methodology deserves close reading because the headline numbers are easy to misquote. </p><p>InferenceMAX runs nightly on hundreds of chips via <strong>GitHub Actions, </strong>sweeping max-concurrency and parallelism to trace the full throughput-vs-interactivity <strong>Pareto frontier</strong> rather than reporting a single point. </p><p>For each run it uses random sequences (no prefix caching, to avoid the workload-dependent complexity prefix caching introduces), an infinite request rate with a capped max-concurrency, and three input/output sequence-length pairs chosen to represent distinct workload classes: <strong>1K in / 1K out (chat), 1K in / 8K out (reasoning), and 8K in / 1K out (summarization)</strong>, with each request&#8217;s input length randomized to 80&#8211;100% of nominal to mimic real traffic variance. </p><p>For Llama 3.3 70B specifically, the default engine is <strong>vLLM</strong> and the Blackwell precision is <strong>NVFP4</strong>; SGLang is the DeepSeek default and TRT-LLM is run where vendors submit configs.</p><p>The InferenceMAX v1 finding on Llama 3.3 70B is unambiguous: &#8220;When it comes to <strong>LLaMA 70B FP4</strong>, B200 significantly outperforms MI355X across all three workload types.&#8221; NVIDIA&#8217;s accompanying announcement quantifies the generational comparison: Blackwell delivers over 10,000 TPS per GPU at <strong>50 TPS/user interactivity</strong> on Llama 3.3 70B, 4&#215; higher per-GPU throughput than H200.</p><p>Three things about how this number is constructed, each of which a careless reader gets wrong. First, <strong>50 tok/s/user is a production interactivity floor, not the single-stream maximum.</strong> </p><p>A B200 at single-stream batch 1 reaches the ~180 tok/s/user ceiling from Part 1.1; the <strong>50 tok/s/user point </strong>sits at a higher concurrency where weight loads amortize across many users, which is exactly the throughput-vs-interactivity trade-off that defines the Pareto curve. </p><p>Second, <strong>the 4&#215; comparison is vs H200, not H100.</strong> H200&#8217;s 4.8 TB/s bandwidth is ~43% higher than H100&#8217;s 3.35 TB/s, so the B200-vs-H100 ratio is correspondingly larger than 4&#215;. Third, the comparison holds <strong>NVFP4 on B200 against FP8 on the Hopper part</strong>, so the 4&#215; already bundles the precision-footprint contribution; it is not a pure-hardware ratio.</p><p>Our derivation decomposes the <strong>B200-vs-H100 ratio</strong> cleanly: 2.3&#215; from bandwidth (<em>7.7 vs 3.35 TB/s</em>) times ~1.69&#215; from the NVFP4 streaming-footprint reduction (<em>42.7 vs 72 GB</em>) gives ~3.9&#215; at the bandwidth-bound decode floor. </p><p>The InferenceMAX 4&#215;-vs-H200 number, adjusted for the H200-to-H100 bandwidth gap, lands in the same region. The decomposition matters more than the point estimate: it tells a deployment engineer that running B200 at <strong>FP8 instead of NVFP4</strong> sacrifices the 1.69&#215; footprint term and recovers only the 2.3&#215; bandwidth term.</p><p><strong>MLPerf Inference v4.1</strong> is the second anchor. NVIDIA&#8217;s August 2024 submission with one B200 GPU on Llama 2 70B reported <strong>10,756 tokens/s server scenario and 11,264 tokens/s offline scenario</strong>, against a per-GPU H100 figure derived by dividing the eight-GPU H100 submission by eight, yielding a reported 4&#215; server / 3.7&#215; offline per-GPU increase. </p><p>Llama 2 70B and Llama 3.3 70B share the relevant architecture (<em>80 layers, hidden 8192, 64 Q heads, 8 KV heads in 8:1 GQA</em>), so the <em>ratio</em> transfers even though the absolute number is for Llama 2.</p><p>The <strong>three paths cluster </strong>around 4&#215;, though on deliberately different baselines that are worth stating precisely. InferenceMAX v1 (production interactivity, NVFP4-vs-FP8): ~4&#215; vs H200. </p><p><strong>MLPerf v4.1 </strong>(offline, peak throughput): ~3.7&#8211;4&#215; vs H100 per-GPU. First-principles (bandwidth-bound decode floor): ~3.9&#215; vs H100. The two H100-baselined numbers agree <strong>tightly at ~3.9&#215;; </strong>the InferenceMAX figure is vs the faster H200, so restated against H100 it would be somewhat <em>larger</em> than 4&#215;, meaning the three are mutually consistent rather than coincidentally equal. </p><p>Convergence of an independent nightly benchmark, a standardized industry benchmark, and a <strong>from-scratch derivation</strong> in the same band, with the baseline differences accounted for, is what separates a defensible figure from a vendor claim repeated.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!gFbx!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F406c05c0-5ca7-49c4-b5e2-2fd11bb9dd42_2193x1263.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!gFbx!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F406c05c0-5ca7-49c4-b5e2-2fd11bb9dd42_2193x1263.png 424w, https://substackcdn.com/image/fetch/$s_!gFbx!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F406c05c0-5ca7-49c4-b5e2-2fd11bb9dd42_2193x1263.png 848w, https://substackcdn.com/image/fetch/$s_!gFbx!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F406c05c0-5ca7-49c4-b5e2-2fd11bb9dd42_2193x1263.png 1272w, https://substackcdn.com/image/fetch/$s_!gFbx!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F406c05c0-5ca7-49c4-b5e2-2fd11bb9dd42_2193x1263.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!gFbx!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F406c05c0-5ca7-49c4-b5e2-2fd11bb9dd42_2193x1263.png" width="1456" height="839" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/406c05c0-5ca7-49c4-b5e2-2fd11bb9dd42_2193x1263.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:839,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:171308,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/199710189?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F406c05c0-5ca7-49c4-b5e2-2fd11bb9dd42_2193x1263.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!gFbx!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F406c05c0-5ca7-49c4-b5e2-2fd11bb9dd42_2193x1263.png 424w, https://substackcdn.com/image/fetch/$s_!gFbx!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F406c05c0-5ca7-49c4-b5e2-2fd11bb9dd42_2193x1263.png 848w, https://substackcdn.com/image/fetch/$s_!gFbx!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F406c05c0-5ca7-49c4-b5e2-2fd11bb9dd42_2193x1263.png 1272w, https://substackcdn.com/image/fetch/$s_!gFbx!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F406c05c0-5ca7-49c4-b5e2-2fd11bb9dd42_2193x1263.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>Chart 4. The throughput-improvement figure does not rest on any single source. InferenceMAX v1 (vs H200), MLPerf v4.1 (vs H100 per-GPU), and the first-principles bandwidth derivation (vs H100) all land in the 3.8-4x band.</em></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Part 3, The cost-per-token economics</h2><p>Two factors determine the cost-per-token improvement: throughput-per-GPU (~4&#215;) and price-per-GPU-hour (varies by cloud, ~2&#8211;3&#215;). The product is the cost-per-token ratio.</p><h3>3.1 The real cost-per-million-tokens math</h3><p>Using Spheron&#8217;s published rates (consistent with the H100 numbers in Issue #1):</p><p>Hardware On-demand $/hr Spot $/hr Throughput at 50 tok/s/user Cost / 1M tokens (on-demand) Cost / 1M tokens (spot) H100 SXM5 (FP8) $2.64 $1.66 ~1,500 TPS $0.49 $0.31 H200 SXM5 (FP8) $4.62 $1.92 ~2,500 TPS $0.51 $0.21 B200 HGX (NVFP4) $6.03 $2.12 ~10,000 TPS <strong>$0.17</strong> <strong>$0.059</strong></p><p>All three per-hour rates are <strong>Spheron&#8217;s published figures</strong> as of 22 May 2026 (H100 $2.64/$1.66, H200 $4.62/$1.92, B200 $6.03/$2.12 on-demand/spot). H100 and H200 throughput at 50 tok/s/user is extrapolated from the SemiAnalysis H200-vs-B200 4&#215; ratio combined with the H100-to-H200 1.43&#215; bandwidth scaling; B200 throughput is the InferenceMAX v1 figure.</p><p>Note the <strong>counterintuitive H200 </strong>on-demand row: at $4.62/hr its cost-per-token ($0.51/M) is actually slightly <em>worse</em> than the H100&#8217;s ($0.49/M), because the ~43% throughput gain from the higher bandwidth does not fully offset the ~75% price premium. </p><p>The H200&#8217;s real advantages are its memory capacity (the KV-wall relief in Part 6.1) and its aggressive spot rate ($0.21/M), not its on-demand cost-per-token. This is exactly the kind of inversion that a vendor &#8220;1.4&#215; faster inference&#8221; headline hides: faster does not mean cheaper-per-token unless the price scales sublinearly with the throughput.</p><p><strong>The honest derived improvement, on-demand vs on-demand, is ~2.9&#215; cheaper per million tokens.</strong> On spot vs spot, where B200 spot pricing is currently aggressive (Spheron $2.12/hr is lower than H100 on-demand), the improvement reaches <strong>~8&#215; cheaper</strong>.</p><p>This is materially different from NVIDIA&#8217;s &#8220;15&#215; lower cost per million tokens&#8221; headline. The 15&#215; number is real, but it compares <strong>rack-scale GB200 NVL72 air-cooled to H100 HGX air-cooled</strong> on MoE workloads (GPT-MoE-1.8T projected throughput, per the NVIDIA datasheet), not single-GPU Llama 3.3 70B. NVIDIA&#8217;s announcement post is precise about this: the cost-per-million-tokens claim applies to &#8220;AI factory economics&#8221; at rack scale, not the single-GPU comparison most readers are actually evaluating.</p><p>The 3&#215; to 8&#215; range is what a deployment engineer should plan around. The 15&#215; number is what a CFO will hear from a vendor presentation.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!p79B!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F67f47707-f7c6-4b28-a69e-f74577eae968_2279x1321.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!p79B!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F67f47707-f7c6-4b28-a69e-f74577eae968_2279x1321.png 424w, https://substackcdn.com/image/fetch/$s_!p79B!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F67f47707-f7c6-4b28-a69e-f74577eae968_2279x1321.png 848w, https://substackcdn.com/image/fetch/$s_!p79B!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F67f47707-f7c6-4b28-a69e-f74577eae968_2279x1321.png 1272w, https://substackcdn.com/image/fetch/$s_!p79B!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F67f47707-f7c6-4b28-a69e-f74577eae968_2279x1321.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!p79B!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F67f47707-f7c6-4b28-a69e-f74577eae968_2279x1321.png" width="1456" height="844" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/67f47707-f7c6-4b28-a69e-f74577eae968_2279x1321.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:844,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:198647,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/199710189?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F67f47707-f7c6-4b28-a69e-f74577eae968_2279x1321.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!p79B!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F67f47707-f7c6-4b28-a69e-f74577eae968_2279x1321.png 424w, https://substackcdn.com/image/fetch/$s_!p79B!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F67f47707-f7c6-4b28-a69e-f74577eae968_2279x1321.png 848w, https://substackcdn.com/image/fetch/$s_!p79B!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F67f47707-f7c6-4b28-a69e-f74577eae968_2279x1321.png 1272w, https://substackcdn.com/image/fetch/$s_!p79B!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F67f47707-f7c6-4b28-a69e-f74577eae968_2279x1321.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>Chart 5. Cost per million output tokens at 50 tok/s/user, from Spheron&#8217;s published per-hour rates times InferenceMAX v1 throughput. The honest single-GPU improvement is ~2.9x on-demand and ~8x on spot, well short of NVIDIA&#8217;s 15x headline, which is a rack-scale GB200 NVL72 MoE-at-BF16 comparison.</em></p><h3>3.2 The migration cost the price-per-token comparison hides</h3><p>Three categories of migration cost do not appear in the $/M-tokens calculation but determine whether the migration actually saves money.</p><p><strong>Infrastructure.</strong> B200 HGX runs at 1,000W per GPU vs H100 SXM5&#8217;s 700W. GB200 in the NVL72 configuration runs at up to 1,200W per GPU. Rack density falls accordingly, and air-cooling becomes marginal at B200&#8217;s power level. Liquid cooling is the production path NVIDIA assumes. Data centers without liquid cooling either run fewer GPUs per rack (reducing the effective $/hr improvement) or deploy in a different facility (capital expenditure not in the per-hour rate). For cloud renters the infrastructure cost is bundled into the per-hour price; for on-premise buyers it is not, and it can easily exceed the GPU hardware cost over a deployment lifetime.</p><p><strong>Software maturity.</strong> vLLM v0.12.0 is the current Blackwell-ready release as of the Llama 3.3 70B recipe page. NVFP4 support landed first in TensorRT-LLM (full); vLLM has shipped what NVIDIA describes as &#8220;early NVFP4 support&#8221; with the second-generation Transformer Engine path; SGLang has it on the roadmap. Production deployment in 2026 means accepting a less-mature software stack than Hopper&#8217;s, with the kernel and scheduler optimizations of Issue #1 (FlashInfer integration, Triton attention, FA3-vs-FA4 selection, CUDA graph capture) still being landed across the three engines for the new architecture. The Hopper stack benefits from 18 additional months of production hardening that Blackwell will need time to accumulate.</p><p><strong>Supply.</strong> Hardware order backlogs through mid-2026 mean on-premise B200 purchases compete with hyperscaler allocations. Cloud rental is the practical path, and cloud pricing on B200 is more volatile than on H100. Spheron&#8217;s $2.12/hr spot rate is real but not guaranteed across providers or time. The H100 ecosystem is mature enough that a one-year reserved instance commits to a known price; B200 commits are quoted but the secondary-market spread is wider.</p><h3>3.3 A worked migration example</h3><p>Take the Issue #1 worked example: a chat product at 50 requests/second peak sustained, 800 input tokens, 400 output tokens. Decode dominates; the output token rate is 50 &#215; 400 = 20,000 tok/s. On H100 FP8 at production interactivity, a single GPU sustained ~1,500 TPS, requiring 14 H100s and costing ~$27,000/month on-demand or ~$17,000/month at spot.</p><p>On B200 NVFP4 at the same interactivity, a single GPU sustains ~10,000 TPS, so the same workload needs <strong>~2 B200 GPUs</strong>. Cost: $6.03/hr &#215; 2 GPUs &#215; 730 hr/month = <strong>~$8,800/month on-demand</strong>, or ~$3,100/month on spot. The on-demand savings vs H100 are roughly $18,000/month; the spot savings are roughly $14,000/month, depending on which side of the comparison gets spot pricing.</p><p><strong>Three caveats</strong> sharpen the picture. First, the 2-GPU B200 deployment has dramatically more KV headroom (~270 GB aggregate budget vs ~16 GB on 2 H100s), so the same hardware can absorb longer sequences or higher concurrency before scaling out further. </p><p>Second, prefill (40,000 in-tok/s required) needs only <strong>one B200 GPU&#8217;s worth of prefill capacity </strong>(~96,000 tok/s ceiling vs the workload&#8217;s 40,000), so the deployment is decode-bound everywhere. </p><p>Third, the autoscaling math improves: 2 GPUs is a less granular floor than 14, but the per-GPU economics make <strong>off-peak headroom</strong> less costly. A peak-to-average ratio of 2&#215; on H100 means paying for 28 GPU-hours per peak-hour-equivalent of demand; on B200 it means paying for 4. </p><p>The autoscaling discontinuities matter less when each GPU is doing more work.</p><h3>3.4 What the per-hour rate actually pays for</h3><p>The single most important fact about inference TCO is structural, and SemiAnalysis states it directly in the InferenceMAX v1 analysis: <strong>colocation rent and electricity together are typically less than 20% of total cost of ownership.</strong> </p><p>The dominant term is the GPU vendor&#8217;s gross margin. In SemiAnalysis&#8217;s words, some vendors charge &#8220;up to 75% gross margins (i.e. a 4&#215; markup over cost of goods sold), while others less than 50% gross margins (i.e. less than 2&#215; cost of goods sold).&#8221;</p><p>This reframes the migration question. A naive analysis treats the B200&#8217;s higher power draw (1,000W vs 700W) as a major cost penalty. It is not: if a B200 delivers <strong>20% fewer tokens per provisioned megawatt</strong> than some alternative, that translates to less than 4% of TCO (20% of the under-20% energy share). </p><p>The migration economics are dominated by the price you pay for the silicon, which is set by vendor margin and supply, not by the power bill. </p><p>This is why the <strong>spot-vs-on-demand spread </strong>(a ~3&#215; swing in the per-hour rate) dwarfs every efficiency consideration in the cost-per-token math: the per-hour rate is mostly margin, and margin is what moves between spot and on-demand.</p><p>The corollary for a buyer: the highest-leverage variable in <strong>Blackwell migration economics</strong> is procurement, not engineering. Securing B200 capacity at the $2.12/hr spot rate rather than the $6.03/hr on-demand rate is a larger cost-per-token lever than any kernel optimization, any precision choice, or any batching strategy. </p><p>The engineering determines the throughput; the contract determines most of the cost.</p><h3>3.5 Power efficiency: the megawatt view</h3><p>For power-constrained deployments (the binding constraint for many AI data centers in 2026), throughput per provisioned megawatt is the relevant metric rather than throughput per GPU. </p><p><strong>SemiAnalysis InferenceMAX v1</strong> measured this directly, though the published Llama-vs-Blackwell power numbers are for gpt-oss 120B rather than Llama 3.3 70B, so they should be read as a Blackwell-generation indicator rather than a Llama-specific figure.</p><p>On <strong>gpt-oss 120B FP4</strong>, an HGX H100 processes ~900,000 tok/s per all-in provisioned megawatt; an HGX B200 processes ~2.8 million tok/s per megawatt, a ~3&#215; generational power-efficiency gain. At a higher interactivity level of<strong> ~180 tok/s/user</strong>, the B200 advantage widens to ~7&#215;, because the H100 falls off its efficiency curve faster at high interactivity than the B200 does.</p><p> (<em>Provisioned megawatt here means all-in utility power including cooling and electrical-distribution overhead, not just GPU TDP, which is the honest denominator for a data-center operator.</em>)</p><p>The <strong>~3&#215; power-efficiency improvement</strong> roughly tracks the ~4&#215; throughput improvement, which is expected: the B200&#8217;s TDP (1,000W) is ~1.43&#215; the H100&#8217;s (700W), and 4&#215; throughput / 1.43&#215; power &#8776; 2.8&#215; efficiency. </p><p>For a Llama 3.3 70B deployment specifically, the power-efficiency improvement should land in the same 3&#215; range by the same arithmetic, but we flag that the published megawatt figures are gpt-oss numbers and the <strong>Llama-specific power measurement</strong> is one of the things Issue #3 will report directly.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!PJzh!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa538b84b-fbb5-4585-9067-dfe8682492e2_2332x1251.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!PJzh!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa538b84b-fbb5-4585-9067-dfe8682492e2_2332x1251.png 424w, https://substackcdn.com/image/fetch/$s_!PJzh!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa538b84b-fbb5-4585-9067-dfe8682492e2_2332x1251.png 848w, https://substackcdn.com/image/fetch/$s_!PJzh!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa538b84b-fbb5-4585-9067-dfe8682492e2_2332x1251.png 1272w, https://substackcdn.com/image/fetch/$s_!PJzh!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa538b84b-fbb5-4585-9067-dfe8682492e2_2332x1251.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!PJzh!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa538b84b-fbb5-4585-9067-dfe8682492e2_2332x1251.png" width="1456" height="781" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a538b84b-fbb5-4585-9067-dfe8682492e2_2332x1251.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:781,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:174161,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/199710189?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa538b84b-fbb5-4585-9067-dfe8682492e2_2332x1251.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!PJzh!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa538b84b-fbb5-4585-9067-dfe8682492e2_2332x1251.png 424w, https://substackcdn.com/image/fetch/$s_!PJzh!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa538b84b-fbb5-4585-9067-dfe8682492e2_2332x1251.png 848w, https://substackcdn.com/image/fetch/$s_!PJzh!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa538b84b-fbb5-4585-9067-dfe8682492e2_2332x1251.png 1272w, https://substackcdn.com/image/fetch/$s_!PJzh!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa538b84b-fbb5-4585-9067-dfe8682492e2_2332x1251.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>Chart 6. Throughput per provisioned megawatt, H100 vs B200, on gpt-oss 120B FP4 (the published InferenceMAX figures are for gpt-oss, not Llama 3.3 70B, and are shown here as a Blackwell-generation indicator). The ~3x gain at production interactivity widens to ~7x at high interactivity because the H100 falls off its efficiency curve faster.</em></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p>
      <p>
          <a href="https://www.thesoftwarefrontier.com/p/the-blackwell-migration-question">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[How Systems Really Fail, Part IV]]></title><description><![CDATA[The complexity problem: why no one designed the system, why no one fully understands it, and why it works anyway.]]></description><link>https://www.thesoftwarefrontier.com/p/how-systems-really-fail-part-iv</link><guid isPermaLink="false">https://www.thesoftwarefrontier.com/p/how-systems-really-fail-part-iv</guid><dc:creator><![CDATA[Lorenzo Bradanini]]></dc:creator><pubDate>Wed, 03 Jun 2026 16:26:41 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!FgTX!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fabd34f5d-0905-4eea-b838-8eb361e2567b_1118x1407.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!FgTX!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fabd34f5d-0905-4eea-b838-8eb361e2567b_1118x1407.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!FgTX!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fabd34f5d-0905-4eea-b838-8eb361e2567b_1118x1407.png 424w, https://substackcdn.com/image/fetch/$s_!FgTX!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fabd34f5d-0905-4eea-b838-8eb361e2567b_1118x1407.png 848w, https://substackcdn.com/image/fetch/$s_!FgTX!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fabd34f5d-0905-4eea-b838-8eb361e2567b_1118x1407.png 1272w, https://substackcdn.com/image/fetch/$s_!FgTX!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fabd34f5d-0905-4eea-b838-8eb361e2567b_1118x1407.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!FgTX!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fabd34f5d-0905-4eea-b838-8eb361e2567b_1118x1407.png" width="1118" height="1407" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/abd34f5d-0905-4eea-b838-8eb361e2567b_1118x1407.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1407,&quot;width&quot;:1118,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2815007,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/190603006?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fabd34f5d-0905-4eea-b838-8eb361e2567b_1118x1407.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!FgTX!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fabd34f5d-0905-4eea-b838-8eb361e2567b_1118x1407.png 424w, https://substackcdn.com/image/fetch/$s_!FgTX!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fabd34f5d-0905-4eea-b838-8eb361e2567b_1118x1407.png 848w, https://substackcdn.com/image/fetch/$s_!FgTX!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fabd34f5d-0905-4eea-b838-8eb361e2567b_1118x1407.png 1272w, https://substackcdn.com/image/fetch/$s_!FgTX!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fabd34f5d-0905-4eea-b838-8eb361e2567b_1118x1407.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>The System Has No Architect</h2><p>The <strong>first essay</strong> argued that distributed systems fail in the spaces between their components.  The second argued that the system you observe is a delayed, partial projection of a system that has already moved on. </p><p>The third argued that the <strong>control loops</strong> you close on that projection cannot stabilise the system you have. </p><p>Each was a<strong> structural failure</strong> (composition, observation, control) and each was localisable: you could point to where the failure lived. This essay is about the failure that is not localised anywhere.</p><p>The first three essays assumed there is a system to fail. Call it S, the object engineers <em>compose</em>, <em>observe</em>, and <em>control</em>. The composition fails because interfaces hide state. The observation fails because dashboards project S into a representable space and discard the rest.<strong> The control fails</strong> because the loop cannot reach the parts of S that matter. </p><p>In all three cases the failures are gaps between S and the operator&#8217;s representation of it. The argument of this essay is that, past a certain scale, <strong>S is not the object the operators think it is.</strong></p><p>Production distributed systems at scale are complex systems in the technical sense the term carries in the work of <em>Perrow, Cilliers, Snowden, Leveson, and Dekker</em>. Their behaviour is not the sum of their parts. </p><p>Their <strong>failure modes are emergent</strong>, properties of the whole that no component possesses in isolation. They have no architect. They cannot be modelled in their entirety by any individual. </p><p>And they sit, by construction, near the edge of failure, because the same <em>competitive pressure</em> that makes them efficient drives them toward the boundary of safe operation.</p><p>The three earlier failures are manifestations of this. The gaps between components cannot be closed because nobody knows where the gaps are. The dashboards <strong>cannot be made complete </strong>because the system is dimensionally larger than any representation. </p><p>The control loops do not close because the plant is not a plant in the textbook sense; it is an emergent process whose dynamics are not contained in any single component&#8217;s specification.</p><blockquote><p><em>Past a certain scale, the system is not the object the operators think it is. And complex systems do not yield to the methods that produced reliable software at small scale.</em></p></blockquote><p>Four incidents, mechanically reconstructed. </p><ul><li><p>The<strong> 2003 Northeast Blackout,</strong> where a race condition in one utility&#8217;s alarm system propagated, through coupling no one had mapped, into a cascade that affected fifty-five million people. </p></li><li><p>The <strong>AWS S3 outage of February 2017</strong>, where one mistyped command exposed a dependency graph no individual had seen end to end. </p></li><li><p>The <strong>Cloudflare WAF outage of July 2019</strong>, where one regular expression, deployed globally in seconds, took down a large fraction of the internet&#8217;s HTTPS. </p></li></ul><p>And, because the thesis is that these failures are structural rather than historical, a fourth: the <strong>Cloudflare outage of November 2025</strong>, where the same company repeated a structurally identical failure six years later for reasons unrelated to the specific bug. </p><p>Four emergent modes: cascade through coupling, concentration through scale, worst case from interaction, and the recurrence that proves the pattern is not an accident.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>What kind of object is a production system?</h2><p>The canonical work is Charles Perrow&#8217;s <em>Normal Accidents</em> (1984), written after Three Mile Island to characterise the systems in which accidents become <strong>statistically inevitable</strong>. </p><p>His framework has <strong>two axes</strong>. Interactive complexity is how many ways the components can affect each other, including ways the designers did not anticipate. </p><p>Coupling is how tightly they are connected in time, whether a failure must propagate immediately or whether there is slack to absorb it.<strong> </strong></p><p><strong>Loosely-coupled</strong>, linear systems (assembly lines, road networks) fail in expected, containable ways. </p><p>Complex, <strong>tightly-coupled systems</strong> (nuclear plants, refineries, the financial system, the internet) fail in ways their operators did not anticipate, and the failures propagate before anyone can intervene.</p><p>Perrow&#8217;s claim is sharp: complex tightly-coupled systems are not made safe by adding safety features. Past a threshold, features add interactions, which add failure modes. The system has accidents as a normal property of operation, not a deviation from it. </p><p>As he restated it in 2012, a <strong>normal accident</strong> is one where everyone tries hard to play safe, but unexpected interaction of two or more failures (interactive complexity) causes a cascade (tight coupling).</p><p>Production distributed systems lie at the far end of both axes. The <em>interactive complexity</em> is enormous: every service has dozens of dependencies, each version-skewed against the others, interacting through coupling paths invisible in the architecture diagrams. </p><p>The coupling is tight: cache TTLs in seconds, retry budgets in tens of milliseconds, timeouts tuned down over years until they barely accommodate the steady state. There is no buffer left.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!QiRC!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb5142ec7-ae70-407c-8168-7d084832d99d_1728x1104.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!QiRC!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb5142ec7-ae70-407c-8168-7d084832d99d_1728x1104.png 424w, https://substackcdn.com/image/fetch/$s_!QiRC!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb5142ec7-ae70-407c-8168-7d084832d99d_1728x1104.png 848w, https://substackcdn.com/image/fetch/$s_!QiRC!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb5142ec7-ae70-407c-8168-7d084832d99d_1728x1104.png 1272w, https://substackcdn.com/image/fetch/$s_!QiRC!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb5142ec7-ae70-407c-8168-7d084832d99d_1728x1104.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!QiRC!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb5142ec7-ae70-407c-8168-7d084832d99d_1728x1104.png" width="1456" height="930" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b5142ec7-ae70-407c-8168-7d084832d99d_1728x1104.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:930,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:161254,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/190603006?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb5142ec7-ae70-407c-8168-7d084832d99d_1728x1104.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!QiRC!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb5142ec7-ae70-407c-8168-7d084832d99d_1728x1104.png 424w, https://substackcdn.com/image/fetch/$s_!QiRC!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb5142ec7-ae70-407c-8168-7d084832d99d_1728x1104.png 848w, https://substackcdn.com/image/fetch/$s_!QiRC!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb5142ec7-ae70-407c-8168-7d084832d99d_1728x1104.png 1272w, https://substackcdn.com/image/fetch/$s_!QiRC!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb5142ec7-ae70-407c-8168-7d084832d99d_1728x1104.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>Figure 1, Perrow&#8217;s two-axis taxonomy. Systems in the upper-right quadrant produce accidents as a normal property of operation. Latency optimisation and dependency growth push production distributed systems monotonically toward the extreme corner.</em></p><p><strong>Three frameworks</strong> sharpen the point. Cilliers distinguished the complicated (<em>many parts, knowable structure, holdable in one head</em>) from the complex (<em>interactions producing behaviours not present in any part and not predictable from the specifications</em>). </p><p><strong>Snowden&#8217;s Cynefin</strong> adds the prescription: in complex contexts cause and effect are clear only in retrospect, so operators must probe, sense, and respond rather than analyse and execute. </p><p><strong>Leveson&#8217;s STAMP </strong>reframes accidents as failures of control over the interactions between components, not failures of the components themselves. </p><p>The compressed claim: the methods that produce reliable software at small scale cannot transfer to large scale, because the object they were built to control is no longer the object that exists.</p><div><hr></div><h2>Cascade: the Northeast Blackout, 14 August 2003</h2><p>At 12:15 EDT, the <strong>state estimator at MISO</strong>, the reliability coordinator for much of the Midwest and Ontario, began diverging from its measurements.</p><p> A state estimator infers load and voltage on every line from noisier telemetry every few minutes; when it converges the operator has a coherent picture of the grid, and when it does not the operator is blind to anything not directly measured. </p><p>An analyst traced the divergence to a tripped Indiana line, fixed the topology by hand, then forgot to re-enable the estimator&#8217;s automatic trigger. </p><p>From 12:37 until 16:04, the window in which the cascade silently assembled itself, MISO&#8217;s contingency analysis was effectively offline: operators could no longer answer &#8220;<em>what happens if line X trips?</em>&#8221; because they had no current model of where the lines were.</p><p>At 13:31 EDT, FirstEnergy&#8217;s Eastlake Unit 5 tripped while carrying 612 MW and 400 MVAr of reactive power. The lost generation should have been absorbed; the lost reactive support depressed voltage across northern Ohio. </p><p>At 14:14 EDT, the alarm processor of GE&#8217;s XA/21 system at <strong>FirstEnergy&#8217;s Akron </strong>control centre deadlocked. The control room was entirely alarm-driven: operators responded to alarms rather than watching the mimic.</p><p> A latent race condition, a deadlock under high event-queue depth, silently stopped the primary alarm server. No error was raised. The backup took over, inherited the same growing queue, and deadlocked too. </p><p>From then until well after the cascade ended, operators believed they were watching a current, stable grid. They were watching a <strong>stale snapshot </strong>from before the cascade started, with no signal it was stale.</p><blockquote><p><em>For three and a half hours the operators were looking at a stale snapshot with no signal that it was stale. The dashboard was not wrong. It was describing a system that no longer existed.</em></p></blockquote><p>At 15:05, 15:32, and 15:41 EDT, three northeast Ohio lines sagged into trees and tripped. Each loss pushed its load onto the survivors, which carried more current, heated, expanded, sagged, and contacted more vegetation. </p><p>The cascade is a <strong>positive-feedback loop</strong> in continuous time: the same physics that produces normal operation produces accelerating failure once a threshold is crossed.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Bod4!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2132200c-2323-4b23-b4a5-20e0c83176bd_1728x864.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Bod4!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2132200c-2323-4b23-b4a5-20e0c83176bd_1728x864.png 424w, https://substackcdn.com/image/fetch/$s_!Bod4!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2132200c-2323-4b23-b4a5-20e0c83176bd_1728x864.png 848w, https://substackcdn.com/image/fetch/$s_!Bod4!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2132200c-2323-4b23-b4a5-20e0c83176bd_1728x864.png 1272w, https://substackcdn.com/image/fetch/$s_!Bod4!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2132200c-2323-4b23-b4a5-20e0c83176bd_1728x864.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Bod4!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2132200c-2323-4b23-b4a5-20e0c83176bd_1728x864.png" width="1456" height="728" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2132200c-2323-4b23-b4a5-20e0c83176bd_1728x864.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:728,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:86454,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/190603006?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2132200c-2323-4b23-b4a5-20e0c83176bd_1728x864.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Bod4!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2132200c-2323-4b23-b4a5-20e0c83176bd_1728x864.png 424w, https://substackcdn.com/image/fetch/$s_!Bod4!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2132200c-2323-4b23-b4a5-20e0c83176bd_1728x864.png 848w, https://substackcdn.com/image/fetch/$s_!Bod4!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2132200c-2323-4b23-b4a5-20e0c83176bd_1728x864.png 1272w, https://substackcdn.com/image/fetch/$s_!Bod4!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2132200c-2323-4b23-b4a5-20e0c83176bd_1728x864.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>Figure 2. The cascade assembles slowly inside a 3.5-hour blind window, then accelerates. The transition is sharp: before 16:05:57 the event was confined to Ohio and recoverable by load-shedding; after it, the cascade propagated at the speed of protective-relay action.</em></p><p>At 16:05:57 EDT the <strong>Sammis-Star line </strong>tripped, not from a tree but from over-current. That was the transition point. After it the cascade became uncontrolled, propagating across the interconnection in milliseconds to seconds per trip. </p><p>Within eight minutes, lines tripped across Ohio, Michigan, Pennsylvania, New York, and Ontario; the <strong>Eastern Interconnection</strong> separated into islands; generators tripped as frequencies diverged from 60 Hz. By 16:13, roughly 508 units at 265 plants were offline and fifty-five million people had lost power. </p><p>The <strong>US-Canada Task Force</strong> estimated four to ten billion dollars in losses and named four causes: two operational (unmaintained vegetation, failure to shed load in time) and two observational (no effective contingency analysis, failure of the monitoring tools).</p><p>The interesting feature is not the bug. The continental cascade was contingent on the interaction of <strong>three unrelated things:</strong> trees in Ohio, an alarm system that worked correctly except under one event-ordering pattern, and a grid coupled tightly enough that a loss in Ohio reached Ontario in eleven minutes. </p><p>None was dangerous in isolation; the danger was <strong>emergent </strong>from the combination. No component was outside its specified parameters when the cascade began. </p><p>In Leveson&#8217;s terms, the regional control structure had not been designed to enforce the constraint that no single utility&#8217;s blindness shall propagate beyond its control area, and the cascade would have happened the same way had the <strong>XA/21 bug</strong> been a different bug producing the same blindness in the same window.</p><div><hr></div><h2>AWS S3, 28 February 2017</h2><p>At 9:37 AM PST, an authorised S3 engineer ran a routine playbook to remove a few servers from the billing subsystem. One parameter was wrong, and the command removed a much larger set. </p><p>A familiar story so far: a <strong>fat-finger event</strong>, an under-validated tool, a blast radius beyond intent. The interesting part is what happened next. </p><p>The removed servers also supported the index subsystem <em>(metadata and location for every object, required for GET, LIST, PUT, DELETE</em>) and the placement subsystem (which depends on the index). </p><p>Both dropped below the capacity they needed and entered a state requiring a full restart.</p><p>Here the <strong>structural problem </strong>surfaced. AWS later wrote that the full restart, relied on since launch, had not been run on the index or placement subsystems in their larger regions for many years. <strong>S3 in us-east-1</strong> launched in 2006, and the metadata had since grown by orders of magnitude. </p><p>The restart procedure, designed at a far smaller scale, took dramatically longer than expected, because each system coming back had to validate a metadata store far larger than the procedure assumed, and the validation scaled non-linearly. Roughly <strong>four hours of impact.</strong></p><p>What failed during those four hours is the substantive part. </p><p>S3 had quietly become a<strong> hidden dependency</strong> for a remarkable fraction of the public internet: Slack, Quora, Trello, Imgur, Medium, Coursera, GitHub release artefacts, Docker Hub, Adobe Creative Cloud, parts of Zillow and Expedia, each having independently chosen S3, each apparently unaware how many others had. </p><p>The cumulative effect was a single point of failure for a non-trivial percentage of the public internet that no organisation had ever ratified. </p><p><strong>The cutting detail</strong>: AWS&#8217;s own Service Health Dashboard depended on S3 to host its status icons, so it could not visually update to reflect the outage of the service it depended on. The icons stayed green because the red ones could not load.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!XT65!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F34b8bf19-c9a0-454c-9c38-02f537bc75dc_1728x1032.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!XT65!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F34b8bf19-c9a0-454c-9c38-02f537bc75dc_1728x1032.png 424w, https://substackcdn.com/image/fetch/$s_!XT65!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F34b8bf19-c9a0-454c-9c38-02f537bc75dc_1728x1032.png 848w, https://substackcdn.com/image/fetch/$s_!XT65!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F34b8bf19-c9a0-454c-9c38-02f537bc75dc_1728x1032.png 1272w, https://substackcdn.com/image/fetch/$s_!XT65!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F34b8bf19-c9a0-454c-9c38-02f537bc75dc_1728x1032.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!XT65!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F34b8bf19-c9a0-454c-9c38-02f537bc75dc_1728x1032.png" width="1456" height="870" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/34b8bf19-c9a0-454c-9c38-02f537bc75dc_1728x1032.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:870,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:172168,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/190603006?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F34b8bf19-c9a0-454c-9c38-02f537bc75dc_1728x1032.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!XT65!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F34b8bf19-c9a0-454c-9c38-02f537bc75dc_1728x1032.png 424w, https://substackcdn.com/image/fetch/$s_!XT65!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F34b8bf19-c9a0-454c-9c38-02f537bc75dc_1728x1032.png 848w, https://substackcdn.com/image/fetch/$s_!XT65!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F34b8bf19-c9a0-454c-9c38-02f537bc75dc_1728x1032.png 1272w, https://substackcdn.com/image/fetch/$s_!XT65!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F34b8bf19-c9a0-454c-9c38-02f537bc75dc_1728x1032.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>Figure 3. The concentration paradox. Each spoke is an independent, individually-rational decision to use the cheapest reliable object store. Nobody ratified the convergence and nobody can see it. The self-loop is the recovery-dependency cycle: the status page for S3 was itself served from S3.</em></p><p>The shape here is not cascade. The failure was <strong>localised to one region</strong> of one service. What made it catastrophic was concentration: an enormous number of independent systems had converged on the same choice without anyone designing the convergence. </p><p>None had chosen to share fate; they had each chosen, separately, the cheapest reliable object store, and that was S3 in us-east-1. The shared fate was <strong>emergent </strong>from the aggregation of independent decisions. </p><p>Network science calls the mechanism <strong>preferential attachment</strong>: independent decisions under similar constraints converge on a few providers, and the providers become single points of failure for the aggregate without anyone designing it.</p><blockquote><p><em>The most reliable service in a category becomes the largest single point of failure for that category, precisely because it is the most reliable. The convergence is rational at the individual level and catastrophic at the aggregate level, and no one can intervene against it because no one can see how many others made the same choice.</em></p></blockquote><p>This is emergence of a different kind from cascade. The blackout&#8217;s was <strong>dynamic</strong>, failures propagating along time-coupled paths in minutes.</p><p> S3&#8217;s was <strong>structural</strong>, dependencies accreting along static paths nobody maintained a record of, over a decade. Both are failures of the whole, in the precise sense that no subsystem was responsible. </p><p>What made the <strong>blast radius</strong> continental was the way thousands of independent decisions had made S3 a chokepoint no individual fully understood. </p><p>AWS&#8217;s fix, <em>beyond safer commands</em>, was to partition the index subsystem into cells so a future restart brings back one cell at a time. </p><p>In <strong>Perrow&#8217;s terms</strong>, that is an explicit attempt to reduce coupling and reintroduce slack between subsystems whose tight coupling had silently emerged from organic growth.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Worst case from interaction: Cloudflare WAF, 2 July 2019</h2><p>At 13:42 UTC, a Cloudflare engineer deployed a minor change to the <strong>WAF Managed Rules</strong>, a new rule meant to improve detection of inline JavaScript used in cross-site-scripting attacks. </p><p>It went out via <strong>Quicksilver</strong>, which propagates configuration to every edge server globally in seconds, by design, so emergency security responses do not wait for a gradual rollout. </p><p>Three minutes later the first page fired. CPU on every Cloudflare edge worldwide had spiked to 100%. Cloudflare&#8217;s network, by 2019 fronting roughly <strong>ten percent of the world&#8217;s HTTPS</strong>, could not process new requests. </p><p>Customer sites returned <strong>502s</strong>; Cloudflare&#8217;s own dashboard, API, and internal tools, all routed through the same edge, went unreachable. At its worst, traffic dropped 82%.</p><p>The rule&#8217;s structural problem was a trailing pattern of the shape <code>.*(?:.*=.*)</code>. A backtracking regex engine handles <code>.*</code> greedily: it matches as much as possible, then, if the rest fails, gives back one character at a time and retries. </p><p>With several unanchored <code>.*</code> constructs in sequence, the number of match positions grows combinatorially in the input length. Against an input that almost matches but diverges late, a pattern of the form <code>.*.*=.*</code> with no equals sign explores on the order of n-choose-2 partition points: </p><p>quadratic, O(n&#178;), for this shape, and exponential, O(2&#8319;), in pathological cases like <code>(a+)+</code>. Quadratic and cubic blow-ups are equally lethal at millions of requests per second.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!MhSm!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8965ae9e-a4f7-408e-8f15-036887e5642c_1728x912.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!MhSm!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8965ae9e-a4f7-408e-8f15-036887e5642c_1728x912.png 424w, https://substackcdn.com/image/fetch/$s_!MhSm!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8965ae9e-a4f7-408e-8f15-036887e5642c_1728x912.png 848w, https://substackcdn.com/image/fetch/$s_!MhSm!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8965ae9e-a4f7-408e-8f15-036887e5642c_1728x912.png 1272w, https://substackcdn.com/image/fetch/$s_!MhSm!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8965ae9e-a4f7-408e-8f15-036887e5642c_1728x912.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!MhSm!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8965ae9e-a4f7-408e-8f15-036887e5642c_1728x912.png" width="1456" height="768" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/8965ae9e-a4f7-408e-8f15-036887e5642c_1728x912.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:768,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:95476,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/190603006?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8965ae9e-a4f7-408e-8f15-036887e5642c_1728x912.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!MhSm!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8965ae9e-a4f7-408e-8f15-036887e5642c_1728x912.png 424w, https://substackcdn.com/image/fetch/$s_!MhSm!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8965ae9e-a4f7-408e-8f15-036887e5642c_1728x912.png 848w, https://substackcdn.com/image/fetch/$s_!MhSm!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8965ae9e-a4f7-408e-8f15-036887e5642c_1728x912.png 1272w, https://substackcdn.com/image/fetch/$s_!MhSm!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8965ae9e-a4f7-408e-8f15-036887e5642c_1728x912.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><p><em>Figure 4. Why one regex took down ten percent of the internet&#8217;s HTTPS. PCRE&#8217;s backtracking engine offers no complexity guarantee, so on inputs the test author never wrote, the same rule that was fast in the suite explored a quadratic-to-exponential number of paths. RE2&#8217;s DFA does one state transition per character: linear, regardless of pattern shape.</em></p><p>This pathology, <strong>catastrophic backtracking</strong>, is a known property of backtracking engines (<em>PCRE, Perl, Python&#8217;s re, JavaScript&#8217;s RegExp</em>) on inputs they were not tested against. </p><p>Cloudflare&#8217;s Lua WAF used PCRE because <strong>PCRE </strong>ships with Lua, and PCRE has no complexity guarantee: it attempts every backtrack until it matches or exhausts the search. </p><p>The rule had passed the test suite. It had even been deployed in &#8220;<em>simulate</em>&#8221; mode, where the rule runs against real traffic but blocks nothing, explicitly to catch this, but simulate mode still executes the regex on each request and the CPU cost was identical. </p><p>Two factors made it worse: the WAF Managed Rules pipeline bypassed Cloudflare&#8217;s normal staged rollout (<em>a deliberate speed-against-safety trade for emergency patches</em>), and a<strong> CPU-time safeguard</strong> that would have caught a runaway regex had been removed by mistake during an earlier refactoring whose explicit goal was to reduce CPU consumption.</p><p>The<strong> recovery was constrained</strong> by the same property that caused the failure: the engineers who needed the kill switch could not reach the control panel, because it runs behind Cloudflare&#8217;s own Access product, which routes through the edge. </p><p>They used a rarely-exercised bypass, diagnosed it by 14:02, and killed the rulesets at 14:09. Total impact: about <strong>twenty-seven minutes</strong>, during which Discord, Feedly, Coinbase, and a large share of global HTTPS returned 502s, an aggregate impact orders of magnitude larger than the duration suggests, because Cloudflare had, like S3, become public infrastructure.</p><p>The form is worst case from interaction. The system has performance regimes not exercised by any input used in development or testing, exposed only by realistic input at scale. </p><p>The regex did not fail under any tested input; the engine did not fail in any benchmark; the pipeline functioned as designed. The fault was emergent from the joint behaviour of a regex, an engine, an input distribution, and a deployment process, each correct in isolation. </p><p>This is what <strong>Hollnagel </strong>calls the underspecification problem: a component&#8217;s spec covers anticipated inputs, the system at scale sees inputs no one anticipated, and the behaviour under those is emergent from the implementation, not contained in the spec. </p><p><strong>Cloudflare&#8217;s fix</strong>, beyond restoring the safeguard and staging rollouts, was to migrate from PCRE&#8217;s backtracking to an engine based on a deterministic finite automaton, which guarantees linear time regardless of pattern shape because every input character causes exactly one state transition. </p><p>That is the move <strong>Perrow </strong>recommended in 1984: where possible, reduce the interactive complexity rather than defend against its consequences.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Recurrence: Cloudflare again, 18 November 2025</h2><p>If the <strong>three preceding incidents</strong> were merely history, a sceptic could argue the field has since learned its lessons. </p><p>The strongest evidence against that reading is that the same company, having written one of the <strong>most-cited post-mortems</strong> in the industry about its 2019 outage, produced a structurally identical failure six years later, for reasons again unrelated to the specific bug.</p><p>On 18 November 2025 at 11:05 UTC, Cloudflare deployed a correct permissions change to a <strong>ClickHouse cluster</strong>, granting users explicit access to metadata for shard tables in a schema called r0. </p><p>The problem was a buried assumption elsewhere. A query feeding the Bot Management system listed a table&#8217;s columns and had always, by assumption, returned only the default database&#8217;s columns. </p><p>After the change, it also returned the r0 schema&#8217;s columns, roughly doubling the rows. That output fed directly into a &#8220;<em>feature file,</em>&#8221; the configuration the <strong>Bot Management model </strong>consumes to score every request as bot or human.</p><p>The feature file, normally stable around sixty features, more than doubled to over two hundred. It is regenerated every few minutes and propagated globally, by design, so the system can react to new bot behaviour. </p><p>The <strong>core proxy</strong> preallocates memory for these features and enforces a hard limit of two hundred. When the bloated file hit production it exceeded the limit, and the Rust code did not handle the error gracefully. </p><p>It panicked: <code>thread fl2_worker_thread panicked, called unwrap on an Err value</code>. Every request through the Bot Management path returned a 5xx. </p><p>The blast radius again included Cloudflare&#8217;s own products and downstream a large slice of the consumer internet: <em>ChatGPT, X, Spotify, Canva, Discord</em>. <strong>Matthew Prince</strong> called it the worst outage since 2019.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!0Euy!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F652c5550-e5fa-4bda-823c-3dc49600272d_1728x816.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!0Euy!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F652c5550-e5fa-4bda-823c-3dc49600272d_1728x816.png 424w, https://substackcdn.com/image/fetch/$s_!0Euy!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F652c5550-e5fa-4bda-823c-3dc49600272d_1728x816.png 848w, https://substackcdn.com/image/fetch/$s_!0Euy!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F652c5550-e5fa-4bda-823c-3dc49600272d_1728x816.png 1272w, https://substackcdn.com/image/fetch/$s_!0Euy!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F652c5550-e5fa-4bda-823c-3dc49600272d_1728x816.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!0Euy!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F652c5550-e5fa-4bda-823c-3dc49600272d_1728x816.png" width="1456" height="688" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/652c5550-e5fa-4bda-823c-3dc49600272d_1728x816.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:688,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:130263,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/190603006?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F652c5550-e5fa-4bda-823c-3dc49600272d_1728x816.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!0Euy!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F652c5550-e5fa-4bda-823c-3dc49600272d_1728x816.png 424w, https://substackcdn.com/image/fetch/$s_!0Euy!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F652c5550-e5fa-4bda-823c-3dc49600272d_1728x816.png 848w, https://substackcdn.com/image/fetch/$s_!0Euy!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F652c5550-e5fa-4bda-823c-3dc49600272d_1728x816.png 1272w, https://substackcdn.com/image/fetch/$s_!0Euy!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F652c5550-e5fa-4bda-823c-3dc49600272d_1728x816.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>Figure 5, Six years apart, the same shape. A correct local change meets a buried assumption (worst case from interaction); a fast-global-propagation pipeline turns one bad artefact into a worldwide event in minutes (the amplifier); and concentration makes the blast radius the consumer internet.</em></p><p>Read against the three earlier incidents, the recurrence is almost eerie. It is worst case from interaction, in the 2019 sense: a query that <strong>behaved one way for years</strong> behaved differently under a correct change, on an input nobody had specified. </p><p>It is concentration, in the 2017 sense, and it landed only weeks after a major AWS us-east-1 outage on 20 October 2025. And the amplifier is the same amplifier as 2019: the fast global propagation path, the very mechanism that gives the <strong>system its responsiveness</strong>, is what turned one bad artefact into a worldwide event in minutes. </p><p>The deepest point is the one a casual reader misses. Cloudflare did learn the 2019 lesson; they migrated regex engines, added staged rollout, restored safeguards. </p><p>None of it prevented 2025, because the 2019 lesson was about regular expressions and the 2025 failure was about a <strong>database permission</strong>, a buried assumption, a hard-coded limit, and an unwrap that should have been a fallback. The two share no component. They share a structure.</p><blockquote><p><em>The 2019 fix did not prevent 2025, because the two incidents share no component. They share a structure. You cannot patch a structure by patching the part that happened to express it last time.</em></p></blockquote><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>The emergent forms, named</h2><p>The incidents are distinct shapes of emergence. Cascade (2003): tight coupling, a <strong>fault propagating faster</strong> than operators can intervene; the reach is a property of the topology, not of any line or processor. </p><p>Concentration (2017, and again 2025): a structural property nobody designed, accumulated through years of independent decisions; the blast radius of a<strong> foundational service&#8217;s failure</strong> is a property of how many systems converged on it, and it grows silently because no organisation can see across all the adoptions. </p><p>Worst case from interaction (2019): regimes not exercised by the inputs the system was tested against, <strong>exposed by the inputs</strong> it actually sees; the failure is in no component&#8217;s specification because no specification covers all realistic inputs.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!f0jk!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F25ddfc78-3722-4789-8209-cd2afd1e0529_1728x720.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!f0jk!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F25ddfc78-3722-4789-8209-cd2afd1e0529_1728x720.png 424w, https://substackcdn.com/image/fetch/$s_!f0jk!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F25ddfc78-3722-4789-8209-cd2afd1e0529_1728x720.png 848w, https://substackcdn.com/image/fetch/$s_!f0jk!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F25ddfc78-3722-4789-8209-cd2afd1e0529_1728x720.png 1272w, https://substackcdn.com/image/fetch/$s_!f0jk!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F25ddfc78-3722-4789-8209-cd2afd1e0529_1728x720.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!f0jk!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F25ddfc78-3722-4789-8209-cd2afd1e0529_1728x720.png" width="1456" height="607" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/25ddfc78-3722-4789-8209-cd2afd1e0529_1728x720.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:607,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:91232,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/190603006?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F25ddfc78-3722-4789-8209-cd2afd1e0529_1728x720.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!f0jk!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F25ddfc78-3722-4789-8209-cd2afd1e0529_1728x720.png 424w, https://substackcdn.com/image/fetch/$s_!f0jk!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F25ddfc78-3722-4789-8209-cd2afd1e0529_1728x720.png 848w, https://substackcdn.com/image/fetch/$s_!f0jk!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F25ddfc78-3722-4789-8209-cd2afd1e0529_1728x720.png 1272w, https://substackcdn.com/image/fetch/$s_!f0jk!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F25ddfc78-3722-4789-8209-cd2afd1e0529_1728x720.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>Figure 6 &#8212; Three shapes of emergence. In production they compose: concentration creates the conditions for cascade; a worst-case interaction triggers a cascade through tightly-coupled subsystems. November 2025 was all three at once.</em></p><p>The three compose in production. Concentration creates the conditions for cascade: when many systems share fate, a fault in the shared component propagates instantly. </p><p>Worst-case interactions trigger cascades through tightly-coupled subsystems: the regex spike that takes down the edge takes down the dashboard the operators need to<strong> push the kill switch</strong>. </p><p>The deeper claim is that the three failures of Parts I, II, and III are themselves emergent properties of the same kind of system: </p><ol><li><p><em>composition gaps appear at interfaces nobody designed end to end; </em></p></li><li><p><em>observation failures appear because the system&#8217;s behaviour is dimensionally larger than any representation, with the dimensions that matter most during novel failures being exactly the ones the projection discarded; </em></p></li><li><p><em>control failures appear because the loops close on a plant whose dynamics are not contained in any single component.</em></p></li></ol><h2>the projection model</h2><p>It is worth making the claim precise, because precision turns a metaphor into a tool. </p><p>Let S be the system as it actually is: the full state, every dependency, every cached value, every in-flight retry, every input it will ever see. S lives in an enormous, <strong>high-dimensional space</strong>, and no human or dashboard holds it. </p><p>What everyone works with is a representation R, obtained by a projection (call it pi) that maps the system into something small enough to fit in a diagram, a metrics store, or one engineer&#8217;s head.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!F39-!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9bcd7644-15db-4908-a176-15dd2e6cf8e9_1728x864.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!F39-!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9bcd7644-15db-4908-a176-15dd2e6cf8e9_1728x864.png 424w, https://substackcdn.com/image/fetch/$s_!F39-!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9bcd7644-15db-4908-a176-15dd2e6cf8e9_1728x864.png 848w, https://substackcdn.com/image/fetch/$s_!F39-!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9bcd7644-15db-4908-a176-15dd2e6cf8e9_1728x864.png 1272w, https://substackcdn.com/image/fetch/$s_!F39-!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9bcd7644-15db-4908-a176-15dd2e6cf8e9_1728x864.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!F39-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9bcd7644-15db-4908-a176-15dd2e6cf8e9_1728x864.png" width="1456" height="728" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/9bcd7644-15db-4908-a176-15dd2e6cf8e9_1728x864.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:728,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:135043,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/190603006?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9bcd7644-15db-4908-a176-15dd2e6cf8e9_1728x864.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!F39-!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9bcd7644-15db-4908-a176-15dd2e6cf8e9_1728x864.png 424w, https://substackcdn.com/image/fetch/$s_!F39-!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9bcd7644-15db-4908-a176-15dd2e6cf8e9_1728x864.png 848w, https://substackcdn.com/image/fetch/$s_!F39-!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9bcd7644-15db-4908-a176-15dd2e6cf8e9_1728x864.png 1272w, https://substackcdn.com/image/fetch/$s_!F39-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9bcd7644-15db-4908-a176-15dd2e6cf8e9_1728x864.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>Figure 7, The unifying picture. Every artefact an operator touches is a point in R. The system lives in S. In the steady state the two agree, because the system spends most of its time in a regime pi captures. An incident is precisely the event of S moving into the part of itself that pi threw away.</em></p><p><strong>The four failures</strong> are four ways pi betrays you. Composition is a failure at the seams of pi: each subsystem is built against its own local projection of its neighbours, and where two such projections meet, the assumptions need not agree (<em>the 2025 query is a perfect specimen</em>). </p><p><strong>Observation</strong> is the claim that pi loses, preferentially, the dimensions that carry the most signal during a novel failure, because a dashboard is built from the dimensions that mattered in past incidents and a novel incident is by definition one whose decisive dimension was not salient before. </p><p><strong>Control</strong> is the observation that the loop closes on R, not S: it can be perfectly stable on R and diverging on S, and the operator sees stability until the divergence reaches a dimension pi still tracks. </p><p><strong>Emergence</strong> is the statement that the dimension of S vastly exceeds that of any R a human or tool can hold, and that no enrichment of R closes the gap, because every dimension you add to the dashboard expands the interaction surface of the system you are charting. </p><p>That is the formal <strong>residue of Perrow</strong>: adding observation is adding components, adding components is adding interactions, and the system you can fully observe is, for that reason, not the system you have.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>The system has no architect</h2><p>What follows, and what engineers trained on smaller systems underestimate, is that production distributed systems at scale have no architect. This is <strong>not about staffing</strong>. It is structural. </p><p>The system is the accumulated artefact of thousands of independent decisions by hundreds of engineers over years, most of whom have left, none charged with maintaining a coherent end-to-end model. </p><p>The architecture diagrams capture, at best, the model of one engineer, on one day, of the slice they were looking at. The whole system is not in any document because it is not in any individual&#8217;s head.</p><p>This is the <strong>Conway&#8217;s Law observation</strong>, after Mel Conway&#8217;s 1968 paper that organisations design systems mirroring their communication structure. The deeper version is Daniel Dennett&#8217;s phrase: competence without comprehension. </p><p>The system serves traffic, accepts payments, delivers content, without any individual comprehending the full mechanism. The competence is real; the comprehension exists nowhere. There is a parallel from political economy. </p><p><strong>Hayek&#8217;s 1945 paper</strong> argued against central planning on epistemic grounds: the knowledge to coordinate a complex economy does not exist in any single mind, but is distributed across millions of agents holding local, tacit knowledge that cannot be efficiently aggregated. </p><p>The argument transfers directly. Conway, Dennett, and Hayek point at the same fact from three lineages, which suggests it is real and not an artefact of one discipline.</p><p>This is<strong> hard for engineers</strong> to accept in proportion to how good they are. The instinct that produced their career, that one can read the code and understand the system, works up to about the size of a single service team. </p><p>Past that it stops, and the engineer who insists on retaining the model loses the ability to operate the system. The mature posture is not to know the system but to navigate it: <strong>Charity Majors</strong> describes moving from understanding systems to interrogating them; the resilience-engineering school calls it coping with complexity, the operator inside the system rather than above it. </p><p>In a <strong>complicated system</strong> you learn the system and then operate it. In a complex system you operate the system and learn what you can, knowing some of what you learn will be obsolete by the time you have learned it.</p><p>The competent on-call engineer at scale is therefore not the one who knows the most, but the one with the best discipline for forming hypotheses, sizing interventions, observing responses, and updating beliefs. </p><p>That discipline is what decides whether an incident lasts twenty minutes or twenty hours.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Drift into failure</h2><p>There is a phenomenon, named by Sidney Dekker in <em>Drift into Failure</em> (2011) and rooted in <strong>Jens Rasmussen&#8217;s 1997 paper</strong>, that explains how complex systems reach the boundary of safe operation without any single decision being responsible. </p><p>A system operates in a state space bounded by three pressures, each a gradient. Economic pressure pushes toward higher throughput at lower cost. Workload pressure pushes toward simpler procedures and less manual intervention. </p><p>The third is <strong>the boundary</strong> of functionally acceptable performance, the safety boundary, beyond which catastrophic failure is statistically expected.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!zQGk!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0c17537b-5af8-4dc8-8551-6ab27cc844b0_1728x1056.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!zQGk!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0c17537b-5af8-4dc8-8551-6ab27cc844b0_1728x1056.png 424w, https://substackcdn.com/image/fetch/$s_!zQGk!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0c17537b-5af8-4dc8-8551-6ab27cc844b0_1728x1056.png 848w, https://substackcdn.com/image/fetch/$s_!zQGk!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0c17537b-5af8-4dc8-8551-6ab27cc844b0_1728x1056.png 1272w, https://substackcdn.com/image/fetch/$s_!zQGk!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0c17537b-5af8-4dc8-8551-6ab27cc844b0_1728x1056.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!zQGk!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0c17537b-5af8-4dc8-8551-6ab27cc844b0_1728x1056.png" width="1456" height="890" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0c17537b-5af8-4dc8-8551-6ab27cc844b0_1728x1056.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:890,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:164174,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/190603006?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0c17537b-5af8-4dc8-8551-6ab27cc844b0_1728x1056.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!zQGk!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0c17537b-5af8-4dc8-8551-6ab27cc844b0_1728x1056.png 424w, https://substackcdn.com/image/fetch/$s_!zQGk!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0c17537b-5af8-4dc8-8551-6ab27cc844b0_1728x1056.png 848w, https://substackcdn.com/image/fetch/$s_!zQGk!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0c17537b-5af8-4dc8-8551-6ab27cc844b0_1728x1056.png 1272w, https://substackcdn.com/image/fetch/$s_!zQGk!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0c17537b-5af8-4dc8-8551-6ab27cc844b0_1728x1056.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>Figure 8: Rasmussen&#8217;s state space. Two boundaries (economic, workload) are quantified, visible, and rewarded. The third, safety, is unmarked and gives no warning at the crossing, because the crossing is statistical, not deterministic. Each locally-rational optimisation removes slack and nudges the operating point along the cost gradient.</em></p><p>The three pressures are <strong>not balanced</strong>. Economic and workload pressure act continuously and visibly: quantified in dashboards, named in reviews, the subject of every planning cycle. The safety boundary is invisible. </p><p>The system gives no warning at the crossing, because the crossing is statistical: a <strong>system past the boundary </strong>does not fail immediately, it fails with elevated probability per unit time, empirically indistinguishable from operating safely until it fails. </p><p>The result is a <strong>steady migration</strong> toward the safety boundary, driven by the visible pressures, the boundary unobserved until it is crossed.</p><p>Applied to production systems: every optimisation that reduces latency or cost is a step toward the boundary. Every timeout tightened for P99, every retry budget cut, every <em>cache TTL shortened</em>, every connection pool sized closer to peak. </p><p>Each is <strong>locally rational</strong>, and the cumulative effect is to remove slack. Slack absorbs perturbations; a system with no slack is tightly coupled, the condition under which interactive complexity becomes accidents. The drift is not a decision to operate unsafely. </p><p>It is many decisions to operate slightly more efficiently, and at no point did anyone authorise the trajectory. </p><p><strong>Per Bak&#8217;s self-organized criticality (1987) </strong>showed that systems under continuous driving organise themselves to a critical point, where a small perturbation can produce an avalanche of any size, with sizes following a power law (<em>the two-dimensional exponent is commonly quoted near 1.2, though it is non-universal and contested; the heavy tail is robust regardless</em>). </p><p>Carlson and Doyle&#8217;s <strong>Highly Optimized Tolerance (1999)</strong> is more directly applicable: systems explicitly optimised for robustness against expected disturbances become fragile against unexpected ones, producing power-law failures without any external driving. </p><p>They are critical because they were engineered to be, the optimisation having consumed every margin against the perturbations the optimiser did anticipate.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!0lM8!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F07c1b494-bd8f-44c6-a36c-6346fc322a9b_1728x864.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!0lM8!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F07c1b494-bd8f-44c6-a36c-6346fc322a9b_1728x864.png 424w, https://substackcdn.com/image/fetch/$s_!0lM8!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F07c1b494-bd8f-44c6-a36c-6346fc322a9b_1728x864.png 848w, https://substackcdn.com/image/fetch/$s_!0lM8!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F07c1b494-bd8f-44c6-a36c-6346fc322a9b_1728x864.png 1272w, https://substackcdn.com/image/fetch/$s_!0lM8!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F07c1b494-bd8f-44c6-a36c-6346fc322a9b_1728x864.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!0lM8!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F07c1b494-bd8f-44c6-a36c-6346fc322a9b_1728x864.png" width="1456" height="728" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/07c1b494-bd8f-44c6-a36c-6346fc322a9b_1728x864.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:728,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:88371,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/190603006?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F07c1b494-bd8f-44c6-a36c-6346fc322a9b_1728x864.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!0lM8!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F07c1b494-bd8f-44c6-a36c-6346fc322a9b_1728x864.png 424w, https://substackcdn.com/image/fetch/$s_!0lM8!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F07c1b494-bd8f-44c6-a36c-6346fc322a9b_1728x864.png 848w, https://substackcdn.com/image/fetch/$s_!0lM8!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F07c1b494-bd8f-44c6-a36c-6346fc322a9b_1728x864.png 1272w, https://substackcdn.com/image/fetch/$s_!0lM8!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F07c1b494-bd8f-44c6-a36c-6346fc322a9b_1728x864.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>Figure 9. Why &#8220;that could never happen at our scale&#8221; is a category error. Optimised systems produce power-law failure sizes, not the thin-tailed distribution intuition assumes. The catastrophic outage is not an outlier off the curve. It is a point on the curve, in the tail the optimisation built.</em></p><p>The implication is <strong>uncomfortable</strong>. The reliability of a mature system is not a function of how careful its operators are. </p><p>It is a function of how much slack they have preserved against the cost pressure that would remove it. </p><p><strong>A team that fights</strong> to keep slack, holding <em>redundant capacity</em>, leaving timeouts loose, refusing to push retry budgets to the minimum, is doing the most important reliability work, and it is the work least visible in the metrics. </p><p>This is also why incidents cluster: with margin, small perturbations are absorbed; at the boundary, the same perturbations cascade. </p><p>Incident frequency is <strong>perturbation frequency</strong> times the probability the system is currently at the boundary, and that probability rises with drift, and drift is monotonic without deliberate effort against it. </p><p>The system gets less safe over time even as no one decides to make it less safe.</p><blockquote><p><em>Drift is the boundary approaching. Cascade is what happens when the boundary is crossed. The system gets less safe over time even as no one decides to make it less safe.</em></p></blockquote><div><hr></div><h2>What operations looks like under emergence</h2><p>The practices that work here are the <strong>post-Perrow tradition</strong> of resilience and site-reliability engineering, tied together by the recognition that no model is reliable and that operations must proceed by probe-and-respond. </p><p><strong>Chaos engineering</strong>, deliberately injecting failures into production, rests on the premise that behaviour under failure cannot be predicted from the components but only discovered by observation, so the system must be perturbed deliberately, with bounded blast radius and operators present.</p><p>Error budgets make <em>Rasmussen&#8217;s invisible boundary</em> visible in the same units as the economic pressure on the other side: define a <strong>service-level objective</strong>, track consumed unavailability as a budget, and slow releases when it is exhausted. </p><p>The arithmetic is sobering:</p><ul><li><p><em>99.9 percent availability allows about 43.8 minutes of downtime per month.</em></p></li><li><p><em>99.95 percent allows about 21.9 minutes; a single cascade exhausts it.</em></p></li><li><p><em>99.99 percent allows about 4.38 minutes; a 27-minute Cloudflare-style event blows two quarters of budget.</em></p></li><li><p><em>99.999 percent allows about 26 seconds per month, essentially no human-in-the-loop budget at all.</em></p></li></ul><p>The budget is the boundary, denominated in minutes, which is why high-availability targets and aggressive release velocity are in genuine, not rhetorical, tension. </p><p><strong>Blameless post-mortems </strong>are a technical practice as much as a cultural one: the information value of a post-mortem is proportional to the accuracy of the reporting, and accuracy is proportional to the safety the reporter feels. </p><p><strong>Ron Westrum&#8217;s</strong> typology (pathological, bureaucratic, generative) formalises this: only generative cultures, where bad news is welcomed because it is operationally valuable, produce the post-mortems complex systems need. </p><p><strong>Game days </strong>and failure injection are variants of one principle: behaviour under stress cannot be modelled, it must be observed, and the observation must be staged before the real failure, because in the real failure operators cannot slow down to learn.</p><p>What unites these is the <strong>acceptance of irreducibility</strong>: the model is incomplete by construction and is updated by observation under controlled perturbation, not by deduction from specification. Operations becomes an empirical discipline, closer to experimental science than to mechanical engineering. </p><p>The practitioners who absorb this invest in observability (<em>not because dashboards reveal the truth, but because they are the only handle on a system you cannot model</em>), in chaos engineering (because the alternative is discovering breakages during real incidents), and in slack (because they have read enough post-mortems to know which systems fail catastrophically). </p><p>The ones who have <strong>not make the opposite choices</strong> and look reasonable doing it: eliminate slack to cut cost, reduce observability to cut noise, avoid chaos engineering because it occasionally causes outages. </p><p>They discover, eventually, that they have built a system both more efficient and more catastrophic when it fails, in proportion to the savings.</p><h3>Operators as the homeostatic mechanism</h3><p>There is a way to name what operators are, structurally, that systems biology names directly. </p><p>In an organism, homeostasis <strong>maintains internal state </strong>against perturbation: temperature, glucose, pressure, pH within a fraction of a unit despite continuous challenge. </p><p>The mechanisms are <strong>not centralised</strong>; they are distributed across hundreds of overlapping negative-feedback loops, none responsible for the overall stability. </p><p><em>No part of the body</em> is in charge of being alive; being alive is what the parts, in aggregate, are doing.</p><p>The operator function is the homeostatic mechanism in this precise sense. Operators are the <strong>negative-feedback loop </strong>that prevents Rasmussen drift from becoming Rasmussen crossing, the slack reintroduced when cost-cutting removes it elsewhere, the compensating force against every gradient that would otherwise push the system over the boundary. </p><p>The<strong> system is stable </strong>not because the engineering is good but because operators continuously absorb the perturbations the engineering does not address. </p><p>The on-call engineer <strong>reverting a deployment</strong> at 03:47 UTC is not interrupting normal operation. They are participating in it. The reverting is what the system is doing, through them, to keep itself alive.</p><blockquote><p><em>The systems do not work. The operators work, and the systems usually fail only when the operators lose the ability to see, infer, intervene, or comprehend.</em></p></blockquote><p>This is structural and testable. Remove the operators and the system enters its statistically expected failure regime within hours to days. </p><p>With<strong> competent operators present</strong>, it runs, degraded but functional, for years against perturbations that would individually crash it. </p><p>The competence is in the loop, not in the system. Yet organisations reward the visible artefacts (<em>clean architectures, well-factored code, comprehensive tests, capacity plans</em>) and underinvest in the invisible work of compensating, absorbing, and quietly maintaining slack, treating it as a cost centre rather than the production function it is. </p><p>The engineering <strong>does not produce reliability.</strong> Reliability is the output of the engineering acted on by the operators, in a loop the organisation rarely acknowledges. </p><p>This also explains why, beyond a threshold, more automation worsens outcomes.<strong> Lisanne Bainbridge&#8217;s 1983 paper</strong> &#8220;<em>Ironies of Automation</em>&#8221; catalogued it: automating the routine cases removes the operators&#8217; opportunity to maintain situational awareness, the very faculty they need when the automation fails on a case it was not designed for. </p><p>The gains accrue early and visibly; the costs accrue late and on the days that matter most. </p><p>The discipline that follows is to treat operators as the load-bearing element, not the engineering: build the observability they need to interrogate the system,<strong> build the chaos practice </strong>they need to keep their model current, and do not automate them out of the loop; automate the routine inside their loop, so they keep the situational awareness they will need when the automation fails.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>Closing the series</h2><p>Four essays. Four structural failures. </p><p><strong>Composition:</strong> distributed systems fail in the spaces between their components, because no individual designed both sides of the interfaces. </p><p><strong>Observation:</strong> the system you see is a delayed, partial, aggregated projection, and the moments when projection and system disagree are the moments that matter most. </p><p><strong>Control:</strong> the loops you close on the projection cannot stabilise the system, because the plant moves faster than the controller can sample, the actuator is in the blast radius, or the feedback path passes through the failure. </p><p><strong>Emergence:</strong> the system is bigger than any of these admits. It has no architect. It sits at the boundary by construction, because the pressure that makes it efficient drives it there, and its failure modes are properties of the whole that no model in any individual head can predict.</p><p>These are not exceptional. They are the structural conditions under which production systems at scale operate. </p><p>The systems do not work because they were engineered to work; they work because the operators, from the SRE on call at 03:47 UTC to the architect <strong>drawing dependency graphs</strong> in a Confluence document nobody updates, are continuously compensating for the gaps the engineering left open. </p><p>What this means for the engineer is twofold. </p><p><strong>The technical</strong>: the practices that increase survivability (slack, observability, chaos engineering, error budgets, blameless post-mortems, out-of-band control, recovery-dependency mapping, drift monitoring) are expensive, and they are what separates systems that fail gracefully from systems that fail catastrophically. </p><p><strong>The epistemic:</strong> the appropriate posture is structural humility, knowing the system is bigger than the model, the model is incomplete by construction, and any incident might be the one that reveals a regime nobody had characterised.</p><p>The pager will go off again. The <strong>dashboards </strong>will be lying again. The interventions will assume conditions that have ceased to hold. The system will have drifted, since the last incident, slightly closer to the boundary. </p><p>The work is the same work, performed by the same kind of people, against systems that none of them <strong>designed </strong>and none of them fully understand. </p><p>The interesting thing is that this works at all. The accidents are kept rare by humans doing a job whose structural conditions make it nearly impossible, and <em>doing it well enough</em> that the rest of the profession can pretend the systems work on their own. They do not. They never have. </p><p>The <strong>serious engineer&#8217;s contribution</strong>, past a certain seniority, is to understand this and act accordingly: to design systems that respect what the operators have to do, build the tools that make their job possible, and write the practices down so the next generation does not learn them from incidents. </p><p>The job is hard. It is also the <strong>most important job</strong> in the engineering function. The systems do not work without it. They never will.</p><div><hr></div><h2>One more thing&#8230;.</h2><p>I wrote a<strong> deep CUDA guide </strong>from exactly this perspective: not isolated tricks, but how to reason about the GPU as a coupled dynamical system whose performance regimes and failure modes (<em>occupancy collapse, memory-bandwidth thrashing, warp divergence, pipeline stalls, register spilling, bank conflicts, tensor-core underutilisation</em>) are structurally the same kinds of seam failures the four essays of this series have described. </p><p>If the framework resonates, the<strong> guide is where it lands in code.</strong></p><p><strong>[Read the CUDA Guide on Gumroad &#8594; <a href="https://lorenzobrada.gumroad.com/l/cuda_mastery">CUDA Mastery</a>]</strong></p><p><em>This is the final essay in the series How Systems Really Fail. The four parts are best read in order, but each can stand alone.</em></p><p></p>]]></content:encoded></item><item><title><![CDATA[The Llama 3.3 70B Benchmark Problem]]></title><description><![CDATA[What a single H100 SXM5 can and cannot do with Llama 3.3 70B at FP8, a first-principles audit of vLLM, SGLang, and TensorRT-LLM, with the deployment decisions that follow]]></description><link>https://www.thesoftwarefrontier.com/p/the-llama-33-70b-benchmark-problem</link><guid isPermaLink="false">https://www.thesoftwarefrontier.com/p/the-llama-33-70b-benchmark-problem</guid><dc:creator><![CDATA[Lorenzo Bradanini]]></dc:creator><pubDate>Thu, 28 May 2026 22:52:32 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!JCws!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F23b71596-8082-40e1-9897-7ca514525a1b_1122x1402.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!JCws!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F23b71596-8082-40e1-9897-7ca514525a1b_1122x1402.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!JCws!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F23b71596-8082-40e1-9897-7ca514525a1b_1122x1402.png 424w, https://substackcdn.com/image/fetch/$s_!JCws!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F23b71596-8082-40e1-9897-7ca514525a1b_1122x1402.png 848w, https://substackcdn.com/image/fetch/$s_!JCws!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F23b71596-8082-40e1-9897-7ca514525a1b_1122x1402.png 1272w, https://substackcdn.com/image/fetch/$s_!JCws!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F23b71596-8082-40e1-9897-7ca514525a1b_1122x1402.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!JCws!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F23b71596-8082-40e1-9897-7ca514525a1b_1122x1402.png" width="1122" height="1402" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/23b71596-8082-40e1-9897-7ca514525a1b_1122x1402.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1402,&quot;width&quot;:1122,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1920887,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/198570949?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F23b71596-8082-40e1-9897-7ca514525a1b_1122x1402.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!JCws!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F23b71596-8082-40e1-9897-7ca514525a1b_1122x1402.png 424w, https://substackcdn.com/image/fetch/$s_!JCws!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F23b71596-8082-40e1-9897-7ca514525a1b_1122x1402.png 848w, https://substackcdn.com/image/fetch/$s_!JCws!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F23b71596-8082-40e1-9897-7ca514525a1b_1122x1402.png 1272w, https://substackcdn.com/image/fetch/$s_!JCws!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F23b71596-8082-40e1-9897-7ca514525a1b_1122x1402.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>Introduction</h2>
      <p>
          <a href="https://www.thesoftwarefrontier.com/p/the-llama-33-70b-benchmark-problem">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[How Systems Really Fail, Part III]]></title><description><![CDATA[The control problem: why the loop you close cannot stabilise the system you have, why every remediation has a half-life, and what it means to act on a system whose state you cannot see.]]></description><link>https://www.thesoftwarefrontier.com/p/how-systems-really-fail-part-iii</link><guid isPermaLink="false">https://www.thesoftwarefrontier.com/p/how-systems-really-fail-part-iii</guid><dc:creator><![CDATA[Lorenzo Bradanini]]></dc:creator><pubDate>Wed, 27 May 2026 21:21:31 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!HYBy!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff98fccdc-1c43-4e0f-a7b9-4379bd28ab22_1402x1122.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!HYBy!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff98fccdc-1c43-4e0f-a7b9-4379bd28ab22_1402x1122.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!HYBy!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff98fccdc-1c43-4e0f-a7b9-4379bd28ab22_1402x1122.png 424w, https://substackcdn.com/image/fetch/$s_!HYBy!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff98fccdc-1c43-4e0f-a7b9-4379bd28ab22_1402x1122.png 848w, https://substackcdn.com/image/fetch/$s_!HYBy!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff98fccdc-1c43-4e0f-a7b9-4379bd28ab22_1402x1122.png 1272w, https://substackcdn.com/image/fetch/$s_!HYBy!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff98fccdc-1c43-4e0f-a7b9-4379bd28ab22_1402x1122.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!HYBy!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff98fccdc-1c43-4e0f-a7b9-4379bd28ab22_1402x1122.png" width="1402" height="1122" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f98fccdc-1c43-4e0f-a7b9-4379bd28ab22_1402x1122.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1122,&quot;width&quot;:1402,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2286644,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.thesoftwarefrontier.com/i/198570858?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff98fccdc-1c43-4e0f-a7b9-4379bd28ab22_1402x1122.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!HYBy!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff98fccdc-1c43-4e0f-a7b9-4379bd28ab22_1402x1122.png 424w, https://substackcdn.com/image/fetch/$s_!HYBy!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff98fccdc-1c43-4e0f-a7b9-4379bd28ab22_1402x1122.png 848w, https://substackcdn.com/image/fetch/$s_!HYBy!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff98fccdc-1c43-4e0f-a7b9-4379bd28ab22_1402x1122.png 1272w, https://substackcdn.com/image/fetch/$s_!HYBy!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff98fccdc-1c43-4e0f-a7b9-4379bd28ab22_1402x1122.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>Intro</h2><p>The <a href="https://open.substack.com/pub/softwarefrontier/p/how-systems-really-fail-part-i?r=3c7w5a&amp;utm_campaign=post-expanded-share&amp;utm_medium=web">first essay</a> argued that distributed systems fail in the spaces between their components, and that those spaces are structurally opaque. <a href="https://open.substack.com/pub/softwarefrontier/p/how-systems-really-fail-part-ii?r=3c7w5a&amp;utm_campaign=post-expanded-share&amp;utm_medium=web">The second</a> argued that the <strong>system you observe</strong> is not the system that exists, that aggregation destroys signal, and that the operator&#8217;s dashboard is a delayed, partial, instrumented projection of a system that has already moved on.</p><p>This one is about what happens next.</p><p>Once you have accepted that the <strong>system is opaque</strong> and that your view of it is incomplete, you still have to act. The pager has gone off. The error rate has climbed from a green 0.02% to a red 12%. Customers are tweeting. Your manager is on the call. </p><p> The runbook has three pages and none of them describe this. You have ninety seconds before the next escalation tier joins, and you have to decide whether to roll back, fail over, shed load, drain a region, or do nothing and let the system <strong>find its own equilibrium</strong>.</p><p>This essay is about that decision. Not the politics of incident response, not the cultural question of blameless post-mortems, but the structural problem underneath: you are closing a control loop on a system whose state you <strong>cannot fully observe</strong>, whose composition you do not fully control, and whose response to your inputs is, in the regime where you most need to act, nonlinear, delayed, and frequently the opposite of what you expected.</p><p><strong>Classical control theory</strong> has names for all of this. The combination is called control under uncertainty, and the bounds it places on what an operator can achieve are not soft. They are mathematical. </p><p>They are the reason a competent on-call engineer with a complete runbook and a working dashboard can still make an outage worse, not through error, but by executing the <strong>textbook intervention </strong>against a system that has, by the time the intervention lands, already entered a regime where the textbook does not apply.</p><p>Three incidents, mechanically reconstructed: <em>Knight Capital&#8217;s forty-five-minute, four-hundred-and-forty-million-dollar loss in 2012</em>, where a human control loop sampling at the speed of decision could not stabilise a software loop running at the speed of order entry;<em> the Facebook BGP withdrawal of October 2021</em>, where the control plane that needed to repair the network had been routed through the network it had just withdrawn from; and the <em>AWS Kinesis outage of November 2020</em>, where the remediation that would have ended the failure could not proceed because it depended on the very subsystem the failure had taken down.</p><p>The pattern beneath all three is the same. The system entered a regime in which the available control inputs were <strong>either too slow</strong>, structurally unable to reach the failing component, or themselves dependent on the failure being already fixed. </p><p>The operators were <strong>not negligent</strong>. They were operating inside a loop whose closure conditions had been silently violated, and the loop did what control loops do when their closure conditions fail: it stopped controlling.</p><p>The interesting question, again, is why this is structural.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>The control loop, mechanically</h2><p>A control loop is a four-stage cycle: a <em>plant</em> whose state evolves over time, a <em>sensor</em> producing measurements, a <em>controller</em> computing corrective inputs against a setpoint, an <em>actuator</em> applying those inputs. </p><p>The plant&#8217;s new state is measured, and the cycle repeats.</p><p>In a textbook loop, all four stages are coupled tightly enough that the system can be analysed as a single dynamical system. The classical results, like <strong>Nyquist stability, Bode gain and phase margins, Lyapunov functions, Kalman observers,</strong> assume this coupling. </p><p>Given a plant with known dynamics, a sensor with bounded noise, a controller with a known transfer function, and an actuator with bounded authority, they tell you whether the closed loop is stable and how it responds to disturbances.</p><p>A production distributed system violates every one of these assumptions.</p><p>The plant is <strong>not one system</strong>; it is a composition of subsystems each with their own dynamics, coupled through interfaces that hide most of the relevant state. </p><p>The sensor is the <strong>observability pipeline</strong> of Part II, with tens of seconds of phase lag and aggregation that destroys precisely the signal the controller needs. </p><p>The controller is split across at least three actors at different sampling rates: <strong>automated systems</strong> (autoscalers, load balancers, schedulers) running at the speed of metric collection; on-call humans running at the speed of cognition under stress; incident commanders running slower still. </p><p>The actuator is whatever combination of <strong>API calls</strong>, configuration pushes, deployment rollbacks, and SSH sessions the operator can bring to bear, each with its own latency, blast radius, and probability of producing the opposite of what was intended.</p><p>The result is a control loop whose stability margin is set by the slowest, noisiest, most delayed component in the chain. In the steady state, this is fine: the loop has <strong>plenty of margin</strong> and corrections are small. In an incident, the margin evaporates, and the loop&#8217;s behaviour is determined by parts of the system no one had thought to characterise.</p><p>There is a precise name for this in control theory. The formal definitions are Kalman&#8217;s (1960): a system is <em>observable</em> if its internal state can be reconstructed from a<strong> finite history of outputs</strong>; <em>controllable</em> if any state can be reached from any other state in finite time by an admissible input sequence. </p><p>In a healthy production system, both hold approximately. In an incident, one or both fails. Observability fails when the failure mode is invisible to the metrics, as in <strong>Slack&#8217;s autoscaler chasing CPU</strong> while threads waited on a degraded network. Controllability fails when the action that would fix the problem is no longer reachable, as we are about to see in three different forms.</p><p>When both fail at once, the operator is, in the precise technical sense, no longer controlling the system. They are watching it. Interventions may correlate with eventual recovery, but the causal chain from action to outcome has been severed.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>When the human loop is too slow: Knight Capital, 1 August 2012</h2><p>The canonical case for control-loop timescale mismatch is not a distributed-systems outage in the conventional sense. </p><p>It is a <strong>financial one</strong>, and the reason it belongs here is that it isolates, more cleanly than any web-scale incident, what happens when the human control loop runs orders of magnitude slower than the software loop it is supposed to govern.</p><p>Knight Capital Americas was, on the morning of 1 August 2012, the largest U.S. retail market-maker, market-making roughly 17% of NASDAQ-listed and 16% of <strong>NYSE-listed stocks</strong>. Its core function was posting bid and ask quotes on thousands of stocks, capturing the spread, managing inventory.</p><p>The platform that did this was <strong>SMARS</strong>, the Smart Market Access Routing System, running for over a decade.</p><p>On 31 July, NYSE was about to launch the Retail Liquidity Program (RLP). Knight had updated SMARS to support RLP order types. </p><p>The update repurposed a flag that since 2003 had activated a piece of dormant code called <code>Power Peg</code>: an old test algorithm, originally designed to buy high and sell low in order to exercise other trading algorithms in a controlled environment, that Knight had stopped using years earlier but never removed from production. </p><p>In 2005, a separate refactor had moved the <strong>cumulative-quantity counter</strong> (the routine that tracked how many shares of a parent order had been filled and was responsible for stopping further child orders once an order was complete) to an earlier point in the SMARS workflow. </p><p>The move disconnected the counter from <strong>Power Peg</strong>, and Knight never retested Power Peg afterwards. In the new RLP code, the flag&#8217;s meaning was repointed at the RLP handler. The deployment was rolled out manually to eight production servers between 27 July and 1 August. Seven of them received the new code. One did not.</p><p>At 9:30 AM Eastern, the U.S. equities market opened. Parent orders flowed into SMARS to be split into child orders and sent to the exchanges. On the <strong>seven correctly-deployed servers</strong>, child orders were generated, sent to NYSE, and matched against the RLP. </p><p>On the eighth, the repurposed flag was being set on incoming RLP-eligible orders, but the code interpreting it was still the old Power Peg algorithm, now without the cumulative-quantity counter that would have throttled it. </p><p>Each parent order on the <strong>eighth server</strong> generated child orders continuously, with no signal back from the fill-confirmation path to indicate the order had been satisfied.</p><p>Over the next forty-five minutes, the eighth server sent more than four million orders into the market in response to 212 customer orders, executing across 154 symbols and ultimately <em>moving 397 million shares</em>. It bought at the offer and sold at the bid hundreds of times per second. Each round trip lost the spread. </p><p>By <strong>contemporary reporting</strong>, Knight&#8217;s losses accumulated at roughly $10M per minute.</p><p>From the perspective of every component except the broken one, the system was behaving correctly. The exchanges were filling the orders. The risk system was receiving the fills. The position-keeping system was updating. </p><p>Knight&#8217;s internal monitoring had generated 97 emails containing &#8220;<em>Power Peg disabled</em>&#8221; between 8:01 and 8:24 AM EST, before the market opened,  but these were not designed as alerts and no one acted on them.</p><p>What did not happen, for forty-five minutes, was anyone stopping the eighth server.</p><p>The reasons map cleanly onto the structure of the control loop. After roughly <strong>twenty minutes of diagnosis</strong> without documented incident-response procedures, engineers reached the conclusion that the issue lay in the new code and reverted SMARS to its previous version on all eight servers. </p><p>This was the opposite of the correct action: the previous version was the one in which the Power Peg flag still activated the broken Power Peg path. The rollback propagated the failure contained on one server onto all of them.</p><p>Eventually the call was made to halt SMARS entirely. By the time the system was actually stopped, at approximately 10:15 AM, Knight had taken positions of approximately $7.65B (net long $3.5B in 80 stocks, net short $3.15B in 74). </p><p>Once unwound, the realised loss was reported by Knight at ~$440M; the SEC&#8217;s enforcement order placed the figure above $460M. </p><p>The firm did not survive in its prior form. By mid-December, less than five months later, Knight had agreed to a <strong>merger with Getco</strong>; the deal closed in July 2013, and the combined entity (KCG Holdings) was itself acquired by Virtu in 2017.</p><p>The point is the loop. SMARS was running an automated control loop generating orders at machine speed, executing against the market, receiving fills, generating more orders. </p><p>The human control loop above it, monitoring positions, raising alerts, halting the system on threshold breaches, was nominally coupled to SMARS through <strong>dashboards and risk limits</strong>.</p><p>In an incident, they decoupled. The position-monitoring metric had a collection interval on the order of a minute. The decision cycle for incident response was five to ten minutes per hypothesis-test iteration. </p><p>The order-generation cycle was <em>milliseconds</em>. The two loops differed by roughly four orders of magnitude. The faster loop accumulated four hundred million dollars of damage in the time the slower one ran three diagnostic iterations.</p><p>This is the structural form. Nyquist&#8217;s sampling argument applies with full force: a control loop sampling at interval $T$ cannot react to disturbances faster than $2T$. Knight&#8217;s human loop sampled at minutes; the plant disturbance was milliseconds. The loop was, by sampling theory, blind to its own plant. </p><p>The crisis simply could not be controlled by the available control structure, regardless of operator competence.</p><p>The lesson the industry encoded after Knight was not that humans should react faster, they obviously cannot, but that any control loop running at machine speed must have a <strong>kill switch at machine speed</strong>: pre-trade risk checks in the order path, position limits enforced before order submission, circuit breakers triggered on order velocity. </p><p>The slow human loop sits above all of this and decides when to <em>re-enable</em> after the automated kill. It does not, anymore, try to be the kill itself.</p><p>The <strong>principle generalises.</strong> Any system whose failure mode propagates faster than the slowest control loop authorised to stop it is, in the precise technical sense, uncontrollable along that axis. </p><p>The mitigation is not faster humans; it is a fast-enough automated cutoff with a slow-enough human override. </p><p>This is the operational meaning of what Marc Brooker calls <em>autonomic behaviour</em>: the component must be capable of saving itself, on millisecond timescales, against failures the<strong> human loop </strong>is structurally too slow to address.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/p/how-systems-really-fail-part-iii?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/p/how-systems-really-fail-part-iii?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><div><hr></div><h2>When the control plane cannot reach itself: Facebook, 4 October 2021</h2><p>The second structural form of control-loop failure is <em>reachability</em>. The diagnosis is correct, the action is <strong>well-understood</strong>, the human loop is fast enough, but the action cannot be applied because the path from controller to actuator runs through the system that has failed.</p><p><strong>At 15:39 UTC on 4 October 2021</strong>, an engineer at what was then still called Facebook executed a routine maintenance command intended to assess backbone capacity. The command was issued through an audit tool whose job was to reject any change that would take too much of the backbone offline at once. </p><p>A bug in the audit tool failed to catch this one. The command withdrew the <strong>BGP </strong>advertisements for every prefix Facebook announced to the rest of the Internet.</p><p>BGP is the <strong>inter-domain routing protocol </strong>that lets autonomous systems tell the rest of the Internet which prefixes they own and how to reach them. </p><p>When Facebook stopped announcing its prefixes, BGP speakers across the Internet, operating standard route-withdrawal semantics, on the order of seconds, removed Facebook&#8217;s routes from their forwarding tables. </p><p>By <strong>Cloudflare&#8217;s </strong>measurements, public resolvers&#8217; cached records for facebook.com had expired by 15:50 UTC. From the outside, Facebook ceased to exist.</p><p>This is, on its own, a recoverable outage. Re-announcing the prefixes is a single configuration push. The question was whether engineers could get that push to the routers.</p><p>They could not.</p><p>The configuration management system ran on Facebook&#8217;s internal network. Facebook&#8217;s authoritative DNS servers, hosted at smaller facilities, had a safety rule: if they could not reach the main data centres, they treated themselves as unhealthy and withdrew their own BGP advertisements. </p><p>When the backbone went down, every DNS server independently concluded that it was isolated and pulled its routes. <strong>Facebook&#8217;s DNS </strong>therefore disappeared from the public Internet as a second-order consequence of the backbone failure. </p><p>And it disappeared from the inside as well, because the same authoritative DNS resolved the <strong>hostnames </strong>of the internal tools engineers would have used to undo the change.</p><p>It got worse. Many internal tools and services engineers would have used to coordinate the response, parts of <strong>Facebook&#8217;s authentication </strong>and communication infrastructure, also depended on the broken backbone or on the now-unreachable DNS. </p><p>Engineers reportedly could not log into internal tools; conference rooms whose locks were on the <strong>same network</strong> would not open; routine communication channels among responders failed.</p><p>It got worse again. Physical access to the data centres was gated by a card-access system whose backend ran on the same internal network. Engineers attempting to physically enter buildings or reach server cages directly found their badges no longer opened the doors. </p><p>Press reports during the incident described engineers using an industrial angle grinder to cut through a <strong>server-cage bar</strong> at the Santa Clara data centre; Facebook later disputed the specifics, acknowledging only that &#8220;<em>some physical barriers had to be worked around.</em>&#8221; </p><p>Either way, a team had to be physically dispatched to a data centre to restore service.</p><p>The total outage was approximately six hours; BGP advertisements resumed shortly before 21:00 UTC. The technical fix could have been completed in minutes if it had been reachable. </p><p><strong>The duration</strong> was determined by the time required to physically reach a console inside the same dependency loop as the failure, restore enough of the internal network to allow remote actions, and only then perform the fix that was, in itself, trivial.</p><p>This is the second structural form: the action that would resolve the failure lies in the <em>unreachable set</em> induced by the failure itself. </p><p>Control theory has a name for the dual notion, a state that cannot be reached from the current state by any admissible input is <em>uncontrollable</em> from that state,  but the <strong>network-engineering </strong>name is more vivid: the control plane was <em>in-band</em>. </p><p>The configuration changes that would repair the data plane had to travel through the data plane.</p><p>The principle Facebook subsequently invested in is <em>out-of-band control</em>. The control plane must reach its actuators by a path that does not depend on the system being controlled. </p><p>In dependency-graph terms, the directed graph of &#8220;<em>X depends on Y to function</em>&#8221; must contain no cycle that passes through the control surfaces of the production system. </p><p>If it does, there exists a failure mode in which those surfaces are no longer accessible, and the system can only be recovered by an<strong> out-of-band action: </strong>physical access, a separate management network, or a kept-current break-glass procedure that shares no infrastructure with normal operations.</p><p>The cost of maintaining true out-of-band control is non-trivial: a second network, separately operated, credentialed, monitored, exercised. </p><p>The<strong> path of least resistance</strong> is always to let the control plane drift back in-band, because in-band is cheaper, easier to operate, and works fine until the day it does not. </p><p>There is a related principle from safety-critical systems, sometimes called <em>recovery independence</em>, that any component whose failure can render the system inoperable must have a recovery path that does not require that component to be operating. </p><p><strong>NASA flight rules</strong> have a version of this. Nuclear plant operating procedures have a version. Most production software systems do not.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share The Software Frontier&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share The Software Frontier</span></a></p><div><hr></div><h2>When recovery depends on the failure being fixed: AWS Kinesis, 25 November 2020</h2><p>The third structural form is the case where the path to recovery passes through the failure itself.</p><p>On 25 November 2020, the Wednesday before American Thanksgiving,  AWS engineers added capacity to the <strong>Amazon Kinesis Data Streams</strong> front-end fleet in us-east-1 between 02:44 and 03:47 PST. </p><p>The first customer-impacting alarms fired at approximately 05:15 PST, with Kinesis error rates climbing through the morning. Kinesis is AWS&#8217;s high-throughput event ingestion service; it underpins <strong>CloudWatch metrics,</strong> AWS Lambda&#8217;s logging path, Cognito&#8217;s analytics path, and a large catalogue of downstream services. </p><p>When Kinesis became unhealthy in us-east-1, a substantial fraction of AWS itself became unhealthy with it.</p><p>The Kinesis front-end fleet handles authentication, throttling, and request routing to the appropriate back-end clusters that own the actual stream shards. </p><p>Each front-end server maintains in memory a <em>shard-map</em>: a cache containing membership data and shard ownership for the back-end clusters. To populate this cache, each<strong> front-end server</strong> creates an OS thread per peer in the front-end fleet, and exchanges shard information over those threads. </p><p>As AWS noted in its post-mortem, fully learning about a newly added fleet member can take up to an hour. The new capacity pushed the per-server thread count past a <strong>configured OS limit</strong>. When this limit was reached, front-end servers could not create the additional threads needed to complete the shard-map cache. </p><p>Cache construction failed, leaving servers with &#8220;<em>useless shard-maps</em>&#8221; (AWS&#8217;s phrase) that prevented them from routing requests to the correct back-end clusters. <strong>Errors began propagating</strong> to downstream callers, and the failure spread across the fleet as more servers crossed the threshold.</p><p>The remediation was familiar: stop the scaling, remove the additional capacity, and restart the fleet. The constraint was that on coming back up, each front-end server had to rebuild its shard-map by communicating with every other<strong> front-end server,</strong> and the resources needed to populate the cache competed with the resources needed to serve requests. </p><p>AWS could only bring servers back in small groups, a few hundred per hour, verifying stability between batches. The first servers re-entered traffic at 10:07 AM PST; Kinesis fully returned to normal at 10:23 PM PST: roughly 17 hours after the first alarms.</p><p>The duration of the <strong>Kinesis </strong>impairment is not the most interesting part. The most interesting part is what was failing while Kinesis was failing.</p><p><strong>CloudWatch</strong> ingested metric data via Kinesis. With Kinesis impaired, its ability to ingest fresh metrics was degraded, so the dashboards customers and AWS engineers used to monitor their systems went dark or stale at exactly the moment they were needed. </p><p>Lambda invocations require publishing metric data to CloudWatch as part of the invocation; as CloudWatch metrics degraded, Lambda&#8217;s local metric agents exhausted their buffers and invocations began to fail. </p><p>Cognito uses Kinesis Data Streams to collect and analyze API access patterns; the path is documented as best-effort, with web servers buffering locally, but as the impairment dragged on, <strong>Cognito web servers </strong>exhausted those buffers and customer authentication started failing. </p><p>AWS&#8217;s own Service Health Dashboard, which would normally have communicated the outage to customers, was itself impaired in its ability to post updates.</p><p>This is the structural form: the path to recovery passed through the failure. The on-call engineers responding to the <strong>Kinesis outage</strong> needed monitoring to verify their interventions were working, and the monitoring depended on Kinesis. </p><p>Customers needed status updates to understand what was happening, and the status page depended on Kinesis. In <strong>dependency-graph terms</strong>, the <em>recovery dependency graph</em> contained a back-edge: a cycle in which X depends on Y to recover, and Y depends on X to function.</p><p>AWS&#8217;s documented mitigations: moving to larger CPU and memory servers (<em>so the fleet needs fewer machines and therefore fewer per-server threads</em>), accelerating the cellularisation of the front-end fleet so any single instance no longer needs a <strong>thread per peer </strong>across the whole fleet, and separating large internal consumers like CloudWatch onto their own partitioned front-end fleets. </p><p>The architectural changes are the substantive ones. Larger servers patch the specific trigger; cellularisation and partitioning defend against the general class of failure in which fleet-wide state synchronisation scales worse than the fleet itself.</p><p>The general principle, restated for the third time:</p><blockquote><p><em>The path from a failed state back to a healthy state must not depend on the failed component being healthy.</em></p></blockquote><p>This is <strong>harder to enforce</strong> than it sounds. Almost every production system has some path of this shape, somewhere. The deployment system that pushes the fix often depends on the very services the fix is repairing. The monitoring that verifies the fix worked depends on the metrics pipeline. </p><p>The communication tools the on-call team uses to coordinate the response depend on internal services that may be in the blast radius of the failure. The discipline is to <strong>identify these back-edges</strong> and either break them or document the manual workaround.</p><p>The <em>recovery dependency graph</em> is distinct from the runtime dependency graph, and almost always more pessimistic. </p><p>A system can have a clean runtime graph and a <strong>deeply cyclic recovery graph</strong>, because the runtime graph captures what depends on what during normal operation, where slow paths, retry loops, fallback caches, and degraded-mode handoffs are all tolerable, while the recovery graph captures what depends on what when something is broken. </p><p>In recovery, those<strong> tolerances vanish</strong>. The path from failed to healthy must be fast and reliable, and a cycle in that path means it may not exist at all.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>The three failure modes, named</h2><p>The three incidents above are not three instances of the same failure. They are three different ways the control loop loses authority over the plant, each with a precise structural cause.</p><p><strong>Knight Capital is </strong><em><strong>timescale decoupling</strong></em><strong>.</strong> The plant ran at milliseconds; the controller ran at minutes. By the Nyquist sampling argument, the controller could not, in principle, react to disturbances at the plant&#8217;s natural frequency. The control structure was, mathematically, the wrong shape for the plant.</p><p><strong>Facebook BGP is </strong><em><strong>unreachable actuator</strong></em><strong>.</strong> The diagnosis was correct and the intervention was correct, but the actuator was in the failure&#8217;s blast radius. There was no admissible input from the controller&#8217;s current state to the recovery state, because the path between them passed through the failed component. The recovery state was, in the formal sense, <em>unreachable</em> from where the operators were standing.</p><p><strong>Kinesis is </strong><em><strong>recovery dependency cycle</strong></em><strong>.</strong> The controller could reach the actuator and the actuator could apply the intervention. But the verification path, knowing whether the intervention had worked, passed through the failure itself. The loop could be opened by the operators but not closed by feedback, which forced recovery to proceed at the speed of careful, blind, incremental probing.</p><p>The three failure modes compose. In a sufficiently bad incident, all three are active simultaneously: the <strong>plant </strong>moves <strong>faster than the controller</strong> can sample, the actuator is partially unreachable, and the feedback path that does reach the controller is itself degraded. </p><p>The control system has lost authority along three axes at once, and the operator is acting through the gaps.</p><p>The deeper claim, the one <strong>Parts I and II </strong>have been building toward, is that these regimes are not exceptional. They are what production distributed systems enter during incidents, by construction of how those systems are composed and observed.</p><p>The healthy operating envelope is the region in which the control loop has <strong>enough bandwidth</strong>, enough reach, and enough feedback to keep the plant on a setpoint. </p><p>Outside that envelope, one or more of those conditions fails, and the operator is no longer controlling the system, they are nudging it and waiting for the system either to find its own equilibrium or to deteriorate further.</p><p>The job of the operator in this regime is <strong>not to control</strong>. It is to <em>survive</em>: keep blast radius bounded, avoid actions that worsen the failure, preserve the option of recovery, and wait for conditions in which control becomes possible again. </p><p>The aviation discipline of <em>aviate, navigate, communicate</em> captures it: maintain altitude first, <strong>locate yourself second</strong>, talk to people third. Reversing that order is how you crash.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share The Software Frontier&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share The Software Frontier</span></a></p><div><hr></div><h2>The OODA loop, and why it fails</h2><p>The framework that incident responders most commonly use, often without naming it, is the OODA loop: <strong>Observe, Orient, Decide, Act</strong>. </p><p>The name and the framework are due to John Boyd, a U.S. Air Force colonel and fighter pilot who began formulating the ideas in the 1950s and 60s as an instructor at the <em>USAF Fighter Weapons School</em> and refined them through a series of briefings in the 1970s and 1980s, most prominently <em>Patterns of Conflict</em> (1986). </p><p>The central claim, which <strong>Boyd&#8217;s briefings</strong> emphasised, is that in adversarial situations, the actor whose decision cycle runs faster than the adversary&#8217;s wins, because they are operating inside the adversary&#8217;s decision cycle: by the time the slower actor has decided what to do, the faster actor has already changed the situation, invalidating the slower actor&#8217;s decision.</p><p>The framework migrated from military doctrine into business strategy, into emergency response, and eventually into software incident response, where it is used to describe the cycle an <strong>on-call engineer</strong> runs during an outage: observe the system, orient the observation against a model, decide on an intervention, act, and observe the result.</p><p>The framework is useful. It is also, in the production-systems context, frequently misapplied, and the way it is misapplied tells you something about why incidents go badly.</p><p>The original Boyd argument applies when the situation is adversarial and roughly symmetric: two actors with<strong> comparable OODA speeds,</strong> where being faster gives you the advantage. The situation an on-call engineer faces is not adversarial in this sense. There is no opponent making decisions. </p><p>The system is not trying to beat them. It is, instead, evolving according to dynamics (<em>autoscaler decisions, retry storms, cache decays, queue accumulations</em>) that have their own timescales, and those timescales are not necessarily compatible with the engineer&#8217;s OODA loop at all.</p><p>In an incident, three OODA loops are running simultaneously, and they do not all have the same period. <strong>The engineer&#8217;s loop</strong> runs at the speed of human cognition under stress: roughly thirty seconds to a few minutes per cycle, slower if multiple humans need to confer. </p><p>The automated control loop (<em>autoscalers, load balancers, schedulers</em>) runs at the speed of metric collection, typically seconds to tens of seconds. The <strong>plant&#8217;s own dynamics</strong>: connection pool exhaustion, queue overflow, cache regeneration; they all run at whatever rate the underlying physics dictate, which can be anywhere from milliseconds (<em>TCP retransmits, GC pauses</em>) to many minutes (<em>cache warming, replication catch-up</em>).</p><p>In Knight Capital&#8217;s case, the engineer&#8217;s loop ran at minutes against a plant operating at milliseconds. In Boyd&#8217;s terms, the plant was hopelessly <em>inside</em> the <strong>engineer&#8217;s decision cycle</strong>: every observation the engineers made was already obsolete by the time they oriented to it, and every action they took landed against a system that had moved orders of magnitude further during the action&#8217;s flight time.</p><p>The reverse case is also possible and is, in some ways, more insidious. An <strong>engineer&#8217;s OODA loop</strong> running <em>faster</em> than the plant&#8217;s natural recovery dynamics produces a different pathology: the engineer observes that the intervention has not yet worked, orients to that as a failure, decides on a new intervention, and acts, before the original intervention has had time to take effect. </p><p>The result is a stack of interventions in flight against a system that is already converging from the first one, with the later interventions arriving as disturbances against the recovery.</p><p>This is the operational form of what control theorists call <em>over-control</em>. The classical example is the shower with a <em>slow-responding mixer</em>: the user adjusts hot, observes no change, adjusts more, observes no change, then receives the full delayed effect of both adjustments and is scalded. </p><p>The loop is too fast for the plant. The mitigation, both in the shower and in production systems, is to wait, to <strong>extend the OODA cycle</strong> until it matches the plant&#8217;s natural response time, even though waiting feels, in the moment, like inaction.</p><p>The discipline this requires is hard. An on-call engineer under pressure does not feel that waiting is the correct action. The reflex, encouraged by the culture of incident response and reinforced by the stress of an outage in progress, is to act.<strong> To do something</strong>. To try the next thing on the runbook. </p><p>The structural argument against this reflex is not that the engineer is wrong to want to help; it is that, in a system with delayed feedback and partial observability, the cost of acting too quickly can exceed the cost of waiting one more OODA cycle for the previous action to land.</p><p>The principle from control theory is: <em>the loop&#8217;s cycle time should be at least the plant&#8217;s settling time</em>. If you act faster than the plant settles, you are stacking interventions against a<strong> system that has not yet responded</strong> to the previous one, and the resulting trajectory is not the sum of the interventions&#8217; intended effects. </p><p>It is the response of a non-linear system to a sequence of disturbances, which is, almost always, worse than the response to any single intervention applied alone.</p><p>The mature on-call discipline, encoded in the better runbooks and the better incident command training, is to <em>act, then wait the settling time, then observe, then decide whether to act again</em>. The settling time is the half-life of the previous action, and <strong>it is system-specific</strong>: deployment rollbacks settle in minutes; cache invalidations settle in seconds; configuration pushes can take longer than either, depending on the propagation path. </p><p>Knowing the settling time of every action available during an incident is a kind of operational knowledge that does not usually live in the runbook, because it <strong>does not exist outside the system</strong> being operated. It lives in the heads of the engineers who have been on-call for that system long enough to have learned it from incidents.</p><p>This is, partly, why senior on-call engineers are so much more effective in incidents than junior ones, even when the junior engineers have the same runbook access and the same training. The<em> senior engineers </em>have an internal model of the plant&#8217;s settling times for each available action, and they pace the OODA loop to match. </p><p>The junior engineers, lacking the model, run the OODA loop at the <strong>natural speed of human cognition</strong> under stress, which is too fast for most production systems and produces the over-control pathology.</p><div class="directMessage button" data-attrs="{&quot;userId&quot;:201922174,&quot;userName&quot;:&quot;Lorenzo Bradanini&quot;,&quot;canDm&quot;:null,&quot;dmUpgradeOptions&quot;:null,&quot;isEditorNode&quot;:true}" data-component-name="DirectMessageToDOM"></div><div><hr></div><h2>The illusion of the runbook</h2><p>Every mature on-call function has runbooks. They are written after incidents, <em>refined over months</em> and years, kept in wikis and indexed by alert. </p><p>The good ones describe, for each known failure mode, the diagnostic steps to confirm it, the intervention to apply, and the verification to perform. The discipline of writing them is one of the visible practices of a mature operations culture.</p><p>The runbook is also, for a specific category of incident, a trap.</p><p>The trap has two forms. The first is that runbooks are written against past incidents. <em>They encode the diagnoses</em> that worked, the interventions that helped, the verifications that confirmed recovery. </p><p>They are, in the language of Part II, an <strong>artefact of monitoring </strong>rather than observability: comprehensive against known failure modes, useless against novel ones. The first time the runbook fails to describe the incident in front of you is the moment you discover this, and the discovery is rarely well-timed.</p><p>The second form is subtler. The runbook describes interventions and assumes those interventions will work the way they did last time. But the system has been changing since the runbook was written. </p><p>The intervention that worked last quarter may now have different downstream effects, because the components it touches have been migrated, the <strong>dependencies have shifted</strong>, the version of the underlying service has rolled forward. </p><p>The runbook, in effect, captures the system as it was; the engineer is operating the system as it is. The two diverge silently, and the divergence is invisible until the runbook produces an action whose effect surprises everyone.</p><p>There is a <strong>specific failure pattern </strong>this produces, common enough to have its own folklore: the engineer follows the runbook, the runbook prescribes restarting the X service, the X service restarts, and the system gets worse rather than better. </p><p>The reason, on inspection, is usually that the X service has acquired new dependencies since the runbook was written, and restarting it now drops more state than it used to, with consequences the runbook author did not anticipate. </p><p>The runbook is not wrong, in the sense that the steps it describes are the steps that worked. The <strong>runbook is out of date</strong>, in the sense that the system those steps were designed for is no longer the system the engineer is operating.</p><p>The mitigation is not to delete runbooks. They are too useful to discard, and the failure modes they handle correctly are far more common than the ones they handle wrongly. </p><p>The mitigation is to use runbooks as <em>hypotheses</em>, not as <em>procedures</em>: the runbook describes what worked last time, which is evidence about what might work this time, but the <strong>engineer must independently verify</strong>, in real time, that the conditions the runbook assumes still hold. The runbook says &#8220;<em>restart X</em>&#8221;; the engineer asks <em>&#8220;is restarting X still safe given the system as it currently is?</em>&#8221; before doing it.</p><p>This is hard to do under the pressure of an active incident, and it is one of the practical reasons that senior on-call engineers are <strong>more cautious </strong>about runbook execution than junior engineers, not less. The juniors, trusting the document, execute the steps. </p><p>The seniors, knowing how the document was written and how the system has changed since, pause to verify before each step. The juniors finish the runbook faster. The seniors finish the incident faster.</p><p>A related, <strong>sharper principle</strong>: the value of a runbook decays with the rate of system change. In a system that does not change, a runbook is durable: the same steps work indefinitely. In a system that changes weekly, with new deployments, new dependencies, new configurations, the half-life of a runbook is on the order of months (at best). </p><p><strong>A two-year-old runbook</strong> in a fast-moving system is closer to historical fiction than to operational guidance. It still has value, but the value is in the model it encodes of how someone once thought about the system, not in the steps it prescribes.</p><div><hr></div><h2>The blast radius principle</h2><p>The third structural property of action in distributed systems, after timescale and reachability, is <em>blast radius</em>: the set of components affected by a given intervention, and the bound on the damage if the intervention is wrong.</p><p>Every available action has a <strong>blast radius</strong>. Restarting a single instance has a small radius, that instance, briefly, plus whatever state it held. Rolling back a deployment has a larger radius, the entire fleet running that deployment, plus the load that flips back to the previous version, plus the downstream effects of running the older code. </p><p>Failing over a region has a radius that may include every customer routed to that region, every dependent service that has to repoint, and every cache that has to be rebuilt. <strong>Dropping a load balancer</strong> has a radius that can encompass the entire service.</p><p>The blast radius principle says: <em>act with the minimum blast radius that can plausibly resolve the failure</em>. If the failure is in one instance, restart the instance, not the fleet. If the failure is in one region, fail over that region, not the global topology. The reason is not just damage control. It is information. </p><p>A <strong>small-radius action</strong> either resolves the failure or does not, and either outcome is informative: it succeeded, which suggests the failure was localised; or it did not, which rules out a hypothesis and constrains the next action.</p><p>A large-radius action, by contrast, may resolve the failure without telling you which part of the action was responsible. If you fail over a region and <strong>the symptoms clear,</strong> you do not know whether the failure was in the region you abandoned, in the load on the region you moved to, or in the path between them. </p><p>You have ended the incident without learning anything that would prevent the next one. The blast radius was excessive for the diagnostic value returned.</p><p>The principle generalises into a hierarchy of interventions, ordered roughly by radius:</p><p>The smallest interventions are read-only: looking at logs, querying metrics, <strong>running diagnostic commands</strong>. These have effectively zero blast radius and are always safe to do first. Most outages benefit from more reading than the operators in the moment feel they have time for.</p><p>The next tier is single-instance: restarting one process, draining one node, removing one server from rotation. These have blast radius limited to the instance, and they are <strong>reversible within seconds</strong>. They are the appropriate first active intervention for almost any failure that has a candidate localised cause.</p><p>The next tier is<strong> service-level:</strong> rolling restart of a fleet, configuration push, deployment rollback. These have blast radius across the service, take longer to apply, and are harder to reverse. They are appropriate when <em>single-instance interventions</em> have ruled out localised causes, or when the failure is observed broadly enough that localising it is itself wasting time.</p><p>The largest interventions are infrastructure-level: regional failover, load shedding at the edge, traffic routing changes, emergency capacity additions. These have <strong>very large blast radii</strong>, and their consequences are often partly unobservable until well after the action. </p><p>They are appropriate only when smaller interventions have failed or are known to be insufficient, and they should be made with the explicit understanding that the system after the intervention will be a different system from the one before, with new failure modes that no one has yet characterised.</p><p>The principle, in compressed form: <em>the appropriate blast radius scales with the certainty of the diagnosis</em>. When you are confident in the cause, you can use a targeted, low-radius intervention. </p><p>When you are uncertain, you have two choices: invest more time in diagnosis to raise the certainty, or accept the higher blast radius of a less-targeted intervention. There is <strong>no third option</strong>. </p><p>Acting with high blast radius on low certainty is how outages turn into multi-region cascades.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>The conservation of risk</h2><p>There is a principle, with versions in the safety-engineering tradition, that goes roughly: <em>the total risk in a complex system tends to be conserved; safety interventions move risk around as much as they reduce it</em>. The folk version is &#8220;what gets safer somewhere gets more dangerous somewhere else.&#8221; </p><p>The more precise versions belong to the risk-homeostasis and risk-compensation literature, most prominently <strong>Gerald Wilde&#8217;s</strong> <em>Theory of Risk Homeostasis</em> (1982), with related work by John Adams, which argues that visible safety improvements are partly absorbed by behavioural adjustments, with the residual risk migrating elsewhere in the system rather than disappearing.</p><p>This applies to incident response with peculiar force. Each intervention an operator makes during an incident is a <strong>safety action</strong>: an attempt to reduce immediate risk. </p><ul><li><p>The intervention, almost always, succeeds at reducing the visible component of the failure. </p></li><li><p>The error rate drops. </p></li><li><p>The dashboards green up. </p></li><li><p>The customers stop tweeting. </p></li><li><p>The intervention has, by every visible measure, worked.</p></li></ul><p>What is harder to see, and what is rarely measured in real time, is what the intervention did to the invisible component. </p><p><strong>Rolling back a deployment</strong> ends the immediate incident but leaves the older code running, which may have its own known issues that the deployment was meant to fix. </p><p>Failing over a region resolves the local failure but loads the target region beyond its tested capacity, potentially priming a second incident. </p><p><strong>Adding capacity </strong>to a saturated service unblocks the immediate queue but increases the surface area for the failure mode that caused the saturation, if the underlying cause has not been addressed.</p><p>The pattern is not unique to software. Aviation has a long literature on accident patterns that begin with a successful response to a minor problem and end with a major accident triggered by the response itself. <strong>Healthcare</strong> has the same literature. </p><p>The general structural shape is: the intervention that resolves Failure A creates the conditions for Failure B, which is rarer, less familiar, and less recoverable. </p><p>The system has been moved from a known failure regime into an unknown one, and the <strong>unknown regime</strong> has its own failures that the operators have not yet learned to recognise.</p><p>The discipline this asks of an on-call function is, again, hard to maintain under pressure. It is the discipline of asking, after every intervention: <em>what did this just change about the system, and what new failure modes have I introduced?</em> </p><p>It is the discipline of treating recovery as a state that itself needs to be monitored, because the <em>post-recovery system</em> is not the same system as the pre-incident system, and its failure characteristics are not yet known.</p><p>In practice, this means that the <em>end</em> of an incident is not when the symptoms clear. The end of an incident is when the system has been observed in its <strong>new configuration</strong> long enough to be confident that the interventions did not introduce a worse failure than the one they resolved. </p><p>This is usually hours, sometimes days, longer than the time the dashboards take to go green. The on-call function that closes incidents at the moment of <strong>symptom resolution</strong> is, structurally, accepting an invisible risk in exchange for ending the call earlier. </p><p>The function that watches the recovered system through at least one full traffic cycle is paying a cost in operator time for the option of catching the second-order failure before it becomes a second incident.</p><div><hr></div><h2>The operator as Bayesian, under pressure</h2><p>Underneath all of the above is a single epistemic structure. </p><p>The operator, during an incident, is performing inference: from observed symptoms, against a prior model of the system, to a posterior estimate of what is wrong and what to do about it. </p><p>The <strong>framework is Bayesian </strong>whether the operator names it that way or not, and the failure modes are the failure modes of Bayesian inference under conditions hostile to it.</p><p>The prior model, what the operator believes the system is, before any incident-specific evidence, is the simulation <strong>Richard Cook </strong>described and that <em>Part II</em> referenced. </p><p>The likelihood, what the operator expects to observe under each candidate hypothesis, is implicit, drawn from training and prior incidents.</p><p>The posterior, what the operator believes after seeing the symptoms, is the <strong>working diagnosis</strong>. The decision follows from the posterior and from the operator&#8217;s loss function over possible outcomes.</p><p>Each step has its failure mode. The prior can be wrong, as in the Cloudflare oscillation in Part I, where the prior placed most of the probability mass on <strong>external attack</strong> and left almost none for gradual rollout against a regeneration cycle. </p><p>The likelihood can be wrong, as in the Slack autoscaler in Part II, where the prior model said high load implies high CPU, and the actual system was producing <strong>high load with low CPU</strong> because the threads were waiting on a degraded network. </p><p>The posterior, conditioned on a wrong likelihood, will be wrong in the same direction.</p><p>The decision, conditioned on a wrong posterior, will be wrong. And,<em> this is the cruellest part</em>, the operator will get feedback from the decision, observe its outcome, and update the model. </p><p>If the <strong>wrong decision happened</strong>, by chance or by partial compensation from other parts of the system, to be followed by recovery, the operator will incorporate that outcome as evidence that the wrong model was right.</p><p> The next time a similar incident occurs, the operator will reach for the same wrong intervention, with higher confidence, because last time it appeared to work.</p><p>This is the <em>correlated false positive</em> problem in operational learning. The operator&#8217;s model is updated by outcomes that are partially decoupled from interventions, but the<strong> decoupling </strong>is not visible to the operator. Recoveries that happen for reasons unrelated to the intervention reinforce belief in the intervention. </p><p>The model drifts, not toward the system&#8217;s actual dynamics, but toward whatever pattern of intervention-and-recovery the operator has happened to experience.</p><p>The mitigation requires explicit discipline, and most on-call functions do not maintain it. The discipline is to ask, after every incident: <em>did the intervention actually cause the recovery, or did the system recover for some other reason that I happened to be present for?</em> </p><p>The blameless <strong>post-mortem culture</strong> is, partly, an attempt to create the conditions in which this question can be honestly asked. It often is not.</p><p>There is a deeper version of the same problem, which is that the operator&#8217;s model is <strong>updated by outcomes</strong> the operator can observe, and the outcomes the operator can observe are filtered by the same observability stack whose limits Part II discussed. </p><p>The model drifts toward whatever the dashboards can see. Failure modes invisible to the dashboards remain invisible to the model, and the operator becomes progressively <strong>more confident</strong> in a model that captures only the observable subspace of the system&#8217;s actual behaviour. </p><p>The model is calibrated, but on a projection that has discarded the dimensions that matter most during novel failures.</p><p>This is, finally, why the systems that recover quickly from incidents are not the systems with the best runbooks or the <strong>best automation.</strong> </p><p>They are the systems whose operators have maintained an active distinction between the model and the system:<em> </em></p><blockquote><p><em>who treat their model as a hypothesis under continuous test</em></p><p><em>who notice when the system surprises them and update accordingly</em></p><p><em>and who are willing, in the middle of an incident</em>, </p></blockquote><p>to admit that the model they have been operating under for years may not apply to the situation in front of them.</p><p>The discipline is <strong>epistemic humility</strong> under pressure. It is the rarest thing in operations, and it is the one that compounds.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesoftwarefrontier.com/p/how-systems-really-fail-part-iii/comments&quot;,&quot;text&quot;:&quot;Leave a comment&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesoftwarefrontier.com/p/how-systems-really-fail-part-iii/comments"><span>Leave a comment</span></a></p><div><hr></div><h2>The discipline of acting under uncertainty</h2><p>The compressed form of this essay, the operational counterpart to Part I&#8217;s <em>what does this depend on that I cannot see</em>, and Part II&#8217;s <em>what is this metric not telling me</em>, is also a question, asked of every intervention before it is applied:</p><p><em>What does this assume about the system, and what happens if that assumption is wrong?</em></p><p>Every action has assumed conditions. Restarting an instance assumes the instance is the<strong> source of the failure</strong>. Rolling back a deployment assumes the previous version is healthier than the current one. </p><p>Failing over assumes the target region can absorb the load. Adding capacity assumes the <strong>bottleneck is capacity</strong>. Each assumption is, before the action, a hypothesis. Each becomes, after the action, a commitment the operator has to live with.</p><p>The discipline is to know, for every available intervention, what assumption it embeds, and to verify that assumption, before applying the intervention. </p><p>The verification can be imperfect; perfect verification is incompatible with the timescale of an outage. </p><p>But the question must be asked, because the alternative is acting on a hypothesis the operator has not consciously formed, and being surprised by the system&#8217;s response in a direction the operator was not prepared for.</p><p>The questions to ask, before any <strong>non-trivial intervention:</strong></p><ul><li><p>What does this action assume about the system?</p></li><li><p>What is the blast radius if the assumption is wrong?</p></li><li><p>What is the smallest intervention that would resolve this if my diagnosis is correct?</p></li><li><p>How will I know whether the action worked, and how long do I need to wait before that signal is reliable?</p></li></ul><p>If this action fails, what state does it leave the system in, and is the next action still reachable from there?</p><p>The discipline is <strong>not to avoid acting</strong>. It is to size the intervention to the certainty of the diagnosis, allow the system time to settle, verify the assumption before escalating, and preserve the possibility of recovery after every step.</p><p>This is what separates<strong> short incidents </strong>from catastrophic ones. Not faster reactions, better dashboards, or thicker runbooks, but the ability to treat every intervention as a probe: an action that changes the system while also revealing something about it.</p><p>During an incident, the system is operating in a regime nobody has fully characterised. The <strong>operator&#8217;s job</strong> is not to force it back through sheer intervention. The job is to keep the system in a state where understanding is still possible, learn from each action, and navigate toward a recoverable regime without collapsing the remaining options.</p><p>This is on-call: control under uncertainty.</p><p>The loops are delayed. The observations are partial. The interventions have non-linear effects. The system being repaired is one no single engineer fully understands.</p><p>At 03:47 UTC, the operator is performing the function the entire architecture <strong>silently assumes</strong> someone will perform, while the architecture itself makes performing it nearly impossible.</p><p>The systems mostly do not work. The operators work, and the systems usually fail only when the operators lose the ability to see, infer, or intervene safely.</p><p><em>The pager will go off again</em>. The dashboards will be wrong again. The runbook will be stale again.</p><p>The work is the same work:</p><p>act carefully, <strong>preserve reversibility,</strong> make the assumption explicit, and already know the next move before committing to the current one.</p><div><hr></div><h2>One more thing&#8230;</h2><p>The structural argument across this series is that distributed systems fail at the seams: composition, <strong>observation</strong>, and control are each independent sources of failure modes no single component owns and no single engineer fully sees.</p><p>The same argument applies, in concentrated form, to GPU programming.</p><p>A<strong> modern CUDA kernel </strong>is itself a tiny distributed system: dozens of streaming multiprocessors, thousands of warps, multiple memory hierarchies, delayed observation through performance counters, and sharply non-linear control through launch configuration and synchronization.</p><p><strong>Correctness is necessary</strong>. Performance lives in how composition, observation, and control interact under the workload you actually have, not the one your benchmark measured.</p><p>I wrote a <strong>deep CUDA guide</strong> from exactly this perspective: not isolated tricks, but how to reason about the GPU as a coupled dynamical system whose performance regimes and failure modes (<em>occupancy collapse, memory-bandwidth thrashing, warp divergence, pipeline stalls</em>) are structurally the same kinds of seam failures this series has been describing all along.</p><p><a href="https://lorenzobrada.gumroad.com/l/cuda_mastery">Read the CUDA Guide on Gumroad</a></p><p></p>]]></content:encoded></item></channel></rss>