The Software Frontier
Subscribe
Sign in
Home
Notes
Chat
Start Here
Archive
Leaderboard
About
Latest
Top
Discussions
Inside Google’s TPU: How it works, and what it costs
We ran it against five TPU generations, from v4 to Ironwood, then followed the money: a rate card that prices bandwidth, an anchor tenant paying 13% of…
6 hrs ago
•
Lorenzo Bradanini
and
Lorenzo Tettamanti
4
2
A wonderful dive into how FP4 works
NVFP4 and MXFP4 multiply the same sixteen codes. The difference is one byte of scale per block, and it is worth 1.6 dB on Gaussian data, 4.2 dB on…
Oct 4
•
Lorenzo Bradanini
and
Lorenzo Tettamanti
4
2
September 2026
Why Cached Tokens Cost 10% and Vanish in 5 Minutes
Every agent resends its whole context at every step, and every provider bills the repeat at a discount with a timer attached. The discount and the timer…
Sep 29
•
Lorenzo Bradanini
and
Lorenzo Tettamanti
3
1
How Mojo Actually Compiles
Ask Mojo 1.1 for sm_90 and it quietly returns sm_90a, hinting that portability no longer lives in the binary. An H100 executable can carry kernels…
Sep 24
•
Lorenzo Bradanini
and
Lorenzo Tettamanti
5
1
How NVIDIA's Data Center Business Actually Works in 2026
NVIDIA booked $89.0 billion of data center revenue in its latest quarter. The company now reports about $530 billion of commitments and maximum…
Sep 20
•
Lorenzo Bradanini
and
Lorenzo Tettamanti
8
1
Unleashing OpenCL’s secrets
A technical account of the standard, its implementations, and where it actually stands in 2026.
Sep 10
•
Lorenzo Bradanini
and
Lorenzo Tettamanti
9
3
August 2026
When Batching Stops Working
Everyone says to raise the batch until you are compute bound. Solve for where that stops and every datacenter GPU of the last four generations goes…
Aug 28
•
Lorenzo Bradanini
and
Lorenzo Tettamanti
4
3
Distilling in depth ROCm: How it Actually Works
AMD ships 2,905,048 tuned GEMM decisions in a public git repository. Zero of CDNA 4’s 75 matrix instructions exist on CDNA 5.
Aug 25
•
Lorenzo Bradanini
and
Lorenzo Tettamanti
4
2
Exploring how Triton actually compiles
Same source, same block sizes, same warp count. An integer the programmer never writes decides whether the loop is pipelined at all, and a second one…
Aug 22
•
Lorenzo Bradanini
and
Lorenzo Tettamanti
7
2
DeepSeek V4-Flash: The Cost of Deciding What to Read
284 billion parameters rebuilt from the published constants, a million-token cache in 3.37 GiB, and the arithmetic showing that 4/5 of the attention…
Aug 10
•
Lorenzo Bradanini
and
Lorenzo Tettamanti
9
2
2
Invite your friends to read The Software Frontier
A warm thank you
Aug 4
•
Lorenzo Bradanini
and
Lorenzo Tettamanti
8
1
How Blackwell’s Tensor Memory Actually Works
Blackwell's largest matrix instruction needs 256 registers per thread. The ceiling is 255.
Aug 3
•
Lorenzo Bradanini
and
Lorenzo Tettamanti
7
2
This site requires JavaScript to run correctly. Please
turn on JavaScript
or unblock scripts