The Software Frontier
Subscribe
Sign in
Home
Notes
Chat
Start Here
Archive
Leaderboard
About
How Blackwell’s Tensor Memory Actually Works
Blackwell's largest matrix instruction needs 256 registers per thread. The ceiling is 255.
Aug 3
•
Lorenzo Bradanini
and
Lorenzo Tettamanti
7
2
Recent posts
View all
How CUDA Binaries Actually Work
A byte-level surgical dissection of the cubin and fatbin formats, and the second encoding nobody documents.
Jul 31
•
Lorenzo Bradanini
and
Lorenzo Tettamanti
5
3
How the NVIDIA Compiler Moat Actually Works
Inside the NVIDIA Compiler Moat: ptxas, SASS, and the 21 Bits Nobody Else Can Write
Jul 27
•
Lorenzo Bradanini
and
Lorenzo Tettamanti
4
3
Kimi K3: 2.8 Trillion Parameters, Four Bits at a Time
Moonshot just announced the largest open-weight model ever built. The parameter count is the headline. The serving stack is the story.
Jul 22
•
Lorenzo Bradanini
and
Lorenzo Tettamanti
6
1
3
Exploring how MLIR works: the compiler rewiring the AI stack
From tensor graphs to machine code, MLIR is quietly becoming the abstraction layer connecting modern AI frameworks to increasingly specialized hardware.
Jul 15
•
Lorenzo Bradanini
and
Lorenzo Tettamanti
8
1
3
Invite your friends to read The Software Frontier
A warm thank you
Aug 4
•
Lorenzo Bradanini
and
Lorenzo Tettamanti
8
1
Subscribe to get deep engineering insights you won’t find elsewhere!
Subscribe
This site requires JavaScript to run correctly. Please
turn on JavaScript
or unblock scripts