The Software Frontier
Subscribe
Sign in
Home
Notes
Chat
Start Here
Archive
Leaderboard
About
DeepSeek V4-Flash: The Cost of Deciding What to Read
284 billion parameters rebuilt from the published constants, a million-token cache in 3.37 GiB, and the arithmetic showing that 4/5 of the attention…
Aug 10
•
Lorenzo Bradanini
and
Lorenzo Tettamanti
9
2
2
Recent posts
View all
How Blackwell’s Tensor Memory Actually Works
Blackwell's largest matrix instruction needs 256 registers per thread. The ceiling is 255.
Aug 3
•
Lorenzo Bradanini
and
Lorenzo Tettamanti
7
2
How CUDA Binaries Actually Work
A byte-level surgical dissection of the cubin and fatbin formats, and the second encoding nobody documents.
Jul 31
•
Lorenzo Bradanini
and
Lorenzo Tettamanti
6
3
How the NVIDIA Compiler Moat Actually Works
Inside the NVIDIA Compiler Moat: ptxas, SASS, and the 21 Bits Nobody Else Can Write
Jul 27
•
Lorenzo Bradanini
and
Lorenzo Tettamanti
4
3
Kimi K3: 2.8 Trillion Parameters, Four Bits at a Time
Moonshot just announced the largest open-weight model ever built. The parameter count is the headline. The serving stack is the story.
Jul 22
•
Lorenzo Bradanini
and
Lorenzo Tettamanti
6
1
3
Invite your friends to read The Software Frontier
A warm thank you
Aug 4
•
Lorenzo Bradanini
and
Lorenzo Tettamanti
8
1
Subscribe to get deep engineering insights you won’t find elsewhere!
Subscribe
This site requires JavaScript to run correctly. Please
turn on JavaScript
or unblock scripts