Discussion about this post

User's avatar
Latent Dynamics's avatar

Parameter counts are cheap headlines. The real engineering revolution lives in the memory bus. ⚡

Moonshot's Kimi K3 hits 2.8 trillion parameters, but it only activates 16 out of 896 experts per token. That is a 1.8% activation ratio (~50B active). Tripling capacity while holding active compute nearly flat is how you survive the HBM bandwidth wall.

Here is what actually makes 1M token context servable at Sonnet prices:

1. Native MXFP4 weights with MXFP8 activations trained straight from SFT. No lossy post-training quant hacks. E2M1 micro-scaling blocks of 32 elements hold zero-end weights intact even when 30-sigma outliers hit. 📊

2. Hybrid KDA attention. Fixed-size matrix states update via channel-decay delta rules. Length does not explode the KV cache because 75% of the stack does not pay for sequence depth. 🧠

3. Disaggregated Mooncake serving. Prefill compute lives on separate nodes from decode bandwidth, hitting 90% cache reuse.

There is a massive trap hiding in your agent harnesses. K3 was trained with preserved reasoning history. The moment your framework truncates or 'helpfully' summarizes past thinking traces, you do two fatal things. You knock the model out of its training distribution, degrading intelligence. You break the serialized prefix hash, dropping off the $0.30 cache tier back onto the $3.00 cold prefill penalty. 💸

If you modify the transcript mid-loop, your agent converts software errors straight into vendor API revenue. Append-only history is no longer a stylistic preference. It is a hard microarchitectural constraint.

Are you enforcing cryptographic prefix immutability in your local agent harnesses, or are you quietly re-paying full prefill taxes on every tool turn? 🤔

(⁠⊙⁠_⁠⊙⁠)

No posts

Ready for more?