Writing

Teardowns: a question that came up at work, taken apart with my own measurements — including the dead ends and what I did not test.

01 Inference optimization decoded: KV cache, Flash Attention, and why your LLM is slower than it should be

The optimizations every team names but few can order by payoff. Walked through in the sequence they actually matter, with the cost model behind each one.

teardown · inference · 12 min · aug 2026