Hi. I am Vatshank. I work in Machine Learning.
And this is an attempt at a technical blog. There is a grand total of one post right now. We'll see how it goes.
Adapting FlashAttention for inference
12 August 2026
Implementing split-KV for FlashAttention 3 in CuteDSL.
1CuteDSL kernel meets vLLM's CUDAGraph capture stream
06 August 2026
Debugging CUDA stream shenanigans.
201 December 2023
Understanding the exact algorithm for efficient computation of self-attention.
3Chinchilla Scaling Laws (Hoffman et al. 2022)
19 June 2026