Notes on machine learning systems

Machine learning,at full resolution.

I take models apart: the bits in the weights, the bytes in memory, the kernels on the GPU. Then I write down what I find.

essays
21
hours of reading
5
first essay
2024

Fig. 1: weights atbits.The lens restores full precision.

Latest essay

Fresh from the workbench

Fig. 2: Causal attention

20 min read

What Kimi K3 asks of the hardware

Reading the Kimi K3 announcement from the systems side. How linear attention changes the KV-cache budget, what 896 experts mean for the interconnect, and why the model ships in MXFP4.

Archive21

Every essay, newest first

All topics →

2026

  1. What Kimi K3 asks of the hardware

    Reading the Kimi K3 announcement from the systems side. How linear attention changes the KV-cache budget, what 896 experts mean for the interconnect, and why the model ships in MXFP4.

  2. GGUF quantization, bit by bit

    Q8_0 to IQ1_S and everything in between. Block layouts, scale hierarchies, codebooks, the importance matrix, and the mixed recipes used by llama.cpp, Unsloth, and DwarfStar.

  3. GRPO is mostly a systems problem

    The GRPO config is small. The hard part is keeping rollout workers, trainers, rewards, and policy versions in the same story.

  4. A fused Lion optimizer kernel with Hugging Face kernels

    I wrote a fused Lion optimizer step with Hugging Face kernels, packaged it with kernel-builder, and tested it on CUDA and Metal/MPS.

  5. TurboQuant: What 3-Bit KV Caches Actually Mean for Your Inference Stack

    A quick look to Google's new quantization method, with a bit of drama

  6. Transformer Attention, Backwards

    You already know what a language model does. Let's reverse-engineer how.

2025

  1. Pipeline Parallelism: Surgery on Models Too Big for the Operating Table

    In our latest article, we explored Data Parallelism.

  2. Data Parallelism: Scaling LLM Training Through Parallel Processing

    In my latest article , I discussed the theoretical memory usage needed for inference and training with LLMs, highlighting the memory cost of each component involved in the process.

  3. The Memory Anatomy of Large Language Models: A Surgeon's Guide

    Picture this: you’re about to deploy a shiny new 70-billion parameter language model, and your colleague asks the dreaded question: “How much GPU memory do we need?”.

  4. From Sequential to Parallel: Your Journey into GPU Programming with Triton

    We all know that GPU programming is hard.

  5. The Transformer's Anatomy: A Deep Dive into the Architecture that Revolutionized Machine Learning

    In the vast landscape of machine learning, few architectures have captured the imagination and transformed the field as profoundly as the Transformer. Like a master anatomist approaching a complex organism, we must carefully dissect each component to understand how this remarkable architecture breathes life into modern AI systems.

  6. Move Fast or Die Slow

    Today’s article steps back from our usual technical deep-dives to examine the strategic importance of ML optimization.

2024

  1. The Machine Learning Surgeon's Guide to Quantization: Precision Cuts for Smarter Models

    As humans, we perceive space and time as a seamless, continuous flow.

  2. The Operating Room Setup

    Undoubtedly, one of the most critical aspects of machine learning is understanding the theory—without grasping how machines learn, you’ll never excel as an ML Surgeon!

  3. Dissecting torch.compile: Surgical Precision in PyTorch Optimization

    You can take a look at the GitHub repository of this blogpost at this link

  4. A quick incision: ten minutes to RAG

    In under 10 minutes, you’ll discover what RAG is, how to build a prototype for it in Python, and—most importantly—the true weight of the infamous Mr. Fat Raccoon.

  5. Performing Kernel Surgery: Profiling CUDA Kernels with NVIDIA Nsight Compute

    Being a Machine Learning Surgeon is not an easy life.

  6. A Machine Learning Surgeon’s Toolkit: Advanced Matrix Multiplication in CUDA

    If you want to learn how to write a CUDA kernel for matrix multiplication, look no further!

  7. Cerebral Cortex and Hippocampus: Understanding the Computational and Memory Design of GPUs

    Just as a surgeon needs to understand anatomy, knowing GPU architecture is key to writing efficient kernels. You don’t need to be a hardware expert—just grasp the basics to unlock their full potential

  8. Hello CUDA: A Surgical Dissection

    CUDA enables developers to harness NVIDIA GPUs for general-purpose tasks. This article guides you to a "Hello, World!" program as a starting point.

  9. An Introduction to Sparsity for Efficient Neural Network Inference

    Sparsity is a solution to reduce the number of parameters and number of operations in Neural Networks, granting outstanding computational speedups and memory savings during inference.