dwarez

Browse by topic

Topics

Posts grouped by the tags already attached to each article.

ai9

attention1

  • What Kimi K3 asks of the hardware

    / Reading the Kimi K3 announcement from the systems side. How linear attention changes the KV-cache budget, what 896 experts mean for the interconnect, and why the model ships in MXFP4.

compilers1

cuda6

efficiency11

gguf1

  • GGUF quantization, bit by bit

    / Q8_0 to IQ1_S and everything in between. Block layouts, scale hierarchies, codebooks, the importance matrix, and the mixed recipes used by llama.cpp, Unsloth, and DwarfStar.

gpu1

grpo1

  • GRPO is mostly a systems problem

    / The GRPO config is small. The hard part is keeping rollout workers, trainers, rewards, and policy versions in the same story.

hardware1

  • What Kimi K3 asks of the hardware

    / Reading the Kimi K3 announcement from the systems side. How linear attention changes the KV-cache budget, what 896 experts mean for the interconnect, and why the model ships in MXFP4.

inference7

kernels1

metal1

ml20

mlsys2

  • What Kimi K3 asks of the hardware

    / Reading the Kimi K3 announcement from the systems side. How linear attention changes the KV-cache budget, what 896 experts mean for the interconnect, and why the model ships in MXFP4.

  • GRPO is mostly a systems problem

    / The GRPO config is small. The hard part is keeping rollout workers, trainers, rewards, and policy versions in the same story.

moe1

  • What Kimi K3 asks of the hardware

    / Reading the Kimi K3 announcement from the systems side. How linear attention changes the KV-cache budget, what 896 experts mean for the interconnect, and why the model ships in MXFP4.

optimizers1

parallelism3

pruning1

pytorch2

quantization3

rl2

scaling3

sparsity1

strategy1

  • Move Fast or Die Slow

    / Today’s article steps back from our usual technical deep-dives to examine the strategic importance of ML optimization.

systems1

  • GRPO is mostly a systems problem

    / The GRPO config is small. The hard part is keeping rollout workers, trainers, rewards, and policy versions in the same story.

training3

transformers1

triton1