LLM Optimizers: From SGD and AdamW to SOAP, Muon, and Scalable Training
Published:
Implementation-first notes on optimizer geometry, AdamW, memory and distributed systems, and matrix-aware methods including SOAP and Muon.
Published:
Implementation-first notes on optimizer geometry, AdamW, memory and distributed systems, and matrix-aware methods including SOAP and Muon.
Published:
Implementation notes on byte-level BPE, fixed-tokenization failure modes, and dynamic byte-level models.
Published:
Today is 2026.3.23, I finally have my personal homepage with the help of Codex!