A ground-up implementation of the transformer — self-attention, multi-head attention, positional encoding, and the full encoder stack — with derivations at each step.
machine learning engineer
monolog_
Research notes on machine learning, AI systems, and the mathematics that make them work — written by someone who builds them from scratch.
latest posts
all posts ->The PyTorch patterns that actually matter when you're training real models — not tutorials, not toy examples.
On starting a research blog, what I build, and why I think ML writing should go deeper than paper summaries.
3 min read