A ground-up implementation of the transformer — self-attention, multi-head attention, positional encoding, and the full encoder stack — with derivations at each step.
tag
#machine-learning
3 posts.
The PyTorch patterns that actually matter when you're training real models — not tutorials, not toy examples.
On starting a research blog, what I build, and why I think ML writing should go deeper than paper summaries.
3 min read