Resources
Video series
-
★ 3Blue1Brown, Neural Networks series — youtube.com/playlist?list=PLZHQObOWTQDNU6R1_67000Dx_ZCJB-3pi. Excellent visual, undergrad level birds-eye view on networks, gradient descent, and backpropagation, GPTs and attention.
-
★ Andrej Karpathy, Neural Networks: Zero to Hero — karpathy.ai/zero-to-hero.html. Much more involved video tutorial on building your own GPT.
-
Stanford CS224N, Natural Language Processing with Deep Learning — lectures on YouTube (youtube.com/@stanfordonline, search CS224N). The standard graduate NLP course; lectures on word vectors, attention, and transformers complement ours with more linguistic motivation.
-
Stanford CS25, Transformers United — web.stanford.edu/class/cs25. Guest lectures by the people who built these systems (Karpathy, transformer authors, interpretability researchers).
-
MIT 6.S191, Introduction to Deep Learning — introtodeeplearning.com. Video lectures.
-
Welch Labs — youtube.com/@welchlabs. One-off videos on the mathematics of deep learning.
Books and lecture notes
-
★ Higham & Higham, Deep Learning: An Introduction for Applied Mathematicians, SIAM Review 61(4), 2019 — arXiv:1801.05894. Forty pages: networks, backprop, and gradient descent with pseudocode.
-
★ Phuong & Hutter, Formal Algorithms for Transformers, 2022 — arXiv:2207.09238.
-
★ Jurafsky & Martin, Speech and Language Processing, 3rd ed. draft — free at web.stanford.edu/~jurafsky/slp3. The standard NLP textbook, aimed at undergraduates.
-
Bishop & Bishop, Deep Learning: Foundations and Concepts, Springer 2024 — free online at bishopbook.com. A probabilistic point of view.
-
Prince, Understanding Deep Learning, MIT Press 2023 — free PDF at udlbook.github.io/udlbook. Lots of pictures.
-
Goodfellow, Bengio & Courville, Deep Learning, MIT Press 2016 — free at deeplearningbook.org. Predates transformers.
-
Telgarsky, Deep Learning Theory lecture notes — mjt.cs.illinois.edu/dlt. Graduate notes: approximation, optimization, generalization with proofs.
-
Jentzen, Kuckuck & von Wurstemberger, Mathematical Introduction to Deep Learning, 2023 — arXiv:2310.20360. Encyclopedic.
Interactive tools and visualizations
- ★ LLM Visualization — bbycroft.net/llm. A 3D, animated walkthrough of every matrix in a small GPT, down to individual weights.
- TensorFlow Playground — playground.tensorflow.org. Train a tiny network in the browser and watch the decision boundary.
- Tiktokenizer — tiktokenizer.vercel.app. Paste text, see exactly how production tokenizers split it.
- The Illustrated Transformer — Jay Alammar, jalammar.github.io/illustrated-transformer. The most-read visual explainer of the architecture.
- Distill.pub — distill.pub. Interactive ML.
- Transformer Circuits thread — transformer-circuits.pub. Anthropic’s interpretability research program.
Code references
- karpathy/micrograd — github.com/karpathy/micrograd. Reverse-mode autodiff in ~100 lines.
- karpathy/nanoGPT — github.com/karpathy/nanoGPT. A complete GPT-2-class training run in ~600 lines.
- karpathy/minbpe — github.com/karpathy/minbpe. Minimal BPE tokenizers.
- PyTorch documentation — pytorch.org/docs; the 60-minute blitz tutorial
Historical and cultural
- Karpathy, The Unreasonable Effectiveness of Recurrent Neural Networks, 2015 — karpathy.github.io/2015/05/21/rnn-effectiveness. Famous pre-transformer blog post.
- Olah, Calculus on Computational Graphs: Backpropagation, 2015 — colah.github.io/posts/2015-08-Backprop. Backprop in six diagrams.
- Brown et al., Language Models are Few-Shot Learners (GPT-3), 2020 — arXiv:2005.14165. Where in-context learning was first observed at scale.
Primary sources, by lecture
Lecture 1 — language models and information theory
- Shannon, A Mathematical Theory of Communication, Bell System Tech. J. 1948 — the founding document; §I.3 already contains character-level -gram text generation.
- Shannon, Prediction and Entropy of Printed English, 1951 — estimates ~1 bit/character for English; your Step 1 model gets ~4.5 and your Step 8 model well under 2. (PDF)
Lectures 2–3 — networks, approximation, backpropagation
- Cybenko, Approximation by superpositions of a sigmoidal function, Math. Control Signals Systems, 1989. The universal approximation theorem.
- Pinkus, Approximation theory of the MLP model in neural networks, Acta Numerica 1999 — the survey with the sharp statements.
- Bengio, Ducharme, Vincent & Jauvin, A Neural Probabilistic Language Model, JMLR 2003 — PDF. The paper our Step 3 model comes from: embeddings + MLP for next-word prediction, in 2003.
- Baydin, Pearlmutter, Radul & Siskind, Automatic Differentiation in Machine Learning: a Survey, JMLR 2018 — arXiv:1502.05767. Forward vs reverse mode, done properly; the “cheap gradient principle” and its history.
- Griewank & Walther, Evaluating Derivatives, SIAM 2008 — the reference monograph on algorithmic differentiation.
Lectures 4–5 — attention and transformers
- ★ Vaswani et al., Attention Is All You Need, NeurIPS 2017 — arXiv:1706.03762. The transformer paper. Read it after Lecture 5 and be struck by how much of modern AI is in these 11 pages — and how much is not explained by them.
- Radford et al., Language Models are Unsupervised Multitask Learners (GPT-2), 2019 — PDF. The architecture our project replicates in miniature.
- Ba, Kiros & Hinton, Layer Normalization, 2016 — arXiv:1607.06450.
- Su et al., RoFormer: Enhanced Transformer with Rotary Position Embedding, 2021 — arXiv:2104.09864. RoPE: positional information as rotations — the most mathematically pleasing of the positional encodings, now standard in open models.
Lecture 6 — optimization and training
- Robbins & Monro, A Stochastic Approximation Method, Ann. Math. Stat. 1951. Where SGD’s convergence theory begins.
- Kingma & Ba, Adam: A Method for Stochastic Optimization, 2014 — arXiv:1412.6980; plus Loshchilov & Hutter, Decoupled Weight Decay Regularization (AdamW), arXiv:1711.05101.
- Glorot & Bengio, Understanding the difficulty of training deep feedforward neural networks, AISTATS 2010 — initialization as variance propagation.
- Jacot, Gabriel & Hongler, Neural Tangent Kernel, NeurIPS 2018 — arXiv:1806.07572.
- Belkin, Hsu, Ma & Mandal, Reconciling modern machine-learning practice and the classical bias–variance trade-off, PNAS 2019 — arXiv:1812.11118. Double descent.
Lecture 7 — tokenization, sampling, scaling
- Sennrich, Haddow & Birch, Neural Machine Translation of Rare Words with Subword Units, ACL 2016 — arXiv:1508.07909. BPE enters NLP.
- Holtzman et al., The Curious Case of Neural Text Degeneration, ICLR 2020 — arXiv:1904.09751. Nucleus sampling, and why greedy decoding produces degenerate text.
- ★ Kaplan et al., Scaling Laws for Neural Language Models, 2020 — arXiv:2001.08361.
- ★ Hoffmann et al., Training Compute-Optimal Large Language Models (Chinchilla), 2022 — arXiv:2203.15556. Power laws clean enough to make an analyst suspicious.
Lecture 8 — alignment and interpretability
- Ouyang et al., Training language models to follow instructions with human feedback (InstructGPT), 2022 — arXiv:2203.02155. The RLHF recipe.
- Rafailov et al., Direct Preference Optimization, NeurIPS 2023 — arXiv:2305.18290. The KL-regularized RLHF objective solved in closed form — a genuinely satisfying derivation.
- Hu et al., LoRA: Low-Rank Adaptation of Large Language Models, 2021 — arXiv:2106.09685.
- ★ Elhage et al., A Mathematical Framework for Transformer Circuits, Anthropic 2021 — transformer-circuits.pub/2021/framework. Attention-only transformers decomposed exactly into interpretable paths; the QK/OV circuit formalism used in our Lectures 4–5 asides.
- Olsson et al., In-context Learning and Induction Heads, Anthropic 2022 — transformer-circuits.pub/2022/in-context-learning-and-induction-heads.
- Elhage et al., Toy Models of Superposition, Anthropic 2022 — transformer-circuits.pub/2022/toy_model. Features as an overcomplete “frame” in activation space; almost-orthogonal vectors in high dimension doing real work.