Series · 80 sessions · ~160 hours
Deep Learning & LLMs From Scratch
An evergreen self-study series that takes you from “I know Python” to “I trained and served my own LLM”. 80 deep 2-hour sessions across 12 modules. Zero DL knowledge assumed on day one. In the Karpathy Zero-to-Hero storytelling register.
Sessions
80
Modules
12
Per session
~2 hr
Total time
~160 hr
M01 — Math for DL from scratch
sessions 1–6- S001NumPy Warm-up — Tensors Are Just Arrays
- S002Vectors, Matrices, and the Geometry of Data
- S003Broadcasting — The Rule That Runs the World
- S004Derivatives, the Chain Rule, and Computation Graphs
- S005Probability and Information — Surprise, Entropy, Cross-Entropy
- S006Optimization Intuition + Numerical Stability
M02 — Neural nets from scratch in NumPy
sessions 7–13M03 — PyTorch fluency
sessions 14–19M04 — Regularization + optimization
sessions 20–24M05 — Convnets + vision
sessions 25–30- S025The Convolution — From Cross-Correlation to Feature Maps
- S026LeNet → AlexNet → ResNet — the three architectures that mattered
- S027Training a Real CNN on CIFAR-10, End to End
- S028Transfer Learning — Frozen Backbones and Fine-Tuning
- S029Data Augmentation — RandAugment, Mixup, CutMix, and Why They Matter
- S030Vision Transformer — Attention All the Way Down
M06 — RNNs, sequences, embeddings
sessions 31–35M07 — Transformers from scratch
sessions 36–43- S036Self-Attention Derived from Scratch
- S037Multi-Head Attention — Why Eight Heads Beat One
- S038Positional Encodings — Sinusoidal, Learned, and a Sneak Peek at RoPE
- S039The Full Transformer Block — Encoder, Decoder, and What Order Everything Goes In
- S040nanoGPT Rebuilt from Scratch
- S041Training Your Decoder on Tiny Shakespeare
- S042KV Cache — Why Inference Is 100× Faster With It
- S043Rotary Embeddings and Flash-Attention Intuition
M08 — Tokenization + data + scaling laws
sessions 44–49M09 — Pretraining your foundation model
sessions 50–57- S050Model Architecture Choices — GPT, Llama, Mistral
- S051Distributed Training Basics — DDP and FSDP
- S052Renting Your First GPU — Runpod / Modal / Vast.ai Walkthrough
- S053Training a 100M-Param Model End-to-End
- S054Loss Curves and Eval Harness
- S055Checkpointing and Resuming
- S056Common Pretraining Bugs and Their Fingerprints
- S057Scaling from 100M to 1B — A Mental Model
M10 — Fine-tuning + alignment
sessions 58–65M11 — Efficient inference + serving
sessions 66–73- S066KV Cache Deep Dive
- S067Quantization — INT8, INT4, GPTQ, AWQ
- S068Pruning and Distillation — When Smaller Beats Quantized
- S069Speculative Decoding — 2–3× Free Speedup
- S070Continuous Batching — The Idea That Made LLM Serving Cheap
- S071Inside vLLM and TGI — PagedAttention, Block Manager, Schedulers
- S072Deploying an LLM API — FastAPI + Streaming + Modal
- S073Latency vs Throughput — Little's Law, TTFT, ITL, and Load Testing