Search Tech Journey

Find topics, journeys and posts

Series · 80 sessions · ~160 hours

Deep Learning & LLMs From Scratch

An evergreen self-study series that takes you from “I know Python” to “I trained and served my own LLM”. 80 deep 2-hour sessions across 12 modules. Zero DL knowledge assumed on day one. In the Karpathy Zero-to-Hero storytelling register.

Sessions
80
Modules
12
Per session
~2 hr
Total time
~160 hr

M01 — Math for DL from scratch

sessions 16
  1. S001NumPy Warm-up — Tensors Are Just Arrays
  2. S002Vectors, Matrices, and the Geometry of Data
  3. S003Broadcasting — The Rule That Runs the World
  4. S004Derivatives, the Chain Rule, and Computation Graphs
  5. S005Probability and Information — Surprise, Entropy, Cross-Entropy
  6. S006Optimization Intuition + Numerical Stability

M02 — Neural nets from scratch in NumPy

sessions 713
  1. S007A Single Artificial Neuron
  2. S008Activation Functions — sigmoid, tanh, ReLU, GELU, SwiGLU, xIELU
  3. S009The Forward Pass for a 2-Layer MLP
  4. S010Backprop by Hand
  5. S011MLP on MNIST — All NumPy, No PyTorch
  6. S012Micrograd — Build a Tiny Autograd Engine
  7. S013Vectorization and Speed

M03 — PyTorch fluency

sessions 1419
  1. S014PyTorch Tensors and Autograd
  2. S015`nn.Module`, `nn.Sequential`, and the Layer Abstraction
  3. S016Datasets, DataLoaders, and Collate
  4. S017The Training Loop, Logging, and Reproducibility
  5. S018GPU and Mixed Precision
  6. S019Debugging — nan, inf, and Overfitting a Single Batch

M04 — Regularization + optimization

sessions 2024
  1. S020Dropout — Why Random Kills the Smart
  2. S021Batch Norm, Layer Norm, and RMSNorm — the normalization tour
  3. S022Weight Initialization — Xavier, He, and μP
  4. S023Schedulers — Warmup, Cosine, One-Cycle, and WSD
  5. S024Gradient Clipping and Accumulation

M05 — Convnets + vision

sessions 2530
  1. S025The Convolution — From Cross-Correlation to Feature Maps
  2. S026LeNet → AlexNet → ResNet — the three architectures that mattered
  3. S027Training a Real CNN on CIFAR-10, End to End
  4. S028Transfer Learning — Frozen Backbones and Fine-Tuning
  5. S029Data Augmentation — RandAugment, Mixup, CutMix, and Why They Matter
  6. S030Vision Transformer — Attention All the Way Down

M06 — RNNs, sequences, embeddings

sessions 3135
  1. S031Word2Vec — Words as Vectors, Vectors as Meaning
  2. S032RNN From Scratch — Modelling Sequences One Step at a Time
  3. S033LSTM and GRU — Gating the Gradient Highway
  4. S034Seq2Seq — Two LSTMs, One Translation, One Bottleneck
  5. S035Why Attention — Bahdanau's Scribble That Ate NLP

M07 — Transformers from scratch

sessions 3643
  1. S036Self-Attention Derived from Scratch
  2. S037Multi-Head Attention — Why Eight Heads Beat One
  3. S038Positional Encodings — Sinusoidal, Learned, and a Sneak Peek at RoPE
  4. S039The Full Transformer Block — Encoder, Decoder, and What Order Everything Goes In
  5. S040nanoGPT Rebuilt from Scratch
  6. S041Training Your Decoder on Tiny Shakespeare
  7. S042KV Cache — Why Inference Is 100× Faster With It
  8. S043Rotary Embeddings and Flash-Attention Intuition

M08 — Tokenization + data + scaling laws

sessions 4449
  1. S044BPE from Scratch — Byte by Byte
  2. S045SentencePiece and Modern Tokenizers
  3. S046Building a Pretraining Dataset (SlimPajama-Style)
  4. S047Chinchilla Scaling Laws
  5. S048Data Mixing and Curriculum
  6. S049Quality Filters and Deduplication in Practice

M09 — Pretraining your foundation model

sessions 5057
  1. S050Model Architecture Choices — GPT, Llama, Mistral
  2. S051Distributed Training Basics — DDP and FSDP
  3. S052Renting Your First GPU — Runpod / Modal / Vast.ai Walkthrough
  4. S053Training a 100M-Param Model End-to-End
  5. S054Loss Curves and Eval Harness
  6. S055Checkpointing and Resuming
  7. S056Common Pretraining Bugs and Their Fingerprints
  8. S057Scaling from 100M to 1B — A Mental Model

M10 — Fine-tuning + alignment

sessions 5865
  1. S058Supervised Fine-Tuning (SFT)
  2. S059Instruction Datasets — Alpaca, Dolly, OpenAssistant
  3. S060LoRA from Scratch
  4. S061QLoRA — 4-bit + LoRA
  5. S062Reward Modeling
  6. S063RLHF with PPO — Intuition + Minimal Code
  7. S064DPO — RLHF without the RL
  8. S065Evaluation and Red-Teaming

M11 — Efficient inference + serving

sessions 6673
  1. S066KV Cache Deep Dive
  2. S067Quantization — INT8, INT4, GPTQ, AWQ
  3. S068Pruning and Distillation — When Smaller Beats Quantized
  4. S069Speculative Decoding — 2–3× Free Speedup
  5. S070Continuous Batching — The Idea That Made LLM Serving Cheap
  6. S071Inside vLLM and TGI — PagedAttention, Block Manager, Schedulers
  7. S072Deploying an LLM API — FastAPI + Streaming + Modal
  8. S073Latency vs Throughput — Little's Law, TTFT, ITL, and Load Testing

M12 — Capstone + wildcards

sessions 7480
  1. S074Multimodal — CLIP → LLaVA
  2. S075Agents and Tool Use
  3. S076Retrieval-Augmented Generation (RAG) from Scratch
  4. S077Mixture of Experts (MoE) Intuition
  5. S078Long Context — RoPE Scaling, YaRN, ALiBi
  6. S079Red-Team + Safety Eval
  7. S080CAPSTONE — Train and Serve Your Own 100M Instruct Model
Sessions 001–003 are fully written as exemplars. Sessions 004–080 are scaffolded and will be enriched in rolling batches — see planning/DL-LLMS-SERIES-DESIGN.md in the repo for the full curriculum table and enrichment plan.