Learning path / 67 published lessons
Deep Learning and LLMs: an inspectable learning laboratory
A broad deep-learning and LLM progression with a verified NumPy spine, advanced framework and systems lessons, source evidence and explicit execution boundaries.
What you’ll work toward
- Derive and test core neural computations
- Navigate all 80 deep-learning topics and the classical ML branch
- Distinguish source-backed concepts, executed CPU examples and unexecuted GPU experiments
Completion is stored on this device only. Nothing is locked; start where it makes sense.
Start this pathBefore the first lesson
- Python functions, loops and dictionaries; basic algebra
These are the starting lesson’s prerequisites, not requirements for every advanced topic below.
How to practise this subject
Attempt each lesson’s exercises before opening the explanation. Reconstruct its main example, change an assumption, and use the stated test boundaries to judge what you have actually checked.
- 01
Arrays, shapes, and the first linear model
Master array ownership and axes, then audit attention, context and quantization claims with exact CPU checks.
- 02
Vectors, Matrices, and the Geometry of Data
Derive geometry, matrix products, SVD and retrieval scores while testing the limits of each analogy.
- 03
Broadcasting — The Rule That Runs the World
Predict shape sharing, memory costs and numerical failure cases in NumPy and transformer patterns.
- 04
Derivatives are local measurements; backprop is bookkeeping
Rebuild the chain rule from local sensitivities through scalar autograd, matrix pullbacks and checkpointing.
- 05
Probability, cross-entropy, and stable logits
Derive information measures and stable probability losses, then distinguish modern language and preference objectives.
- 06
Optimization: steps, momentum, AdamW, and honest experiments
Trace optimizer updates and numerical safeguards without mistaking toy convergence for universal model performance.
- 07
One neuron: logistic regression, stable loss, and its boundary
Train and audit a sigmoid neuron, verify exact framework parity, and prove and repair its XOR limitation.
- 08
Activation Functions — sigmoid, tanh, ReLU, GELU, SwiGLU, xIELU
Derive and test activation functions, full gradient chains, gated parameter budgets and source-corrected xIELU.
- 09
The Forward Pass: Shapes, XOR and Parameter Counts
Compute an MLP by hand, construct a strict-margin XOR classifier, and distinguish actual architecture counts from universal claims.
- 10
Backprop by Hand: Nine Gradients and a Verified XOR Network
Derive all nine ReLU-network gradients, preserve the tanh XOR experiment, and verify finite differences against CPU PyTorch.
- 11
MLP on MNIST: A Complete Local-Only NumPy Recipe
Load existing local IDX files, implement stable multiclass loss and backpropagation, and separate a synthetic contract from real digit evaluation.
- 12
Micrograd: Build and Test a Scalar Autograd Engine
Implement complete scalar reverse-mode differentiation, shared-edge accumulation, stable cross entropy and a four-point neural network.
- 13
Vectorization: Measure Speed, Shapes and Memory
Compare exact scalar and array programs, derive einsum contractions and FLOP counts, and measure performance without promised speedup factors.
- 14
PyTorch Tensors and Autograd: Leaves, VJPs and Safe Updates
Port NumPy mathematics to raw PyTorch tensors and verify gradient accumulation, cotangents, graph lifetime and mutation controls.
- 15
nn.Module: Registration, Composition and Checkpoint Contracts
Structure models with registered parameters and buffers, distinguish eval from no-grad, and verify trusted state/config round-trips.
- 16
Datasets, DataLoaders and Collate: Coverage Before Throughput
Test indexed sampling, ragged collation, worker/rank sharding and local stream resume while treating optional cloud adapters as unexecuted.
- 17
The Training Loop, Logging, and Reproducibility
Build and verify a resumable training state machine with RNG, sampler, optimizer, scheduler, EMA and safe checkpoints.
- 18
GPU and Mixed Precision
Separate device placement, operation-specific mixed precision, loss scaling and measured performance without numerical myths.
- 19
Debugging — nan, inf, and Overfitting a Single Batch
Diagnose training with exact objective contracts, finite checks, tiny-batch fits, negative controls and profiling.
- 20
Dropout and generalization — masks, scaling and controlled experiments
Derive dropout moments, preserve leakage-resistant evaluation, test mask variants and choose regularization by controlled experiments.
- 21
Batch Norm, Layer Norm, and RMSNorm — the normalization tour
Derive normalization axes, affine and buffer contracts, implement BN/LN/RMSNorm and test train/eval and calibration failures.
- 22
Weight Initialization — Xavier, He, and μP
Derive fan and second-moment initialization, test activation and gradient propagation, and understand bounded width-transfer parameterizations.
- 23
Schedulers — Warmup, Cosine, One-Cycle, and WSD
Implement and verify warmup, linear, cosine, one-cycle and WSD clocks, resume behavior and fair schedule comparisons.
- 24
Gradient Clipping and Accumulation
Derive weighted accumulation and global clipping, flush unequal tails and distinguish AMP, DDP and activation-checkpointing contracts.
- 25
Convolution: local shared operators, shapes and gradients
Compute original convolution examples and general output shapes Implement multi-channel forward and backward accumulation Compare receptive fields, pooling and modern convolution designs.
- 26
CNN architectures: LeNet, AlexNet, residual learning and modern backbones
Build LeNet-like and ResNet18 models with exact parameter counts Distinguish degradation from vanishing gradients Design fair architecture and local-data comparisons.
- 27
CIFAR-10 training: a complete local pipeline and controlled experiments
Build a complete local CIFAR training and selection pipeline Test normalization, parameter groups and weighted evaluation Diagnose training failures with controlled ablations.
- 28
Transfer learning: frozen features, fine-tuning and buffer discipline
Implement local-checkpoint linear and fine-tuning regimes Freeze parameters and BatchNorm buffers deliberately Compare transfer objectives and source-target assumptions.
- 29
Data augmentation: tested mixing, task assumptions and calibration
Implement mixup and clipped-area CutMix with exact targets Compute calibration and soft-target loss controls Choose task-valid augmentation and controlled ablations.
- 30
Vision Transformers: patch equivalence, position and attention cost
Prove patch projection output and gradient equivalence Build and test a small complete Vision Transformer Explain position interpolation, parameter counts and quadratic cost.
- 31
Word2Vec: complete skip-gram training, geometry and honest evaluation
Train complete skip-gram negative sampling on bundled text Derive lookup and binary-loss gradients and sampling contracts Query and evaluate embedding geometry without semantic guarantees.
- 32
RNNs from scratch: complete BPTT, character generation and structured recurrence
Implement complete NumPy BPTT and verify every coordinate Train and generate with a complete character RNN locally Distinguish clipping, truncation, gates, attention and affine scans.
- 33
LSTM and GRU — Gates, State and Gradient Paths
An ordinary RNN repeatedly applies a nonlinear state transformation. Its time derivative is a product of **time-varying Jacobians**, not just the recurrent matrix's eigenvalue raised to a power. LSTM introduced a separately carried cell state; the la
- 34
Seq2Seq — A Complete Length-Aware Encoder–Decoder
A sequence-to-sequence model defines a conditional distribution over a target sequence given a source. In the fixed-state recurrent version, an encoder reads the source and passes its final state to an autoregressive decoder. For an LSTM the interfac
- 35
Why Attention — Query-Dependent Source Access
A fixed-state encoder-decoder must route all source information through its terminal state. Attention retains the sequence of encoder representations and lets each decoder step retrieve a different weighted context. This removes the requirement that
- 36
Attention is a weighted sum with a precise mask
Calculate attention numerically, derive its softmax backward pass, and test causal and permutation properties.
- 37
Multi-Head Attention — Shapes, Specialization and KV Budgets
A single attention head already forms a distribution over many keys; it is not forced to pick exactly one token. Multi-head attention gives several independently parameterized distributions and value projections, then mixes their results. At fixed mo
- 38
Positional Encodings — Sinusoids, Learned Tables and RoPE
Unmasked self-attention with shared tokenwise projections and no position signal is **permutation-equivariant**: permuting input rows permutes output rows. It is not invariant unless an additional permutation-invariant readout removes row order. Thus
- 39
Transformer Blocks — Full Architectures and an Analytic Decoder
Attention mixes information across token positions. A position-wise feed-forward network transforms features with the **same weights at each position**. Normalization sets a branch's input or output scale, and residual addition carries an explicit sk
- 40
nanoGPT Rebuilt — A Complete Tested Decoder
The useful exercise in nanoGPT is tracing a compact decoder-only model end to end: configuration, embeddings, fused QKV attention, pre-norm blocks, final norm, tied output projection, optimizer groups and generation. Its public model/train files are
- 41
Training a Decoder on Local Text — Tiny Shakespeare Recipe
A small character-level decoder is a useful way to inspect tokenization, shifted targets, causal loss, optimizer state and generation. Tiny Shakespeare is one possible **already-local** corpus, not data bundled or downloaded by this audit. This lesso
- 42
KV caching: correct positions, masks, memory and reuse
Implement and test a complete cached decoder, derive cache memory, and distinguish paging, sharing, compression and eviction.
- 43
RoPE and FlashAttention: relative positions and online softmax
Derive rotary attention and tiled normalization, test CPU parity, and separate positional extension from GPU kernel performance.
- 44
Byte-pair encoding: merges, Unicode and artifact contracts
Build and inspect deterministic byte BPE, correct weighted merge traces, preserve Unicode bytes and freeze tokenizer artifacts.
- 45
SentencePiece and unigram: lattices, normalization and compatibility
Derive unigram marginal likelihood and soft EM, test segmentation, and distinguish tokenizer algorithms from toolkit and normalization choices.
- 46
Pretraining data: auditable extraction, grouping and token shards
Design a provenance-aware pretraining pipeline with tested filters, grouped splits, boundary-safe windows and validated token shards.
- 47
Chinchilla scaling: fit assumptions, budgets and serving cost
Derive compute-constrained allocation, correct the numerical budget, and separate fixed-ratio heuristics from fitted and inference-aware scaling.
- 48
Data mixing: sampling, exposure, proxy ablations and curriculum
Implement weighted shard sampling, measure domain exposure, run a complete synthetic ablation and explain DoReMi and token-clock schedules.
- 49
Quality and deduplication: MinHash, LSH and measured decisions
Implement probabilistic near-duplicate retrieval with exact verification, audit learned quality policies and screen contamination with explicit limits.
- 50
Architecture choices: derive and test a 100M decoder
Implement pre-norm, RMSNorm, RoPE, GQA and SwiGLU; reconcile exact model counts, cache budgets and modern architecture claims.
- 51
Training a 100M model: complete local recipe and budget
Build a local-shard training state machine with exact token budgets, validation, persistence and honest resource-dependent execution limits.
- 52
Loss curves and evaluation: likelihood, tasks and controls
Preserve the exact synthetic decoder experiment while restoring token-weighted metrics, downstream evaluation, uncertainty and contamination audits.
- 53
Checkpointing and resuming: consistent state and durable publication
Implement trusted local restart snapshots, test continuation and failures, and derive retention and distributed recovery boundaries.
- 54
Pretraining bugs: hypotheses, counterexamples and tests
Diagnose stalled, nonfinite, drifting or misleading training runs through bounded interventions rather than loss-curve folklore.
- 55
Supervised fine-tuning: objective, templates and local recipes
Derive assistant-only supervision, test masks and conditioning gradients, and specify pinned local SFT with template and regression checks.
- 56
Instruction datasets: provenance, normalization and mixtures
Audit exact dataset identities and rights, reconstruct conversation trees, and derive reproducible example and supervised-token sampling weights.
- 57
Adaptation with LoRA: low-rank updates and masked supervision
Derive low-rank adapter gradients, inject and test a complete layer, and audit versioned SFT, merge and variant workflows.
- 58
QLoRA: NF4 storage, gradients and local integration
Derive block-scale memory, implement a CPU NF4 teaching quantizer, and distinguish adapter gradients, paging and merge behavior.
- 59
Reward Modeling: Pairwise Likelihood, Data and Failure Modes
Derive preference likelihoods, test reward gradients and pooling, and audit data, calibration and reward hacking.
- 60
Evaluation and Red-Teaming: Protocols, Uncertainty and Safety
Build reproducible capability, chat, domain and red-team evaluations with correct metrics, overlap analysis and uncertainty.
- 61
KV Cache Deep Dive: Correctness, Memory and Serving Budgets
Test exact cached decoding and grouped-query attention, derive byte budgets, and distinguish paging, precision and eviction.
- 62
Pruning and Distillation: Exact Masks and Correct Token Losses
Implement exact pruning masks and shifted masked distillation, derive temperature gradients, and evaluate compression tradeoffs.
- 63
Continuous Batching: Scheduling, Cache Admission and Capacity
Trace static and continuous batches, enforce cache admission, derive Little’s law and diagnose latency-throughput tradeoffs.
- 64
Inside vLLM and TGI: Paging, Versioned Schedulers and Operations
Implement a reference-counted block allocator, inspect pinned vLLM internals, and prepare honest serving and observability recipes.
- 65
Deploying an LLM API: Bounded FastAPI Streaming and Lifecycle
Preserve SSE framing, bound admission, authenticate and clean up streaming requests, and design deployment, retry and usage contracts.
- 66
RAG as a testable evidence pipeline, not a truth switch
Implement chunking, exact retrieval and rank fusion, integrate evidence packing and evaluate a complete RAG pipeline.
- 67
Capstone: a reproducible model-and-evidence laboratory
Complete a reproducible CPU laboratory or a fully specified 100M training and release project with explicit evidence gates.