Learning path / 13 published lessons
Deep learning architectures and training
Follow neural-network mechanisms from tensor and gradient semantics through architectures, training, adaptation and efficient inference. Keep the mathematical model separate from hardware-specific performance claims.
What you’ll work toward
- Compute memory and communication under explicit dtype and topology assumptions
- Trace and audit complete DDP and FSDP2 reference lifecycles
- Design gradient-equivalence, scaling and checkpoint-recovery checks
- Choose GPU capacity and compare complete provider contracts without stale prices
Completion is stored on this device only. Nothing is locked; start where it makes sense.
Start this pathBefore the first lesson
- Python and arrays; preceding topics in the learning progression
These are the starting lesson’s prerequisites, not requirements for every advanced topic below.
How to practise this subject
Track tensor shapes and gradients through a small worked case. State which quantities a local model checks and which require a real framework, accelerator or distributed run.
- 01
Distributed Training Basics — DDP and FSDP
Derive distributed-training memory and communication costs, implement rank-aware reference loops, and distinguish DDP, ZeRO and FSDP2.
- 02
Renting Your First GPU — Runpod / Modal / Vast.ai Walkthrough
Choose GPU resources from measured workloads, compare service lifecycles, and plan safe transfer, monitoring, checkpointing and teardown.
- 03
Scaling from 100M to 1B — A Mental Model
Calculate compute, memory and elapsed-time ranges across model scales, interpret published receipts, and plan infrastructure from measured constraints.
- 04
RLHF with PPO — Intuition + Minimal Code
Derive PPO clipping and GAE, execute a complete tiny CPU loop, and audit language-model trainer, reward and modern RL workflows.
- 05
DPO — RLHF without the RL
Derive DPO from KL-regularized rewards, test masked log probabilities and gradients, and compare preference objectives and trainer contracts.
- 06
Quantization — INT8, INT4, GPTQ, AWQ
Compute quantization error and metadata costs, distinguish compression algorithms and kernels, and design a complete local-artifact benchmark.
- 07
Speculative Decoding — Acceptance, Residual Sampling and Cost
Prove speculative sampling, implement bounded greedy and stochastic decoders, and compare proposal families with a measured runtime cost model.
- 08
Latency vs Throughput — Little's Law, TTFT, ITL, and Load Testing
Define honest streaming metrics, derive capacity and test SLO measurement with bounded load-test references.
- 09
Multimodal — CLIP → LLaVA
Derive contrastive learning and visual-token interfaces, train a synthetic dual encoder and design a complete multimodal evaluation.
- 10
Agents and Tool Use
Build a bounded tool controller with validated proposals, independent authorization, recovery and task-level evaluation.
- 11
Mixture of Experts (MoE) Intuition
Implement sparse expert routing, derive parameter and memory accounting, and test load-balancing and dispatch assumptions.
- 12
Long Context — RoPE Scaling, YaRN, ALiBi
Derive RoPE extension, YaRN scaling and ALiBi, then evaluate long-context systems without confusing acceptance with recall.
- 13
Red-Team + Safety Eval
Design controlled safety tests, separate harm from refusal, and connect system controls to evidence-based release decisions.