Dinesh’sLearning Lab
← All learning paths

Learning path / 67 published lessons

Deep Learning and LLMs: an inspectable learning laboratory

A broad deep-learning and LLM progression with a verified NumPy spine, advanced framework and systems lessons, source evidence and explicit execution boundaries.

What you’ll work toward

  • Derive and test core neural computations
  • Navigate all 80 deep-learning topics and the classical ML branch
  • Distinguish source-backed concepts, executed CPU examples and unexecuted GPU experiments
Loading this browser’s progress…

Completion is stored on this device only. Nothing is locked; start where it makes sense.

Start this path

Before the first lesson

  • Python functions, loops and dictionaries; basic algebra

These are the starting lesson’s prerequisites, not requirements for every advanced topic below.

How to practise this subject

Attempt each lesson’s exercises before opening the explanation. Reconstruct its main example, change an assumption, and use the stated test boundaries to judge what you have actually checked.

  1. 01beginner · 33 min

    Arrays, shapes, and the first linear model

    Master array ownership and axes, then audit attention, context and quantization claims with exact CPU checks.

  2. 02beginner · 18 min

    Vectors, Matrices, and the Geometry of Data

    Derive geometry, matrix products, SVD and retrieval scores while testing the limits of each analogy.

  3. 03beginner · 15 min

    Broadcasting — The Rule That Runs the World

    Predict shape sharing, memory costs and numerical failure cases in NumPy and transformer patterns.

  4. 04beginner · 19 min

    Derivatives are local measurements; backprop is bookkeeping

    Rebuild the chain rule from local sensitivities through scalar autograd, matrix pullbacks and checkpointing.

  5. 05beginner · 18 min

    Probability, cross-entropy, and stable logits

    Derive information measures and stable probability losses, then distinguish modern language and preference objectives.

  6. 06beginner · 17 min

    Optimization: steps, momentum, AdamW, and honest experiments

    Trace optimizer updates and numerical safeguards without mistaking toy convergence for universal model performance.

  7. 07beginner · 17 min

    One neuron: logistic regression, stable loss, and its boundary

    Train and audit a sigmoid neuron, verify exact framework parity, and prove and repair its XOR limitation.

  8. 08beginner · 17 min

    Activation Functions — sigmoid, tanh, ReLU, GELU, SwiGLU, xIELU

    Derive and test activation functions, full gradient chains, gated parameter budgets and source-corrected xIELU.

  9. 09beginner · 20 min

    The Forward Pass: Shapes, XOR and Parameter Counts

    Compute an MLP by hand, construct a strict-margin XOR classifier, and distinguish actual architecture counts from universal claims.

  10. 10intermediate · 20 min

    Backprop by Hand: Nine Gradients and a Verified XOR Network

    Derive all nine ReLU-network gradients, preserve the tanh XOR experiment, and verify finite differences against CPU PyTorch.

  11. 11beginner · 20 min

    MLP on MNIST: A Complete Local-Only NumPy Recipe

    Load existing local IDX files, implement stable multiclass loss and backpropagation, and separate a synthetic contract from real digit evaluation.

  12. 12intermediate · 20 min

    Micrograd: Build and Test a Scalar Autograd Engine

    Implement complete scalar reverse-mode differentiation, shared-edge accumulation, stable cross entropy and a four-point neural network.

  13. 13intermediate · 20 min

    Vectorization: Measure Speed, Shapes and Memory

    Compare exact scalar and array programs, derive einsum contractions and FLOP counts, and measure performance without promised speedup factors.

  14. 14beginner · 20 min

    PyTorch Tensors and Autograd: Leaves, VJPs and Safe Updates

    Port NumPy mathematics to raw PyTorch tensors and verify gradient accumulation, cotangents, graph lifetime and mutation controls.

  15. 15beginner · 20 min

    nn.Module: Registration, Composition and Checkpoint Contracts

    Structure models with registered parameters and buffers, distinguish eval from no-grad, and verify trusted state/config round-trips.

  16. 16beginner · 20 min

    Datasets, DataLoaders and Collate: Coverage Before Throughput

    Test indexed sampling, ragged collation, worker/rank sharding and local stream resume while treating optional cloud adapters as unexecuted.

  17. 17beginner · 16 min

    The Training Loop, Logging, and Reproducibility

    Build and verify a resumable training state machine with RNG, sampler, optimizer, scheduler, EMA and safe checkpoints.

  18. 18intermediate · 15 min

    GPU and Mixed Precision

    Separate device placement, operation-specific mixed precision, loss scaling and measured performance without numerical myths.

  19. 19intermediate · 15 min

    Debugging — nan, inf, and Overfitting a Single Batch

    Diagnose training with exact objective contracts, finite checks, tiny-batch fits, negative controls and profiling.

  20. 20intermediate · 19 min

    Dropout and generalization — masks, scaling and controlled experiments

    Derive dropout moments, preserve leakage-resistant evaluation, test mask variants and choose regularization by controlled experiments.

  21. 21intermediate · 16 min

    Batch Norm, Layer Norm, and RMSNorm — the normalization tour

    Derive normalization axes, affine and buffer contracts, implement BN/LN/RMSNorm and test train/eval and calibration failures.

  22. 22intermediate · 15 min

    Weight Initialization — Xavier, He, and μP

    Derive fan and second-moment initialization, test activation and gradient propagation, and understand bounded width-transfer parameterizations.

  23. 23intermediate · 17 min

    Schedulers — Warmup, Cosine, One-Cycle, and WSD

    Implement and verify warmup, linear, cosine, one-cycle and WSD clocks, resume behavior and fair schedule comparisons.

  24. 24intermediate · 17 min

    Gradient Clipping and Accumulation

    Derive weighted accumulation and global clipping, flush unequal tails and distinguish AMP, DDP and activation-checkpointing contracts.

  25. 25intermediate · 20 min

    Convolution: local shared operators, shapes and gradients

    Compute original convolution examples and general output shapes Implement multi-channel forward and backward accumulation Compare receptive fields, pooling and modern convolution designs.

  26. 26intermediate · 20 min

    CNN architectures: LeNet, AlexNet, residual learning and modern backbones

    Build LeNet-like and ResNet18 models with exact parameter counts Distinguish degradation from vanishing gradients Design fair architecture and local-data comparisons.

  27. 27intermediate · 20 min

    CIFAR-10 training: a complete local pipeline and controlled experiments

    Build a complete local CIFAR training and selection pipeline Test normalization, parameter groups and weighted evaluation Diagnose training failures with controlled ablations.

  28. 28intermediate · 20 min

    Transfer learning: frozen features, fine-tuning and buffer discipline

    Implement local-checkpoint linear and fine-tuning regimes Freeze parameters and BatchNorm buffers deliberately Compare transfer objectives and source-target assumptions.

  29. 29intermediate · 20 min

    Data augmentation: tested mixing, task assumptions and calibration

    Implement mixup and clipped-area CutMix with exact targets Compute calibration and soft-target loss controls Choose task-valid augmentation and controlled ablations.

  30. 30intermediate · 20 min

    Vision Transformers: patch equivalence, position and attention cost

    Prove patch projection output and gradient equivalence Build and test a small complete Vision Transformer Explain position interpolation, parameter counts and quadratic cost.

  31. 31intermediate · 20 min

    Word2Vec: complete skip-gram training, geometry and honest evaluation

    Train complete skip-gram negative sampling on bundled text Derive lookup and binary-loss gradients and sampling contracts Query and evaluate embedding geometry without semantic guarantees.

  32. 32intermediate · 20 min

    RNNs from scratch: complete BPTT, character generation and structured recurrence

    Implement complete NumPy BPTT and verify every coordinate Train and generate with a complete character RNN locally Distinguish clipping, truncation, gates, attention and affine scans.

  33. 33intermediate · 20 min

    LSTM and GRU — Gates, State and Gradient Paths

    An ordinary RNN repeatedly applies a nonlinear state transformation. Its time derivative is a product of **time-varying Jacobians**, not just the recurrent matrix's eigenvalue raised to a power. LSTM introduced a separately carried cell state; the la

  34. 34intermediate · 20 min

    Seq2Seq — A Complete Length-Aware Encoder–Decoder

    A sequence-to-sequence model defines a conditional distribution over a target sequence given a source. In the fixed-state recurrent version, an encoder reads the source and passes its final state to an autoregressive decoder. For an LSTM the interfac

  35. 35intermediate · 20 min

    Why Attention — Query-Dependent Source Access

    A fixed-state encoder-decoder must route all source information through its terminal state. Attention retains the sequence of encoder representations and lets each decoder step retrieve a different weighted context. This removes the requirement that

  36. 36intermediate · 24 min

    Attention is a weighted sum with a precise mask

    Calculate attention numerically, derive its softmax backward pass, and test causal and permutation properties.

  37. 37intermediate · 20 min

    Multi-Head Attention — Shapes, Specialization and KV Budgets

    A single attention head already forms a distribution over many keys; it is not forced to pick exactly one token. Multi-head attention gives several independently parameterized distributions and value projections, then mixes their results. At fixed mo

  38. 38intermediate · 20 min

    Positional Encodings — Sinusoids, Learned Tables and RoPE

    Unmasked self-attention with shared tokenwise projections and no position signal is **permutation-equivariant**: permuting input rows permutes output rows. It is not invariant unless an additional permutation-invariant readout removes row order. Thus

  39. 39advanced · 20 min

    Transformer Blocks — Full Architectures and an Analytic Decoder

    Attention mixes information across token positions. A position-wise feed-forward network transforms features with the **same weights at each position**. Normalization sets a branch's input or output scale, and residual addition carries an explicit sk

  40. 40intermediate · 20 min

    nanoGPT Rebuilt — A Complete Tested Decoder

    The useful exercise in nanoGPT is tracing a compact decoder-only model end to end: configuration, embeddings, fused QKV attention, pre-norm blocks, final norm, tied output projection, optimizer groups and generation. Its public model/train files are

  41. 41intermediate · 20 min

    Training a Decoder on Local Text — Tiny Shakespeare Recipe

    A small character-level decoder is a useful way to inspect tokenization, shifted targets, causal loss, optimizer state and generation. Tiny Shakespeare is one possible **already-local** corpus, not data bundled or downloaded by this audit. This lesso

  42. 42intermediate · 18 min

    KV caching: correct positions, masks, memory and reuse

    Implement and test a complete cached decoder, derive cache memory, and distinguish paging, sharing, compression and eviction.

  43. 43intermediate · 18 min

    RoPE and FlashAttention: relative positions and online softmax

    Derive rotary attention and tiled normalization, test CPU parity, and separate positional extension from GPU kernel performance.

  44. 44intermediate · 18 min

    Byte-pair encoding: merges, Unicode and artifact contracts

    Build and inspect deterministic byte BPE, correct weighted merge traces, preserve Unicode bytes and freeze tokenizer artifacts.

  45. 45intermediate · 17 min

    SentencePiece and unigram: lattices, normalization and compatibility

    Derive unigram marginal likelihood and soft EM, test segmentation, and distinguish tokenizer algorithms from toolkit and normalization choices.

  46. 46intermediate · 20 min

    Pretraining data: auditable extraction, grouping and token shards

    Design a provenance-aware pretraining pipeline with tested filters, grouped splits, boundary-safe windows and validated token shards.

  47. 47advanced · 15 min

    Chinchilla scaling: fit assumptions, budgets and serving cost

    Derive compute-constrained allocation, correct the numerical budget, and separate fixed-ratio heuristics from fitted and inference-aware scaling.

  48. 48intermediate · 16 min

    Data mixing: sampling, exposure, proxy ablations and curriculum

    Implement weighted shard sampling, measure domain exposure, run a complete synthetic ablation and explain DoReMi and token-clock schedules.

  49. 49advanced · 19 min

    Quality and deduplication: MinHash, LSH and measured decisions

    Implement probabilistic near-duplicate retrieval with exact verification, audit learned quality policies and screen contamination with explicit limits.

  50. 50advanced · 18 min

    Architecture choices: derive and test a 100M decoder

    Implement pre-norm, RMSNorm, RoPE, GQA and SwiGLU; reconcile exact model counts, cache budgets and modern architecture claims.

  51. 51advanced · 14 min

    Training a 100M model: complete local recipe and budget

    Build a local-shard training state machine with exact token budgets, validation, persistence and honest resource-dependent execution limits.

  52. 52advanced · 18 min

    Loss curves and evaluation: likelihood, tasks and controls

    Preserve the exact synthetic decoder experiment while restoring token-weighted metrics, downstream evaluation, uncertainty and contamination audits.

  53. 53advanced · 15 min

    Checkpointing and resuming: consistent state and durable publication

    Implement trusted local restart snapshots, test continuation and failures, and derive retention and distributed recovery boundaries.

  54. 54advanced · 14 min

    Pretraining bugs: hypotheses, counterexamples and tests

    Diagnose stalled, nonfinite, drifting or misleading training runs through bounded interventions rather than loss-curve folklore.

  55. 55advanced · 16 min

    Supervised fine-tuning: objective, templates and local recipes

    Derive assistant-only supervision, test masks and conditioning gradients, and specify pinned local SFT with template and regression checks.

  56. 56advanced · 18 min

    Instruction datasets: provenance, normalization and mixtures

    Audit exact dataset identities and rights, reconstruct conversation trees, and derive reproducible example and supervised-token sampling weights.

  57. 57advanced · 16 min

    Adaptation with LoRA: low-rank updates and masked supervision

    Derive low-rank adapter gradients, inject and test a complete layer, and audit versioned SFT, merge and variant workflows.

  58. 58advanced · 15 min

    QLoRA: NF4 storage, gradients and local integration

    Derive block-scale memory, implement a CPU NF4 teaching quantizer, and distinguish adapter gradients, paging and merge behavior.

  59. 59advanced · 20 min

    Reward Modeling: Pairwise Likelihood, Data and Failure Modes

    Derive preference likelihoods, test reward gradients and pooling, and audit data, calibration and reward hacking.

  60. 60advanced · 20 min

    Evaluation and Red-Teaming: Protocols, Uncertainty and Safety

    Build reproducible capability, chat, domain and red-team evaluations with correct metrics, overlap analysis and uncertainty.

  61. 61advanced · 20 min

    KV Cache Deep Dive: Correctness, Memory and Serving Budgets

    Test exact cached decoding and grouped-query attention, derive byte budgets, and distinguish paging, precision and eviction.

  62. 62advanced · 20 min

    Pruning and Distillation: Exact Masks and Correct Token Losses

    Implement exact pruning masks and shifted masked distillation, derive temperature gradients, and evaluate compression tradeoffs.

  63. 63advanced · 20 min

    Continuous Batching: Scheduling, Cache Admission and Capacity

    Trace static and continuous batches, enforce cache admission, derive Little’s law and diagnose latency-throughput tradeoffs.

  64. 64advanced · 20 min

    Inside vLLM and TGI: Paging, Versioned Schedulers and Operations

    Implement a reference-counted block allocator, inspect pinned vLLM internals, and prepare honest serving and observability recipes.

  65. 65advanced · 20 min

    Deploying an LLM API: Bounded FastAPI Streaming and Lifecycle

    Preserve SSE framing, bound admission, authenticate and clean up streaming requests, and design deployment, retry and usage contracts.

  66. 66advanced · 18 min

    RAG as a testable evidence pipeline, not a truth switch

    Implement chunking, exact retrieval and rank fusion, integrate evidence packing and evaluate a complete RAG pipeline.

  67. 67advanced · 19 min

    Capstone: a reproducible model-and-evidence laboratory

    Complete a reproducible CPU laboratory or a fully specified 100M training and release project with explicit evidence gates.