Dinesh’sLearning Lab
← All learning paths

Learning path / 22 published lessons

Machine learning and statistical reasoning

Study how models learn from data: regression and classification, optimization, regularization, trees and evaluation. Follow the split and fitting boundaries that keep an experiment honest.

What you’ll work toward

  • Specify prediction-time features, targets and the population an evaluation measures
  • Design train, validation and test boundaries for independent, grouped and temporal data
  • Explain loss versus metric with numerical examples and baseline comparisons
  • Diagnose preprocessing, target, time, entity and model-selection leakage
Loading this browser’s progress…

Completion is stored on this device only. Nothing is locked; start where it makes sense.

Start this path

Before the first lesson

  • Python functions and NumPy arrays
  • Means, variance and basic probability

These are the starting lesson’s prerequisites, not requirements for every advanced topic below.

How to practise this subject

Choose a baseline and metric before comparing models. Identify leakage, inspect a failure slice, and distinguish a training improvement from evidence of better generalization.

  1. 01intermediate · 24 min

    The ML Mental Model — Features, Labels, Train/Val/Test

    Design an honest supervised-learning experiment, diagnose five leakage mechanisms, compare losses and metrics, and test a small regression pipeline without downloads.

  2. 02intermediate · 20 min

    Linear Regression from Scratch: Solvers, Geometry and Evidence

    Derive least squares, test gradients and projection geometry, recover its history, and compare stable and iterative solvers.

  3. 03intermediate · 15 min

    Logistic Regression — Sigmoid, Cross-Entropy, from Scratch

    Derive linear log-odds, stable binary cross-entropy and its gradient; separate calibration, ranking and validation-selected decisions.

  4. 04intermediate · 15 min

    Regularization — L1, L2, Elastic Net

    Derive ridge, lasso and elastic net; verify normalization, soft thresholding, fold-safe tuning and the AdamW distinction.

  5. 05intermediate · 16 min

    Bias–Variance Trade-off & Learning Curves

    Derive squared-error bias, variance and noise; measure independent refits and bootstrap limits, and interpret learning curves carefully.

  6. 06intermediate · 16 min

    Decision Trees — Gini, Entropy, Splits

    Calculate weighted impurity gains, implement a greedy tree, trace predictions and diagnose leaf-size, pruning and representation limits.

  7. 07intermediate · 16 min

    Random Forest & Bagging

    Build a per-split randomized forest, derive bootstrap coverage and covariance limits, and audit out-of-bag evaluation and feature importance.

  8. 08intermediate · 16 min

    Gradient Boosting — XGBoost, LightGBM

    Derive gradient and regularized Newton tree updates; run bounded XGBoost and LightGBM experiments with validation-controlled stages.

  9. 09intermediate · 16 min

    Evaluation Metrics — P/R/F1/ROC/PR/AUC

    Derive confusion-matrix, ROC and average-precision metrics; select thresholds without test leakage and account for prevalence and capacity.

  10. 10intermediate · 19 min

    Feature Engineering — Encoding, Scaling, Missing

    Build a fold-safe heterogeneous feature pipeline; test encoding, missingness, learned state, target cross-fitting and persistence boundaries.

  11. 11intermediate · 15 min

    Imbalanced Data — SMOTE, Class Weights, Thresholds

    Separate ranking, calibration and threshold decisions under imbalance. Apply weighted learning and fold-local resampling without leakage.

  12. 12intermediate · 15 min

    Model Selection — CV, Hyperparameter Tuning, Optuna

    Choose split units and time boundaries that match deployment. Distinguish nested evaluation from hyperparameter selection.

  13. 13advanced · 15 min

    Perceptron & Activation Functions

    Derive perceptron updates and the XOR impossibility proof. Compare activation derivatives and failure modes.

  14. 14advanced · 15 min

    Multi-Layer Perceptron — Forward Pass

    Trace forward-pass shapes and count every parameter. Explain initialization and approximation assumptions.

  15. 15advanced · 15 min

    Backpropagation — Derived by Hand on a 2-Layer Net

    Derive affine and activation gradients including biases. Check complete NumPy gradients against numerical and autograd references.

  16. 16advanced · 15 min

    Optimizers — SGD, Momentum, Adam, RMSprop

    Implement and compare SGD, momentum, RMSprop and Adam updates. Derive startup correction and decoupled weight decay.

  17. 17advanced · 15 min

    PyTorch Fundamentals — Tensors, Autograd, nn.Module

    Use tensor storage, dtype, device and gradient contracts. Build registered modules and a validation-safe training loop.

  18. 18advanced · 15 min

    Regularization in DL — Dropout, BatchNorm, Weight Decay

    Derive dropout moments and distinguish normalization axes. Explain coupled versus decoupled decay and train-eval state.

  19. 19advanced · 45 min

    CNNs — Convolution, Pooling, ImageNet Architectures

    CNNs — Convolution, Pooling, ImageNet Architectures: mechanisms, worked examples, assumptions and source-backed corrections with explicit experiment boundaries.

  20. 20advanced · 45 min

    RNNs & LSTMs — Sequences & the Vanishing Gradient

    RNNs & LSTMs — Sequences & the Vanishing Gradient: mechanisms, worked examples, assumptions and source-backed corrections with explicit experiment boundaries.

  21. 21advanced · 45 min

    Embeddings — word2vec, GloVe, Contrastive Learning

    Embeddings — word2vec, GloVe, Contrastive Learning: mechanisms, worked examples, assumptions and source-backed corrections with explicit experiment boundaries.

  22. 22advanced · 45 min

    Transfer Learning & Fine-Tuning Classical DL

    Transfer Learning & Fine-Tuning Classical DL: mechanisms, worked examples, assumptions and source-backed corrections with explicit experiment boundaries.