Learning path / 26 published lessons
Language model engineering
Connect language-model mechanics to retrieval, evaluation, tool use, adaptation and serving. Treat data boundaries, failure analysis and measurable outcomes as part of the design.
What you’ll work toward
- Distinguish byte, character, word and subword coverage
- Hand-count BPE pairs and implement reversible ranked byte merges
- Separate WordPiece inference, unigram scoring and SentencePiece configuration
- Test Unicode, whitespace, unknown IDs and template boundaries
Completion is stored on this device only. Nothing is locked; start where it makes sense.
Start this pathBefore the first lesson
- Python and arrays; preceding topics in the learning progression
These are the starting lesson’s prerequisites, not requirements for every advanced topic below.
How to practise this subject
Define a testable task and a failure case. Separate model output from authorized actions, examine retrieval evidence, and specify how quality and latency would be measured.
- 01
Tokenization — BPE, WordPiece, SentencePiece
Tokenization — BPE, WordPiece, SentencePiece: mechanisms, worked examples, assumptions and source-backed corrections with explicit experiment boundaries.
- 02
Attention Intuition — Sequence Bottlenecks and Soft Retrieval
Attention Intuition — Sequence Bottlenecks and Soft Retrieval: mechanisms, worked examples, assumptions and source-backed corrections with explicit experiment boundaries.
- 03
Q/K/V Math — Scaled Dot-Product Attention Derived
Q/K/V Math — Scaled Dot-Product Attention Derived: mechanisms, worked examples, assumptions and source-backed corrections with explicit experiment boundaries.
- 04
Multi-Head Attention — Parallel Views
Multi-Head Attention — Parallel Views: mechanisms, worked examples, assumptions and source-backed corrections with explicit experiment boundaries.
- 05
Positional Encoding — Sinusoidal, Learned, RoPE
Prove attention equivariance, derive sinusoidal and rotary position geometry, and test offsets without confusing computable positions with length generalization.
- 06
Full Transformer Architecture — Encoder + Decoder
Assemble causal and encoder-decoder Transformers, derive residual and normalization behavior, count parameters, and train a tiny CPU character fixture.
- 07
Encoder (BERT), Decoder (GPT), Enc-Dec (T5) — When Each
Separate Transformer families, pretraining losses and output heads; work through BERT corruption and T5 sentinels and inspect local-only inference interfaces.
- 08
LLM Sampling — Greedy, Beam, Top-k, Top-p, Temperature
Derive temperature, nucleus truncation and search; implement edge-case-safe CPU sampling and translate policies into versioned API contracts.
- 09
Scaling Laws — Chinchilla, Compute-Optimal Training
Derive budget-conserving language-model allocations, reconcile Chinchilla and Kaplan, and compare training-only and lifetime-cost objectives.
- 10
Efficient Attention — Flash, Sparse, Linear
Derive online softmax, distinguish exact execution from sparse and kernel operators, and verify CPU parity while specifying a separate GPU-backend proof.
- 11
Prompting — Zero-Shot, Few-Shot, Chain-of-Thought, ReAct
Design testable prompts, bounded tool controllers and evaluation contracts; distinguish published reasoning results from offline fixtures and generated explanations.
- 12
RAG I — Chunking Strategies & Indexing
Build a versioned RAG index with source ranges, tokenizer budgets, stale-chunk removal and access controls; test a complete local lexical-index lifecycle.
- 13
RAG II — Retrieval, Hybrid Search, Reranking
RAG II — Retrieval, Hybrid Search, Reranking: mechanisms, worked examples, assumptions and source-backed corrections with explicit experiment boundaries.
- 14
Vector Databases — pgvector, HNSW, IVF
Vector Databases — pgvector, HNSW, IVF: mechanisms, worked examples, assumptions and source-backed corrections with explicit experiment boundaries.
- 15
LLM Agents — Function Calling, Tools, Planning
LLM Agents — Function Calling, Tools, Planning: mechanisms, worked examples, assumptions and source-backed corrections with explicit experiment boundaries.
- 16
Multi-Agent Orchestration — LangGraph, CrewAI Patterns
Multi-Agent Orchestration — LangGraph, CrewAI Patterns: mechanisms, worked examples, assumptions and source-backed corrections with explicit experiment boundaries.
- 17
LLM Evaluation — LLM-as-Judge, RAGAS, Golden Sets
LLM Evaluation — LLM-as-Judge, RAGAS, Golden Sets: mechanisms, worked examples, assumptions and source-backed corrections with explicit experiment boundaries.
- 18
Fine-Tuning — LoRA, QLoRA, PEFT, When NOT to Fine-Tune
Fine-Tuning — LoRA, QLoRA, PEFT, When NOT to Fine-Tune: mechanisms, worked examples, assumptions and source-backed corrections with explicit experiment boundaries.
- 19
LLM Serving — KV Cache, Batching, Speculative Decoding
LLM Serving — KV Cache, Batching, Speculative Decoding: mechanisms, worked examples, assumptions and source-backed corrections with explicit experiment boundaries.
- 20
Multimodal LLMs — CLIP, VLMs, Audio, Video
Multimodal LLMs — CLIP, VLMs, Audio, Video: mechanisms, worked examples, assumptions and source-backed corrections with explicit experiment boundaries.
- 21
Google and Kaggle Gen AI study notes — the opening foundations
An opening-reading notebook, not a completed five-day course: distinguish recurrence, attention, autoregressive generation, prompting, retrieval and fine-tuning.
- 22
A deep-study prompt with examples, evidence and an audio version
Build a bounded study request, keep sources separate from narration, and evaluate generated explanations instead of mistaking fluent detail for verified knowledge.
- 23
How transformers actually attend — a weighted lookup you can calculate
Follow queries, keys and values through a numeric attention head, test causal masking, and separate multi-head capacity from guaranteed interpretability or long-context quality.
- 24
Image classification: from a scalar model to a tested experiment contract
Derive loss and softmax, plan controlled Fashion-MNIST experiments, and follow complete local TensorFlow references for training, callbacks and save/load.
- 25
Reading MEDEC — detection, localization and correction are different tests
Audit the MEDEC benchmark's data construction, exact reported scores, prompt format and metric limitations without treating a clinical NLP benchmark as deployment validation.
- 26
Stanford CS229: Machine Learning Course
A source-reviewed CS229 orientation with worked linear algebra, probability, optimization and runnable algorithm fixtures.