Learning path / 23 published lessons
Data Platform & Orchestration
Trace data from relational queries and storage through batch processing, streams, orchestration and lakehouse maintenance. Focus on grain, state, late data and repeatable recovery.
What you’ll work toward
- A data platform is a chain of independently testable contracts
- A pipeline has green task statuses but a dashboard double-counts sales. Which contract should be inspected first?
Completion is stored on this device only. Nothing is locked; start where it makes sense.
Start this pathBefore the first lesson
- Basic programming and the preceding concepts in this study sequence
These are the starting lesson’s prerequisites, not requirements for every advanced topic below.
How to practise this subject
Start with a tiny dataset containing duplicates, NULLs and late corrections. Predict the output, rerun a partition, and account for exactly which effects are idempotent.
- 01
The Relational Model — Tables, Keys, Normalisation
Design keys, dependencies and historical facts; normalise worked schemas and test real SQLite constraints.
- 02
SQL Basics — SELECT, WHERE, ORDER BY, LIMIT
Query rows safely with explicit NULL, alias, ordering and pagination semantics, verified against exact SQLite results.
- 03
Joins — INNER, LEFT, RIGHT, FULL, Anti-Join
Choose join semantics, preserve measure grain and compare physical algorithms without fan-out or NULL mistakes.
- 04
Aggregations — GROUP BY, HAVING, Subqueries
Compute honest grouped metrics, exact medians and weighted rollups; distinguish missing populations and approximation contracts.
- 05
Window Functions — the Game-Changer
Rank deterministically, reason about peers and frames, and compute calendar-aware deltas and session boundaries.
- 06
CTEs & Recursive Queries
Compose named SQL stages and safely traverse graphs with explicit cycle, reachability and truncation semantics.
- 07
Indexes — B-Tree Intuition, When to Add
Design and validate query-driven indexes; distinguish access paths, output work, visibility and concurrent-build safety.
- 08
Transactions & ACID — Isolation Levels, MVCC
Protect transaction invariants with precise isolation, retry and durability contracts; test local rollback and bounded anomaly models.
- 09
Query Planning — EXPLAIN, Execution Plans, Tuning
A query returning ten rows can still scan millions, build a large hash table, or repeat a cheap lookup thousands of times. An index is one possible remedy, not a diagnosis. The objective is to read PostgreSQL 16 plans, distinguish
- 10
NoSQL Landscape — KV, Document, Column, Graph
“NoSQL” is an umbrella, not a transaction or performance contract. Key-value, document, wide-column and graph describe different ways to organize access; products often span families. The goal is to explain all four, model Cassand
- 11
Data Modelling — Dimensional, Data Vault, OBT
A warehouse needs an agreed meaning for one row before it needs a fashionable schema. This lesson designs a sales star, distinguishes Kimball, Inmon, Data Vault and one-big-table (OBT) approaches, implements historical customer at
- 12
Batch vs Streaming — Mental Model & Use Cases
“Real time” is a requirement to quantify, not an engine to buy. The goal is to classify workloads by input boundedness, freshness and repair semantics, then justify batch, micro-batch or record-at-a-time processing with explicit a
- 13
Spark — RDD, DataFrame, Jobs/Stages/Shuffles
Spark tuning begins with an execution model, not “add executors.” This lesson traces DataFrame expressions through logical/physical plans, jobs, stages, tasks and exchanges, then reasons about skew, broadcast, caching and small fi
- 14
Kafka — Topics, Partitions, Consumer Groups
Kafka is a partitioned retained log with independent readers. Treating offset commits as deletion, replication as invulnerability, or a transaction flag as protection for every downstream effect causes design errors. This lesson t
- 15
Stream Processing — Watermarks, Windows, Exactly-Once
A loop reading Kafka is not yet a correct stateful processor. This lesson defines time, window membership, progress, emission, state cleanup and sink effects separately. It targets **Spark 3.5.7** and **Flink 1.18** where their co
- 16
Orchestration — Airflow, DAGs, Retries, Backfills
Orchestration coordinates dependencies, scheduling, retries and historical runs. It does not make arbitrary writes atomic or correct. This lesson targets **Airflow 2.10.5** and preserves the full daily-order pipeline, sensor, retr
- 17
dbt — Models, Tests, Docs, Warehouse-Native ELT
Structure staging, intermediate and mart models with explicit grain; Use ref, source, macros and generic/singular test contracts; Choose materialisations and reconcile late changes, updates and deletes.
- 18
Lakehouse — Delta / Iceberg / Hudi, ACID on Files
Explain atomic snapshot membership and validated commit retries; Trace Delta versions and distinguish compaction from retention; Compare Delta, Iceberg and Hudi compatibility and maintenance.
- 19
Data Quality — Freshness, Volume, Schema, Distribution
Check eligible freshness, volume, schema, validity and distribution; Distinguish unknown checks from passing data and route by consequence; Design consumer-specific publication gates and seasonal baselines.
- 20
Governance & Cost — Lineage, PII, Attribution
Trace column dependencies and enforce classification separately; Reconcile direct, shared and unallocated cost against a declared total; Design justified retention and verifiable authorised erasure workflows.
- 21
Apache Airflow - open source orchestration engine
Design retry-safe Airflow tasks without mistaking DAG ordering for a database transaction. A source-bounded lesson with worked reasoning, failure analysis and explicit runtime limitations.
- 22
Data Infrastructure for AI and Experimentation at Scale
Trace events, time-valid features, recommendation decisions and experiment outcomes through a source-reviewed data platform with executable SQL and statistical fixtures.
- 23
Overall Engineering Clarity — Data, Distributed Systems and AI (Deep Dive)
Trace compute, commit, indexing and authorization boundaries across a data/AI platform. Mechanisms, worked examples, failure analysis and complete practice answers.