Dinesh’sLearning Lab
← All learning paths

Learning path / 23 published lessons

Data Platform & Orchestration

Trace data from relational queries and storage through batch processing, streams, orchestration and lakehouse maintenance. Focus on grain, state, late data and repeatable recovery.

What you’ll work toward

  • A data platform is a chain of independently testable contracts
  • A pipeline has green task statuses but a dashboard double-counts sales. Which contract should be inspected first?
Loading this browser’s progress…

Completion is stored on this device only. Nothing is locked; start where it makes sense.

Start this path

Before the first lesson

  • Basic programming and the preceding concepts in this study sequence

These are the starting lesson’s prerequisites, not requirements for every advanced topic below.

How to practise this subject

Start with a tiny dataset containing duplicates, NULLs and late corrections. Predict the output, rerun a partition, and account for exactly which effects are idempotent.

  1. 01beginner · 18 min

    The Relational Model — Tables, Keys, Normalisation

    Design keys, dependencies and historical facts; normalise worked schemas and test real SQLite constraints.

  2. 02beginner · 16 min

    SQL Basics — SELECT, WHERE, ORDER BY, LIMIT

    Query rows safely with explicit NULL, alias, ordering and pagination semantics, verified against exact SQLite results.

  3. 03beginner · 17 min

    Joins — INNER, LEFT, RIGHT, FULL, Anti-Join

    Choose join semantics, preserve measure grain and compare physical algorithms without fan-out or NULL mistakes.

  4. 04beginner · 16 min

    Aggregations — GROUP BY, HAVING, Subqueries

    Compute honest grouped metrics, exact medians and weighted rollups; distinguish missing populations and approximation contracts.

  5. 05beginner · 17 min

    Window Functions — the Game-Changer

    Rank deterministically, reason about peers and frames, and compute calendar-aware deltas and session boundaries.

  6. 06beginner · 17 min

    CTEs & Recursive Queries

    Compose named SQL stages and safely traverse graphs with explicit cycle, reachability and truncation semantics.

  7. 07beginner · 16 min

    Indexes — B-Tree Intuition, When to Add

    Design and validate query-driven indexes; distinguish access paths, output work, visibility and concurrent-build safety.

  8. 08intermediate · 18 min

    Transactions & ACID — Isolation Levels, MVCC

    Protect transaction invariants with precise isolation, retry and durability contracts; test local rollback and bounded anomaly models.

  9. 09intermediate · 20 min

    Query Planning — EXPLAIN, Execution Plans, Tuning

    A query returning ten rows can still scan millions, build a large hash table, or repeat a cheap lookup thousands of times. An index is one possible remedy, not a diagnosis. The objective is to read PostgreSQL 16 plans, distinguish

  10. 10intermediate · 20 min

    NoSQL Landscape — KV, Document, Column, Graph

    “NoSQL” is an umbrella, not a transaction or performance contract. Key-value, document, wide-column and graph describe different ways to organize access; products often span families. The goal is to explain all four, model Cassand

  11. 11intermediate · 20 min

    Data Modelling — Dimensional, Data Vault, OBT

    A warehouse needs an agreed meaning for one row before it needs a fashionable schema. This lesson designs a sales star, distinguishes Kimball, Inmon, Data Vault and one-big-table (OBT) approaches, implements historical customer at

  12. 12intermediate · 20 min

    Batch vs Streaming — Mental Model & Use Cases

    “Real time” is a requirement to quantify, not an engine to buy. The goal is to classify workloads by input boundedness, freshness and repair semantics, then justify batch, micro-batch or record-at-a-time processing with explicit a

  13. 13intermediate · 20 min

    Spark — RDD, DataFrame, Jobs/Stages/Shuffles

    Spark tuning begins with an execution model, not “add executors.” This lesson traces DataFrame expressions through logical/physical plans, jobs, stages, tasks and exchanges, then reasons about skew, broadcast, caching and small fi

  14. 14intermediate · 20 min

    Kafka — Topics, Partitions, Consumer Groups

    Kafka is a partitioned retained log with independent readers. Treating offset commits as deletion, replication as invulnerability, or a transaction flag as protection for every downstream effect causes design errors. This lesson t

  15. 15intermediate · 20 min

    Stream Processing — Watermarks, Windows, Exactly-Once

    A loop reading Kafka is not yet a correct stateful processor. This lesson defines time, window membership, progress, emission, state cleanup and sink effects separately. It targets **Spark 3.5.7** and **Flink 1.18** where their co

  16. 16intermediate · 20 min

    Orchestration — Airflow, DAGs, Retries, Backfills

    Orchestration coordinates dependencies, scheduling, retries and historical runs. It does not make arbitrary writes atomic or correct. This lesson targets **Airflow 2.10.5** and preserves the full daily-order pipeline, sensor, retr

  17. 17intermediate · 20 min

    dbt — Models, Tests, Docs, Warehouse-Native ELT

    Structure staging, intermediate and mart models with explicit grain; Use ref, source, macros and generic/singular test contracts; Choose materialisations and reconcile late changes, updates and deletes.

  18. 18intermediate · 20 min

    Lakehouse — Delta / Iceberg / Hudi, ACID on Files

    Explain atomic snapshot membership and validated commit retries; Trace Delta versions and distinguish compaction from retention; Compare Delta, Iceberg and Hudi compatibility and maintenance.

  19. 19intermediate · 20 min

    Data Quality — Freshness, Volume, Schema, Distribution

    Check eligible freshness, volume, schema, validity and distribution; Distinguish unknown checks from passing data and route by consequence; Design consumer-specific publication gates and seasonal baselines.

  20. 20intermediate · 20 min

    Governance & Cost — Lineage, PII, Attribution

    Trace column dependencies and enforce classification separately; Reconcile direct, shared and unallocated cost against a declared total; Design justified retention and verifiable authorised erasure workflows.

  21. 21intermediate · 19 min

    Apache Airflow - open source orchestration engine

    Design retry-safe Airflow tasks without mistaking DAG ordering for a database transaction. A source-bounded lesson with worked reasoning, failure analysis and explicit runtime limitations.

  22. 22advanced · 55 min

    Data Infrastructure for AI and Experimentation at Scale

    Trace events, time-valid features, recommendation decisions and experiment outcomes through a source-reviewed data platform with executable SQL and statistical fixtures.

  23. 23advanced · 119 min

    Overall Engineering Clarity — Data, Distributed Systems and AI (Deep Dive)

    Trace compute, commit, indexing and authorization boundaries across a data/AI platform. Mechanisms, worked examples, failure analysis and complete practice answers.