Data Platform & Orchestration
A data platform is a chain of independently testable contracts. A source-bounded lesson with worked reasoning, failure analysis and explicit runtime limitations.
Editorial review: · What review means
Stored in this browser only. No account, no sync. Clearing browser data removes your record.
By the end, you should be able to
- A data platform is a chain of independently testable contracts
- A pipeline has green task statuses but a dashboard double-counts sales. Which contract should be inspected first?
Bring with you
- Basic programming and the preceding concepts in this study sequence
Listen to this article
Browser / device speech · no paid TTS integration. Voice quality depends on your device.
Choose a local device voice to avoid a remote speech service. This site adds no TTS service, account or API calls.
Checking browser speech support…
Pause saves your segment; resume repeats that short segment. Changing voice or speed pauses playback. Stop resets to the beginning. Progress counts finished text segments, not audio time. Leaving or hiding this page stops or pauses speech.
What gets read aloud?
Reads the article body as it appears when you press Listen. Navigation, controls and closed sections are skipped. Expand a section, then Stop and Listen to include it. Code and equations get brief notices; figures use available labels or captions, not their visual details. This narration does not teach omitted mathematics or replace reading examples on the page.
For better sound at no added site cost, try installed English voices, including enhanced voices offered by your device. We cannot guarantee a best voice on every browser. Use Stop or your device’s audio controls if its speech engine misbehaves.
In this article · 15 sections
A data platform is a chain of independently testable contracts
The event spine preserves identity and replay position; orchestration schedules retryable transformations; SQL models define grain and business meaning; serving projections expose data under freshness and access contracts. None of these responsibilities disappears merely because all components share a cloud vendor. Begin with one source and one consumer, then add tools only when a concrete contract requires them.
A useful sequence is relational/query foundations, batch/stream semantics, Kafka, Spark, Airflow, dbt, table formats, quality and governance. Use the SQL foundations to establish grain and invariants, then connect the same data through ingestion, transformation and serving. For each transition record input version, output identity, replay behavior and ownership. A conceptual vendor-backed design is valid learning; a green local fixture is narrower evidence than a deployed pipeline. The links below lead to both forms and state that distinction.
Failure exercise and worked answer
A pipeline has green task statuses but a dashboard double-counts sales. Which contract should be inspected first?
Compare the invariant, failure and recovery
The data grain, join multiplicities and idempotency/replay semantics. Successful execution does not establish business correctness.
Who this is for
You can write basic SQL and Python and want to reason about scheduled runs, retries, replay, late data and another team's consumers. A local pipeline that returns without error has not yet demonstrated stable business meaning. The prerequisite is being able to name the input row and expected output, not having already operated a cloud cluster.
The path
The original five-part composition route remains useful. Read its short executable boundaries first, then use the broader foundations below when a question exposes a missing concept. Reading duration varies; the map does not promise mastery in a fixed number of minutes or that every linked article has automatic previous/next navigation.
Part 1 · The event spine
Kafka replay and offsets separates a log position from a downstream effect. Follow with topics, partitions and consumer groups for the operating model. Produce a crash table for effect-before-offset and offset-before-effect. Partition keys determine the ordering/distribution contract; a consumer group alone cannot prevent a duplicate database effect. Increasing partition count is an operational change with routing/order consequences, not an absolutely irreversible mathematical choice.
Part 2 · Orchestration
Airflow architecture and TaskFlow connects parsing, dependencies, data intervals, retries and output publication. The orchestration lesson extends the operating decisions. Deliver a task contract naming its input partition, stable output identity, validation and retry behavior. DAG ordering is not a database transaction; the task must make its external effect retry-safe.
Part 3 · The query layer
ADX and KQL supplies the original telemetry-query branch; the time-window lab contrasts aligned bins with rolling windows. Write the intended grain and denominator before interpreting a dashboard. A fast query with a multiplied join or omitted late events is still wrong. ADX service latency is an environment- and query-dependent measurement, not a promise that a particular partition setting yields one-second results. These are related query lessons, not claims that SQL and KQL are identical.
Part 4 · Composition
Data infrastructure for AI and experimentation connects ingestion, transformations, feature history and measurements. Trace one event from trusted identity to a prediction or experiment outcome. Record which timestamps represent occurrence, availability and measurement maturity, and how a correction propagates. A green orchestration run cannot establish causal business lift or leakage-free training by itself.
Part 5 · The wider view
Engineering clarity is a cross-stack reference, not a prerequisite to read linearly in one sitting. Bring one concrete invariant—for example, “a replay must not double-count accepted usage”—and compare its database, queue, cache and application failure boundaries. Record unanswered assumptions rather than treating a diagram or vendor list as an implementation guarantee.
Foundation and operating branches
- Apache Airflow - open source orchestration engine — Design retry-safe external effects; a DAG is not an external transaction.
- Data Infrastructure for AI & Experimentation at Scale — An AI platform review must separate prediction, identity and causality.
- Overall Engineering Clarity — Data, Distributed Systems and AI (Deep Dive) — A cross-stack mental model with explicit boundaries.
- The Relational Model — Tables, Keys, Normalisation — Dependencies, keys and the fact being represented.
- SQL Basics — SELECT, WHERE, ORDER BY, LIMIT — Logical query order is not a physical execution promise.
- Joins — INNER, LEFT, RIGHT, FULL, Anti-Join — Join grain before join algorithm.
- Aggregations — GROUP BY, HAVING, Subqueries — Aggregation denominators and empty populations.
- Window Functions — the Game-Changer — Peers, frames and deterministic ranking.
- CTEs & Recursive Queries — Recursive traversal needs cycle semantics, not only a depth cap.
- Indexes — B-Tree Intuition, When to Add — Index cost includes visibility, output and maintenance.
- Transactions & ACID — Isolation Levels, MVCC — Isolation protects invariants only inside its defined scope.
- Query Planning — EXPLAIN, Execution Plans, Tuning — Read measurements without turning EXPLAIN into a hazard.
- NoSQL Landscape — KV, Document, Column, Graph — NoSQL names do not specify a transaction contract.
- Data Modelling — Dimensional, Data Vault, OBT — Declare fact grain before dimensions or dashboard columns.
- Batch vs Streaming — Mental Model & Use Cases — Choose freshness and repair semantics before an engine.
- Spark — RDD, DataFrame, Jobs/Stages/Shuffles — Separate logical operations from physical exchanges.
- Kafka — Topics, Partitions, Consumer Groups — Offsets, acknowledgements and application effects.
- Stream Processing — Watermarks, Windows, Exactly-Once — Watermarks and end-to-end effect boundaries.
- Orchestration — Airflow, DAGs, Retries, Backfills — Design retry-safe external effects; a DAG is not an external transaction.
- dbt — Models, Tests, Docs, Warehouse-Native ELT — Incremental is a correctness policy, not merely a speed flag.
- Lakehouse — Delta / Iceberg / Hudi, ACID on Files — Atomic snapshots and retention are separate contracts.
- Data Quality — Freshness, Volume, Schema, Distribution — Data quality checks need eligible populations and failure policy.
- Governance & Cost — Lineage, PII, Attribution — Lineage assists governance but does not enforce deletion.
How to read this series
Complete the five-part route once with one example event. If joins or aggregates are unclear, work the relational/query branch first. If event order or retry behavior is unclear, work the Kafka/streaming branch before a larger platform diagram. Keep a four-column notebook: input contract, output invariant, failure injection, and evidence. A conceptual design backed by versioned docs is a legitimate result; deployment reproduction is a different claim.
Composition checkpoint and solution
A sale of 12 units is delivered twice. An Airflow run succeeds both times, and a dashboard reports 24 units. Which fix belongs where? Would adding lineage automatically correct it?
Trace identity before adding another platform service
Keep the same sale identity on transport replay. At the authoritative sink, atomically accept that identity and its effect once, or rebuild an idempotent partition from a stable source snapshot. Check whether joins multiply the one accepted sale before aggregation. Lineage can help locate dependencies and affected outputs but does not enforce uniqueness or repair the metric. Reconcile the result against the accepted event ledger and rerun with the same inputs.
What this series does not cover
The route links Spark execution and dimensional modelling rather than excluding those topics from the broader curriculum. It is not an exhaustive database/vendor comparison, a cloud setup transcript or an operational certification. Cross-lane KQL/ADX lessons retain their own version/runtime boundaries; this hub does not certify every linked article's deployment.
How to study
Work the derivation before selecting products. State one failure scenario, a recovery policy and the test boundary for every design. Cloud execution is not required to understand a documented mechanism; it is required before claiming that deployment was reproduced.
Sources and review boundary
Primary passages were checked on 2026-10-07. Versioned sources below delimit the claims; mutable pages are captured in the corrective evidence report. No cluster, cloud service, load benchmark, provider send or model inference was executed for this review.
- Apache Kafka 4.1 design — Exactly-once effects outside Kafka require cooperation with the destination system.
- Airflow 2.10.5 best practices — Retry-safe tasks use stable partitions and idempotent writes, not latest/now inputs.
- PostgreSQL 16 constraints — CHECK does not reject NULL automatically; separate constraints enforce required data.
- OpenLineage object model — OpenLineage separates runtime run events from static job/dataset lineage metadata.
Pause / Recall / Apply
Can you explain it without the page?
Close the example. Reconstruct the core idea, then change one assumption. Mark complete when you’re ready; you can always undo it.
Stored in this browser only. No account, no sync. Clearing browser data removes your record.