Series · 5 parts
Data Platform & Orchestration
The plumbing under every ML system: streaming with Kafka, orchestration with Airflow, fast analytics with Kusto, and the experimentation infrastructure that ties it together — ending with a broad engineering-clarity deep dive.
- 1Kafka 101 for ML engineersTopics, partitions, consumer groups — the parts of Kafka that actually matter when you put ML features behind it.
- 2Apache Airflow - open source orchestration engineArchitecture of Apache Airflow, how DAGs help design complex flows and dependencies, and how we can leverage Apache airflow to train a ML Model and monitor.
- 3Exploring Azure Data Explorer and Best PracticesA self-sufficient deep-dive on Azure Data Explorer (ADX/Kusto) — architecture, the KQL language from zero to advanced, ingestion patterns, performance/cost levers, and operational best practices.
- 4Data Infrastructure for AI & Experimentation at ScaleA comprehensive deep-dive into the data backbone powering ML, personalization, experimentation, and GenAI on modern streaming platforms
- 5Overall Engineering Clarity — Data, Distributed Systems and AI (Deep Dive)A long-form, primary study companion. Internals, flows, decision trees, code, and Q&A with reasoning across Spark, lakehouse, graphs, search, LLMs, RAG/agents, distributed HLD, governance, modeling, SQL, JVM, Python, K8s and CI/CD.