Learning path / 11 published lessons
Systems laboratory: make the failure observable
A practical route from capacity estimates to replay-safe data, permission-aware retrieval and evidence-based release gates.
What you’ll work toward
- Turn a service promise into a measurable contract
- Follow a request and its evidence through nine failure boundaries
- Produce an inspectable capstone rather than a diagram-only design
Completion is stored on this device only. Nothing is locked; start where it makes sense.
Start this pathBefore the first lesson
- Basic arithmetic and service request paths
These are the starting lesson’s prerequisites, not requirements for every advanced topic below.
How to practise this subject
Attempt each lesson’s exercises before opening the explanation. Reconstruct its main example, change an assumption, and use the stated test boundaries to judge what you have actually checked.
- 01
Capacity estimates that survive an overload question
Derive traffic, storage, bandwidth, working set and backlog with explicit units, sensitivity analysis and exact retained laboratory programs.
- 02
Token buckets: bursts, clocks and the atomic boundary
Implement and test one token bucket, then identify what changes when several servers enforce the same quota.
- 03
Replay is a feature; duplicate effects are a design choice
Choose partition keys and independent groups for hot, warm and cold ML pipelines. Mechanisms, worked examples, failure analysis and complete practice answers.
- 04
A metering ledger that survives a replay
Put event identity and usage aggregation in one transaction, then test rollback, duplicates and conflicting payloads.
- 05
Kusto and KQL — pipelines, safe queries and honest windows
Read KQL pipelines, reason about joins and aggregation, parameterize safe telemetry queries, and distinguish fixed bins from rolling windows.
- 06
Historical features need two clocks
A point-in-time join lab that rejects future events and late backfills while preserving explicit TTL semantics.
- 07
Recommendation quality starts before ranking
An executable retrieve-score-rerank pipeline that exposes candidate recall and diversity trade-offs.
- 08
RAG begins with an evidence boundary
Build permission-filtered retrieval and a deterministic evidence packet before asking a model to synthesize an answer.
- 09
Release gates need denominators, not reassuring scores
Combine deterministic checks, calibrated model judges, multimodal evidence and statistically explicit release gates without replacing human accountability.
- 10
Kimi K3: architecture, derivations and a pinned-source audit
Derive KDA state updates, depth attention and latent expert routing; reconcile pinned Kimi architecture claims and execution limits.
- 11
Capstone: an evidence service with replay-safe usage
Compose tenant-scoped retrieval, explicit abstention and a durable-event invariant into one executable failure lab.