82 articles & lessons
Systems & Infrastructure
Articles and learning notes on systems & infrastructure.
The 6-Month Learning Plan · 130 Sessions
Evergreen 6-month curriculum covering Python → Math → DSA → Databases → Data Engineering → Backend → Systems → Distributed → SRE → Security → Classical ML → Deep Learning → Transformers → LLMs → System Design. Assumes zero background, one concept per session.
25 min readShip It — Deploy, Package, Sustain
Getting things in front of people: local scripts to a live API, Docker on Azure, a Windows app from C++ compilation to the store, and the learning system that keeps you shipping.
7 min readDesign an AI Chat Product — RAG + Agents + Serving
Capstone: designing an AI chat product from the client SSE stream all the way to the LLM serving fleet. Streaming, RAG, agent orchestration, evaluation, safety, cost, and the incidents (ChatGPT title leak, Claude prompt injection, Copilot rate-limit fallout) that shape every mature deployment.
60 min readDesign a Newsfeed / Recommender — Pull vs Push, Ranking
The highest-QPS ML application in the world. Fanout-on-write vs fanout-on-read, the celebrity problem, the two-stage candidate → rank funnel, and the feedback loops that eat naïve recommenders.
55 min readDesign a Chat System — WebSockets, Delivery, Presence
Realtime bidirectional messaging at billions/day scale — WebSockets, routing tables, at-least-once delivery, ordering, presence, offline pushes, and the reconnect thundering herd that took Slack down.
55 min readDesign a URL Shortener — the Classic Warm-Up
Simple on the surface, deep underneath. Unique ID generation, sharding, redirect latency budget, and the anti-abuse story every serious shortener has to solve.
55 min readSystem Design Framework — Reqs, Capacity, HLD, Deep-Dive
The four-step framework every senior reviewer scores against — how to structure a 45-minute system design review, when to push back, when to accept a number, and the anti-patterns that kill senior engineers in the first ten minutes.
55 min readOWASP Top 10, Secrets Mgmt & Threat Modelling
The security bugs everyone ships (and the ones attackers exploit) — walk the OWASP Top 10 for 2021, wire up a real secrets manager, and threat-model a service in 30 minutes.
55 min readTLS 1.3, PKI & Cert Lifecycle
The green padlock, demystified — key exchange, certificate chains, revocation, and why 60 % of production outages the past decade were expired certs.
55 min readAuthN vs AuthZ, Sessions & Password Storage
The two questions every request has to answer: who are you, and what are you allowed to do? Plus how to store passwords without ending up in a HIBP breach dump.
55 min readIncident Response — Runbooks, Postmortems, On-Call
Turning outages into learning. The IMOC roles, the blameless postmortem template, and why the best incident responses feel boring.
55 min readSLIs, SLOs & Error Budgets — the SRE Math
How to measure reliability without lying to yourself. The formulas Google's SRE team uses to decide when to ship features vs when to freeze deploys.
55 min readPrometheus, Grafana, OpenTelemetry — Hands-on
The default open-source observability stack. Scrape, store, query, visualise, and instrument — all with the CNCF tools every serious team uses.
55 min readThe 3 Pillars — Metrics, Logs, Traces
The senses of a running system. What each pillar is good at, what it's terrible at, and why you need all three — not two, not one — to debug production.
55 min readMulti-Region — Active-Passive, Active-Active, Failover
Surviving whole-region outages without lying to your users. The three architectures, the two RTO/RPO knobs, and why 'active-active' is often a marketing term.
55 min readMessage Queues — SQS, RabbitMQ, Kafka as Queue
Decouple producers from consumers so the two never have to be up at the same time. The difference between a queue and a log, and why picking the wrong one is a two-year rewrite.
55 min readSharding & Partitioning Strategies
Split your data before it splits you. Hash vs range vs directory sharding, the hot-partition problem, and why 'just add a shard key' is the wrong answer for 90% of teams.
55 min readConsensus — Paxos & Raft Intuition
How a group of machines that can crash, lag, or lie by omission agree on the same value. The algorithm powering etcd, ZooKeeper, Kafka, CockroachDB, and every serious cluster's brain.
55 min readConsistency Models — Linearizable, Sequential, Eventual
The hierarchy of promises a distributed system makes about what you'll read after you write. Pick the wrong rung and you either lose money or lose latency.
55 min readReplication — Leader/Follower, Multi-Leader, Leaderless
Three ways to keep copies of your data in sync — leader/follower (Postgres, MySQL), multi-leader (active-active), leaderless (Dynamo, Cassandra). Sync vs async, replication lag, and the split-brain problem.
55 min readCAP & PACELC — the Actual Trade-Offs
The three-letter theorem everyone quotes wrong, and the five-letter one that fills in the gap. Consistency, Availability, Partition-tolerance — and what happens the 99.9% of the time your network is fine.
55 min readInfrastructure as Code — Terraform / Bicep Basics
Click-ops doesn't survive contact with production. Terraform + Bicep — providers, state, plans, modules — the ~15 concepts that let you rebuild your whole cloud from a git repo.
55 min readAzure Cloud — Identity, Storage, Networking, App Service
The four pillars of every real Azure workload — Entra ID + RBAC, Storage & Cosmos, VNets & Private Endpoints, App Service & Container Apps. What to pick and why.
55 min readKubernetes II — ConfigMaps, Secrets, HPA, Network Policies
The next four objects you'll use every week — ConfigMaps for tuning, Secrets for credentials, HPA for autoscaling, NetworkPolicies for zero-trust inside the cluster.
55 min readKubernetes I — Pods, Deployments, Services
Kubernetes without the mysticism — Pods run containers, Deployments keep N of them alive, Services give them a stable IP. Ship a real app to a local cluster in 25 minutes.
55 min readDocker — Images, Layers, Dockerfile, Networking
The 90 minutes that end ‘works on my machine’ forever. Layers, images, containers, networking, and the ten Dockerfile lines that separate a 2 GB toy from a 60 MB production image.
55 min readCDN — Edge, Cache Hierarchies, Cache-Control
How your static asset travels 40 ms to Sydney instead of 400 ms — edges, origins, cache hierarchies, and the four Cache-Control directives you'll set every day.
50 min readCaching — Cache-Aside, Write-Through, TTLs, Invalidation
The oldest performance trick in the book, done right. Cache-aside vs write-through vs write-back, TTLs, invalidation patterns, and the two hardest problems in computer science.
55 min readNetworking II — Load Balancers L4 vs L7, Reverse Proxies
How one hostname fans out to a hundred servers without dropping a packet — the L4 vs L7 decision, health checks, sticky sessions, and the reverse-proxy patterns that run every real web system.
50 min readNetworking I — TCP/IP, DNS, Sockets
How bytes actually cross the internet — from getaddrinfo() to the 3-way handshake to the router hop that decides your latency. The mental model every backend engineer needs before they can debug ‘why is this slow?’
55 min readOS Basics — Processes, Threads, Memory, FDs
The four abstractions the OS gives you and why every senior debugging story starts with one of them. Processes, threads, virtual memory, and file descriptors — with real strace output, htop shots, and the incidents each one caused.
55 min readReading Docs & Effective Googling — the Meta-Skill
The single skill that separates a 10× engineer from a 1× engineer: knowing how to find the answer, fast, without asking a human. Man pages, official docs, GitHub source-diving, and search queries that actually work.
45 min readThe Command Line — bash, pipes, grep, jq
Live in the terminal without fear. The pipes, filters, and text-crunching muscle memory that separates an engineer from a button-clicker — and the 20 commands you'll use every single day for the rest of your career.
55 min readGit & GitHub — Commits, Branches, PRs
The 90 minutes that turns Git from a scary black box into a save-point machine you trust with your career. Real commits, real branches, real PRs — the muscle memory every senior engineer runs on.
50 min readDev Environment — Linux/WSL, Terminal, VS Code
The 90 minutes that saves you 90 hours. Real environment, real editor, real terminal — no ‘works on my machine’ for the next six months.
45 min readR26 · Week 26 Recall & Drill
Week 26 revision: design reviews scoring process rather than recall, why truncated hashes collide, the difference between an open socket and a delivered message, celebrities breaking fanout, and the assumptions AI products violate.
32 min readR25 · Week 25 Recall & Drill
Week 25 revision: agent boundaries as lossy serialisation, judge bias as the real problem rather than subjectivity, fine-tuning teaching behaviour rather than facts, decoding as a bandwidth problem, and images compressed to bounded tokens.
32 min readR24 · Week 24 Recall & Drill
Week 24 revision: reasoning text as compute rather than explanation, the chunk as the atomic unit of retrieval, why keyword search survives, index knobs and recall as a dial, and multiplicative error in agent loops.
32 min readR23 · Week 23 Recall & Drill
Week 23 revision: why every part of the block is load-bearing, masks as the structural difference between families, decoding knobs at the logit level, compute-optimal as a joint optimum, and exactness versus approximation in efficient attention.
32 min readR22 · Week 22 Recall & Drill
Week 22 revision: tokens are not words, attention as soft dictionary lookup, why the scaling constant is a square root, why more heads is not more capacity, and why defined-at-a-position is not trained-at-a-position.
32 min readR21 · Week 21 Recall & Drill
Week 21 revision: regularisers interact rather than stack, convolutions as structural priors, why gated cells only mitigate vanishing gradients, what cosine similarity actually measures, and preprocessing as the silent transfer killer.
32 min readR20 · Week 20 Recall & Drill
Week 20 revision: depth is not decoration, the clean softmax gradient, backpropagation computes but does not update, what momentum and adaptive scaling each fix, and autograd as the real difference from arrays.
32 min readR19 · Week 19 Recall & Drill
Week 19 revision: boosting attacks bias where bagging attacks variance, why the ROC curve flatters imbalanced data, leakage-proof pipelines, thresholds before resampling, and the selection bias in tuned scores.
32 min readR18 · Week 18 Recall & Drill
Week 18 revision: cross-entropy from maximum likelihood, why the default threshold is arbitrary, L1 versus L2 geometry, reading learning curves, and why bagging alone is not enough.
32 min readR17 · Week 17 Recall & Drill
Week 17 revision: authorisation per resource not per login, TLS 1.3 and the chain of trust, injection as a parsing problem, the ML lifecycle scaffold, and least squares by hand.
32 min readR16 · Week 16 Recall & Drill
Week 16 revision: RTO and RPO driving topology, which observability pillar answers which question, PromQL and cardinality, error budgets as budgets, and mitigate-before-diagnose.
32 min readR15 · Week 15 Recall & Drill
Week 15 revision: replication topologies and split-brain, the six-rung consistency ladder, Raft's real difficulty, hot partitions and consistent hashing, and at-least-once queues.
32 min readR14 · Week 14 Recall & Drill
Week 14 revision: declarative reconciliation over imperative starts, config and secrets injection, cloud identity and network isolation, Terraform state as authoritative mapping, and CAP versus PACELC.
32 min readR13 · Week 13 Recall & Drill
Week 13 revision: TCP as a byte stream, L4 versus L7 load balancing, cache-aside and stampedes, CDN cache keys and Vary, and Docker layers as processes not VMs.
32 min readR12 · Week 12 Recall & Drill
Week 12 revision: resource modelling and idempotency keys, GraphQL's N+1 and cost limits, Protobuf field tags and RPC shapes, delegated authorisation versus identity, and Linux processes, memory, and file descriptors.
32 min readR11 · Week 11 Recall & Drill
Week 11 revision: dbt as a compiler not an engine, lakehouse metadata layers over Parquet, the four data-quality pillars, lineage-driven governance and cost, and HTTP caching.
32 min readR10 · Week 10 Recall & Drill
Week 10 revision: bounded versus unbounded data, Spark stages and shuffles, Kafka as a log rather than a queue, watermarks and exactly-once effect, and idempotent orchestration.
32 min readR09 · Week 9 Recall & Drill
Week 9 revision: B-tree selectivity and composite index order, isolation levels and write skew, reading execution plans, the four NoSQL families, and star schema grain.
32 min readR08 · Week 8 Recall & Drill
Week 8 revision: SQL logical execution order, joins as filtered Cartesian products, aggregate NULL semantics, windows that preserve rows, and recursive CTEs.
32 min readR07 · Week 7 Recall & Drill
Week 7 revision: sort trade-offs beyond Big-O, the two binary search templates, DP preconditions versus caching, greedy proofs and backtracking pruning, and normalisation to 3NF.
32 min readR06 · Week 6 Recall & Drill
Week 6 revision: LIFO vs FIFO as a scheduling choice, recursion's three ingredients, tree traversal orders, heaps as partial order, and graph traversal with weights.
32 min readR05 · Week 5 Recall & Drill
Week 5 revision: distributions and Bayes' rule, what a p-value actually claims, two-pointer and sliding-window templates, hashmap internals, and linked-list surgeries.
32 min readR04 · Week 4 Recall & Drill
Week 4 revision: discrete math as the language under SQL, vectors and cosine similarity, matrices as transformations, the chain rule, and gradient descent from scratch.
32 min readR03 · Week 3 Recall & Drill
Week 3 revision: exception discipline and pdb, import resolution and packaging, pytest fixtures and the TDD loop, type hints that no runtime enforces, and Big-O as a growth rate.
32 min readR02 · Week 2 Recall & Drill
Week 2 revision: the iterator protocol behind for-loops, argument grammar and closures, data structure Big-O choices, class attribute lookup, and inheritance vs composition.
32 min readR01 · Week 1 Recall & Drill
Week 1 revision: reproducible dev environments, Git's object model, shell pipelines, doc-hunting strategy, and Python's reference-based memory model.
32 min readDesigning for Scale — Requirements to Review
A twenty-part path through system design: the discipline of constraints and estimation first, then eight classic designs built from those constraints, then the data and ML systems underneath modern products, closing with two end-to-end design reviews.
9 min readDesigning for Scale · End-to-End Design Review II
A second live design review, tearing apart an event ticketing system. Handling massive, instantaneous write spikes, distributed transactions, and inventory locks without destroying the database.
20 min readDesigning for Scale · End-to-End Design Review I
Pulling the components together. How to trace a requirement from population estimates through capacity math to a defensible caching strategy, using a live design review as the frame.
20 min readDesigning for Scale · LLM-as-a-Judge and Release Gates
How to ship generative AI to production without shipping a liability. Breaking subjective prompts into deterministic grading rubrics, automated evaluation pipelines, and multimodal scoring.
25 min readDesigning for Scale · RAG and Vector Search
Why LLMs hallucinate, how Retrieval-Augmented Generation grounds them in reality, and the architecture required to execute semantic search over millions of documents in milliseconds.
25 min readDesigning for Scale · Model Serving
Why wrapping a PyTorch model in a Flask API is a prototype, not a production system. Handling GPU saturation, dynamic batching, and the difference between CPU and GPU scaling.
20 min readDesigning for Scale · The Feature Store
Bridging the gap between data engineering and machine learning. How to serve features for model training offline, and serve those exact same features for inference in five milliseconds online.
25 min readDesigning for Scale · Metering and Billing
Why billing systems cannot drop a single event, the difference between at-least-once and exactly-once processing, and how to build idempotent pipelines that survive crashes without double-charging users.
25 min readDesigning for Scale · The Data Lakehouse
Why the data warehouse and data lake converged. Moving from expensive, proprietary compute-storage monoliths to open table formats like Iceberg, Hudi, and Delta Lake.
25 min readDesigning for Scale · Change Data Capture (CDC)
Why dual-writes fail, how the transaction log is the only true source of state, and the architecture required to stream database changes to search indexes and caches without losing data.
20 min readDesigning for Scale · Real-Time Analytics
How to count billions of events in real-time without crushing your database, using stream processing, time-window aggregations, and Lambda architecture.
25 min readDesigning for Scale · Notification Systems
Why sending a push notification is not a fire-and-forget API call. Handling rate limits, provider outages, deduplication, and the retry queues required to make delivery reliable.
22 min readDesigning for Scale · Typeahead and Search Autocomplete
How to return query suggestions in under fifty milliseconds while the user is still typing, using Tries, offline aggregations, and edge caching.
25 min readDesigning for Scale · Newsfeed and Fan-out
Why reading from a database to render a feed is too slow, and how the fan-out-on-write model precomputes millions of feeds in memory before users even ask for them.
25 min readDesigning for Scale · Chat and Real-time Communication
Why HTTP fails for real-time delivery, how WebSockets change the load balancer math, and the architecture required to deliver a message to a million users simultaneously.
25 min readDesigning for Scale · Distributed Rate Limiter
How to stop abuse without stopping legitimate traffic. An engineering breakdown of Token Bucket, Leaky Bucket, Sliding Window algorithms, and how to execute them across a fleet of servers without locking Redis.
25 min readDesigning for Scale · URL Shortener
The classic system design starting point. It looks like a toy problem until you have to guarantee collision-free generation at scale while keeping read latency under ten milliseconds.
25 min readDesigning for Scale · CAP and the Consistency Spectrum
Why strong consistency is a latency penalty you choose to pay, why eventual consistency is not a defect, and how to use the CAP theorem to end distributed systems arguments.
22 min readDesigning for Scale · The Storage Decision Tree
Databases are not religions. Choosing between relational, wide-column, document, and blob storage by mapping the access pattern before looking at the engine.
20 min readDesigning for Scale · Estimation That Constrains
Back-of-envelope arithmetic is only worth doing when a number rules something out. How to derive requests per second, storage growth, bandwidth and working set — and how to tell a constraining estimate from a decorative one.
20 min readDesigning for Scale · Requirements to Architecture
From a one-line prompt to a defensible architecture: separating functional from non-functional requirements, quantifying who/what/how-many, and refusing to draw a single box until the constraints are on the board.
22 min read