← All topics

82 articles & lessons

Systems & Infrastructure

Articles and learning notes on systems & infrastructure.

  1. The 6-Month Learning Plan · 130 Sessions

    Evergreen 6-month curriculum covering Python → Math → DSA → Databases → Data Engineering → Backend → Systems → Distributed → SRE → Security → Classical ML → Deep Learning → Transformers → LLMs → System Design. Assumes zero background, one concept per session.

    25 min read
  2. Ship It — Deploy, Package, Sustain

    Getting things in front of people: local scripts to a live API, Docker on Azure, a Windows app from C++ compilation to the store, and the learning system that keeps you shipping.

    7 min read
  3. Design an AI Chat Product — RAG + Agents + Serving

    Capstone: designing an AI chat product from the client SSE stream all the way to the LLM serving fleet. Streaming, RAG, agent orchestration, evaluation, safety, cost, and the incidents (ChatGPT title leak, Claude prompt injection, Copilot rate-limit fallout) that shape every mature deployment.

    60 min read
  4. Design a Newsfeed / Recommender — Pull vs Push, Ranking

    The highest-QPS ML application in the world. Fanout-on-write vs fanout-on-read, the celebrity problem, the two-stage candidate → rank funnel, and the feedback loops that eat naïve recommenders.

    55 min read
  5. Design a Chat System — WebSockets, Delivery, Presence

    Realtime bidirectional messaging at billions/day scale — WebSockets, routing tables, at-least-once delivery, ordering, presence, offline pushes, and the reconnect thundering herd that took Slack down.

    55 min read
  6. Design a URL Shortener — the Classic Warm-Up

    Simple on the surface, deep underneath. Unique ID generation, sharding, redirect latency budget, and the anti-abuse story every serious shortener has to solve.

    55 min read
  7. System Design Framework — Reqs, Capacity, HLD, Deep-Dive

    The four-step framework every senior reviewer scores against — how to structure a 45-minute system design review, when to push back, when to accept a number, and the anti-patterns that kill senior engineers in the first ten minutes.

    55 min read
  8. OWASP Top 10, Secrets Mgmt & Threat Modelling

    The security bugs everyone ships (and the ones attackers exploit) — walk the OWASP Top 10 for 2021, wire up a real secrets manager, and threat-model a service in 30 minutes.

    55 min read
  9. TLS 1.3, PKI & Cert Lifecycle

    The green padlock, demystified — key exchange, certificate chains, revocation, and why 60 % of production outages the past decade were expired certs.

    55 min read
  10. AuthN vs AuthZ, Sessions & Password Storage

    The two questions every request has to answer: who are you, and what are you allowed to do? Plus how to store passwords without ending up in a HIBP breach dump.

    55 min read
  11. Incident Response — Runbooks, Postmortems, On-Call

    Turning outages into learning. The IMOC roles, the blameless postmortem template, and why the best incident responses feel boring.

    55 min read
  12. SLIs, SLOs & Error Budgets — the SRE Math

    How to measure reliability without lying to yourself. The formulas Google's SRE team uses to decide when to ship features vs when to freeze deploys.

    55 min read
  13. Prometheus, Grafana, OpenTelemetry — Hands-on

    The default open-source observability stack. Scrape, store, query, visualise, and instrument — all with the CNCF tools every serious team uses.

    55 min read
  14. The 3 Pillars — Metrics, Logs, Traces

    The senses of a running system. What each pillar is good at, what it's terrible at, and why you need all three — not two, not one — to debug production.

    55 min read
  15. Multi-Region — Active-Passive, Active-Active, Failover

    Surviving whole-region outages without lying to your users. The three architectures, the two RTO/RPO knobs, and why 'active-active' is often a marketing term.

    55 min read
  16. Message Queues — SQS, RabbitMQ, Kafka as Queue

    Decouple producers from consumers so the two never have to be up at the same time. The difference between a queue and a log, and why picking the wrong one is a two-year rewrite.

    55 min read
  17. Sharding & Partitioning Strategies

    Split your data before it splits you. Hash vs range vs directory sharding, the hot-partition problem, and why 'just add a shard key' is the wrong answer for 90% of teams.

    55 min read
  18. Consensus — Paxos & Raft Intuition

    How a group of machines that can crash, lag, or lie by omission agree on the same value. The algorithm powering etcd, ZooKeeper, Kafka, CockroachDB, and every serious cluster's brain.

    55 min read
  19. Consistency Models — Linearizable, Sequential, Eventual

    The hierarchy of promises a distributed system makes about what you'll read after you write. Pick the wrong rung and you either lose money or lose latency.

    55 min read
  20. Replication — Leader/Follower, Multi-Leader, Leaderless

    Three ways to keep copies of your data in sync — leader/follower (Postgres, MySQL), multi-leader (active-active), leaderless (Dynamo, Cassandra). Sync vs async, replication lag, and the split-brain problem.

    55 min read
  21. CAP & PACELC — the Actual Trade-Offs

    The three-letter theorem everyone quotes wrong, and the five-letter one that fills in the gap. Consistency, Availability, Partition-tolerance — and what happens the 99.9% of the time your network is fine.

    55 min read
  22. Infrastructure as Code — Terraform / Bicep Basics

    Click-ops doesn't survive contact with production. Terraform + Bicep — providers, state, plans, modules — the ~15 concepts that let you rebuild your whole cloud from a git repo.

    55 min read
  23. Azure Cloud — Identity, Storage, Networking, App Service

    The four pillars of every real Azure workload — Entra ID + RBAC, Storage & Cosmos, VNets & Private Endpoints, App Service & Container Apps. What to pick and why.

    55 min read
  24. Kubernetes II — ConfigMaps, Secrets, HPA, Network Policies

    The next four objects you'll use every week — ConfigMaps for tuning, Secrets for credentials, HPA for autoscaling, NetworkPolicies for zero-trust inside the cluster.

    55 min read
  25. Kubernetes I — Pods, Deployments, Services

    Kubernetes without the mysticism — Pods run containers, Deployments keep N of them alive, Services give them a stable IP. Ship a real app to a local cluster in 25 minutes.

    55 min read
  26. Docker — Images, Layers, Dockerfile, Networking

    The 90 minutes that end ‘works on my machine’ forever. Layers, images, containers, networking, and the ten Dockerfile lines that separate a 2 GB toy from a 60 MB production image.

    55 min read
  27. CDN — Edge, Cache Hierarchies, Cache-Control

    How your static asset travels 40 ms to Sydney instead of 400 ms — edges, origins, cache hierarchies, and the four Cache-Control directives you'll set every day.

    50 min read
  28. Caching — Cache-Aside, Write-Through, TTLs, Invalidation

    The oldest performance trick in the book, done right. Cache-aside vs write-through vs write-back, TTLs, invalidation patterns, and the two hardest problems in computer science.

    55 min read
  29. Networking II — Load Balancers L4 vs L7, Reverse Proxies

    How one hostname fans out to a hundred servers without dropping a packet — the L4 vs L7 decision, health checks, sticky sessions, and the reverse-proxy patterns that run every real web system.

    50 min read
  30. Networking I — TCP/IP, DNS, Sockets

    How bytes actually cross the internet — from getaddrinfo() to the 3-way handshake to the router hop that decides your latency. The mental model every backend engineer needs before they can debug ‘why is this slow?’

    55 min read
  31. OS Basics — Processes, Threads, Memory, FDs

    The four abstractions the OS gives you and why every senior debugging story starts with one of them. Processes, threads, virtual memory, and file descriptors — with real strace output, htop shots, and the incidents each one caused.

    55 min read
  32. Reading Docs & Effective Googling — the Meta-Skill

    The single skill that separates a 10× engineer from a 1× engineer: knowing how to find the answer, fast, without asking a human. Man pages, official docs, GitHub source-diving, and search queries that actually work.

    45 min read
  33. The Command Line — bash, pipes, grep, jq

    Live in the terminal without fear. The pipes, filters, and text-crunching muscle memory that separates an engineer from a button-clicker — and the 20 commands you'll use every single day for the rest of your career.

    55 min read
  34. Git & GitHub — Commits, Branches, PRs

    The 90 minutes that turns Git from a scary black box into a save-point machine you trust with your career. Real commits, real branches, real PRs — the muscle memory every senior engineer runs on.

    50 min read
  35. Dev Environment — Linux/WSL, Terminal, VS Code

    The 90 minutes that saves you 90 hours. Real environment, real editor, real terminal — no ‘works on my machine’ for the next six months.

    45 min read
  36. R26 · Week 26 Recall & Drill

    Week 26 revision: design reviews scoring process rather than recall, why truncated hashes collide, the difference between an open socket and a delivered message, celebrities breaking fanout, and the assumptions AI products violate.

    32 min read
  37. R25 · Week 25 Recall & Drill

    Week 25 revision: agent boundaries as lossy serialisation, judge bias as the real problem rather than subjectivity, fine-tuning teaching behaviour rather than facts, decoding as a bandwidth problem, and images compressed to bounded tokens.

    32 min read
  38. R24 · Week 24 Recall & Drill

    Week 24 revision: reasoning text as compute rather than explanation, the chunk as the atomic unit of retrieval, why keyword search survives, index knobs and recall as a dial, and multiplicative error in agent loops.

    32 min read
  39. R23 · Week 23 Recall & Drill

    Week 23 revision: why every part of the block is load-bearing, masks as the structural difference between families, decoding knobs at the logit level, compute-optimal as a joint optimum, and exactness versus approximation in efficient attention.

    32 min read
  40. R22 · Week 22 Recall & Drill

    Week 22 revision: tokens are not words, attention as soft dictionary lookup, why the scaling constant is a square root, why more heads is not more capacity, and why defined-at-a-position is not trained-at-a-position.

    32 min read
  41. R21 · Week 21 Recall & Drill

    Week 21 revision: regularisers interact rather than stack, convolutions as structural priors, why gated cells only mitigate vanishing gradients, what cosine similarity actually measures, and preprocessing as the silent transfer killer.

    32 min read
  42. R20 · Week 20 Recall & Drill

    Week 20 revision: depth is not decoration, the clean softmax gradient, backpropagation computes but does not update, what momentum and adaptive scaling each fix, and autograd as the real difference from arrays.

    32 min read
  43. R19 · Week 19 Recall & Drill

    Week 19 revision: boosting attacks bias where bagging attacks variance, why the ROC curve flatters imbalanced data, leakage-proof pipelines, thresholds before resampling, and the selection bias in tuned scores.

    32 min read
  44. R18 · Week 18 Recall & Drill

    Week 18 revision: cross-entropy from maximum likelihood, why the default threshold is arbitrary, L1 versus L2 geometry, reading learning curves, and why bagging alone is not enough.

    32 min read
  45. R17 · Week 17 Recall & Drill

    Week 17 revision: authorisation per resource not per login, TLS 1.3 and the chain of trust, injection as a parsing problem, the ML lifecycle scaffold, and least squares by hand.

    32 min read
  46. R16 · Week 16 Recall & Drill

    Week 16 revision: RTO and RPO driving topology, which observability pillar answers which question, PromQL and cardinality, error budgets as budgets, and mitigate-before-diagnose.

    32 min read
  47. R15 · Week 15 Recall & Drill

    Week 15 revision: replication topologies and split-brain, the six-rung consistency ladder, Raft's real difficulty, hot partitions and consistent hashing, and at-least-once queues.

    32 min read
  48. R14 · Week 14 Recall & Drill

    Week 14 revision: declarative reconciliation over imperative starts, config and secrets injection, cloud identity and network isolation, Terraform state as authoritative mapping, and CAP versus PACELC.

    32 min read
  49. R13 · Week 13 Recall & Drill

    Week 13 revision: TCP as a byte stream, L4 versus L7 load balancing, cache-aside and stampedes, CDN cache keys and Vary, and Docker layers as processes not VMs.

    32 min read
  50. R12 · Week 12 Recall & Drill

    Week 12 revision: resource modelling and idempotency keys, GraphQL's N+1 and cost limits, Protobuf field tags and RPC shapes, delegated authorisation versus identity, and Linux processes, memory, and file descriptors.

    32 min read
  51. R11 · Week 11 Recall & Drill

    Week 11 revision: dbt as a compiler not an engine, lakehouse metadata layers over Parquet, the four data-quality pillars, lineage-driven governance and cost, and HTTP caching.

    32 min read
  52. R10 · Week 10 Recall & Drill

    Week 10 revision: bounded versus unbounded data, Spark stages and shuffles, Kafka as a log rather than a queue, watermarks and exactly-once effect, and idempotent orchestration.

    32 min read
  53. R09 · Week 9 Recall & Drill

    Week 9 revision: B-tree selectivity and composite index order, isolation levels and write skew, reading execution plans, the four NoSQL families, and star schema grain.

    32 min read
  54. R08 · Week 8 Recall & Drill

    Week 8 revision: SQL logical execution order, joins as filtered Cartesian products, aggregate NULL semantics, windows that preserve rows, and recursive CTEs.

    32 min read
  55. R07 · Week 7 Recall & Drill

    Week 7 revision: sort trade-offs beyond Big-O, the two binary search templates, DP preconditions versus caching, greedy proofs and backtracking pruning, and normalisation to 3NF.

    32 min read
  56. R06 · Week 6 Recall & Drill

    Week 6 revision: LIFO vs FIFO as a scheduling choice, recursion's three ingredients, tree traversal orders, heaps as partial order, and graph traversal with weights.

    32 min read
  57. R05 · Week 5 Recall & Drill

    Week 5 revision: distributions and Bayes' rule, what a p-value actually claims, two-pointer and sliding-window templates, hashmap internals, and linked-list surgeries.

    32 min read
  58. R04 · Week 4 Recall & Drill

    Week 4 revision: discrete math as the language under SQL, vectors and cosine similarity, matrices as transformations, the chain rule, and gradient descent from scratch.

    32 min read
  59. R03 · Week 3 Recall & Drill

    Week 3 revision: exception discipline and pdb, import resolution and packaging, pytest fixtures and the TDD loop, type hints that no runtime enforces, and Big-O as a growth rate.

    32 min read
  60. R02 · Week 2 Recall & Drill

    Week 2 revision: the iterator protocol behind for-loops, argument grammar and closures, data structure Big-O choices, class attribute lookup, and inheritance vs composition.

    32 min read
  61. R01 · Week 1 Recall & Drill

    Week 1 revision: reproducible dev environments, Git's object model, shell pipelines, doc-hunting strategy, and Python's reference-based memory model.

    32 min read
  62. Designing for Scale — Requirements to Review

    A twenty-part path through system design: the discipline of constraints and estimation first, then eight classic designs built from those constraints, then the data and ML systems underneath modern products, closing with two end-to-end design reviews.

    9 min read
  63. Designing for Scale · End-to-End Design Review II

    A second live design review, tearing apart an event ticketing system. Handling massive, instantaneous write spikes, distributed transactions, and inventory locks without destroying the database.

    20 min read
  64. Designing for Scale · End-to-End Design Review I

    Pulling the components together. How to trace a requirement from population estimates through capacity math to a defensible caching strategy, using a live design review as the frame.

    20 min read
  65. Designing for Scale · LLM-as-a-Judge and Release Gates

    How to ship generative AI to production without shipping a liability. Breaking subjective prompts into deterministic grading rubrics, automated evaluation pipelines, and multimodal scoring.

    25 min read
  66. Designing for Scale · RAG and Vector Search

    Why LLMs hallucinate, how Retrieval-Augmented Generation grounds them in reality, and the architecture required to execute semantic search over millions of documents in milliseconds.

    25 min read
  67. Designing for Scale · Model Serving

    Why wrapping a PyTorch model in a Flask API is a prototype, not a production system. Handling GPU saturation, dynamic batching, and the difference between CPU and GPU scaling.

    20 min read
  68. Designing for Scale · The Feature Store

    Bridging the gap between data engineering and machine learning. How to serve features for model training offline, and serve those exact same features for inference in five milliseconds online.

    25 min read
  69. Designing for Scale · Metering and Billing

    Why billing systems cannot drop a single event, the difference between at-least-once and exactly-once processing, and how to build idempotent pipelines that survive crashes without double-charging users.

    25 min read
  70. Designing for Scale · The Data Lakehouse

    Why the data warehouse and data lake converged. Moving from expensive, proprietary compute-storage monoliths to open table formats like Iceberg, Hudi, and Delta Lake.

    25 min read
  71. Designing for Scale · Change Data Capture (CDC)

    Why dual-writes fail, how the transaction log is the only true source of state, and the architecture required to stream database changes to search indexes and caches without losing data.

    20 min read
  72. Designing for Scale · Real-Time Analytics

    How to count billions of events in real-time without crushing your database, using stream processing, time-window aggregations, and Lambda architecture.

    25 min read
  73. Designing for Scale · Notification Systems

    Why sending a push notification is not a fire-and-forget API call. Handling rate limits, provider outages, deduplication, and the retry queues required to make delivery reliable.

    22 min read
  74. Designing for Scale · Typeahead and Search Autocomplete

    How to return query suggestions in under fifty milliseconds while the user is still typing, using Tries, offline aggregations, and edge caching.

    25 min read
  75. Designing for Scale · Newsfeed and Fan-out

    Why reading from a database to render a feed is too slow, and how the fan-out-on-write model precomputes millions of feeds in memory before users even ask for them.

    25 min read
  76. Designing for Scale · Chat and Real-time Communication

    Why HTTP fails for real-time delivery, how WebSockets change the load balancer math, and the architecture required to deliver a message to a million users simultaneously.

    25 min read
  77. Designing for Scale · Distributed Rate Limiter

    How to stop abuse without stopping legitimate traffic. An engineering breakdown of Token Bucket, Leaky Bucket, Sliding Window algorithms, and how to execute them across a fleet of servers without locking Redis.

    25 min read
  78. Designing for Scale · URL Shortener

    The classic system design starting point. It looks like a toy problem until you have to guarantee collision-free generation at scale while keeping read latency under ten milliseconds.

    25 min read
  79. Designing for Scale · CAP and the Consistency Spectrum

    Why strong consistency is a latency penalty you choose to pay, why eventual consistency is not a defect, and how to use the CAP theorem to end distributed systems arguments.

    22 min read
  80. Designing for Scale · The Storage Decision Tree

    Databases are not religions. Choosing between relational, wide-column, document, and blob storage by mapping the access pattern before looking at the engine.

    20 min read
  81. Designing for Scale · Estimation That Constrains

    Back-of-envelope arithmetic is only worth doing when a number rules something out. How to derive requests per second, storage growth, bandwidth and working set — and how to tell a constraining estimate from a decorative one.

    20 min read
  82. Designing for Scale · Requirements to Architecture

    From a one-line prompt to a defensible architecture: separating functional from non-functional requirements, quantifying who/what/how-many, and refusing to draw a single box until the constraints are on the board.

    22 min read