back to blog
systemsbeginner 25m read
The 6-Month Learning Plan · 130 Sessions
Evergreen 6-month curriculum covering Python → Math → DSA → Databases → Data Engineering → Backend → Systems → Distributed → SRE → Security → Classical ML → Deep Learning → Transformers → LLMs → System Design. Assumes zero background, one concept per session.
The plan: 130 atomic sessions across 16 modules, 26 weeks, 6 events/week (Mon–Fri new + Sat revision). Assumes zero background. Each session = 60 minutes. Sunday off.
How to use this
- Follow in order. Every session lists prerequisites; if you can't answer them in 30 seconds each, back up.
- 60 minutes per new session. 5-min intuition → 15-min visual → 20-min hands-on → 10-min production reality → 10-min quiz.
- Saturdays are for revision. Blank-page recall, hands-on redo, Anki cards.
- Sundays are off. Rest, catch up, or work on a side project applying the week's concepts.
The 16 modules
M00 · Setup & Tools (S001–S004, 4 sessions)
- S001 🧰 Dev Environment — Linux/WSL, Terminal, VS Code — Baseline setup so you never hit an environment wall.
- S002 🔀 Git & GitHub — Commits, Branches, PRs — Version control the way real teams use it.
- S003 ⌨️ The Command Line — bash, pipes, grep, jq — Live in the terminal and never fear it again.
- S004 🔍 Reading Docs & Effective Googling — the Meta-Skill — The single skill that unblocks every other one.
M01 · Python Foundations (S005–S014, 10 sessions)
- S005 🧠 Python Variables & Types — Mental Model of Memory — What actually happens when you write x = 5.
- S006 🔁 Control Flow — if/else, loops, comprehensions — Every program is a giant decision tree.
- S007 🎯 Functions — arguments, scope, closures — The atomic unit of reusable code.
- S008 📦 Data Structures — list, tuple, dict, set (when to use what) — Pick the right container and your code writes itself.
- S009 🏗️ Classes & Objects — the OOP Mental Model — Bundling data with behaviour, the way OOP actually clicks.
- S010 🧬 Inheritance, Composition & Polymorphism — When to inherit, when to compose, why 'favour composition' is a rule.
- S011 🐛 Errors, Exceptions & Debugging with pdb — Failures are data — learn to read them.
- S012 📚 Modules, Packages, Virtualenvs, pip & uv — How real projects are organised and shipped.
- S013 ✅ Testing with pytest — TDD Workflow — Tests turn code from hope into evidence.
- S014 🏷️ Type Hints, mypy, dataclasses & pydantic — Modern Python that scales past 1000 lines.
M02 · Math Foundations (S015–S022, 8 sessions)
- S015 📈 Big-O Notation — Reasoning About Scale — How to talk about performance without measuring.
- S016 🔢 Discrete Math — Sets, Logic, Combinatorics, Graphs — The vocabulary of CS.
- S017 ➡️ Linear Algebra I — Vectors, Dot Product, Geometry — The language ML speaks.
- S018 🔲 Linear Algebra II — Matrices, Transforms, Eigenvalues — Neural nets are stacked matrix multiplies.
- S019 📉 Calculus I — Derivatives & Chain Rule — Rate of change is the heart of learning.
- S020 ⛰️ Calculus II — Gradients & Gradient Descent from Scratch — The one optimisation algorithm behind all of ML.
- S021 🎲 Probability — Random Variables, Distributions, Expectation — Uncertainty, quantified.
- S022 📊 Statistics — CLT, Hypothesis Testing, Confidence Intervals — How to make claims from noisy data.
M03 · Data Structures & Algorithms (S023–S034, 12 sessions)
- S023 🧵 Arrays & Strings — Indexing, Slicing, Two-Pointer — The workhorse data structure.
- S024 🗺️ Hashmaps & Sets — Hash Functions, Collisions — O(1) lookup is the trick behind 40% of interview problems.
- S025 🔗 Linked Lists — Singly, Doubly, When They Win — Pointers made explicit.
- S026 📚 Stacks & Queues — LIFO/FIFO in Practice — Two shapes of order.
- S027 🌀 Recursion — Call Stack, Base Case, Worked Examples — Functions that call themselves — and why that's not scary.
- S028 🌳 Trees & BSTs — Traversal (BFS/DFS) — Hierarchical data everywhere.
- S029 🎋 Heaps & Priority Queues — 'Give me the smallest/largest thing' in O(log n).
- S030 🕸️ Graphs — Representation, BFS, DFS, Shortest Path — Networks, relationships, roads — all graphs.
- S031 🔀 Sorting — Merge, Quick, and When to Trust the Built-in — Understand the trade-offs so you know when to worry.
- S032 🎯 Binary Search — the Pattern Behind 100 Problems — Sorted + bisect = O(log n).
- S033 🧩 Dynamic Programming — Memoisation & Tabulation — Break the problem down, cache the answer.
- S034 🪜 Greedy & Backtracking — When to Use Each — Two more decision-making patterns.
M04 · Databases & SQL (S035–S044, 10 sessions)
- S035 📋 The Relational Model — Tables, Keys, Normalisation — Why tables + relationships beat spreadsheets.
- S036 🔤 SQL Basics — SELECT, WHERE, ORDER BY, LIMIT — The universal data language.
- S037 🔗 Joins — INNER, LEFT, RIGHT, FULL, Anti-Join — Combining tables the right way.
- S038 ∑ Aggregations — GROUP BY, HAVING, Subqueries — Turn rows into insight.
- S039 🪟 Window Functions — the Game-Changer — Ranking, running totals, moving averages without joins.
- S040 🧱 CTEs & Recursive Queries — Compose complex SQL like Python functions.
- S041 📑 Indexes — B-Tree Intuition, When to Add — The single biggest performance lever.
- S042 🔒 Transactions & ACID — Isolation Levels, MVCC — Correctness under concurrency.
- S043 🗺️ Query Planning — EXPLAIN, Execution Plans, Tuning — How the database actually runs your query.
- S044 🧭 NoSQL Landscape — KV, Document, Column, Graph — When relational isn't the answer.
M05 · Data Engineering (S045–S054, 10 sessions)
- S045 🧮 Data Modelling — Dimensional, Data Vault, OBT — How analytics warehouses are shaped.
- S046 🌊 Batch vs Streaming — Mental Model & Use Cases — The two shapes of data pipelines.
- S047 🔥 Spark — RDD, DataFrame, Jobs/Stages/Shuffles — The distributed compute engine that runs everything.
- S048 📮 Kafka — Topics, Partitions, Consumer Groups — The event backbone of modern systems.
- S049 ⏱️ Stream Processing — Watermarks, Windows, Exactly-Once — Time is the hard part of streaming.
- S050 🎼 Orchestration — Airflow, DAGs, Retries, Backfills — Scheduling pipelines like an adult.
- S051 🧪 dbt — Models, Tests, Docs, Warehouse-Native ELT — Version-controlled SQL that scales.
- S052 🏞️ Lakehouse — Delta / Iceberg / Hudi, ACID on Files — The convergence of warehouse and lake.
- S053 🎚️ Data Quality — Freshness, Volume, Schema, Distribution — How to know your data is trustworthy.
- S054 💰 Governance & Cost — Lineage, PII, Attribution — The unsexy work that keeps companies out of court.
M06 · Backend & APIs (S055–S059, 5 sessions)
- S055 🌐 HTTP Fundamentals — Verbs, Status Codes, Headers, Caching — The protocol every web system speaks.
- S056 🧬 REST API Design — Resources, Versioning, Idempotency — APIs that scale past v1.
- S057 ◈ GraphQL — Schema, Resolvers, N+1, When to Pick It — REST's flexible cousin.
- S058 ⚡ gRPC & Protobuf — When RPC Wins — Binary + typed = fast.
- S059 🔐 AuthN & AuthZ — OAuth 2.0, OIDC, JWT — The two things you can't afford to get wrong.
M07 · Systems & Infrastructure (S060–S069, 10 sessions)
- S060 🖥️ OS Basics — Processes, Threads, Memory, FDs — What the OS actually does.
- S061 📡 Networking I — TCP/IP, DNS, Sockets — How bytes cross the internet.
- S062 ⚖️ Networking II — Load Balancers L4 vs L7, Reverse Proxies — Distributing traffic without dropping it.
- S063 ⚡ Caching — Cache-Aside, Write-Through, TTLs, Invalidation — The oldest performance trick, done right.
- S064 🌍 CDN — Edge, Cache Hierarchies, Cache-Control — Move bytes close to users.
- S065 🐳 Docker — Images, Layers, Dockerfile, Networking — Package your app so it runs anywhere.
- S066 ⚓ Kubernetes I — Pods, Deployments, Services — Orchestrating containers at scale.
- S067 🎯 Kubernetes II — ConfigMaps, Secrets, HPA, Network Policies — Making K8s production-ready.
- S068 ☁️ Azure Cloud — Identity, Storage, Networking, App Service — The cloud you actually use.
- S069 📐 Infrastructure as Code — Terraform / Bicep Basics — Click-ops doesn't scale.
M08 · Distributed Systems (S070–S076, 7 sessions)
- S070 ⚖️ CAP & PACELC — the Actual Trade-Offs — Beyond the meme, what CAP really means.
- S071 🔁 Replication — Leader/Follower, Multi-Leader, Leaderless — How systems stay up when a node dies.
- S072 📐 Consistency Models — Linearizable, Sequential, Eventual — The hierarchy of promises.
- S073 🗳️ Consensus — Paxos & Raft Intuition — How machines agree.
- S074 🧩 Sharding & Partitioning Strategies — Split your data before it splits you.
- S075 📨 Message Queues — SQS, RabbitMQ, Kafka as Queue — Decouple producers from consumers.
- S076 🌏 Multi-Region — Active-Passive, Active-Active, Failover — Surviving whole-region outages.
M09 · Observability & SRE (S077–S080, 4 sessions)
- S077 🔭 The 3 Pillars — Metrics, Logs, Traces — The senses of a running system.
- S078 📊 Prometheus, Grafana, OpenTelemetry — Hands-on — The default stack.
- S079 🎯 SLIs, SLOs & Error Budgets — the SRE Math — How to measure reliability without lying.
- S080 🚨 Incident Response — Runbooks, Postmortems, On-Call — Turning outages into learning.
M10 · Security (S081–S083, 3 sessions)
- S081 🔑 AuthN vs AuthZ, Sessions & Password Storage — The two questions every request has to answer.
- S082 🔐 TLS 1.3, PKI & Cert Lifecycle — The green padlock, demystified.
- S083 🛡️ OWASP Top 10, Secrets Mgmt & Threat Modelling — The 10 ways your app will get hacked, and how to stop it.
M11 · Classical ML (S084–S095, 12 sessions)
- S084 🧠 The ML Mental Model — Features, Labels, Train/Val/Test — What ML actually is.
- S085 📈 Linear Regression from Scratch (numpy) — The 'hello world' of ML.
- S086 🎯 Logistic Regression — Sigmoid, Cross-Entropy, from Scratch — Classification, first principles.
- S087 🎚️ Regularization — L1, L2, Elastic Net — Fighting overfitting.
- S088 ⚖️ Bias–Variance Trade-off & Learning Curves — The fundamental ML trade-off.
- S089 🌳 Decision Trees — Gini, Entropy, Splits — The most interpretable model.
- S090 🌲 Random Forest & Bagging — Trees + randomness = surprisingly strong.
- S091 🚀 Gradient Boosting — XGBoost, LightGBM — The tabular-data champion.
- S092 📏 Evaluation Metrics — P/R/F1/ROC/PR/AUC — Accuracy is almost never the right metric.
- S093 🛠️ Feature Engineering — Encoding, Scaling, Missing — The unsexy work that wins Kaggles.
- S094 🎭 Imbalanced Data — SMOTE, Class Weights, Thresholds — Fraud, churn, disease — all imbalanced.
- S095 🎛️ Model Selection — CV, Hyperparameter Tuning, Optuna — Picking the winner honestly.
M12 · Deep Learning (S096–S105, 10 sessions)
- S096 ⚡ Perceptron & Activation Functions — The neuron, first principles.
- S097 ➡️ Multi-Layer Perceptron — Forward Pass — Stacking neurons into a network.
- S098 🔙 Backpropagation — Derived by Hand on a 2-Layer Net — The chain rule that makes deep learning work.
- S099 🚴 Optimizers — SGD, Momentum, Adam, RMSprop — How to actually train the thing.
- S100 🔥 PyTorch Fundamentals — Tensors, Autograd, nn.Module — The tool you'll live in.
- S101 🧊 Regularization in DL — Dropout, BatchNorm, Weight Decay — Keeping deep nets from memorising.
- S102 🖼️ CNNs — Convolution, Pooling, ImageNet Architectures — Vision, cracked.
- S103 🌀 RNNs & LSTMs — Sequences & the Vanishing Gradient — Sequences before transformers took over.
- S104 📐 Embeddings — word2vec, GloVe, Contrastive Learning — Turning things into vectors.
- S105 🎓 Transfer Learning & Fine-Tuning Classical DL — Standing on giants' shoulders.
M13 · NLP & Transformers (S106–S115, 10 sessions)
- S106 🔤 Tokenization — BPE, WordPiece, SentencePiece — How text becomes numbers.
- S107 👀 Attention Intuition — Why RNNs Failed, Why Attention Won — The idea that changed everything.
- S108 🧮 Q/K/V Math — Scaled Dot-Product Attention Derived — The equation that powers ChatGPT.
- S109 👁️ Multi-Head Attention — Parallel Views — Why one attention isn't enough.
- S110 📍 Positional Encoding — Sinusoidal, Learned, RoPE — How transformers know word order.
- S111 🏗️ Full Transformer Architecture — Encoder + Decoder — Putting it all together.
- S112 🎭 Encoder (BERT), Decoder (GPT), Enc-Dec (T5) — When Each — The three families of LLMs.
- S113 🎲 LLM Sampling — Greedy, Beam, Top-k, Top-p, Temperature — How the model picks the next word.
- S114 📈 Scaling Laws — Chinchilla, Compute-Optimal Training — Why bigger models keep winning.
- S115 ⚡ Efficient Attention — Flash, Sparse, Linear — Making attention scale.
M14 · LLMs & Applications (S116–S125, 10 sessions)
- S116 💬 Prompting — Zero-Shot, Few-Shot, Chain-of-Thought, ReAct — The interface layer of the LLM era.
- S117 📚 RAG I — Chunking Strategies & Indexing — The '80% of LLM apps' pattern, part 1.
- S118 🔎 RAG II — Retrieval, Hybrid Search, Reranking — Making retrieval actually work.
- S119 🧭 Vector Databases — pgvector, HNSW, IVF — The infra behind RAG.
- S120 🤖 LLM Agents — Function Calling, Tools, Planning — LLMs that do things, not just talk.
- S121 🕸️ Multi-Agent Orchestration — LangGraph, CrewAI Patterns — When one agent isn't enough.
- S122 ⚖️ LLM Evaluation — LLM-as-Judge, RAGAS, Golden Sets — Measuring quality of a non-deterministic system.
- S123 🎛️ Fine-Tuning — LoRA, QLoRA, PEFT, When NOT to Fine-Tune — Adapting foundation models.
- S124 🚀 LLM Serving — KV Cache, Batching, Speculative Decoding — Making inference cheap enough to ship.
- S125 🎨 Multimodal LLMs — CLIP, VLMs, Audio, Video — Beyond text.
M15 · System Design (S126–S130, 5 sessions)
- S126 🏛️ System Design Framework — Reqs, Capacity, HLD, Deep-Dive — How to structure a 45-min interview.
- S127 🔗 Design a URL Shortener — the Classic Warm-Up — Simple problem, deep trade-offs.
- S128 💬 Design a Chat System — WebSockets, Delivery, Presence — Real-time at scale.
- S129 📰 Design a Newsfeed / Recommender — Pull vs Push, Ranking — The core of every social product.
- S130 🏁 Design an AI Chat Product — RAG + Agents + Serving — Capstone: put the whole 6 months together.
Weekly revisions
One revision session per week, covering the 5 new sessions of that week.
- R01 · Week 1 revision — S001–S005
- R02 · Week 2 revision — S006–S010
- R03 · Week 3 revision — S011–S015
- R04 · Week 4 revision — S016–S020
- R05 · Week 5 revision — S021–S025
- R06 · Week 6 revision — S026–S030
- R07 · Week 7 revision — S031–S035
- R08 · Week 8 revision — S036–S040
- R09 · Week 9 revision — S041–S045
- R10 · Week 10 revision — S046–S050
- R11 · Week 11 revision — S051–S055
- R12 · Week 12 revision — S056–S060
- R13 · Week 13 revision — S061–S065
- R14 · Week 14 revision — S066–S070
- R15 · Week 15 revision — S071–S075
- R16 · Week 16 revision — S076–S080
- R17 · Week 17 revision — S081–S085
- R18 · Week 18 revision — S086–S090
- R19 · Week 19 revision — S091–S095
- R20 · Week 20 revision — S096–S100
- R21 · Week 21 revision — S101–S105
- R22 · Week 22 revision — S106–S110
- R23 · Week 23 revision — S111–S115
- R24 · Week 24 revision — S116–S120
- R25 · Week 25 revision — S121–S125
- R26 · Week 26 revision — S126–S130
Timeline
- Start: Monday 2026-07-27 · 20:30 IST
- End: Saturday 2027-01-23 · 11:00 IST
- 156 total events (130 new + 26 revisions)
Design principles
- Assume zero background — every concept from first principles.
- One concept per session — mastered, not skimmed.
- Explicit prerequisites — no back-references without setup.
- Every session has runnable code — no hand-wavy pseudo-code.
- Every session tested with 'explain out loud' — if you can't teach it, redo it.
This plan supersedes the previous 36-topic / 48-session / 24-topic series. Old URLs redirect to the equivalent atomic session in this plan.