23 articles & lessons
Language Models
Articles and learning notes on language models.
Multimodal LLMs — CLIP, VLMs, Audio, Video
How language models learned to see, hear, and (barely) watch — CLIP dual-encoders, vision-language models, and the ways they still hallucinate about things you can see in the image.
55 min readLLM Serving — KV Cache, Batching, Speculative Decoding
The three ideas that turned LLM inference from $0.20/query to $0.001/query — KV cache, continuous batching, speculative decoding — plus the memory-bandwidth story nobody tells juniors.
55 min readFine-Tuning — LoRA, QLoRA, PEFT, When NOT to Fine-Tune
The engineering behind adapting foundation models. Why LoRA works, how QLoRA squeezes a 70B model onto a single GPU, and the five situations where you should NOT fine-tune.
55 min readLLM Evaluation — LLM-as-Judge, RAGAS, Golden Sets
How you measure a non-deterministic, unbounded-output system without lying to yourself. Golden sets, judge models, RAGAS, and the biases that ruin every naive eval.
55 min readMulti-Agent Orchestration — LangGraph, CrewAI Patterns
When one LLM agent isn't enough — the three patterns (hierarchical, peer-to-peer, graph) that structure multi-agent systems, how LangGraph and CrewAI implement them, and the coordination failures that eat 2 AM pagers.
50 min readLLM Agents — Function Calling, Tools, Planning
How an LLM stops just talking and starts doing — the function-calling protocol, tool schemas, planning strategies (ReAct, Plan-Execute, Reflexion), and the failure modes every agent framework fights.
50 min readVector Databases — pgvector, HNSW, IVF
How dense retrieval actually works at 10M+ vectors — HNSW graph traversal, IVF partitioning, product quantisation, and when pgvector beats Pinecone (and vice versa).
50 min readRAG II — Retrieval, Hybrid Search, Reranking
The other 80% of RAG performance — dense vs sparse vs hybrid retrieval, cross-encoder reranking, query rewriting, and the failure modes each one fixes.
50 min readRAG I — Chunking Strategies & Indexing
The 80% of RAG performance you get before you even do retrieval — chunk sizes, overlap, semantic vs fixed splitting, metadata, and the indexing choices that make or break your bot.
50 min readPrompting — Zero-Shot, Few-Shot, Chain-of-Thought, ReAct
The four prompting patterns that actually move the needle — with the failure modes, decision rules, and Anthropic/OpenAI-recommended templates for each.
50 min readEfficient Attention — Flash, Sparse, Linear
The three families of tricks that broke attention's O(N²) memory wall — Flash-Attention's tile-based recomputation, sparse patterns (BigBird/Longformer), and linear-time approximations. What each buys you, what it costs, and which one to reach for.
50 min readScaling Laws — Chinchilla, Compute-Optimal Training
The empirical laws (Kaplan 2020, Chinchilla 2022) that predict how big a model to train given a fixed compute budget — and why the '20 tokens per parameter' rule made LLaMA-3 possible.
50 min readLLM Sampling — Greedy, Beam, Top-k, Top-p, Temperature
How the model picks the next token — and why the same model can be a boring bureaucrat or a wild storyteller depending on 4 sliders you control at inference.
50 min readEncoder (BERT), Decoder (GPT), Enc-Dec (T5) — When Each
The three families of Transformer architectures — what each is optimised for, why BERT never generates, why GPT never fills in blanks well, and when T5's encoder-decoder still wins in 2026.
50 min readFull Transformer Architecture — Encoder + Decoder
Everything you've learned, snapped together. Attention + FFN + residual + layernorm × N — the block that powers every LLM, and why encoder-only vs decoder-only vs encoder-decoder each earn their place.
55 min readPositional Encoding — Sinusoidal, Learned, RoPE
Attention is permutation-invariant — it doesn't know word order. Positional encoding fixes that. From the paper's sinusoidal trick to modern RoPE, the encoding that quietly powers LLaMA and every 2024 LLM.
50 min readMulti-Head Attention — Parallel Views
Why one attention head is not enough. How 8-128 heads let a Transformer look at the same tokens from different angles simultaneously — with almost the same compute cost.
50 min readQ/K/V Math — Scaled Dot-Product Attention Derived
The Transformer's atomic operation, derived from scratch. Where Q, K, V come from, why we scale by √d_k, and how backprop through attention actually works.
55 min readAttention Intuition — Why RNNs Failed, Why Attention Won
The 2017 paper that killed RNNs. What 'attention' actually means, the bottleneck it fixes in seq2seq, and why parallelism made Transformers eat NLP.
50 min readTokenization — BPE, WordPiece, SentencePiece
The unglamorous layer that decides what your model can even see. BPE, WordPiece, SentencePiece — how the choice of tokenizer costs OpenAI millions and breaks non-English languages.
50 min readApplied LLMs — Origins to Production
A read-in-order path through large language models: the 80-year origin story, how attention actually works, retrieval augmentation, prompting that survives contact with real problems, audio and clinical case studies, and shipping on Azure AI Foundry.
8 min readWhere LLMs Came From, and What They Actually Are
The 80-year origin story, the five ideas that actually matter, and a verified watch-list to go deeper — no maths required.
18 min readKimi K3 From Zero to Deep — Math, Architecture, Training, and Systems
A zero-to-deep Kimi K3 course: prerequisite math and terminology, Transformers and MoE, KDA, AttnRes, vision, RL, training systems, inference, benchmarks, and a verified YouTube learning path.
90 min read