← All topics

23 articles & lessons

Language Models

Articles and learning notes on language models.

  1. Multimodal LLMs — CLIP, VLMs, Audio, Video

    How language models learned to see, hear, and (barely) watch — CLIP dual-encoders, vision-language models, and the ways they still hallucinate about things you can see in the image.

    55 min read
  2. LLM Serving — KV Cache, Batching, Speculative Decoding

    The three ideas that turned LLM inference from $0.20/query to $0.001/query — KV cache, continuous batching, speculative decoding — plus the memory-bandwidth story nobody tells juniors.

    55 min read
  3. Fine-Tuning — LoRA, QLoRA, PEFT, When NOT to Fine-Tune

    The engineering behind adapting foundation models. Why LoRA works, how QLoRA squeezes a 70B model onto a single GPU, and the five situations where you should NOT fine-tune.

    55 min read
  4. LLM Evaluation — LLM-as-Judge, RAGAS, Golden Sets

    How you measure a non-deterministic, unbounded-output system without lying to yourself. Golden sets, judge models, RAGAS, and the biases that ruin every naive eval.

    55 min read
  5. Multi-Agent Orchestration — LangGraph, CrewAI Patterns

    When one LLM agent isn't enough — the three patterns (hierarchical, peer-to-peer, graph) that structure multi-agent systems, how LangGraph and CrewAI implement them, and the coordination failures that eat 2 AM pagers.

    50 min read
  6. LLM Agents — Function Calling, Tools, Planning

    How an LLM stops just talking and starts doing — the function-calling protocol, tool schemas, planning strategies (ReAct, Plan-Execute, Reflexion), and the failure modes every agent framework fights.

    50 min read
  7. Vector Databases — pgvector, HNSW, IVF

    How dense retrieval actually works at 10M+ vectors — HNSW graph traversal, IVF partitioning, product quantisation, and when pgvector beats Pinecone (and vice versa).

    50 min read
  8. RAG II — Retrieval, Hybrid Search, Reranking

    The other 80% of RAG performance — dense vs sparse vs hybrid retrieval, cross-encoder reranking, query rewriting, and the failure modes each one fixes.

    50 min read
  9. RAG I — Chunking Strategies & Indexing

    The 80% of RAG performance you get before you even do retrieval — chunk sizes, overlap, semantic vs fixed splitting, metadata, and the indexing choices that make or break your bot.

    50 min read
  10. Prompting — Zero-Shot, Few-Shot, Chain-of-Thought, ReAct

    The four prompting patterns that actually move the needle — with the failure modes, decision rules, and Anthropic/OpenAI-recommended templates for each.

    50 min read
  11. Efficient Attention — Flash, Sparse, Linear

    The three families of tricks that broke attention's O(N²) memory wall — Flash-Attention's tile-based recomputation, sparse patterns (BigBird/Longformer), and linear-time approximations. What each buys you, what it costs, and which one to reach for.

    50 min read
  12. Scaling Laws — Chinchilla, Compute-Optimal Training

    The empirical laws (Kaplan 2020, Chinchilla 2022) that predict how big a model to train given a fixed compute budget — and why the '20 tokens per parameter' rule made LLaMA-3 possible.

    50 min read
  13. LLM Sampling — Greedy, Beam, Top-k, Top-p, Temperature

    How the model picks the next token — and why the same model can be a boring bureaucrat or a wild storyteller depending on 4 sliders you control at inference.

    50 min read
  14. Encoder (BERT), Decoder (GPT), Enc-Dec (T5) — When Each

    The three families of Transformer architectures — what each is optimised for, why BERT never generates, why GPT never fills in blanks well, and when T5's encoder-decoder still wins in 2026.

    50 min read
  15. Full Transformer Architecture — Encoder + Decoder

    Everything you've learned, snapped together. Attention + FFN + residual + layernorm × N — the block that powers every LLM, and why encoder-only vs decoder-only vs encoder-decoder each earn their place.

    55 min read
  16. Positional Encoding — Sinusoidal, Learned, RoPE

    Attention is permutation-invariant — it doesn't know word order. Positional encoding fixes that. From the paper's sinusoidal trick to modern RoPE, the encoding that quietly powers LLaMA and every 2024 LLM.

    50 min read
  17. Multi-Head Attention — Parallel Views

    Why one attention head is not enough. How 8-128 heads let a Transformer look at the same tokens from different angles simultaneously — with almost the same compute cost.

    50 min read
  18. Q/K/V Math — Scaled Dot-Product Attention Derived

    The Transformer's atomic operation, derived from scratch. Where Q, K, V come from, why we scale by √d_k, and how backprop through attention actually works.

    55 min read
  19. Attention Intuition — Why RNNs Failed, Why Attention Won

    The 2017 paper that killed RNNs. What 'attention' actually means, the bottleneck it fixes in seq2seq, and why parallelism made Transformers eat NLP.

    50 min read
  20. Tokenization — BPE, WordPiece, SentencePiece

    The unglamorous layer that decides what your model can even see. BPE, WordPiece, SentencePiece — how the choice of tokenizer costs OpenAI millions and breaks non-English languages.

    50 min read
  21. Applied LLMs — Origins to Production

    A read-in-order path through large language models: the 80-year origin story, how attention actually works, retrieval augmentation, prompting that survives contact with real problems, audio and clinical case studies, and shipping on Azure AI Foundry.

    8 min read
  22. Where LLMs Came From, and What They Actually Are

    The 80-year origin story, the five ideas that actually matter, and a verified watch-list to go deeper — no maths required.

    18 min read
  23. Kimi K3 From Zero to Deep — Math, Architecture, Training, and Systems

    A zero-to-deep Kimi K3 course: prerequisite math and terminology, Transformers and MoE, KDA, AttnRes, vision, RL, training systems, inference, benchmarks, and a verified YouTube learning path.

    90 min read