
The paper that made diffusion affordable.
Why it mattersLatent diffusion is why high-resolution image generation became trainable outside big labs: compress away imperceptible detail first, run the diffusion model in latent space, and condition with cross-attention on text or boxes.

Showed the large models of the era were badly undertrained: for a fixed compute budget, model size and training tokens should scale roughly in step, not size alone.
Why it mattersIt reset how everyone spends a training budget.

Fine-tuning with human feedback made a 1.3B model preferred over 175B GPT-3, establishing RLHF as the step that turns a language model into something that follows instructions.
Why it mattersThe gap between a raw language model and an assistant is this paper: alignment to human preference beat raw scale for usefulness, with a 1.3B InstructGPT preferred to 175B GPT-3.

This is part 2 of what to do when facing a limited amount of labeled data for supervised learning tasks.
Why it mattersIf you're training models with scarce labels and a tight annotation budget, this walks through active learning methods for choosing which samples are worth labeling.

Showed that prompting a model to produce intermediate reasoning steps unlocks arithmetic and symbolic reasoning that appears only above a certain model scale.
Why it mattersThe finding that asking for the working, not just the answer, changes what a model can solve — and that the effect only emerges above a certain scale.

When facing a limited amount of labeled data for supervised learning tasks, four approaches are commonly discussed.
Why it mattersA clear walkthrough of the main approaches for training when labeled data is scarce, giving practitioners concrete semi-supervised learning strategies to squeeze more value from unlabeled datasets.

[Updated on 2022-03-13: add expert choice routing .] [Updated on 2022-06-10]: Greg and I wrote a shorted and upgraded version of this post, published on OpenAI Blog.
Why it mattersA rigorous, well-organized reference on the parallelism strategies (data, tensor, pipeline, MoE, expert-choice routing) and memory-saving techniques that make large-scale model training feasible.

[Updated on 2021-09-19: Highly recommend this blog post on score-based generative modeling by Yang Song (author of several key papers in the references)]. [Updated on 2022-08-27.
Why it mattersA rigorous, continuously-updated technical walkthrough of diffusion models — from the DDPM math to consistency models and latent diffusion.

Freezes the pretrained weights and trains small rank-decomposition matrices instead, cutting trainable parameters by orders of magnitude with no added inference latency.
Why it mattersIt is why fine-tuning is something you can do on your own hardware.

The goal of contrastive representation learning is to learn such an embedding space in which similar sample pairs stay close to each other while dissimilar ones are far apart.
Why it mattersA grounding on how contrastive learning builds embedding spaces where similar samples cluster and dissimilar ones separate.

Encodes position by rotating query and key vectors, so attention depends on relative distance — the position scheme nearly every current open model uses.
Why it mattersRotary embeddings are why modern models handle long contexts at all, and why context-extension tricks work the way they do.

Large pretrained language models are trained over a sizable collection of online data. They unavoidably acquire certain toxic behavior and biases from the Internet.
Why it mattersIf you're deploying LLMs in production, this breaks down concrete techniques for controlling toxic and biased generation—the safety layer that separates a demo from a shippable product.

Routes each token to a single expert rather than combining several, making trillion-parameter sparse models trainable at roughly the cost of a dense one.
Why it mattersThe paper that made mixture-of-experts practical, by routing each token to exactly one expert instead of blending many.

[Updated on 2021-02-01: Updated to version 2.0 with several work added and many typos fixed.] [Updated on 2021-05-26.
Why it mattersA deep technical survey of how to steer language model outputs — covering guided decoding, prompt tuning, and unlikelihood training — for anyone who needs finer control over what an LLM generates than raw prompting provides.

Although most popular and successful model architectures are designed by human experts, it doesn’t mean we have explored the entire network architecture space and settled down with the best option.
Why it mattersExplains how model architectures can be discovered automatically rather than hand-designed by experts, covering the search space, strategy, and performance estimation that underpin NAS.

Exploration versus exploitation in deep reinforcement learning.
Why it mattersA structured survey of modern exploration methods in deep RL — count-based bonuses, curiosity/forward-dynamics, and exploration via disagreement.

Combines a parametric seq2seq generator with a non-parametric dense retrieval index, letting a model draw on knowledge that can be updated without retraining.
Why it mattersThe original formulation of what the industry now just calls RAG: pair a generator with a dense retrieval index so knowledge lives in a store you can update rather than in weights you must retrain.

[Updated on 2020-02-03: mentioning PCG in the “Task-Specific Curriculum” section. [Updated on 2020-02-04: Add a new “curriculum through distillation&rdquo.
Why it mattersA structured deep-dive into how curriculum learning accelerates and stabilizes RL training — covering task ordering, procedural content generation, and distillation-based curricula.

Established that language-model loss falls as a smooth power law in model size, data and compute — the result that turned scaling from a hunch into a budgeting exercise.
Why it mattersThis is where 'bigger reliably means better' got its evidence, and where the curve's shape — smooth and predictable across orders of magnitude — first let labs forecast a model's performance before training it.

A survey of self-supervised representation learning, covering contrastive predictive coding and the momentum-contrast family including MoCo, SimCLR and BYOL.
Why it mattersA deep, well-organized survey of self-supervised representation learning methods (CPC, MoCo, SimCLR, BYOL) that gives engineers the conceptual grounding to build and choose embedding/pretraining approaches for AI-native products.

Shares a single key-value head across all query heads, shrinking the KV cache that dominates memory during incremental decoding.
Why it mattersIt named the real bottleneck in generation: not the arithmetic, but the size of the key-value cache carried per token.

Stochastic gradient descent is a universal choice for optimizing deep learning models. However, it is not the only option.
Why it mattersExplains how evolution strategies can optimize objectives where gradients are unavailable or unreliable.

In my earlier post on meta-learning , the problem is mainly defined in the context of few-shot classification.
Why it mattersA clear, technical walkthrough of meta-RL methods for training agents that generalize to unseen tasks quickly — useful background for anyone building adaptive or few-shot-capable RL agents.

Deep RL is too sample-hungry to train on real robots, so models are trained in simulation and fail on the reality gap.
Why it mattersIf you're training RL policies in simulation and struggling to deploy them on physical robots, this breaks down how domain randomization over physical parameters (friction, mass, damping) closes the reality gap and where naive sim modeling fails.

[Updated on 2019-05-27.
Why it mattersA rigorous walkthrough of why overparameterized deep networks generalize instead of overfitting, covering the Lottery Ticket Hypothesis, intrinsic dimension, and generalization bounds.

[Updated on 2019-10-01: thanks to Tianhao, we have this post translated in Chinese !]
Why it mattersA rigorous, well-organized primer on meta-learning that walks through metric-based, model-based, and optimization-based approaches (including MAML and memory-augmented networks) with the math and intuition.

GANs and VAEs never explicitly learn the density of real data because the integral is intractable.
Why it mattersA rigorous walkthrough of normalizing flows and models like RealNVP and Glow that explicitly learn tractable data likelihoods — useful for practitioners who need exact density estimation rather than the implicit distributions of GANs/VAEs.

A tour from the plain autoencoder through denoising, sparse and contractive variants to the variational autoencoder and beta-VAE.
Why it mattersA single, rigorous walkthrough of the autoencoder-to-VAE lineage — including VQ-VAE, VQ-VAE-2, and TD-VAE.

[Updated on 2018-10-28: Add Pointer Network and the link to my implementation of Transformer.] [Updated on 2018-11-06.
Why it mattersA rigorous, implementation-backed walkthrough of attention and the Transformer architecture that grounds engineers in the mechanics behind modern LLMs, with code links for self-attention, Pointer Networks, and Neural Turing Machines.

A hands-on implementation walkthrough of deep reinforcement learning models in TensorFlow against OpenAI Gym.
Why it mattersBridges the gap between deep RL theory and working code, walking through actual TensorFlow + OpenAI Gym implementations of models most tutorials only describe abstractly.

A survey of policy gradient algorithms in reinforcement learning, from the basic theorem through actor-critic variants including SAC, D4PG, TD3 and SVPG.
Why it mattersA comprehensive, math-first survey of policy gradient methods that walks through the derivations and design tradeoffs of REINFORCE, A2C/A3C, DDPG, TD3, SAC, PPO, IMPALA and more in one place.

[Updated on 2020-09-03: Updated the algorithm of SARSA and Q-learning so that the difference is more pronounced. [Updated on 2021-09-19.
Why it mattersA rigorous, foundational walkthrough of core RL algorithms (SARSA, Q-learning, policy gradients) that grounds the concepts increasingly relevant to agent training and RLHF-style fine-tuning.

The exploration-exploitation dilemma stated as the multi-armed bandit problem.
Why it mattersA clear, code-backed walkthrough of bandit algorithms (epsilon-greedy, UCB, Thompson sampling) that maps the exploration/exploitation tradeoff to real problems like ad selection and A/B testing.

Human vocabulary comes in free text.
Why it mattersA clear walkthrough of how free-text words become numeric vectors — from one-hot encoding to learned embeddings — grounding the intuition behind the embedding models and vector search that agentic engineers rely on daily.

Naftali Tishby's information bottleneck applied to deep learning, using information theory to describe how a network's representations grow and transform over the course of training.
Why it mattersIt unpacks Tishby's Information Bottleneck framework and the two-phase (fitting then compression) view of DNN training, giving practitioners an information-theoretic lens on generalization that most engineering-focused writeups skip.

A walk from the original generative adversarial network to Wasserstein GAN.
Why it mattersA rigorous, well-illustrated walkthrough of why vanilla GANs are unstable and how Wasserstein distance fixes the gradient/convergence problems — useful grounding for anyone building or debugging generative models.

The 2017 paper that dropped recurrence and convolution for self-attention alone, introducing the Transformer — the architecture every large language model still builds on.
Why it mattersEvery model you use descends from this architecture, and reading it is how the rest of the stack stops being magic.
An index of the vibe-coding frontier. Corrections welcome.