Machine Learning from Human Preferences

Authors

Sang T. Truong

Andreas Haupt

Sanmi Koyejo

Updated

September 23, 2026

Introduction

Machine learning is increasingly shaping various aspects of our lives, from education and healthcare to scientific discovery. A key challenge in developing trustworthy intelligent systems is ensuring they align with human preferences. Learning from human feedback offers a promising solution to this challenge. This book introduces the foundations and practical applications of machine learning from human preferences. Instead of manually predefining the learning goal, the book presents preference-based learning that incorporates human feedback to guide the learning process, drawing insights from related fields such as economics, psychology, and human-computer interaction. Throughout, we emphasize not only the methods themselves but also their assumptions, limitations, and the conditions under which they can and cannot be applied—understanding when a model fails is as important as understanding when it works. By the end of this book, readers will be equipped with the key concepts and tools needed to design systems that effectively align with human preferences.

The book is intended for researchers, practitioners, and students who are interested in integrating machine learning with human-centered application. We assume some basic knowledge of probability and statistics, but provides sufficient background and references for the readers to follow the main ideas. The book also provides illustrative program examples and datasets. The field of machine learning from human preference is a vibrant area of research and practice with many open challenges and opportunities, and we hope that this book will inspire readers to further explore and advance this exciting field.

We hope with the present book to both allow more use of human preferences in machine learning, and new data modalities as Artificial Intelligence systems become increasingly important.

Stanford, May 2025, THK

Structure of this book

The book has three parts which introduce foundational models, present learning paradigms, and discuss societal considerations.

Chapter 1 lays the mathematical groundwork for the rest of the book. It covers random preference models, types of comparison data (binary rankings, accept-reject, lists), and deterministic and stochastic utility models including the Rasch model, Bradley-Terry, and Gaussian Processes. A central theme are axioms that allow counterfactual predictions, like the Independence of Irrelevant Alternatives (IIA) axiom.

Chapter 2 studies how to infer utility from preference data. It covers maximum likelihood estimation, Bayesian inference via MCMC and Laplace approximation, and online learning through Elo ratings as stochastic gradient descent. The chapter also addresses regularization, model selection via cross-validation, and modern optimization methods, with a case study on LLM preference learning. It also considers active data collection with the goal of efficient preference elicitation. Fisher information quantifies how much each comparison teaches us, enabling optimal query selection via A-optimal, D-optimal, and E-optimal design criteria. The chapter includes applications to robotic trajectory learning and active DPO for language model alignment.

Chapter 3 considers how machines and humans interact in a common setting when the machine has an objective that is based on human behavior. We study this under sequential interactions with identified and anonymous humans. It introduces bandit approaches, optimization against fixed human policies, and cooperative inverse reinforcement learning.

Chapter 4 brings in normative judgment, and documents how we can “invert” preference from behavior. The chapter covers behavioral biases, style control, surrogates, and paternalism.

Chapter 5 examines preference aggregation when building AI systems for multiple people. It introduces social choice theory, Arrow’s and Gibbard-Satterthwaite’s impossibility theorems, and ways to escape these impossibilities through domain restrictions and scoring rules. It connects the Borda count to DPO, introduces Nash Learning and how to aggregate preference in a principled manner when disagreement is structured.

How to engage with this book

There are four models of reading, and teaching with, this book. Chapter 1 is foundational to all of the book, so is part of all of these pathways.

  • For practitioners and those teaching applied AI content, we recommend Chapters 1, 2, and 3, which can be used as a sequence in an early graduate course on Machine Learning. This sequence covers foundations, learning methods, and decision-making with human preferences.

  • For people with background in discrete choice, we propose to skim Chapter 1, and study Chapters 2 and 3. These studies allow readers to integrate machine learning in their studies of discrete choice, demand models, and Industrial Organization.

  • For those with deep background in machine learning, we propose to study Chapters 2, 3, and 4. These chapters maximize the amount of machine learning covered, and are suitable for a deep learning-based course on machine learning.

  • For those interested in the societal and theoretical foundations of machine learning from comparisons, we recommend Chapters 1, 4, and 4. Chapter 1 establishes the modeling foundations, Chapter 4 studies the normative foundations of preference learning, and chapter 5 studies aggregation. This pathway is suitable for a course on Computation and Society.

For instructors

This book is designed to support course instruction, and each chapter includes lecture plan callouts with suggested timing. A few notes for instructors adopting this book:

  • Pacing: The book can be covered in a 10-week quarter (as in Stanford’s CS329H) or a 15-week semester. For a quarter: allocate approximately 1 week for chapter 1, 3 weeks for chapter 2, 3 weeks for chapter 3, 1.5 weeks for chapter 4, and 1.5 weeks for chapter 5. For a semester, expand each chapter proportionally.
  • Assumptions and limitations: Throughout the book, callout boxes marked with warnings highlight key assumptions and their limitations. We encourage instructors to treat these as first-class content—understanding when models fail is as pedagogically important as understanding when they work.
  • Assessment: Each chapter includes exercises at three difficulty levels, discussion questions for class participation, and quick check questions for self-assessment. Problem sets are available on the Teaching Materials page.

Prior knowledge

The book assumes knowledge of the fundamentals of statistics, linear algebra and machine learning. Many example code excerpts are written in python, and make experience in the python programming language valuable for readers.

Additional Materials

Every chapter has problems for readers and slides for teaching of the material available. See the Teaching Materials page for all slides, problem sets, and suggested course pathways.

Acknowledgments

Initial versions of this book were compiled as lecture notes to the class CS329H: Machine Learning from Human Preferences at Stanford University taught in Fall 2023 and Fall 2024. We thank Rehaan Ahmad, Ahmed Ahmed, Jirayu Burapacheep, Michael Byun, Akash Chaurasia, Andrew Conkey, Tanvi Deshpande, Eric Han, Laya Iyer, Adarsh Jeewajee, Shreyas Kar, Arjun Karanam, Jared Moore, Aashiq Muhamed, Bidipta Sarkar, William Shabecoff, Stephan Sharkov, Max Sobol Mark, Kushal Thaman, Joe Vincent, Yibo Zhang, Duc Nguyen, Grace Sodunke, Ky Nguyen, Mykkel Kochenderfer, and Zoë Hitzig for helpful conversations.

Citation

Thanks for reading our book! We hope you find this book useful in your research and teaching.

BibTeX citation:
@book{mlhp,
  author    = {Truong, Sang and Haupt, Andreas and Koyejo, Sanmi},
  title     = {{Machine Learning from Human Preferences}},
  year      = {2025},
  publisher = {Stanford University},
  doi       = {},
  note      = {}
}
For attribution, please cite this work as:
S. Truong, A. Haupt, and S. Koyejo. 2025. Machine Learning from Human Preferences. Stanford University.