5 useful things you'll learn in my new post-training textbook (shipping now!)
Lambert’s new book, “Human-Feedback Reinforcement Learning: Alignment and Post-Training Large Language Models”, has been published by Manning and is now available for delivery. Written based on the content of his long-running documentation website, the book aims to explain post-training methods, trade-offs, and common misconceptions in an intuitive way, suitable for readers with a bachelor’s degree in computer science or higher. The book covers topics such as rejection sampling and result-reward models, includes explanations of reinforcement learning algorithms like PPO by about 25%, and provides complete courses, code libraries, and model comparisons. The book is currently available at a 50% discount; the discount code is PBLambert, valid until August 19.