Prime Day

Como cliente Amazon Prime obtén 3 meses de Audible gratis

Diseño de la portada del título Reinforcement Learning from Human Feedback

Reinforcement Learning from Human Feedback

LLM alignment and post-training

Muestra

Suscríbete a la prueba gratuita para poder disfrutar de este libro a un precio exclusivo para suscriptores

Pagar 13,29 € con prueba
Después de los 30 días, 9,99 €/mes. Cancela tu siguiente plan mensual cuando quieras.
Disfruta de más de 90.000 títulos de forma ilimitada.
Escucha cuando y donde quieras, incluso sin conexión
Sin compromiso. Cancela tu siguiente plan mensual cuando quieras.

Reinforcement Learning from Human Feedback

De: Nathan Lambert
Narrado por: Julie Brierley
Pagar 13,29 € con prueba

Después de los 30 días, 9,99 €/mes. Cancela cuando quieras.

Compra ahora por 18,99 €

Compra ahora por 18,99 €

"Reinforcement Learning from Human Feedback: LLM alignment and post-training" helps you understand how modern AI models can be adapted to better match the needs and expectations of their users. Rather than surveying the vast field of reinforcement learning, elite AI researcher Nathan Lambert concentrates exclusively on RLHF and its immediate importance to post-training generative AI models.

This compact book gets right to the point. Early chapters establish the training overview, explain instruction fine-tuning, and build reliable reward models. The middle chapters transition into the heart of alignment, exploring core policy gradient algorithms, Direct Preference Optimization (DPO), and inference-time scaling. Later chapters tackle the messy reality of data, guiding you through preference data collection, synthetic data generation, and the nuances of function calling.

As you go, you will see how these post-training methods work. You will explore common failure modes, such as qualitative over-optimization, reward hacking, and the unreliability of external evaluation comparisons. Difficult concepts like KL regularization, proximal policy optimization, and generative reward modeling are clarified with hands-on experiments.

The book’s seventeen short chapters lay out the core material, while supplements like vocabulary definitions, compute cost management, evaluation variance, and training performance tracking appear in handy appendixes.

About the listener:

For established engineers, AI scientists, and students trying to get a practical foothold in AI model alignment.

About the author:

Dr. Nathan Lambert is a leading AI researcher known for leading post-training at the Allen Institute for AI. With previous experience at HuggingFace, DeepMind, and Meta, he is a passionate advocate for open models. His work focuses on increasing access to, and the understanding of, AI technology—empowering listener to contribute to the advancement of AI outside closed corporate labs.

PLEASE NOTE: When you purchase this title, the accompanying PDF will be available in your Audible Library along with the audio.

©2026 Manning Publications (P)2026 Manning Publications
Historia y cultura
adbl_web_anon_alc_button_suppression_t1
No hay reseñas aún