James Le 9/17/2026

13 Lectures on RLHF: What I Took Away From Nathan Lambert's Course

Read Original

A reward model reads its score off a single token. You push the whole response through a transformer, take the hidden state sitting at the EOS position, run it through a scalar head, and the number that comes out is what the policy will spend the rest of training chasing. One number. For an entire answer.Almost every complaint people have about RLHF (the hedging, the refusals nobody asked for, the

0 comments
13 Lectures on RLHF: What I Took Away From Nathan Lambert's Course

Comments

No comments yet

Be the first to share your thoughts!

Browser Extension

Get instant access to AllDevBlogs from your browser

Top of the Week

No top articles yet