13 Lectures on RLHF: What I Took Away From Nathan Lambert's Course
Read OriginalA reward model reads its score off a single token. You push the whole response through a transformer, take the hidden state sitting at the EOS position, run it through a scalar head, and the number that comes out is what the policy will spend the rest of training chasing. One number. For an entire answer.Almost every complaint people have about RLHF (the hedging, the refusals nobody asked for, the
Comments
No comments yet
Be the first to share your thoughts!
Browser Extension
Get instant access to AllDevBlogs from your browser
Top of the Week
No top articles yet