
Sign up to save your podcasts
Or


Controlled decoding (CD) is a novel off-policy reinforcement learning method that uses a value function called a prefix scorer to steer autoregressive generation towards high reward outcomes. CD is effective in controlling language models and can handle multiple rewards without additional complexity. It can also be applied in a blockwise fashion at inference-time, making it a promising approach for aligning language models.
https://arxiv.org/abs//2310.17022
YouTube: https://www.youtube.com/@ArxivPapers
TikTok: https://www.tiktok.com/@arxiv_papers
Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692476016
Spotify: https://podcasters.spotify.com/pod/show/arxiv-papers
By Igor Melnyk5
33 ratings
Controlled decoding (CD) is a novel off-policy reinforcement learning method that uses a value function called a prefix scorer to steer autoregressive generation towards high reward outcomes. CD is effective in controlling language models and can handle multiple rewards without additional complexity. It can also be applied in a blockwise fashion at inference-time, making it a promising approach for aligning language models.
https://arxiv.org/abs//2310.17022
YouTube: https://www.youtube.com/@ArxivPapers
TikTok: https://www.tiktok.com/@arxiv_papers
Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692476016
Spotify: https://podcasters.spotify.com/pod/show/arxiv-papers

953 Listeners

1,971 Listeners

438 Listeners

112,700 Listeners

10,063 Listeners

5,531 Listeners

214 Listeners

51 Listeners

99 Listeners

473 Listeners