
Sign up to save your podcasts
Or


Controlled decoding (CD) is a novel off-policy reinforcement learning method that uses a value function called a prefix scorer to steer autoregressive generation towards high reward outcomes. CD is effective in controlling language models and can handle multiple rewards without additional complexity. It can also be applied in a blockwise fashion at inference-time, making it a promising approach for aligning language models.
https://arxiv.org/abs//2310.17022
YouTube: https://www.youtube.com/@ArxivPapers
TikTok: https://www.tiktok.com/@arxiv_papers
Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692476016
Spotify: https://podcasters.spotify.com/pod/show/arxiv-papers
By Igor Melnyk5
33 ratings
Controlled decoding (CD) is a novel off-policy reinforcement learning method that uses a value function called a prefix scorer to steer autoregressive generation towards high reward outcomes. CD is effective in controlling language models and can handle multiple rewards without additional complexity. It can also be applied in a blockwise fashion at inference-time, making it a promising approach for aligning language models.
https://arxiv.org/abs//2310.17022
YouTube: https://www.youtube.com/@ArxivPapers
TikTok: https://www.tiktok.com/@arxiv_papers
Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692476016
Spotify: https://podcasters.spotify.com/pod/show/arxiv-papers

977 Listeners

1,993 Listeners

443 Listeners

113,121 Listeners

10,254 Listeners

5,576 Listeners

221 Listeners

51 Listeners

101 Listeners

475 Listeners