
Sign up to save your podcasts
Or


As a continuation of Episode 238, I explain some effective and fun attacks to conduct against LLMs. Such attacks are even more effective on models served locally, that are hardly controlled by human feedback.
Have great fun and learn them responsibly.
References
https://www.jailbreakchat.com/
https://www.reddit.com/r/ChatGPT/comments/10tevu1/new_jailbreak_proudly_unveiling_the_tried_and/
https://arxiv.org/abs/2305.13860
By Francesco Gadaleta4.2
7272 ratings
As a continuation of Episode 238, I explain some effective and fun attacks to conduct against LLMs. Such attacks are even more effective on models served locally, that are hardly controlled by human feedback.
Have great fun and learn them responsibly.
References
https://www.jailbreakchat.com/
https://www.reddit.com/r/ChatGPT/comments/10tevu1/new_jailbreak_proudly_unveiling_the_tried_and/
https://arxiv.org/abs/2305.13860

31,971 Listeners

7,584 Listeners

1,706 Listeners

1,091 Listeners

623 Listeners

585 Listeners

823 Listeners

301 Listeners

99 Listeners

9,161 Listeners

207 Listeners

306 Listeners

5,512 Listeners

228 Listeners

1,104 Listeners