
Sign up to save your podcasts
Or


In this post, I'll list all the areas of control research (and implementation) that seem promising to me.
This references framings and abstractions discussed in Prioritizing threats for AI control.
First, here is a division of different areas, though note that these areas have some overlaps:
This division is imperfect and somewhat [...]
---
Outline:
(01:43) Developing and using settings for control evaluations
(13:51) Better understanding control-relevant capabilities and model properties
(28:21) Developing specific countermeasures/techniques in isolation
(36:01) Doing control experiments on actual AI usage
(40:43) Prototyping or implementing software infrastructure for control and security measures that are particularly relevant to control
(41:58) Prototyping or implementing human processes for control
(46:50) Conceptual research and other ways of researching control other than empirical experiments
(49:37) Nearer-term applications of control to practice and accelerate later usage of control
---
First published:
Source:
Narrated by TYPE III AUDIO.
By LessWrongIn this post, I'll list all the areas of control research (and implementation) that seem promising to me.
This references framings and abstractions discussed in Prioritizing threats for AI control.
First, here is a division of different areas, though note that these areas have some overlaps:
This division is imperfect and somewhat [...]
---
Outline:
(01:43) Developing and using settings for control evaluations
(13:51) Better understanding control-relevant capabilities and model properties
(28:21) Developing specific countermeasures/techniques in isolation
(36:01) Doing control experiments on actual AI usage
(40:43) Prototyping or implementing software infrastructure for control and security measures that are particularly relevant to control
(41:58) Prototyping or implementing human processes for control
(46:50) Conceptual research and other ways of researching control other than empirical experiments
(49:37) Nearer-term applications of control to practice and accelerate later usage of control
---
First published:
Source:
Narrated by TYPE III AUDIO.

26,354 Listeners

2,444 Listeners

9,124 Listeners

4,154 Listeners

92 Listeners

1,595 Listeners

9,902 Listeners

90 Listeners

506 Listeners

5,471 Listeners

16,051 Listeners

540 Listeners

131 Listeners

95 Listeners

521 Listeners