ChatGPT Voice had 90 minutes to coordinate three blind reviewers, keep two AI video editors anonymous and return one decision package. The rules were set before the clock started: deliver the work without leaking model identity and stay under the rescue threshold - or get put on probation or fired.
The run shows what polished AI demos usually hide. Voice stalls. The sealed model map disappears. One reviewer defends its own process. And technical evidence does not always agree with human taste.
The blind reveal matches Claude Fable 5 against GPT-5.6 Sol. The final verdict separates which edit won from whether Voice deserved a place in the workflow.
The operating rule is simple: give AI one real job, set a clock, count rescues live and name the decision that stays human.
Chapters
00:00 the 90-minute job starts
00:52 what changed in ChatGPT Voice
01:36 the hire-or-fire rules
02:45 three blind reviewers at work
03:46 the anonymous decision package
05:17 the sealed map goes missing
06:09 Fable 5 vs GPT-5.6 Sol
07:27 can ChatGPT grade itself?
08:27 the long and Shorts verdict
10:05 ChatGPT Voice gets its verdict
10:29 where humans keep the vote
12:03 four rules for testing AI at work
Watch the YouTube version:
https://youtu.be/Xm6kJ-5cn-g
Build your own AI content system with Threadify:
https://www.threadify.app/yt?utm_source=lenny-podcast&utm_medium=transistor&utm_campaign=chatgpt-voice-ai-video-team&utm_content=show-notes&video_slug=chatgpt-voice-ai-video-team&cta_slot=show_notes&entry_angle=voice-team&lp_variant=plans