In this episode of Inference, I sit down with Mati Staniszewski, co-founder and CEO of ElevenLabs, to explore the future of AI voice, real-time multilingual translation, and emotionally rich speech synthesis. We dive into what still makes dubbing hard, how Lex Fridman's podcast was localized, and what it takes to preserve tone, timing, and emotion across languages. Mati shares why speaker detection in noisy rooms is tricky, how fast their models really are (70ms TTS!), and the deeper strategy behind partnering with creators and enterprises to show – not just tell – what the tech can do.
What needs to happen for natural, free-flowing multilingual conversations to become reality? Mati says: give it two or three years. Watch to learn more!
Mati Staniszewski, co-founder and CEO at ElevenLabs
Website: https://elevenlabs.io/
https://www.turingpost.com/p/mati
0:00 Real-time voice translation
0:11 Language barriers and AI
0:29 Why ElevenLabs started
0:45 Preserving emotion in translation
1:06 Tech challenges in real-time translation
2:32 Speaker diarization and emotional nuance
3:04 Speech-to-text to LLM to TTS pipeline
5:51 Concrete examples: healthcare & customer support
7:05 Real-time AI dubbing use cases
8:02 Lex Fridman podcast dubbing challenge
13:01 Audio model performance & latency
14:44 Conversational AI & multimodal future
16:57 Product vs research focus at ElevenLabs
20:42 Why ElevenLabs didn't open source (yet)
21:28 Strategy: creators, enterprises & brand building
Turing Post is a newsletter about AI's past, present, and future. Publisher Ksenia Semenova explores how intelligent systems are built—and how they’re changing how we think, work, and live.
Sign up: Turing Post: https://www.turingpost.com
Mati: https://x.com/matistanis
ElevenLabs: https://x.com/elevenlabsio
Turing Post: https://x.com/TheTuringPost
Ksenia: https://x.com/Kseniase_
TuringPost: https://www.linkedin.com/company/theturing...
Ksenia: https://www.linkedin.com/in/ksenia-se
SUBSCRIBE TO OUR CHANNEL, SHARE YOUR FEEDBACK