
Sign up to save your podcasts
Or


The Future is Spoken presents Jonathan Bloom as this week’s guest.
Jonathan is a Senior Conversation Designer for Google, focusing on the Google Assistant. Jon was previously the UX Research Lead for Jibo, Inc., creators of the social robot of the same name. Jon was also Senior Voice User Interface Manager for Nuance Communications, where he sat on Nuance’s Innovation Steering Committee. Over his 20-year career, Jon has designed graphic, speech, and multimodal user interfaces for robots, IVR’s, dictation software, cars, and mobile applications. Jon holds a Ph.D. in Cognitive Psychology from the New School for Social Research.
Starting with their own experiences, they end up discussing Standardizing Voice Experiences.
Tune in now!
Conversation Highlights:
[00:03:39]: Jonathan's journey from Nuance to Google
[00:10:45] : Multimodal Design and how it helps
[00:13:03] : Why do people expect Emotionally Intelligent Digital Assistants?
[00:17:52]: Does anthropomorphism lead people to expect more?
[00:22:37]: People are expecting more from the assistance. Where do we draw that balance?
[00:31:40]: How much is too much personality?
[00:45:17]: Script writing keeps Jonathan inspired
[00:48:00] Jonathan encourages new conversation designers to read these books
Cathy Pearl's - Designing Voice User Interfaces
Learn more about Jon at
- LinkedIn
- Twitter
If you enjoyed this episode of The Future is Spoken Podcast, then make sure to subscribe to our podcast.
Follow Shyamala Prayaga at @sprayaga
The Future is Spoken presents Jon Stine as this week’s guest.
Jon Stine is the Executive Director of The Open Voice Network (OVN), a non-profit global association dedicated to bringing the benefit of standards to the world of artificial intelligence-enabled voice assistance. The OVN is a Directed Fund of The Linux Foundation.
He brings to this role more than 30 years of global leadership in the commerce and technology industries.
Jon’s retail industry knowledge was first shaped in the womenswear apparel business, as he headed sales of a well-known national brand to leading US department and specialty stores. In 2000, he joined the Intel Corporation to create and head its first global outreach to the retail and consumer goods industry. In the years that followed, he was a co-founder of the Metro Group Future Store Initiative in Germany, and the Pan-Pearl River Delta Initiative that first brought digital transparency to the China-to-US supply chain.
He joined Cisco Systems retail-CPG consulting team in late 2006, and later headed Cisco’s North America consulting practice for retail-CPG. In 2014, he returned to Intel as the Global Enterprise Sales General Manager for the retail, hospitality, and consumer goods industries. He stepped away from Intel in 2019 to build The Open Voice Network.
Through the years, Jon has worked directly with customers across the Americas, Western and Central Europe, Middle East, India, China, and Japan, as well as delivery partners in hardware, software, services, and consulting.
Jon resides in Portland, Oregon, USA.
Starting with their own experiences, they end up discussing Standardizing Voice Experiences.
Tune in now!
Conversation Highlights:
[00:04:05]: Story of Open Voice Network Inception
- Listen how a group of met over coffee and ended up starting open voice network.
[00:06:39]: Voice is at the same stage as User Experience was before 2006.
- Jon and around 180 volunteers are working together to bring voice standards to life.
[00:07:19]: Value of Standards
[00:12:05] How do we make Voice trustworthy?
[00:13:29] Voice Standards needs regulation like Accessibility to empower users
[00:16:06] What is being done to make Voice more Ethical?
[00:17:32] Privacy and Security. How do we protect it? And what values might we promote?
[00:27:12] Voice is a crowed space, will users select Voice assistant based on Standards?
[00:33:43] Designers and Developer can all add value to Open Voice Network.
Learn more about Jon at
- Open Voice Network
If you enjoyed this episode of The Future is Spoken Podcast, then make sure to subscribe to our podcast.
Follow Shyamala Prayaga at @sprayaga
The Future is Spoken presents Jeff Adams as this week's guest. Jeff has been leading prominent speech & language technology research for more than 20 years. Until 2009, he worked at Nuance / Dragon, where he was responsible for developing and improving speech recognition for Nuance's "Dragon" dictation software. He presided over many of the critical improvements in the 1990s and 2000s that brought this technology into the mainstream and enabled widespread consumer adoption.
After leaving Nuance, Jeff joined Yap, a company specializing in voicemail-to-text transcription. He assembled a strong team of 12 speech scientists who, within two brief years, were able to beat all competitors on an unbiased test set. They also matched the performance of a competitor who used (off-shore) human transcription.
Yap's success caught the interest of Amazon, who wanted to jump-start their new speech & language research lab. Upon acquisition, Jeff led efforts to build one of the industry-leading speech & language groups. His Amazon team developed products such as the Echo, Dash, and Fire TV. Jeff left Amazon in 2014 to found Cobalt Speech and Language.
Starting with their own experiences, they end up discussing crafting natural conversations for the bot.
Tune in now!
Conversation Highlights:
[00:28] The journey to Voice…..
● Jeff has been working on speech technology for almost 26 years now. He started with a small speech company in Boston.
● He ended up working at Amazon on Alexa before it was launched.
● Cobalt work with companies that are looking for speech-related technologies. It is a company that licenses technology and also customizes it.
[03:58] Can anyone design natural conversations?
● Jeff explains that it is an art to designing natural conversations. One way to approach this is to assume the system as a human.
● He also talks about designing a system that can cater to everyone's needs.
● Giving the user what they are looking for is what matters! Jeff touches on building a system that responds appropriately to all the different ways people ask for something.
[15:10] Creating a bot despite the lack of resources?
● Jeff divulges different ways of creating a natural voice application despite the lacks of resources.
● A slow launch with a lot of beta testing is the key.
[18:41] NLP v/s NLU
● Jeff explains the difference between Natural Language Processing and Natural Language Understanding.
● NLP is a broad umbrella term referring to any computerized processing of human language, while NLU is a subset.
● What are the uses of NLU?
● Spoken Language Understanding is Automatic Speech Recognition(ASR) + NLU.
● Jeff also touches on ensuring a system that understands what the user says.
[34:04] The secret sauce for the ASR systems to work better!
● Jeff elaborates on the different approaches for the ASR systems to work together.
● How can we design a speech recognition system that understands the users in the natural environment?
[41:41] What are the Best Practices while designing a speech system?
[44:15] Must Listen
● Jeff's piece of advice for someone trying to get into Speech Recognition.
Learn more about Jeff at
● Or email him at [email protected].
If you enjoyed this episode of The Future is Spoken Podcast, then make sure to subscribe to our podcast.
Follow Shyamala Prayaga at @sprayaga
The Future is Spoken presents Elaine Lee as this week's guest.
Elaine Lee is a designer specializing in artificial intelligence and machine learning, with a focus on ethical AI. She's currently a Principal Product Designer at Twilio on the AI team, and previously led the design of eBay's AI-powered shopping assistant on Facebook Messenger and Google Assistant.
Starting with their own experiences, they end up discussing the building a trustworthy voice assistant
Tune in Now!
Conversation Highlights:
Learn more about Elaine at
If you enjoyed this episode of The Future is Spoken Podcast, then make sure to subscribe to our podcast.
Follow Shyamala Prayaga at @sprayaga
The Future is Spoken presents Rupal Patel as this week's guest. Rupal is the Founder and CEO of VocaliD, a voice Al company that creates unique synthetic voices. Unlike conventional methods, VocaliD's award-winning technology generates high-quality, natural-sounding voices within hours, not months. They leverage cutting-edge machine learning techniques, proprietary Voice blending algorithms, and our crowdsourced Voicebank dataset to enable brands and individuals to be heard in a voice that is uniquely theirs. Vocal Id is a spin-out from her research lab at Northeastern University. She is a tenured professor in the Department of Communication Science and Disorders and the Khoury College of Computer Sciences.
Starting with their own experiences, they end up discussing synthetic voices.
Tune in Now!
Conversation Highlights:
[00:23] The journey to Synthetic voices…..
● Rupal works on making customized synthetic voices for individuals as well as for companies. She started the mission to create voices for people who couldn't speak.
● She also explains how the world of Voice is touching the sky right now.
[03:32] Identifying the problem…….
● Rupal explains the reason behind creating the Vocal ID. She divulges the problems she identified while researching people with speech impairment.
● People with limited speech capabilities still can control the prosody of their Voice.
● What does it take to create a natural-sounding voice?
[11:47] Tuning the prompt according to your need.
● Rupal speaks about the different ways to tune the prompt for the pitch or speed or even the tone. The end-to-end synthesis methodologies allow controlling Speech differently.
● They have also started implementing a new method to make a change at the word level. She is also excited about some of the style modifications.
[20:09] The importance of Natural Sounding Voice
● She elaborates that almost every way we are consuming information is through our ears. Because of so much audible capability, you need to have a natural voice.
[22:18] What secret skill do you need to enter the 'Text to Speech' world?
● She touches on the skills you need to enter the world of Speech and design natural sounding voices.
● Linguistics is becoming the heart of Voice.
[26:11] Researching is the most crucial aspect of everything.
● Rupal explains that apart from doing experiments on building up the voices and making them sound more natural, they are also doing listening perception experiments to understand how consumers have different preferences for Voice.
● She also touches on how they ensure that the quality remains up to par and the operating system's role in amplifying the quality.
[40:36] How is Vocal ID different from others?
● Vocal ID is focused on customized Voice as supposed to specific libraries that other companies possess.
● Machine learning can get you to 90% of the way, but you will require an understanding of Speech to reach that last mile.
[46:40] Must Listen
● Rupal's piece of advice for someone trying to get into the world of Voice.
Special Reminder:
Celebrate The Diversity of Human Voices! Will You Share your Voice?
Join others from around the world in sharing the gift of Voice. Register today.
Learn more about Rupal at
● Vocalid.ai
● Vocalid.ai/voicebank
If you enjoyed this episode of The Future is Spoken Podcast
The Future is Spoken presents Marco Pasqua as this week's guest. Marco is Co-founder of LIKE ventures and an award winning speaker and Entrepreneur.
Starting with their own experiences, they end up discussing the act of accessibility and inclusion in Voice Design.
Tune in Now!
Conversation Highlights:
[00:02:11]: 2008 Lay off which lead to a career in Accessibility
[00:06:06] LIKE Venture and Marco's Journey towards building an accessible future
[00:13:02] Accessibility and Meaning Access are two different things
"for example, for a building to have a ramp outside a door and then say, Oh, people can get inside the front door. Right. But then ask yourself, where is that ramp located? Is it actually at the front door or is it at the back of the building near a trash can or near the dumpster? So that you're saying, well, by the way, we have accessibility, but it means that you have to come through a different entrance, but thats ok.
Right. Because you're just like everybody else. No meaningful access is when every single person who's expected to use a product or service can use it in exactly the way it's intended without having to feel like they have to adapt themselves to meet the function of the product. But rather that the product is already thinking about how everyone, whether you're nine or 90 can use it."
[00:18:22]: Accessibility by Design
[00:23:04]: Making Voice Application more accessible
[00:45:11]: Must Listen
Learn more about Marco at
If you enjoyed this episode of The Future is Spoken Podcast, then make sure to subscribe to our podcast.
Follow Shyamala Prayaga at @sprayaga
The Future is Spoken presents Maaike Coppens as this week's guest. Maaike is an international Conversation Design Expert, speaker and co-author of the Voice UX Workbook. With a successful background in linguistics and UX design, Maaike found conversational design the perfect blend.
Over the last couple of years, Maaike has worked with award-winning agencies, large enterprises, and innovative companies worldwide. Eager to be part of a team with an opinionated take on Conversation Design, Maaike recently joined Greenshoot Labs - a chatbot and applied AI agency in the UK - as their Head of Conversational UX.
Tune in Now!
Conversation Highlights:
[02:54] The journey to Conversational UX….
● Maaike explains that we are constantly flipping the coin in the wrong way when designing experiences.
● She had an academic path in linguistics. She then got deeper into UX. When conversational designing became more widespread, it was an ideal mix for her.
[05:35] Linguistics as a career?
● Maaike shares her perspective on choosing Linguistics as a career.
● It's great to see how conversation design has helped her think differently about human to human Conversations.
[08:44] The most important traits that a conversational Al must-have.
● Maaike explains that being flexible and being inclusive are essential traits. She also feels that relatability is also a necessary aspect for conversational Al or voice assistant.
● Accessibility and inclusivity are related, but they are not mutually exclusive.
[21:34] Enabling the disabled with Voice....
● She elaborates that when you are designing Voice assistants, empathy is critical. It is necessary to get those insights and research before doing any empathy exercises.
● Empathy becomes very important when backed up by user research.
● She still sees loads of difficulties around language proficiency. She says that instead of finding the next big thing, we should perfect the ones available.
[35:05] Reality of Voice Assistants
● She feels that if you adapt your human behaviour to technology, it is a sign that the technology is already broken.
● The industry is more reactive than being proactive in designing an inclusive solution. When amazon launched smart speakers, they never thought about disabled people.
[42:30] How can we design an inclusive solution?
● Maaike divulges what the designers can do to design an inclusive solution. She further instructs to get out of the bubble.
● She speaks that listening is the one vital skill that you need to have as a conversational designer.
[47:02] No Al without IA(Information Architecture)
● Maaike elaborates that Information Architecture is a significant part of conversational designing. It is about mapping the information available to you.
● Why do we forget the basics of UX?
● With the Information Architecture, researching and bringing in some of the UX processes into conversational designing can significantly change the experience.
[54:50] Must Listen
● Maaike's piece of advice for someone trying to get into the world of Voice.
Learn more about Maaike at
● GreenShoot Labs
If you enjoyed this episode of The Future is Spoken Podcast, then make sure to subscribe to our podcast.
Follow Shyamala Prayaga at @sprayaga
The Future is Spoken presents Greg Bennett as this week’s guest. Greg is Conversation Design Principal at Salesforce, leading the company’s first dedicated Conversation Design practice since its inception. As a linguist, Greg focuses his work on empowering businesses to create chatbots that feel natural and helpful, build user trust, and meet customer expectations for conversational behavior. Greg works with Salesforce’s Product teams, customers, and partners to tailor their conversation designs for cross-cultural differences across channels and user populations, as well as how to effectively express personality or conversational style.
Starting with their own experiences, they end up discussing about designing conversational bots for enterprise.
Tune in Now!
Conversation Highlights:
[00:32] From Linguistics to Conversational AI
[12:26] Benefits of using Enterprise Bots
[16:34] Creating your Bot…..
[37:43] Choosing the right options!
[46:08] Ensuring to stick with your branding traits
[50:16] Measuring the User's Trust
"The Future is Spoken" presents Dr Maria Aretoulaki as this week's guest. Maria is a Voice First Veteran, having been designing Voice User Interfaces (VUIs) since 1996, long before Voice Assistants, but also before telephone self-service and Speech IVRs were mainstream. She got into Voice Design through a Post-Doc in Spoken Dialogue Management for Speech Recognition applications, after getting a PhD and an MSc in NLP and Machine Learning and, earlier, a BA (Hons) in Linguistics & English.
She has held Senior VUI | Voice Design and Technical Project Management positions both in Academia and in Industry in the UK and Germany. In 2008, she founded her own VUI & Conversational Design Consultancy, DialogCONNECTION Limited. In that time, Maria has provided her VUI Design and Speech Recognition expertise to organizations in Europe, the US and Asia, including APPLE, SAMSUNG, VODAFONE, SKY, TALK TALK, EMIRATES NBD, the NHS and the European Commission. Her Voice designs, call flows, dialogue scripts and tuning recommendations have saved her clients up to $10 million annually.
Originally from Greece, Maria has spent the last 30 years between the UK and Germany, and has provided multilingual and culturally appropriate Voice Design services for English, German, French, Italian, Spanish and Greek.
Tune in now!
Conversation Highlights:
[00:25] From Linguistics to NLP to VUI Design
● Maria originally studied Linguistics and English Literature in Greece, so she naturally wanted to go and study and work in the UK.
● She was looking for Sponsorship for her Masters, when she bumped into the field of Machine Translation and NLP.
● During her PhD studies in 1993, she discovered the world of Artificial Neural Networks and got fascinated by their potential and decided to apply them to Automatic Text Summarization.
● It was through a Post-Doc in Spoken Dialogue Management for Speech IVRs that she got into the world of Voice and Speech Recognition back in 1996. From then on, she has been a VUI Designer!
[07:20] SIRI comes into play!
● Maria explains how the iPhone was like having a full computer in our pockets and how SIRI was the beginning of a new era, making Speech Recognition and Voice mainstream.
● She feels very proud about the Voice field, which she considers like her "baby" growing up to be an adult!
[09:40] Explainable VUI
● Maria coined the term "Explainable VUI" amid the myriad of Voice applications and Voice Assistant skills / actions / capsules designed by programmers or marketing people.
● "Explainable VUI" means to design a Human-Computer interface bearing in mind both the complexities and imperfections of human language and the limitations of the technology (ASR / speech recognition / NLP).
● A lot of her work with various companies and organizations creating VUI Designs from scratch or reviewing existing ones is carefully crafting system prompts.
● She stressed the importance of knowing how the background technology works.
[52:14] Must Listen
● Maria's pieces of advice for aspiring Conversational Designers and people new in this field on how to start, learn, get ahead, and flourish.
Learn more about Maria here:
● Company website, DialogCONNECTION Ltd http://dialogconnection.com/
● LinkedIn https://www.linkedin.com/in/aretoulaki-
● Twitter https://twitter.com/dialogconnectio and https://twitter.com/ar3toul4ki
● Blog https://aretoulaki.wordpress.com/
The Future is Spoken presents Obaid Ahmed as this week's guest. Obaid is founder and CEO of Botmock, the leading chatbot design collaboration platform. Before Botmock, Obaid co-founded a design consultancy firm OAK where he led the technology and design teams to deliver over 400 projects. Today, Obaid concentrates on building on the success of Botmock to make it a real driver in building the next generation of conversational apps.
Starting with their own experiences, they discuss how the startups can have their VUI and share some powerful tips for aspiring Conversational Designers.
Tune in now!
Conversation Highlights:
[00:16] What is Botmock?
● Obaid is a software developer by training and had spent some time working at Blackberry and IBM. And then he ran a consulting agency where they built all sorts of software for mobile and web for many different companies.
● Botmock is a design and prototyping tool. Essentially, they enable teams who build conversational experiences to go from an idea from planning to developer handover. Understand what exactly will be built, how the experience will unfold eventually, and be able to test it before they go into the development phase.
● Obaid unfolds whether Botmock is designers centric tool or a developer-centric tool?
[05:31] Conversation designing is a team sport.
● Obaid explains how startups(Conversation Design company) should go about building a conversation design team?
● He elaborates on the importance of research. Many teams that they work within enterprise settings do have very specific people who are doing training data research.
● He also explains how they identify different utterances. The best thing is to look at what the user's already saying in terms of data sets.
● Successful bots can bring the users slowly into their world and teach them how to interact further.
[14:36] What about the startups?
● Obaid speaks about how the startups can have their voice automation despite the lack of data.
● He recommends creating some prototypes. Don't worry about the depth of those prototypes, but make some samples. Come up with some use cases that you think are your customers are looking to tackle.
● He also sheds light on the need to build early and then test as soon as possible, and explains how they can have an in-depth analysis?
[26:15] Tackling the problems
● Obaid touches on how the conversation designer gets to the data to synthesize it?
● One of the critical things that are different and hard to build from a design perspective, especially Conversational design is usually in production after they roll chatbox out. So there's very little in-depth analysis they can do at a design level.
● What kind of skill sets should the writer have?
[31:56] The importance of Linguistics and Psychologists in Conversational designing
● Obaid elaborates that Language plays a significant role in making something easy. He also explains how they are very crucial to have in Conversational designing, especially the linguistics.
● Emoji in voice conversations? (Text + Pitch = Satisfied Conversation)
[37:46] It's all about Text to Speech tuning!
● Obaid speaks that they are working on some of the advance sorts of customization features as well. But it depends on the engine to engine, and there's a lot of variations on it, and soon they would probably be able to give more control to designers.
● He also touches on what' Engine to Engine' actually mean?
[40:07] Who is responsible for tuning the prompts?
● Obaid also explains if the designers and writers should know about the technology in order to drive successful communications and collaboration.
[43:49] Must Listen
From the publisher's feed