VUX World

VUX World

By Kane SimmsTechnology
Download on the App Store

VUX World episodes

  • What is text-to-speech and how does it work with Niclas Bergström

    Every voice assistant needs three core components: Automatic Speech Recognition (ASR), Natural Language Understanding (NLU) and Text-to-Speech (TTS). We've already covered what Automatic Speech Recognition is and how it works with Catherine Breslin and in this episode, we're covering the latter, text-to-speech.


    To guide us through the ins and outs of TTS, we're joined by Niclas Bergström, a TTS veteran and co-founder of one of the largest TTS companies on the planet, Readspeaker.


    Text-to-speech is the technology that gives voice assistants a voice. It's the thing that produces the synthetic vocal sound that's played from your smart speaker or phone whenever Alexa or Siri speaks. It's the only part of a voice assistant that you'd recognise. The other core components, ASR and NLU, are silent.


    And, given how we're hard wired for speech - a baby can recognise its mother's voice from the womb - how your voice assistant or voice user interface (VUI) sounds is one of the most important parts of it.


    A voice communicates so much information without us necessarily being aware. Just from the sound of someone's voice, you can infer gender, age, mood, education, place of birth and social status. From the sound of someone's voice, you can decide whether you trust them.


    With voice assistants, voice user interfaces, or any hardware or software that speaks, choosing the right voice is imperative.


    Some companies decide on a stock voice. One of Readspeaker's 90 voices or perhaps Amazon Polly. Others create their own bespoke voice that's fit for their brand.


    We see examples of Lyrbird's voice cloning and we hear Alexa speak every day, so it's easy to take talking computers for granted. Because speaking is natural and easy for us, we assume that it's natural and easy for machines to talk. But it isn't.


    So in this episode, we're going to lift the curtain on text-to-speech and find out just exactly how it works. We'll look at what's happening under the hood when voice assistants talk and see what goes into creating a TTS system.


    Readspeaker is a pioneering voice technology company that provides lifelike Text to Speech (TTS) services for IVR systems, voice applications, automobiles, robots, public service announcement systems, websites or anywhere else. It's been in the TTS game for over 20 years and has in-depth knowledge and experience in AI and Deep Neural Networks, which they put to work in creating custom TTS voices for the world's biggest brands.


    Links

    Visit Readspeaker.com to find out more about TTS services

    And Readspeaker.ai for more information on TTS research and samples

    Hosted on Acast. See acast.com/privacy for more information.

    54 min
  • What is text-to-speech and how does it work with Niclas Bergström

    Every voice assistant needs three core components: Automatic Speech Recognition (ASR), Natural Language Understanding (NLU) and Text-to-Speech (TTS). We've already covered what Automatic Speech Recognition is and how it works with Catherine Breslin and in this episode, we're covering the latter, text-to-speech.


    To guide us through the ins and outs of TTS, we're joined by Niclas Bergström, a TTS veteran and co-founder of one of the largest TTS companies on the planet, Readspeaker.


    Text-to-speech is the technology that gives voice assistants a voice. It's the thing that produces the synthetic vocal sound that's played from your smart speaker or phone whenever Alexa or Siri speaks. It's the only part of a voice assistant that you'd recognise. The other core components, ASR and NLU, are silent.


    And, given how we're hard wired for speech - a baby can recognise its mother's voice from the womb - how your voice assistant or voice user interface (VUI) sounds is one of the most important parts of it.


    A voice communicates so much information without us necessarily being aware. Just from the sound of someone's voice, you can infer gender, age, mood, education, place of birth and social status. From the sound of someone's voice, you can decide whether you trust them.


    With voice assistants, voice user interfaces, or any hardware or software that speaks, choosing the right voice is imperative.


    Some companies decide on a stock voice. One of Readspeaker's 90 voices or perhaps Amazon Polly. Others create their own bespoke voice that's fit for their brand.


    We see examples of Lyrbird's voice cloning and we hear Alexa speak every day, so it's easy to take talking computers for granted. Because speaking is natural and easy for us, we assume that it's natural and easy for machines to talk. But it isn't.


    So in this episode, we're going to lift the curtain on text-to-speech and find out just exactly how it works. We'll look at what's happening under the hood when voice assistants talk and see what goes into creating a TTS system.


    Readspeaker is a pioneering voice technology company that provides lifelike Text to Speech (TTS) services for IVR systems, voice applications, automobiles, robots, public service announcement systems, websites or anywhere else. It's been in the TTS game for over 20 years and has in-depth knowledge and experience in AI and Deep Neural Networks, which they put to work in creating custom TTS voices for the world's biggest brands.


    Links

    Visit Readspeaker.com to find out more about TTS services

    And Readspeaker.ai for more information on TTS research and samples

    Hosted on Acast. See acast.com/privacy for more information.

    54 min
  • The Rundown: Talking volcanos finding gaps you can Read Along to
    This week's top stories:

    Google’s Read Along 

    Vernacular.ai raises $5.1 million led by Exfinity Ventures, Kalaari Capital

    HR AI system, Paradox AI gets 40m funding for replacing the ‘boring’ jobs

    ConverseNow has closed a $3.25 million seed funding round led by Bala Investments

    Omilia, a conversational artificial intelligence platform developer, has raised $20 million in a funding round led by Grafton Capital

    Getting the tone right - automated copy generation has to be retrained in a time of crisis

    Stores may use voice assistants to transform shopping, retail consultant says

    In a world fearful of touch, voice assistants like Amazon's Alexa, Apple's Siri are making our lives easier

    Startup adjusts medical voice assistant for a Zoom world

    France launches AI voice assistant to help coronavirus patients

    Katy Perry announces new album on Alexa

    Spirent approved to test AVS products

    How speech recognition techniques are helping to predict volcanoes’ behaviour

    The Information by James Gleik

    MIDI Sprout

    Learn guitar on Google Assistant

    Hosted on Acast. See acast.com/privacy for more information.

    54 min
  • The Rundown: Talking volcanos finding gaps you can Read Along to
    This week's top stories:

    Google’s Read Along 

    Vernacular.ai raises $5.1 million led by Exfinity Ventures, Kalaari Capital

    HR AI system, Paradox AI gets 40m funding for replacing the ‘boring’ jobs

    ConverseNow has closed a $3.25 million seed funding round led by Bala Investments

    Omilia, a conversational artificial intelligence platform developer, has raised $20 million in a funding round led by Grafton Capital

    Getting the tone right - automated copy generation has to be retrained in a time of crisis

    Stores may use voice assistants to transform shopping, retail consultant says

    In a world fearful of touch, voice assistants like Amazon's Alexa, Apple's Siri are making our lives easier

    Startup adjusts medical voice assistant for a Zoom world

    France launches AI voice assistant to help coronavirus patients

    Katy Perry announces new album on Alexa

    Spirent approved to test AVS products

    How speech recognition techniques are helping to predict volcanoes’ behaviour

    The Information by James Gleik

    MIDI Sprout

    Learn guitar on Google Assistant

    Hosted on Acast. See acast.com/privacy for more information.

    54 min
  • Voice technology and music with Dennis Kooker and Achim Matthes

    Music has been the top use case on smart speakers pretty much from the beginning. Having any song you like at your beckoning call makes playing music around the house easier than ever. And households that play music out loud are, apparently, happier households. 

    It doesn't require too much thought, either. So, discoverability isn't as much of a challenge as with skills, actions and services. If you want to play some Michael Jackson, just ask. 

    Having said that, music consumption habits are advancing. According to Pandora, more people are listening to up-beat, exercise music during lockdown, presumably to exercise to given the gyms are shut. And more people are listening to more ambient music, too, as well as child friendly playlists. People spending time at home and using their music service to relax and entertain the kids respectively. 

    And there's a growing trend moving away from listening to artists and towards listening to playlists. Random compilations of different tunes grouped around a theme. And with smart speakers, we're seeing an insight into people's contexts with the music they ask to play. For example 'play BBQ music' might not be something you'd try and find on Spotify, but you might ask for it from your smart speaker. 

    In the age of playlists, mood music and music on demand, how does a record label make sure that its catalogue of music is found and played on smart speakers? Well, that's what we're going to find out in this episode. 

    In this episode: voice strategy at Sony Music

    We're joined by Dennis Kooker, President, Global Digital Business and US Sales, and Achim Matthes, Vice President, Partner Development, at Sony Music Entertainment. Dennis and Achim walk us through how Sony Music is thinking about voice, some of the behavioural trends they're seeing play out, how they make sure that, when you ask for a Sony Artist song, you get what you've asked for, what's involved in music discoverability, what trends they're seeing and where they see music and voice heading in future. 

    Hosted on Acast. See acast.com/privacy for more information.

    1 hr 7 min
  • The Rundown: Talking elevators and clever Blenders
    Stories covered:

    Talking elevators and the Scottish elevator sketch

    Contact centre voice biometrics

    Google Assistant's voice match

    Facebook Chatbot, Blender, can talk about anything. See some sample dialogues and try it out.

    Juniper research

    NPR Smart Audio Report

    Pandora's listening habit changes

    AI conference on Aminal Farm

    Claire Mitchell on VUX World

    Hosted on Acast. See acast.com/privacy for more information.

    56 min
  • Voice technology for kids with Dr. Patricia Scanlon

    Dustin and Kane are joined by Patricia Scanlon, CEO SoapBox Labs, to discuss how its specialist speech technology for kids is being used and scaled across the globe to help kids learn to read and more.


    In this episode: voice tech for kids

    Imagine being able to have your child read to an iPad and have it tell them how they’re doing. Whether they’re pronouncing the words right and encouraging them to improve.

    Imagine, as a parent or teacher, being able to report on different child’s progress so that you can focus on the real areas that need improving.

    Well, this is what Soapbox Labs enables you to do with its specialist speech technology which you can use to build bespoke applications specifically for kids.

    You might be wondering 'why would I need speech technology specifically for kids?' Well, kids have totally different voices to adults. Their pitch is higher, they don't always pronounce words properly and it changes across the ages. A 5 year old's voice is different to a 10 year old's voice. 

    Most of the speech recognition systems you'll be familiar with are trained on adult voices and don't tend to work too well for kids voices. 

    In this episode, we expand on this and more with a deep discussion on just why voice technology for kids is so important, how the solution was created, what makes it unique and how you can use it to create life changing applications that help kids all over the world learn and entertain themselves. 

    We discuss use cases in education, such as learning to read or learning a new language, as well as leisure, such as speech recognition in toys.

    After listening or watching this episode, you'll not only be equipped with the knowledge you need to create effective voice applications for kids, you'll also have a new appreciation for just how important this kind of technology is, what kind of opportunity exists in creation educational solutions for kids, but also just how important all of this is. 

    About Patricia Scanlon

    Patricia Scanlon is the founder and CEO of SoapBox Labs, the award winning voice tech for kids company. Patricia holds a PhD and has over 20 years experience working in speech recognition technology, including at Bell Labs and IBM.

    Patricia has been granted 3 patents, with two pending. She is an acclaimed TEDx speaker, and in 2018 was named one of Forbes "Top 50 Women in Tech" globally.

    In 2013, inspired by the needs of her oldest child, Patricia envisioned a speech technology to redefine how children acquire literacy. She has successfully raised multiple rounds of both public and private funding to bolster research and product development, and her technical approach has been independently validated by the world's top three academic authorities on speech recognition.

    SoapBox Labs is based in Dublin and has a world class team of 22 employees.

    Links

    Get a free 90 trial of the SoapBox Labs API:

    Visit soapboxlabs.com

    Email [email protected]

    Hosted on Acast. See acast.com/privacy for more information.

    48 min
  • Ethical design with Microsoft's Deborah Harrison

    Deborah Harrison was the very first writer on the Cortana team and defined the personality of Microsoft's digital assistant that is used on over 400m surfaces globally.


    In this talk, we chat to Deborah about the importance of personality and persona in conversational experiences, and the critical responsibility of ethical design.


    Links

    Follow Deborah on Twitter

    Hosted on Acast. See acast.com/privacy for more information.

    55 min
  • Measuring your share of voice with Andy Headington

    Voice search is growing. As smart speaker adoption increases and people use the voice assistants on their phones more and more, voice is playing an increasing role in more customer journeys.

    But what's your share of voice? How often does your company 'rank' for the search terms you'd like to rank for?

    Amazon and Google are tight-lipped about voice search volumes and don't offer any way of tracking voice search performance for websites, skills or actions.

    Thankfully, Andy Headington and his team at Adido, has created a tool that lets you do just that.

    In this episode, we speak to Andy about how you can measure your share of voice, and how you can find ways of identifying insights that will enable you to improve you voice search performance.

    Links

    Try Share of Voice 

    Hosted on Acast. See acast.com/privacy for more information.

    48 min
  • What is automatic speech recognition and how does it work? With Catherine Breslin
    What is speech recognition and how does it work?

    Automatics speech recognition (also known as ASR) is a suite of technology that takes audio signals containing speech, analysis it and converts it into text so that it can be read and understood by humans and machines. It's the technology that makes voice assistants like Amazon Alexa able to understand what a user says.

    There's obviously a whole lot more to it that than, though. So, in today's episode, we're speaking to one of the most knowledgable and experienced speech recognition minds the world has to offer, Catherine Breslin, about just exactly what's going on under the hood of automatic speech recognition technology and how it actually works.

    Catherine Breslin studied speech recognition at Cambridge, before working on speech recognition systems at Toshiba and eventually on the Amazon Alexa speech recognition team where she met the Godfather of Alexa, Jeff Adams. Catherine then joined Jeff at Cobalt Speech where she currently creates bespoke speech recognition systems and voice assistants for organisations.

    In this episode, you'll learn how one of the fundamental voice technologies works, from beginning to end. This will give you a rounded understanding of automatic speech recognition technology so that, when you're working on voice applications and conversational interfaces, you'll at least know how it's working and then be able to vet speech recognition systems appropriately.


    Presented by Sparks

    Sparks is a new podcast player app that lets you learn and retain knowledge while you listen.

    The Sparks team are looking for people just like you: podcast listeners who're also innovative early adopters of new tech, to try the beta app and provide feedback.

    Try it now at sparksapp.io/vux

    Hosted on Acast. See acast.com/privacy for more information.

    53 min

About VUX World

From the publisher's feed

Interviews with the best brains in AI, sharing how to improve customer experience and business operations using emerging AI technologies such as voice AI, conversational AI, NLP, Large Language…

More shows like VUX World

Pivot by New York Magazine

Pivot

9,620 Listeners

Decoder with Nilay Patel by The Verge

Decoder with Nilay Patel

3,154 Listeners

Pod Save America by Pod Save America

Pod Save America

87,077 Listeners

The Diary Of A CEO with Steven Bartlett by DOAC

The Diary Of A CEO with Steven Bartlett

8,502 Listeners

Off Menu with Ed Gamble and James Acaster by Plosive

Off Menu with Ed Gamble and James Acaster

2,657 Listeners

Hard Fork by The New York Times

Hard Fork

5,554 Listeners

Huberman Lab by Scicomm Media

Huberman Lab

29,187 Listeners

The Mel Robbins Podcast by Mel Robbins

The Mel Robbins Podcast

19,271 Listeners

The Rest Is Politics: Leading by Goalhanger

The Rest Is Politics: Leading

781 Listeners

People of AI by Google

People of AI

11 Listeners

Prof G Markets by Vox Media Podcast Network

Prof G Markets

1,448 Listeners