Skip to content

May 18, 2023

Voice Interface Technology: What You Need to Know

What a voice user interface is, the three technologies behind it, how it got here since 1779, and how businesses put it to work today.

Voice interface technology is everywhere, and it may well be in your home. Voice assistants such as Alexa, Siri and Google Assistant control more than 3 billion devices (opens in a new tab), a figure expected to more than double by 2023. These familiar names are the public face of voice user interfaces, or VUIs.

But VUIs are not only for smart speakers. The technology improves essential business processes, from hands-free control on the production line to booking a meeting room while on the move. Here is what business decision makers need to know about voice interface technology: what it is, what it does, and how it can help a company reach its goals.

Voice user interface: a definition

In computing, the user interface (opens in a new tab) is the hardware and the software that let a person interact with a machine. It can include a keyboard, a mouse or a touchscreen, along with the software that draws the on-screen elements to click, drag or type into.

The personal computers of the early 1980s were controlled through a text-only interface (opens in a new tab). Users had to type very specific text commands to make the machine do anything at all. Graphical user interfaces (GUI), such as the one on the Macintosh, Apple’s 1984 revolution (opens in a new tab), replaced those demanding commands with visual icons that users could manipulate with a mouse. That is how the desktop metaphor we still work with today was born.

Like the graphical user interface, and like the command line before it, the VUI gives users a new way to pass commands to digital devices, but this time without a screen, a keyboard or a mouse. In short, a VUI can be defined as a technology that lets people interact with digital devices using their voice.

The elements of a voice user interface

A pure voice user interface takes its input and delivers its output using speech alone. You can compare it with a bimodal user interface, which combines voice interaction with another medium, such as text shown on a screen. The smart television, which lets you turn the volume down with a spoken command, is an example of a bimodal user interface. It is a voice-controlled device, but it will still draw the volume bar on the screen, and that bar will go down as you lower the volume.

For now, we will keep to the end-to-end VUI, a system that takes spoken commands and answers those commands with synthetic speech. In a voice-only VUI, three technologies come together to create an ever more natural interaction between people and their tools:

  1. Automatic speech recognition. The first task of the VUI is to transcribe the spoken command into a machine-readable format, usually text. In the very early days of the VUI, around the middle of the 2000s, automatic speech recognition was limited to a defined list of commands, and the first engines were easily thrown off by the modulation, the tone and the accent of the person speaking. That is no longer the case, as we will discuss in the third point of this list.
  2. Text to speech. A voice-controlled device turns a spoken command into text, carries out that command and prepares an answer, a written text answer. A text-to-speech engine turns that text into synthetic speech to close the loop of the interaction with the user. Text-to-speech quality varies widely, even in the voice user interfaces of today, from robotic voices stripped of emotion to the warm, realistic voices of the ReadSpeaker range.
  3. Artificial intelligence (AI). The first VUIs were not easy to use. They stumbled over the subtle variations of accent and dialect from one person to the next. The pre-written text-to-speech answers sounded crackly and inhuman, and were often hard to follow. Artificial intelligence solves these problems. Powerful neural networks learn from real human speech, and so improve recognition over time. This kind of AI-based automatic speech recognition is called natural language understanding (NLU), and it is what lets Alexa recognise that “play my favourite playlist” and “let’s listen to some music” mean the same thing. On the text-to-speech side, deep learning produces voice models that reproduce the subtle variations of a user’s language to create speech with a far more human tone, even reflecting the user’s dialect where that applies. This is called natural language generation (NLG).

But while artificial intelligence is transforming both automatic speech recognition and text-to-speech engines, these remain two very different technologies. When user interface providers design an interface for voice, they need at least two partners: a company that builds automatic speech recognition systems, and another that specialises in text to speech.

Looking for a text-to-speech provider for a custom VUI? Read our customer stories and see what working with ReadSpeaker looks like.

A short history of voice interface technology

Voice user interface technology only entered the home when Apple launched its voice assistant Siri on the iPhone 4S in 2011 (opens in a new tab). But the roots of the VUI run much deeper, with automatic speech recognition and text to speech each following a separate path.

According to the International Computer Science Institute, automatic speech recognition was born in 1952, when Bell Labs unveiled a device called Audrey (opens in a new tab). Audrey could understand the spoken digits from zero to nine with 99 percent accuracy, which limited its use to dialling telephone numbers by voice command. The device also cost a fortune and filled a rack close to two metres tall. Audrey was not a consumer product, but it served as a proof of concept.

A decade later, at the World’s Fair, IBM lifted the lid on the “Shoebox” (opens in a new tab), a machine able to understand 16 words of English. In 1971, the U.S. Defense Advanced Research Projects Agency (DARPA) began work on Harpy, the first speech recognition system able to understand a vocabulary of more than 1,000 words. All through the 1970s and 1980s, however, automatic speech recognition stayed strictly outside the consumer space.

Everything changed in 1990, when a company called Dragon Systems released a limited automatic speech recognition program aimed at the general public. Seven years later, Dragon shipped the first recognition software able to understand complete sentences: Dragon NaturallySpeaking. Doctors still use an updated version (opens in a new tab) of that product today as a hands-free dictation system.

In the 2010s, progress in natural language understanding gave rise to the first generation of voice assistants, and IBM’s Watson system took part in the famous American quiz show Jeopardy (opens in a new tab). Today, NLU lets speech recognition systems pick up the subtle differences of spoken language, creating a more natural interaction between devices and the people using them.

Synthetic voice technology goes back even further than automatic speech recognition. In an interview on the Alpha Voice podcast (opens in a new tab), Niclas Bergström of ReadSpeaker looks back on a history of speech synthesis that starts in 1779, with a machine that produced a synthetic voice (opens in a new tab) built from a system of reeds and resonators.

From the late 1920s, Bell Labs began experimenting with electronic speech synthesisers, which led a decade later to the invention of engineer Homer Dudley: the Voder (opens in a new tab), the first fully working speech synthesis machine.

The first true text-to-speech system appeared in Japan in 1968, Niclas Bergström explains. The 1970s saw an explosion of speech synthesis technology and of major commercial systems, such as Texas Instruments’ Speak and Spell or Ray Kurzweil’s range of reading machines for people with visual impairments.

In the 1990s, text to speech drove the growth of interactive voice response (IVR), the automated computerised telephone systems still in use today.

1999 is the year ReadSpeaker was founded, and it quickly became the first company to bring text to speech to cloud computing systems. That innovation let developers building for voice fold text to speech easily into standalone software and, later, into mobile applications. Today ReadSpeaker continues to advance text-to-speech technology as a pioneer in the use of powerful neural networks, a technology that makes the VUI more dynamic and easier to use on an ongoing basis. Here are some of the ways companies use the VUI today to add value.

Voice user interface examples in today’s businesses

While the most familiar VUIs are those of mobile phones and smart speakers, businesses use voice interface technology to make collaboration easier, to multiply the occasions on which they can express their brand, to improve the user experience for their customers, and much more. Here are a few examples of voice user interfaces in practice:

  • Manufacturers use the VUI to control production lines and to adopt the local industrial Internet of Things while keeping their hands on their tools.
  • Teachers use VUI devices in the classroom (opens in a new tab) that answer student questions, provide information instantly and even act as an aid to language teaching (opens in a new tab).
  • In healthcare, professionals value hands-free dictation systems that simplify the creation of medical records.
  • Adding a VUI to server-based IT systems lets employees book meeting rooms, move appointments and record notes in a safe, closed system, without touching a single terminal.
  • Some companies deliver voice assistant services designed for business. Synqq, for instance, is a smart note-taking application that uses NLU to record meetings and highlight the important moments, such as discussions around the measures to be put in place.
  • Conversational AI platforms such as MindMeld give companies a starting point when they want to bring a VUI into their own customer service systems.

As these examples suggest, businesses use VUIs in two main ways: in the office, to simplify internal processes, and in their products, to create a better user experience. In either case, a dedicated voice can strengthen recognition, loyalty and engagement between the company and the person listening. See how ReadSpeaker powers VUIs and other text-to-speech applications.

Do you need neural text to speech for a voice user interface?

ReadSpeaker custom voices are almost impossible to tell apart from a human voice. They are built to measure to match your brand. They are available in more than 30 languages, with others in preparation. Voice interfaces offer none of the traditional visual identifiers, such as logos and graphic guidelines. That means one brand is set apart from another by the voice itself, and ReadSpeaker is well placed to help with that.

Whether you choose a dedicated custom voice or a standard voice, ReadSpeaker text-to-speech services are ideal for anyone designing voice user interfaces. Our solutions run online, on your server, or even offline, embedded in a device. Every ReadSpeaker text-to-speech solution has been built by teams of engineers, linguists and powerful neural networks, and we have been doing this since 1999. Contact us today to discuss how we can help you design and put in place voice interface technology for your mission-critical systems.

Related articles

Find your ReadSpeaker solution

Opening the conversation…

Search

    Searching…

    No results for

    Try another wording, or start from one of these.

    Search is not available right now.

    You can still reach the pages below.