Amazon , has unveiled Nova Sonic, a new generative AI voice modelthat can natively process voices and produce natural-sounding speech. The company believes it can compete with similar models from OpenAI and Googleand conversation quality.

Nova Sonic is available through Bedrock, Amazon’s developer platform for building enterprise AI applications (via a new bi-directional streaming API). In a press release, Amazon called Nova Sonic “the most cost-effective” AI voice model on the market (about 80% cheaper than OpenAI’s GPT-4o).
Nova Sonic components already power Alexa+, Amazon's upgraded digital voice assistant.
See also: Amazon introduced the new AI agent “Nova Act”
Compared to competing AI voice models, Nova Sonic excels at routing user requests to different APIs, according to Amazon. This helps Nova Sonic “know” when to retrieve real-time information from the web, analyze a proprietary data source , or take action in an external application — and use the appropriate tool to do so.
Amazon says that during a conversation, the AI voice model Nova Sonic waits to speak “at the right time,” taking into account pauses and interruptions from the interlocutor. It also creates a text transcript of the user’s speech, which developers can use for various applications.
Amazon claims that the Nova Sonic is less prone to speech recognition errors than comparable models, meaning it's relatively good at understanding a user's intent even if they mumble, mispronounce, or are in a noisy environment. On Multilingual LibriSpeech, a test that evaluates speech recognition across languages and dialects, the Nova Sonic achieved a word error rate (WER) of just 4.2% (averaged across English, French, Italian, German, and Spanish). That means that about four out of every 100 words the model missed differed from a human transcription.
See also: Ocelot: Amazon Web Services' (AWS) quantum computing chip

In Augmented Multi Party Interaction, which measures loud interactions with multiple participants, Amazon’s AI voice model Nova Sonic was 46.7% more accurate in terms of WER than OpenAI’s GPT-4o model. Nova Sonic also has industry-leading speed.
Executive Rohit Prasad says Nova Sonic is part of Amazon’s broader strategy to develop AGI (artificial general intelligence), which the company defines as “AI systems that can do anything a human can do on a computer.” Prasad says Amazon plans to release more AI models that can understand different media, including images, video and voice.
See also: Amazon Project Kuiper: Starlink competitor launches satellites
Amazon's new AI voice model Nova Sonic is another step towards more natural human-machine communication. It promises realistic pronunciation, intonation, emotions and tries to sound like a real person. This means that it could be used to significantly improve various services we use every day (Audiobooks, Virtual characters in games, Customer Service, etc.).
Source: techcrunch.com
