A team of NVIDIAworking on generative AI has created a special audio tool (Fugatto), which allows users to control audio output and create new voices, sounds, and music.

There are already some artificial intelligence models that can compose a song or modify a voice, but none can offer something completely new, according to the company.
Fugatto (short for Foundational Generative Audio Transformer Opus 1), creates or transforms any mix of music, voices, and sounds described by prompts, using any combination of text and audio files.
See also: Microsoft: Launches Recall and adds new AI capabilities
"For example, it can create a musical excerpt based on a text prompt, remove or add instruments from an existing song, change the accent or emotion in a voice — even let people produce sounds that have never been heard before," the company says in its blog post.
“Audio perception” of sound
“We wanted to create a model that understands and produces sound like humans,” said Rafael Valle, director of applied audio research at NVIDIA, as well as an orchestra conductor and composer.
Supporting numerous audio production and transformation tasks, NVIDIA's AI tool, Fugatto, appears to be the first fundamental generative AI model that exhibits capabilities that arise from the interaction of skills learned through training and the ability to combine free-form instructions.
Fugatto usage examples
Music producers could use NVIDIA's Fugatto to quickly create original tracks or refine a song idea, experimenting with different styles, voices, and instruments. They could also add effects and improve the overall sound quality of an existing track.
See also: Microsoft: New AI capabilities to improve search in Windows
An advertising agency could use Fugatto to make an existing campaign better suited to multiple regions or situations, by applying different accents and emotions to the voices.
Language learning tools could also leverage the AI tool to use whatever voice the speaker chooses.
Video game developers could use the model to modify pre-recorded elements in their game to match the action as users play the game.
NVIDIA explains that to produce the result, the model uses a technique called ComposableART and combines commands, studying them separately. For example, a combination could be: a text spoken with a sad emotion in a French accent.
The model's ability to distinguish commands gives users fine-grained control over text commands, in this case the severity of the pronunciation or the degree of sadness.
“I wanted to allow users to combine features in a subjective or artistic way, choosing how much emphasis they give to each,” said Rohan Badlani, an AI researcher who designed these aspects of the model.

"In my tests, the results were often surprising and made me feel a bit like an artist, even though I'm a computer scientist," said Badlani.
The model also produces sounds that change over time, a feature it calls temporal interpolation. It can, for example, create the sounds of a thunderstorm in an area with louder thunder that slowly fades away. It also gives users granular control over how the soundscape evolves.
See also: Adobe: New Generative AI capabilities in Illustrator and Photoshop
Furthermore, unlike most existing AI models, which can only recreate the training data they have been exposed to, “Fugatto allows users to create soundscapes never seen before, such as a storm receding at dawn with the sound of birds singing.”
Fugatto: A special AI tool
This innovative technology offers exciting possibilities for artists, content creators, and professionals in other fields, allowing them to push the boundaries of creativity in unprecedented ways. Leveraging advanced machine learning techniques, users can seamlessly combine different audio elements to produce unique compositions. This could revolutionize industries ranging from music production to gaming and cinema, where sound design plays a crucial role in enhancing the immersive experience.
Source: blogs.nvidia.com
