
via Microsoft
What you need to know
- Microsoft recently released an AI tool called VALL-E that can create convincing replications of people’s voices.
- The tool uses just a 3-second recording as a prompt to generate content.
- VALL-E can replicate the emotions of a speaker, differentiating it from several AI models.
Microsoft recently released an artificial intelligence tool known as VALL-E that can replicate people’s voices (via AITopics). The tool was trained on 60,000 hours of English speech data and uses 3-second clips of specific voices to generate content. Unlike many AI tools, VALL-E can replicate the emotions and tone of a speaker, even when creating a recording of words that the original speaker never said.
A paper out of Cornell University used VALL-E to synthesize several voices. Some examples of the work are available on GitHub.
The voice samples shared by Microsoft range in quality. While some of them sound natural, others are clearly machine-generated and sound robotic. Of course, AI tends to get better over time, so in the future generated recordings will likely be more convincing. Additionally, VALL-E only uses 3-second recordings as a prompt. If the technology was used with a larger sample set, it could undoubtedly create more realistic samples.
At the moment, VALL-E is not generally available, which may be a good thing as AI-generated replications of people’s voices could be used in dangerous ways by threat actors and others with malicious intent.
Windows Central take: Impressive but scary
While VALL-E is undoubtedly impressive, it raises several ethical concerns. As artificial intelligence becomes more powerful, the voices generated by VALL-E and similar technologies will become more convincing. That would open the door to realistic spam calls replicating the voices of real people that a potential victim knows.
Politicians and other public figures could also be impersonated. With the speed social media travels and the polarity of political discussions, it’s unlikely that many would stop to ask if a scandalous recording were genuine, as long as it sounded at least somewhat authentic.
Security concerns also come to mind. My bank uses my voice as a password when I call. There are measures in place to detect voice recordings and I’d assume the technology could sense if a VALL-E voice was used. That beings said, it still makes me uneasy. There’s a good chance that the arms race will escalate between AI-generated content and AI-detecting software.
While not a security concern, some have brought up the fact that voice actors may lose work to VALL-E and competing tech. While it’s unfortunate to see people lose work, I don’t see a way around this. If VALL-E reaches a point where it can replace voice actors for audio books or other content, companies are going to use it. That’s just the reality of technology advancing. In fact, Apple recently announced a feature that uses AI to read audio books.
Like any technology, VALL-E will be used for good, evil, and everything in between. Microsoft has an ethics statement on the use of VALL-E, but the future of its usage is still murky. Microsoft President Brad Smith has discussed regulating AI in the past (via GeekWire). We’ll have to see what measures Microsoft puts in place to regulate the use of VALL-E.
Original Article: Microsoft’s VALL-E can imitate any voice with just a three-second sample
More from: Microsoft Research
The Latest Updates from Bing News
Go deeper with Bing News on:
VALL-E
- How AI is putting President Macron at the heart of the French pension protests
Amidst the ongoing troubles in France over the retirement reforms and the social unrest that has decried from President Macron’s controversial bill, Internet users have turned to AI to get creative.
- As voice-cloning becomes easier, take this one step with your family members to stay safe
Scammers are using AI voice cloning technology to trick you into thinking your loved ones are in urgent distress and need your money.
- Ubisoft unveil AI dialogue-writing tool, prompting debate among developers
Game Developer reported on a GDC talk from Ubisoft La Forge researcher Ben Swanson that sheds more light on the tool, emphasising that Ghostwriter still requires developer input and that it’s mostly ...
- Cybercriminals are using AI voice cloning tools to dupe victims
The FTC described such a scenario amid the rise of AI-powered tools like ChatGPT and Microsoft's Vall-E, a tool the software company demonstrated in January that converts text to speech.
- Speak a Foreign Language in Your Own Voice? Microsoft’s VALL-E X Enables Zero-Shot Cross-Lingual Speech Synthesis
Not so long ago, text-to-speech (TTS) outputs were disappointingly deadpan and robotic. The leveraging of deep neural networks in recent years has dramatically transformed TTS, enabling conditioning ...
Go deeper with Bing News on:
AI-generated content
- Backlash over Levi’s AI-generated clothing models to ‘increase diversity’
Brand’s idea to create realistic computer-generated images of different body types is likened by critics to ‘digital blackface’ ...
- ChatGPT, Bing, Bard, Or Claude: Which AI Chatbot Generates The Best Responses?
Shown below is the second draft. Need inspiration for your content strategy? AI chatbots can help you get started in the right direction. It’s important to note that AI-generated content is not unique ...
- How is AI changing the way we write and create?
Since late last year, artificial intelligence platforms like ChatGPT have become a growing topic of conversation on college campuses, with students using the technology for everything from class ...
- How to use GPTZero to check for AI-generated text
GPTZero can tell you whether a document, report or other item was possibly written by a human or by AI. Here’s a step-by-step guide on using GPTZero for this purpose. With the popularity of ChatGPT, ...
- AI-generated videos have arrived, and they’re evolving fast
AI can now generate videos simply from your description of what you want to see. It's like Dall-E or Stable Diffusion for video.