What you need to know
- Microsoft recently released an AI tool called VALL-E that can create convincing replications of people’s voices.
- The tool uses just a 3-second recording as a prompt to generate content.
- VALL-E can replicate the emotions of a speaker, differentiating it from several AI models.
Microsoft recently released an artificial intelligence tool known as VALL-E that can replicate people’s voices (via AITopics). The tool was trained on 60,000 hours of English speech data and uses 3-second clips of specific voices to generate content. Unlike many AI tools, VALL-E can replicate the emotions and tone of a speaker, even when creating a recording of words that the original speaker never said.
A paper out of Cornell University used VALL-E to synthesize several voices. Some examples of the work are available on GitHub.
The voice samples shared by Microsoft range in quality. While some of them sound natural, others are clearly machine-generated and sound robotic. Of course, AI tends to get better over time, so in the future generated recordings will likely be more convincing. Additionally, VALL-E only uses 3-second recordings as a prompt. If the technology was used with a larger sample set, it could undoubtedly create more realistic samples.
At the moment, VALL-E is not generally available, which may be a good thing as AI-generated replications of people’s voices could be used in dangerous ways by threat actors and others with malicious intent.
Windows Central take: Impressive but scary
While VALL-E is undoubtedly impressive, it raises several ethical concerns. As artificial intelligence becomes more powerful, the voices generated by VALL-E and similar technologies will become more convincing. That would open the door to realistic spam calls replicating the voices of real people that a potential victim knows.
Politicians and other public figures could also be impersonated. With the speed social media travels and the polarity of political discussions, it’s unlikely that many would stop to ask if a scandalous recording were genuine, as long as it sounded at least somewhat authentic.
Security concerns also come to mind. My bank uses my voice as a password when I call. There are measures in place to detect voice recordings and I’d assume the technology could sense if a VALL-E voice was used. That beings said, it still makes me uneasy. There’s a good chance that the arms race will escalate between AI-generated content and AI-detecting software.
While not a security concern, some have brought up the fact that voice actors may lose work to VALL-E and competing tech. While it’s unfortunate to see people lose work, I don’t see a way around this. If VALL-E reaches a point where it can replace voice actors for audio books or other content, companies are going to use it. That’s just the reality of technology advancing. In fact, Apple recently announced a feature that uses AI to read audio books.
Like any technology, VALL-E will be used for good, evil, and everything in between. Microsoft has an ethics statement on the use of VALL-E, but the future of its usage is still murky. Microsoft President Brad Smith has discussed regulating AI in the past (via GeekWire). We’ll have to see what measures Microsoft puts in place to regulate the use of VALL-E.
Original Article: Microsoft’s VALL-E can imitate any voice with just a three-second sample
More from: Microsoft Research
The Latest Updates from Bing News
Go deeper with Bing News on:
- Generative AI & Social Media: Creating entertainment, news and advertising content just got easier. Competition gets tougher too
From curated AI that has been tailoring our social media feeds for long we are entering the age of creative AI which is simple to use, customisable and easy to integrate to various platforms. The ...
- Can AI make you a musical star? We used Voicify and ChatGPT to find out
As vocal clones of music’s biggest names go viral, the Financial Times’ pop critic Ludovic Hunter-Tilney embarked on an unlikely quest to replicate his favourite singer’s voice.
- Microsoft News Roundup: Surface Duo 3, Apple apps on Windows … – Windows Central
No offers foundWhen you purchase through links on our site, we may earn an affiliate commission. Here’s how it works.A report on the Surface Duo 3 ...
- Can AI make me a musical star?
As vocal clones of music’s biggest names go viral, the FT’s pop critic embarks on an unlikely quest to replicate his favourite singer’s voice ...
- Jourdan To Add Mining Claims To Flagship Vallée Project
C$100,000; 2,040,816 common shares of the Company at a deemed price of $0.0735 per share, being the 10-day volume weighted average price of the shares on the TSX Venture Exchange (“ TSXV ”) as ...
Go deeper with Bing News on:
- CNET's new guidelines for AI journalism met with union pushback
The storied tech publication is promising to disclose when artificial intelligence generates a portion of a story's text, but that's about it..
- CNET is overhauling its AI policy and updating past stories
In a memo shared today, CNET outlines how it could use AI systems in its journalism in the future. The policy promises that no stories will be entirely produced by an AI tool.
- Instagram, Google, and TikTok should warn users about AI-generated content and build safeguards, EU official says
Vera Jourova, an EU official, said companies should "clearly label" apps which could spread AI disinformation and build "safeguards," per Bloomberg.
- Stack Overflow Moderators Stop Work in Protest of Lax AI-Generated Content Guidelines
Moderators of Stack Overflow, the go-to Q&A forum for programmers, have announced today they will be going on strike citing the company’s prohibition on moderating AI-generated content on the platform.