Meta introduces Voicebox, new generative text-to-speech AI

AI summary of the linked article

Meta has announced Voicebox, a generative AI speech model that can perform tasks it was not specifically trained for, such as editing, sampling and stylizing speech, through in-context learning. Voicebox can produce high-quality audio clips, remove noises like car horns or a dog barking from pre-recorded audio while preserving the content and style, and generate speech in six languages: English, French, German, Spanish, Polish and Portuguese. According to the announcement, it can match a voice from an audio sample as short as two seconds for text-to-speech generation, and it can produce a reading of text in a different language from the sample speech.

Recommended by 2 curators
Characters remaining: 10,000

comment guidelines