Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

I had planned to play around with TorToiSe[1] next weekend and already watched some videos. There it looks like all you have to do, is to offer you own voice samples to the system and no separate training seems to be required. TorToiSe is slow to synthesize, so it doesn't beat the 3 seconds but can anyone confirm that these models really don't need an extra training phase to clone a voice?

[1] https://github.com/neonbjb/tortoise-tts



That’s correct. However, the 3 seconds refers to the minimum amount of audio required for the reference audio, not how long it takes the mode to synthesize.


Correct, that's what zero-shot means, no training steps.


I've been voice-cloning using tortoise-tts and I'm very happy with the results, but it is indeed very slow. It's also free and open-source though.




Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: