I had planned to play around with TorToiSe[1] next weekend and already watched some videos. There it looks like all you have to do, is to offer you own voice samples to the system and no separate training seems to be required. TorToiSe is slow to synthesize, so it doesn't beat the 3 seconds but can anyone confirm that these models really don't need an extra training phase to clone a voice?
That’s correct. However, the 3 seconds refers to the minimum amount of audio required for the reference audio, not how long it takes the mode to synthesize.
[1] https://github.com/neonbjb/tortoise-tts