| HN Mirror

Y	Hacker News new \| ask \| show \| jobs


	by lhl 996 days ago
	For neutral sounding very fast/efficient voices, I find Coqui TTS VITS models to be very good. For slower, more expressive voice or voice cloning I think the Coqui TTS XTTS is good (or you can look at the mrq/tortoise-tts). I'm still awaiting a StyleTTS2 implementation. The audio samples sound top notch: https://styletts2.github.io/

1 comments

modeless 995 days ago

You're in luck, the code dropped 6 hours ago :) https://github.com/yl4579/StyleTTS2

Looks promising, I'm going to check it out too! MIT license, even! If it's fast enough for real time, it could be the new best option. The paper claims faster inference than VITS...

link

lhl 995 days ago

Ha awesome! I just checked the repo literally before I posted and it was still empty, thanks for the heads up, will give it a spin now.

link

lhl 995 days ago

Just a followup for those interested, inference implementation notes and comparison clip between StyleTTS2, TTS VITS, and XTTS: https://fediverse.randomfoo.net/notice/AaOgprU715gcT5GrZ2

link

modeless 995 days ago

Wow you got it working so fast! I'm still stuck in package manager hell trying to debug a million little issues.

link

lhl 995 days ago

In my post I link to my issue where I outline what I needed to do from a clean mamba env that might help.

Pytorch nightly (I use for cuda-12) doesn't work w Python 3.12, but if you stick w 3.11 or 3.10 you should be ok. Rest was just w/o version numbers if you're on a clean venv should be fine, however there's a bug in the Utils lib that requires a 1-line fix if you're trying to inference (also linked). nltk was the only dependency not listed so not bad compared to most code drops!

link

modeless 995 days ago

I spent a couple of hours debugging why jupyter's debugger wasn't working right, so not exactly related to the code. I did also find and fix that utils bug you mentioned. But my current issue is that phonemizer won't find espeak even though I set the environment variables that are supposed to work. I'll figure it out eventually...

Thanks for writing up your experience! Good to know it works! And it's fast!

link