|
|
|
|
|
by yjftsjthsd-h
3 days ago
|
|
Couple highlights: > Complete local text-to-waveform speech synthesis under 10M parameters. In case, like me, you hoped "complete" voice might mean both stt and tts. Not to speak poorly of it, just clarifying. > English only, with one fixed male voice. This is not zero-shot voice cloning. (And then a bunch of statements on limitations that I read as 'quality can be spotty but if you play with it it should be fine') But like. In <10M params I'm not judging:) |
|