> Complete local text-to-waveform speech synthesis under 10M parameters.
In case, like me, you hoped "complete" voice might mean both stt and tts. Not to speak poorly of it, just clarifying.
> English only, with one fixed male voice. This is not zero-shot voice cloning.
(And then a bunch of statements on limitations that I read as 'quality can be spotty but if you play with it it should be fine') But like. In <10M params I'm not judging:)
this is extremely encouraging for individuals/small companies being able to train pareto-frontier TTS models (specifically compute required to run vs quality of model output)
I think we need more neurons in the human brain than parameters in this model for speech. I wonder what it says about the human brain vs LLM efficiency.
here my implementation with speech dispatcher and server: https://github.com/skorotkiewicz/inflect-speechd
thanks for shearing!