Hacker News new | ask | show | jobs
by cdr 21 days ago
It is indeed a common weakness of TTS models.

Unfortunately it makes it unsuited for my use case, which is almost entirely single words, as I don't particularly want to deal with stitching/segmenting input/output.