Hacker News new | ask | show | jobs
by 100ms 20 days ago
Is it responsive to personality settings? I actively don't want fake AI girlfriend, but I do get a ton of value out of voice mode. Looking forward to trying this but hoping it's not a creepy overdone mess (like Sesame). Expectations are they'll keep doubling down on fake AI girlfriend approach because the thing I want probably wouldn't drive engagement anywhere nearly as well
2 comments

I too found this with their previous attempts.

I have my Chat personality settings stripped right down to no-fluff. I'd want voice to be more akin to the Star Trek computer, and less akin to as you said an AI friend, but previously it was tuned too personable/friend-like.

Star Trek computer voice model is something I have yet to encounter, and I've looked repeatedly :) It's not about a specific voice, it's the fact they managed to capture "I am a utility" perfectly in the voice. Our modern friends do not want to be thought of as a utility, but to engender trust and agency all of their own and that's a huge problem for me.
I thought to try voice cloning with dots.tts ( https://huggingface.co/spaces/rednote-hilab/dots.tts ), the result is pretty good, but likely wouldn't be fast enough to use on a quasi-realtime basis:

Input clip: https://vocaroo.com/19QtEPtwTjOS

Prompt text: There are 14 varieties of tomato soup available from this replicator. With rice, with vegetables, Bolian style, with pasta specify hot or chilled.

Output: https://vocaroo.com/1f3XuQQoSzwB

I think the request here is not about sounding like Majel Barrett but in keeping the output extremely terse and unobtrusive.

There's been a few studys showing that novices love LLM output that's long, but experts hate it. As an example, I've been tasked with using some agentic PM tool to write specs, and it keeps generating these huge page long outputs with "HBR voice" bolded summaries of paragraph long bulletpoints. I.e.:

> Right-size hard, and watch the one open-ended edge. Endorse the DRI's simplifications wholesale: drop the runbook-per-alert mandate (keep 1–2 diagnostic-only runbooks for the high-priority set), and ride durability on the existing weekly incident + monthly operational reviews — no new governance. The single scope-creep risk is the coverage strand (gaps are defined by absence); bound it to gaps evidenced by real, already-missed customer-facing outages, not a proactive gap hunt. Curing ownership gaps (e.g. foo-bar, no clear owner) is finite in-scope work.

There's dozens of these every iteration. I can't imagine trying to deal with that via voice, I would just zone out after the second sentence.

When the voice models start rambling, I think of C-3PO. When Uncle Owen told 3PO to shut up and 3PO said, "Shutting up, sir."

Or even some scenes where Data did something similar.

It's funny to now experience it.

AHAHAHA. HBR Voice. Thanks man, you captured it perfectly. Finally I have a name for that.
I'm just going to say, Wow. That's a pretty incredible result. I had no idea you could clone a voice with such a short snippet of audio these days.

I've had this dream of talking to the Enterprise-D computer since I was 8 years old. Midlife-crisis me still has that dream, but hooked up to Home Assistant so it can actually do useful things too. A couple months back, I went looking around for "clean" samples of Majel's voice as the computer but didn't have a lot of luck. Even though there are three television series and several movies, pretty much all of them have some amount of background noise, bleeps, bloops, or warp core thrum. (As this one does.) There may be modern ways to clean those up without affecting her voice much, but I haven't dug into that yet.

There are a few audiobooks narrated by Majel Barrett but obviously her role as the computer was proper voice acting and so the books would not be good source material.

There were also a few games/CD-ROM (Omnipedia) with some samples, but they did not bother to post-process them for that lofi Enterprise-D computer feel. Can _probably_ be replicated fairly faithfully after the fact, but I only know a _little_ about audio post-processing.

According to her son (Rod Roddenberry), Majel sat down in a studio and recorded audio samples specifically for the purpose of having her voice cloned someday for future Star Trek stories. However, those haven't been released publicly. (And likely never will, but I can dream, can't I?)

Edit: I played with your samples and the time it takes to generate the output audio is pretty brutal. Too slow for interactive use. Maybe that's a limitation of the HF-hosted app, though.

If you read their GitHub README and the code, it's possible to separate and cache the voice cloning step, and they have some streaming variant with low first chunk latency. I only played with it via Hugging Face. It is astonishingly good with certain rare accents, and it responds well to longer input clips
That is pretty incredible with such a small amount of input. Makes me want to check out dots.tts myself!
I just tell them to talk like Jarvis. I tried the "cold logic" prompt but it turned them into a bunch of assholes.
I noticed that too, I guess the content in the training data suggests that when some text self-identifies as "cold, hard logic" it additionally assumes a mean tone. Additionally it makes responses more likely to disagree even when given inputs that are more or less valid. A while ago I tested out giving some idea to a model with the "hard logic" prompt, then taking its own output, reverting the conversation and giving it its own statement again, making it disagree with the very same statement it just made itself. This was mostly just humorous, and doesn't indicate too much, after all both statements might've been wrong in some way, but the statement itself was moreso an opinion with no fully right answer, showing this tendency to disagree.

I guess when we tell models to use "cold logic" they don't interpret that purely definitionally, but moreso with what this statement usually implies, which is often a disagreement between two people, one trying to leverage supposed logic to disagree and denigrate the other party, oftentimes actually arguing out of emotion and not the logic they claim to use. This probably occurs enough to give model responses a mean tint. That's my theory on why this could occur at least.

For many years, I've wanted ED-209 (robocop) voice from something like espeak or similar. Still can't find anything good.

Not for chat, just as a way to make notification messages that sound like ED-209.

One popular speech synth from back in the day, I believe it was WillowTalk, had a voice called Colossus, which sounded like the voice module of the computer from Colossus: The Forbin Project. This voice was used for that of CATS in the famous "All your base are belong to us" Flash video.

Another WillowTalk voice was a clone of DECtalk's Perfect Paul good enough to be used as the voice in the MC Hawking rap recordings.

I am having a hell of a time googling for WillowTalk, do you know any links/places where I can learn more? Maybe even download something?
It came from a company called Willow Pond Software. That seems to narrow searches down.

WillowTalk is apparently still of interest to Half-Life modders because another WillowTalk voice, possibly a clone of DECtalk's Huge Harry, was used as the Black Mesa VOX facility-wide announcement system in the original Half-Life.

I think that's mostly just a frequency shift :) You could probably recreate it with another model and some effects on top. Also, why the hell not for voice mode haha.
You can using chatterbox and a couple voice to voice models. Sorry I don’t recall the whole process but you can YouTube Star Trek computer voice AI.
I explicitly prompt my small local models to channel that kind of energy. It's one of few ways to get them to just spit it out without yapping on about temperature- and humidity-appropriate activities when I just asked if it'll rain this week.

You do have to do it carefully and not literally say "Star Trek" or else it'll start yapping about main shields being at full strength.

Well you remember how Captain Kirk had his computer speak to him in a special female voice, and he was "married" to the ship ;)
It’s 2026 what the customer wants is probably an indication of what not to do if you’re a hyperscaler.

Whoever ends up actually winning after the crash will be the ones who figure out the needs of users, not the needs of the financial machinations of the companies trying to grab land.