Hacker News new | ask | show | jobs
by qznc 35 days ago
https://mikeveerman.github.io/tokenspeed/?rate=750&mode=thin...

This is what 750tps looks like, I guess.

3 comments

You get used to it. I don't even see the code. All I see is blonde.. brunette.. redhead.
That’s an awful visualization. I can skim code quite quickly, but not when it shows up one character at a time in a small window, modem style.

At least that site should draw out a full page then start replacing that page with the next, starting from the top and working downwards, repeating each time it hits the bottom.

This is how tools like claude code and chat prompts output their tokens, so I'd say it's actually a pretty good visualisation.
Not for prefill. I suppose if you just want to imagine what generation speed looks like in the current generation of TUIs, it’s an okay visualization.
That's exactly what it looks like in the tools I use most (opencode and codex), so for that purpose it's a pretty good visualization.
> I can skim code quite quickly

are you by any chance hyperlexic? interested to hear more about this, like how fast is considered fast

Just to think what this will look like in a couple of years.
Hopefully like this (but smarter): https://chatjimmy.ai/
This is genuinely confusing to my senses. The future is going to be so strange/neat/me unemployed.
> strange/neat/me unemployed

I'm not sure if that's what you were going for, but I read it as if it were written by The Board in the game Control, and found myself with the appropriate level of existential dread.

and I haven't played that game, so I read it in Ralph Wiggum's voice.. which also feels appropriate.

I'm in danger.

We love/help/replace you
The future is totally illegible to me. I love these AI models, but I feel like I'm going to be jobless within 10 years.

Anomie is at an all time high right now.

10 years? An optimist, I see.
Yeah. It keeps catching me off guard that it answered me already.
Why is the insane speed of 13KTPS of this site is not more on the the top of the AI conversations?
Because there's been nothing to discuss since their announcement. Their API access immediately closed due to overwhelming demand and they didn't fab newer models than Llama3 yet.

Probably they will make bank selling to HFT for a while.

Because I just tested it and it took 3-4 clarifications before it actually gave a correct response vs gemini/google search. It's not great, but good.

I'd rather wait 3x as long.

It's pretty well known by now.
I asked it for a block of C++ code and it hit 14,189 tok/s. I assume it cached someone else's session?
Wow.. what?! How is this so fast?! Where can I read more?
Funnily enough, pasting your comment straight into Jimmy leads to a... Funnily suboptimal answer that does not answer the question.

As someone else already contributed, this is driven by a Canadian startup taalas that basically makes chips that are llms, so everything is very fast but also, baked into the chip. Once this kind of stuff is a commodity in like 10 years, our world will be very, very different.

Taalas HC1 AI uses Llama 3.1 8B, but takes up a massive 53B transistors and 815mm2 on TSMC N6 (nearly at the reticle limit of 858mm2). N2 is a little less than 3x as dense (110MTr/mm2 vs 313MTr/mm2).

This chip would still be 272mm2 on N2 which is an eye-watering $30k/wafer and bigger than a 9950x or Nvidia 5070.

This just isn't feasible. Some of the latest-gen LLMs seem to have 5-10T parameters or about 1000x more. I don't know that taping out just one chip makes economic sense let alone the 300-1000 chips required for a cutting-edge model. Things like continuing education so your model knows about the latest NPM packages or world news is super important, but seems like it would require new chips.

There are a TON of uses for an 8B parameter models on the edge, but this is WAY too big to put on the edge of anything. Something like a 10mm2 100m parameter voice model might be feasible on the edge, but only for expensive devices, but most of those are TSMC 28nm (up to 29MTr/mm2) or GF FDX22 (up to 40MTR/mm2) which would increase the AI chip to the point where it would absolutely dominate the BOM.

> Things like continuing education so your model knows about the latest NPM packages or world news is super important, but seems like it would require new chips.

They probably have a few ideas around that. Me, personally, I'd have one main expensive chip (replaced every 10 years, or whatever), with a secondary cheap chip in front of it that gets replaced every year or so.

The secondary chip could act the way RAG does, or perhaps both chips together can act as LoRA.

Either way, 99.999% of the knowledge is static, you just need to fine-tune the weights with that remaining 0.001% knowledge, which can be done using RAG or LoRA on a much smaller (thus cheaper) disposable chip.

Yeah, they're clearly just starting out and just shipped their very first proof of concept. But to me, their plans seem generally reasonable https://taalas.com/the-path-to-ubiquitous-ai/, and like I wrote, if this kind of thing succeeds and could become some kind of cheaply producible commodity component, I think there's huge value in that. Alas, maybe not as a frontier model replacement, but say 10 years from now you can drop a cheap raspberry pi like device in your Lan and have a fast local engine for things like text sentiment analysis, text summarisation, voice recognition, basic vision and things like that, that would be pretty exciting to me (but maybe as you outlined, impossible in practice)
That’s why this stuff should be a government mega project ultimately.

It is not market viable but it is sure as heck revolutionary. Like an atomic bomb but including more… peaceful uses.

That’s exactly where government should take rein like with ISS etc. However the models are too rapidly advancing for now for it to make sense

the flash models have fallen in size at least between deep seek models. Is there a limit to the shrinking capacity of the models?
This caused me to have some sense what blistering fast AI actually is. What it means for the future is a question that remains.
Damn that is crazy.
This is the reaction every time it's posted, and deservedly so.
Not opening here... HN killed?
What

How?

Which model is behind it?

It’s pure silicon. Llama3.
What does that even mean? There must be a software stack somewhere feeding the inputs to the model. Do you mean the weights are baked in sillicon?
hugged to death?
I started with a 2400baud modem, I've seen how this goes
Sometimes I visualize a setup like this [0], based on 2D art by Simon Stålenhag. Someone has their home robot sitting on a desk connected to their old PC with thick cabling, dumping endless lines of each subsystem's <think> logs to diagnosis why it did something weird earlier in the day. Systems pushing 750+ tokens per second per subsystem might even be considered on the slow side for realtime tasks by then.

[0] https://www.therookies.co/entries/39513

Probably not. Everyone will still need a lot of reasoning tokens and tool calls. Running the tests for every round is tiring but must be done.
Imagine a Beowulf cluster of these…
That's a name I haven't heard in a while.
First post?
The user has many comments and updoots if you look at their profile.
We’re being silly and spamming Slashdot spam comments :p (“imagine a Beowulf cluster” and “first post?”)
I always think of Furbies because of that geocities (memories!) site.
Probably will not be looking at text like this in a few years.
probably something like this https://sb0xw.csb.app/