Hacker News new | ask | show | jobs
by TiredOfLife 12 days ago
Kinda already exists.

demo https://chatjimmy.ai/

https://news.ycombinator.com/item?id=47103661

1 comments

Hallucinates on the first question I ask, as 90% of these models that try to take shortcuts.
You’re expecting the wrong thing. The demo demonstrates the insane inference rate of dedicated hardware. Iirc it’s llama 3 or something. Not a very good model by today’s standards. But it runs at 16k tokens per second, an order of magnitude above the competition.

Imagine what’s possible if you had GLM-5.2 turned into a hardware chip like this.

It's a custom 3 bit quant of Llama 3.1 8B and other shortcuts. The quant is not good. Their newer arch switches to standard 4 bit quants, should be far better!