Hacker News new | ask | show | jobs
by jmillikin 23 days ago
That's the idea behind Taalas (https://taalas.com), except as silicon rather than ROM. They run a demo at https://chatjimmy.ai/ which serves an old open weights model (Llama 3.1 8B) at something like 15,000 tokens per second.