Hacker News new | ask | show | jobs
by yieldcrv 22 days ago
I think local inference will be fast enough

There is so much happening in that scene, where tokens/sec double or 10x

So I could see the same hardware doing 20 tokens/sec on a large model suddenly doing 200 tokens/sec in the future, a better device in the future doing 500 tokens/sec, while having vision models baked in, audio models etc

Users wont consciously switch to local, they will just have it and use it