Hacker News new | ask | show | jobs
by sylware 7 days ago
I wonder how long it would take a "normal" coding prompt to go thru a "1T-2T" models on a medium performance consumer desktop.

Hours? Days?

1 comments

https://www.youtube.com/watch?v=9TyJ9s26ylc

acoording to this video, around 30t/s, if in 2027 some 256gb ai box arrive, the performance will be even better

I'll watch this video once I am on my yt-dlp rig. Hopefully, those are 1T-2T params models.

If "AI" becomes ubiquitous, it may end up pushed out of the 'standard economy' and made sorta an 'utility' (to speak using US words). Because, you will have those able to pay for performant AI... and the other ones. To mitigate that, running locally at decent (but still slow) speed, frontier models, may become critical.

I say that: I have never fooled around with AI since they are gated with 'whatwg cartel' web engines.

I did watch your vid: it is for quantized deepseek models. I would look for full weight inference for maximum quality.

For instance, do we have similar benchmarks on the latest frontier open weight KIMI model at several tera-weights (I would be interested at coding prompts)?