Hacker News new | ask | show | jobs
by khanhnguyen8386 1 day ago
Can't wait to run this at 0.02 tokens/sec on my CPU so I can get a response just in time for next month.
1 comments

Unfortunately I don't think modern consumer CPUs are physically capable of addressing enough RAM to even load the model into memory. We'd have to wait for some random person to make an extremely quantized version before we could reach those blazing speeds