Hacker News new | ask | show | jobs
Show HN: I wrote a 1-bit WebGPU runtime to run a 1.7B LLM in the browser (aidekin.com)
5 points by stfurkan 22 days ago
1 comments

Wow this is a cool prototype and looks amazing. I am surprised some level of baseline intelligence survives this kind of aggro quantization. Do you think 300mb initial download is ok for something like quick website where I want to ask quick support question? Are you planning to have hosted fallback to answer q while download is happening?
Thank you :) Currently I am not planning to have hosted fallback Q&A but it's a nice idea. It should only download the ~300mb initially once and then use the cached model.