Hacker News new | ask | show | jobs
by jboss10 22 days ago
For people who saw this and might want a recomendation, I like running a tiny qwen model with llama cpp. Qwen2.5 coder 0.5B or 1.5B (not the instruct version)

On a modern-ish GPU these should run really fast with little latency. They cost nothing and don't send your data to anyone.