|
|
|
|
|
by jboss10
22 days ago
|
|
For people who saw this and might want a recomendation, I like running a tiny qwen model with llama cpp. Qwen2.5 coder 0.5B or 1.5B (not the instruct version) On a modern-ish GPU these should run really fast with little latency. They cost nothing and don't send your data to anyone. |
|