Hacker News new | ask | show | jobs
by Tepix 36 days ago
If you want decent performance (more than say 20 tokens/s) for your dev team, you absolutely do need all of the model in VRAM.