Y
Hacker News
new
|
ask
|
show
|
jobs
by
Tepix
36 days ago
If you want decent performance (more than say 20 tokens/s) for your dev team, you absolutely do need all of the model in VRAM.