Hacker News new | ask | show | jobs
by aae42 22 hours ago
In my admittedly little experience running a tiny LLM service on a single 5090 for friends, I would guess a fair bit. It depends on how many tok/sec you're targeting and how much memory you're allowing be used for context (and their context). Definitely more than 1, MAYBE an order of magnitude.