|
|
|
|
|
by aae42
22 hours ago
|
|
In my admittedly little experience running a tiny LLM service on a single 5090 for friends, I would guess a fair bit. It depends on how many tok/sec you're targeting and how much memory you're allowing be used for context (and their context). Definitely more than 1, MAYBE an order of magnitude. |
|