|
|
|
|
|
by embedding-shape
9 days ago
|
|
> For distributed inference of smaller LLMs and diffusion models that fits in one consumer GPU rather than splits on multiple machines But isn't what everyone want almost the complete opposite of this? For smaller LLMs you can already run them with consumer hardware, sometimes consumer hardware 2x, but it's the rest of the stuff where distributed inference could actually "democratize" access, as almost no one has 500GB VRAM available at home, or whatever the latest SOTA local models take nowadays. |
|