Hacker News new | ask | show | jobs
by fergus00 19 days ago
https://blog.doubleword.ai/fast-sglang-starts yeah, this is part of the idea. If you get cold starts down to seconds or subseconds, then you can run many models multiplexed on the same GPUs
1 comments

I have read a Linkedin post in french recently with a university lab (one big machine serving hundreds of users) mentioning that they could swap models on demand. Will try to find back the link.