Hacker News new | ask | show | jobs
by Xalutiono 8 days ago
Whats the prefered LLM runtime to use? vLLM?

Tips & Tricks on parameters/settings?

What happens at peak? Do people have to wait now? Increase of latency?

1 comments

We use vllm as it generally has the best ecosystem support. Parameters are largely dependent on what type of requests you are serving (concurrency, input/output ratios, cached hit patterns). We've never had a limitation at the tokenizer step. Limitations at peak tend to manifest more on slower time in vllm doing prefill or decode though we actively try and minimize this.