Hacker News new | ask | show | jobs
by navbaker 40 days ago
Yeah, I was a bit baffled by the author complaining about cache prefixes getting destroyed when more than one user hit the model, but then continuing to use llama.cpp instead of switching to vLLM.