Y
Hacker News
new
|
ask
|
show
|
jobs
by
verdverm
46 days ago
it's a prompt cache invalidation bug that causes all input to be reprocessed instead of getting preloaded
There are other reasons to prefer vllm to llama-cpp as well