Hacker News new | ask | show | jobs
by mycall 10 days ago
DGX does parallel inference with llamacpp/vLLM much better than AMD GPU at 128GB VRAM.