Hacker News new | ask | show | jobs
by javaeeeee 302 days ago
This calculator estimates the GPU memory needed to run LLM inference. Select the model size and precision (FP32 - FP4) to get a quick memory range estimate.