Hacker News new | ask | show | jobs
by josh-wrale 8 days ago
I agree with this. I have a dual rtx4090 machine, a 128gb m5 max mbp, and a dgx spark. RedHatAI/gemma-4-26B-A4B-it-FP8-dynamic on the dual rtx4090 machine under vLLM absolutely slays at token speed. I'm frustrated with how hard it is to realize goodness on the dgx spark.