Hacker News new | ask | show | jobs
by matt_daemon 46 days ago
> Hardware (minimum): 1× H100 @ FP8

Cool to see this but seems like it would be pretty expensive to run

2 comments

This is a 30B parameter model with 3B active. It should run performantly on a Mac with > 48GB RAM at 8bit precision.
Well that is like 3 USD/hour if you run it on a rented gpu
4-bit quantized 30B-A3B MoE models can run at something like 21 tokens/sec on a several year old AMD CPU.