|
|
|
|
|
by throwawayffffas
8 days ago
|
|
> whoever burns their models to ASICs fastest. There is already custom hardware see cerebras. GPUs have a lot of slack there is at least one lab that had a (small 8b) model generate almost 3000 tokens per second on a MI300X for a talk, instead of the typical software stack that did maybe 100ish tokens per second. High bandwidth flash storage is in the works, i.e hard drives with TBs of storage and over 1 TB per second of read speeds. Meaning that in a couple of years you may be able to buy a card with 40-90GBs of HBM and 4TB of HBF
and run a 3T model locally at a reasonable speed for 10-20k as opposed to a cool mil. |
|
There is no "may" here. You will see this.
It's always difficult to see it from the present, but we're not at some end stage in hardware development; we're still on the same curve our predecessors also couldn't see: they couldn't imagine that there would be high performance computers carried in our pockets, with staggering amounts of storage and compute, putting to shame the machines they filled rooms with.