Hacker News new | ask | show | jobs
by nylonstrung 10 days ago
There's going to be a golden age of GPGPU compute in the next few years once A100/H100 are fully obsolete for running frontier models efficiently and the price plummets

It will be perfect for stuff like GPU-accelerated query engines, "classical ML" and every other CPU-based workload that could conceivably be offloaded to GPU

3 comments

There's nothing stopping you doing this now.

You can get used 16GB P100s on AliExpress for ~$100 if you want obsolete GPUs. Allegedly new AMD BC 250s are only slightly more.

I've looked at this some but I already have a GTX1070 which is only supported upto CUDA 11.9.

That's precludes some interesting modern optimizations out of the box. I've spend a lot of LLM tokens backporting some things, but I'm really not sure the hassle is worth it.

New hardware is just better. I think in maybe 5 years when supply and demand are back in equilibrium we are going to have some killer technology for decent prices, and 15yo H100s won't look attractive.

> You can get used 16GB P100s on AliExpress for ~$100 if you want obsolete GPUs. Allegedly new AMD BC 250s are only slightly more.

In a recent Gamer Nexus video with Level 1 Tech, they mention V100s are also quite useful for many applications that use FP64: stuff four in a workstation, and many PhD candidates would be quite happy with the throughput they can get for certain scenarios.

Linux hackers will be finding all kinds of crazy uses for hardware that now costs $100,000 and in 5 years will be available as scrap.

This is assuming that there is no big, big disruption to the semiconductor industry (e.g. TSMC getting attacked), in which case... well, I am gonna treat each stick of RAM I currently have like it's a faberge egg.

GPGPU? General Purpose GPU? If embarrassingly parallel CPU algorithms weren't offloaded to the GPU previously, why would the A100/H100 price drop make a difference? We had cheap GPU in the past and we still left plenty of performance on the table with CPU programs because they were easier to build.

Is the idea that previously maintaining GPU programs was expensive whereas now AI makes it cheap? If so, I could buy that line of reasoning.

Maybe relatedly, I expect (hope) the hardware manufacturers will ramp up supply in the meanwhile which would also put downward pressure on GPUs. Right now though this hardware crunch is making me sad, not even because of GPUs but also because of general memory / disk.

GPGPU programming has become significantly easier now and the payoff is bigger (better hardware), due to the immense investment in this due to ML/AI.
What are the best tools for this?
We've not had cheap GPUs with this much VRAM before, though. Might be an interesting change, though I also doubt it personally.
As noted in my other comment you can get obsolete GPUs (P100s, BC250s) with lots of RAM on AliExpress now. It hasn't proven revolutionary.
is 16Gb "lots of RAM" in an LLM world? How many of these would you need to stack on a motherboard to inference a decent size model?
At least 4, probably 6+.
BC250 shares memory with system. 16 GB GPUs have long been available to consumers for a small premium. What's never been readily available before is 40+ GB of HBM.
Those don't have a lot of RAM though. Not like these 40, 80GB ones.
I am rooting for literally this... Blogged here: After "AI": Anticipating a post-LLM science & technology revolution https://www.evalapply.org/posts/after-ai/

TL;DR.

> I, for one, welcome the coming age of the post-LLM-datacenter-overinvestment-bust-fueled backyard GPU supercomputer revolution.

> The Big Question is…

> Who is cultivating the option to snap up and repurpose vapourised datacenter investments at fire sale prices, soon as the "datacenter debt" cometh calling?