Hacker News new | ask | show | jobs
by rsync 25 days ago
"GLM 5.2 is "almost Opus," and it needs at least 8xH200s for comfortable inference ..."

What is the behavior if one were to run GLM 5.2 with only a single H200 ?

Would it fail to run at all, or would it just run so slowly as to be unusable ?

I would like to prove out the build, and concept, of a SOTA model locally, but then backfill the rest of the GPUs in 18-24 months when they cost significantly less ...

1 comments

> in 18-24 months when they cost significantly less ...

going to need you to sit down for this one...

Say more. My expectation is that the current gen of gpus will start being replaced by the next gen, and then it may be possible to get used ones that are still within their useful life at lower prices. My expectation is also that memory vendors are likely to increase production, which will drive those prices down eventually. Maybe not over the next 18-24 months though.
The only thing that diminishes the value of a GPU right now is unsupported features with outsized value during inference and/or training (like FP4 support) and it takes time for those features to actually take off

And labs are fully leaning into pricing for intelligence, so their margins are improving very quickly (which allows them to pay even more for existing compute)

I'd be shocked if current prices aren't the bottom for the next 18-24 months.

Many newer Chinese lab models are releasing with int4 native weights. Latest NVIDIA generation GPUs have a hard time with this and can actually be slower than previous generations. This may make Blackwell depreciate faster than other recent generations.
That's not a real problem, hardly any 3rd party was running native weights anyways so they'll get quantized to NVFP4
Given the prices and shortages, I'd think people would keep and use the current gen stuff till it drops dead. It may not be as good but it's paid for and given the prices for next gen stuff, it's probably worth using for another cycle or two.
I sat in a meeting 7 weeks ago where senior leaders said they expect token prices to drop significantly over the next 6 months, and we should all be using as much AI as possible; our team goal was set to use more tokens.

This week, we are banned from using anything more expensive than opus 4.6 and encouraged to use sonnet (but not sonnet 5! Thats expensive!) or lower for daily tasks to help manage costs.

Weeks ago, they gave exactly the same justification as you just gave; and it makes sense!

…but maybe not over the next 6 months.

> Maybe not over the next 18-24 months

Maybe not. Probably not, I guess.

A lot of money has been invested on the expectation that the current gen of hardware is going to reap a colossal profit, and the capex to replace it, is vanishing into investor skepticism as we speak.

It seems like most people have a very very low ability to forecast long horizon change in the current environment, but, in general… it seems like until demand drops, the chances of prices dropping is dubious; at best we get a price war with chinese models or a bubble pop; and even then, there are plenty of startups lurking to snap up cheap hardware.

For individuals, the horizon for buying cheap AI capable compute doesn’t seem close, at all, to me.