Hacker News new | ask | show | jobs
by SwellJoe 2 days ago
I think there's a ton of room for efficiency improvements in how the models are built and run, and I think OpenAI has both prioritized that work and figured out a lot of the tactics (and borrowed some from the Chinese models like DeepSeek and Kimi, which have published a lot of their research and tactics for running big models fast on minimal hardware).

I think there's also a new generation of hardware in the past year or so tuned specifically for LLM workloads, where it was almost an accident that GPUs worked to run LLMs before. So, while there's still this ridiculous shortage of hardware, what is being delivered is much faster and cheaper to run for these specific workloads.

I wasn't expecting it to happen from the US vendors, though, as they've spent so much capital to get to where they are they need to make huge margins on inference to pay it all back. I expected the Chinese models who're running much leaner operations to be the "frontier" on costs (and they have been). But, I'm glad to see OpenAI joining the "cheap and cheerful" models party. There's a lot of work in that area of capability. Probably most work people are doing falls into that area of capability.