Hacker News new | ask | show | jobs
by preommr 2 days ago
> Starting today, GPT‑5.6 Luna, our fastest and most affordable model, will cost 80% less,

I don't have the words.

I genuinely thought we were in a stage where we were plateauing and going in for 5-10% improvements over months. Seeing spikes like this makes me question about where the floor really is.

18 comments

When model intelligence reliably hits 90%-95% of current day knowledge worker tasks, they are going to burn those weight to silicon and we will see another 10X improvement in price/performance frontier.

The dynamic GPU clusters will be used for the 5% of tasks, and pushing out the frontier. Also there will be a set of knowledge tasks that are not done today (because they are too difficult for most knowledge workers), that will start being done in the future.

Burning the weights into silicon would be many orders of magnitude increase, not just 10x. It's kind of crazy that this hockey stick the AI hype bros talk about seems more and more every day like it might be real
https://taalas.com/ has done it already for a wildly obsolete model. 14000 tokens per second.

https://chatjimmy.ai/ is their interactive. Tiny context, very dumb, but absurdly fast. Imagine this as a tool call for claude code for trivial changes - the tool call from the harness takes longer than the execution.

Wow! You weren't kidding,

I just tried it too and 14,098 tokens in .05 seconds, I barely blinked and it was done. There was no typing at all appearing on the screen. It just showed up.

https://chatjimmy.ai/chats/01dc66a4-4b1b-4dea-bb5f-926855e37...

That link isn't bringing up your chat, FYI. It just shows the default new chat state.
Sure enough. I just tested it in incognito, and it is blank. I thought it would work because I wasn't logged in to Jimmy AI. It still shows for me though in my regular session so they must just have a cookie that makes it stick.

Here's a copy and paste prompt if somebody wants to just test it real quick to see what I saw:

Write a story about the fastest monkey who ever lived, his name is Jimmy and he is an AI superbot monkey that is part cyborg primate. He can travel through time and is psychic.

Wow. This is absolutely wild. I didn't expect that.

If we get to anywhere near this speed for the equivalent of the current models... I don't even know what to think about that future.

Speed is the metric I'm currently most interested in. The models are smart enough.

Once speed significantly increases I think we're going to see some interesting downstream effects. The three things I currently spend the most time waiting on are LLM API requests, Rust compile times, and nix derivations. As AI latency approaches zero I think we're going to start taking a hard look at whether slow-compiling languages are adding enough value over Golang, Typescript, or even dynamic languages to be worth the slowdown.

what do you mean "if"? of course we will, and the models will be smarter as well
Yeah. A model smarter than Fable by, say, 50 human IQ point equivalents, running in some kind of autoregressive process continously on whatever goals you put into the loop, on baked silicon, would be maybe my low end for the potential in the relatively near future.
Holy crap, I was not prepared for how fast it responded. I just wrote "Just wanted to see how fast you are! Can you write me a quick story about a tiger who lives inside a block of cheese the size of a house?"

I pressed Enter, and the response was instant.

> Generated in 0.037s • 14,205 tok/s

This is unbelievable.

It's crazy. Are they doing any precomputing as you type, I wonder if you paste a block of text is it the same speed.
The magic is in the fact that they essentially have an ASIC llm device. There is no other trickery. The problem they will face is that it is actually locked in silicon, so upgrading models will be difficult, and likely require new hardware each time.
I pasted and instantly hit enter on this prompt: "I generated a filter set using REW v5.31.3 using real-world sweep tone measurements from the room I'm listening in . How can I use it as my MacOS output equalizer so that my spotify music is adjusted for this room and speakers"

and it gave a very reasonable answer in non-perceptible time.

No, it really does take ~0.03s to generate the answer. Try your browser's developer tools and watch the requests.
It’s crazy that you asked the very question whose answer spawned the sub thread, within the sub thread.
For what its worth the frontier lab models can surely be a lot faster if they wanted them to be but theyre supply constrained so theyre doing stuff like multi tenancy. Since you cant self host them no one outside the labs really knows speed as a solo tenant
You can kind of get a sense by running these things at home - I'm currently running Laguna. One interesting thing is that per stream doesn't actually slow down that much with multiple concurrents, because the bottleneck remains the memory bandwidth until pretty significant request depths, and then eventually you hit the GPU's limits. It's one of the big forces that pushes for centralization in this stuff, the fixed costs to run one are huge, the marginal costs of additional tenants, relatively small.
"Stochastic gradient descent algorithm in Haskel"

"LMS algorithm in bash"

Just barfed it up lol.

Amazing.

Just wow, I made a similar request. Result: Generated in 0.042s • 14,201 tok/s

This is crazy.

I'd like to imagine the things that can be done with this speed and the current frontier models.
Fully interactive games where you can talk to every NPC by text or voice and have an LLM drive the story (with your own meta prompts to guide it, if you so wish). Maybe even have them generate assets on the fly too.

I’m still trying to figure out coding agents. I can’t even begin to imagine the things it would enable. Even the most mundane ideas like LLMs-in-HiFreq-trading have huge implications.

Truth, it feels like we're in the dial-up age of LLMs right now. And this Jimmy AI is fiber.
Seriously, if Fable or even Opus was this fast that would be a real game changer.
Wow! Responded essentially instantaneously to my prompt:

"I need a short, 4000 word essay on the the difference between star wars and Star Trek universes from the perspective of graduate level scientific work."

See Cerebras and Groq as well.
Answers like 8Bish model it appears, but instant response
hugged?
And don't underestimate how much money Google, Microsoft, Amazon and Meta still have to spend on this tech.

Blocking Fable for sure made it very politicl a lot sooner than i expected it to happen.

and because China already has massive problems of getting access, they are pushing it on hardware too like what Huawai did without EUV.

It seems China is already able to do DUV a lot sooner than others expected.

> It seems China is already able to do DUV a lot sooner than others expected.

That's the media and in particular US KOLs of all sorts driving the wrong impression of China and other places. China and many other places for example have fast public transport that the US doesn't and can't even imagine today. They're not behind.

China's DUV still isn't that production grade (mass produce-able) so don't get that hyped up the wrong way (in a different direction).

The whole China-is-behind with tech and in particular semi wasn't that they can't. The truth is they spent decades in internal politics and corruption. That all got solved with the bans, so thank the bans! Jensen even said the bans were bad.

Yep it will be ASICs and DSPs all over again. Orders of magnitude changes.
So which shovels companies are the ones to watch for burnt in silicon models ?
Imagine the price of a $9 million NVL72 dropped to about $100, used 4 orders of magnitude less power, was the size of ARM cpu, be bundled with pretty much any electronic device, and ran as fast a frontier AI is today.

That's about how disrupting DSPs were to the industries they arose out of (over a very long time frame).

How would that disrupt the industry?

$CBRS - Cerebras Systems

There is other in the space, Groq and Sambanova are both private companies attempting to develop their own technology.

Cerebras and AMD collabrated on some Helios system stuff.
I'm curious, how hard/expensive it is to burn a really large model into silicon, and why aren't we doing this already?

Or, when we will start doing this, who's going to be able to do that in scale?

I'm seeing the TAALAS example, but it's only an 8B model, suggesting some real limitations parameter wise. And for 2.5kW?

I can't speak for all cases, but the AI space is seeing improvements month by month, so it is beneficial to wait until it settles (a model becomes the standard in intelligence/price) before designing and mass producing an "LLM ASIC" of said model.

The big AI labs won't do that unless they are forced to, as they want you to spend more money on the big, expensive, frontier models (so they can live up to their valuation), so it's more likely that you will see this on smaller open weights models.

Not an expert on this but wouldn’t this be possible with something similar to an FPGA?
My understanding is that for FPGA the issue is either it eats all your gates on internal memory if you interleave, or it takes forever to load everything between the SRAM on the board and the actual FPGA component over a bus, last time I looked into it.
Weights can be baked into silicone or programmed into hardware ala FPGA, but the context will always be dynamic.

High speed SRAM is where the $$$ is

What does that mean though? Like some kind of a ROM memory ?
Stacked ROM can, in theory, be a lot denser than anything that depends on a capacitor and refresh cycle.

I don't think it would be that difficult to manufacture compared to other process tech. HBM is really hard to do compared to other memory types.

> When model intelligence reliably hits 90%-95% of current day knowledge worker tasks, they are going to burn those weight to silicon

Google is already working on a similar idea but more "flexible".

Explain.
I'm not who you responded to and I don't have any info on Google. Nor can I explain in detail due to NDAs. But multiple major players are working on something along the lines of what the parent is alluding to.

The "edge" AI landscape (in particular, what you can do with ~5W) is going to be nuts in about 18 months.

How will this affect the newly build data centers? What effect do you think it will have on memory prices?
My uneducated guess says, not much. For running massive models you still need a ton of high-bandwidth interconnects between many individual chips/GPUs/etc since you need to do math across a few TB worth of weights. That's simply going to require more power (and more die area in I/O, and therefore more cost). Being able to run small models in tiny power envelopes is incredibly useful to people, but I believe it will be covering a different niche than what datacenters can provide. Likewise, you'll still need crazy amounts of high-end memory to populate whatever goes in these datacenters.

The only thing that will crash prices is reduced demand (duh) or, more interestingly, increased production. In particular, if CXMT is able to get their DDR5 fabs up to a reasonably high yield, that could add some downward price pressure (as could government subsidies). As well, if Micron/Kingston/Hynix think that CXMT is going to start cutting into their market share, they might be willing to either increases supply or drop prices. Unfortunately CXMT looks to be taking quite a while to get their new fab up to max capacity so that may take a year+ before anything manifests.

If you're interested in following the (publicly available) info on these sorts of things, check out what companies like Axelera, DeepX, and MemoryX are doing today and have on their roadmaps, as well as the sorts of chips/SoCs Qualcomm, Kinara (now NXP), and Ambarella currently have announced (or have on the market). And remember, that pretty much all of these chips on the market today were in initial development more or less when ChatGPT first launched. If you knew what you knew today (or a year ago) about what requirements current- and next-generation models would have (from a silicon perspective), what might you do differently? Think for instance, host system interconnects, amount and speed of on-package or on-die memory, image/video decode capabilities, int8 vs fp8 vs fp16 vs bf16 compute units, etc. And, consider that most "AI" stuff in development a few years ago was all 15nm or 12nm - because who was gonna pay big money to get fab capacity at 3nm to run some object detection models? So most of the stuff on the market today is on very old nodes and therefore not super power efficient.

Google Frozen v2 chips hardwires the underlying architecture but leaves the weights to be configurable.
To be fair we don't really know in terms of prices what's real and what's just investor subsidised attempts at market capture at this point. It could well be OpenAI's attempt to drown Anthropic while they've got the halo product if they feel they've got deeper pockets.
Enterprises implemented spending caps and inference providers are lowering prices. Seems they are jockeying for market share.
Partnerships then consolidation comes next...
I wouldn't be surprised if they still had some margins since cheaper models are much harder to nail the accurate sizes off, and you still pay 2x for 1M context window.

But if this is even at 400B size it's insanity those inference prices, maybe 10-20% margins, if it's higher I would like to know is it their own chips or maybe they have accurately sized the model to fit on exactly a B300?

Could be a lot of magical things we can only speculate, but from here there likely isn't another 60-70% margin, like I have heard people claim, I would definitely be willing to bet on that.

Could still be a healthy 10-30% margin. Especially with Terra.

We can guess based on the decisions of other inference providers who serve these models.
Do you mean if other providers will cut their prices in turn?
Yes. For example, third-party inference providers serve DeepSeek V4 Flash just as cheaply as DeepSeek themselves, if not even more so. This is very strong evidence that the low price of the model is not subsidized.
I bet they’ll announce new fundraising soon. Someone gave them a top up to try and corner the market.
Hard to believe numbers. I don't mean that as a critique, but literally I am so impressed. Even if the model is a few percent lower for performance but is 80+% cheaper than competitors and is a US company hosted on US based hyperscaler clouds this is kind of a no brainer. Hard for most businesses to justify otherwise.
DeepSeek v4 Pro & MiMo v2.5 Pro (Opus 4.6 quality models for code) are insanely cheap for agent-driven work due to their super low cached-input prices ($0.0036/mtok) [0]. For Luna, the cached-input price drop isn't disclosed in TFA, but the pricing page puts it at $0.02/mtok, & that's 5x more expensive.

[0] I am constantly surprised how much work pay-as-you-go with DeepSeek / MiMo will get done. I've barely crossed $2 each in a month of use (~200m tokens).

Absolutely. Although DeepSeek started announcing "Peak valley" pricing which started making me nervous. I have spent $50 usd in July on deepseek and for that much spend I got SO MUCH mileage.

I feel perfectly content in using pay as you go pricing with deepseek. On the other hand, although Anthropic's models used to be my bread and butter for personal work, they are simply too expensive to reach for these days.

This is exactly the model that DeepSeek V4 Flash followed, and it's been insanely successful as a result, even though it's not frontier.
it's 80% less cost, not 80% in efficiency gains, could be that Luna was overpriced to begin with, we don't have much info on the models themselves.

Assuming the efficiency gains are real, I feel like something has to give, maybe worse quality due to aggressive quantization/kv cache compression?

Let's suppose each models was subsidized at 70%, so that we only pay 30% of the cost. They would loose much more money per token on the more powerful models. It's in their interest to encourage the use of the less expensive models. Let's say they increase Luna subsidies at 90%. They would still "save" relative to the use of the more expensive models.
> Let's suppose each models was subsidized at 70%, so that we only pay 30% of the cost.

why on earth would you suppose that?

Because their leaked financials point that way.
Source?
Not true at all. Do you really believe that generating 1m output tokens on sol really cost 100 usd?

These companies are printing money on inference. The only issue is the capex of expending on new servers but they are generating TONS of income.

If the frontier labs are indeed facing threats from open weight models, we would first see this in pricing pressure from the base model (like Luna) first.
High-performing open weight models being released recently, and your customers looking into working with multiple providers as a result, are a great reason to drop prices on your non-frontier offerings.

Although I'm sure there are some efficiency gains, the technology is too new and labs are scrambling to release too quickly to think that the low-hanging optimization fruit has been picked already.

Been using OpenAI models since ada/babbage/curie/davinci and at least from my own experience, their APIs feel the same.

If you use Codex it's different, the harness has a lot to do with it and there's definitely been changes including recently.

Something can be overpriced and still lose money.
Prices have consistently gone down 90% every 18 months like clockwork for about 5 years for the same level of quality. There is no end in sight for at least another generation.

There is ZERO reason to believe models 1/10th the size of frontier are completely capped on intelligence and impossible to get smarter.

They have consistently compressed the intelligence of larger models.

You'll see it first on the small end, when they stop being able to compress intelligence, you know that will slowly bubble up and up the chain to larger and larger models.

There's no evidence we've reached that at the bottom.

In fact there's very good reason to think intelligence has much more room to be compressible - the human brain, for one.
I think there's a ton of room for efficiency improvements in how the models are built and run, and I think OpenAI has both prioritized that work and figured out a lot of the tactics (and borrowed some from the Chinese models like DeepSeek and Kimi, which have published a lot of their research and tactics for running big models fast on minimal hardware).

I think there's also a new generation of hardware in the past year or so tuned specifically for LLM workloads, where it was almost an accident that GPUs worked to run LLMs before. So, while there's still this ridiculous shortage of hardware, what is being delivered is much faster and cheaper to run for these specific workloads.

I wasn't expecting it to happen from the US vendors, though, as they've spent so much capital to get to where they are they need to make huge margins on inference to pay it all back. I expected the Chinese models who're running much leaner operations to be the "frontier" on costs (and they have been). But, I'm glad to see OpenAI joining the "cheap and cheerful" models party. There's a lot of work in that area of capability. Probably most work people are doing falls into that area of capability.

5-10% over months would still be quite crazy.

But yeah I do'nt want to know what Kimi 3 is pushing buttons inside Anthropic, OpenAI and Google.

Besides any floor: For every year the tokens get faster and cheaper, we will see new things like properly working AI factories which mimic expert teams. A lot more parallism as well.

According to https://deepswe.datacurve.ai/, Luna at Max at it's previous cost was comparable in both performance and cost to Sol at High.

With an 80% reduction in cost that becomes a ridiculous outlier in efficiency.

Vera Rubin will be hitting racks very soon, and this is purported to have a 10x improvement in token throughput per megawatt. Of course, old chips don't get replaced with new chips overnight, but I don't think we're anywhere near the floor yet.
AMD MI400 series is already shipping to customers (basically everybody) and it is crazy fast (8x to 10x faster than the previous gen and beats published Vera numbers in FP8, loses in FP4) and 432 GB per chip. 72 chip unified rack architecture (Helios) already shipping and projected to also beat Vera in NVL72.

MI500 series is supposedly already taping out and they're claiming massive increases (we'll find out end of 2027 prob).

Absolutely true although at some point it's not just raw numbers but also the kernels that run matmuls and there seems (from outsider perspective) to have been more optimization in the cuda kernels
In a data center that is power constrained but not space constrained they could build out new racks and flip the power from the old racks. Wonder if this will lead to moderately used server GPUs on the secondary market someday.
I don't think they overengineered a DC like this.

Besides Nvidia Hardware is still sold out and super expensive. Not a single Nvidia consumer GPU got cheaper at all, Nvidia DGX Spark got more expensive too.

It will be swooped of the market the second it hits the market.

This is gonna put Sonnet 5 in a really awkward spot.
Sonnet and Haiku were already in an awkward spot, likely by design.

Anthropic's big marketing push this year has been entirely focused on getting people to use Opus via a Claude Code subscription, to the point that Sonnet is almost viewed as the poor man's alternative, and from what I've seen, almost nobody uses it.

Actually, here's an interesting project for all the vibe coders looking for their next front page post: scrape a ton of commits from GitHub with Co-Authored-By: Claude and figure out what the percentage split between Opus/Fable/Sonnet is. I'm willing to bet it's less than 10% Sonnet.

>figure out what the percentage split between Opus/Fable/Sonnet is.

This may be misleading, since I suspect many are using a blend through sub-agents. I tend to bias for Fable to orchestrate and Opus for implementation via sub-agents.

When I'm paying for it, Sonnet. When work is paying, Opus 5, then Fable if Opus gets confused.
I've found that Claude nearly always claims to be Opus in the commit message, regardless of the actual model making the commit.
Opus 5 is not strong enough as the top-of-stack model, and feels idiotic after a week or two of heavy Fable usage, to the point where I'm paying for Usage Credits to keep using Fable rather than having to slum it with Opus.
I use sonnet as a smart grep and haiku never and that’s only when I have to use Anthropic at all
Luna is comparable to Haiku, not Sonnet.
Totally untrue. Luna and Sonnet 5 are very comparable: https://artificialanalysis.ai/#intelligence

Luna is an extremely strong model.

> Luna is an extremely strong model.

By benchmarks, which sadly is a poor measure. Yes Luna is a good model under certain circumstances. Whether it is great for general usage is another story. Sonnet is definitely better when prompts are more vague and it needs to decide things. Luna generally sticks to things very strictly and goes off in bad ways.

I'm actually finding Luna to be an extremely strong model in practice.

I almost exclusively use it with xhigh or max effort, but when run like that it's been an incredibly cheap little workhorse for most development work. I'm still leaning on Sol for planning and debugging, but when it's time to start pumping out code I've been leaning into Luna (Max) and I've been enjoying it! And that was before the price drop, it's going to feel practically free at this point

Yes, I've seen this too and how Luna xhigh is so good that Terra doesn't really serve a purpose because beyond that you can continue at Sol medium. This can be the most cost efficient way, and especially now!
Yes but there is a big, big market for subagents to consume lots of tokens cheaply and condense information up to parent agents. Luna would not be my choice for planning. But an explorer to comb through a codebase to find relevant parts? Or for enterprise retrieval, where it needs to search across many different types of data to see where to focus efforts for a smarter model? Or to wake up periodically to evaluate some conditions and determine if a bigger model should be spun up? Definitely.

I've previously found flash (for all the hate it gets) to be good for these kinds of things. Haiku was fine but it's ancient.

> Yes but there is a big, big market for subagents to consume lots of tokens cheaply and condense information up to parent agents.

That's again not some "intelligence factor" here. Different agents work for different use cases. Luna wins some. Terra wins some. Sonnet wins some. Flash was really good at exploring.

So I'm not sure what your point is? There's a big market for everything. Even within the market you describe it's likely not a Luna-size fits all either.

In my real world use Luna is as useful to me as Sonnet. And it gets stuff done faster and follows my instructions more closely.
Anthropic basically downgraded all their tiers when they released Mythos, no nah, Sonnet 5 is the successor to Haiku 4.x
> Seeing spikes like this makes me question about where the floor really is.

You mean they increased the price and then cut it back and now it is amazing?

Luna had a price hike vs mini (its previous replacement). The cut now just puts it back in that ball park.

Not that this isn't good news, but what's impressive?

Had to ctrl+f for someone saying this.

I typically do lots of mini calls for research (100s of millions or something in that ball park). Newer models made that absolutely impossible, and the fact that the older ones are starting to get deprecated made me switch to e.g. deepseek for some of my runs. We'll see if I move back after this.

Luna's now cheaper than 5.4-nano (for output tokens). That's a significant improvement.
Totally possible that humans aren't actually that intelligent.
As well as the existing intelligence being swayed by emotions, hormones, circadian rhythms, stress, peer pressure, propaganda, and survival instincts.
As if ALL OF THAT doesn't represent inherent and crucial elements of judgement, and therefore"intelligence" itself.

We are not purely rational creatures, thank God. Sometimes those "limiting factors" you listed -- stress, peer pressure, hormones -- are crucial elements of informing the problem solving process and arriving at a decision or a solution that actually works.

All an LLM can do is fulfill a prompt, no matter how misguided, backwards, or incomplete that prompt actually was.

"Go jump off a bridge." Hmm. Dying makes me stressed out. I'm not gonna do that.

IMO we won’t see AGI until the models are sitting inside bodies with senses and gut biomes and opposable thumbs and the like, all influencing the context window in real time. Intelligence is a full-bodied experience
That's a bit like saying a tail is swayed by a dog, as if it could exist without one, or would have anything to do if it did.
Why does it matter? This is completely based on data produced by humans.
They have no (other) equivalent to nano, so it makes sense that it’s much cheaper now. It may have been better, but it was also hell of a lot more expensive.
more like they were facing pressure from chinese models, and dropped prices and now their margins are squeezed
there is a ton of downward price pressure from Chinese open weight models
this type of thing usually means you are the product
I don't see how this follows. The cost of nails has fallen by 95% over the last century. It's because the cost of manufacturing has fallen. Not because they are selling the information of nail consumers.

Tokens are not normal software, because they have marginal cost, and I think people who are used to software economics really struggle with this. With token generation there really can be manufacturing cost efficiencies where one producer is just straight up better at serving product at a lower marginal cost.

> The cost of nails has fallen by 95% over the last century

No it hasn't!

A century ago, some nails cost 2.5% of disposable income, and now the same nails cost 2.3% - only a little cheaper.

The cost of nails has remained remarkably consistent for a century. The problem is that you have ignored the depreciation of money.

Let's assume California prices and income and pick a bigger retail package of nails as you might use for building a house. The numbers used to calculate percentages: in 1926 a 50lb keg of 4" nails was $2.75 and median after tax income might be $108 per month. In 2026 a 50lb carton of 4" nails is $106 and income might be $4,516. Albeit I assume nails are now more readily available and the quality of nails is likely better; and perhaps I should have compared galvinised nail prices.

These are the API rates. They don’t train on prompts from the API. In what sense are you the product?
Until I stop getting downvoted for asserting that these guys are profitable per token, HN is going to continue being pikau face shocked at easily predictable things that any serious AI researcher would tell you, and has been telling you for years now!!!
They over purchased hardware.

This is very likely priced below recovering the cost of the hardware but still above operating expenses.

That’s ridiculous. Every major AI lab is compute constrained. That’s exactly why nvidia is worth trillions today. If OpenAI had a single extra GPU they’d be using it to run another training cycle or experiment for their next model.
What evidence is there?

I have no idea either way but one thing that detracts from these threads is folks claiming things as a fact without evidence.

sama literally just said they wish they had bought more. the price drops are almost certainly due to good old fashioned hardware innovation (wafer scale with cerebras) and optimizing hardware development based on model architecture and inference costs. other inference providers will try to do the same if they can.

https://www.youtube.com/watch?v=XDB5beon4DY&t=4m20s