Hacker News new | ask | show | jobs
by gyomu 16 days ago
In 5-10 years an Apple Watch will run a Fable level model locally. I don’t think we (hackers) should worry too much about token cost inflation. The current wave of providers, that’s another story.
6 comments

Apple Watch with 1TB of vram with the size of well.. a watch.

Amazing story. If we make such leap in semiconductor field, it will be bigger than anything we have done till now. And all of that in 10years!

No, it won't. We moved about order of magnitude that from 1990 to 2000.

The thing is, it needs demand to drive it. Laptops have been roughly the same spec for the last 10 years because we don't need them to be bigger; there's no demand for a 16Tb RAM laptop because we don't have anything that could possible need that much RAM. Until LLMs came along, and we all want to run them locally, and so now there is a market for 16Tb laptops. So we'll invent the tech to make that happen.

Right but 2000 to 2010 didn't have similar progress, and especially 2010 to 2020 didn't. Sure, things have gotten better but not as much as the 1990 to 2000 leaps.

And yes, laptop specs haven't changed much and this is partially because the need for spec changes wasn't present, but also during the last 20 years there has been tremendous pressure for efficiency in datacenters.

Despite that, dennard scaling is dead since 20 years. There are physical limits. Already now, the wear effect of electrons jumping is present, and it will only get worse as things scale towards smaller sizes.

There are some benefits to be had, e.g. one can etch models into chips directly so you can pack them more closely, and run more inference on Tensor like chips, but that gives you maybe one order of magnitude improvement in total, at most. Also, of course nobody does that when each 2-6 months a new model comes out.

The thing that we did in 1990-2000 was adopt new standards as the old ones became blocks on progress.

I had a friend working in optical computing back in the late 80's that would wax lyrical about how optical computing was vastly superior to silicon back then. But it never took over because silicon worked well enough.

If we've hit the limits of silicon then there are other options. We would need to reinvent huge chunks of our tech stack, and that is incredibly expensive, but if the demand is there, we'll do it. The demand has never been there.

The original claim from the parent comment was running a Fable-level comment within a decade. Even if you're right about whether it's possible that another model could support that level physically, do you really think that we'll figure it out and ramp up the infrastructure to profitably sell on come consumer hardware anywhere close to that soon?
Well, we did do similar stuff back then. All it takes is money ;)
If it happens, that’ll be proof enough for me that LLM assisted science is giving economic returns!
How will it do that without burning your skin? I don't think we're seeing exponential decrease in compute cost. Right now it costs a lot of money, power, and heat to run Fable.
I gave you a upvote simply because I see your prediction just as likely as anyone else's here, meaning nobody has any idea what this space will look like in 10 years.

I hope everyone reads these LLM threads like your post, complete shots in the dark because nobody here will get close to predicting what the environment will look like, even the "insiders".

No, Apple Watch won't have the compute or the ram needed in 5-10 years. That will require 8-9 doublings in RAM and even more in compute.
You can’t seriously think an Apple Watch will have hundreds of gigs of vram in 10 years?
Maybe he thinks the model will be smaller.
I doubt it can get < 100gb. Isn't fable like 10 trillion params?
The math doesn't math. Wearables have to be small and body temperature, which is not something that fable likes.
why would you need a local fable at that point? AGI will surely solve all the problems in the world at that point
I’m kidding but.. it’ll be great if in 10 years we look back on this and your joke has become a prophecy.
I'm somewhat serious -- if you think AI will scale that well, you can't really make predictions like that

I personally don't think the weight efficiency will improve that much; if anything big does happen, I expect it to be about scalable architectures and continual learning

great???