Hacker News new | ask | show | jobs
by rmunn 14 days ago
Given the power draw of the GPUs and RAM needed to run those local LLMs, I don't see them being comfortable to hold in one's pocket anytime soon. :-) Not until some not-yet-imagined breakthrough in heat dissipation is made. But a thin client that talks to the local LLM on your private network, that's already possible today. So yes, within the next 5-10 years I fully expect I'll actually want to use an AI assistant on my phone. Currently I go through the settings and turn everything related to Google Assistant off, and do that again after every Android update. But once it's talking to a model that I control, instead of a model controlled by the world's largest advertising company, I'll feel the opposite way about it.

P.S. If "clown-based LLMs" was an autocorrect-assisted typo for cloud-based, it was inspired. :-) If you typed that on purpose, it was also inspired. Have an upvote.

1 comments

Power consumption per unit of work keeps going down, too. Our pocket supercomputers are astoundingly efficient compared to what it used to take to get the same work done at relatable points in the past. We'll get there.

Before we get there, we'll have network connectivity to our desktop AI boxes at home. That's pretty good, too.

And if this clown-bot[1] boom is a bubble (as I believe it is), then it's just a matter of time before it pops. The blast radius is unknown, but at one end it seems likely to result in something between potentially-idle fabs that have already been mostly paid for. Idle factories are bad and it makes sense to avoid that even if it is expensive, so this means cheap hardware.

At the other end, it means failed companies with fabs that are sold for pennies on the dollar alongside a resolute unwillingness amongst the investors (who just had their asses handed to them) to start on Boom 2.0. This latter scenario would mean golden age of cheap hardware.

I think we'll be fine.

[1]: Yeah, clown is on purpose. It fits as a replacement in any technical parlance where "cloud" would be used. :)

I'm of the same opinion re: the bubble. The dotcom bubble held on longer than I thought it would (I had been predicting that it would burst in 1999), so I'm not sure I should be making predictions as to when the rental-LLM bubble will burst. But I'm sticking with the hardware I have until memory prices finally fall again. Whether that's only because of spare fab capacity as you mentioned, or because one of the major companies has gone bankrupt and their already-purchased hardware is being sold off at fire-sale prices (perhaps to other companies, who then reduce their new-hardware orders accordingly, which again results in spare fab capacity), I'm going to buy more RAM when it's cheap.

... And I've just been agreeing with you on all points with this comment, haven't I? Time to wrap up the discussion, then: once total agreement is reached then there's not much point in continuing the discussion unless someone has something new to say, and I don't have anything new to say on this topic. See you around.