Hacker News new | ask | show | jobs
by nkozyra 16 days ago
No.

But do you need to run every small problem through a 10B-30B model?

We're smashing ants with hammers most of the time. We're asking frontier Opus/Fable models to classify text and build frontend code.

Once we start dissecting these problems into smaller discreet tasks and having the big reasoning models do the tough stuff, we suddenly have an economical system. Not for the company hoping for a big IPO, but for the end user.

4 comments

I do about 10 google search queries for every 1 opus/gpt prompt. For google, I don't actually open pages anymore 9 out 10 times; I rely on the AI summary. It's fast and accurate; the trick is that you learn where the boundary is of what you can ask it. Querying information the small model is great at.

Then there might be slow, batch tasks. I can see myself getting 1T of slow RAM one day (in a few years?) and having a slow onsite GLM5.2 doing batch jobs that would be wasteful of my subscription limits, plus sensitive but boring things, such as bookeeping and general admin.

I'd like to to read all my email and al quarterly reporting. But that would have to be a good local model, probably a model simmilar to whatever google search uses, which seems just correct unless you throw serious challenges at it.

The big hammers buy you more confidence and need less supervision. In pure task execution you _might_ smash the ant with a small surgical hammer, but if you absolutely need it smashed, that's when people still reach for the big hammer. It buys more confidence.
> do you need to run every small problem through a 10B-30B model? ... We're asking frontier Opus/Fable models to classify text

Actually probably yes: text analysis (magazine articles) by LLMs in the ~30b .. ~120b range failed miserably (and also randomly - the rare cases of proper interpretation occurred among the failure cases) with the main public models of around one year ago, tried extensively.

So, yes, you can employ an ~80IQ only if you will expect the related quality.

I don’t think you need a 10-30b model for most smartphone use cases.

But I meant to counter gp’s claim that “I can run a 27b model on an iPhone” is kind of pointless and disingenuous. Yes I’m sure someone will come up with a way to run a “27b model” at 0.1 bit quantization on an Apple Watch pretty soon misses the whole point of saying a model is “27b” in capability.

Achieving a parameter count is not the point. And is almost meaningless