|
|
|
|
|
by daveguy
19 days ago
|
|
> Often yes. In this case, it's more like they get upset when someone says something factually wrong, and then defensively changes the goalposts. Oh give me a break. Show me one example of 1) any knob twisting that makes the underlying model better. or 2) any example of the AI providers twisting those knobs to do anything other than degrade performance for their own bottom line or safety. The current post says: "it would be expected for a better model to use different amounts of brevity if it gets better at determining the appropriate amount." When no, the model cannot "get better". It doesn't determine any appropriateness of response realtime except for the weights baked into it from the beginning and whatever context it can muster. If you cram enough guidance that it doesn't decide to ignore maybe you can make it more brief. But it (the model) can do none of those things. LLM models are literally stupid by design. |
|
Your comments are conflating multiple kinds of “smart” and “better”. You’re right that if all the inputs are exactly the same, it takes a new model to improve (ignoring non-determinism). But the knobs and context and harness change the inputs, and they do improve output, contrary to your claim. You’re failing to capture the distinction between what the model itself does and how the harness can boost the model’s performance. It is legitimately valid and fair to call improved performance “better”, no matter where it comes from.
This all gives me the feeling you might not have experience with or understand what’s happening in today’s harness development, and the degree to which it may be as important as the weights. There are in fact a lot of things you can do to improve a model’s performance on tasks & benchmarks, without changing the model weights. @coldtea mentioned a bunch, but the harness feedback loop, internal prompts, system prompts, skills, and requests for a model to try harder, and verify and validate it’s output all lead to improved performance, all without retraining.
I agree LLMs are stupid; they’re statistical token predictors. But somehow statistical token prediction is amazing and works much better than we imagined. The talking points about LLMs being stupid token predictors are fading now because they lack explanatory power for how good the models have become. The big surprise here isn’t about LLMs. It’s about language, and how much “thinking” and intelligence is contained in language. We don’t have a good grasp on where the line is between language and intelligence. LLMs have crushed the Turing Test into dust, and yet we don’t consider them intelligent. They often appear to understand what you ask thoroughly, can re-state it in different words, they can correct your misunderstandings or add nuance you didn’t see. All this because that’s what humans do and LLMs talk like humans.