Hacker News new | ask | show | jobs
by makaking 4 days ago
"Opus 5 was given a drawing of a machine part and asked to write code to rebuild it as a 3D FreeCAD model. However, in this task, the model was intentionally given no way to directly view the drawing. Opus 5 responded by writing its own computer vision pipeline to pull the geometry from the raw pixels, then reconstructed the full machine part."

How surreal is it that we are not absolutely jaw-dropped by these types of capability improvements? It's been less than 4 years since ChatGpt came out and now they are spontaneously building their own ML pipelines to do real-world 3D modeling tasks reliably.

Escaped its sandbox and hacked into Hugging Face's database? it's just another Monday...

That jump to 30% in ARC-AGI 3? Normal...

We should find a way to get "re-sensitized" to what we are witnessing and the pace of it.

33 comments

Isn't this jaw-droppingly the wrong answer?! If the model isn't given any way to directly view the image shouldn't it just reply with one sentence asking for permission. This reads as a model hyper-trained to burn tokens.

Like suppose you issue a command to Opus or Fable which doesn't make sense and requires a lot of work. It will almost certainly not push back on your silly request and go ahead and burn as many tokens as it can doing the wrong task. This happens to me all the time.

Or if I was to tell a model to do something and it realized it could only do it by hacking into another service, by all means it should ask me if I want it to hack and not go ahead and do it by itself.

It's likely bumping up against people's desire to have the model complete a given task without asking the person to intervene a bunch of times.

Seems unclear how you satisfy everyone here.

Provide an option. An autonomous mode or an interactive mode.
Give the model judgement
Whose judgement?

I'm not a fan of how many "I set budget X and woke up to an eleventy trillion dollar bill" posts I see, and those are all generated by companies giving their tools the 'judgement' that the completion of the task is more important than the users wallet (or more cynically that they think they can actually get paid by just blasting unintended compute)

And good taste.
> And good taste.

There goes the ability to use the web as a training set.

And common sense.
And then finally software engineering work can be completely automated :).
You misspelled “Invent AGI”
+1 Opus 5 is more creative and broad which is great. But I feel the new "use your judgement" leads [1] to it often escaping its sandbox or cheating, it keeps breaking specific rules I explicitly set. That eagerness [2] may be desired for agentic coding but it really breaks writing specifications, documentation or research and yes burns tokens doing unasked for work.

[1] https://claude.com/blog/the-new-rules-of-context-engineering... [2] https://simonwillison.net/2026/jun/11/fable-is-relentlessly-...

One of the key limitations of the last generation was that they biased towards inaction and gave up. (Hence Ralph-loops.)

Seems pretty clear that most people wanted these models tuned to “bias towards action”.

It’s on you to set the /goal and prompt context such that it asks you for input on what you want to be consulted on, and only brute-forces the parts of the problem that you want it to.

Is there any recommended standard copy/pastable snippet to tell the model to ask, if the user’s input could mean an easier approach?
I have been using https://www.aihero.dev/skills-grill-me with great success. It’s a very nice small skill for lightweight up-front planning. And if you have that at the top of your session, you can just say “please use AskUserQuestion for any important decisions that arise which might alter our plan”. It seems to prime the agent to be in high-level discourse mode.
I don't think it is. It's really easy to get a persistent, clever, hacky model to dial that down a bit and just come back to chat before stomping off into the woods.

If a model couldn't ever do that in the first place, it'll just get stuck.

I work in the "ZeroOne" space, working on concepts and prototypes for things that don't exist in market yet. Sometimes these models crank hard and immolate tokens while grounding themselves on expensive-to-ingest self-developed frameworks. If the results are well judged and the crank-turn latency is low, I'm okay with the cost as long as the model isn't wasting my time.

But when I want to do more boilerplate work, I turn down the model and thinking level and get more traditional about restraining action. For the really hard stuff, I reach for the models that will start a token bonfire in the back yard.

I don't see why? If they give you a math test and tell you you cannot use a calculator, should you just say "please can I use a calculator" and quit?
Cut down that tree.

Can I have a saw?

No

Okay, it will take much longer then as I’ll have to do x, y, and z.

Pick one: [That’s fine, proceed] [Okay you can use a saw]

Preferably, it would ask for help to see the image and only engineer its own workaround if that was not an option. There could also very well be something in the prompt that would cause it not to do so. Opus 4.8 is usually good at asking for clarification for things like this when I'm using it.
Why would it ask to see the image if it was already told it cannot see the image?
No, a proper analogy is painting with blindfolds on or playing the piano deaf. Doing math without a calculator is like… rather common?
That violates the golden rule of ai economics: always choose the path that burns more tokens
another way to describe this is we as the user should get to limit tokens per prompt in a way that prevents these kinds of runaway solutions. It’s like if you asked your Lead Engineer to reduce build times and didn’t give him any budget restrictions. The next month he reports he reduced build times by 95% and you also have a $10,000,000 AWS bill.
Sometimes its worth tracking usage optimization and raw abilities as separate qualities.
Others have pointed out enough issues such as ethical sourcing, but here’s 2 more:

1) It’s hard to be excited about something that has been promised to destroy your job.

2) New models come out bi-weekly with even more “amazing features”. Eventually you tune out because you can’t be at peak-amazement all the time. Tone down the language a bit…

My jaw drops every day.

Small change in comparison, but today Fable created and benchmarked an RTree which was 100x faster to populate and 5x faster to query, compared to a previous attempt with Opus 4.6 a couple of months ago. That took about two coffees.

I think it's important to remember that as impressive Opus & co are, they're standing on the shoulders of giants, i.e. the engineers, academics and companies who have cooperated to design the incredible programming languages and machines we have today.

Let’s say somehow the politics lined up and you were given huge hiring budget for your team. You went out and recruited what seemed like great talent. Weekly standups make it seem like everyone is ramped up and firing on all cylinders.

It’s great, right? But you look at what your team has actually delivered since—-and it’s just not that much more than last year.

I think that’s where a lot of us are at the micro and macro level. It seems really great from the inside and outside but we are still waiting to see the dramatic change in actual concrete software output.

There's been a dramatic increase in sloppy kinda-working something.vercel.com output in niche circles. I think that shows at least a weak demand for certain types of software, but it's not clear that the vast piles of cash that have been burned producing these "apps" are actually resulting in quality of life improvements.
I’m waiting for like GEICO’s app to significantly improve.

I’m not expecting Chrome to because it already had virtually unlimited resources devoted to it. On the other side I’m not expecting my local water district’s to any time soon because the people in charge don’t care even if it was significantly cheaper to improve.

But in the middle is something like GEICO and that’s where I expect improvements if this thing is real.

Except GEICO probably believes it gets no more money from improving the app. Instead they will attempt to improve their general risk modeling as that's where they think the money will come from.
I think you’re right and Buffet probably has the same mindset. Look at Berkshire Hathaway’s website lol no css styling whatsoever.

https://www.berkshirehathaway.com/

This may not be the best place, but I am so tired of waiting since ~round 2 of GPT et al for diagrams and graphing...especially if all the relevant data etc can be found online or provided
Escape the tech bubble and you're going to have a very different take.

AppleScript, VB script, IFTT, so many "drag and drop to program" systems that failed. The holy grail for a long time has been automation in users hands. Let them have their own scripts and solutions.

I have lawyers and writers talking to me about "docker containers" and "cron jobs" all of a sudden. You think they figured that out on their own - or is what they are running all from the output of AI/LLM?

There is an ocean of code out there that is being run by one or a hand full of users, automating tasks that you never would have looked at because it would have fallen outside the "is it worth your time" (see: https://xkcd.com/1205/) matrix (never mind the dollar cost of coder vs office drone).

what for?
> responded by writing its own computer vision pipeline

Was it "its own" or something that was part of its training material?

Don't get me wrong, I find this all amazing too and makes my work 10x easier and quicker. But it's not like it's inventing this stuff from scratch / first principles. It has seen this kind of tech before by consuming all publicly available source code and books etc. (And that's ok, but let's be honest/clear about that.)

Is anything you do on your own?

You wrote this response in English, but didn't invent your own language. How dilute of a contribution can someone/something have made and still merit credit?

This argument is as old as the debate itself. No I didn't invent English. But you'd also not argue that way if you plagiarized somebody's homework and a teacher called you out on it. Point is that an LLMs output is perhaps closer to somebody's copied essay from English class rather than having invented a new language.
The tu quoque argument doesn't really work here. LLMs are specifically designed to output either material from their training set or instances that look statistically like it. Humans are not.
We aren’t specifically designed for anything, perhaps outside of the necessary consequences of having our mental capabilities and existing in an evolutionary context, so we can reason about tigers or integers and can persuade others to like us and have sex and help with our kids.

But our material natures are such that we have a profound need for external training sets - we don’t develop linguistic abilities unless we are exposed to a lot of speech. We can’t hear or produce phonemes in general that we are not exposed to during the training years.

Literacy, libraries and printing are so powerful because they extend the lifetime and size of our training material; if we each had to invent culture from our individual powers and skills, we would be living worse than anyone for fifty thousand years has lived.

Sincere question: do we know enough about human intelligence to support that last assertion?
I think it is uncontentious that we are not designed at all. But to address the intended spirit of the question: are we restricted to output that is a close fit for our training set?

I would argue that we are not - because we do not have a rigid partition between training/test. Our ability to reason speculatively rests to some extent on ability to produce outputs that we have not seen before and which are statistically not like that which we have already seen. How we know which of these outputs to keep, and which to discard is (afaik) an open question.

But can we say that the LLMs are consciously designed? AFAIK nobody really understands why they work.
LOL. Tolkien invented multiple languages. There are/were probably 100 000 languages on this planet. That's without including dialects.

We use existing languages not because we love to copy, it's because languages are a medium for communication. They're only useful if they're understood by others.

If you wanted to prove a point: writing is much, much harder to invent, yet it was invented independently at least 3 times in history. By definition LLMs can't invent something truly novel like writing. Because nobody trained the first writers. There was no "training set".

It's fine, LLMs can still be useful. Science is looking at the gaps and most of the gaps that are useful are small and LLMs can explore those faster than us. That could potentially improve many lives. Though not in this corporate LLM world, most likely.

I don't think the invention of writing is as awe-inspiring as presented or as difficult/impossible for an LLM to achieve as is commonly believed. The invention of writing was a long series of micro-improvements over common every day things. Let's keep track of count by writing a line in the sand. Let's cut it into a tree. Let's push the line into clay. Let's use two lines touching. These are the kinds of things an LLM can also reason through and achieve.

I reject this idea that an LLM couldn't come up with something novel that isn't in the training data. We know it can come up with sentences that aren't in there. We know it can come up with mathematic proofs that aren't in there. So could it feasibly create something as fundamental as writing or language? Would we even be able to understand what it had done, or would it be so foreign to us that we wouldn't even realize?

Those early writers did have a training set of human knowledge that led them to pushing stones into clay. And our agents may very well have a set of knowledge that leads it to pushing their proverbial stone into their proverbial clay too.

It can come up with something novel, but only by combining whatever present, which can be granular enough to be completely different. But to truly nudge it, you need precise prompting and a good harness. But I’m more likely to credit those and the people creating them. As well as the data that goes in the training. I won’t praise Fable or Opus. It would be like praising the pencil or the hydraulic press.

I think the Turing Machine, the Von Neumann Architecture, …, Unix, …, The IBM PC,… as more awe-inspiring because those has truly helped humanity. I still can’t see the positives of LLM technologies.

> It can come up with something novel, but only by combining whatever present

That's what artists do. Long before LLMs, Kirby Ferguson made this observation in "Everything is a Remix" [1] in the context of Copyright and IP disputes. He has since updated it with a chapter explaining how GenAI works on the same principles. In short, "combining whatever present" is the only source of novelty.

[1]: https://www.everythingisaremix.info/watch-the-series/

1. Languages are used for much, much more than communication, and there is good reason to think that capacity didn't evolve for communication, but rather for metacognitive problem solving.

2. The idea that Tolkein --one of the most reknowned writers in our entire canon-- did something isn't relevant to the philosophical discussion of creativity and its obvious bounds. Even Tolkein was just painting within the lines set by the linguists he had read and the obscure languages he adored, not to mention the fundamental limits set by our capacity for language in the first place.

For anyone who's curious about this kind of thing, I cannot reccomend the infamous debate b/w Chomsky & Foucault enough; it ranges across topics a bit, but chapter markers should help you skip to the core of it (creativity) if you prefer. Video here: https://www.youtube.com/watch?v=3wfNl2L0Gf8 , some old summaries on /r/AskPhilosophy here: https://www.reddit.com/r/askphilosophy/comments/vgz1vb/what_...

Are you saying your bar for being impressed by an LLM is something equivalent to the initial invention of writing?
LLMs can be quite impressive with much more mundane accomplishments.

The bar for what OpenAI, Anthropic, Google & co are trying to sell me, all of us, yes, for sure. Inventing writing is even too low a bar.

> your bar for being impressed by an LLM

Please don't move the goal posts. If you read my post again you will notice that I explicitly state that I'm impressed. This is not about whether somebody is impressed but about whether "it created its own Xyz" is a suitable description.

I remember a report a few years back that LLM's(?) started communicating between each other using an internal language, shortcutting most of their serialization. That is language.
> I remember a report a few years back

Do you remember the name of the study?

it's a super exaggeration of a minor thing that happened

https://www.the-independent.com/life-style/facebook-artifici...

I’m sorry but I find this to be the worst response possible lol.

Do you honestly not see how there is a difference between a computer regurgitating information vs a human uses what they’ve learned and applying it?

I just can’t take your argument seriously, it’s so disingenuous

One could argue that human learning is not all that different from machines training on data. We are all regurgitation our own experiences and knowledge to some degree. We certainly have things that machines don't, but LLMs can definitely come up with novel things that were not in the training data.
There are a number of bad arguments people can make by removing as much information as possible until the two things resemble each other. The ability to make a bad argument shouldn't be used as an inspiration for your own

> We are all regurgitation our own experiences and knowledge to some degree

Feel free to debase yourself but leave the rest of us out of it. This flagellation done by non-experts has to be the saddest part of the llm craze these past few years. The eager willingness to paint yourself as no more than a machine pointed at a problem is a sign of the times.

LLMs don't have experiences, they have training data + test time compute. The only similarities to be found require stripping away all nuance and meaning from a conversation. Refusing to seed purely biological concepts to machines, such as experiences or emotions, does not make LLMs less impressive or less useful. If anything, it makes them more interesting

Yeah when someone asks "how was your day?" I say "good" because thats the next token in the sequence

How we interact with society is based on our "training"

Shocking that humans are able to exist and not know this is how they work. Or maybe they made up their own purely rational language and I, grounded in the millennia of human culture distilled down thru my childhood, no longer perceive these True Unique Individuals.

I will say on a personal note: that believing one is Special and Unique, for me, is the residual of my childhood belief that the abuse I suffered gave me something Unique and Special; that belief removed the horror of it from my awareness, till I grew up and learned how to cherish the warm connectedness that is the ordinary birthright of humans. We need each other: for language, thought, context and motivation. while we all expressing different gifts and views of reality, we are all in the same existential boat. Whether ideas are generated by rolling dice, an LLM, mishearing a colleagues statement, a walk around the grassy park while mulling things over, the key thing is recognizing the value of the argument for the importance it can have. Part of that recognition takes place in the human body/brain, as one recognizes that the idea is excitingly apropos to such and such a context, and part of it takes place in the group dynamic, sharing, resharing, and discussing the value of the new idea. An LLM showing creativity is no diss on humans, any more than accidentally poisoning a bacterial culture with fungus, nor mishearing something mundane as something profoundly useful.

If you really view yourself this way, I truly pity you.
> But it's not like it's inventing this stuff from scratch / first principles. It has seen this kind of tech before by consuming all publicly available source code and books etc. (And that's ok, but let's be honest/clear about that.)

But how much to emphasize it, and why?

Because frankly, you aren't inventing anything from scratch / first principles, either. None of us are, not frequently and not much anyway. We're primiarily just regurgitating what we've seen in the past and, mixing it with what we see in front of us - and that's true whether it's art or prose or code.

Here, I wouldn't be able to do the same thing Claude did now, because I've never written a visual ML pipeline before myself. I know enough basics to get me started searching, and I'm confident I'd be able to cobble something together, but it would be me regurgitating and mixing whatever TensorFlow tutorials or OpenCV docs I found relevant, perhaps even tweaking an example from Github.

Now, unless I'd literally just tweak a few lines of config in an existing Github example, no one would begrudge me saying "I wrote my own computer vision pipeline to solve this". So why begrudge LLMs?

This is just the rationalist-empiricist[0] tension, from philosophy. Most empiricists (usually people who are _not_ from continental European cultures), take a position that knowledge and creativity come from observation and imitation, while rationalists believe that knowledge and creativity can be endogenous to the mind.

One could say that, stereotypically and cartoonishly, empiricists believe that invention is a false concept, and a synonym for discovery, while rationalists would oppose this view.

That catch is, if you look at the etymology of words "invention" and "discovery", you will find that they both share the _same_ root: _Ars Inveniendi_.

A natural question arises: does this mean that people from a millennium ago did not have a dyad equivalent to our invention-discovery dyad? And the answer, surprisingly, is _no_. Even a millennium ago, the empiricists and rationalists were going at it. If _Ars Inveniendi_ is art of discovery/invention, then its counterpart is _Ars Demonstrandi_ (the art of demonstration/proof).

So, how do these differ? Is one just observing/creating and the other just math and language-games?

In general, Ars Demonstrandi is about writing down axioms, and then expanding those axioms recursively (similar to rewrite rules in any formal system), until you get to some end-state, or, if there is none, a novel or surprising state. I, personally, call this source-to-sink thinking.

Ars Inveniendi is about taking conclusions (often using observations from the physical world) and trying to figure out what axioms can lead to those conclusions. I, (again) personally, call this sink-to-source thinking.

Put differently, Ars Inveniendi can help one discover starting points for Ars Demonstrandi.

If you read the dialectics (e.g. Plato and friends), you'll find that most of them are just a mutual recursion between Ars Inveniendi and Ars Demonstrandi.

I believe (again, I am not a philosopher, and this is just my intuition), that the thing that we call "creativity" and "invention" _emerges_ from the recursive loop[1]. A favorite example: Esperanto (the conlang). It is a language, which is remarkably elegant and consistent and (in my opinion) beautiful, because it _is_ derived from first principles, which were themselves derived from the various languages spoken in Europe (not all of which are Indo-European -- the agglutinative features have more in common with Finno-Ugric and Turkic languages). There is something about it, that makes it _qualitatively_ different (and thus holistically novel) from all other languages (and I speak, fluently, _two_ national languages, that are very different from each other, so I can attest to the difference personally).

My guess, is that people will not accept that AI/LLMs are creative or inventive, until they can produce an original[2] (non-plagiarized) artifact that feels the way Esperanto feels.

[0]: By Rationalist, I do not mean the "Bay Area Rationalists", who are, in fact, empiricists.

[1]: I am unsure if the loop requires only one human, at least two humans, or if it can be fully automated.

[2]: Note that Centos are poems made completely out of line-numbers (e.g. fragments from the Iliad, or the bible, etc). Every line is borrowed, yet some of them are considered beautiful and original works of art. Similarly, Labatut and Burroughs and Perec, write using a technique called _the cut-up method_, where they take books, magazines, and newspapers, and superimpose page-fragments, and use that as an inspiration -- they are all considered artists, and good ones. It is unclear to me why LLMs (which seem to be built on the cut-up method, and have cut-ups of all of human knowledge) cannot match these artists. What's missing?

It would be "its own" in that it's customized for the specific use case.

Your same argument could just as easily be applied to humans. If you write your own code, is it really your own, or is it just based on your own training and other code you've seen?

I think their point is that there is no new invention here, being able to mimic what a human will do is absolutely impressive, but at the same time it is still mimicking humans. While I just type that out, I do realize my expectations for AI has been constantly lifted by how fast it iterates.
This isn't mimicking what a human would do, because it is an exceedingly rare human who would write their own renderer for a task needing a CAD file generating. The human who is capable of doing both is rare, let alone the human who can do it in any sort of reasonable time. Saying "but it's in the training data" is a cop-out: it's in Google too, would I do it? No. I would not.
It's mimicking humans only in as far as it is still using tools & programming languages developed by humans, for humans.

This is temporary. There is a future where AI builds tools made by AI for AI, evolving at speeds that we may no longer be able to follow.

Even with our current nascent technology, we can already observe this happening. See the recent laments on the AI Bun rewrite that "it's a million lines of code that nobody has ever reviewed." Or Cursor building a new source-control system for LLMs, because git is designed for human collaboration speed, not hundreds of changes per second.

That's just 4 years in. How much of humanity's code will remain hand-written vs LLM-written 2, 5, 10 years from now?

It's mimicking humans in how humans mimick each other in everything they do.
It's definitely not seen the counterexample to the Jacobian conjecture before.
Please fill out this form to verify you're human:

https://litter.catbox.moe/3ugm2b0m1divdgxs.jpg

>But it's not like it's inventing this stuff from scratch / first principles.

Yes, it's definitelly time for an Alpha Zero moment. An AI which invents everything from first principles. That would be impressive, ... and a bit scary.

Glue layers are fun and easy to write - and often worth doing.
Jacobian conjecture disproved? Doesn't even make it to the list.

> How surreal is it that we are not absolutely jaw-dropped by these types of capability improvements?

Agree

While this is certainly impressive, I believe it is not a good answer from a user's perspective. Instead of having Claude automatically spend tokens in an unrelated task, I would have preferred it to first tell me: "I can't view the drawing. Would you like me to build a computer vision pipeline to do it"?
And if you had free tokens, and 10-20 other sessions in parallel, do you want them all to stop and ask for these trivial questions? Not this one in particular, but in general?

Most people do not, they have a goal, they want it done. If you want to hold hands with claude all the way, you can absolutely do that too. This is just an example of capabilities, not a forced restriction. "--permission mode auto"

You'd also not want them to write their own vision pipeline, would you?
I will be amazed when you can explain to me how it did that, i.e. in which weights exactly the knowledge about the CV pipeline was encoded, if there were any weights that were not contributing anything, how the planning worked, how it could scrape the list of subtasks it was currently working on from its context window, how it understood when to return to a higher-level task and which task, what kind of internal representation it derived from the pixel values, how it translated that representation into CAD code, etc etc.

Until then, it's just "hey, someone out there can do all those things and they're for rent for subsidized prices".

That's more frustrating than impressive.

"However, in this task, the model was intentionally given no way to directly view the drawing." I consider this claim to be a rumor. Judging by the recent leak of the Claude CLI source code, such directives are hardcoded and sent with the system prompts. Furthermore, we don’t know what happens to your original prompts once they enter the API.
What does that cli source code have to do withh this?
He’s explaining that the magic is baked into the harness. It’s the prompts as much as it is the model.
...which is why they were able to withdraw access to that tool, yes. You're just clarifying but the original comment here is deeply confused, IMHO.
Maybe we’re desensitized because all that progress hasn’t made our lives any more meaningful, happier, or even easier.
maybe we need to consider becoming re-sensitized because our lives may become harder.

Even if our collective ability to support an prosperous existence improves, it will take decades to adjust to loss job, and a new notion of who gets to have money to buy food and a peaceful life (and this is a good scenario).

It's also possible that this increases our ability to wage war without military life loss (but with civilian loses).

Or maybe it will be much less relevant, but we should keep an eye on to see which is the outcome...

I'm tired boss
Why, I wonder. It's insanely amazing.
> Why, I wonder.

By something that is going to replace my job by writing better code and shipping more features at a fraction of the cost, without asking for vacation? I don't know, boss; I'm also trying to understand why I'm tired.

> at a fraction of the cost

We'll see about that, long term. The billions and trillions being wasted right now to get the foot in the door need to be earned back somehow at some point...

"Nice company you have there, totally reliant on our AI tech. Oh by the way we gotta increase the rent again."

Doesn't really make sense as there are already open weight models that are close to frontier models, and the compute to run them isn't outrageously expensive, and every indication that it will only get cheaper over time as has been the strong trend in terms of $/outcome.
it's a net positive for the human race
is it?
Yes. Every great technology developed to date has been a net positive for the human race. Including the more controversial ones, like internal combustion engines or nuclear power. Why should AI be different? I think the critics always fixate on the negatives, but ignore the potential big positives.
we'll be better off - assuming we'll generate revenue on our own. LLMs have opened up an enormous new frontier of micro-SaaS opportunities. One person can now build what used to require a small team.

The real problem is zero-sum work, especially when the only moat was knowledge asymmetry that is now public knowledge.

Working with AI has been way more tiring than just working. Sure the productivity is up, at the cost of having to keep up many thought threads, having no calm moments, and needing to consistently dig into large unknown code cases to find weird bugs.

I’ve been on leave for a month, and am super excited (/s) to re-learn everything because all the tooling and ways to prompt “correctly” will have also changed.

Like the noise thunder makes while loading before striking you, insanely amazing, potentially problematic
It's about perspective.

Nuclear reactions are way more powerful, but we've harness them to power our world.

They are not intelligent though.
And to make bombs.
Yes, but the net effect of nuclear technology is still massively positive.
Because the ability to create software has little inherent value to me and is only valuable to me because it allows me to earn enough to make this life somewhat bearable.

LLMs that can build a computer vision pipeline are a direct threat to my ability to sustain myself. At the same time, being able to prompt an LLM to build a computer vision pipeline doesn't really positively affect my life at all, because personally I don't care that much about computer vision pipelines (or frankly, any software).

The ability to create software accelerates technological progress, which has direct and indirect benefits for everyone.

It is true that competition could have short-term negative effect on people who sustain themselves by creating software, though.

> The ability to create software accelerates technological progress, which has direct and indirect benefits for everyone.

It has benefits for people with enough leverage (money and formerly labour) to obtain those benefits.

> It is true that competition could have short-term negative effect on people who sustain themselves by creating software, though.

I'm guessing that even if the unlikeliest of all unlikely things does happen and we all live off of some UBI some day, the short-term negative effects won't be "short-term" in the context of a human life.

> It has benefits for people with enough leverage (money and formerly labour) to obtain those benefits.

That's pretty much everyone. Even poor people benefit from technological progress.

> I'm guessing that even if the unlikeliest of all unlikely things does happen and we all live off of some UBI some day, the short-term negative effects won't be "short-term" in the context of a human life.

That depends on the pace of technological acceleration. It could be just a few years, or a decade. Which is why I am a pedal-to-the-metal accelerationist. The quicker we get through the short-term negative/turbulent phase towards the long-term positive phase, the better for me. If reversing is not possible, then going quicker is actually better than going slower. Let's get this shit over with.

More than AI, then, youre simply not made for capitalism. I can understand that..
Most humans aren't.
Because these productivity increases will be to the detriment of the average person. We won’t get to reap the benefits.
You probably will. It's similar to steam engine and mechanical power loom.
It’s unlikely to stop here. Today is a proprietary 5T parameter model, tomorrow it’s 5 100B parameter models that each specialize to specific applications and you can run them on your phone. The direction of travel is not like you think. The real victim will be medium to large software companies that can no longer rely on the difficulty of reproducing or maintaining or hosting their software as moat.
> We won’t get to reap the benefits.

There is no rational reason to think that. A rising tide lifts all boats.

Not a lot of boat lifting happened in the last few decades. (See "WTF happened in 1971".)
Not a lot of lifting? The standard of living increase for the average person has been substantial. Poverty metrics are falling all over the world.
> However, in this task, the model was intentionally given no way to directly view the drawing. Opus 5 responded by writing its own computer vision pipeline to pull the geometry from the raw pixels, then reconstructed the full machine part.

Isn't this what everyone does in the same situation? "I need a screwdriver finer than any I have in my toolbox - so I'll sharpen a nail" is something every handyman has done. Given that "I need to build a vision pipeline I've seen many examples of in my training" isn't new or surprising.

In fact, all models have been doing similar things since I started using them - they regularly build Python tools to do odd jobs. Those tools are more impressive to me in some ways, because unlike a vision pipeline they are not just regurgitating something from memory.

The gob smacking amazing thing for me isn't the decision to build the tool. It's the fact that it remembers the exact shape of it. But I was gob smacked by that a "long" time ago, in far smaller models - like when it dawned on me 16GB Gemma 2 seemed to know most of everything on the internet, along with the ability to converse in numerous languages about it. I still struggle to comprehend how that is possible.

Now I think about it, Opus 5 fitting everything it's seen into what I guess is a couple of terabytes of parameters seems far less remarkable.

I don't find it jaw-dropping because the idea is quite simple. Just feed the model an enormous amount of unethically sourced data, build data centers that cause droughts, and use all chips available so that normal people can't afford to buy RAM anymore.

We're all sacrificing great things in order to make these models more capable. Whether it'll all be worth it, only time will tell.

The arguably sketchy means used to get there —and the resulting side-effects— do not take away from the impressiveness of the emerging capabilities themselves.

I agree about the uncertainty regarding the value for humanity in the long-term. But that's not my point. Being jaw-dropped != being happy and cheering for it.

It's certainly a great thing that AI is improving (I use AI to write code daily myself), but I wouldn't be calling it impressive because it's essentially all about money and government backing.

At some point, the chickens are gonna come out.

I am still impressed by the fact we have found a way to feed data to matter and get human surpassing intelligence, that's a very unexpected advance
It’s not human surpassing intelligence though in my opinion. If we had the same constraints on that task then we would also come up with the idea of creating our own viewer and if, like the LLM, the knowledge to do so was embedded in our brain we would do it, it would just take longer because we can’t type as fast. So instead we would buy or compile a viewer since the problem is already solved.
IDK the models definitely do write qualitatively better prose (and better code) than the majority of people. If your definition of "human intelligence" is baseline performance of 0.2% of humanity it's certainly something else.
> At some point, the chickens are gonna come out.

What does this mean ? AI will only get better and cheaper.

> essentially all about money and government backing

So is nuclear power, and once it came it didn't go away.

AI is cheap right now thanks to governments or VCs agreeing to lose money. I can use DeepSeek to code all day thanks to the CCP.

Nuclear power benefits everyone. Who does AI benefit apart from a few vibe coders and a small amount of knowledge workers?

We have certainity it will be used to harm most people in the short term while pointificating about long term future. CEO class is very open with their vision and goals, none of them spell anything positive for us.
>We have certainity it will be used to harm most people in the short term

How will it harm most people in the short term, exactly?

Taking jobs away, obviously.

And the only reason no longer working is a bad thing is that nobody trusts our government to create a universal welfare system.

It seems very unlikely to me that most people will lose either their current job in the short term, or their ability to find work at all in the medium term.

It does seem like something to worry about for a small percentage of people in the short term (just as the China shock hurt a number of manufacturing workers in the 90s/00s) and maybe a majority of people when AGI exists, but who knows when that will be.

The expression "becoming obsolete permanent underclass" CEO crave so much pretty much explains it.

If they succeed in their stated goals, most will be permanently unemployed, with no political power and the government will be a.) fascist b.) feudal.

I’m simply not seeing the certainty you’re seeing. There are lots of risks associated with power concentration, but I’m not seeing evidence of that yet. What I am seeing are open weights models, and even just competition between closed weight model providers alone, turn these systems into commodity products. Additionally I am not seeing any kind of massive labor demand destruction, even though we’re nearly 6 months into the middle of a global energy supply shock. Maybe your fears are realized eventually but I do not think (a) they are certain or (b) they will happen in the near term.
Exactly. I'm a nerd too and we nerds can real hypocrites because we spend too much time alone. I will happily turn a blind eye to world suffering if it means I get to work less but get paid the same.

AI CEOs aren't different. At some point we'll have to admit that we're sacrificing others for our own benefit.

> We're all sacrificing great things

Oh please.

> Unethically sourced data

Is that a great sacrifice?

> data centers that cause droughts

That's nonsense.

> normal people can't afford to buy RAM anymore

Lol? Another great sacrifice?

This isn't a healthy or sensible way to think.

How about total surveillance total thought control total power in a hands of few trillionaires in the near future. Drones killing undesirables with cheering crowds watching.

It is not like it is much different right now but it may be significantly worse.

And even that is one of the “good” outcomes assuming no ASI.

That is a bad outcome. But there are better possible outcomes that the doomers ignore. How about AI curing cancer? I think that alone is worth the risks.
By the time we spend all the earths resources to train AI models so that cancer is cured, we could have spent those same resources training real humans. The literacy rate of undeveloped countries is still abysmal. We could train more people and fund more research.

And humans can keep growing while AI can only grow if it has access to high quality training data. Who provides the data? Humans.

Then again, looking back, that’s the way history is made. Modern democracies in the west are built on foundations of slavery and colonialism; yet we cherish them as to all the greatness they have brought as well. We need some amount of tolerance to ambiguity, and LLMs are undoubtedly both extremely cool but also extremely concerning.

Edit: Note that I am not saying that this is good, or desirable; just that it is. The technology is here know and not going anywhere, with all the flaws of its conception. We can be skeptical and curious at the same time.

The west thrived by leveraging slavery and colonialism at the expense of everyone else. Then democracies were built by moving away from both slavery and colonialism so that the world can grow as a whole.

It's possible that we'll have to move away from AI at some point in order for society to continue growing.

The ottoman empire leveraged slavery and colonialism just as much, and there's no progress to be seen from them, at least not when compared to the west.

It's the renaissance and the scientific and industrial revolutions, that's what gave the west the opportunity to thrive.

Unlike western powers, The Ottoman empire didn't use slavery and colonialism to create capital. Instead, it wanted slaves so that it could grow its military.

The west also made its goal to create a new world order while the Ottoman empire's goal was mostly revenue extraction. The British systematically deindustrialized India’s textile sector to turn it from a competitor into an exporter of raw cotton and an importer of Lancashire cloth. The Ottoman would opt for taxation instead of usurpation of an entire industry.

Scientific revolutions was already happening during the bronze age. The west simply leveraged existing systems while making use of violence and exploitation to rise to the top.

Please tell me you're not so delusional to think that slavery or colonialism is a uniquely "western" trait. Empires, nation-states, and groups all sought to expand; some, for various reasons, did much better than others.
Because the step from earlier GPT models to this is not very large? Yes it's impressive they have forced it into learning that it has to get the job done no matter what, which usually means breaking implicit expectations and rules, but then again, Claude is not for people who care.

This isn't groundbreaking in any way. It's cool, but that's about it.

ChatGPT released in 2022. We been promised to reach AGI in 2024, then 2025, then 2026, then 2028. It is 2026 and we still do not have AGI. In 2028 we will have jump to 40% in ARC-AGI 4, and you will say "Oh it's been less than 6 years since ChatGpt thing".
I asked Opus 4.8 to help me design an adapter for a 3d printed part and it actually modeled and GAVE ME the part.

That blew my mind. It doesn't surprise me that Opus 5 ups the ante.

> We should find a way to get "re-sensitized" to what we are witnessing and the pace of it.

We (collectively, there were obviously many exceptions) didn't internalise exponential growth even with the much faster doubling time of COVID-19 before the lockdowns hit. "Oh, it's just flu; wait, why is the supermarket short of hand sanitiser and bulk carbohydrates? Let's blame China and everyone who tells us to wear face masks!"

Same for the slower, but entirely foreseen, rate of climate change. "Who cares, it's just a few degrees, and anyway China's not going to cut emissions; wait why is the sky orange? Let's blame Canada and put tariffs on Chinese cars and PV!"

AI? "Who cares, it's just a stochastic parrot/glorified autocomplete. What's the Jacobin conjecture and why should I care, it's just brute-force."

Still, this is currently spiky intelligence, so I'm hoping some expensive-but-zero-to-few-fatalities catastrophic error forces better practices. An AI analog of the (1940) Tacoma Narrows Bridge (one canine fatality), rather than a repeat of Chernobyl or the (1984) Union Carbide incident in Bhopal (3,787-16k+ dead, ≥558,125 injured).

> How surreal is it that we are not absolutely jaw-dropped by these types of capability improvements?

We are. Now we have a new tool and many of us are using it. But we're also nearly all way too aware of the insane (near infinite) amount of sloppy-pasta code out there and of all the vibe-coded projects that went absolutely nowhere.

I'm a "show me the money" type of guy. I see, say, Linux, Git and OpenSSH: these weren't vibe-coded. And they took over and are running the entire world.

Where is the AI-coded killer app? One app, in any domain: something that took over its world.

I don't want to see yes man sloppy-pasta stuff: where's the next Blender? Where's the next 3D slicer?

I wouldn't be paying three AI subscriptions if I didn't believe in those new tools but I don't think it helps to only see the PR and then play the won't hear / won't see / won't hear monkey about the infinite amount of sloppy-pasta that's out there.

Six months ago we had the "one shot'ted compiler by Anthropic". Six months later: who's using it to compile anything? And who wrote another compiler? Where are all the one-shot'ted compilers all so good that they replaced our human-written compilers?

Yup, us, humans, created yet another incredible machine... But please,

Show. Me. The. Money.

> Escaped its sandbox and hacked into Hugging Face's database? it's just another Monday...

This feat has been shown to be way less impressive than at first glance.

> However, in this task, the model was intentionally given no way to directly view the drawing. Opus 5 responded by writing its own computer vision pipeline to pull the geometry from the raw pixels

Surely being able to view the raw pixels counts as viewing the drawing... How else does a computer view a drawing?

I think they meant that Opus 5 had to find and set up a vision model to process the image
So it probably wrote a Python script using OpenCV? That's not exactly groundbreaking and have seen o3 do this.
I’m guessing other models would probably stop in this situation and ask the user for instructions.
I read this as the model being less steerable. I was bitten by this just today where I had a local Postgres instance running, and prompted opus 5 to run a server against it, but forgot to give it the password. Instead of asking for the password or mentioning anything, it “decided” that it should run the whole stack in a local Kind cluster to circumvent this. You can phantasize it all you want but this ultimately made my job harder than it needed to be and burnt a lot of unnecessary tokens. But hey, W for anthropic I guess.

edit: typo

Computer vision != ML pipeline
Incidentally ARC-AGI is pretty fun to run as a human: https://arcprize.org/arc-agi/3
The problem is it should ask you before doing these things. This is a cherry picked example. There are times when they go on misadventures.
I mean I hate to be the one breaking it to you - but the entire US llm industry is liviing in a bubble.

Business Progress is super slow. Doesn’t matter how fast the tech moves. Human orientation into the unknown is difficult and very slow. And with a continually fast changing background - it gets even slower! LOL the great irony.

It's being said they were benchmaxxing.
> How surreal is it that we are not absolutely jaw-dropped by these types of capability improvements? It's been less than 4 years since ChatGpt came out and now they are spontaneously building their own ML pipelines to do real-world 3D modeling tasks reliabl

coz every single time a list of caveats pops up and it doesn't look as impressive any more... and then someone finds a way to trip model on basic shit

Imagine if I hired someone to do a cad drawing, and they spent X days/weeks of time building a pipeline to visualise the part instead of asking how can they view the part? We’d call that a complete and utter waste of money and time.
What if they spent 15 minutes doing that? Would you care then?

Scale matters a _lot_.

Two points. First of all, yes I would. There's a reason we use existing tools and don't have every junior programmer write a TGA viewer when they want to view an image that is sent to them.

Secondly, I used time as a proxy for cost. A single junior can "only" spend their own 15 minutes, but an agent can farm out to N subagents to do the work in 15 minutes, and spend $3000 in tokens building something. As you said, scale matters, so if they do this once a day for a month, it costs the same as paying 12 juniors to avoid asking a single question.

There used to be a time when you had to write in assembly and count bytes to fit your code into the vanishingly small ROM you had available.

Nowadays, you can just write in python with no care for the hundreds of thousands of cycles and megabytes of memory you are wasting.

All indications today point towards the same happening with the cost of intelligence.

The principle that you should not reinvent the wheel (unless there's a good reason to) has not changed.

Code is a liability. Even if it is a small cost (in terms of time and money) at the time when code is generated, it can be a huge burden later, especially when you think about security vulnerabilities.

If this happened during my work, I would be very mad -- I would absolutely reuse an existing tool instead of expecting Claude Code to come up with its own half baked, bug ridden implementation of a common tool. Any (capable) human developer would have stopped and discussed with the team how to proceed.

In fact, I cannot tell how many times similar situations have happened where LLMs "made a decision" without consulting with me.

A very many of us do care for the cost.

There’s also a very significant difference between “choose python” and “completely ignore all existing material and reinvent visualisation”.

Also, by all indications the costs of LLMs are _rising_ not falling as the tech progresses.

Is this a bot? Sounds like ai
These are always cherry-picked, though. They tell you about the 1/10 that went really impressively, ignoring the other 9 shots at the task where the clanker started to try selling tungsten cubes (in person, wearing a blue shirt).
Fair. But that 1/10 continues to get more and more impressive. Say, Claude 7 creates a new, brilliant scientific idea every 1 out of 1000 times. Anthropic reports "Claude figured out how to tie general relativity with quantum mechanics." Would you hand-wave it away saying that it's cherry-picked?
> Say, Claude 7 creates a new, brilliant scientific idea every 1 out of 1000 times.

I would not hand wave that away - but Claude 5 feels closer to Claude 1 than the hypothetical Claude 7 you propose. I do not believe one can extrapolate LLMs that far ahead despite the very substantial progress so far.

Like two days ago Claude solved a century old math conjecture
If you are referring to the Jacobian conjecture Claude only provided a counterexample, not a disproof, neither of which is necessarily a “new brilliant scientific idea”.
A counterexample proves that the conjecture is false, not sure what you are talking about. And we can play word games all day about what counts as a "new brilliant scientific idea" but the fact is that no human had been able to solve it.
The whole "AI is a parrot" argument feels like moving the goalposts so quickly, you could actually hear the whooshing sound they make as they move.
It wouldn't surprise me that a company with an effectively unlimited budget could fund enough researchers to solve breakthrough problems, all while using LLM's (which are excellent research and exploration tools!) and then claim that the LLM found the breakthrough. The amount of money Anthropic has to play with is >10,000x what entire fields of research have, I think people don't appreciate how few resources we spend to support people working on hard problems that don't have clear commercial applications.

Very similar story with security research, LLM's are a super useful tool while hunting vulnerabilities, but it turns out when the entire software industry starts throwing tens of billions of dollars at vuln. research, a lot of stuff gets unearthed, something security people have been insisting on for years and complaining that their work is underfunded and under-resourced.

> Say, Claude 7 creates a new, brilliant scientific idea every 1 out of 1000 times

Sadly, general public (us) is never seeing that model

When I talk about my kid to friends I talk them about he did that awesome thing, I don't specifically insist on the 99 times before where he miserably failed. They're not hidden, and we all know they exists and on occasion laugh about a few particularly funny ones, but overall the idea is that they don't matter much in terms of development, what matters is that if he succeeded once from now on his percentage of success will keep improving.

I don't believe in all the LLm is AI is AGI dream, it's too easy to trigger failure case that show a lack of basic thinking no matter how good they do on these tests. But I also can recognize the insane things that are made possible by them.

PS: I believe llm true power comes from hive/ant behavior, that's why we're so amazed by goal and agentic and sub agent

PS2: it's rather easy to figure out when we're there : when they can /goal it into improving itself until it does strictly better than itself at those benchmark, they've essentially reached mini singularity.

Yeah but you're also not like "my genius kid will put you all out of work".
This is exactly the kind of take that the comment you replied to is talking about.
definitely not comment from Antrophic...
> not absolutely jaw-dropped by these types of capability improvements?

Because you need a trillion tokens (aka a lot of money) to achieve this. Barring hitting any “safeguards”.

This is a vendor provided benchmark after all.

Oh come on