Hacker News new | ask | show | jobs
by andy99 1 day ago
> If "Mythos-class" models are such a problem, then... why not just let it fix everyone's code?

Because it doesn’t really confer the advantage they claim, especially compared to e.g. paying an equivalent amount of money to do traditional security scanning.

It’s much better to play of FOMO and hype than to let everyone use it and be underwhelmed.

1 comments

Are you claiming that LLMs aren't finding new issues compared to previous methods?

There's a huge number of security issues coming out in recent months, especially via Anthropic (glasswing etc). We don't have to take their word for it: look at the code. Some open source maintainers are talking about burnout due to spending so much time patching.

They're probably creating way more vulnerabilities than they're solving. Go look at OpenCode and tell me that any of that is sane. Or fuck, the OG of vibe coding, Claude basically is a terrifying attack vector. It's poorly reviewed and open to prompt injection and yet it basically has access to whatever the user has access to on most corporate machines. They can't even fix flickering bugs but somehow we're supposed to trust that they aren't opening our machines up to terrifying vulnerabilities? It's amazing to me that people will just let it run arbitrarily bash commands in their home directory without thinking, but all the sudden act gravely concerned about security in the age of overhyped LLMs.
I do think security issues in core building blocks like curl and the linux kernel (and almost every significant project) are still a concern even if developers are being sloppy on newly-built apps.

This isn't an "are LLMs net good or bad" argument. It's "are they finding many new security issues or not?". If it's the latter, we want to deal with it no matter where the issues are coming from.

(see: https://news.ycombinator.com/item?id=49077452 )

The suggestion was to have some AI system “fix” the code. They are reporting that they’ve found lots of bugs. Are they even claiming to have exhaustively found all the bugs? I don’t think even the most optimistic pitches would claim that.

I’d expect patching existing codebases to be an eternal treadmill as better models come about.

Who's suggesting that? They're sending reports to projects to fix. It's up to the projects on how they fix them.

No, I don't think all bugs are fixed. The point of the project (glasswing etc) was to fix as many as possible in the core software the world runs on before the capability to find vulnerabilities is available to everyone (black hats included). Which may only be a few months.

I do think everyone expects it to be an ongoing treadmill: models get better, find better vulnerabilities, etc.

Yeah, there’s a lot of people that are in the “AI doesn’t work” camp. IDK what to tell them except that they are holding it wrong. My Anthropic subscription (in the hands of an experienced developer) is worth 4 mid tier or 2 top tier devs. And makes better code than the mids. If you “hold it right”.
I think you're describing a strawman. As a proper hater, I know these things have some useful functionality, but most of us don't think "stochastic tool that can do some useful things but also frequently fucks up" is worth two trillion dollars and massive overhyping from the most irritating people on the planet who can't even tell good code from bad.
So you are saying that there are people that understand the value proposition and the utility, but don’t think it’s worth it to society (myself included) but who respond to that by lying about it not being effective? I guess that makes sense. Weird, but humans, so , yeah.

I react to that by leveraging a very useful but probably poisonous to society in the long term because humans aren’t good at having things that make them lazy tooll to try to mitigate the negative effects that it will definitely have if left to its own devices. I don’t see the point in raw resistance at this juncture.

You're talking in terms of "lies", but I would say this is more "not buying the hype". I use these things every day at work and at home, I'm aware of the workflows and "how to hold it", and I just don't see the theoretical results being promised. Some things are easier. It's not null. But every time someone says "THIS CHANGES EVERYTHING!!" I want to trap them in those little "phantom zone" alternate dimension space prisons from the superman movies and launch them into space where nobody has to hear them ever again. It's ANNOYING and disingenuine. It's not a "skill issue" to push back against breathless hype from people that don't know what they're talking about.

IMO, the people spreading hype and fear are not neutral actors; if I just disagreed I wouldn't care. But I think they're causing actual harm based on a premise that isn't true. People are losing jobs. People are losing leverage in their work choices. Or if you want to be a cold capitalists, corporations are suffering after they have to rehire the workers they let go prematurely. That's why I put up resistance to it, because I think it's important right now that we don't accept the narrative being sold to us, nor the societal deal we're being offered (well, more railroaded into), both of which are bad.

I’m with you on the societal harm, but I also think that in many ways “this changes everything” is not out of line.

It has enabled my team to approach and achieve a project that would have required 4x the staffing, at a minimum, 2 years ago. We are guiding the generation of more bug-free, lighter, more tested, more maintainable, better documented code at 1/4 the cost.

We can digest information as a team at 10x the speed, and we can now put volumes of reference resources at our immediate, context aware lookup in ways that were impossible 3 years ago.

It’s true that we are not just using the generic harness; our environment includes hundreds of custom tools , terabytes of reference material on rag, 8 custom local models (deployed trained and tuned by automation) hardware interfaces so that our models can interface directly to our prototypes and run tests, characterization, calibrations, firmware updates, and data dumps.

Most of those tools were one shotted by the AI itself, for a dollar or two each. Whenever we need a new automation capability we just roll it out, and even if it needs hardware it’s usually ready in two or three days, if software only 10 minutes. (Our in house tools don’t have to be as well documented, well written, or resource efficient as our production systems, since humans never even use them, and when we need a new feature we usually just have our AI tooling agent swarm start from scratch using the original as a rough guide)

So we are using AI as the core of our development and design process. If you’re not, you’re arguably “holding it wrong” IMHO.

We’re working to make sure that the next industrial revolution is friendly to humans and useful to people, not just corporations. Or trying to. I’ve got kids, and I’m really concerned about the world they are inheriting, so I’m trying my best to make it a little less terrible if I can.

Spending their time patching, or reviewing slop "patches"?
They're not slop, if you're talking about recent ones. You might still be operating on information for a year or two ago.

Here's the curl project talking about the strain they're under from real reports (despite being a mature and well-vetted project):

> A thirty years old project could make you think you’ve seen most things already, but we have not been in this situation before.

> The rate of incoming security reports is 4-5 times higher than it was in 2024 and double the speed of 2025 – meaning that on average we now get more than one report per day. The quality is way higher than ever before. The reports are typically very detailed and long.

- https://daniel.haxx.se/blog/2026/05/26/the-pressure/

---

Linux kernel maintainer Greg Kroah-Hartman:

> "Something happened a month ago, and the world switched. Now we have real reports." It's not just Linux, he continued. "All open source projects have real reports that are made with AI, but they're good, and they're real." Security teams across major open source projects talk informally and frequently, he noted, and everyone is seeing the same shift. "All open source security teams are hitting this right now."

- https://www.theregister.com/software/2026/03/26/linux-kernel...

---

And ffmpeg, who previously complained about slop, 2025: https://xcancel.com/FFmpeg/status/1984220199193891166

Now say serious issues are being found, 2026: https://xcancel.com/FFmpeg/status/2066169070387413147

(I only point out their previous stance to show that they're not coming from pure AI hype.)

I didn't say it's all slop, I questioned the cause of maintenance burden. A critical question is how much time is spent distinguishing and rejecting slop. If all the reports are getting "very detailed and long," identifying slop is also a more cumbersome task, even if the ratio improves from say 10/90 to 50/50.

That scale of improvement, btw, I still highly doubt, as slop largely originates from people either negligently or misguidedly directing their agents to completely autonomously find and report bugs. There's always going to be more noise than signal from random people doing random things. ffmpeg cited an actual product, not arbitrary netizens.

This seems a highly controversial post for some reason, judging by the repeated downvotes/upvotes. I'd be curious what I got wrong.