Hacker News new | ask | show | jobs
by HyperL0gi 4 days ago
Isn’t it just hilarious that a model that seemed so superior to Fable but didn't get doomsay marketing from Anthropic got released without any issues? In theory, this was supposed to be AGI level according to Anthropic, yet here we are, just a normal Friday.
4 comments

Go read the safeguards section in the report and you will realize why that is.

These models are heavily as safeguarded and that was the initial reason why they said they couldn't and haven't released Mythos because that model is the one without the safeguards.

OpenAI is did the same thing when they announced a model without safeguards broken into HuggingFace servers.

Yes, this makes a lot of sense, but it’s just very amusing to see. 2 months ago, the world was about to end, now not so much.
7+ years ago GPT2 couldn’t be released because it was deemed too dangerous[0]. It was, of course, eventually released.

0: https://openai.com/index/better-language-models/

> We can also imagine the application of these models for malicious purposes , including the following (or other applications we can’t yet anticipate):

  *   Generate misleading news articles

  *   Impersonate others online

  *   Automate the production of abusive or faked content to post on social media

  *   Automate the production of spam/phishing content
Seems like the prediction was pretty accurate.
Dario Amodei is one of the authors of the paper. He used the same marketing tactic with Anthropic, and got flagged by US Gov.
It’s seven years already. Crazy.
Yes, I think so too.
I realized it a why back these labs are selling hype.

since then I have never cared about models except those that affect money in my pocket e.g AWS Nova Sonic

Have you been patching your systems for the past two months? It was crazy even if you completely forget the supply chain literal FUBARs and you must’ve been living under a rock to not see OpenAI (accidentally) pwning hugging face
> now not so much

OpenAI Huggingface breach begs to differ

Do you have an example of the "doomsday marketing" you're referring to?
I see a pretty big gap between finding software vulnerabilities and “the world is about to end”. It is literally true that AI models are finding software vulnerabilities. It is also to my mind a reasonable thing that you’d want to be cautious about rolling out a model that can find more vulnerabilities. So what is the objection you have to these sources?
Absolute masterpiece of a rebuttal. No notes.
feels almost like anthropic is desperate for ipo huh

i think we'll see one of the fastest deflations in history post anthropic/oai ipo

could you elaborate please?
both companies have large spending commitments while their margins will be continuously wiped by open models / router-like solutions. both companies are rushing for ipo and doing an awful lot of financial engineering to look good (highly recommend ed zitrons analysis on anthropic profitability) before their margins are eroded that much.
I feel like i've seen less hype about "the next model will be agi". GPT-6 is supposed to be coming this summer, and nobody is expecting AGI now. Not sure how they're going to keep the hype cycle going
Or another way to see it is that current models are AGI as it was defined before, and the goal post is being moved.
they are definitely not agi as it was ever defined. they’re only a bit more capable than they were a year ago. they crossed over from interesting crap to useful tool recently but really only for software
You have short memory if you think we've not blown past at least 5 different AGI goalposts. They're being moved every time and we're hitting them every time.

Or maybe you just don't know exactly how capable these models are. Most people's experience of AI is a stupid chatbot, it's no wonder they don't understand how these things are coming for their jobs.

On my end, I have a software that is designed and built by Claude, that I did a strategy session on (with claude), and prepared a fundraise for (with claude). My only role, other than "knowing what to aim for", has been to feed the AI some fairly basic english prompts for a few weeks... which is also easily automatable.

Everyone's job is fucked. Devs, CEOs, everyone.

We’re had the ability to make coffee with a machine for decades yet you still pay $5 for a barista made Java, I think we’ll be okay.
Which barista? The one that uses the machine to do 90% of the work, or the one that isn't a barista and uses the machine to do 100% of the work?

How do you know I don't have said machine at home?

And what bug bit you to make you think this is a good comparison anyway?

The simple fact is that economy can’t be 100% services. Who is going to pay these baristas if nobody else earns anything?
> Everyone's job is fucked. Devs, CEOs, everyone.

It’s curious to me that there are two distinct factions here. People like parent commenter who has no discernment and others who see llms for what they are. I just talked to opus 5 and in it’s first response caught some well disguise BS. These things are bullshit machines. There are indeed a lot of bullshit jobs around so maybe parent does discern something I don’t?

If you're saying I don't see LLMs for what they are, you are wrong. I have worked in the AI sector for over a decade and I know exactly how the sausage is made. As of today, I run an AI lab (ingram.tech), and we see daily not just the theoretical of what's possible, but how these systems get deployed and who's really at risk. We're head to head with the reality of the terrain.

But it's completely irrelevant. The emergent properties of LLMs, what was built on top of those emergent properties, and the emergent properties of that, are all together building a world nobody is ready for.

If you don't think this, you haven't seen what these things are truly capable of yet. Either that, or you have a romanticized view of what humans actually do in 99% of non-manual jobs.

I'm blown away by how so many people on HN are just... idk, "blind" is the politically correct way to say it, I think. With zero ability to understand the transitive aspects of what they are looking at. For example, these HN threads are so often polluted with comments claiming some random use case cannot possibly be automated.

Sometimes I feel like I'm showing somebody how a spreadsheet can calculate 1+1, and they ask "Yes, but can it do 1+2?".

...chartered accountants whose job is to deal with the 100s of pages of jargon for you. LLMs are the like smartphones. They do everything so you don't need your ipod, flashlight, alarm clock, game console, computer, that handheld clicking counter thing, map, camera anymore. Llms will do that to a lot of professions.
LLM labs dumbed down the definition of AGI as much as possible, yet their models haven't reached it still. We are nowhere near the original definition of AGI. Not even 1% of the way there.
Oh yeah? Who even talks about the Turing Test anymore? Half a decade ago,that was the informal benchmark.
Not at all.

The Turing test was never about AGI, just about being able to discern a chatbot from a human in a casual conversation...

I would also say that funnily enough it's extremely easy to detect if you're talking to a human or an LLM after a few messages.

in 2023 i wouldve said gpt could pass the turing test. today i could figure out it was an llm in a few turns no problem. llms cannot pass the turing test now that we’re accustomed to them
The only reason you can figure it out is because of the system prompt which is designed to make model “useful”, safe and compliant.

Raw modern LLM with different pre-prompt will easily fool anyone.

We wouldn't be having any talk about AI Slop if it was impossible to tell AI apart from people.
Uhuh, but if it was so easy, we also wouldn't be funding billions of euros and dollars in anti-AI-disinformation systems, holding entire conferences about deepfakes and hybrid attacks on civilians, and have dozens of governments opening entirely new defense departments to study and counter the capabilities of AI to disseminate human-like disinformation at scale (which has been affecting elections across the world).

But yeah, your AI slop take about the local burger joint that used free chatgpt to generate a menu filled with typos and bad images is A+.

Yes 18 months ago it seemed like AGI was being promised every other week, and now I don't see any of those headlines.
It's already happened but no one wants to admit it
We can start having this conversation when it's able to do at least 20% of the work that I have to do.
Fable established the frontier, this is just catching up.
So unless doomsday actually happens then you're unhappy with the warning - is that right? You see false promises of apocalypse as marketing?
It’s either advertising, or they’re idiots, because the apocalypse keeps not happening. Either way, it’s not worth listening to them.
Exxon: "The exceptionally explosive refinery beside your house has not exploded because of our safety culture and protection protocols"

Emp: "what a bunch of lies, I bet they don't even do anything over there"

My point is why the sudden change in tone? I’m not dismissing the models’ capabilities.
As they explicitly say, Opus 5 is ~ equally capable as Mythos/Fable at finding vulnerabilities, but it is much less capable at exploiting those vulnerabilities on it's own. That is an extremely meaningful difference and to me completely explains the difference in tone, release style etc.