Hacker News new | ask | show | jobs
by theplumber 14 days ago
The private AI companies should be forced to release the models as open weights with a license (I.e no commercial use) due the high risk they present and the data they basically steal from everyone to train their models. This should be the safety push not the regulatory capture that Dario is trying.
7 comments

> the data they basically steal from everyone

What's the precedent?

All I see if a bunch of people who haven't loaded an ad in 20 years and have a 5TB collection of pirated movies and music suddenly decrying that LLM's doing next token prediction over a dataset is theft.

You don't get to change your mind 20 years into "I'm never going to pay for anything binary" ethos (cough look at what this post is cough) that has dominated the internet for decades. If people are genuinely upset about LLMs training on all available data without compensation, all I can say is "Reap what you sow".

It would be all fine and good if they kept it non-commercial, but when you start charging money for it that's when you're stepping out of bounds. If they want to benefit from fair use, they should contribute their results back to the public domain. Getting it for free and then commercializing it is what's unethical.
>Getting it for free and then commercializing it is what's unethical.

Like all those unethical software devs taking salaries for turning stackoverflow into products? I suppose the blood still flows since now they just use LLM output?

Any way you try and slice LLM morality, you end up with "It's bad because they are not me" reasons. "When I monetize information I get for free, it's good, when they monetize information they get for free, it's bad"

Ok then can the feds bust the executives of these companies like they did with Kim Dotcom? It’s not like Dario trained his AI in the bedroom or do a small business by letting people scrapping IP work to train their models. He is actively stealing and selling the stolen work and even preaches us safety measures: how to protect us from our own knowledge that he stole!
I am not convinced by your retort, which essentially boils down to “some people pirate, so its ironic some people, maybe different, are mad at dario!”

  "Reap what you sow"
It is kind of incredible that you're not focused at all on the copyright holders, instead focusing on random tech people you had online disagreements with.

Artists and writers got screwed first by piracy, then by generative AI. They didn't sow anything. They just got reaped.

And the only thing the copyright hypocrites are "reaping" is a feeling of hypocrisy. Congrats for pointing that out. Your comment is simply myopic.

> Artists and writers got screwed first by piracy, then by generative AI.

That's too glib. Artists don't have a right to money that people won't spend.

It's not all down to bad actors. Although it's true that musicians got screwed when their distributors switched to a subscription model. Writers got screwed by the Amazon monopsony that crammed down publisher's margins.

Mostly what happened is that media technologies changed, professional creators had less control over production, and the money flowed upstream leaving them with a smaller niche.

All of their output used to be gatekept by media companies who had a lock on publishing. The internet reduced distribution costs to zero. Suddenly anybody could reach an audience.

Now generative AI has dropped the cost of content creation to the basement. It's so much easier to write a blog post or make an illustration.

That isn't cannibalizing content. It's a new way of making. It takes a lot of the value out of creating those kinds of things. The customer experiences this as a reduction in cost.

These changes happen with every new gadget. Photographers have to compete against everybody with a phone in their pocket. Now they just do weddings.

My grandfather was a portrait painter in the 1940s, a profession that was already moribund from the proliferation of photography studios. He gave up and became an insurance salesman.

Maybe he got screwed. The photography studios are all gone too. Because anybody can make a portrait using the phone in their pocket. At zero cost.

> the data they basically steal from everyone

Agreed. Especially since now competitors have more difficulties getting the same advantage. They don't have to do so immediately and perhaps not their specific tuning. But the weights of the raw training data at least should be publicised.

> no commercial use

Why? It's not like they did the hard work. It's disgraceful that this kind of commons enclosure has been allowed in the first place.

Wow, thank you for this! I'd forgotten about enclosure for some years.

> The law locks up the man or woman / Who steals the goose from off the common / But leaves the greater villain loose / Who steals the common from the goose.

> release the models as open weights with a license (I.e no commercial use) due the high risk they present

For some kinds of risk (ex: walking people through on how to make infectious bioweapons) an open-weights approach would increase risk.

Don't buy the idea that the only thing that's standing in the way of people making bioweapons is the censoring of LLMs.

If you are motivated enough to assemble all the kit you'd need, and actually do it, then you should be motivated enough to find the knowledge to do it without chatgpt etc.

I'd imagine the main thing that's stopping people is biological weapons are a terrible tool to do what most people want to do - which is target specific enemies.

So while the materials you'd need to build such a thing are much more accessible than say nuclear material, it's much less attractive - but if already had somebody mad enough to try - it's already possible.

> Don't buy the idea that the only thing that's standing in the way of people making bioweapons is the censoring of LLMs.

It's not a matter of a one thing standing in the way of people making bioweapons: there's a long chain of actions one would need to complete, and many places for the chain to fail. Access to expertise can reduce the chance of failure at many of these steps, and AIs can increasingly substitute for human expertise: https://securebio.org/benchmarks/

> If you are motivated enough to assemble all the kit you'd need, and actually do it, then you should be motivated enough to find the knowledge to do it without chatgpt etc.

If someone was motivated enough to assemble the kit you need for anthrax and actually distribute it then you might expect them to also be motivated enough to find the knowledge to identify an appropriate strain, but in fact this is where Aum Shinrikyo failed in the 1990s: https://en.wikipedia.org/wiki/Aum_Shinrikyo#Incidents_before...

Lack of knowledge is one of many factors that can lead to failure, and LLMs make it less of a barrier.

> biological weapons are a terrible tool to do what most people want to do - which is target specific enemies

Except:

1. There are also people out there who want to kill everyone. They're sufficiently rare that no one with the motive has also had the means, but as technological progress keeps lowering the bar the risk of motive and means intersecting increases.

2. This didn't stop the Soviets. They did an enormous amount of very dangerous research with minimal logical application.

>This didn't stop the Soviets. They did an enormous amount of very dangerous research with minimal logical application.

Ultimately what stopped them ( and everybody else - let's face it it wasn't just the soviets ), is it's a bad idea.

I think you way over estimate how hard the knowledge bit is in a biological weapon ( if you just want to kill indiscriminately - specific/controlled targeting a whole different ball game ).

The barrier is motivation and kit ( though the kit is much easier to access that say the stuff you need to build a nuke ).

The idea that all that's stopping some disaffected Joe sitting on his sofa at home and suddenly deciding to make a biological weapon is he can't get a set of instructions from ChatGPT is to miss the point.

Ultimatelty it all comes down to are you motivated enough - if you are then the knowledge is out there with or without an AI summary.

> I think you way over estimate how hard the knowledge bit is in a biological weapon ... if you just want to kill indiscriminately

That's a surprising claim: usually when I make this argument skeptics say that the knowledge barrier is so high that an LLM won't help enough!

I'd also be curious to hear what you think of the Aum Shinrikyo case, since that seems to me to be straightforwardly a failure of knowledge.

> The idea that all that's stopping some disaffected Joe sitting on his sofa at home and suddenly deciding to make a biological weapon is he can't get a set of instructions from ChatGPT is to miss the point.

I agree motivation is a huge barrier, in the sense that almost no one would do it. But to keep catastrophic bio attacks from happening as the knowledge barrier decreases it's not enough that most people wouldn't have the motivation. If 0.0001% of people have the motivation and 0.0001% have the means then we're probably ok, but if 0.0001% of people have the motivation and 0.1% have the means then we're not (0.0001% * 0.1% * population > 1).

In an earlier phase of AI I might have agreed but now the publicly available data they’ve trained on is increasingly useless for the frontier.

RL/post-training is now a much bigger part of it, with that being largely proprietary and expensive work that the open source model just won’t fund.

Good idea! I'd never heard that. Very interesting.

But wouldn't this help china build models just as good as ours immediately? Wouldn't it make the investment in training a model worth a lot less?

I'd rather they published the training data.

Then we can verify that there's nothing nasty hiding in it.