Hacker News new | ask | show | jobs
by linkregister 6 days ago
It matters because the closed-source frontier labs spend lots of money on human data (RLHF / RLAIF with human oversight). Moonshot is accused of circumventing these costs. Frontier labs add research costs into their inference pricing. If the market doesn't permit them to sustain sufficient pricing to have a positive cash flow, then their business prospects become weaker and they risk insolvency. Furthermore, other leveraged companies are at risk.

The reason why the United States government is weighing in is because it's in the national interest of the US to have supremacy in "AI".

Legality or lack thereof is one of many data points about whether a thing is noteworthy.

Moonshot performing distillation is rational from their point of view. Reducing costs is in the interest of businesses. It's also rational for frontier labs and the US government to add obstacles to this process.

As consumers this is probably a positive development.

4 comments

My parents put in countless hours and tens of thousands of dollars into raising me to the point where I could write an answer on StackOverflow

And OpenAI scraped and distilled that answer and gave me nothing

What does your story have to do with Moonshot AI? Do you think they didn't also use the same corpus? Bizarre
And now people such as myself have access to open weight models with that information. I wasn't lucky enough to have parents put me through school, and LLMs have absolutely helped me further educate myself and play "catch up" on opportunities others have been given. So, the net effect has been (and is continuing to be) a democratization of information.
You could have gained that stuff prior to LLMs. The leg up you're describing is free information on the internet, not AI. AI just makes it a little easier to find, while also crushing the original sources in the process. (Even if it had a broken culture, is stack overflow even going to exist in a year? Where are they going to train on going forward?)
Democratization of information, but Sam Altman gets a $100B net worth and I'm still broke :)

I would prefer some sort of democratiziation of the money made from the democratization of information as well

Agreed, i wish that it would have done more than change who gets rich off rent-seeking behavior surrounding the knowledge that others created, instead it just consolidated that from many gatekeepers to a few.

I can at least take some measure of pleasure in the fact that it has generally lessened the roadblocks in gathering information. I am still displeased that there are any gatekeepers of humanity's combined knowledge

It’s even worse than before

Businesses like these used to public at reasonable valuations. You could ride with them to trillion dollar valuations and grow your own fortune too. Everyone has a story of buying Apple or Google or Amazon stock and making millions

Now they’re going live at trillion dollar valuations and by the time you get in, all the upside has already gone (see Spacex IPO)

Not only did they steal all human data, they also made sure that the upside was only limited to themselves and their cronies

Not to be an arse, but didn't you have access to Stack Overflow with all questions/answers prior to LLMs?
Yes, but time is finite
Isn’t that the same argument they are making for replacing human labour?

Circumventing costs.

There are many frames that one can place upon this issue. They do not contradict the other. There are moral framings (stole the internet so go eff yourselves, is one), but so is national security, and so is the doomer recursive self improvement risk, and then there is the framing purely on what this implies for future AI training.

I mainly focus on the last.

It will be hard for a frontier lab to justify spending the compute and data curation needed to advance AI further if that expenditure can be assimilated into your competitor's products within months/weeks. So reality will present labs with three choices:

A. Cease spending massive amounts of money and compute improving those models.

B. make those improved models more difficult to distill from, either through some regulatory regime, or some technical solution, which seems unlikely to me.

C. making the best models available only to select partners and government.

In all these potential outcomes, China, which lacks compute that U.S. labs enjoy, will likely stop seeing massive improvements in their AI models. Improvements to be sure, but right now they are enjoying gains from distillation AND their own model innovations, and these potential outcomes would largely stop one of those sources.

Do you get mad at your computer for replacing clerical workers? What does this nonsense comment have to do with the issue at hand?
Those were told "find another job" and in fact they were able to find different jobs.

AI companies are gleefully bragging and "making humans obsolete", "permanent underclass" and 40% unemployment rates they plan to create.

They pushed to replace people years BEFORE their technology even can produce that work.

So, you know, it is not the same. But also in fact, clerks did disliked when occasionally arrogant claimed to replace them while pushing unfinished software that dont quite work yet.

Oh, so mass theft is okay as long as American companies are doing it
copyright infringement is not theft, even if right holders often claim it is.

part of the definition of theft is that the original owner is deprived of it, which does not apply to copyright infringement.

You can only argue with damages from the perspective of potential profits, still not theft though.

https://en.wikipedia.org/wiki/Theft

So having tons of AIs quoting various literary works and reproducing knock-offs of them has a positive effect on those books' sales?

I think you're wrong: there is absolutely damage to the authors and publishers from what the AI companies have done.

? I literally said that, how am I wrong?

> You can only argue with damages from the perspective of potential profits, still not theft though.

Damages are not deprival of ownership. They're conceptually related but orthogonal

Also there was no moral judgement from my end, I just pointed out that an incorrect word is being applied. It's just not theft - by definition. But language is a fluid concept and definitions change over time. As people keep misusing it, it will eventually lose its original meaning. Which may have already happened for you, but this change hasn't been settled yet as can be seen from looking at the official definitions of the term, which as of today still mention the criteria

If you steal an unpopular product from a store, the damage is also only to "potential profits", so how does that differ? It's entirely possible no one would have purchased the product and it would have eventually been discarded/destroyed.

Or with services, if a barber cuts your hair and then you run away without paying them, do you not consider that theft, even though there's no change in ownership occurring?

This is almost funny to me, because in many jurisdictions, software companies sure invested a lot of effort into painting people copying software as thieves. In Germany, they (the software producer lobby, and later politicians influence by the former) even coined and spread the term "Raubkopie", which you could roughly translate as "robbed copy", i.e., that's one step worse than "theft", as a robbery in Germany legally means " theft accomplished by force or intimidation". So, yeah: like putting a knife to the throat of someone while you copy the software.

So, after literally decades of investing into advertising campaigns, lobbying to politicians to pass harsher and harsher laws against software "thieves and robbers", now that big tech are doing it, suddenly we are supposed to consider it with more nuance?

Ahhh... no thank you sir. I really enjoy them drinking their own kool-aid.

Moreover, reading a copyrighted book and learning from it is not theft.
Generating a set of weights is not learning.
would you say that airplanes don't fly because they don't flap their wings? it's possible to achieve the same things with different approaches.
I would say airplanes fly, but I wouldn't say that submarines swim. Things have a bit more nuance, and the field of "learning" isn't as well understood as the ML proponents claim it is.
That is a strong statement. I guess you are telling Machine Learning to go fuck itself.
No, Machine Learning is an unfortunate name for a well documented process for creating black-box classifiers. The process is good, the name is not.
And what, to your mind, would classify something as learning? I assume that your position is not the hard "only humans/living creatures can learn"
Great! Neither is distillation then.
Never said it was. Still, understanding to what extent the ability of Chinese labs to keep up to western models with much less compute needs to be understood.
Machines are not humans.
Reread my comment and look for a value judgement on my part. The final sentence is probably a good clue as to my opinion.
Chatgpt routinely cites and uses papers I don't have access to because they're behind a paywall. I don't think OpenAI is paying for all that copyright. That's in my opinion way more serious.
Yes the fact that the scientific literature - created largely on the back of the tax payer - isn't open to all free of charge by force of law is a travesty. A cartel should not get to charge for access to the bulk of human knowledge. That is indeed a far more important issue than whether or not Moonshot violated the Anthropic ToS, possibly committing mass fraud in the course of doing so.

I mean honestly if they did that why should I care? I'm happy to see copyright violated in a manner that leads to the creation of new technology. IP law exists strictly for the benefit of society and by all appearances AI is an incredibly powerful tool.

Also while I'm at it libgen is a gift to humanity. Information wants to be free. Spreading and preserving knowledge is generally one of the most wholesome activities anyone can undertake as far as I'm concerned.

Your statement is orthogonal to my comment. Why reiterate the schadenfreude / fairness comment already stated several dozen times in this thread?
Who do you think is paying $100K+ for "Enterprise" access to Anna's Archive?