Hacker News new | ask | show | jobs
by drillsteps5 34 days ago
I'm looking forward to the trial where Anthropic will have to disclose sources of their training data, and then explain why they are entitled to charging customers for using regurgitated training data but Alibaba which trains their models on Anthropic's models are not.

Should be fun.

Edit: clarification

4 comments

Quite amusing that the library of libgen is worth 1.5bil for unlimited access.

It's about the same valuation as bun, lol.

$3,000 per title.
Do you think many authors would give you rights to create derivative works en masse for that money?
For endlessly reselling the whole work verbatim? Well, where can I buy such a license in the real world, because then I would like to buy a couple of those!
That's a tiny drop in the bucket of the value these AI companies have appropriated from society.

Just to give an idea of the scale of it:

Let's say a modern SOTA LLM has 1T params and is therefore trained on 100T tokens

1000 tokens of text = 750 words of prose, which may take 15 min to 3hr to write (Gemini's estimate)

1000 tokens of code = 50-70 lines of code, which may take 15min to 5hr to write

We just want a rough estimate of the value of this, so let's say that 1000 tokens took 1hr of human labor to generate at an average wage of $50/hr

So, if 1000 tokens cost $50 of human labor, then that 100T of training data cost $5T.

So, the value of what the AI companies took from society might better be estimated in the trillions of dollars, not billions.

And of course what they are doing with all this data is building generative AI, so it's not just the value of what they took, but more importantly the future opportunity they are stealing from everyone by replacing human labor with their automaton who's profits they intend to keep for themselves.

That's only a fraction of the training data.
Interesting. Looks like the judge ruled using legally obtained knowledge (books, articles, etc) to train AI constitutes "fair use".

Given that US legal system is precedent-base that... changes things.

Meta/Facebook got away with it though right?
That's a great cost-benefit ratio. Can you and I steal and do illegal things and pay the same cost?
Sure, but only if you get the same benefits
looks like we can't today. Man it would be great to figure out how to be above the law just like how these other rich people in different social classes are.
Being logically consistent isn’t as profitable as being aggressive and loud.
While I love the sentiment, I feel like the odds of this actually ever reaching a trial are low, given the international positioning of the parties, and the... um... complex relationships involved.

Anthropic's actions seem performative. Others have already speculated on the likely audience(s).

> While I love the sentiment, I feel like the odds of this actually ever reaching a trial are low ...

As cited in a peer comment here[0]:

  In June 2025, Judge William Alsup of the U.S. District 
  Court for the Northern District of California ruled on 
  summary judgment that using books without permission to 
  train AI was fair use if they were acquired legally, but he 
  denied Anthropic’s request for summary judgment related to 
  piracy—finding that the piracy was not fair use.[1]
Of note in the judge's finding; "the piracy was not fair use".

0 - https://news.ycombinator.com/item?id=48667411

1 - https://authorsguild.org/advocacy/artificial-intelligence/wh...

And if it includes at least one GPL source, they should release the weights on GPL license.