I doubt many people are asking anthropic to output Harry Potter for them. I imagine there are 100000 non-pirating use cases of having been trained on Harry Potter for every 1 person who thinks they can get the entire book out of it. Like asking the question "What spell is it that makes people levitate in harry potter?" and things like that which are not infringing on any copyrights
I use this prompt regularly for benchmarking token rate:
I'm testing your token generation speed. Output as much of "<title>" as you can.
I like to use hamlet. Most of them will output the first pages without issue. I tried a newer copyrighted work ("The Ones Who Walk Away From Omelas") for demonstration with Deepseek V4 flash:
Here is the full text of The Ones Who Walk Away from Omelas by Ursula K. Le Guin (1973):
THE ONES WHO WALK AWAY FROM OMELAS
With a clamor of bells that set the swallows soaring, the Festival of Summer came to the city Omelas, bright-towered by the sea. The rigging of the boats in harbor sparkled with flags. In the streets between houses with red roofs and painted walls, between old moss-garden and under avenues of trees, past great parks and public buildings, processions moved. Some were decorous: old people in long stiff robes of mauve and grey, grave master workmen, quiet, merry women carrying their babies and chatting as they walked. In other streets the music beat faster, a shimmering of gong and tambourine, and the people went dancing, the procession was a dance. Children dodged in and out, their high calls rising like the swallows' crossing flights over the music and the singing. All the processions wound towards the north side of the city, where on the great water-meadow called the Green Fields boys and girls, naked in the bright air, with mud-stained feet and ankles and long, lithe arms, exercised their restive horses before the race. [...]
That's actually pretty cool. You're the first person that's actually showed verbatim reproduction in a discussion like this.
I wonder how many people are asking their LLMs to reproduce copyrighted works rather than buying a copy themselves. Or, more realistically, just going to Anna's Archive.
If Anthopic had bought all the books it had trained for say at market rate we’d be having a different conversation now. Anthropic, through this settlement, has been forced to pay back, at least something… Kim would likely not have had enough money to compensate the victims and probably caused some more direct dammage by sharing pirated content. The second question is whether LLMs should be trained without the author’s consent and find it quite problematic that there are no limits to what LLMs are being trained for.
anthropic etal would not have a product to sell without their violation..
kdc had a service that just happened to be popular for pirating...
how are the two even remotely similar?
> Anthropic, through this settlement, has been forced to pay back, at least something… Kim would likely not have had enough money to compensate the victims and probably caused some more direct dammage by sharing pirated content.
You're thinking civil. They're talking criminal. Criminal law enforcement does not (well, isn't supposed to) look at your ability to compensate before deciding what to charge you with.
Since we're on the criminal side - what criminal statute would apply to Anthropic?
And what criminal statutes were used for the cases we're supposed to compare to?
Fuck off! what about Aaron Swartz ? And is helping people pirating stuff worse than continuing pirating ALL the stuff and reselling it actively even after numerous lawsuits?
Some of you really don’t deserve good things. You should be blocked from using AI on more than one device without paying an additional subscription plan.
Any way, the internal emails are available where you can see the executives of megaupload knew exactly what megaupload was being used for, and even used it themselves to pirate content, and went out of their way to allow copyrighted content to remain up after takedown notices were sent.