Hacker News new | ask | show | jobs
by rw2 28 days ago
Should charge AI for training on top of it or get them to donate. A small amount can fund them easily.
5 comments

That would be a trap. It's healthier for a non-profit to have many small funders than a few large ones.
exactly, the only reason Mozilla exists today is as a legal shield against an anti-browser monopoly suit against Google. that's the product they sell, and Google is paying hundreds of millions per year for this valuable service
I thought google pay Mozilla so they don't set the default search engine to something else (they same way Google pay Apple for Safari) and so Google continues to dominate and makes money of ads.
That cannot be the reason: Firefox has a market share of 3%...
If Google didn't pay hundreds of millions, Microsoft would.

If Google just wanted them to exist and didn't care about profiting off of the search traffic they wouldn't partner with Mozilla.

Papers submitted to arXiv under its most permissive license should always be free, as in beer, speech, freedom. For researchers that contribute to it, that is the intention for a reason. It is to serve public and corporate good without restriction.

This isn't me siding with AI companies by the way; it's a slippery slope argument.

The papers can make it free, they can just build a convenient search/retrieval API layer so Open AI and Anthropic don't pay 1-2 engineers a year at 1-2m to crawl and index that info for training.

The AI services have an option then to pay for this service, support a open service, or write their own crawler. I think if every open AI request didn't just do a web search but a more targeted arXiv search the results would be better.

> It is to serve public and corporate good without restriction.

Sometimes those two are in conflict, such that it will not be possible to satisfy both simultaneously.

Part of the promise of open access and open science is that the information is free and open to all. Including robots.

I submit to open things because I want my material to be openly available. If I wanted restrictions, I would submit to gated journals.

It would be nice if the corporations pay for the bandwidth. As arXiv corpus hosted on GCP, network charge cannot be small.
as if they would pay.... they would pirate the contents as they already did
Engineers to crawl that content cost money too
They’ve never paid for any content?