Hacker News new | ask | show | jobs
by subarctic 17 days ago
I don't run one of these sites that has these issues so I'm really not aware of this problem. How can it be that sites are getting overwhelmed with scrapers that are just looking for training data? You only need to scrape it once to train a model, so shouldn't there be less traffic from this than there is from search engines?

On the other hand if the article is wrong and the traffic is coming from other ai uses (like an agent visiting pages on behalf of a user) then that would make sense.