Hacker News new | ask | show | jobs
by guardiangod 3 days ago
“You are trying to kidnap what I have rightfully stolen, and I think it quite ungentlemanly."
2 comments

Google crawling respects robots.txt and doesn't break capcha. It is easy to tell Google to piss off. SerpAPI fully relies on end-user proxies distributed like malware (in LG tv apps for instance), it has no other way it could function because it exclusively ingests data from sources that tell it to stop. If you wanted to scrape a bot friendly site, you wouldn't need SerpApi
I'm vaguely sympathetic to this argument.

But only vaguely. Google uses its monopoly position in advertising to basically ensure that you allow them to scrape your site (or if not you personally, the majority of revenue driving sites). They have the benefit of being allowed by default.

They also then scrape again at the user level for users operating chrome.

They also conveniently ignore global blocks for their adsbots (you have to specifically name them to block them).

If you're not Google, you likely don't have this luxury.

My preference would be that governments force search indexes to be public. The exact mechanisms for this can be debated.

And they do specifically also scrape sites anonymously, ostensibly to ensure you don't serve different content to GoogleBot than users (though this is in my experience unreliable at best, and that's even before we get to the "do we trust that that's the only scope?").
The question I have that no one has really answered is...what kind of crawling have the current frontier labs done historically and what are they continuing to do now for training? Inference can follow rules easily, but the training is a big black box that mostly gets headlines for books getting slashed but that's not the only source of data is it?
This sums up every single bit of AI "progress" since 2022.