Hacker News new | ask | show | jobs
by jofzar 7 days ago
We had googlebot blast a random customer system and almost cause an outage, this is when I first learnt that google will use it for AI training also. It's honestly kind of frustrating also because you then search on it and theres (was) nothing on how you are meant to "correctly" tell google to fuck off, and not use it like that.
2 comments

The vast majority of Googlebot user agents are lying. Real Googlebot is pretty well behaved in my experience. You should use reverse DNS or IP lists to check: https://developers.google.com/crawling/docs/crawlers-fetcher...
If you're using Cloudflare, set up a security rule to block requests that have "Googlebot" in the UA and are not recognised by CF as a real bot.
Google's web scraping functionality has been acting as a ddos for more than two decades. I've seen literally hundreds of reports of them attacking websites and taking them down, where there's nothing you can do but accept the traffic, or get delisted

This is unfortunately nothing new. There's no correct way to tell them to fuck off, they do not care, and they never will do. People have even taken them to court over this

If a site cannot handle traffic from the real Googlebot that is a serious issue with the site itself since it's actually pretty conservative

Also I should note there are lots of fake Googlebots...

Indeed, I've got a site which gets a lot of bot traffic and google bot is pretty sensible compared to a lot of other mainstream bots.
Yes, it really does not make that many requests. In fact lots of site owners struggle with having it not crawl and index their site enough
It is mostly, but it doesn't take a lot of googling to find sites getting ridiculous amounts of traffic from googlebot on google IPs. Its one of the most common complaints about google's search indexing
Lots of people abuse Google Cloud to get a "Google IP" for a fake Googlebot. Why don't you show me a single screenshot from Google Search Console showing a high number of requests to a site that would be counted as a DoS? All requests from the official Googlebot are logged there so if it's such a common problem it must be very easy for you to show me this.
Still waiting for a single shred of evidence ;)