I'm a bit on the fence about this, but leaning towards the "bad luck". I'm sure there's a large swathe of nuance that I'm missing, but my simplistic view is: If they don't want to be scraped then they don't get included in "the thing" which, at minimum, is a data point for consumers to make consumer decisions about.
It strongly depends on how scraped data is being used.
If your 'cat' hasn't been able to catch their 'mouse' then the cat needs to get smarter, or look for alternative sources of mice, or the cat should be considered 'unviable'.
Have you approached them to get access to their data? Have you explained to them how your service can benefit their business?
(I generally come from a position of suspicion as to why someone wants to scrape data that the owner goes to certain lengths to protect, but then I'm also an 'information wants to be free' kinda person, but the Internet is increasingly an untrustworthy place, so security is overruling narrative).
> I generally come from a position of suspicion as to why someone wants to scrape data that the owner goes to certain lengths to protect
Let's try this for example. Suppose you want to create a price comparison site.
A lot of the major retailers don't want these, because they want the customer going to their site when they want to buy something, not to the price comparison site that tells the customer which retailer has the best price and then might not always be them.
If the comparison site doesn't have the prices from retailers who sometimes have the best price then it can't serve its function -- customers still have to check sites manually. And comparison sites are pro-consumer whereas blocking them is anti-competitive, so the ones doing the scraping are the good guys.
Suppose someone wants to create a price comparison site, but is unable to access all the data that would serve their customers in finding the lowest price.
The I guess the site is either useless, or can be used in combination with customer's own research that includes sites not included in the comparison. The site is still useful, it's just not exhaustive. There's nothing much in the world that's exhaustive, it's all a range of percentages.
Also, maybe the source that isn't scrape-able becomes less popular as a result of not being included in the price comparison list. If they're always more expensive, then neither you nor they have any advantage in listing on your site. If they're always cheaper, then your site may have no purpose to serve.
Everyone seems to want _everything_ despite _everything_ not being necessary to provide a service. The lack of a certain set of data may itself be a data point, and should be used as marketing material for or against the company resisting the scraping.
Just because residential proxies are somewhat of an 'easy answer' doesn't mean they're right. Use more creativity and imagination! Or follow a new idea. It's not like price comparison sites are a revolutionary idea.
(I'm a consumer who happily does significant research before a buying decision, going to various sites and checking prices, warranties, model numbers, reviews, availability, delivery times, and all the shit. And prefer doing the research myself than trusting a price comparison site, so I'm not the target market, thus my bias in my commentary and opinion on that topic).
> so the ones doing the scraping are the good guys.
Everyone thinks they're the good guys. If there's profit involved then there's always an element of delusion.
> The I guess the site is either useless, or can be used in combination with customer's own research that includes sites not included in the comparison. The site is still useful, it's just not exhaustive.
There is a significant difference between having the prices for 85% of retailers because you haven't figured out how to get pricing from 15% of them and having the prices for 5% of retailers because the biggest ones block you from getting their prices. The amount of work you can save the consumer, and thereby the usefulness of the service to them, changes by a dramatic amount.
> Also, maybe the source that isn't scrape-able becomes less popular as a result of not being included in the price comparison list. If they're always more expensive, then neither you nor they have any advantage in listing on your site. If they're always cheaper, then your site may have no purpose to serve.
Now consider the possibility that they sometimes have the best price and sometimes don't.
> The lack of a certain set of data may itself be a data point, and should be used as marketing material for or against the company resisting the scraping.
Suppose you go to the price comparison site to look for the best price, the best price in its data set is $22, then you go to Amazon and they have it for $19 because it doesn't have their prices. The customer then starts checking Amazon in addition to the price comparison site. Their price is only actually lower 15% of the time, but another 75% of the time it's exactly the same, so the customers who don't want to keep checking two sites start checking only Amazon instead of only the price comparison site, at which point Amazon gets to charge them a higher price on 10% of stuff, or a higher price on more than 10% once people have been trained not to check. And prevent them from patronizing a random competitor when their price is exactly the same
Causing that to happen is the reason they don't want their prices in the comparison site. If not being listed there hurt them then they wouldn't be trying to prevent scraping.
> I'm a consumer who happily does significant research before a buying decision, going to various sites and checking prices, warranties, model numbers, reviews, availability, delivery times, and all the shit. And prefer doing the research myself than trusting a price comparison site, so I'm not the target market, thus my bias in my commentary and opinion on that topic
Consider also that you may be able to justify doing this when buying electronics, or choosing which brand of something to buy on a recurring basis, but if you just want the best price on the product you already know you want, it's crazy to spend hours to save cents. But completely sensible to switch to a search box that would save you 10 cents on every $2 purchase by giving you price-sorted results from multiple retailers who all sell the same products, if it's allowed to exist.
That's cutting the tether from a _lot_ (there's no way to overstate this) of useful information, but that's potentially one of the great things about the open AI/LLM models, is that all that info is baked in there.
Disclaimer: As far as I understand it. Please educate me if I'm way off the mark.
Also, how much of the content of reddit is in Common Crawl? (same disclaimer applies to this comment)
I also understand 'the archival mindset', I hoard a bunch of data. But I also understand the logarithmic graph of futility.
Government website data should be openly available one way or another (at least within the country, and with reasonable security provision against obvious maliciousness).
If it has to be scraped then there may be other problems (which include lack of resources to make the data API accessible).
Apologies for the double post, but I've got an alternate perspective:
Do you allow the data you've scraped to be scraped? Do you share it as freely as you desire the 'companies that want to hide it' would? Or do you consider the scraped data is 'hard earned reward for effort' and therefore has value that others should subscribe to your service for?
It strongly depends on how scraped data is being used.
If your 'cat' hasn't been able to catch their 'mouse' then the cat needs to get smarter, or look for alternative sources of mice, or the cat should be considered 'unviable'.
Have you approached them to get access to their data? Have you explained to them how your service can benefit their business?
(I generally come from a position of suspicion as to why someone wants to scrape data that the owner goes to certain lengths to protect, but then I'm also an 'information wants to be free' kinda person, but the Internet is increasingly an untrustworthy place, so security is overruling narrative).