| Unfortunately, a "full" text crawl of the internets is a YUUUGE amount of data to manage Maybe instead of a problem, there is an opportunity here. Back before Google ate the intarwebs, there used to be niche search engines. Perhaps that is an idea whose time has come again. For example, if I want information from a government source, I use a search engine that specializes in crawling only government web sites. If I want information about Berlin, I use a search engine that only crawls web sites with information about Berlin, or that are located in Berlin. If I want information about health, I use a search engine that only crawls medical web sites. Each topic is still a wealth of information, but siloed enough that the amount of data could be manageable to a small or medium-sized company. And the market would keep the niches from getting so small that they become useful. A search engine dedicated to Hello Kitty lanyards isn't going to monetize. |
[1] https://en.wikipedia.org/wiki/Searx [2] https://asciimoo.github.io/searx/ [3] https://stats.searx.xyz/
featuring the semantic map of [4] https://swisscows.ch/
incorporating [5] https://curlie.org/ and Wikipedia and something like Yelp/YellowPages embedded in Open Streetmaps for businesses and points of interest, with a no frills interface showing the history (via timeslide?) of edits.
Bang! Done!