Hacker News new | ask | show | jobs
by o_m 10 days ago
I've subscribed to Kagi for three of years now. Kagi is still great, but it doesn't feel like "Google ten years ago" anymore. I don't think it is because Kagi got worse, rather it is the web that has become worse. The content nowadays kinda sucks.
5 comments

Yeah, it's become quite noticeable. The amount of clearly LLM generated sites that I've blocked just keeps going up and up. The only advantage Kagi has in that regard is that blocks are account wide, regardless of what device you use
I would be interested in seeing that block list if you have it somewhere
Kagi nowadays has a feature called "SlopStop" that you can opt into. It's basically a global block list for LLM generated websites, etc that is shared across all users who opt into and anyone can report sites/content to it.

Kagi of course reviews it but generally it's good enough that I don't see much LLM generated content in the results anymore.

I was extremely pleased to login to my Kagi account and see that SlopStop was enabled by default. Thanks for the heads-up on this, and thanks to Kagi of course for enabling it.
> Kagi of course reviews it

They had quite some lag at the beginning. I made reports from the very day they announced it, and none of them were processed last month. But they are now! (Though they rejected half of them)

I had the same experience. I don't know why they haven't invested more in this feature. Heck, if they classified every site in their index as "AI/not AI" using machine learning, and had a large allowlist of "known-good" sites, that would take care of most of the problem (I am aware of the drawbacks of text-only AI classification, but they could use other indicators, like the site's publication date, domain name, domain registry date, etc.). I think for right now, the "state-of-the-art" for non-AI search is going to entail maintaining large whitelists/blacklists of websites that are crowd-sources by users (e.g. what you can find in some GitHub repositories). I was hoping that's what SlopStop was going to be, but it is not nearly as effective as it needs to be.
Discussed on HN back in January: https://news.ycombinator.com/item?id=46716806
Do they review false positives? Do site owners have an avenue to challenge being falsely blocked?
They started marking a lot of my reports as not slop when it’s very obviously ai garbage. Stopslop is nice in theory if they actually handle reports correctly. It’s exhausting to spend time trying to help clean up their results only to be told no.
I doubt that my list would be of much use to you, since a lot of the sites relate to specific hobbies.

However, as eloisius mentioned, you can view the most commonly blocked sites. Going through that list is a good start: https://kagi.com/stats?stat=insights

I didn't know that page existed, thanks! I was amused to see that my blocklist already includes most of the sites that appear there.
If you’re a Kagi user you can view the most-blocked list and adopt it
- start with this website https://www.xjavascript.com/blog/ block it everywhere
But this is also why Google "sucks now". The thesis of these competitors was that Google got worse because of bad product decisions or implementation failures. But I found that to be a very questionable thesis from the start.
> The thesis of these competitors was that Google got worse because of bad product decisions or implementation failures.

But they did though. Google results got significantly worse after ads started lagging in revenue. So they hindered search and increased focus on the ad platform which shot up profits in 2019 I believe.

The web becoming more inhospitable is true. But Google has the tools, the data, and the engineers to wade through the bs and return good results. They just rather have bad results and good ads because that is where their bread is buttered.

A lot of competitors are working with a fraction of the budget, teams and data. So asking a small team to do proper AI detection, to return smaller indexed sites with high quality info over mass linked SEO abused hellscapes is very complicated, but asking Google to figure it out is proportional to their talent, expertise and resources.

It just not going to move their bottom line, so search is de prioritised compared to ads (some might even argue its made worse on purpose to make ads look even better)

This is the thesis, I understand the thesis, I remain skeptical of it. In my view, search sucks now primarily because there just isn't much worth searching for on the open web anymore. I think the "it's Google's fault" thesis is mostly an unwillingness to admit this.
To large extend, that content stopped dying after Google deprioritized it in search results. People stopped being able to find it, sites/communities stopped having views, authors gave up.

I stopped caring about my blog about that time - after I was unable to google my own postd. I remembered the name and some of the content, but it was just not coming up in the search. There was also noticeable drop in views. I kind of concluded it is pointless when it is not coming up in searches unless I SEO hard.

I could be convinced of this narrative. Like, if someone writes a book about this, with good evidence of what caused this to happen, I might find that persuasive.

But to me, what happened was social media and video becoming the locus of attention. I don't think people stopped putting content on independent websites because their traffic from search results went down, I think they just started putting content on Facebook and Twitter and Instagram and YouTube and TikTok instead. To me, the whole way the web is used just changed over this period of time, in ways that mostly had very little to do with Google's search algorithm.

You remain skeptical of something proven through leaked emails?

https://www.searchenginejournal.com/google-execs-scheme-to-i...

The thing those emails demonstrate is not the thing I'm skeptical of!
I don't see why it can't be both. The web has less good content AND Google has altered their search results such that results are worse generally for us, but better for them and their bottom line.
I'm not saying it isn't both, what I'm saying is that the thesis of the newer Google competitors is that the problem is mostly that Google has chosen to suck, so they can do a lot better just by choosing to not suck. But I think that thesis is wrong, that it is mostly that the web sucks now, and that these search engines are ultimately likely to be unsatisfying because they can't actually solve that underlying problem.

But if I'm wrong, that would be great! The comment that started this thread made me think that it seems like I'm not wrong though.

I think the counterargument would be that the web sucks in large part because of the way it’s been shaped by Google. Most of the hyperwordy blogspam is driven by sites trying to optimize their SEO for Google. So websites like Kagi could potentially push the underlying web in a positive direction
¿Por qué no los dos?
Google can both be worse because of the web results declining in quality and terrible product decisions. It is of course much easier to see one of the two, and make an alternative search interface possibly using different rating criteria. It just can’t solve the entire issue, because spam is also absolutely cheap to mass produce now.
This is absolutely true and yet I know from experience the product forums for years were people screaming at Google for stripping away features

And I know they removed the boolean + operator because of Google+

Yes I agree. I don't begrudge competitors trying to do better at search at all. Maybe they'll succeed. I've just always been skeptical that they will be able to make significant progress on this problem, because I think the problem is ultimately garbage in garbage out.
The percentage of quality content on the web has gone down but it still exists. I usually use Marginalia Search (https://marginalia-search.com/) when I'm researching something and not finding any good results on Kagi or Google.
kagi sorts by recency and thats the best. i dont just want this months results but whats been update.

of course i dont know how they calculate but hopefully they SEO proof against listicles that just increment date stamps with the same stale content

imagine if we had a browser like kagi with its own engine from scratch where you are not the product and it charged you a monthly subscription for privacy free browsing
What's wrong with Firefox and AdNauseum?
You're the product in Firefox, just a different kind than in chrome. Ad Nauseam is silly because ad companies know they are fake clicks and don't count them. Ads aren't the only thing you'd need to block from the internet.
There is zero evidence Ad Nauseum does not work, but there are plenty of signals it does. Google banned it from the Chrome store and setting any threshold below one hundred percent click rate just uses random.

Even if it is somehow detectable it's just Ublock origin and not every ad network is worth billions of dollars

Have you bought ads and then browsed that site with Ad Nauseum and then discovered that you paid for the ads?
i dont follow em closely but wasnt there a new firefox ceo who wanted to push AI everywhere
All my hopes lay with Ladybird now
I mean, Orion exists (on Mac at least, although I heard there's a Linux version now?). Might not have its own engine built from the ground up, but WebKit is pretty much the next best thing.
> although I heard there's a Linux version now?

Yes, there's a Linux beta. Discussion from a few days ago on HN: https://news.ycombinator.com/item?id=48970894.

Link to the beta: https://orionbrowser.com/platforms/linux