Hacker News new | ask | show | jobs
by poetril 10 days ago
I feel like half of what I do on HN is talk about Kagi. I switch to Kagi in 2023, and haven't looked back. Between vim keybinds for navigating search, explicit AI opt-in, the ability to curate search (block sites, increase others), and more. It really lets me feel like I can control my search tool, and at least so far its felt like our interests are aligned, I give them money and they make sure search is as good as possible for me.
5 comments

They're still ultimately scraping other engines pages to get the results though. The last time this was discussed an engineer on their team told me that they were building their own index internally, but I couldn't get any additional information about that and it doesn't appear to be in their docs.
Why is it bad if they're using other engines to produce results?
I like paying to invest in a new sustainable ecosystem. Using SERP API doesn't feel like that to me. See Brave for a counterexample.
My pushback on this idea and your previous comment in this thread is that a search engine is a service being provided to me.

I pay Kagi, they give me search results. It’s not an ecosystem. None of their competitors are all that much of an ecosystem, either.

If Kagi goes out of business tomorrow it’s really not my problem, with the exception of the prorated balance of the subscription.

The only thing that really matters is that running searches in Kagi is a better experience in Kagi than Google or Bing or DuckDuckGo or most of the rest. That’s what makes me pay for it.

I don’t really need to know how the sausage is made (except for the privacy aspect).

This. We are all too worried about the future. Like today's "I won't use Codeberg because tomorrow they could ban all repositories that contain Chinese characters." Well if that happens then you move. We all die in the end anyway.
You undervalue our collective momentum.

The speed with which many of us respond when a tech company does something really stupid is a key reason that things are not even shitter.

I think that ship has sailed [0].

If Microsoft hasn’t managed to build a meaningful search index after 17 years of sinking billions in it, there’s a good chance that no one ever will.

[0]: https://blog.kagi.com/waiting-dawn-search (cf. section _The problem: A search monopoly_)

I guess you can't make such ecosystem before growing sufficiently large. And for that you need to rely on something for some time first.
It is a bit weird that a ‘search engine’ is using other companies search results. They aren’t the only one doing this though.
I wouldn't be surprised to hear that this was always the plan for Kagi. Google has been rather distracted and complacent in their, former, core business. If Kagi is bootstrapping their indexing to emerge on the other end of this as the next "good" search, then good on them.

I'll continue to support them because it's a great experience.

Why reindex everything? You just need to show the relevant information
The search engine is a lot more than the index. And they have their own index
If engine A uses the results of engine B, and engine B uses the results of engine A, there is a problem :-).
What tier do you use for it? I would like to use Kagi, but it seems really expensive for the lower tiers. I may be overestimating how much I use search but it seems like it is pretty limiting.
Do they let you chose to not send your search queries to say Yandex or Google?
No, they don't, and I had a feedback request back in September 2024 about that:

https://kagifeedback.org/d/4727-option-to-choose-or-exclude-...

I ended up quitting Kagi over that and built my own search engine for my own use. But if the comments here are correct that Kagi are trying to build their own index that replaces all the others, that would be great news. It seems lots of people still like Kagi regardless of that too.

For anyone trying to build their own, turns out SQLite can take you much further than it seems you should ever be able to. A metasearch layer fills in for everything else.

I am trying this. I noticed I only browse a subset of sites a lot. Also HN works wonders for finding links. But how do you get around all those bot measures/cloudflare nowadays? It seems only google IPs get a special pass.
Yep, discovering that you only need a subset of sites is the key. In my case, I still had my entire Firefox browsing history in its SQLite cache, so I tried to index every page I'd ever visited, and had a list of the domains I actually visit. It's the same insight Brave Search had - they used to be a metasearch engine, but they built their own index until they could serve 99% of queries directly, even though their index was vastly smaller than Google or Bing.

I struggle with indexing too. I don't do any crawling, just indexing from sitemap.xml files or URLs I add to the queue manually. Also I'm indexing from my own residential PC, not a hosted VPS. Something that helped me was Cloudflare's new crawler API. For all the sites where Cloudflare blocks you, just use their API instead:

https://developers.cloudflare.com/browser-run/quick-actions/...

https://developers.cloudflare.com/browser-run/quick-actions/...

Also for a quick index bootstrap, the Curlie database can be useful. It's the old DMOZ directory, the open Yahoo competitor, with 1.4 Million websites. A lot of the links are now very outdated, but it can be useful, and it's less than 500MB when converted to an SQLite database.

Yes, Kagi is building its own index, and try to mix it more and more if it finds results from it.

More details are at: https://help.kagi.com/kagi/search-details/search-sources.htm...

Also, for every query, it shows a short infoline telling the index distribution.

Unfortunately, that's nearly identical to the text from September 2024, with the only change being an increase in the number of external search sources:

https://web.archive.org/web/20240901052114/https://help.kagi...

I'm aware of their Teclis index, I used to be a paying API customer. Teclis is very small. It's primarily an index of indie blog websites (smallweb) mostly crawled via RSS feeds. That itself is a very cool idea, and a great supplemental index to have! But it isn't the kind of index that could ever replace their Brave / Yandex / SerpAPI dependencies. The Teclis API is even supplemented with results from Marginalia Search, a much larger index created by one person with less funding.

There's some technical info here on how Teclis is indexed, eg Readability.js for content extraction & Elasticsearch for the full-text indexing:

https://teclis.com/

Kagi's short infoline about index distribution sounds new to me though. I would love to see a blog post from Kagi about that & where they're at with query coverage.

I can't imagine they don't. I just run my own searxng and turn on the engines I want.
I tried Kagi a year or two ago but it just didn't have most results. Google always had what I was looking for.

I might try it again sometime.

I’ve found the opposite. Most notably was when I joined an outage call at work. The call had been going for a few hours. They knew the issue, but didn’t have a fix. There were a dozen or two people on the call all searching for options (presumably all using Google). I went to Kagi and found an answer within a couple minutes that we ended up rolling out.

I’m just one person, and that’s just one incident, but it happened pretty early on after I switched to Kagi and had a pretty big impact on my perception of it.

I had been trying to get away from Google for many years, and with other options like DDG, I always felt I needed to go to Google for various searches, but so far I haven’t been back to Google at all in several years now.

I dont use it anywhere near as much because it felt like I was hitting their limits too quickly, not sure if that has changed.
I've subscribed for years. I've never hit a limit once.
Over the last 2 months I’ve started to hit AI limits. I don’t feel like I’m using it more than before, but my usage went from $1-2/month to $9.50… where I’m actively avoiding usage to not hit my $10 limit halfway through the month.
Do you know what their upper limit is? I average 1900 queries/month so I'm curious where they drew the line for you.
Curious what your distribution is, as a ~400 person that thought they did more.
I do between 2,000 and 3,000 searches per month and have never hit a limit.