Hacker News new | ask | show | jobs
by enraged_camel 23 days ago
The problem is that the remaining 10% can bite you in bad ways.

I was in Cotswolds, UK a couple of months ago. For those of you who don't know, it's a rural region known for its "chocolate-box" villages and honey-colored limestone architecture. Basically, you go from village to village, most commonly via bus, taking in the sights and doing touristy stuff.

When planning the trip, my sister used ChatGPT, which helpfully (and relatively quickly) found the bus schedules and times for each hop.

Midway through the day, though, we ran into a huge problem: it turns out bus schedules are different on Sundays, and more limited. Which meant we couldn't actually go to our primary destination (the Model Village), and had to cut the trip short.

Yes, ChatGPT was quick and pleasant to use, but missed a crucial detail.

Afterwards I tried it with Opus and it did not make the same mistake.

3 comments

Arguably I'd call that the 90%. In my case, answering the restaurant question correctly with "Rishi" in my tests was the sole intent and 90% of the problem. All the models "helpfully" added extra junk about the closure, dates, quotes, etc and many of them got these details wrong--the 10% or extra crap not central to the question.

If the central question was "what is the bus schedule on `day`" and the model screws that up, it gets a fail in my book.

Also curious if Google Maps gets the timetables correct (assuming it has them).

Semi-related, I also discovered that the default web search/fetch tools are pretty primitive and Exa MCP annihilates them. I ended up doing some comparisons with Claude Code comparing built-in server-side to Exa and to a Python MCP that used SearXNG for search and Exa was a clear winner and Python+SearXNG ended up coming out roughly the same after a few cycles of letting Claude optimize the Python code and adjust SearXNG settings. Ultimately it landed on this (making some changes to optimize returning relevant context directly in the search results so the model didn't need an additional web fetch call) https://gist.github.com/nijave/604c43e3e0fdcd60f5280d3a6b109...

> Midway through the day, though, we ran into a huge problem: it turns out bus schedules are different on Sundays, and more limited. Which meant we couldn't actually go to our primary destination (the Model Village), and had to cut the trip short.

Why trust an LLM with information like bus schedules? They fuck up things like this routinely.

This likely comes down to how it accessed the bus schedules (i.e. web search tool) and not intelligence.

You need to add the actual bus schedule to context somehow (research agent, custom tool or just dump in prompt) and even the simpler modern models will be able to do the planning.

Tool usage competency is part of overall intelligence. If the model can't get the information it needs, it must clarify that in the response.
This isn't tool usage competency, it's tool quality and/or luck. Regular web search is not good for grounding if you want accurate results. You can ask the model to make a tool for getting bus schedules and then use it only then you are comparing apples to apples in this case.
If the model can't get the information it needs to accurately answer the question, it must surface that risk to you instead of guessing. This is part of model intelligence and tool use competency. Fable and to a lesser extent Opus is very good at this
This would be hallucination rate and neither of those models excels in this area and in fact Grok does, or at least prior version.
It's both. Some models, given web search and web fetch may only run a search and assume the summary text is correct and blindly return it. Others will validate by running a web fetch and checking the whole page contents. Even better, the model will run multiple searches and fetches to cross check the information. The best models will attempt to verify whether a source is authoritative or not and try to only return authoritative results.

I have an example here: https://gist.github.com/nijave/2873b8b10d8c732e46264237b0755...

Tldr; all the Claude models had identical tools and some used them efficiently and verified data while others did a crap job and hallucinated responses. Additionally, Exa MCP tools generally worked better even on older/smaller model (Llama)

Unless I'm missing something I think to a large degree you're just comparing system prompts.

If I add "Research the question extensively" to your prompt at the end I get the correct answer from Haiku and Sonnet Med on first try and I've reproduced the original prompt not returning the answer.

Unfortunately every other run now gets your gist in results.

It could be variations in system prompt but I'm not sure how "change the user prompt to encourage tool usage" proves that either way