Hacker News new | ask | show | jobs
by customguy 14 hours ago
For a one time coin flip, sure. For a lot of them, depending on the stakes, that can be very, very wrong.

https://www.google.com/search?q=why+99.99+accuracy+is+not+en...

It doesn't take a lot of thought experiments to realize this. Imagine if every bite of food we eat had a 0.01% chance to turn into something instantly lethal in our mouth. Average lifespans would be reduced measurably, and apart from anxiety, we'd develop all sorts of strategies and laws around that. E.g. absolutely NO eating for airplane pilots. You wouldn't go on a date to have dinner, dancing and sex, you'd go dancing and have sex, and then have breakfast. People would modify their jaws and stomachs so they could eat less, but bigger chunks of food. It would be a whole thing!

And that's not even talking about water changing on us, or a tiny chance of getting sucked into the toilet whenever we use it, and a lot of other things where going from damn near 100% to 99.99% would change everything for the worse, by so much.

1 comments

ziofill's claim was that "A direction [in an LLM's embedding vector space] that is 99.99% accurate" is fine for practical purposes, not that 99.99% is fine for the chance of any given bite of food not killing you or similar hypotheticals - you'd want a few more 9s there.

To justify relevance of inability to correctly answer liars-paradox-type questions ("what won't your response to this be?"), the article suggested the way LLMs are used in practice is dependant on them being entirely accurate truth oracles:

> > as a truth-oracle [...] is how these things will be used practically by the vast majority of people. They are already replacing standard Google search results

But for the replacement to make sense they just need to be more accurate than what they're replacing (ignoring other factors like convenience and cost) - in this case standard Google search results and knowledge box which were obviously not 100.0% accurate.

I think you are making a lot of assumptions on what "practical purposes" even mean in that case.

Please think carefully and then try to tell me whether or not some Trump-administration government agency would not in the "99.99% reliable" circumstance just plug an LLM into the nukes and funnel worldstate input into it and have it make the decision "is it time to fire the nukes?" over and over again each second.

I argue that that would require far more nines than even "will food turn to poison in my mouth" would.

> plug an LLM into the nukes and funnel worldstate input into it and have it make the decision "is it time to fire the nukes?" over and over again each second [...] I argue that that would require far more nines than even "will food turn to poison in my mouth" would.

Sure - but (even assuming that's a practical purpose) the point is it that it doesn't need to be a 100% accurate truth oracle, which is all the article's argument prohibits. If the current human chain of command has 99.99999994% accuracy, then 99.99999995% accuracy is an improvement and not ruled out by the argument.