Hacker News new | ask | show | jobs
by clnq 1162 days ago
I agree for the most part, but I wish to underscore the primary function inherent in each tool. For a LLM, it is to generate plausible language. For a bolt, it is to offer structural integrity. For a car, it is to provide mobility. Should these tools fail to do what they were designed for, we can rightfully deem them as defective.

GPT was not primarily made to produce factual statements. While factual accuracy certainly constitutes a desirable design aspiration, and undeniably makes the LLM more useful, it should not be expected. Automobile designers, for example, strive to ensure safety during high-speed collisions, a feature that almost invariably benefits the user. However, if someone uses their car to demolish their house, this is probably not going to leave them satisfied. And I don't think we can say the car is a lemon for this.

1 comments

> For a LLM, it is to generate plausible language.

LLMs are not being sold as delivering only plausible language. Is it the craftsman's fault when the salesman lies?

Who (and with links showing it) is selling it as delivering more than that?

From https://simonwillison.net/2023/May/27/lawyer-chatgpt/ (which has the dates)

> Mar 1st, 2023 is where things get interesting. This document was filed—“Affirmation in Opposition to Motion”—and it cites entirely fictional cases! One example quoted from that document (emphasis mine):

GPT-4 wasn't available until March 14th ( https://openai.com/research/gpt-4 ), so we're dealing with 3.5 here.

Where was 3.5 advertised to do more than generate plausible language in a conversation and who is making those claims?

The very first limitation listed on the ChatGPT introduction post is about incorrect answers - https://openai.com/blog/chatgpt. This has not changed since ChatGPT was announced. OpenAI is advertising that it will generate more than plausible language.

I think you are barking up the wrong tree here. As much as I understand your scepticism, OpenAI have been very transparent about the limitations of GPT and it is not truthful to say otherwise.

A definition of "plausible" is "apparently reasonable and credible, and therefore convincing".

In what limit does "apparently reasonable and credible" diverge from "true"?

We'd make the LLM not lie if we could. All this "plausible" language constitutes practitioner weasel words. We'd collectively love if the LLMs were more truthful than a 5-year-old.