Hacker News new | ask | show | jobs
by Ferret7446 16 days ago
There has to be a signal to detect it.

Take this sentence: Bob went to the store to buy milk.

Was that AI generated or not? There simply isn't a signal there. The problem isn't noise, the problem is, is there even a signal to begin with.

Sure, you might be able to recognize the quirks of a specific LLM just as you recognize the quirks of a particular person, but as the number of LLMs proliferate, then the signal turns into noise. (The signal isn't buried by noise, it becomes noise. The signal no longer has any discriminating power.)

3 comments

The article itself explained that it was much easier to classify text as human or LLM generated than to have "human" as just a category along with all the different LLMs as it's likely the LLMs are distilled from each other, creating a unique footprint.

If a signal is weak, it might not even appear in every sentence, but that doesn't mean it doesn't exist. For instance, I don't recall ever consciously using an em dash, but you'll probably need an entire paragraph to find one in LLM-generated text.

My own sense of whether text is generated is partially based on its sheer length - humans typically don't bother writing so much.

but.... the LLMs are actually all trained on approximately the same stuff, and tend to have similar quirks. In the human world, writers develop recognizable voices, which are detectable and classifiable (as in the article we have all supposedly read). Furthermore, we don't necessarily care about telling one LLM from another, just that they aren't human. That's different from trying to identify one human amongst a sea o fhumans, or one bot form within a sea of bots.
Statistical power comes from having many samples. I agree that having just one sample doesn't take you very far.