Hacker News new | ask | show | jobs
by dofm 15 days ago
> Text is simply not information dense enough to be able to decode some arbitrary signal of provenance from it.

This does not sit well with personal experience and I wonder if it is just one of these questions of AI people being unaware of the level of skill that exists in domains they think have been automated.

It is of course possible that my tendency to spot LLM-written text has much to do with the way that it sounds like an averaged Californian college student to my British grammar-school-educated ears, as so many of the situations where I am encountering AI text are Brits using it without apparently realising they are giving themselves away.

But I know people who don't have particular technical skills in this sphere or a grammar-school background who also have an uncanny knack for pointing out LLM-written text.

> Words, no, the signal is far too sparse and we are well into undetectable sophistication with today's models, let alone tomorrow's.

I especially don't think this is true. Will they be able to do it in the future? Maybe. Is it possible to prompt a current cloud LLM to write in a way that is obvious? Yeah. (IMO Gemma 4 writes less detectably than most of them!)

But my instinct is that someone with any facility for language is going to be better than chance at spotting LLM-written text once it is three or four paragraphs long. So I think it should be possible in principle to train machine learning systems to detect those patterns.

2 comments

If you can train a system to detect these patterns, presumably you can train systems not to generate text which matches them?

I do struggle at times with thinking my own writing looks like AI. But I’m an average Californian who went to college half way between SF and LA…

> If you can train a system to detect these patterns, presumably you can train systems not to generate text which matches them?

I don't know. I mean, it feels like the systems that would detect them are likely qualitatively different to the machines that make them.

One of the things that feels obvious to me is that LLMs are always going to write in a new way, because words do not get all that close to perfectly conveying the inner thoughts of competent writers. Competent writing is always a battle to find the better word, or even to create it.

So sure, you could add another adversary that the generator has to satisfy, but "this sounds like a machine wrote it" is only an observation; it's not a prescription for not writing like a machine.

Maybe it's never going to be possible.

> I do struggle at times with thinking my own writing looks like AI. But I’m an average Californian who went to college half way between SF and LA…

:-)

You guys do just sound a certain way, in the same way Brits sound a certain way to you I expect. But I think the reality is that the final stage of training LLMs was largely done in a Californian voice and with rather Californian communication objectives.

(Though equally I think much of what I am detecting is more Madison Avenue than Palo Alto)

Perhaps mistral will save us all from sounding like Californians. But sure it’s grand, you know yourself :-)

(Haven’t lived in California in a long time now)

> Perhaps mistral will save us all from sounding like Californians.

Perhaps ;-)

> This does not sit well with personal experience and I wonder if it is just one of these questions of AI people being unaware of the level of skill that exists in domains they think have been automated.

I suspect the difficulty here lies more with your reading of the quoted sentence. British grammar school education, for all the years it devotes to the enterprise, does not always succeed in teaching reading comprehension.

You seem to be treating two rather different propositions as though they were one and the same. If text in general is not sufficiently information dense to support decoding some _arbitrary_ signal of provenance, that hardly establishes that no _specific_ passage can carry distinctive markers of provenance.

For example, you can recognize the unmistakable cadence of the California undergraduate. Impressive. Alas, even in your own example, when your British friends are "giving themselves away", you resort to an external signal, beyond the text, to determine provenance! That is, unless the text itself is claiming that its author is British (like the bots who claim they're John Horsetrader from Arkansas oblast).

When you have to decide whether a 2010s era SAT essay was from a SAT prep book author or an LLM prompted to write such an essay, you will struggle to distinguish one from the other. Not all texts have provenance signals. This is what it means for text to simply not be information dense enough to be able to decode some arbitrary signal of provenance from it.

> I suspect the difficulty here lies more with your reading of the quoted sentence. British grammar school education, for all the years it devotes to the enterprise, does not always succeed in teaching reading comprehension.

Well aren't you a genuine delight?

> Alas, even in your own example, when your British friends are "giving themselves away", you resort to an external signal, beyond the text, to determine provenance!

A possibility I addressed in the actual text you are responding to, where I started the sentence with "It is of course possible" and continued to clarify that "...so many of the situations…" I encounter it are localised.

It's almost like I was expressing just such an awareness of the limits of my assertion, isn't it?

Erstwhile elsewhere whilst among the midst of the unbeknownst, someone is singing along to Taylor Swift songs I would not recognize. In a democracy of ideas, it is the lightning and not the cloud that makes it thunder. You Britons do sound smart.