Hacker News new | ask | show | jobs
by akersten 12 days ago
Text is simply not information dense enough to be able to decode some arbitrary signal of provenance from it. Sure you might be able to detect today's tells (particular sentence structures preferred by Claude, phrases, etc) to get you some arbitrary chance percentage it was machine generated, but it's a bad fiction to perpetuate that any of this is anything more than tarot card reading.

Images, absolutely, there are tell-tale artifacts from today's generators that simply aren't emitted by "natural" paths to create them, and you can "detect AI" with high confidence (for now). Words, no, the signal is far too sparse and we are well into undetectable sophistication with today's models, let alone tomorrow's.

21 comments

The article discusses a technique by which the author achieves high accuracy at detecting AI written text. Unless you have a problem with their experimental method, this is the opposite of tarot card reading.

> we are well into undetectable sophistication with today's models

The article directly contradicts this, as do you, in your previous paragraph: "Sure you might be able to detect today's tells". The article is literally about a technique that detects today's tells.

Your comment is mostly expressing doubt that this technique will work reliably in the future, but it's framed as opposition to the article, which it's not: the article is about detecting today's AI-written text, at which it seems to be quite successful.

It does achieve high accuracy but I think given the context when one wants to know this information, plagarism for research papers and college/highschool essays and work, it's unfortunately not good enough.

My neighbour is a teacher. She has a really good idea which of her students uses AI to do their homework but 80% accuracy is not good enough. She'd need to be able to prove it with certainty.

The context in which the author wants to know this information is that they enjoy reading fiction, but not low-effort AI fiction, so they built a tool to filter out some of the low-effort AI fiction. Not every application is such a high-stakes affair that anything less than perfect isn't good enough.
not really, a strong suspicion is enough to motivate assigning an extra paper and pen in person test to a student, and then you can fail them on that result.
That seems pretty unfair. Why not make the original test pen and paper then? (Or at least a typewriter, offline computer, etc - my handwriting is awful)
Originally the idea was that a take-home essay would allow students to work at their own pace, study their own way, and produce something interesting. But if more than half will just prompt an AI and learn nothing, then I agree you should proctor all assessments.
Whether a text was written by a human or not is just a single bit of information. So you can't rule out its detectability a priori, since even the shortest text contains more information than that.

As long as LLMs are used to write texts humans wouldn't want to write if they could help it (that's why they're getting an LLM to do it, after all), they'll remain detectable. Even if the reasoning might end up equivalent to "This looks like spam; no human in their right mind would write this spam by hand if they could get an LLM to write it, therefore it's most likely written by an LLM."

That's like saying whether or not you're going to fall in love this year is just one bit of information, so you might be able to read it from astrology. Yeah, sure, it might happen for some people with a certain star sign. But across the population there is zero reason to believe that there is a) any significant correlation and b) enough data variation in to even distinguish classes of humans.
Indeed you cannot rule out astrology on information-density grounds. Astrology involves quite a lot of information, the problem is that it's mostly unrelated to the outcomes of interest. To get back to the information-density of text, "I love you" doesn't contain a lot of information, but it does contain the one bit you care about, because someone who loves you is more likely to say it than someone who doesn't.

So if you want to determine whether something was written by a human or by AI, to do better than chance it's enough for there to be a difference in the probabilites of a human writing it and AI writing it, respectively. Whether the resulting accuracy is good enough for a particular use case is another matter. 99% is pretty good odds for love and pretty bad odds for "am I going to survive today?" Hopefully there won't be a death penalty for posting AI-generated content.

The 80% accuracy from the article would be one reason to believe there's significant correlation, no?
I could probably find quite a lot of people who will tell you astrology is 80+% correct for them. Would you believe them or wait for an independent analysis? There are other AI "detector" systems out there that claim 99% accuracy. But independent research always found that they are actually garbage once used on real data. It's all in how you pick your tests. It's also funny to see how people on places like HN will easily dismiss stuff astrology, but fall for the exact same patterns when used in tech-y applications.
Well I don't think the position of planets when you're born has a large correlation to how your life will go.

On the other hand, how an AI writes will have a big correlation to whether the written text would likely be written by an AI.

The latter is more of a direct relationship.

Maybe A and B are not correlated, and Y and Z are? What pattern are people falling for here?

It's not really about the planets. It's about what other people believe about the planets.

You could work for a boss that's a Leo, and he/she believes only Leos deserve to get promoted, or that Leos and Scorpios should never be assigned together on a shared project. Your life and career trajectory under this boss could be totally different, depending on whether you were born a Leo or not.

Certainly "not all bosses" applies, but it's not really that farfetched or uncommon either. The point is that a correlation does exist, but it's a social one and not a physical or astronomical one (and it's also often a self-reinforcing/self-fulfilling one: Leos who read what astrology says Leos should do, may end up choosing to behave more like that).

So in the LLM example, it may not really matter much what physical markers of provenance or physical correlations there are, as social beliefs or perceptions about suspected provenance may be the strongest correlation anyways (in terms of impact and outcomes).

> Whether a text was written by a human or not is just a single bit of information

I doubt this models reality well at all. If I write the first paragraph, and AI writes the second; a float seems to model that better. If you choose to collapse a float into a bool, I don't think you can make useful conclusions based on that bit?

> since even the shortest text contains more information than that.

I also don't think that's how information theory and bits of information works...

> Whether a text was written by a human or not is just a single bit of information. So you can't rule out its detectability a priori, since even the shortest text contains more information than that.

This is word salad, a complete non sequitur.

> to write texts humans wouldn't want to write if they could help it (that's why they're getting an LLM to do it, after all)

Er, that's obviously not true.

Not all humans are in their right minds, unfortunately.
This is exactly the point I saw in a recent x post, that building anti-bot detection was incredibly difficult because some people exhibit bot like behavior.

Blizzard employee once told me anti-botting in WoW was extremely challenging due to the number of real people that acted identically to bots.

Every assumption was invalidated: - unbelievable # of consecutive hours played - consistently repetitive patterns of movement and clicks - farming patterns that aren’t considered fun (“why would anyone do that”) - solo, no external engagement - goes on for months

The problem with botting is many humans ARE bots

https://x.com/IceSolst/status/2076372992959959493

How do they know those were real people? Were they livestreaming their face and talking about what they were doing the whole time?
It is much harder to tell one from the other, and for oneself, than it often seems on the surface.
>As long as LLMs are used to write texts humans wouldn't want to write if they could help it (that's why they're getting an LLM to do it, after all), they'll remain detectable.

Come on, that's circular reasoning.

You definitely can rule out the general case a priori. If the problem were possible, for every text there would be a unique provenance label “human” or “ai”. But since humans and machines have both written many texts, it is not possible.

As an example, you could imagine a giant lookup table that deterministically mapped every text ever written to “human” or “AI”. You would very quickly run into situations where the labels conflict for the same piece of text.

The data is statistically inseparable which makes it impossible to classify from text alone.

That just proves that perfect classification is impossible. Classification doesn't need to be 100% accurate to be useful.
It’s worse. If the data was separable in this way, you would equally be able to train an AI to mask those signs.
IIRC, the big names in LLMs have no real interest in cloaking the LLM-nature of the text, Google adds deliberate watermarks to text, OpenAI developed a watermark for text but reportedly arent't actually using it.
Considering [1], I’m going to challenge that their techniques are currently even mildly effective. Given the absolute academic malpractice these papers are pushing, I’m calling BS; while they want to watermark it, they clearly aren’t actually able to. For images. Which are drastically easier than text.

Their interest is irrelevant in the face of technical impossibility. And that’s before you get into other people who don’t care and will just build adversarial tools to bypass the attempted watermarks. It’s a losing useless battle. Google and OpenAI engage in it to try to catch competitors when there’s a lawsuit or to try to clean their datasets clean.

But it’s absolutely unusable for something like “did someone cheat”.

[1] https://hackerfactor.com/blog/index.php?/categories/1-Image-...

> Their interest is irrelevant in the face of technical impossibility.

I'm responding to "If the data was separable in this way, you would equally be able to train an AI to mask those signs.": yes, if you wanted to you could, the big names clearly don't consider masking to be a priority.

> But it’s absolutely unusable for something like “did someone cheat”.

This is the one case where I'd most expect it to succeed:

I suspect most of the people who do want to cloak-to-cheat, don't have the skills to do so; I also suspect most of them are so unaware of what they don't know that they won't even ask an LLM to write cloaking software for them.

> But since humans and machines have both written many texts, it is not possible.

Maybe you meant "many humans have used AI when writing texts"? Your stated reason that they can't be separated because there are many texts of each kind is nonsensical, you clearly need to supply more reasons than "there are many".

"Text is simply not information dense enough to be able to decode some arbitrary signal of provenance from it...it's a bad fiction to perpetuate that any of this is anything more than tarot card reading."

Not true at all. Pangram is highly effective and has a very low false positive rate.

The post here is impressive for a small project, it looks like they independently thought of one of the core ideas Pangram uses of creating twins to compare.

You can see how it works here: https://arxiv.org/pdf/2402.14873

So, if the decision from Pangram determined, on every assignment, if you would be expelled from university for plagiarism, would that be acceptable to you regardless of how you actually did the work?

If you would not be okay with that, what level of consequence would be acceptable for the output from this tool?

Even if Pangram was blessed by God to be 100% accurate no, your argument is a strawman. The reliability of the software has nothing to do with the principle behind "software should never make a management [legal / disciplinary / etc] decision." So no consequence from the tool, but perhaps it can be used as evidence in an academic integrity hearing. Maybe the university equivalent of probable cause. I am not knowledgeable enough to make a firm determination.

FWIW if I were a student I would definitely be using Track Changes or version control, etc etc, to make clear my work was human-written. Which sucks.

That’s a different point.

I’d want detectors to be as accurate as possible, false positives of 1 in 10000 seems like a good starting point. I believe their results have been independently tested.

And as a separate matter, any tool for evaluating students should be applied fairly, safely, and with adequate human review and due process.

You need good tools and good oversight.

Due process should never just become a checkbox item. To deal with lives and livelihoods justly, you need appeal pathways and meaningful liability exposure for the processors.

Plagiarism and cheating sucks for everyone. Worth solving.

>And as a separate matter, any tool for evaluating students should be applied fairly, safely, and with adequate human review and due process.

Agreed, that's a fair and reasonable stance.

The reason I asked is that I have a hard time understanding the point of these tools. When it comes to education, it can be a matter of learning objectives. But outside that, what's the point?

The prediction from the tool is pointless for deciding on copyright or contract issues, and other text should be judged on its correctness or applicability to the task.

If all the tool is good for is "maybe this student cheated, but only an in-depth investigation would maybe prove it", it isn't a very useful tool, because it's more straightforward to just mandate that evidence is submitted regardless of what the tool says. On top of that, even the lack of evidence of manual work isn't good proof of using LLMs.

I’m personally interested in it as part of the research to improve LLM writing. Detecting “AI voice” is part of understanding what’s wrong with it in the first place and how to improve it.

But yeah, in general I think you’re right, the actual utility is pretty niche.

>>> "Text is simply not information dense enough to be able to decode some arbitrary signal of provenance from it...it's a bad fiction to perpetuate that any of this is anything more than tarot card reading."

>> Not true at all. Pangram is highly effective and has a very low false positive rate.

> So, if the decision from Pangram determined, on every assignment, if you would be expelled from university for plagiarism, would that be acceptable to you regardless of how you actually did the work?

What point are you arguing? Something having a high success rate does not necessarily translate to treating it as a 100% success rate.

I explained what I was thinking about when asking that here: https://news.ycombinator.com/item?id=48940887
There are two problems, false positives and changing the LLM's pattern.

It's really easy to have a false positive and false positives can be very harmful if the person using the detector isn't aware of that risk.

It's also very easy to change the pattern of LLM output. You can provide basic prompting that will significantly change the structure of the output. For example, having it utilize the Wikipedia article on signs of AI writing and avoid everything it describes. https://en.wikipedia.org/wiki/Wikipedia:Signs_of_AI_writing

"It's really easy to have a false positive"

Not really. The false positives for the SOTA detector are very very low.

"It's also very easy to change the pattern of LLM output."

Not in a way that can reliably avoid detection. The problem is the patterns are baked into the distribution itself. It's smoothed over, so it becomes difficult to prompt your way out of that.

Wrong. Effective sampling (I.e high temperature like temp 10) with the corresponding sampler stack that enables this coherently destroys all attempts to detect it. There are many more ways like this involving manipulating the logprobs
I’m not sure what you’re saying I’m wrong about.

The comment I was responding to, about changing LLM output, referred to prompting, not temp/sampling tricks. I’m not aware of Pangram being beat by clever prompting. There’s some interesting work on creative writing using contrastive prompt techniques, but I haven’t seen it tried as evasion.

Even if you control temp and sampling, they’re not magic. If you raise the temperature too much writing can go to hell, so you may beat the detector but end up with junk. There are some ways to mitigate such a quality drop like raising temperature in conjunction with min-p, but still, I haven’t read any research that shows it getting good results at anything close to 10.

Now you want to get more clever and manipulate logprobs…well ok, you could come up with elaborate strategies designed to evade specific detection methods. But I don’t see that getting done as a weekend project while maintaining writing quality. And if it does happen there’s no guarantee the detector can’t train on its characteristics and start an arms race.

As min-p approaches 1, the temperature you can get away with approaches infinity. Also more modern samplers like top-n-sigma are explicitly designed to get away with temperature of infinity.
Not convinced just using top n sigma is going to beat a SOTA detector. The fingerprints Pangram uses should in principle be able to detect style above the token selection level.

Other papers have tried to beat it with temperature and it didn’t work, although I haven’t seen anyone try insane levels.

Give it a shot and let me know if you have any success.

With sufficient information you can derive a signal even in the presence of overwhelming noise. Assuming the noise is not perfectly correlated with the signal this is always possible.

Schemes like GPS, CDMA and DSSS are based upon this concept. GPS in particular is quite impressive in its ability to recover information that is received below the thermal noise floor.

There has to be a signal to detect it.

Take this sentence: Bob went to the store to buy milk.

Was that AI generated or not? There simply isn't a signal there. The problem isn't noise, the problem is, is there even a signal to begin with.

Sure, you might be able to recognize the quirks of a specific LLM just as you recognize the quirks of a particular person, but as the number of LLMs proliferate, then the signal turns into noise. (The signal isn't buried by noise, it becomes noise. The signal no longer has any discriminating power.)

The article itself explained that it was much easier to classify text as human or LLM generated than to have "human" as just a category along with all the different LLMs as it's likely the LLMs are distilled from each other, creating a unique footprint.

If a signal is weak, it might not even appear in every sentence, but that doesn't mean it doesn't exist. For instance, I don't recall ever consciously using an em dash, but you'll probably need an entire paragraph to find one in LLM-generated text.

My own sense of whether text is generated is partially based on its sheer length - humans typically don't bother writing so much.

but.... the LLMs are actually all trained on approximately the same stuff, and tend to have similar quirks. In the human world, writers develop recognizable voices, which are detectable and classifiable (as in the article we have all supposedly read). Furthermore, we don't necessarily care about telling one LLM from another, just that they aren't human. That's different from trying to identify one human amongst a sea o fhumans, or one bot form within a sea of bots.
Statistical power comes from having many samples. I agree that having just one sample doesn't take you very far.
Signal is easier to detect with more data to work with.

Largely AI generated books are a vastly different situation than a one paragraph homework assignment. But multiple rounds of homework assignments would change the accuracy.

> Text is simply not information dense enough to be able to decode some arbitrary signal of provenance from it.

This does not sit well with personal experience and I wonder if it is just one of these questions of AI people being unaware of the level of skill that exists in domains they think have been automated.

It is of course possible that my tendency to spot LLM-written text has much to do with the way that it sounds like an averaged Californian college student to my British grammar-school-educated ears, as so many of the situations where I am encountering AI text are Brits using it without apparently realising they are giving themselves away.

But I know people who don't have particular technical skills in this sphere or a grammar-school background who also have an uncanny knack for pointing out LLM-written text.

> Words, no, the signal is far too sparse and we are well into undetectable sophistication with today's models, let alone tomorrow's.

I especially don't think this is true. Will they be able to do it in the future? Maybe. Is it possible to prompt a current cloud LLM to write in a way that is obvious? Yeah. (IMO Gemma 4 writes less detectably than most of them!)

But my instinct is that someone with any facility for language is going to be better than chance at spotting LLM-written text once it is three or four paragraphs long. So I think it should be possible in principle to train machine learning systems to detect those patterns.

If you can train a system to detect these patterns, presumably you can train systems not to generate text which matches them?

I do struggle at times with thinking my own writing looks like AI. But I’m an average Californian who went to college half way between SF and LA…

> If you can train a system to detect these patterns, presumably you can train systems not to generate text which matches them?

I don't know. I mean, it feels like the systems that would detect them are likely qualitatively different to the machines that make them.

One of the things that feels obvious to me is that LLMs are always going to write in a new way, because words do not get all that close to perfectly conveying the inner thoughts of competent writers. Competent writing is always a battle to find the better word, or even to create it.

So sure, you could add another adversary that the generator has to satisfy, but "this sounds like a machine wrote it" is only an observation; it's not a prescription for not writing like a machine.

Maybe it's never going to be possible.

> I do struggle at times with thinking my own writing looks like AI. But I’m an average Californian who went to college half way between SF and LA…

:-)

You guys do just sound a certain way, in the same way Brits sound a certain way to you I expect. But I think the reality is that the final stage of training LLMs was largely done in a Californian voice and with rather Californian communication objectives.

(Though equally I think much of what I am detecting is more Madison Avenue than Palo Alto)

Perhaps mistral will save us all from sounding like Californians. But sure it’s grand, you know yourself :-)

(Haven’t lived in California in a long time now)

> Perhaps mistral will save us all from sounding like Californians.

Perhaps ;-)

> This does not sit well with personal experience and I wonder if it is just one of these questions of AI people being unaware of the level of skill that exists in domains they think have been automated.

I suspect the difficulty here lies more with your reading of the quoted sentence. British grammar school education, for all the years it devotes to the enterprise, does not always succeed in teaching reading comprehension.

You seem to be treating two rather different propositions as though they were one and the same. If text in general is not sufficiently information dense to support decoding some _arbitrary_ signal of provenance, that hardly establishes that no _specific_ passage can carry distinctive markers of provenance.

For example, you can recognize the unmistakable cadence of the California undergraduate. Impressive. Alas, even in your own example, when your British friends are "giving themselves away", you resort to an external signal, beyond the text, to determine provenance! That is, unless the text itself is claiming that its author is British (like the bots who claim they're John Horsetrader from Arkansas oblast).

When you have to decide whether a 2010s era SAT essay was from a SAT prep book author or an LLM prompted to write such an essay, you will struggle to distinguish one from the other. Not all texts have provenance signals. This is what it means for text to simply not be information dense enough to be able to decode some arbitrary signal of provenance from it.

> I suspect the difficulty here lies more with your reading of the quoted sentence. British grammar school education, for all the years it devotes to the enterprise, does not always succeed in teaching reading comprehension.

Well aren't you a genuine delight?

> Alas, even in your own example, when your British friends are "giving themselves away", you resort to an external signal, beyond the text, to determine provenance!

A possibility I addressed in the actual text you are responding to, where I started the sentence with "It is of course possible" and continued to clarify that "...so many of the situations…" I encounter it are localised.

It's almost like I was expressing just such an awareness of the limits of my assertion, isn't it?

Erstwhile elsewhere whilst among the midst of the unbeknownst, someone is singing along to Taylor Swift songs I would not recognize. In a democracy of ideas, it is the lightning and not the cloud that makes it thunder. You Britons do sound smart.
> but it's a bad fiction to perpetuate that any of this is anything more than tarot card reading.

Hard disagree. LLMs (especially base ones, that only received pre-training) can produce output that is undistinguishable from human writing (because that's what they were trained to do).

But commercial chat models are specifically tuned in a way that maximizes user engagement. It's that specific tuning that is very easy to spot when reading AI slop, and that's not surprising that it's easy to spot automatically either. And I don't think that's going to change anytime soon, unless their incentives change.

(We can say exactly the same thing about man-made stuff optimized for a specific purpose, like stock photography, clickbait titles or industrial food: they aren't stereotypical because their creator lacks the skill to make them otherwise, they are like that because that's what works best).

They're also designed to not offend anybody, so their output tends to be very bland even compared to the most milquetoast of human beings. I was only surprised once when ChatGPT responded with an enthusiastic "hell yes" seemingly organically, but 99.9% of the time these AI services clearly are instructed and trained to provide flavorless word vomit. I don't think there's a technical reason why an LLM couldn't produce totally convincing output, but internet grifters don't need to go through that trouble. It's like how most phone, email, and social media scams come off as completely transparent to most of us, but that's the whole point; we're not the target audience of the scams. Readers looking for substance, nuance, and real opinions aren't going to notice if something with written by an LLM – unless there are some cliche punctuation tells.
When DANmode bypasses were a common thing the LLMs would drift significantly far from corporate speak.

But that's the point of corporate speak, you tend not to say thing that may offend your clients and deprive the company of future revenue. Of course there are some companies that make their living being 'counter-culture' and saying what they want, but they are a small percentage of all revenue.

> especially base ones

Did you actually try them? I did.They generated even more "slopey" text than instruction-tuned ones.

But was it content indistinguishable from someone learning the language being used, for instance?
It does mean that this will have a drift problem if it's just trained on the idiosyncrasies of model fine tuning. That's fine! But it is something to be aware of.
> But commercial chat models are specifically tuned in a way that maximizes user engagement. It's that specific tuning that is very easy to spot when reading AI slop, and that's not surprising that it's easy to spot automatically either.

There are two problems with this.

The first is that it would still misclassify human-authored text written under the same incentive, and most people have various incentives to "maximize engagement".

And the second is that then people would just make other models that are tuned for defeating that sort of classifier, which would be used whenever the classifier is being used.

All of that may be true, but pangram currently has a false positive rate of about 1 in 10000, and this has been tested by feeding in thousands of texts written before 2020.

That may not last if AI companies start trying to build models that fool it, but for the time being at least, modern models do have strong tells.

>and this has been tested by feeding in thousands of texts written before 2020.

And these text didn't train the model in the first place? I just want to ensure clarity on that.

>pangram currently has a false positive rate of about 1 in 10000

Says Panagram.

The problem with just looking at old text is language is a living thing. Say for example I make up the world 'oklambroahaha' right today. Both humans and AI pick up that word and start using it. Now lets say the model says that anything that uses oklambroahaha is 100% AI, you can't just point and say, "well my detection AI is correct on things 20 years old, so it's right skibbidy toilet 6/7".

There is a ton of evidence that use of AI changes the way we speak and write, so it will just turn these AI detectors into bullshit generating classifiers.

You can get an arbitrarily low false positive rate by sacrificing against false negatives. It's trivial to make it zero, just classify everything as human-generated. Meanwhile a false negative rate of even 1% is a pretty big problem since someone can easily use LLMs to generate 100x the volume of text and then use whichever ones make it through the classifier.

And that's before anyone even tries to get the LLM to generate a different style of text. Or for that matter creates a "style model" that rephrases text.

You don't really need a style model - current models are very good at doing "style transfer" of a model text onto whatever it has written if you just have it do it chunk by chunk. It takes more to prevent it from being detectable by good detectors, but it does remove a lot of the worst tells.
The point being that you wouldn't need the developers of the most popular models to themselves be trying to fool classifiers because their output could be run through an independent special purpose one designed to remove the tells the classifier is looking for, and the special purpose one wouldn't need to be made by anyone with the resources to create a good general-purpose model since it only has to do that one thing.
Pangram won't know how much AI written text they fail to detect, though, and detectors is a great tool to adjust methods of generating less AI-sounding text.
> The first is that it would still misclassify human-authored text written under the same incentive, and most people have various incentives to "maximize engagement".

The thing is, humans are significantly worse at maximizing numerical goals than computers.

> And the second is that then people would just make other models that are tuned for defeating that sort of classifier, which would be used whenever the classifier is being used.

Anyone can already do that right now, just grab unsloth studio and fine-tune your local Gemma, but nobody cares. People posting slop content don't care if pangram or I flag their slop with certainty, they are using the easiest option, which is commercial chat models. And given this segment of user doesn't care, the provider have zero incentive to provide a dedicated stealth model for that purpose.

> The thing is, humans are significantly worse at maximizing numerical goals than computers.

I'm not sure this is even the right premise.

Existing LLMs try to maximize engagement, and they often write in a particular style that has tells, but these two things are not necessarily related. Over-using em-dash or whatever isn't the thing that maximizes engagement.

So the two problems really are, what happens to the actual humans whose writing style is a close match for what a given generation of LLMs output? And, what stops LLMs from using a different style when someone wants to fool the classifier?

> People posting slop content don't care if pangram or I flag their slop with certainty, they are using the easiest option, which is commercial chat models.

They don't care as long as the consequences of identifying it are immaterial, but in that case what's the point of classifying it? Whereas if they need to fool the classifier some threshold percentage of the time in order for enough of their spam to get through, they're going to care.

> Over-using em-dash or whatever isn't the thing that maximizes engagement.

It's the thing that minimizes the loss during the RLHF phase, and the RLHF phase is the one that's aimed at maximizing engagement (it's literally trained on that).

> what happens to the actual humans whose writing style is a close match for what a given generation of LLMs output?

If a human, for instance because its writing gets polluted by reading too much AI slop, matches the style of an LLM closer than a certain threshold, then his own writing is going to be flagged as well. Whether it's an actual problem or merely a theoretical one is an open question. (unlike OpenAI and Anthropic, humans writers do have an incentive to avoid being flagged as AI).

> And, what stops LLMs from using a different style when someone wants to fool the classifier?

In theory: nothing. In practice if you fine-tune your own model: nothing. In practice with commercial models: the interests of the model making company.

> And, what stops LLMs from using a different style when someone wants to fool the classifier?

Websites have pretty much stopped using ad-blocker-blockers, it seems that it's not a fight worth fighting for them. Does that mean that ad-blockers are useless?

Most people don't even care about ads, I don't think they care about slop either, that's why there's slop posts and obnoxious websites that are unreadable without an ad blocker. A slop blocker used by 10-20% of the internet users wouldn't change the calculation more than ad blockers did.

> It's the thing that minimizes the loss during the RLHF phase, and the RLHF phase is the one that's aimed at maximizing engagement (it's literally trained on that).

I don't think RLHF is the biggest reason its style is the way it is.

A lot of it is that it's trained on everything they could get their hands on, which includes domain-specific literature and books that go all the way back to the advent of writing, and then will pick up habits that are common in some specific domain or in 19th century literature etc. that are less common in most modern writing when no attempt is being made to do otherwise.

Do you really think that RLHF humans were requesting more em-dash?

> Websites have pretty much stopped using ad-blocker-blockers, it seems that it's not a fight worth fighting for them. Does that mean that ad-blockers are useless?

Websites have pretty much stopped using them because they realized readers with ad blockers will stop using the site sooner than stop using their ad blocker, and since websites have a network effect, it's better to let a minority of readers block ads when having them makes it more likely they'll distribute links to the site. And because it's the user who controls the browser for web pages, which gives ad blockers a decisive advantage.

> Most people don't even care about ads, I don't think they care about slop either, that's why there's slop posts and obnoxious websites that are unreadable without an ad blocker. A slop blocker used by 10-20% of the internet users wouldn't change the calculation more than ad blockers did.

Sites don't want users to use ad blockers, but having a user with an ad blocker is still better for them than not having the user at all, because of the network effect.

Whereas many sites don't want slop at all, and then if slop detectors work they'll put them in the site itself and block the slop for 100% of users. At which point the slop generators have a 100% incentive to find a workaround instead of a 20% incentive, which is different.

I mean, back when I was spam filtering setting up a simple Bayesian classifier was easy. Train it on your spam and ham and it worked damned good. "Mission Accomplished".... until it wasn't. Spam rates started climbing and it started getting harder than ever to filter them.

There is always an incentive to get spam to bypass filters, so as your filters increase in accuracy, those attempting to pass said filters adjust their behaviors.

Spammers/cheaters/whateverers will at least just use a second pass filter that uses one of these 'ai scoring' systems to beat said AI scoring systems. So while it's worthwhile to do it at this moment, this window will rapidly close.

I don't think it's a very good remark, as there's significantly less email spam than 20 years ago.

Another example is ad-blocker-blocker. There was a little bit of an arm race between ad blockers and advertisers in the middle of the 2010s, but it didn't last long. Advertisers mostly just decided not to care about ad-blockers.

>Advertisers mostly just decided not to care about ad-blockers.

Directly not to care because they lost in court.

And yet the biggest advertizer on Earth (Google) decided to change their browser to make adblocking far more difficult. That or they say "just use an app, oh and turn on notifications". I'm not exactly sure who you think won the arms race there, but it seems like we the user did not.

There is significantly more spam than 20 years ago, just less of it reaches your inbox. This is a very important distinction as the cost of spam filtering is just as high as ever. On top of that most people have given up on their own email servers and instead depend on Google/Microsoft to do it for them. This allows these companies to have an overwhelming influence on email on the internet, to the point they can send spam with near impunity, and where if your system does it will be nuked from orbit by their systems.

And much like now Google supplies both the email spam, and the solution to the spam, they'll gladly supply the LLMs spam and the LLM solution while applying their 'flavor' of what's allowed to the entire internet.

> Text is simply not information dense enough to be able to decode some arbitrary signal of provenance from it. Sure you might be able to detect today's tells (particular sentence structures preferred by Claude, phrases, etc) to get you some arbitrary chance percentage it was machine generated, but it's a bad fiction to perpetuate that any of this is anything more than tarot card reading.

This is simply untrue, and completely divorced from reality.

Tarot card readings have literally zero predictive success. Last I checked, LLM-detection had a +90% success.

Sure it is; we do it all the time, and then we modify each other's etc, etc; english we speak today was spoke yesterday waspake the same in yesteryears; we have no trouble dating english or other languages to a time.

A better argument is people themselves are just too influenced by reading that they'll sound like LLMs in a couple of years.

i think one thing overlooked by this perspective is that many of a detectors adversaries are not that sophisticated. so despite this i think it is a useful thing to try to do. particularly when people are trying to do fraud which will often having to use abliterated models and generally trying to be as economical in their efforts
So you’re saying that the linked article’s findings are implausible? Is the article fake, then, in your opinion?
If you have access to the detector, you can formulate a generative solution that avoids being flagged. Which gets me wondering why don’t model providers do that? There must be something about that that destroys semantic weights somehow.
Why would sounding human be a goal rather than a byproduct of trying to communicate efficiently?
That is manager/executive/manager speak, real people don’t speak like that (unless they are in the aforementioned roles).
I just feel like, from the POV of AI companies, that reducing the amount of em dashes they use to "blend in" more and talk more human-like, for the sake of being less detectable, wouldn't be a big priority.
One of the big use case is cheating on your essay assignment.
See the discussion on https://news.ycombinator.com/item?id=48837460

You can absolutely still tell.

It depends on how much text. For example, chardet often falls down on short strings, but 1K characters it nails it.
> ... but it's a bad fiction to perpetuate that any of this is anything more than tarot card reading

Most people's issue with AI-generated llmish however is not that it's AI-generated. It's its insufferable tone.

So if we get to a point where we have to read tea leaves (an image you seem to appreciate) to determine if it's llmish or not, we'll have won by then.

Really: it's that full-on asshole tone I (and many others) want to see disappear from blogs, comments, LinkedIn, etc.

My broad feeling is that if we generally independently identify the insufferable tone, and we absolutely can, so can a sufficiently trained machine-learning model.

It's an aside, but my biggest problem with trying to get up to date and learn about LLMs is how much of the documentation, blog writing, and tutorial material has obviously been written by no-one. It is just so much harder to read (and, like generative AI slop generally, curiously much harder to recall later).

obviously, a universal model doesn't exist since the signals are non-stationary but it's way better than what tarot reading
I don't know, the thing about most text slop is how little effort goes into disguising it (for now, anyway). I'm sure anyone dedicated can go undetected, but it's the really low-effort stuff that's generally the problem. If you can catch some of it, that's something at least.
pow(n,m) where n is alphabet size and m is number of characters is very dense.
This sounds like it was edited by an llm.
The best method is, as always, an anti-privacy method.

Simply track all citizens' writing patterns throughout their life, from cradle to grave, then diff with any given text's signature--you'll know if it was human written or not.

Better--opt in--install a "personal text signature" on your devices, sign things that you wrote yourself with it.

But I suppose that's just like the image provenance chips on cameras.

Either way father fascism is more with us than ever, praise him!