|
|
|
|
|
by Tade0
11 days ago
|
|
The article itself explained that it was much easier to classify text as human or LLM generated than to have "human" as just a category along with all the different LLMs as it's likely the LLMs are distilled from each other, creating a unique footprint. If a signal is weak, it might not even appear in every sentence, but that doesn't mean it doesn't exist. For instance, I don't recall ever consciously using an em dash, but you'll probably need an entire paragraph to find one in LLM-generated text. My own sense of whether text is generated is partially based on its sheer length - humans typically don't bother writing so much. |
|