|
I wanted to be sure I wasn't lying to myself. So, I built this: https://prose-or-con.com/ I can still detect AI-written prose with pretty high confidence (I average about 85% accuracy after a bunch of rounds while testing, though I'm rushing because I'm testing), but it has gotten harder. Opus 4.8, GLM 5.2, GPT 5.5, and DeepSeek v4 Pro, are able to fool me sometimes. Mistral writes so weird it almost trips me up, but then I go, "oh, that's just Mistral". Qwen sometimes accidentally switches to Chinese (filtered out of the game, just a thing I found in testing). I'm doubling the corpus, currently, with a focus on the frontier models that have been able to fool me more often, so it'll probably get a little harder. GPT 4o is disgusting, just horrible saccharine nonsense (and, I guess that's the model that's caused the most psychosis with its sycophancy). And, I have to concede that the best models prose has gotten better, especially when they're prompted to write in a specific style, like literary fiction or an encyclopedic article about a specific topic. Still flowery and verbose, though, and never really getting to the point (the AI doesn't want anything, and it shows). |