Pangram's marketing always reminds me of Anchorman's Sex Panther cologne: "They've done studies, you know. Sixty percent of the time, it works every time."
Did you read the paper you are linking to? It records a 0% false positive rate for evaluation of human-authored controls, and fewer than 5% of the hybrid and humanized papers had their AI levels overestimated by pangram. It is not completely clear from the data, but it seems that depending on whether the n=2 overestimates were “100% ai generated” assessments, the study you link says pangram’s 100% AI assessments were right anywhere from 95-100%
My citation was correct. To question to ask yourself: For my use case, is it okay that Pangram can't reliably tell the difference between "100% AI generated" and "AI assisted"?
Your citation was incorrect. The data also shows that pangram reliably tells the difference between those two cases. The statistic you made up or hallucinated is not supported by your link, and you’re being disingenuous here.
Thank you for pointing that out. But it still doesn’t support what you said, which is that only 65% of pangram’s fully AI claims are accurate. The “65% strict accuracy” rating is followed in the same sentence by a “97.5% inclusive accuracy” rating for the same reasons I pointed out above: pangram is systematically underrating the percentage AI in content, and it is tuned against producing false positives. The 32.5% in between the inclusive and strict accuracy ratings is entirely due to it estimating that 12 fully ai generated works were 60-80% AI, not the reverse. You can see that in table 4.