Hacker News new | ask | show | jobs
by himata4113 28 days ago
the classifier is very picky, working with languages such as C I get cyber refusals and random reasoning_extraction errors.
1 comments

It's gone from bad to unbelievably bad. Just a prompt with the single word "virus" was enough to get downgraded to O4.8 before the ban.

"Tell me a story about a man named CVE." gets me downgraded now.

"CVE-" gets me downgraded and broken reasoning with no response.

"Giant crane fly, 2 feet wide" gets me downgraded.

Is it likely that it keeps track of what you said recently and decides whether you might be referring to things you've previously been downgraded for?

I haven't been downgraded once, either a few weeks ago for the three days it was live, or since I got it back today.

>Is it likely that it keeps track of what you said recently and decides whether you might be referring to things you've previously been downgraded for?

I have no special knowledge here, it feels rather unproductive for me to speculate.

Out of curiosity, if you're comfortable trying any of them, do any of the above prompts cause you to get downgraded?

Good question! I'll give them a shot when I'm back to my computer.
I went last to first, and I was not blocked.
Thanks for trying and for sharing results here. This is a pretty interesting data point and suggests something I don't think I've read about: the possibility that safeguards may not be account-invariant; that two different users might be getting downgraded or blocked differently for the same prompts.

This raises some really interesting ethical questions in my mind, but I suppose I need to do more reading and research on this before anything else.