Hacker News new | ask | show | jobs
by purpleidea 2 days ago
I would pay significantly more to use these models if there was a legal contract that guaranteed they weren't ever terfing them and some way to prove that.
3 comments

Don't you realize how many people are using these in production and rely on responses in very specific formats and have logging and metrics in place to detect any issues and if they were just randomly nerfing models they would cause major issues in production that tons of people would immediately notice? The easy way to prove this is to set up your own evals and then run them daily
https://marginlab.ai/trackers/claude-code/

Has been pretty accurate when pre-release quality or bugs pop up. In theory you could run your own.

What?
nerfing*