|
|
|
|
|
by weird-eye-issue
2 days ago
|
|
Don't you realize how many people are using these in production and rely on responses in very specific formats and have logging and metrics in place to detect any issues and if they were just randomly nerfing models they would cause major issues in production that tons of people would immediately notice? The easy way to prove this is to set up your own evals and then run them daily |
|