I certainly agree. As we lose the ability to properly "test" these AI systems, we must focus more on how they behave. A system can ace every benchmark in existence and still be profoundly untrustworthy. I'm interested to hear from others if there are benchmarks that are immune to reward hacking or saturation.