This is why the world model approach is so important. It allows you to feed back the prediction accuracy of the model to itself at training time, enabling it to predict (to some degree) its own uncertainty. If you jump through a couple of hoops you can also do this at run time to give it “spidey sense” that something’s not right with current inference.
RMSE is just an extrapolation from the training data. If the data is wrong because the world changed, any model (parametric or not) can be confidently incorrect.
I gave up on Grok. It is going the way of Tesla, SolarCity, gigabattery and Autopilot. Now on GLM5.2 via Open router hosted in SG. Mimo is also good. Their agent is so convenient and Deepseek level cheap. Quality a bit behind GLM5.2. But then Mythos is myth, technically GLM is high up there in quality but on lower end pricing.