Arena can definitely be benchmaxxed a bit, if you try. The distribution of prompts there is very different than usage by regular coders. E.g., lots of requests for one-shot games from scratch. So if you fine-tuned your model to be great at making fun one-shot games from underspecified prompts, your coding model might look better than it is (on general tasks, at least).
I work at OpenAI, and am happy to say we don't try to juice our scores here, as doing so would be counterproductive and make Arena a worse signal for everyone.
There's a ton of arenamaxxing going on (especially from facebook), though I don't disagree that it's one of the better actual benchmarks.
Always fun to ask them to recreate classic demoscene effects (sadly they're still pretty bad at generating music, though at least claude seems to create decent synths).
I keep trying to get them to recreate the fluid+particle stuff from Agenda Circling Forth etc., but even giving them the blog posts describing the implementation (and screenshots) they're still pretty bad.
I work at OpenAI, and am happy to say we don't try to juice our scores here, as doing so would be counterproductive and make Arena a worse signal for everyone.