The ExploitGym paper evaluated several frontier models on the bench and reported that "Different models find different exploits" [1], so it seems most plausible that the "test solutions directly from Hugging Face’s production database" [2] which GPT-internal found were authored by Mythos (or some other LLM with complementary strengths), and placed in some internal HF repository when creating the ExploitGym paper/leaderboard.