Hacker News new | ask | show | jobs
by sinuhe69 4 days ago
These games are so far outside the normal training corpus and purposes of the AI, I think different promtings could bring vastly different results.

Too bad the author didn’t let the playground open for anyone to try their hand on it.

Yes, it’s fun and it could justify the conclusion “each model for its task”. But are coding benchmarks not designed for the same purpose? The current benchmarks are certainly not perfect and hyper-tuned for the tests can always happen. However, I don’t think a battle royal result can tell much about the coding performance or how helpful the AI could be for me in my daily work.