Hacker News new | ask | show | jobs
by drusepth 20 hours ago
One of the best features I built into our game studio's AI router (IMO) was to take a random sampling of generations and have them generate across ALL models (that we support, at least) so we, the developers, can browse outputs and get a sense each model's output. Seeing them all side by side for generations you're already familiar with makes it feel like significantly less cognitive load.

The better I get at recognizing what each model is good/bad at, the more I'm glad we're taking the time to choose specific models for specific prompts -- and the more I wouldn't trust a generic router to efficiently route for me.

1 comments

This is a problem we're looking at atm. We want to start running evaluations as we're having problems with consistency across releases even just for individual agents.