Hacker News new | ask | show | jobs
by zvolsky 21 days ago
The misleading prompt cases are inadvertently testing the model's ability to filter out noise from its instructions. That could be a benchmark on its own. The correct response is to flag the inconsistency and ask for clarification.