Hacker News new | ask | show | jobs
by aiagenta2z 2 days ago
I have a question about how did the project evaluate the technical English performance compared to other skills/MCPs and with LLM w/o skills? In the Github repo, there is a summary of "measured: 6 Claude models × 8 tasks × 2 conditions, 96 runs", Does it means the percentage of violations in the words? If that's the case, the measurement might be a little bit strange weather the STE measure is already in the prompt? STE violations per 100 words ▼ 72.9% (every model won). The bench or evaluation should not be in the same prompt, like eval/test.