Hacker News new | ask | show | jobs
by mk89 33 days ago
At my company someone has introduced an internal tool that should help understand and give a "score" to design documents from teams.

Needless to say, this tool gives scores exactly like the article mentions. Same document, same LLM, same prompt, and different results. It becomes even more ridiculous once you switch to other models, or if you ask a model to review the work of another model.

I am not sure why we insist on making LLMs do the work they are not supposed to do and/or in a way they are not supposed to do.

The worst part is that people are aware of the problem but they just ignore it and consider it as "a reference number, just to have an understanding".

If it were like that, it would be less of a problem. The issue comes from the fact that eventually someone without enough knowledge will trust the output (so X points out of Y is how it is), or someone will stop challenging the output and consider it for their process - like in this unfortunate case of hiring.

At a certain point, people who don't know what they are doing give a tool that doesn't know what its doingto people who don't know what they are doing. A pure mess. And everyone has to comply and applaud. If you go against, you are against AI.

This is what I hate the most about AI. Not the tool, but the shortcuts we're willing to take to justify its existence.