| But you can't read the original! That's why you built the pipeline in the first place! Only people who can read both versions can have a grounded opinion on whether or not the translation is accurate. C'mon man, I know criticisms of your baby feel like attacks on you, but take a step back. You must see that "the name of the author of the pipeline is public knowledge" is not a sensible response to "nobody who can read both versions has attested to its accuracy in public". Right now, everybody believes that LLMs can't self-correct (see https://arxiv.org/abs/2310.01798) and that they hallucinate when asked to do OCR tasks. It doesn't matter if that's true or not, that's the prevailing belief you're working against. If you want to convince people your pipeline can self-correct, you're going to need to supply evidence. I see two possibilities: (1) Pay someone to audit one of the translations publicly. Crucially they need to answer the question: does the pipeline introduce inaccuracy? If it does then it really is worse than useless, because it tells lies. Clunky, inconsistent, partial translations would still be better than no translation - it's the risk of hallucination that's the killer. (2) Are you doing test runs on similar documents that have also been translated by humans? (Preferably very recent translations). Publishing those test runs would allow poor uneducated slobs like me to line up a human translation and your machine translation and see for ourselves that your pipeline works. This is meant as helpful advice - I'm trying suggest paths that respond to valid critique with something other than bluster. I wish you and your project well. |
These are the kind of comments I also typically get regarding regressions. Problems with mostly mechanically repairable aspects of the translation, but not with the quality of the translation itself. From people who know of what they speak. That I do not put them on the record is not relevant, except to people like you. They are real, and their attestations are real.