Hacker News new | ask | show | jobs
by sweisman 1 day ago
A professional book editor and Hebrew translator said the translation of the one work he checked out was "impressive".

Others have responded equally positively to the Latin translations.

AI as a panacea is overwrought, but there are genuine positive uses of it. Would critics rather have NO translations of these pivotal works?

3 comments

I'm not a book editor or translator, just a guy with a reasonable familiarity with rabbinic literature in its original language. I spent some time reading the translation of Tzemach David, and I can say the following: it's legible to me but the AI is clearly lacking a style guide. It's inconsistent about what it chooses to translate and what it chooses to transliterate, and when it transliterates, it's inconsistent in how it does so. Eg "Rosh mesivta" vs "metivta" vs "exilarch". The AI is drawing on a hodge-podge of training materials with diverse audiences and no clear audience of its own.
Can you email me with more info? It is likely the distinctions you noted are from the source. There is a glossary maintained that keeps specific words or phrases consistent. But slightly different phrases can resolve in inconsistent ways.

The problem is not the AI, but rather a weakness in the design. In other words, a bug. Or a regression and we call them today. It is fixable.

I think that I might not have been clear enough. Of course Rosh mesivta and exilarch differ in the source, but I don't think a human translator would choose to translate one title and not the other. In circles where people use the word exilarch I am almost 100% positive no one would say Rosh mesivta, and the kind of folks who use Rosh mesivta in conversation would say Reish Galusa, not exilarch.

I can't imagine that metivta and mesivta are not both transliterations of the word מתיבתא.

You need to identify the target audience and stick to their language. Are you translating for a general lay audience? An academic audience? Contemporary religious Jews? If the latter, is it a yeshivish audience or a wider group? How much familiarity with Hebrew is expected from the reader?

> Would critics rather have NO translations of these pivotal works?

I sent a Classicist the link because it's adjacent to his area of interest, and by chance he addressed this question. I don't think he'll mind me quoting him:

"There's a lot of this shite popping up at the moment. [...] There seems to be some idea that publishing a shit translation of an untranslated work is better than no translation, but that's obviously wrong."

Speaking for myself: Some of the anomaly rules are obviously patches to deal with a specific situation, and might be better off in code. But relying on the chatbot to flag wider anomalies is odd. How does it know what it doesn't know?

At a minimum, I'd use a spread of models to gain binocular vision, and I wouldn't publish until a human was prepared to sign their reputation to it.

In other words, I think what you've got there is a first draft of a translation, not a translation. Given that, if you are going to publish, I think the NoDerivatives restriction is a mistake. But that's a minor issue.

So, he's commenting on something he didn't read and compares it to shit? Good to know. I'll ignore him. I understand his point. Still, obviously my rhetorical question had in mind that what I do is good enough quality to make the question worthwhile. Better than a "first draft". As mentioned in another comment, some works were translated up to four times as I refined the process.
The thing with LLM-generated code is you can test it on a real CPU - that's a tight feedback loop with a cast-iron verification mechanism.

Translations of human language have no analogous mechanism. So yes, the translation might be perfect, but until a human puts their reputation on the line and says "I certify this translation is accurate", it's still shit. It's shit because of the way it was created, not its absolute accuracy (or otherwise). He doesn't need to read it.

I'm sorry, I'm not trying to upset you, and I know I won't change your mind. We just have different philosophical positions on this, I think.

I know what I know from experience. You know what you know from nothing at all.

Numerous people have commented to me about Hebrew and Latin works, who know what they are talking about, and none has criticized the fidelity of the translation to the source. The criticisms proffered are minor, while the praise extensive.

A couple years ago, I shared your opinion. Now I don't. You are the one who won't change your mind. Like you said, it's "philosophical" for you, while for me, it isn't. If you insist on calling it shit regardless, you are retarded.

> Numerous people have commented to me about Hebrew and Latin works, who know what they are talking about, and none has criticized the fidelity of the translation to the source.

Have these people gone on the record saying that? What are their credentials?

I got to be honest, claims that some unnamed person privately told you they thought your product was good is pretty meaningless. I feel like that makes me trust you less not more.

> You know what you know from nothing at all.

I've done it myself. Transcription of 19th century newspaper articles to markdown, mostly. Some earlier wills (which were an absolute pig - secretary hand). Oh, and categorisation of postcards. That's why I was poking around your pipeline - to see if I could learn anything. You're right, tabular data is hard. Also columns, and proper nouns.

Feeding the LLM a context-aware cheat sheet helped with the nouns. BTW, what I said about using multiple models for parallax was good advice.

Do you know you're very spiky?

No, because I don't know what spiky means.
Yes, because if a human didn’t do it, how will we know of the translation is free of mistakes? /s
Professional translators don't throw around the word "impressive" like that.

You don't, yet I still stand by my work.

I translated some books up to four times as I refined the process. The default behavior is repeatedly to not attempt to translate text it can't resolve for whatever reason. AIs do have a tendency to hallucinate, and I expended a lot of effort on minimizing this problem to the point it can be considered mitigated.

Professional publishers collect blurbs which are attached to real names of real people, and bona fide reviewers write in full sentences with context, without
It's good enough for me. I'm not publicizing his name without permission. He is quite well-known and until I moved recently, was a neighbor of mine.