|
|
|
|
|
by flir
3 days ago
|
|
> You know what you know from nothing at all. I've done it myself. Transcription of 19th century newspaper articles to markdown, mostly. Some earlier wills (which were an absolute pig - secretary hand). Oh, and categorisation of postcards. That's why I was poking around your pipeline - to see if I could learn anything. You're right, tabular data is hard. Also columns, and proper nouns. Feeding the LLM a context-aware cheat sheet helped with the nouns. BTW, what I said about using multiple models for parallax was good advice. Do you know you're very spiky? |
|
Hebrew is also much worse than Latin, which uses Arabic numerals. Hebrew conventionally uses the letters for numeric representation too.
I am getting good (and improving results), but the token cost is heavy.