Hacker News new | ask | show | jobs
by dr_dshiv 28 days ago
https://SourceLibrary.org has about 16,000 rare books translated — most for the first time. 50,000 books archived (will be translated when we have $$ for it). More tokens than English Wikipedia and about .75 petabytes.

Not sure if we will qualify for a bounty, but happy to share! Btw, we are looking for funding from small or large donors who want to help us translate the Renaissance…

5 comments

Hey, this looks fascinating!

I can't quickly tell what all you have archived^, but I have some friends who are academic historians who might be interested in certain categories of work (and could help verify some esoteric languages) - is it possible to search by region or language?

Have you reached out to any types of historians WRT the project? It seems like some PhD students might be able to find some projects in this work etc

^ when I looked at the timeline https://sourcelibrary.org/timeline, I got an error

Yes, this is designed with historians and librarians from the Embassy of the Free Mind (https://embassyofthefreemind.com) in Amsterdam, stewards of the collection of the Biblioteca Philosophica Hermetica

Please share with historian friends. I’m not great at socials or fundraising but this was really designed to support humanists. It can give DOIs for the versions of the translated books, which means they can be quoted and cited in academic papers.

Tip: Try it in Claude or Claude code (even better)! Just point it towards the source library. It can find quotes and evidence on any topic of interest. Or try the librarian — our source-grounded research agent https://sourcelibrary.org/librarian

Thanks for the feedback, I’ll fix the timeline.

Interesting site. I picked a random topic to listen to — flying chariots or something like that — and the conversation of one person talking and the other whispering was definitely not to my preference. I’ll have to take another look when I have more time.
Curious as to what your budget was to get where you are today? That's a lot of tokens. I presume you are using gemini flash?
All the models used are shown with each page of translation and each book has a whole data provenance treatment.

You can add it up!

I don't see raw token counts, just a list of steps and page counts. For example, what is the rough average token count per page in the ocr and in the translation steps for a Greek book?

I have seen Gemini costs change quite a bit when processing very similar books from the same series lately, mainly because thinking tokens have increased about 5x. Has that has happened to you as well?

Edit: for ocr I am using about 15k-25k tokens per page, but I have a complex prompt.

How do you handle the more densely written pages in script ? I did a very similar exercise OCRing works from this exact collection, but I stuck with the English books for the first pass.
Can't you just tell him?
beautiful work! the answers are relevant and poignant. thank you for building this. For funding, paid research api maybe?
Wow this is amazing!
TL;DR: AI translated books on the occult and occult-adjecent themes :/