I download only title, description, thumbnail, common og fields.
My index is very lean, I think. I have 2m of pages crawled.
https://github.com/rumca-js/Internet-Places-Database
It has tags, and votes support.
Recently I also launched my first fdroid app
https://github.com/rumca-js/OfflineWebSearch
https://f-droid.org/en/packages/io.github.rumcajs.offlineweb...