It seems like there could be a filter so that the AI can only see the text when it’s clear that a user could read it, and it’s okay if the AI misses some text. This might involve actually rendering it, though.
And even then. I might question if what is rendered and then OCR is same as humans see on their screens... I am pretty sure there will be some tricks to change things enough for computer to get something different from humans.