Hacker News new | ask | show | jobs
by reasonableklout 18 days ago
It's not the same thing. For example, given a GUI with a titlebar, title, subtitle, text, and buttons, a human can instantly understand spatially the relationship between these items. But a naive OCR of such a GUI would be a flat stream of text that loses a ton of information.
1 comments

But that’s not how models handle images either. They spatially segment and reason about title bars, placement, etc.