Hacker News new | ask | show | jobs
by master_crab 18 days ago
A lot of these are visual-heavy tests that often require first person sight to confirm results. Considering GLM isn’t multimodal, that might explain why it did better on the calculator question and not much else.