Hacker News new | ask | show | jobs
by killingtime74 39 days ago
Yes, it could just make one call to a multimodal llm to describe the scene