Hacker News new | ask | show | jobs
by teiferer 8 days ago
Lots of words about multi-modal but then this:

> our mission to develop real-world visual intelligence

Visual is mono-modal, isn't it?

2 comments

its doing video, audio, images and motion. I think that counts as multimodal.
Indeed, but audio isn't exactly "visual", is it?

And I'm not sure video, images and motion are actually 3 different things. Images are just still motion and videos capture motion, so it's really just "video and audio" of which one is visual, the other is not, thus my confused/surprised comment.

Is this really the value-add comment you’re going with?
Says the one who posts this comment?