Y
Hacker News
new
|
ask
|
show
|
jobs
by
octember
29 days ago
Cool idea, but keyframes are not videos. Motion, object permanence, are not things Claude can infer from a set of images. Nice demo though!
2 comments
fzysingularity
29 days ago
Exactly! We experimented with a whole bunch of video encoding techniques for LLMs here:
https://vlm-run.github.io/mm/encoders/#video
link
sawjet
29 days ago
I have been going through this with claude and qwenvl3:8b this week. Both are pretty decent at inferring context and analyzing contact sheets. Finding high visual interest moments with a mixture of coarse and fine keyframes.
link
octember
28 days ago
Might be time to check gemma :)
link