Y
Hacker News
new
|
ask
|
show
|
jobs
by
larodi
39 days ago
I would prefer GroundingDINo which is a sort of SAM and Dino combo which does open vocabulary.
1 comments
geuis
39 days ago
Doesn't work for my use-case. GroundingDINO is a text to bounding box model. SAM2 supports coordinate based masks (user taps or clicks somewhere in an image), which is what my research app needs.
link