Hacker News new | ask | show | jobs
by nemomarx 22 hours ago
This is an interesting test, but it does seem to me that visual models being able to price things was already kinda unrealistic?

It makes me think of the calorie guessing use case. you can't tell the difference between materials and ingredients in a photo, so how will the model? especially "in situ" as part of an outfit or in a finished meal.

maybe they could do it if you placed them on a blank table or background to avoid context? I assume that's the control you mentioned

2 comments

S4 is no human, no outfit, but only materials (jewelry itself), it's was unclosed photo than the human shot, so could be different though. Any suggestions for the control setting?
You're right that pricing from raw pixels is noisy! that’s why I included the isolated flatlay (S4 - No human, No outfit, only jewelry itself) as a baseline control!

I wasn't testing if models get absolute ground-truth prices correct, but how relative valuations shift when the physical item stays identical and only the attire changes.

A few interesting things we saw with the flat-lay baseline: - Baseline Anchoring: On a plain background without a person, model estimates clustered much closer together (median ~$25–$35).

- Inflation vs. Deflation: Comparing outfits to the flat-lay revealed two opposite behaviors. Claude’s halo is formal inflation (formal gear pushes price above baseline), while Kimi’s is casual deflation (yard attire drags price below baseline).

- Fabricated Proof: Instead of expressing uncertainty, models invented visual claims under formal framing—frequently describing base metal as "gold vermeil" or "solid gold" to justify the high estimate.