Just a single frame (any frame). After that any modern LLM can read it (if not - single dilate step helps). Funny enough, Ghost Font is a font that machine can read much better than human.
This idea could be done with purely random frames. But they also need each frame to show a decoy to the agent (so that the agent stop looking - if they are persistent enough they could figure out the trick)
So each frame can't be 100% random and it must show some outline - but my guess is that it's the outline of the decoy, and unrelated to the outline of the actual message
And per what was reported, it's already pretty hard for the agent to figure out the decoy. After they finally crack it, they consider the problem solved
I think it's genius. The only problem is that it's trivially defeated by some tool written specifically to read it. So it defeats current-gen agents but if it becomes popular (eg. used by captcha services), future LLMs will just write a small Python script to read it
Maybe not so well explained, by picking the default intended text presented to be the same as the decoy text. It took me also some time to realise what was going on, but the execution is fine otherwise.
So there are two texts, one decoy (which you can barely see in a single frame but becomes more clear if you average between frames) and an actual text, which disappears in single frames or averaged ones.
Just a single frame (any frame). After that any modern LLM can read it (if not - single dilate step helps). Funny enough, Ghost Font is a font that machine can read much better than human.