Hacker News new | ask | show | jobs
by throwa356262 19 days ago

    As of today, Ploy’s agent runs on GPT-5.6 Sol, the flagship tier of the model family OpenAI released this morning. 

Wait a moment, did they make the switch based on half a days of playing with Sol? Are these companies ran by teenagers?
5 comments

hah - we actually skew staff, senior staff.

We have been testing GPT 5.6 for about a week as a preview model through a YC relationship, providing them feedback on the model. Our evals run in github CI and we can run them all in about 15 minutes against our eval bench of 115+ web design and marketing related jobs that ploy.ai specializes in.

then after we toggled it on (through a posthog feature flag) we actively monitored for failures.

I came from running Webflow, which powers > 1% of the internet so trying my best to relay all of that knowledge to ploy to power more % of the internet!

This is super impressive - both the pedigree of the team and the approach you took.

The funny thing is - when I first saw ploy, I didn't take it very seriously since so many of the signals that used to signify quality (decent design, copy, hard technical problems) are easy to fake. Plus the "grow while you sleep" space is crowded with weak players.

I wonder what the new markers of quality will be, which would separate the hand-crafted (to the extent possible) work v/s slop.

agree that many of those markers of quality are now low signal. Ultimately we let our customers vote with their wallet

Funnily enough, we spent a long time on our brand. From our launch video that has human actors, to our product details. I believe a distinctive, high quality, well implemented brand is still a hallmark of a strong product or service.

There’s every possibility they got some amount of early access to evaluate GPT5.6, precisely so they could write an article like this.
at this point, it's pretty easy to create evals/benchmarks, and then run the latest model on them.

LLMs are so easy to swap out, so having good benchmarks/evals are pretty useful.

Even then, a lot of the time the model improvements are so obvious that you don't even need an eval.

I would expect they have production based datasets they evaluate new models against.
I would not expect that. They wouldnt have missed mentioning it if that was the case. Its mostly driven by vibes.
>For four months, no frontier model beat Claude Opus in our production evals

What do you think that is referring to?

AI psychosis is still at all time high