Hacker News new | ask | show | jobs
by brryant 19 days ago
hah - we actually skew staff, senior staff.

We have been testing GPT 5.6 for about a week as a preview model through a YC relationship, providing them feedback on the model. Our evals run in github CI and we can run them all in about 15 minutes against our eval bench of 115+ web design and marketing related jobs that ploy.ai specializes in.

then after we toggled it on (through a posthog feature flag) we actively monitored for failures.

I came from running Webflow, which powers > 1% of the internet so trying my best to relay all of that knowledge to ploy to power more % of the internet!

1 comments

This is super impressive - both the pedigree of the team and the approach you took.

The funny thing is - when I first saw ploy, I didn't take it very seriously since so many of the signals that used to signify quality (decent design, copy, hard technical problems) are easy to fake. Plus the "grow while you sleep" space is crowded with weak players.

I wonder what the new markers of quality will be, which would separate the hand-crafted (to the extent possible) work v/s slop.

agree that many of those markers of quality are now low signal. Ultimately we let our customers vote with their wallet

Funnily enough, we spent a long time on our brand. From our launch video that has human actors, to our product details. I believe a distinctive, high quality, well implemented brand is still a hallmark of a strong product or service.