|
I remember being blown away by o1-o3 family of models finally stringing together coherent agentic tool calls to write and execute scripts semi-reliably for workloads in the several minutes before they would start hallucinating/flailing. GPT 5 was a bit ahead of that, but barely Now we take for granted that the latest models can juggle between multiple browser tabs, applications, databases, simulators, docker etc to write, execute, e2e test and deploy full-stack applications over hours managing up to dozens of subagents, relatively untouched, without taking down prod even 1% of the time Not only this, but in the GPT 5.0 era, agents had 0 taste. Nothing looked good. It was the agentic version of the twitter bootstrap era, but worse somehow. Now, I would argue the average agent frontend beats the average human frontend. This isn't even getting into 3D applications in the GPT 5 era Anyway, the models now reliably execute more than a human can fit into their own context. It's magic |
Once we have something that experiences a desktop interface more like a human does, an entire swathe of tooling that has heretofore been nigh-impossible to automate moves into the fold, and that'll be another explosion of folks finally getting to join the agentic workflow world on their industry specific apps...