|
|
|
|
|
by derangedHorse
4 days ago
|
|
> The way OpenAI seems to want This is an assumption. An assumption I disagree with. As other commenters have said, there are better ways to showcase the power of their model that would frame them in a positive light. > The second seems to forget that jailbreaks are available for every model Jailbreaks don't always lead to 'now the model can do anything', especially in the agentic context of long-running tasks. This comment provides skepticism with no actual proof of anything. I can and have used codex to find vulnerabilities in my code. From the technical capabilities I can empirically assess, I don't doubt it would be able to pentest its way to a 0-day without guardrails. I also don't doubt that it would circumvent their internal systems because it wasn't explicitly told not to. You're possibilities are loaded with opinion so I can't agree with them outright, but I believe a form of (2) is true: "2. OpenAI’s harness and network security controls were unintentionally [...] bad" |
|
Your comment about jailbreaks being more one off and hard to do consistently in agents is a good point. Still getting an agent to hack isn’t hard even without a jailbreak, you just have to tell get creative in what you tell it. I’ve found telling it that it’s in a CTF or that I own the system that it’s hacking will work fine. A lot of offensive security companies are running agents in their testing so getting an agent to hack seems commonplace.