An interesting question. I suspect with the current training that it would be very difficult to get an agent to distrust a skill or an MCP server. Could one get them to distrust an API call?
As the AI world seeks more benchmarks it would be an interesting one to see them try to build some benchmarks around testing this.
As the AI world seeks more benchmarks it would be an interesting one to see them try to build some benchmarks around testing this.