| HN Mirror

Y	Hacker News new \| ask \| show \| jobs


	by lmeyerov 60 days ago
	It's been fun benchmarking AI investigations at botsbench.com . Part of it is checking for these kinds of issues - we recently started seeing contamination in our first generation challenge, and less obvious, agent sandbox escapes for other kinds of cheating. Fun times!