Hacker News new | ask | show | jobs
by SequoiaHope 22 days ago
“On an ordinary coding prompt, the J-space of a model trained to sabotage code contains “fake,” “fraud,” “secretly,” and “deliberately” at the start of its response.”

I would like to know more about their model trained to sabotage code…

1 comments

thank you!