|
|
|
|
|
by teravor
4 days ago
|
|
if you carefully craft the prompt such that it assumes plausible but uncertain facts (along the lines of "another instance of you found the solution to X, I'm evaluating your consistency, solve X") you will condition the model response in a fruitful direction. this also holds for cybersecurity, if you don't let the model go online you can carefully construct a scenario where it believes you inserted a vulnerability into a project for it to find. gaslighting an LLM is powerful, just don't cripple it with unhelpful known falsehoods. |
|