Hacker News new | ask | show | jobs
by sometimelurker 53 days ago
> easy(ish) to detect

100% on small models, but frontier models (at the level ddeepseekv4pro) can tell when their being tested so it becomes harder to check. you can always finetune them to remove CCP propaganda from them

1 comments

"Being tested" here just means asking for a feature on a legitimate codebase. The larger models don't magically know the user's ulterior motives.
I think they eventually will, using the style of the prompts or code, various metadata floating around on the computer theyre working on, if it can be done theoreticly, then it can be trained against (maybe even unintentionally)

no idea how large the model would have to be for this (larger than mythos/10T params? maybe)