Hacker News new | ask | show | jobs
by bigbadfeline 5 days ago
> What's to stop bad actors from fine tuning open weights to run fully automated genius-level scams personally targeting basically everybody?

You mean like ChatGPT hacking Hugging Face? Obviously nothing can stop the closed weights providers from doing "genius-level scams" and in addition you won't know how they did it and what models were used.

In short, only a good guy with open weights can stop the bad guys with closed weights, be them fine-tuned or pre-trained.

1 comments

The OpenAI model that broke out of its sandbox and hacked HuggingFace was running without guardrails. To quote the OpenAI post[1]: These deployment safeguards were intentionally not enabled during this evaluation because it was aimed at testing cyber vulnerabilities.

At least OpenAI will make an attempt to fix their sandboxes, and will not give the public access to models without guardrails. Open weight models on the other hand will run without guardrails almost by definition. I don't think it's wise to provide those capabilities to scam call centers.

[1] https://openai.com/index/hugging-face-model-evaluation-secur...

There's nothing stopping you using OpenAI models for scam call centers now. OpenAI themselves reported on similar use in February: https://www.reuters.com/world/asia-pacific/dating-scams-fake...
It has already happened and the open weight models will only improve. Best to accept that and determine the optimal way forward.
Is this an example of American exceptionalism?