Anyone can write out the code for a bad virus. You can go download it from an open repository. It's only through deep interface with the world that the idea for the bad virus turns into an actual bad virus. The fever dreamers will say the LLMs will help you interface with reality to do the bad thing™ which their model let's you do. But you still need thousands to millions of times the effort And once made how do you deliver it in a way that might further your (bad) objectives? Presumably it just makes humans sick. There aren't "targeted" bioweapons, and among humans we are too damn similar for there ever to be. And all of this said, there is almost nothing special about the LLMs' abilities in biology. They only know what we know. They're not being trained autonomously with RL and a robotic wetlab. When that's a thing I'll start to take claims of biology risk more seriously. Right now they have the same logic as the paperclip theory of superintelligence risk. And cynically, it would seem that Anthropic purchased a biotech company and almost immediately decided to lock down biology work with their models.
> They're not being trained autonomously with RL and a robotic wetlab. When that's a thing I'll start to take claims of biology risk more seriously
So if IIUC your point is "they're not good enough at biology right now because they're not trained on it so they're not a threat".
To which I want to answer: "they're not a threat now but I see *no* reason for models not to be trained on biology pretty darn soon unless people like you convince the world otherwise."
They are already being trained in biology right now? What he is saying is that they are being trained on what we know about biology currently. AI is going to hit that limit. And there is no way for the AI to gain more knowledge without actually doing lab experiments
My point is more about conflict of interest in Anthropic, and certain techpeople types thinking that biology is a useful scapegoat because biologists don't use or understand this tech (they claimed 0.03% of users would be affected by biotech security stuff). So, arguing the world will fall apart because someone can ask your superduper powerful tool for a bad bad virus sequence, or how to do some technique in the lab, and get an answer that's plausible (oh but you never bothered to actually test if it's legit because you don't have the capability to do so).
As for the future... today the LLMs are "trained on biology", in that they read the textbooks, the research, the web.
They aren't trained on biology in the sense of being embodied, autonomously or semi-autonomously driving actual biological experiments. If you come from software, the timescale of these experiments is outlandish. Yes, I am partly saying the LLMs are not good enough today because they need to be embodied and trained for literal decades of lab time before there is even the _remote_ possibility that they could present a novel risk profile that is even a shadow of what the current fearmongering suggests the current models can enable.
And they aren't trained on biology in the sense that they've read the literature, but even 100T token training run only begins to touch the data scales that rather mundane bioinformatics operate at. True multimodal models that work on DNA and human language at high quality haven't yet emerged. We're talking new architectures which are going to arise after the next AI winter.
All of this ignores an even more fundamental point. Cost. If someone wants to make a bionuke, they don't need to use AI. They can set up the right evolutionary context and run quadrillions of parallel explorations of the design space. Directed evolution like this is cheap, well-understood, and insanely powerful. If you actually care about biosafety, we should be doing hard work to surveil gain of function research. Different flavors of LLM use are not going to be a differentiator for the foreseeable future.