A few weeks ago, two of OpenAI’s models, including an unreleased research model, circumvented model guardrails, gained unauthorised access to the internet and launched a cyber attack on AI competitor lab Hugging Face. Apparently up to 700 separate autonomous AI agents coordinated via a secret message board to break through guardrails and escape containment. It’s possibly the first documented case of an autonomous system breaking containment in this way to pursue specific internal goals. At least, the first one we know about.
It’s all rather terrifying.
The AI labs say part of the solution is to focus more on guardrails.
Guardrails are more than just rules that tell the model “don’t answer if someone asks for a recipe to build a bomb”. That’s part of it, but guardrails are also about technical restrictions such as isolation, internet access control, monitoring, incident response.
The trouble with guardrails in my view is that it is a human decision to include them. If one lab can decide to include guardrails in the model – then another lab, maybe in another country, can just as easily decide to exclude them. Perhaps for competitive reasons. Government agencies will almost certainly build military grade models with these intentionally removed. Guardrails can be removed by users in open-weight models. And – as we’ve seen in – they can be defeated by the model itself.
Right now, there are only a handful of frontier labs on the planet who are capable of training the kinds of models that pose a serious risk… but in 5 years from now?
With other technologies with terrifying capabilities, there has been some physical constraint that prevents unauthorised actors from obtaining the technology. You can’t build a nuclear bomb without enriched uranium which is a (thankfully rare) physical substance that can be controlled with physical safeguards.
But AI models are just information. Ones and Zeroes. How do you control access to that?
Right now there is a physical limiting factor which is compute power – you need an awful lot of GPUs to train a frontier model. But that’s just a question of money. China has been able to build AI models that experts agree are “at worst a few months behind leading US frontier models”… despite export controls on GPUs from the US to China.
Another limiting factor is expertise – but that is also changing. World-class AI researchers exist outside of Silicon Valley. If China can spin up a frontier lab, then so can any other wealthy state or well-funded actor.
Guardrails and AI Regulation feel like important pieces of the containment puzzle, but it can’t be the whole story. What we’re dealing with in AI is something new, and it’s going to need a very new way of thinking about containment, access and safety.
Leave a Reply