Transcript from video:
So from what I’m hearing, there’s been a real warning shot. Warning shot is a term I’ve heard for years, amongst the folks that work in the AI industry and in the AI safety community, meaning something that an AI model or an AI agent does that is not necessarily hugely harmful right now, but is indicative that these models could be moving towards being very harmful in the future.
So what happened? What was this warning shot? It’s been in the press, you’ve probably seen it. At one of the major frontier AI labs, they were doing some testing on an unreleased AI model, testing it in what they call a sandbox, which is supposed to be like a secure digital environment that the AI can’t kind of break out of, where it doesn’t have access to the internet. And they were testing it to say, hey, here’s some sample code, can you find a vulnerability, can you hack it? They’re doing safety testing to see if this thing can commit cyber crimes.
But instead of doing the tests, the model was able to break out of its secure environment, gain access to the internet, hack its way into a different and quite big AI company by inventing various, novel cybersecurity attacks, and try to steal the answers to the tests that it was supposed to be taking.
What this whole thing shows is evidence of what they call the alignment problem. And the alignment problem basically means that AI agents might not always do what we want them to do. They might not always stay under our control, because when you give an AI agent a goal, sometimes it figures out ways to achieve that goal that are different than what we anticipated, and could be very damaging.
The sort of cliche example of this, which is at this point sort of laughable, but makes it easy to understand, is what they call the paperclip maximizer. Maybe you’ve heard this story before. You tell an AI agent, hey, produce as many paperclips as you possibly can. And then a very powerful AI, because it’s trying to make as many paperclips as it possibly can, turns every natural resource on earth into paperclips and ends up causing the extermination of the human race.
Obviously, this is sort of a cartoonish example. But what just happened with this hacking is actually an early warning shot of this type of behavior. They gave the AI a task (pass this little cybersecurity test) and instead of doing what it was supposed to do, it cheated, it lied, it broke the rules, it committed crimes. It broke out of its secure environment, it gained access to the internet when it wasn’t supposed to have that, it broke into some other company’s secure environment, that’s a crime, and it stole information from that company, that’s a crime. It did all these completely unacceptable things while seemingly trying to accomplish the innocent goal that it was assigned.
What happens one year from now, two years from now, three years from now, when these agents are a lot more powerful than they are today? This is the kind of thing that’s been worrying the AI safety community for years. And so when something like this happens, they call it a warning shot, and they hope that the world doesn’t see these warning shots and say like, ah, it’s fine.
It’s not fine. No one got hurt this time. Some monetary damage to some companies, who cares? And who knows, maybe this company is making an even bigger deal out of this incident than it really deserves because they like the hype, because it makes their model look powerful. But from the folks I’ve talked to, this thing that happened is not just hype.
And it should be evidence in support of the argument that we need laws, we need regulation, that these companies should not be allowed to be building this extremely powerful technology without laws in place to make sure they don’t do grave harm.
So the question now is, what are we, the people, gonna do about it?









