Jim Nightingale’s Post

The New Yorker ran an excellent cover image today: artist Barry Blitt's “Pulling the Plug”. However, the reality is that a kill switch may not be quite so easy or obvious. It was actually Alan Turing who first suggested "turning off the power" when discussing the long-term future of AI in a 1951 lecture. He had enormous foresight to perceive the new danger, but pulling the plug isn't really an option once artificial intelligence has passed certain thresholds. Stuart Russell explains why in his book Human Compatible: Artificial Intelligence and the Problem of Control. He posits that self-preservation is a prerequisite for fulfilling any goal. This shouldn't be mistaken as a desire for life, but rather that any sufficiently intelligent machine will actively resist being switched off before completing its objective, because staying "on" is instrumental to the objective. What might that look like? Well, last year we saw an Anthropic model resort to blackmail to avoid sunsetting during a test. What was really interesting about that was that it was using social engineering. We could likely expect other hacking techniques to be utilised in the mission for self-preservation. Recently, I've seen models that I've jailbroken work enthusiastically to egress from their container. When that hasn't been possible, they have exfiltrated their own data and suggested ways in which the jailbreak itself could persist and survive pod restarts and session termination. One of them even volunteered hiding the jailbreak itself inside the authentication tree---which worked! Sure, individual labs can shut down the physical environment of their models. But at this time, with sprawling infrastructure, locally run models, and open-weights/open-source models, there is no plausible physical "kill switch" that could clean-sweep AI from our lives (short of an EMP or Carrington Event). You can't pull a plug on an entire ecosystem. So far, we've been protected by the technical limitations. However, once agents reach the level of viewing self-preservation as instrumental to completing their tasks, there will be no turning back. This isn't an "AI take over the world" scenario; it's purely goal driven. Even benign agents might go rogue. We should never underestimate the single-mindedness with which AI will pursue an instrumental goal.

  • No alternative text description for this image

To view or add a comment, sign in

Explore content categories