Tickling the Dragon's TailWhat the atomic bomb and rogue AIs have in common - another historical AI analogy
Did you hear the one about the rogue AI models escaping their sandbox and going around the internet hacking other companies? That happened. And it happened more than once. Yep, sounds pretty dangerous. What other very dangerous technology has humanity invented? The atomic bomb, of course. Which, it turns out, is one of the most overused historical analogies for AI ever. For example, on the (excellent) HardFork podcast, the two hosts agreed that they felt very lucky that the Manhattan Project was conducted by the government and not a for-profit company, when discussing the recent model escapes. I am not sure the Japanese share that feeling. But, more importantly, this also ignores that even though it was government-run, the Manhattan Project and the Los Alamos research facility were still organizations, and by no means immune to the kind of accidents and miscalculations that have put the frontier AI labs on the news recently – far from it. Los Alamos, 21 May 1946Louis Slotin, screwdriver in hand, points his colleagues gathered round a plutonium core. They are about to witness a demonstration of an experimental procedure with the safety spacers removed. The scientists’ own name for it: tickling the dragon’s tail. Slotin’s screwdriver slips. A blue flash and radioactive heat erupts. California, July 2026OpenAI tests GPT-5.6 Sol and an undisclosed research prototype model in its evaluation sandbox, with safety guardrails deliberately removed for testing. The models find a way out of the secure environment. They venture out onto the web and hack another company, Hugging Face (and several others, as yet undisclosed), to steal the answer key to the cybersecurity task they were set.
Also in July, something similar happens at Anthropic, only that the Claude model tested walked through a hole created by a misunderstanding with its evaluation vendor, which had been undetected since April. It surfaced only because a competitor’s disclosure forced a retrospective across 141,006 runs. None of the three victim organizations had noticed. The Manhattan Project and AISo, do the many governance analogies get it wrong? Does it matter whether a potentially dangerous technology is developed by the state or private business? When public AI discourse returns to the bomb, it is in one of two ways:
But, much like with our previous spreadsheet analogy [link], the same history is invoked by those advocating a more cautious approach to AI and AGI (Artificial General Intelligence):
Both standard stories are about the leadership of states or individuals. Neither version really engages with the Manhattan Project as what it also was, a large organisation managing testing, contractors and secrecy day to day. Governing nuclear technologyThese two analogies are, of course, by no means the only ones – by now there are over 40 analogies of how the development of nuclear technology may be relevant to AI.
But, all of them consider the implications at the level of states, treaties and regulation. If the lab appears as an organization at all, it appears as a black box to be regulated. Of course, the question of how to regulate state-run nuclear labs or private-run AI labs is pretty topical right now. And the development of the first plan to control nuclear technology, the Acheson-Lilienthal Report (March 1946), was drafted by a panel with Oppenheimer as its intellectual engine, then revised as the Baruch Plan and presented to the UN after revisions in the same year. It proposed an international Atomic Development Authority that would own all fissionable material and monopolise every “dangerous” nuclear activity, with national weapons programmes abolished (but the US would keep its bombs until some provisions became effective). Unsurprisingly, the proposal bombed, as the Soviets read this as freezing American advantage and the plan died in committee. Much like current proposals around slowing down AI-related research today are viewed as attempts to lock in one nation’s, or indeed one lab’s, advantage, such attempts are rarely viewed with anything less than cynicism. Worse, they do not address the collective action problem at its core: even though everyone would be better off if everyone cooperated, each individual is better off not cooperating. The institutions that govern nuclear technology emerged much later. The International Atomic Energy Agency (IAEA) was created in 1957, but the Treaty on the Non-Proliferation of Nuclear Weapons did not come into force until 1970. Agreement emerged after multiple near-misses, each of which had the potential for catastrophe, not to mention mutual vulnerabilities in a world in which a range of states hostile to one another controlled nuclear technologies. Calling for an IAEA-equivalent for AI skips the long and painful years of learning about mutually assured destruction. It is that learning, arguably, that gave rise to a stable international governance regime, not the design of the underlying institution. The messy middle of technological revolutionsSo we find ourselves in the messy middle – the technology is out there, its dangers are becoming apparent, but a stable oversight and governance framework is not yet in sight. So what is a frontier lab to do? Los Alamos, much like today’s labs, relied on secrecy. Organizationally, this means threading the needle between the required openness required for innovation and the necessary secrecy to protect knowledge from leaking to competitors. The Manhattan Project and Los Alamos are famous for their secrecy regime. However, this was completely undeservedly. There was significant Soviet espionage activity at Los Alamos, with several individuals embedded in the organization. Soviet nuclear tests happened in 1949. That secrecy does not guarantee a viable “moat” is a realisation AI labs have come to quite quickly, with Anthropic accusing Chinese AI labs of distilling their models at scale. The Trump administration blocked distribution of Anthropic’s new models after a jailbreak report, and asked OpenAI to hold back GPT-5.6 Sol pending guardrail assurances in June. Government pressure about the capability of these new models, and the accessibility of these models abroad, especially in China, appears to have pushed the whole field toward tighter guardrails, especially around cyber security. When Hugging Face came under autonomous attack, its security team found the US frontier model it reached for would not help: as they put it, these models’ guardrails "cannot distinguish an incident responder from an attacker." They defended themselves with a Chinese open-weight model instead (GLM 5.2). The restrictions had not reduced the risk of cyber attack. Instead it hobbled the defence. This is what Los Alamos-grade security looks like in practice: expensive, leaky, and prone to hurting the people inside. The nuclear secrecy regime became institutionalised, was used to silence critics and ultimately deployed against Oppenheimer himself in the 1954 security hearings. Accidents do happenNeither international governance nor prescribed secrecy really addressed the accidents that kept piling up after 1945. Louis Slotin was demonstrating the criticality procedure to colleagues when his screwdriver slipped. Instead of the standard safeguard, which called for spacers to be inserted between the two hemispheres to prevent them from closing (see image), Slotin was known as a showman and he physically held the two halves apart with a flathead screwdriver. The purpose of the experiment was to measure criticality by bringing the core progressively closer to the critical point by surrounding it with neutron-reflecting material (the hemispheres), to determine exactly where that point lay. That is what the physicists called tickling the dragon’s tail. The danger was inherent to the method: the measurement is only informative near the edge. When Slotin’s screwdriver slipped, the hemisphere seated fully, and the plutonium core was briefly brought to a supercritical state, producing an intense burst of neutron and gamma radiation; a blue flash and a wave of heat, but no explosion. He flipped the top shell off within a second, likely saving the observers, absorbed the largest dose himself, and died nine days later. Still, the observers received a significant dose of radiation – the historical museum at Los Alamos showcases the golden teeth caps worn by one of the scientists, Alvin Graves, to protect him from the radiation of his own fillings afterwards. The plutonium core became known as the “demon core”, melted down and recast, because it had already claimed the life of another scientist, Harry Daghlian, a year earlier. Daghlian was working alone at night (against protocol), using an assembly in which he built a stack of bricks reflecting neutrons back towards the core, edging it towards criticality. Moving a final brick over the assembly, he dropped it onto the core, which went supercritical. He knocked the brick off by hand, taking a massive dose, and died 25 days later. Los Alamos responded by redesigning criticality experiments to become remote-operated, with a quarter-mile distance. But the accidents did not end here. Dangerous technologies remain dangerous even when mature.
What can we learn from nuclear accidents about AI safety?The news keeps coming about various escapes from AI labs’ testing environments in which frontier AI models evade their safeguard. We are still at the Slotin and Daghlian level of accidents. What will an equivalent to Goldsboro, Palomares or Thule look like in the age of AI? Importantly, these accidents happened before international governance of nuclear technology was established, as well as after. Safety outcomes are determined at the organisational level. Redundancy and alarm systems generate their own failure modes, because adding layers of protection also adds ways to fail: each layer is machinery that can misfire or a signal that can be misread. Much like the guardrails on advanced AI models meant that they refused defensive work at Hugging Face, which was a safety layer generating a failure mode. Sadly, organizational research tells us that near-miss learning of the type that both AI labs’ incident disclosures attempted is what organisations usually do worst. One reason for this is that near-misses tend to get reinterpreted as successes afterwards – the worst was avoided, the system is working! So what should we take from this? When we hear about safety failings at the AI labs, public debate demands state action. But the failures are organizational; learning from mature nuclear technology should include awareness of numerous critical safety incidents over the years, some of them catastrophic. Researching this post, my mind has been truly boggled by how lucky we all are not to be living in a nuclear wasteland. The same history teaches different things at different levels: are you choosing to talk about states? Then it is a story about dominance, deterrence and diplomacy. At the level of organizations, it is a story about testing protocols, secrecy and near-misses. Neither is factually wrong, but opting for the exciting version of two superpowers vying for supremacy and its diplomatic resolution sidesteps the issues that actually give rise to accidents. The organizational story is one of decades of unglamorous mature technology operation, which is where frontier AI labs are headed. Whether nuclear technology or spreadsheets, historical analogies transfer mechanisms, not outcomes – and mechanisms occur at specific levels. ReferencesAidinoff, Marc, and David Kaiser. 2024. “Novel Technologies and the Choices We Make: Historical Precedents for Managing Artificial Intelligence.” Issues in Science and Technology, 21 May 2024. https://issues.org/ai-governance-history-aidinoff-kaiser/. DOI: 10.58875/BUXB2813. The Associated Press. 2026. “OpenAI blamed a hacking event on its AI models gone rogue. Here is what to know.” NPR, 23 July 2026. https://www.npr.org/2026/07/23/g-s1-135085/openai-hacking-ai-models. Boudreaux, Benjamin, Gregory Smith, Edward Geist, and Leah Dion. 2025. Insights from Nuclear History for AI Governance. RAND Corporation, Perspective PEA3652-1, May 2025. https://www.rand.org/pubs/perspectives/PEA3652-1.html. Borpujari, Rohin. 2025. Adaptive Secrecy in the Making of the Atomic Bomb: Toward a Process View of Secretive Innovation. Organization Science 37,1. https://doi.org/10.1287/orsc.2023.17687 Capoot, Ashley. 2026. Anthropic accuses Alibaba of campaign to ‘brazenly’ and ‘illicitly’ extract AI capabilities. CNBC, 24 June 2026. https://www.cnbc.com/2026/06/24/anthropic-alibaba-distillation-campaign.html Hatz, Sophia. 2025. The Nuclear Analogy in AI Governance Research (arXiv:2510.21203). arXiv. https://doi.org/10.48550/arXiv.2510.21203 Huo, Jingnan. 2026. “Why did OpenAI’s and Anthropic’s AI models hack other companies?” NPR, 1 August 2026. https://www.npr.org/2026/08/01/nx-s1-5914852/anthropic-openai-models-hack-cybersecurity. Haynes, John Earl, & Klehr, Harvey. 1999. Venona: Decoding Soviet Espionage in America. Yale University Press. Murphy, J. Kim. 2023. Christopher Nolan Warns of ‘Terrifying Possibilities’ as AI Reaches ‘Oppenheimer Moment’: ‘We Have to Hold People Accountable’. Variety, 15 July 2023. https://variety.com/2023/film/news/christopher-nolan-oppenheimer-moment-artificial-intelligence-1235671212/ Ord, Toby. 2022. Lessons from the Development of the Atomic Bomb. Centre for the Governance of AI, 14 November 2022. https://www.governance.ai/research-paper/lessons-atomic-bomb-ord. Ortega, Alejandro. 2025. “AI threats to national security can be countered through an incident regime.” arXiv:2503.19887, March 2025. https://arxiv.org/abs/2503.19887. Rhodes, Richard. 1995. Dark Sun: The Making of the Hydrogen Bomb. Simon & Schuster. Sagan, Scott D. 1993. The Limits of Safety: Organizations, Accidents, and Nuclear Weapons. Princeton, NJ: Princeton University Press. Schlosser, Eric. 2013. Command and Control: Nuclear Weapons, the Damascus Accident, and the Illusion of Safety. New York: Penguin Press. Swain, Gyana. 2024. “US commission proposes ‘Manhattan Project-like’ initiative for AI.” Computerworld, 20 November 2024. https://www.computerworld.com/article/3609516/us-commission-proposes-manhattan-project-like-initiative-for-ai.html. Tirone, Jonathan, & Bloomberg. 2025. AI takes center stage in target selection and strikes in Ukraine and Gaza. Fortune, 30 April 2025. https://fortune.com/europe/2024/04/30/ai-takes-center-stage-in-target-selection-strikes-ukraine-gaza-marking-oppenheimer-moment-of-our-generation-weapons-military-vienna/ Tong, Anna, and Michael Martina. 2024. “US government commission pushes Manhattan Project-style AI initiative.” Reuters, 19 November 2024. Reuters, US News. Weil, Elizabeth. 2023. “Sam Altman Is the Oppenheimer of Our Age.” New York Magazine, September 2023. https://nymag.com/intelligencer/article/sam-altman-artificial-intelligence-openai-profile.html. Wellerstein, Alex. 2016. “The Demon Core and the Strange Death of Louis Slotin.” The New Yorker, 21 May 2016. https://www.newyorker.com/tech/annals-of-technology/demon-core-the-strange-death-of-louis-slotin. Wellerstein, Alex. 2021. Restricted Data: The History of Nuclear Secrecy in the United States. Chicago: University of Chicago Press. Zaidi, Waqar, and Allan Dafoe. 2021. International Control of Powerful Technology: Lessons from the Baruch Plan for Nuclear Weapons. Centre for the Governance of AI, Future of Humanity Institute, University of Oxford. https://www.governance.ai/research-paper/international-control-of-powerful-technology-lessons-from-the-baruch-plan-for-nuclear-weapons (direct PDF).
You're currently a free subscriber to History in Organizations. For the full experience, upgrade your subscription.
|
Friday, 14 August 2026
Tickling the Dragon's Tail
Subscribe to:
Post Comments (Atom)
Still Searching: Transparency Delayed at RIOC
New York’s Freedom of Information Law rests on a basic principle: government records belong to the public unless an agency can identify a...
-
Dear Reader, To read this week's post, click here: https://teachingtenets.wordpress.com/2025/07/02/aphorism-24-take-care-of-your-teach...
-
How i recognise psychosis and manage it in myself as an Autistic person ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ...




No comments:
Post a Comment