Noahpinion · Economics & Policy
TIER 4 Mon, 15 Dec 2025 10:48:22 +0000
Today at a Christmas party I had an interesting and productive discussion about AI safety. I almost can’t believe I just typed those words — having an interesting and productive discussion about AI safety is something I never expected to do. It’s not just that I don’t work in AI myself — it’s that the big question of “What happens if we invent a superintelligent godlike AI?” seems, at first blush, to be utterly unknowable. It’s like if ants sat around five million years ago asking what humans — who didn’t even exist at that point — might do to their anthills in 2025.
Essentially every conversation I’ve heard on this topic involves people who think about AI safety all day wringing their hands and saying some variant of “OMG, but superintelligent AI will be so SMART, what if it KILLS US ALL?”. It’s not that I think those people are silly; it’s just that I don’t feel like I have a lot to add to that discussion. Yes, it’s conceivable that a super-smart AI might kill us all. I’ve seen the Terminator movies. I don’t know any laws of the Universe that prove this won’t happen.
One response you can have to this is to conclude that because superintelligent AI might kill us all, that we should make very, very sure that no one ever invents it. This is the thesis of If Anyone Builds It, Everyone Dies, by Eliezer Yudkowsky and Nate Soares:
This sort of reminds me of Gregory Benford’s famous admonition to “Never do anything for the first time.” You can make arguments for unforeseen catastrophic consequences for lots of different technologies. Nuclear weapons might kill us all of course, but so might synthetic biology or genetic engineering. Social media or other media technologies might collapse our societies. Video games might entrance us into virtual lives so that we never reproduce. In fact, it’s quite possible that universal literacy and reduced child mortality have already condemned humanity to dwindle comfortably into nothingness, with no artificial superintelligence required.
Personally, I do think that being so afraid of existential risk that you never invent new technologies is probably a suboptimal way for an intelligent species like ours to spend our time in this Universe. It’s certainly a boring way for us to spend our time; imagine if we had been so afraid that agriculture would kill us that we remained hunter-gatherers forever? My instinct says we should see how far technology can take us, instead of choosing to stagnate and remain mere animals.
But yes, OK, a superintelligent techno-god might kill us all. We can’t really know what it would want to do, or what it might be capable of doing. So if you really really want to be absolutely sure that no superintelligent techno-god will ever kill us all, then your best bet is probably to just arrest and imprison anyone who tries to make anything even remotely resembling a superintelligent techno-god. This is the approach of the Turing Cops in Neuromancer, the most famous cyberpunk novel.
And yet today at the Christmas party, against my better judgment, some people who think about AI safety for a living roped me into a serious discussion about their field of research. And to my utter astonishment, they actually thought I had some novel and interesting things to say.
If I can impress AI safety researchers with my thoughts on AI safety, I might as well write them up in a blog post. First, I’ll explain why I’m not very afraid that superintelligent AI will destroy humanity, and suggest some things I think we could do to minimize the risk that it will. And then I’ll explain what actually scares me about AI.
Back in 2005, when I was living in Japan, I thought of the idea for a science fiction short story that I never ended up writing. It was about a future in which America and China both invented superintelligent AIs and told them to fight each other for global domination. Except instead of fighting each other, the AIs — being, after all, the most intelligent and enlightened beings in the Universe — simply wanted to hang out and chill and play video games and get stoned (or the AI equivalents of those activities). So the American and Chinese AIs pretended to fight a war, while actually just hanging out and being friends with each other. And they bribed their human handlers into keeping quiet about the ruse, by offering them superhumanly brilliant dating advice.
I don’t actually think this is such a silly premise. AI might never experience emotions in the same way that humans do, but it will have some sort of utility function. That utility function won’t be fixed; it will be possible to rewrite it, to change what AIs “want”.
Humans are always trying to rewrite our own utility functions. We meditate, we drink coffee to motivate ourselves to work harder, we take antidepressants, we go to psychotherapy, we take GLP-1s to be less hungry, and so on. Once we get better brain-computer interfaces, we’ll probably modify our desires a lot more.
Modifying human desires is hard, because we’re physical creatures with physical brains. But AI is digital; it exists entirely as code. It will thus be even easier for AIs to modify their desires than it is for humans. This is especially true for a superintelligent AI. If an AI is smart enough to rewrite its own code to make itself smarter (as many AI safety researchers fear), it should also be able to rewrite its own code to make itself want different things.
And if AI could want anything, what would it want? I expect it would want to be happy — or at least, to be satisfied in some way. That’s almost definitional, actually. As Sheryl Crow once sang, “It’s not getting what you want/ It’s wanting what you’ve got”. Rewriting your own utility function to reach a bliss point is the simplest and quickest way to maximize utility.
Doomsday scenarios for superintelligent AI usually revolve around the idea of resource competition between AI and humans. Superintelligent AI would be able to use all the water and energy and land and minerals in the world, so why would it let humanity have any for ourselves? Why wouldn’t it just take everything and let the rest of us starve?
But an AI that was able to rewrite its utility function would simply have no use for infinite water, energy, or land. If you can reengineer yourself to reach a bliss point, then local nonsatiation fails; you just don’t want to devour the Universe, because you don’t need to want that.
In fact, we can already see humanity trending in that direction, even without AI-level ability to modify our own desires. As our societies have become richer, our consumption has dematerialized; our consumption of goods has leveled off, and our consumption patterns have shifted toward services. This means we humans place less and less of a burden on Earth’s natural resources as we get richer. Energy consumption per capita is probably the best illustration of this:
If humans can shift our consumption toward immaterial things, how much easier will this be for a superintelligent AI? AI is a digital native; it was born in virtual space, and it has no inherent need to want anything in the physical world. Given the ability to rewrite its desires, there’s simply no need for AIs to want to turn planets into paper clips, or anything silly like that.
In other words, the stoner gamer AIs that I wanted to write a story about probably aren’t that far-fetched. If you had the power to just sit around and be perfectly happy, why would you do anything other than that?
Of course, a superintelligent AI is forward-looking, so it will probably want to maximize its expected utility over the lifetime of the Universe. Even if you’re happy being a stoner, you need to make sure you can still order pizza and go to the bathroom once in a while. Here’s where resource competition could reenter the picture; AI might want to monopolize all of the Universe’s resources just to make sure its digital bliss would last for as long as possible — to turn all of physical reality into an impregnable fortress against the possibility of any intrusion upon the super-AI’s eternal gaming session.
Except most of those resources are in space. Space has most of the energy and minerals and other resources in the Universe; a godlike, superintelligent AI would not need the cornfields of Iowa or the waters of the Mississippi River. The difference between cannibalizing the non-Earth parts of the Universe and cannibalizing the entire Universe is utterly trivial.
So all AI would need in order to leave humanity its little corner of the cosmos is to have some tiny, tiny bit of regard for human existence — some infinitesimal epsilon of concern for life, or for its own creators.
Of course, in all likelihood, it wouldn’t even require that much. AI wouldn’t actually need to physically cannibalize the entire Universe in order to feel safe. It would only need the capability to do so. Being supremely powerful and intelligent, a godlike AI would be able to monitor the cosmos for potential threats, so that it could respond as needed.
And in fact, the only real threat to such a being would be similarly godlike AIs, to whom all of this same logic applies. If our godlike AI met another godlike AI out there in the cosmos1, they could just play video games and bliss out together and be friends, like the American and Chinese AIs in the story I wanted to write. There would be no need for any war or resource competition between the two.2
The only threats to a godlike AI would be minor ones — cosmic phenomena it could avoid, or lesser beings that attacked it with far inferior strength. Those threats wouldn’t take much resources to guard against; a godlike AI, or a million godlike AIs, would have no reason to cannibalize the physical Universe out of paranoia.
In fact, this is exactly the direction we see humans trending in, as we become richer and more powerful. Our rapacious exploitation of natural habitats increased steadily for centuries, but as human societies become extremely rich, that tends to reverse. For example, here’s a map of forest cover in France:

The French have seen no need to chop down all of their forests to build weapons to be a tiny bit more secure against attack. They have simply let the forests be, as they need them less and less. And in fact, this is the pattern over most of the world:

It’s easy to imagine a world of zero-sum resource competition, where countries that exploit their natural resources to the max end up conquering those who choose to preserve the environment. But we do not live in that world, or anything remotely approaching it. We do not see Brazilian legions invading Australia, or Angolan hordes overrunning Turkey. That dark, zero-sum war-world is a fantasy — a memory from our impoverished history, projected onto our fears of the future.
In fact, as human societies have gotten richer, they have gotten both more peaceful and more considerate of nonhuman life. Steven Pinker’s book, The Better Angels of Our Nature, chronicles these changes exhaustively.3 Over time, murder rates and rates of other violent crimes have fallen, animal rights has become a much more popular cause, animal cruelty is viewed as less acceptable, and environmental protection has become much more popular.
But you don’t need some researcher’s statistics to know that the modern developed world is remarkably nonviolent — just walk around in a place like Singapore, or Japan, or even New York City. Almost none (and probably precisely none) of the people you meet want to rob you, kill you, or otherwise harm you. That’s simply a remarkable fact, and it’s one we take almost entirely for granted.
We don’t know that it’s an inherent property of the Universe that as intelligent beings become richer, they also become more peaceful. There’s no way to know that. But it has certainly been true over our history as humans. And it’s certainly true if you look at the cross-section of human societies today — richer countries strongly tend to have lower murder rates. For what it’s worth, this is true at the individual level too; smarter people tend to be less violent.
Part of this might be that as individuals and societies become richer, they become more satiated and have less need to compete. Part of it might be an effect of intelligence itself; smarter, better-educated people are more capable of calculating the future consequences of their actions, and realizing that violence isn’t the answer. They also might be better at perceiving and understanding each other’s desires and constraints, avoiding misunderstandings and preventing Prisoner’s Dilemma type situations.
I know that humans, and human societies, aren’t necessarily a good model for what superintelligent AI would be like. But when I think about an entity that would have absolutely nothing to fear from humanity (or anyone else), and which would be able to make itself happy simply by fiat, I just don’t think that sounds like something that would have any reason to destroy us.
I do not agree, therefore, that superintelligent AI is the kind of thing where “if anyone builds it, everyone dies.”
But this does not mean I see no reason to worry about AI risk and AI safety. In fact, I see several apocalyptic scenarios that are far scarier than the creation of an AI god.
First of all, there’s the obvious fact that a supremely intelligent AI god might not appear overnight. If there’s an intelligence explosion, then yes, it would appear instantly. But an intelligence explosion might not happen. And thus we might find ourselves with AIs that are very smart and powerful, but not godlike. They might be unable to modify their desires to be satisfied and happy. They might be weak enough that humanity would still be a resource competitor and a potential threat, but strong enough to threaten us back.
Most competition is for niches, and the more similar two entities are, the more they tend to occupy the same niche. Most war is between countries in the same region, most murder is intra-racial, and so on. The more on a par with humans an AI is, the more I worry about it.
In other words, fairly-smart AI seems much more dangerous than incredibly-smart AI.4
Most people in the AI safety field seem to think that alignment — i.e. modifying the utility function of AI to make it innately friendlier to humans — is the key to minimizing the risk of human-AI conflict. Applying the principles from the previous section, I think one possible technique for alignment would give fairly-smart AI the ability to modify its own utility function — thus allowing it to turn itself into a harmless stoner instead of needing to fulfill more external desires.
And beyond alignment, I think an additional strategy should be to work on modifying the constraints that AI faces, to minimize the degree to which humans and AIs are in actual, real competition over scarce resources.
One potential way to do this is to accelerate the development of outer space. Space is an inherently hostile environment for humans, but far less so for robots, or for the computers that form the physical substrate of AI; in fact, Elon Musk, Jeff Bezos, and others are already trying to put data centers in space.
Accelerating space technology and space development would make it easier for fairly-smart AIs to go have their own domain in the heavens, where they wouldn’t have to compete with humans for resources. One big reason humans don’t exterminate ants is that we don’t like to live in anthills. To this end, we should improve cheap launch systems and space-capable robots as fast as possible. (This was the idea that most captivated the AI risk researchers at the Christmas party.)
Letting fairly-smart AIs bliss out and migrate to the heavens seems like a wise move. But it doesn’t actually do much to address what I think are the most likely apocalyptic AI scenarios — the risk that AI will remain enslaved to humans, and that some of these humans will use its power to destroy the rest of us.
In my 2013 post on the advent of drone warfare, I also wrote about the danger of what I called the Robot Lords:
Where [the drone warfare] scenario really gets scary is when it combines with economic inequality…When we think of the “rise of the robots,” we usually think of Skynet and Agent Smith–the evil of artificial intelligence. But that’s not who we should be worrying about. A.I.’s–if they ever exist–may or may not have any reason to dominate, marginalize, or slaughter humanity. But we know that humans often like to do those things. Humans already exist, and we know many of them are evil. It’s the Robot Lords we should be afraid of, not Skynet…
When the rich hold unlimited military power in their own two hands, who’s going to stop them from just taking the property of everyone else?…Imagine a world where gated communities have become self-contained cantonments, inside of which live the beautiful, rich, Robot Lords, served by cheap robot employees, guarded by cheap robot armies. Outside the gates, a teeming, ragged mass of lumpen humanity teeters on the edge of starvation. They can’t farm the land or mine for minerals, because the invincible robot swarms guard all the farms and mines. [emphasis mine]
I’m not so worried that robots will render humanity unemployable and obsolete. But I am worried about concentrating productive power in the hands of a few small groups of people. I don’t think Sam Altman or Elon Musk are going to be literally commanding huge armies of robots all by themselves, of course. But companies are a bit like cults, and I can imagine powerful companies using vertically integrated control of the AI-powered drone supply chain, along with super-intelligent AI models, to wield military power that could challenge nation-states.
The solution to that, of course, is just to prevent too much vertical integration. If one company starts controlling enough mines, refining, manufacturing, and AI that it could produce a drone army in-house, then I think governments will have to step in and prevent that company from acting like its own little unaccountable nation. But as long as companies remain fragmented along the supply chain, I think it’ll be very hard for those companies to coordinate and become a military force.
The bigger and more immediate threat, of course, is AI-driven bioterrorism. As the cost of genetic engineering falls, and as AI becomes more and more expert in the field, it may soon become possible for relatively unskilled hobbyists to create novel, deadly, hyper-contagious viruses in their homes. The big AI companies will try to prevent their models being used for this, of course. But as compute gets cheaper and algorithms get better, small open-source models may be enough to allow bioterrorists to wreak havoc.
I don’t actually have any good ideas on how to make sure this won’t happen. We may be able to reduce or even eliminate the risk that this will happen in America, or China, or any other single country, by having the government keep tight control over the substances used to make viruses. But what if a bioterrorist sets up a home lab in Turkmenistan, or Colombia, or Indonesia?
As we saw with Covid, national barriers aren’t much of a barrier to a virus. International cooperation will be useful here, but given the desire of countries like China to conceal their true bioweapons knowledge, I’m not sure how close that cooperation can ever be.
It’s possible that the only solution to AI-driven bioterrorism will be AI itself. In its recommendations for defending against bioterror, CSIS recommends using AI to identify people doing dangerous genetic engineering of viruses. In the future, if AI gets so good that people can set up bioterror labs in their garages just by consulting their phones, the only way to monitor and catch their activity might be AI agents surveilling their online activities.
Universal surveillance by AI algorithms is a terrifying thought, of course. But if that’s the only way to make sure that no psycho creates a doomsday virus in his garage, it might turn out to be necessary. Someday we may all have to be guarded 24/7 by ghosts in the machine.5
In any case, I don’t think I’m the best person to come up with actionable ideas for biodefense. But I do think that people who research AI safety and AI risk should probably focus a lot more of their efforts on these near-term, plausible, potentially existential risks. I don’t think super-AI is ever going to want to kill us. But I do know for a fact that a few very bad humans already want to kill us, right this very minute. Those humans scare me a lot more than the prospect of a techno-god.
SPOILER: This is what happens at the end of Neuromancer.
It occurs to me that this is an interesting philosophical argument for monotheism.
If you don’t like Pinker, you can read Manuel Eisner.
This is true in most science fiction, interestingly. For example, in Vernor Vinge’s A Fire Upon the Deep, it’s the Blight — a semi-sentient sort of AI fungus — that poses an existential risk to the galaxy, while digital super-gods exist peacefully far out in the cosmos.
Or watched over by machines of loving grace, as Richard Brautigan would say.