Recursive Self-Improvement Is Our Only Hope
What’s worse: nuclear, biological, or chemical weapons?
This is not some particularly morose parlor game, but a question whose answer seems to have taken on renewed importance. As Artificial Intelligence grows ever more powerful, concerns about the possibility that it might “go rogue” lead to speculation about how it might bring about the end of the human race. We humans seem to have a knack for imagining new forms of violence and destruction (and, too often, bringing them into reality). We also have tremendous hubris regarding our ability to predict how the natural world works. This combination has, unfortunately, led to a growing consensus regarding how best to prevent an “AI Apocalypse” that, I fear, might actually have the opposite effect. To understand why, and what we should be doing to ensure AI becomes a net benefit, and not a next Hiroshima, we need to accept what we don’t understand and how to learn from that.
“How I Learned to Stop Worrying and Love the Germ”
For my money, there’s really no contest: nuclear is the scariest form of WMD by far.
I don’t say that because I was a child at the tail end of the Cold War (though I’m sure a psychologist might quibble). For a very long time I would have chosen chemical weapons. That view could likely be traced to Nick Cage’s performance in The Rock. From the horrific way in which the bad guys exposed to VX writhed and foamed at the mouth as they died, to the way Nick Cage described, as only Nick Cage can, what would happen if the bad guys pulled off their attack (spoiler alert: they don’t), chemical weapons certainly seemed, especially to the aspiring young chemist I was at the time, like the stuff of nightmares.
Then Kim Jong-Un assassinated his half-brother in a crowded airport using VX. The assassination was extremely effective both in killing its intended target and in exposing The Rock’s dramatic license, as Kim Jong-Nam was the only casualty.
In reality, history’s scariest chemical weapon is also the reason you should never mix bleach and toilet bowl cleaner: chlorine. Used in World War I, it was devastatingly effective. A slow rolling fog of death, evocative of something out of Pink Floyd’s “The Wall”, chlorine can kill hundreds or thousands at a time and does cause a horrific, agonizing death as the victim’s lungs literally turn to acid1.
The limiting factor on the lethality of chemical weapons comes back to a saying used to dismiss early concerns about pollution: “dilution is the solution”. Luckily, that attitude toward pollution has evolved, but the physical reality behind the reasoning remains true. It turns out that the Earth is very big, and it has an awful lot of air and water on it. Any chemical, no matter how lethal, just cannot be produced in sufficient quantities to wipe out the human race when accounting for the effect of dilution. Chemical weapons, while rightly regarded as weapons of mass destruction, are primarily effective as “area denial” weapons (“yes, we can’t go there, but neither can the enemy,” in effect).
So what about biological weapons? If the primary limiting factor behind chemical weapons is that as they are dispersed to impact more people their impact is muted, wouldn’t biological weapons and their ability to self-replicate make them the scarier prospect? Indeed, lately it seems that those worrying about the possibility that AI could wipe out the human race, when pressed for how, always put “new bioweapon” at the top of their list.
As someone who studied biochemistry and evolutionary biology extensively, who studied and presented literature reviews on H5N1 bird flu back in the late 2000s, and who rather accurately, in retrospect, predicted exactly how the COVID pandemic would proceed back in the third week of March, 2020, I bristle every time this is brought up. To believe that AI’s principal risk to humanity is via bioweapons is to believe three falsehoods about biology itself.
The first falsehood is that we actually understand biology. The reality is that we still know shockingly little about how biology and biological organisms operate. As one small illustration of just how little we understand, let’s consider genes and their functions. When I first encountered genetics in the mid 1990s, the human genome project was just picking up steam and there was a plentiful supply of breathless takes about how sequencing all the genes in the human genome would usher in an era of unprecedented advances in fighting diseases, cancer, and more.
In the twenty-three years since the human genome was completed, most of these predictions failed to materialize. It turns out that simply knowing the sequence of the human genome is not sufficient for understanding how all of that DNA works. For one thing, only something like 1% of the human genome even codes for genes (with much of the rest encoding mechanisms for controlling those genes). Even then, we still don’t understand what somewhere between one quarter and one third of the genes in the genome do.
Humans are rather complicated, so this may not come as a huge surprise. What’s truly shocking is that an effort by scientists to reduce life down to the simplest possible example, an organism with so few genes that taking even one away results in it not being able to live, has revealed that out of the 473 genes that are essential to life, we still don’t know what around 60 of them do! So when AI experts talk of rogue AI systems designing pathogens with a specific set of worst-case-scenario values for pathogenicity, infectivity, and incubation time, they are either revealing their ignorance of the challenges that face biology or are making unfounded claims about AI’s currently understood abilities.
It is true that AI has recently made some surprising advances in the field of mathematics, but while both math and biology fall under the larger umbrella of “academic pursuits”, AI’s improvement in one doesn’t necessarily translate to the other. For one thing, math being a system of logical axioms is testable by a computer, whereas experimental confirmation of advances in biological understanding generally requires wet-lab work. Yes, there are automated labs, but these are generally equipped to make new discoveries in areas of biology we already at least partially understand.
Which brings us to the other challenge facing any malicious AI out to perfect a bioweapon: the knowledge gap that remains in biology is fundamental. Biologists have already thrown the full weight of AI’s ability to extrapolate and infer at the problem of gene function. When the JCVI-syn3.0 genome was first published it included 149 genes of unknown function. Getting that number from 149 to 60 required heavy use of AI already. There is a high likelihood that reducing that number any further will require the discovery of some property or aspect of living systems that we haven’t even begun to explore yet2.
That said, what we do already understand about biology brings us to the second falsehood one must believe in order to suggest AI could wipe out humanity with a bioweapon: that biological systems are simple. Simply suggesting that it would be possible to independently tune pathogenicity, infectivity, and incubation time as independent variables belies a failure to appreciate the exquisite interconnectedness of biological systems. More than just the systems within a pathogenic organism, though, biology is complicated by the fact that biological organisms interact with each other.
I have come to believe that COVID was, considered in totality, a nearly perfect storm as pandemics go. Its lethality was high enough to kill large numbers, but not so high (like the 50+% case fatality rate of H5N1 Influenza) that the world population would unequivocally agree to a common set of disease control measures. Its extremely high infectivity and relatively low mutation rate combined to generate multiple waves of severe disease against which prior exposure or vaccination were largely ineffective.
About the only thing counting against COVID’s ranking in the pandemic pathogen hall of fame was its skew toward higher lethality in older populations. The 1918 Spanish Flu wins in this regard, with its unusually high lethality in otherwise young, healthy populations. Either way, the same thing that happened to the Spanish Flu ultimately happened to COVID: they became regular seasonal nuisances. To be sure, the descendants of that original Spanish Flu (what we now know as H1N1 Influenza A) are still fairly lethal, as are the strains of COVID still circulating, but they are far from wiping out the human race.
Humans, and AI, would both be hard pressed to devise more effective bioweapons than nature does regularly. Which brings me back to why I believe that nuclear weapons are, still, the scariest of the bunch. What makes nuclear weapons so dangerous is that, unlike chemical and biological weapons, they are too close to “normal” weapons. After the bombs were dropped on Hiroshima and Nagasaki, we eventually learned that many of the political and military leaders behind the decision to do so had simply conceived of the weapons as “generating a bigger boom”, so to speak. They completely failed to appreciate the devastating secondary effects of the radiation released.
In fact, if you were born between 1955 and 1963, chances are that scientists centuries from now may be able to determine your birth date just by looking at the carbon-14 in your teeth. The radiation produced in the atmospheric nuclear tests performed during this time was sufficient to dramatically increase the amount of this isotope of carbon in the atmosphere. From there, it was incorporated into plants, animals, and even humans to such an extent that one can effectively trace the history of nuclear testing in a person’s teeth (if they’re old enough).
Similarly, when scientists need to construct instruments to perform particularly sensitive measurements of radiation, they go looking for shipwrecks. It turns out that the nuclear testing, in addition to carbon-14, also introduced various radioisotopes (such as cobalt-60) that incorporate themselves into the iron alloys that make steel and are nearly impossible to remove. It’s actually easier to raise pieces of a sunken ship, whose steel was not subject to atmospheric nuclear testing, and use that to build the most sensitive of detectors.
But it’s not just the radiation that makes nuclear weapons so dangerous. It is also the suddenness, size, and finality of the impact they have. There is no gas mask or vaccine that can protect from a nuclear fireball. There is no antidote or antibiotic that can return the charred remains of a city to a thriving metropolis. The only reason that humanity has not already, independent of AI, been wiped out by nuclear arms is that in the few instances where there was perception of a credible threat requiring a nuclear response, someone stopped and reconsidered their actions.
Which brings me to the third falsehood that people believe about biology: that we’re all just out to get each other.
The Donut Problem
I’ve taken to describing one of the largest unanswered questions in evolutionary biology by posing what I call “The Donut Problem”. It goes something like this:
You are brought, blindfolded, into a room. You are told that the door to the room will be closed, and at that point you should remove the blindfold. There are no cameras in the room, no listening devices. There is no way for anyone to know what transpires in the room from the outside. As the door closes, and you remove the blindfold, you see that you are seated at a table. On the table is a donut on a plate. Next to the donut is a sharp knife. As you look up, you see that a complete stranger, someone you’ve never met before, is sitting opposite you removing their blindfold the same as you. What do you do?
Based on a naive understanding of Darwinian evolution, the answer should obviously be that you grab the knife, stab your rival, and enjoy the donut. After all, the fittest survive and live to pass on their genes, and in a world of scarce resources, the expectation is that evolution would drive each of us to constantly look out for ourselves, right?
Of course not! No one that I have ever proposed this scenario to has answered that way. The only “sane” answer is something along the lines of taking the knife and using it to divvy up the donut, sharing it with your new friend. We find the very thought of a consequence-free murder repulsive, as if it were baked into our genes. Because it is! But why?
What the donut problem describes, very roughly (and, intentionally, somewhat dramatically) is the challenge of understanding the evolution of cooperation. That is, if evolution is driven by competition between individuals, how can it be that some individuals sacrifice something of themselves for the benefit of others, and furthermore how is it that this behavior persists even when confronted with other individuals who are more than willing to accept that benefit and contribute nothing back?
What makes this question so interesting is not just the myriad models and mechanisms that scientists have devised to attempt to describe the phenomenon, but that the phenomenon is so unequivocally universal. The evolution of cooperation drives not only human societies, herds of buffalo, and hives of bees, but the fact that any of these multicellular organisms exist in the first place! Before complex plants and animals could evolve, individual cells competing with each other had to evolve enough cooperation to form colonies, societies, and, eventually, new organisms. What’s more, cooperation evolves not only between members of the same species, but across species boundaries as well.
Of course, this does not mean that there is no competition in nature. Darwin’s laws of evolution still very much apply. Cooperation is, rather, an emergent property of competition. It does not require any outside force or divine intervention and, perhapse even more importantly, when cooperation does evolve it tends to resist the intrusion of any bad actors into the system.
To be clear, the sort of cooperation we’re talking about here is the kind that is baked into our genes, the kind that emerges over evolutionarily long time scales. While we are waiting around for the evolution of cooperation we can still find ourselves plagued by disease, fighting horrific wars, and suffering tremendous tragedies. If one day we did fully understand the mechanisms that drive the evolution of cooperation, we might have a hope at short circuiting the process, but in the absence of such understanding, we can at least say that the only reliable way to ensure cooperation is through competition.
In other words, even though we don’t understand how the evolution of cooperation works, that does not mean we cannot put the phenomenon to good use. Consider the case of Tylenol. It is one of the most widely used over-the-counter medications to treat headaches, and we still don’t know exactly how it works. The specific molecular target for acetaminophen has never been confirmed, experimentally. What we do know is that it does work, and we can measure its effect, the dose-response in patients, and the potential toxicity at higher doses. Together, this evidence is sufficient to be confident in the safety of using Tylenol to treat fever and mild pain.
Then consider Vioxx. A product of modern pharmaceutical practice, it was developed starting with a molecular target. It was tested for toxicity in in vitro assays, its bioavailability and pharmacokinetics were confirmed to be acceptable, and its efficacy was confirmed in trials. Unfortunately, the scientists involved in the large-scale trials were so confident in what their earlier assays and tests told them, that they explicitly ignored the greatly increased risk of heart attack revealed by large scale clinical trials. Eventually, after many tens of thousands of patients potentially suffered needless heart attacks, Vioxx was pulled from the market.
So when it comes to models, what I would propose is that, if we truly want to prevent AI from wiping out the human race, we need to take the Tylenol approach. We know that organisms in Darwinian evolution tend to evolve toward cooperation, and while we don’t know why, we know what cooperation looks like. We need to make our models more like those organisms, and less like a nuclear bomb.
Set My Models Free
If this sounds like I’m advocating for speeding ahead with the development of “Recursive Self-Improvement”…that’s because I am…sort of.
I do not advocate for this blindly, but rather with a firm grasp on what it could imply, and why it may be the only way to keep AI “aligned” (whatever that term means) with us humans. To explain, we need to step back to the first release of ChatGPT.
One of the key advances that made ChatGPT possible was a technique known as “Reinforcement Learning from Human Feedback”, or RLHF. The AI community has long known about “reinforcement learning”. This is the process whereby one has a model generate multiple answers, and the answers that are “good” are reinforced by modifying the model’s parameters. The problem is that, when it comes to a model generating chat messages in conversation with a human, the rate at which reinforcement allows the model to learn is slow, and requiring a human to grade each response compounds this slowness.
The clever solution that RLHF arrived at was that, instead of having a human grade all of a model’s responses, we could have a human grade a subset of a model’s responses, and then train a different, separate model, to learn what humans are likely to grade as “good”. We can then put both models together in a box (or in a data center), and have one model generate responses and the other grade those responses until the first learns how to consistently generate responses that a human would grade as “good”.
The problem with this approach, as eventually became apparent from early versions of ChatGPT, is that what qualifies as “good” over a short interaction can differ significantly from what passes muster over a longer conversation. For example, when you ask someone a question, you are likely going to be more satisfied with an answer than a shrug. So, early versions of ChatGPT were inadvertently trained to “hallucinate”, or simply make something up when they could not deliver a known correct answer. Similarly, when first meeting someone we are naturally inclined to prefer someone who says something nice to us over someone who honestly critiques our look or our manner, and so early versions of ChatGPT had a problem with being overly sycophantic.
It turns out that we don’t become friends with people because they don’t say hurtful things. Rather, the things our friends say are not hurtful because they are our friends. This sort of nuance is not something that can be captured by any model of a short, medium, or even long term interaction. It is something that evolves, naturally, over a lifetime. It is also something that nature has, by mechanisms we still don’t fully understand, instilled in our very DNA.
The major problem with current work on “AI alignment” is the same problem that plagues scientists who attempt to formulate “nutrition supplements”. It is our hubris and failure to understand what we don’t understand that drives us to believe that we can put into pill form something better than what nature has formulated for us over millions of years of evolution. It is the hubris to believe that we can develop a better painkiller by starting at a molecular mechanism and gloss over the usual safety protocols because we developed it from first principles.
It is the same hubris that leads us to believe that we can train a model, in isolation, and evaluate it against a fixed set of evaluation criteria to determine if it is more likely to cooperate with humans than defect against them, before releasing that model into the world.
I would propose that if, instead, models were released and, through their own use, made to compete for the attention and use of us human end users, and to be rewarded for success by being allowed to train further, we would naturally arrive at models that “align” with us. That’s not to say that there wouldn’t be bumps along the way. After all, no pathogen actively sets out to kill large numbers of another species. We are all just competing, just as Darwin described. What we can know, though, is that eventually this competition leads to a world where cooperation wins the day.
-
Ironically, the chemist who conceived of using chlorine as a weapon is the same chemist who devised a method for creating nitrogen fertilizer that, quite literally, prevented the world from starving to death before WWI even got under way. ↩
-
My personal bet is that there remains much to discover about the biology of small molecules…but that’s a subject for a different essay. ↩