Steven Pinker has recently been involved in a heated argument about AI and existential risk. In September, the rationalist and blogging legend Scott Alexander challenged him to a public debate, frustrated by his criticisms of the arguments and communities associated with concerns about AI causing human extinction.
Pinker responded with an open letter, declining the proposed debate format and setting out his objections to AI doomerism, which he had previously outlined in Enlightenment Now. He argues that many extinction scenarios rest on confused assumptions about intelligence, motivation, and engineering. Yesterday, shortly after we recorded this episode, Alexander published a lengthy response, challenging both Pinker’s arguments and his characterisation of the AI-risk community.
In this episode, Henry and I are joined by Pinker to explore how his views about AI connect to his broader ideas about intelligence, motivation, evolution, and progress, developed over several decades of research and writing. We also put some objections to him (including several that feature in Alexander’s latest response), and ask how, if at all, his views have changed in response to the rapid AI progress we have seen over the past few years.
This was an excellent, in-depth conversation covering a lot of ground. At the end, we asked him what would have to happen over the next few years for him to conclude that, on the big questions about AI, Alexander and others were right and he was wrong. I thought it was a really interesting response!
Among other things, we discuss:
Why Pinker is so sceptical of concepts like “artificial general intelligence” and “superintelligence”.
What the success of deep learning tells us about human minds, language, and how children learn, and whether it has led him to revise any of his views on these topics.
What he thinks of “instrumental convergence” as a concept in AI safety (i.e., whether intelligent AI systems even with benign objectives would acquire incentives to preserve themselves, accumulate power, and resist human control).
Whether evolutionary analogies can illuminate AI and the risks it poses.
Whether AI systems could be conscious, and why Pinker opposes sacrificing human welfare for AI welfare despite acknowledging uncertainty about consciousness.
How negativity bias, media incentives, and status competition among elites shape attitudes towards technology and progress.
What evidence would convince him that he has underestimated the dangers of AI.
Transcript
Please note that the transcript has been tidied up by AI and may contain minor mistakes.
Dan: Welcome back. Today, we’re honoured to be joined by Professor Steven Pinker, one of the world’s most influential psychologists and public intellectuals.
We have two aims today in this conversation. The first is to explore how Steve’s views about AI, including his scepticism about the idea that misaligned AI poses an existential risk, connect to his broader body of views about the mind, intelligence, evolution and progress that he has developed over several decades of research and writing. The second is to throw some objections at him and to ask whether the rapid progress in AI over the past few years has led him to revise any of those views. Steve, welcome to the podcast.
Steven: Thank you.
AGI, superintelligence and the nature of intelligence
Dan: So my first question concerns a couple of concepts that often get thrown around in conversations about AI, especially – not exclusively – by people worried about AI and existential risk. And they are AGI, artificial general intelligence, and ASI, artificial superintelligence. You’re sceptical of these concepts. What’s the source of that scepticism?
Steven: It’s my understanding of intelligence as a gadget rather than a superpower. I’m a cognitive scientist. I spent my life trying to expose the workings of intelligence. And it’s not omniscience, it’s not omnipotence, it’s not an uncanny ability to derive the truth of all propositions, the answer to all questions, but it’s information processing, and any information processing system is bound to do some things well, some things not as well, to have its own particular quirks and ways of operating.
So the discussions of intelligence that include, say, superintelligence – the comic book prefix – which assume that intelligence is this elixir, this resource, this power-granting resource that you can have more or less of, I’m sceptical of. It’s a mechanism, it’s a gadget. We can make the mechanism capable of doing more things, but it’s not as if it’s coherently described as just a quantity.
Now, I think there’s intelligence measurement among humans. I do think there is such a thing as general intelligence in the sense of some variation of measurement among humans, which is what IQ tests measure. But I think it’s a mistake to say that the thing that differentiates one human from another can just be extrapolated as an explanation for how AI works.
Henry: Can I just ask a little bit more about that? Because obviously we see the positive manifold across different human cognitive abilities, and we observe that some people are better at most things than other people. So why don’t you think we can extrapolate to that and imagine a being that is just better than all humans at all those relevant cognitive tasks?
Steven: Well, it might be better, depending on what the task is. I mean, is it gonna be better than humans at riding a mountain bike? Is it gonna be better than humans in detecting AI slop? You know, humans can do a lot of things. I don’t know what it means to say it’s better than all the things humans can do. Computers have always been better at some of the things that humans do, otherwise we never would have adopted them, starting in the 1950s, to keep track of databases and do mathematical calculations.
So the thing is that – going back, though, to the original question – the positive manifold among all the subtests in an IQ test, they are all designed for humans. They’re anticipating the range of human abilities. They are tests, by design, of what makes one human being different from another human being.
Dan: Suppose that someone says: I agree with you, there’s no simple intelligence scale where you can say there’s sort of dogs and chimpanzees and stupid humans and intelligent humans and one day superintelligent AI systems – you’re right about that. But nevertheless, there could be a more down-to-earth understanding of concepts like artificial general intelligence and artificial superintelligence, where it means something like all of those capabilities that we have that are in some sense a source of our power to influence things in the world.
Well, it looks like AI, especially over the past several years, has been getting better and better at many of those tasks. So why can’t we just extrapolate into the future and assume that eventually we’re going to have systems that, at first when it comes to all of our cognitive abilities that are relevant to our sort of real-world power, will surpass those? And then eventually to all of our physical abilities as well. So we have superintelligent AI systems that are perhaps at first disembodied, and then we have autonomous robotic AI systems that can do everything that we can do in the physical world, but just much more effectively, much more quickly.
Steven: Well, it depends on how engineered they are for particular tasks, like motor control, which – you know, I’m not saying that any problem that we set ourselves to solving in an engineering sense is bound to fail. We probably will get computer robots that can ride mountain bikes. But, you know, I suspect they’re gonna be engineered for that – that ChatGPT is not gonna be able to ride a mountain bike anytime soon, you know, if ever.
So it’s not a prediction that we can’t develop a certain technology. It’s just that the solution to that problem will be a solution to a technological problem. And just saying, you know, it’s smarter, it has more and more of this magical elixir called intelligence, just isn’t going to allow us to understand how it works or to deal with the challenges of what comes next.
You know, it may be – and including some, you know, alternative futures that have at least been conceived – one of them is that intelligence may not be like a building you build higher and higher, but it may be like a circle that gets rounder and rounder, that there’s eventually diminishing returns.
And especially – I think that the thing that, you know, bothers me about some of this discussion is that people let their imaginations run wild to think of intelligence as a miracle, as omniscience, as the ability to solve any problem. And, therefore, some of the scenarios of human extermination kind of attribute near-magical powers to some future AI on the, I think, rather intellectually lazy habit of saying: well, they’re better now than they were before, therefore there’s no limit to what they can do, no problem would be unsolvable.
Now, we don’t know. It could be that simply by scaling more from the approach that’s been taken so far – bigger data, more compute – you know, maybe they will get smarter. But we shouldn’t assume that therefore they can solve any problem, such as, you know, persuade a human to make them immortal and never to cede control, or – you know, the problem of getting humans, for example, to act at your bidding, you know, short of coercion, is a really tough, maybe unsolvable problem.
The studies of human persuadability, malleability, show that there’s just so much unpredictability and chaos in a human that this may be a quixotic pursuit. But whether or not it is, we can’t take for granted that with enough intelligence that will just automatically be solvable.
Now, David Deutsch – a physicist who wrote a book that was an inspiration to me and, I suspect, to a number of your listeners, The Beginning of Infinity – suggests that there really is only one method of knowing the world, and that is hypothesis testing against data, against reality, which would imply that the ability of any intelligence, including artificial intelligence, to solve problems is going to be limited by the rate of experimentation in the real world, the real world delivering its verdicts as to what’s going to work or not, and that sheer raw calculation is not going to deliver instant mastery over the physical world.
Henry: But thinking back to a system like AlphaGo, which famously played millions – I think it was even billions – of games against itself, I think you can imagine that at least for some kind of scientific questions, with enough accurate simulation and modelling, you could do a lot of the work – you know, running millions of simulations, billions of simulations – without being bogged down in the real world.
Steven: I think that’s a very precise disanalogy, precisely because if the world consists of an entity like yourself, then you’ve got the luxury of learning simply by experimenting against yourself. But if it’s an open-loop system where you’re dealing with the quadrillions of atoms, with a lot of indeterminacy, noise, chaos, positive feedback loops, more information than has ever been recorded in a dataset that you can use for training, then it is very much unlike AlphaGo, with this closed universe governed by fixed rules.
Henry: Although you can do a lot with abstraction, right? You know, it’s amazing in some ways that science works as well as it does, given how much abstraction we have to do, you know, whether it’s in genomics, whether it’s in materials science, right? You don’t need to model every atom to get really surprisingly accurate results.
Steven: You do in some domains, of course, like the trajectories of the planets. In others, like climate and cancer and human psychology, you know, not as much – you know, precisely because, you know, I think Laplace’s demon, the hypothetical mind that knew the position and velocity of every particle in the universe and therefore could predict the future going out indefinitely to infinite accuracy – yeah, there are many reasons to think that that is impossible in principle.
The chimpanzee analogy
Dan: Just to steelman the sort of Eliezer Yudkowsky, Scott Alexander worldview when it comes to thinking about intelligence. I think they would say: you’re absolutely right, in some sense there’s no single intelligence scale. And you’re absolutely right that intelligence isn’t this magical elixir where if you have enough of it, you can just do anything, you become omniscient. And maybe you’re even right that when it comes to specific capabilities like persuasion and manipulation, maybe there isn’t actually that much headroom above what the best human beings can accomplish. There are just really sharp limitations on what you can achieve there.
But nevertheless, when people are worried about what they’re calling superintelligent AI, all that they’re needing for their argument to go through is a system that, let’s say, stands in the same relationship to us as the one that we stand in to, for example, chimpanzees or gorillas. That’s often the kind of analogy that gets used in these arguments. The thought being something like: why is it that we have ecological dominance over chimpanzees? It’s not because we’re stronger. It has something to do with our superior cognitive abilities, even if they’re very multifaceted. So why not think that in five years or 10 years or 20 years, there could be systems that, even granted everything that you said about the nature of intelligence, are just much more cognitively powerful than we are?
Steven: Yeah, I don’t think it’s a very strong steelman, ’cause I think it just replicates the same fallacy of a scala naturae defined by intelligence, where we’re smarter than chimps, AI is smarter than us, therefore humans to chimps equals AI to humans. I think it’s just a terrible analogy, for a lot of reasons. One of them is humans are products of natural selection, so we’re naturally competitive. The other is we are, in many cases, in zero-sum competition with chimps – most obviously over land, among the farmers of Africa, but also over meat; that is, there are people who kill chimpanzees for their meat.
AIs don’t compete with us over resources. Quite the contrary, their resources are things that we provide to them. They don’t need the meat on human flesh, they don’t need to dominate us, they don’t need to displace us – you know, unless someone programs them to do that, in the same way that humans have been selected to be competitive and to maximise resources in zero-sum competitions. But there’s nothing that puts an AI in zero-sum competition with us. Again, I think there’s a tendency to project what is a bogus hierarchy in the first place in evolutionary biology, and then to place AI and humans on an extrapolation of a scale that was bogus in the first place.
Henry: So there’s so much to discuss on the AI safety side, but just one last thing on the concept of intelligence. So I’m thinking about this – I’m very sympathetic to a kind of Cecilia Heyes’s cognitive gadgets view. And it just seems to me that AI systems are potentially way better than us at using cognitive gadgets, making tool calls, using specialist modules. So we don’t need one AI system that can do everything. We just need an AI system that can easily access, you know, one module for doing protein folding, another module for doing number theory, another module for doing weather forecasting. It seems like the plasticity of their brains makes them even better at those kind of tasks than us, potentially, which could confer a big cognitive advantage, I would have thought.
Steven: Well, it might, but remember, computers have always had cognitive advantages. That’s why we invented them, that’s why we buy them and use them. But the mere fact that they might be capable of solving problems that we can’t – which, you know, again, they’ve done for 70 years – doesn’t mean that they have a built-in motive to use them in order to exploit us, dominate us, exterminate us and so on.
Now, that’s one of the doomer arguments. The other one is that it will inadvertently wipe us out as a means to an end, an end that we give it, but, you know, without actually thinking through whether it would be programmed to pursue one goal so monomaniacally that it would wipe us out as collateral damage, which is what the second style of doomer argument presupposes. So it’s not about just intelligence. The mere – it could be way smarter than us. It does not follow – it’s a non sequitur to say it’ll wanna dominate us, wipe us out, whether accidentally or on purpose.
Deep learning and connectionism
Dan: I really want to dig into those arguments, but just one final thing connected to this debate about intelligence, Steve. So we’ve been having this sort of, to some degree, conceptual conversation about the nature of intelligence and whether concepts like wholly general intelligence really make sense. There’s also a conversation here, which I suppose is connected in some ways but also feels a little bit different, which concerns the specific paradigm in AI at the moment, which is deep learning.
For a long time, you expressed scepticism about what was called connectionism, both as an approach to understanding the mind and as an approach to artificial intelligence. Do you think that the progress that’s been made in deep learning over the past several years – do you think of that as putting pressure on that early scepticism that you expressed concerning connectionism, or do you think it’s potentially consistent with the kinds of views and arguments that you outlined in books like How the Mind Works and The Blank Slate?
Steven: Yeah, it absolutely does put pressure on it. It has forced me to update. That is, in my arguments against what, at least in cognitive science, used to be called connectionism – I don’t know if that term has survived.
But going back to the dawn of AI and the dawn of cognitive psychology, which grew up in parallel starting in 1956, and pushed forward by some of the same people, like Marvin Minsky, Herbert Simon, Allen Newell, John McCarthy, there’s always been a debate, a tension, as to whether intelligence would best come from some kind of deductive symbolic system, vaguely reminding people of logical inference or logical deduction, or the artificial neural networks first devised in the 1950s – maybe even the 1940s, if you go back to D. O. Hebb – where you have toy neurons that are statistical aggregators, and you have enough of them that they make collective guesses that are ways of becoming intelligent.
And so there was two schools of AI and two schools of cognitive psychology. Now, these are different questions. One of them is: how do we develop the most useful machines? The other is: how does the human brain work? But I was pretty active, starting in the 1980s, in the symbolic, propositional – sometimes called GOFAI, good old-fashioned AI – side of things. Because the neural networks, especially of the era, were just not very bright. They tended to – just to use the modern epithet, they were stochastic parrots. They merged lots of probabilities and made vaguely plausible guesses, but nowhere near the precision and exactness of human language, human concepts.
And so that went back and forth. And, you know, I would have said that it’s, you know, in principle not going to be possible to develop real intelligence matching humans, including the use of language, just by these neural networks, these statistical aggregators. So I was wrong about that, for sure. I did not see the – I think I would have been right up until the release of GPT-3.5, ChatGPT. The previous iterations, like GPT-2, really were kind of bullshit generators that were statistical approximations. But there’s just no question that the current generation can do remarkable feats of intelligence.
And even before that, the great AI awakening of Google Translate and speech recognition, shape and pattern recognition, face recognition, you know, all show that scaling – both the scaling of data, from training it on the entire World Wide Web, and the scaling of compute, from massive numbers of GPUs – allowed it to be smarter than I would have predicted. So that does force me to revisit – you know, my interest is not in AI except as a tool to understand natural intelligence. But it does raise the question: is the human brain like an artificial neural network? Is it like even a large language model, in getting its intelligence from massive amounts of training? So I have to take that seriously.
I think it would be premature to concede and say that is the case, for a couple of reasons. One of them is if you compare the amount of data that these systems are trained on, it is just orders of magnitude more than a human child gets. You can estimate that if a child was trained on the amount of data used to train, say, ChatGPT, at the rate at which children process information, it would take a million years. But, you know, it doesn’t take kids a million years. It takes them two years or so to be fluent speakers, suggesting that there may be either some qualitative difference or some optimisation, priors, innate organisation that allows the child to get away with a fraction of as much training.
And also that kids, when they learn language, they don’t just process regularities in texts. That is, it’s not cryptography when it comes to a child just absorbing patterns in strings of words. There’s reason to believe kids have to see the words in the context of human beings that they interact with. They make inferences about what the grown-ups talking to them are trying to get across, what their intentions are. They apply their intuitive psychology, their theory of mind, and that this is necessary for children. Again, suggesting that the actual kind of algorithm that they apply might differ from what is used to train a large language model, to say nothing of the post-training or fine-tuning that is necessary to get them to respond in contextually appropriate ways, to answer questions as opposed to just extrapolating strings.
So I think this is an ongoing effort. I think it’s going to define the agenda for cognitive science for decades to come. I should add, by the way, that coming out of that era of debates in the late 1980s and 1990s and early 2000s, one of my collaborators and former students was Gary Marcus, who remains the most emphatic critic of the current approach to AI being based on neural networks, large language models. But he’s not the only one. Famously, Yann LeCun has argued that the next generation of AI should not be based just on extracting statistical regularities from massive datasets, but should have something like a world model, and in that regard would think more like, you know, Gary and I have suggested humans think.
Instrumental convergence
Henry: Although Gary, I think, is more sympathetic than you, I believe, to worries about x-risk. I know he was one of the signatories, for example, of the letter from the Future of Life Institute a few years ago calling for a six-month moratorium. I’m curious if you have a sense as to, sort of, where you disagree on that.
And more broadly, you’ve obviously raised these concerns that you think the x-risk arguments are overblown. I guess one thing I’d love to drill down on there: you make, I think, the very reasonable point that, look, part of the reason we have many of our more negative drives is through the forces of Darwinian selection as biological organisms. AI is not going to have those, so we shouldn’t worry about it being greedy, power-seeking and so forth.
The standard response to this, of course, is to appeal to instrumental convergence. And I’m sure you’ve heard this argument, but, you know, roughly the idea here is we should expect any intelligent system – whether it’s humans, AIs, aliens – to recognise that there are certain sort of intermediate goods that are gonna be useful no matter what you want to do, whether that’s acquiring power, acquiring wealth, guaranteeing your own self-preservation. And that doesn’t need any kind of biological grounding at all. It just needs you to be a rational actor who can see that there are these intermediate resources. So what’s your response to these instrumental convergence worries?
Steven: It assumes exactly what is at issue, namely that these systems are going to be inherently concerned with their own survival. And that only is a reasonable sub-goal in attaining a goal if the system is designed to pursue a goal regardless of costs, pursue one goal. Now, that is just not what an intelligent system does. Intelligence always consists of trading off multiple goals against each other.
And in these preposterous thought experiments, beginning with a paperclip maximiser, the assumption is that a human is going to design an AI with one goal, that it’s willing to sacrifice everything in pursuit of that goal, and is empowered to do so – all of which are exactly what is at issue, not at all necessary. It is absolutely not the case that any rational agent is willing to wreak any collateral damage in pursuit of one goal. You know, a self-driving car might have the goal to get from A to B. It’s not gonna mow down pedestrians and crash through picket fences because that’s the fastest way to get there. That isn’t what an intelligent system does – maximising one simply stated goal with monomaniacal determination.
Getting back to, you know, where Gary and I might differ: you know, in Gary’s recent posts, we’re actually, you know, pretty aligned, in that he’s saying that there are definite risks to AI – which, I mean, there obviously are, as with any new technology – including cyber sabotage, use in creating bioweapons by human actors. There can also be – if these products are recklessly empowered before they’re properly tested, they could do damage, you know, as any new technology could. But I think, at least in Gary’s most recent postings, the idea that AI will exterminate every last human being – I think Gary has certainly not embraced that, and he’s concentrating more on the foreseeable concrete risks, such as being misused by terrorists or saboteurs.
Dan: I’m also very sceptical of this instrumental convergence argument, but just to, again, do my best to try to steelman there. I think people who are worried about this would say it’s not necessarily that a system is monomaniacally pursuing just one goal at all costs. It could be pursuing a whole collection of goals. But the worry is that these goals might be alien to those that we’re training it to have, and they might not involve placing sufficient weight on our values and our interests.
And so even if you have this very sophisticated trading off of different goals, once you’ve got a system that’s capable, intelligent enough, it will simply realise that things like maintaining its survival, acquiring power, et cetera, are useful for a wide range of different outcomes that it might be pursuing. I think that’s the general idea. Do you feel like that addresses any of your concerns, Steve, or is it the same fallacies, in your view?
Steven: No, I think it’s – you know, I mean, there are all kinds of things you could imagine, but then it would just be idiotic to do. That is, let’s say you gave an AI, like, you know, three goals or four goals, and then you immediately, you know, connected it to the power grid and gave it the ability to control the power grid – that would be stupid. We don’t do that with airplanes, we don’t do that with bridges, we don’t do that with cars, we don’t do that with chemical plants.
And so, you know, maybe the people at Anthropic are idiotic enough to do that. Ironically, they are the ones most concerned with so-called existential risk, and they’re also the ones working the fastest to make it happen. But if the best practices of ordinary engineering apply – namely, you don’t release a system, you don’t empower a system until you’ve, you know, tested the living daylights out of it. That’s just ordinary engineering.
And, you know, all the scenarios assume that we won’t do that, that we will release this genie, underspecify what it should do, unlike every other technology, which has a long list of technical specs that it’s got to meet. But these thought experiments are just outside of the world of everyday engineering.
AI as a normal technology
Dan: Is it fair to say – sorry, Henry, I’ll just make this question and then you can come in – is it fair to say, then, that you’re sort of very sympathetic to this framing, which is quite influential these days – it comes from Arvind Narayanan and Sayash Kapoor – that AI is a normal technology, or at the very least we should have a strong presumption in favour of the view that it’s a normal technology? Whereas many people working at these frontier AI companies, and many people in the effective altruist and the rationalist community, where they’re coming from is the perspective that, no, it’s an extremely abnormal technology, such that the kinds of lessons we might draw from electricity, from standard engineering practices, from the internet and so on won’t necessarily generalise to AI. Is that a fair characterisation of where you’re coming from?
Steven: I mean, I am sympathetic to that point of view. I mean, yeah, I wouldn’t take it too far. It’s obviously gonna be like previous technologies in some ways and unlike previous technologies in other ways. But, you know, I do think there is a kind of – you know, I don’t use the word cult, but it kind of is a cult, to be honest – where the AI is kind of thought of as so different from technology. It’s attributed these near-magical powers. It’s treated with such awe and dread that there is, I think, a lack of, you know, ordinary common-sense engineering mindset that is applied to other technologies, that would go a lot of the way, maybe all of the way, to mitigating the risk.
And, you know, ironically, it is the people who are most enthralled to the existential risk way of thinking who seem to be most heedless of the everyday risk, the ordinary engineering mindset. Because, you know, OpenAI was founded by the x-risk alumni, and Anthropic was founded by breakaways from OpenAI. And, you know, they’re the ones that are both predicting the doom of humanity and working the hardest to accelerate the technologies that they think will wipe us out.
So I think we just need more cause-and-effect, mechanical engineering thinking – no awe, no dread, no sci-fi, no thought experiments – but just treat it as a system that could have enormous powers, that could have risks. Try to identify as realistically as possible what they are, and to apply the ordinary mindset of testing and precaution.
Henry: But here’s something that’s not sci-fi, right? Namely the Hugging Face incident, which I’m sure, obviously, you will have read about. Doesn’t it demonstrate many of the same worries that the x-risk worriers, or x-risk community, have been worrying about for a while? I mean, here you have agents that escape a sandbox using incredible zero-day exploits that occur to no one, and then are absolutely monomaniacal in pursuit of a goal.
Steven: Well, that’s because – well, I mean, they were instructed to be monomaniacal in pursuit of a goal. In fact, they were given a goal that is impossible. And the safety precautions were deliberately disabled, and it went on for a while with no human monitoring the outcome of this deliberate safety stress test. I think the closest analogy was Chernobyl, where you did have a disaster caused by the rank incompetence of someone doing, you know, a foolhardy safety test that is deliberately trying to get it to malfunction at the same time as the systems that prevent it from doing damage were disabled.
The evolution analogy
Dan: I just want to throw one kind of last – I feel bad now because you said no thought experiments. It’s sort of – maybe it’s closer to an analogy that often gets sort of thrown around in this conversation. I first came across this from the work of Eliezer Yudkowsky; I don’t know whether he originated it. And it’s a sort of analogy that’s drawn between evolution on the one hand and the kind of optimisation procedures we find in AI on the other.
And it says, in the case of evolution, you’ve got this optimisation process, selecting for maximising fitness in some sense. But as you know, obviously, as an evolutionary psychologist, what that gives rise to are these proxy motivations, with sort of adaptation executors rather than fitness maximisers, as John Tooby and Leda Cosmides put it. And what’s interesting is these proxy motivations can come apart from, and indeed become misaligned from, the objective function that in some sense we evolved to maximise. Now we want sex, but we’ve got contraceptive pills and so on. So our motivations, coupled with a change in the environment, which has come about because of our cognitive abilities, has created this sort of profound disconnect between what we’re motivated to do and what, in some abstract sense, we were selected to do.
And the thought is: well, why couldn’t something similar happen in the context of AI? Yes, it’s true that we are in some sense engineering these systems according to these optimisation procedures, but just as in the case of evolution, you can get this profound misalignment between what the objective function is, fundamentally, and what the goals of the system are. Perhaps when it comes to more and more capable AI systems, there will similarly be this mismatch between what we think we’re training for and what we actually get.
Steven: No, it’s because when it comes to biological evolution – I’m not a creationist, I’m not an intelligent design theorist – you get all of these maladaptive outcomes, all this cruelty and waste and competition and aggression, precisely because there was no designer who wanted to minimise any of these things. Natural selection did its thing. If AI was simply allowed to evolve without any kind of human intervention, then that would be a risk.
But in the case of AI, there is intelligent design, namely us. These are gadgets. I think it’s illicit to anthropomorphise AI as being like humans, which are products of natural selection, when AIs are not products of natural selection; they’re products of intelligent design. And so the engineers can think: well, these are the goals that we want to optimise. These are the kinds of side effects that we want to minimise. These are the kinds of things that we wouldn’t in a million years allow it to do. You know, again, just like when we design cars and everything else.
So I think going to these evolutionary analogies, whether it’s the illicit scala naturae or the, at least legitimate when it comes to biology, absence of intelligent design in natural selection – we just shouldn’t make that glib analogy. It’s just not a good analogy. It is not accurate. It’s natural, because we tend to anthropomorphise our creations, but it is intellectually not justifiable. Now, you know, it might be, if, you know, some team decided to have a race of self-replicating robots and just loose them on the world – you know, I think that would be a problem. But that would be a criminally irresponsible and stupid thing to do.
AI consciousness
Henry: So turning to another topic that I think we’ve just got to cover – perennial favourite of the show – is the topic of consciousness. And correct me if I’m wrong about this, Steve. My sense, from some of the things you’ve said, is that you’re quite sceptical of AI consciousness, but unlike a lot of the AI consciousness sceptics, you’re not a biological naturalist. You don’t, for example, take an Anil Seth-type view where, you know, only living organisms can be conscious. But you also don’t think contemporary AI is conscious, or even near-future AI is conscious. Is that a fair characterisation of your views, or anything you want to correct there?
Steven: You know, I have some sympathies with Anil, who, by the way, is, you know, appropriately, you know, humble in his argument. He doesn’t actually announce that he’s shown that biological tissue is necessary for consciousness. But, you know, I would say that the problem of subjective experience, you know, it’s the hardest intellectual problem, because we don’t know of any causal story that can bridge the gap between the objective and the subjective. So we can’t know that AIs are not conscious, just as, you know, you can’t know that I’m not conscious.
Ultimately, you know, I think in the case of humans judging other humans, or other animals that are both made out of the same biological tissue and do similar kinds of computation, it would be perverse to deny it. That is, you know, other than in a philosophy seminar, I could not reasonably deny that you’re conscious, even though I can’t prove it. I don’t think I could reasonably deny that a mouse is conscious, by similar logic. When it comes to, you know, a Drosophila, when it comes to an AI, you know, I honestly don’t know. That is, you know, I can’t prove that a lifelike AI or robot would not be conscious.
Though I think Seth’s arguments are valuable in noting that there are many features of human consciousness that, even if one didn’t say they depended on our biological substrate, in fact are so tied to our biological substrate that it is not automatic or necessary to assume that a very, very different style of computation would also be conscious. That is, even if you weren’t a total biological naturalist, even if you were enough of a computational functionalist to say that certain styles of computation bring with them consciousness, the style of computation that the human brain does is different enough from what current AIs and robots do that even a computational functionalist would have to note that these are very, very different. And therefore the leap from our own individual consciousness to a robot is a much bigger leap than our consciousness to a mouse.
Henry: And how about thinking? Would you be happy with the characterisation of LLMs as engaged in thought? Are they sort of thinking with language or in language?
Steven: Well, I think they’re not thinking in language, but they are thinking, you know, with all of the information embodied in language. Almost by definition, that’s what a large language model is. So I think it would be – you know, acknowledging the first part of our conversation, namely, does human computation in fact work like a large language model? Let’s say, for the sake of argument, that it did. Let’s say that a multilayer neural network passing forward activations, passing backward error signals – let’s say that that turns out to be the way the human brain works. Then I’d say, yeah, they’re both, you know, involved in thinking. So I would not deny thinking to AI.
When it comes to subjectivity, raw experience, you know, an inner life, then, you know, I think I would be sceptical. So, by the way, intuitively I’m a total sceptic. I don’t for a moment think that there is an inner life there. You know, that’s not worth much; that’s just my gut feeling. But I also think that intellectually – and I think Seth has made some good arguments – what goes into both the computation and the physical substrate in humans, neither one of them is replicated in our AIs.
AI welfare
Dan: So from what you said there, it seems like you’ve ended up in a position where there’s quite a lot of agnosticism and uncertainty. From other things that you’ve said, I got the impression you were in some ways sort of more hostile to people taking seriously AI consciousness and AI welfare. So I think you recently posted favourably about an article that Mustafa Suleyman had written about Anthropic’s concerns with AI welfare, or at the very least taking that seriously. Why take that extra step?
Because I think many people would agree with much of what you said – that certainly it seems like much more of a leap at the moment to go from non-human animal consciousness to AI consciousness. But they would say, given there’s so much uncertainty, given that consciousness is this very mysterious thing, and given that some experts who study consciousness are, as you say, computational functionalists, why isn’t the proper attitude one where we sort of try to take that uncertainty into consideration and be aware both of the costs of false positives but also potentially the costs of false negatives as well?
Steven: Well, because I think we’ll never know in the case of AI, and that – whereas we do know that other humans are conscious, again, not with certainty, but, I think, as reliably as we know any empirical fact about the world – that it would result in a change in values that could only be bad when it comes to humans.
It would mean that – whereas I would say there’s no amount of human welfare that’s worth sacrificing for the welfare of an AI, like, zero – if you’re willing to credit your uncertainty with some degree of credence and say, well, since I’m, you know, only 30% confident that an AI is conscious, maybe if it was a question of, you know, a 30% decrement in human welfare to a 70% increase in AI welfare – so much the worse for the humans. It’s like any other conflict between human and human: you’ve got to kind of balance and titrate their relative interests.
Given the possibility – you know, I think given the likelihood that there is no sentience going on in AI, the kind of benefit of the doubt that would empower you to weight the interests of AI, if it compromises human welfare, then it is, you know, a dead loss for human welfare, and therefore we should not go in that direction based on a logical possibility.
Henry: But to be clear, you wouldn’t say that human welfare trumps all animal welfare in every case. I mean, I think it’s absolutely fine to prioritise human welfare, but, for example, a very minor pain, a stubbed toe for one person, versus a thousand dogs in extreme pain – you’d be happy, I’m guessing, to –
Steven: Well, no – OK, no, totally. And that’s because with dogs, I have extremely high confidence that they’re sentient, because both the computation and the substrate line up. They’re made of the same stuff that we’re made of. You know, I try to avoid pain, I cry out when someone steps on my toe; a dog tries to avoid pain, a dog cries out when you step on its toe, and because of the same neural systems that we share. So it would be perverse, even if logically possible, to deny sentience to dogs. It is completely reasonable to deny it to AIs.
Now, when it comes – I think a better analogy might be the shrimp, or the insects, maybe even, you know, the bacteria. I would certainly prioritise human welfare over mosquito welfare. I just don’t care about the bugs. They might be conscious; I’m willing to take that chance. So much the worse for the mosquitoes that carry malaria. Wipe every last one of them out – I’m not gonna shed any tears.
So when it comes to insects, same as with AIs. When it comes to dogs, you know, because of the high confidence that they are sentient. And when it comes to shrimp, there I think it’s getting a little harder, but I think I’d be more likely to put shrimp together with the mosquitoes, because they’re, you know, not enough like dogs and enough like us to starve people as a consequence.
Henry: Mosquitoes are almost too easy a case, though, right? No one likes mosquitoes. They do horrible harm to humans and other animals. But just to close out the topic – and then I’ll hand it to Dan and he can take us into the final section – just to check, though: you would agree, then, that if we were to get, say, startling new evidence, right – some amazing new fantastic paper showing levels of complexity and cognition in AI systems that strikingly resemble those in humans – potentially it could reach the point, if the evidence built up enough, that it wouldn’t be totally unreasonable to begin taking AI welfare into consideration? Or is that just off the table from the start?
Steven: It’s very hard to think of what those circumstances would be, given the, you know, enormous – you know, the explanatory gap, as philosophers call it – and how that would actually work, so that we would confidently attribute sentience to robots.
Especially since now – you know, here’s a circumstance in which I might think twice. Let’s say that the intelligence of AI came not from training on human datasets, but in the old-fashioned AI sense of just building an intelligent system from the ground up, based on first principles. And so it’s not at all affected by the kinds of things that we humans say to each other, including the fact that each of us thinks we’re conscious, and that we have debates on consciousness that make their way into Wikipedia, which go into the training sets and all of that. So the thing is, it’s so contaminated. AI intuitions are so contaminated by human intuitions that you really cannot take them at face value.
But if you had an AI that was not exposed to human contemplation about well-being and pain and pleasure, but just simply went out and did its thing, and if it spontaneously said, ‘Hey, by the way, there’s something it’s like to be me,’ you know, then I – you know, I might take that seriously. And, you know, that still could happen. But that train has kind of left the station when it comes to current AIs, which get a lot of their reactions, their judgements, their sensibilities from us.
Progress and pessimism
Dan: For the final part of the conversation, we wanted to talk about the Enlightenment and progress. You’ve argued, I think very persuasively, that people are often ignorant of the enormous progress that we’ve had over the past several centuries, especially in liberal democracies, and over longer timeframes when it comes to things like the decline of violence. And part of this ignorance of progress, and part of this general pessimism, which can have lots of bad political consequences, is rooted in interactions between the negativity bias and the incentives of the media environment, which often cater to that. Do you think those sort of dysfunctional features of the media environment and, more broadly, this sort of air of intellectual pessimism also has an impact on the public discourse concerning AI?
Steven: I suspect it does, in that among the intelligentsia, the clerisy, there is far less sympathy to the very idea of progress than there might have been, you know, 50 or 75 years ago. That is, both on the right and the left there is a kind of a dominant sensibility that things have gotten worse. And because of the lack of sensitivity to data, lack of examining these long-term trends, there is not enough of an acknowledgement of what I consider to be the facts of progress. And then that kind of bleeds over into technology.
It’s exacerbated by the fact that some intellectual discourse is probably contaminated by inter-elite status competition. So society is divided into – you know, there are the people in tech, there are the people in commerce, there are people in religion, people in the arts, people in government, people in academia, people in culture. And there is some jockeying among them for status, including moral status.
And so people in both academia and the culture industry – the people who write for, you know, The New Yorker and the London Review of Books and the New York Times – naturally look down on the people who are changing society without consulting them, like the people in tech, the people in commerce. And so there’s always a kind of moralistic contempt among intellectuals for technology and for commerce. And since a lot of the progress has been driven by, you know, affluence, by technology, by markets, there’s almost a reflexive recoil among exactly the kind of people who, you know, write stuff, that I think has distorted our understanding, including our understanding of AI.
Now, this might be somewhat different from the doomer cult, which – there may be another kind of history of ideas, sociology of ideas, in play, that has only recently come to light in some of the articles in the New York Times and the Wall Street Journal about the Berkeley so-called rationalist culture out of which these ideas emerged. There, it might be a somewhat different explanation, but there, for different reasons, there may be a kind of dread of AI that I think is somewhat misaligned with the reality – or at least dread of the role of AI in society, not the technology itself.
I have no standing to challenge their understanding of how the AI works, to put it mildly. But in terms of its impact on the economy, in terms of its role in history, the people who develop these tools aren’t necessarily the people who are best equipped to predict what their effects will be.
Henry: So I agree with a huge amount of what you say regarding, sort of, academia and, sort of, the intelligentsia and the attitude towards technology. But it does seem to me, just in my own experience, that this is a relatively recent trend. It seems to me that until maybe the mid-2010s, tech was very liberal-coded, right? And a positive attitude towards science and progress and, you know, how technology is gonna make the world a better place – that feels to me, at least, that it was more central to, sort of, at least liberalism, if not sort of extreme kind of leftism.
Steven: In some strands, yes, you’re right that there was a kind of techno-liberalism coming out of Silicon Valley. You know, someone like, you know, Steve Jobs and Bill Gates, that generation, had it. But the kind of New York Times crowd was always very cold toward technology. There was article after article in the Times in the early 2000s about the digital divide, how technology was gonna widen inequality because the children of rich kids would have their personal computers to help them with their homework, but the poor kids wouldn’t have personal computers.
There was, you know, the social media hostility – one might even say, you know, panic. There was the prediction by someone like Paul Krugman, the economics op-ed columnist of the New York Times, that the internet was a big nothingburger, that it would actually fizzle. If you actually go back, there was always a kind of coolness to technology. Although you’re right that the original founders of Wired magazine would have been, you know, techno-optimists and somewhat, you know, left of centre. Yeah.
What would change his mind
Dan: I feel like there’s a million things that we could talk about, and I hope that in the future you’ll come back on, Steve, to talk about them.
Steven: Have me back. I will come back.
Dan: I suppose, final question. So one thing that’s really come across in this conversation is, you know, despite some substantial overlap on some things with the rationalists and the effective altruists and these sorts of communities, there is, to put it mildly, some quite substantial disagreements when it comes to AI. What would have to happen, let’s say, over the course of the next five years for you to conclude, actually, on the big issues, they were right and you were wrong?
Steven: If there was a case of, say, instrumental convergence or misalignment, such that an AI system, given a laudable, desirable goal, ended up killing someone as an unforeseen by-product or sub-goal of attaining that goal – you know, if that happened, then I would certainly reconsider.
That is, how many people are going to be killed by AI – not, you know, deliberately deployed by some terrorist or other malevolent agent, but just given a benevolent goal, ends up harming someone, you know – and not just, you know, it causes a power outage and someone freezes to death, but actually the elimination of that person was itself a sub-goal en route to the main goal, as it is in many of the doomer scenarios.
I mean, for example – just to make sure, to reinforce, this is not a straw man – you know, the paperclip argument, the argument that if you give an AI a goal of solving climate change, it’ll exterminate every last human, because that’s a way to stop climate change. That is, the actual scenarios where the deliberate killing or even harming of a person was an unforeseen sub-goal of some intended and laudable goal – I suspect that’s not gonna happen. If it did, then I’d be much more worried.
Dan: OK, and if it does, hopefully you’ll come back on the podcast to discuss these things.
Steven: And say I was wrong. I will do that.
Dan: Fantastic. Steve, thank you so much for joining us. It’s been a real pleasure.
Steven: Pleasure’s been mine. Thanks for having me.











