0:00
/
Generate transcript
A transcript unlocks clips, previews, and editing.

Navigating the Intelligence Explosion (with Fin Moorhouse)

Explosive technological progress, the "grand challenges" on the path to superintelligence, space expansion, and the ethics and politics of artificial minds

What would happen if the amount of intelligence and ingenuity devoted to science and technology exploded over a short period of time, compressing a century of technological progress into one decade?

In this episode, Henry and I speak with Fin Moorhouse, one of the most interesting thinkers on how we should understand - and prepare for - the transition to truly transformative artificial intelligence.

Last year, Fin and Will MacAskill published “Preparing for the Intelligence Explosion”, which explores what might happen once AI systems start to dramatically expand the amount of intelligence in the world. In my view, this is one of the most important essays to master if you want to understand just how crazy the world might become in the near future due to rapid AI progress, and the huge challenges and opportunities this will create.

Fin currently works as a researcher at Google DeepMind, but he joins us in a personal capacity. The views he expresses are his own.

Among other things, we discuss:

  • What an “intelligence explosion” actually means and how likely it is.

  • Whether AI must substitute for human researchers, rather than merely augmenting them, to trigger explosive growth.

  • Why current AI systems may lack research taste - and whether scientific progress depends on ambition, status-seeking, and irrational confidence in one’s own ideas that purely truth-seeking AIs would lack.

  • The role of experiments, learning by doing, institutions, embodiment, incentives, and large-scale social coordination play in technological progress.

  • Robotics, self-replicating capital, and the alarming possibility of an industrial explosion.

  • Techno-optimism and techno-solutionism.

  • Why solving alignment isn’t sufficient to ensure a flourishing and enlightened future with superintelligent AI.

  • Why superintelligence might make questions of space governance extremely important.

  • Fin’s deflationary “illusionist” view of consciousness and Henry’s objections to it.

  • Whether we need to settle metaphysical controversies about consciousness before we address ethical and political questions about AI welfare, rights, and moral status.

I really enjoyed this conversation, which, in addition to introducing Fin’s perspective on a range of questions, includes a wide range of disagreements and debates as well!

Fin is one of the smartest and most thoughtful people working on questions that are shockingly under-researched given their monumental importance.

Links and further reading

Transcript

  • This transcript is lightly AI-edited and may contain minor mistakes.

Dan: Welcome back. Today we’re joined by Fin Moorhouse. Fin is one of the most insightful thinkers working on how we should understand — and also how we should navigate — the transition to a world with truly transformative artificial intelligence. Fin, last year you published an excellent article with Will MacAskill on preparing for the intelligence explosion. What’s the central argument of that essay?

Fin: Yeah, thanks for having me, both — fan of the pod. The argument is: first of all, something like an intelligence explosion — and I can say what that means — seems really quite plausible, potentially really quite soon. Secondly, that could just have very transformative effects, and effects in lots of different domains that matter — so not just isolated to one very kind of neat, easily analysable kind of transformation. And then lastly, that we should be trying to do stuff to prepare for this kind of wave of transformations that we might expect from AI. And we should be trying to do that now.

And as a corollary, there isn’t one magic bullet here, right? In the same way that when you think about the Industrial Revolution, it’s very hard to think of any particular thing that any particular person could have done to make things go better — but it still might have paid to do lots of things and think about lots of different upshots earlier rather than later.

Dan: So you mentioned there at the beginning this term intelligence explosion. Do you want to say a little bit more about how you’re understanding that term? Because it carries quite a lot of baggage, I think, in terms of how people understand it. And I think the way in which you unpack it in the essay is actually very interesting, and in some ways requires less baggage, or sort of fewer assumptions, than some uses of that term.

Fin: Yeah, so we didn’t mean anything precise by that term — somewhat lazily, but also somewhat deliberately. Roughly speaking, I think what we had in mind was just a radical, dramatic, fast influx of machine intelligence in the world, especially when compared to the kind of sum total of human intelligence and ingenuity. So the thought is that at some point — again, in a hand-wavy sense, I’m not trying to pin this down precisely right now — but at some point, the kind of sum total of just intelligence represented by AI and non-human thinking could just surpass all of us combined. Now, there are different ways that could happen, which we can talk about — different ways to really try to operationalise that — but that’s the rough idea.

Henry: So the intelligence explosion is specifically an explosion in non-human intelligence, right? Because I think you could say: look, the amount of intelligence in our civilisation, in the world, has been rapidly increasing as we augment our intelligence. Even, you know, the internet is an example of a massive intelligence augmentation. So, you know, I could imagine someone saying: look, the intelligence explosion has been going on for thousands of years. Like, how is this different?

Fin: Yep — that’s a totally fair point. And actually maybe it’s kind of not necessary to — I used the word non-human, and maybe I shouldn’t have — because maybe in some ways you can think about AIs as augmenting, or as tools for, people. I think the distinctive thing then is that the kind of influx will one way or another come from machines, rather than the number of people, for example, or people learning to be smarter on their own. And you’re right that in some sense, you know, our modern history is a kind of drawn-out intelligence explosion. What we’re trying to emphasise is the speed, which I think — again, very plausibly, but not necessarily — could be unprecedentedly fast: like, many, many times faster than it’s ever been before.

Dan: So the thought is, at the moment, in terms of increasing the amount of intelligence in the world — and your focus in the paper, at least initially, is the amount of intelligence which is going into research, R&D, the kind of research that underpins technological progress — at the moment, we’re being heavily bottlenecked by just the size of the human population, which is not expanding at a really, really rapid rate. So I think you give an estimate in the paper: something like the rate of growth of research effort rooted in human intelligence is, like, four or five per cent per year, or something like that.

And the thought is: well, we’ve now got this other source of intelligence, which comes from AI. And at the moment, AI in terms of its capabilities doesn’t really match everything that the human mind can do when it comes to research. But at a certain point, you’re going to get AI systems that at the very least will have kind of parity with human capabilities. So now you’ve got this other source of intelligence, and because we can and because we will scale up the amount of that kind of machine intelligence very, very quickly, you get this, like, rapid expansion in the amount of machine intelligence, in a very kind of short period of time relative to what we’re used to on the kind of human timescale — you get this just explosion of the kind of research that then powers lots of technological progress. Is that a fair summary of what you argue in the paper?

Fin: That is such a good summary that I’m gonna struggle to improve on it, but I’ll try almost paraphrasing it at least. I think part of the thought here is — I mean, it sounds a little silly to say out loud, but why should we actually care about AI progress? One reason is because there are, you know, direct effects from AI — people directly interacting with LLMs — that has, you know, social epistemic upshots and so on. But those are not the only effects. Actually, I think that at least as important are the downstream effects, and in particular the kind of technological change, the new kinds of ideas, that AIs could generate, right? In the sense that you don’t actually have to interact directly with the AIs that are generating new technologies, in the same way that you don’t have to meet the researchers who develop technologies that influence your life profoundly. So that seems like a place to start thinking about what AIs could be doing.

Henry: Can I just, before we get too into details — there is a worry I have here that you’re sort of using intelligence as a mass noun, as though it’s something that we can cleanly measure. And, you know, look, I realise in the human case we do have actually pretty clean operationalised measures of things like fluid g. You know, people argue about IQ, but it’s fairly predictive. But I think as soon as we start — particularly as soon as we start moving outside the human case, right? Like, if you give me a crow and you give me a dog and say which one has more intelligence inside their head? That seems like an ill-posed question.

Fin: Yep, totally agree. So I hopefully don’t want to imply that I think intelligence is easy to measure. I mean, you can think about the kind of intensive sense — how smart are the smartest AI systems? — I mean, notoriously, famously kind of ambiguous and rough to operationalise. Also the kind of extensive sense — like you said, you know, intelligence as a kind of lump noun. Yep — I absolutely only intend to mean this in a very hand-wavy and extremely vague sense. And I think it’s important and useful to try to kind of pin down different particular ways to operationalise. But I think so far I’m trying not to go in for any kind of particular conception.

But — so, you know, if that kind of sounds somewhat reasonable — one way to get a handle on this question, like Dan was alluding to, is just to ask what drives technological progress historically. And you can write, as many have done, you know, bookshelves about the details. But a very high-level account that I find very compelling, from a kind of economics perspective, is: look, if you need a one-factor explanation of technological progress, it is in terms of the amount of research effort at a time, right? Which you can proxy by just the number of researchers. Charles Jones, the economist, is kind of best known for pushing on this view in the kind of mid-90s. Moreover, it becomes harder to make technological progress as you make more of it, because you pick the low-hanging fruit first. And so you kind of need to parameterise, like, the rate at which ideas get harder to find. But also, historically, we’ve just piled on more and more researchers. And so roughly we’ve sustained about a kind of constant rate of tech progress, at least in terms of measured productivity.

The reason that’s kind of convenient and nice is not just because it’s borne out fairly well empirically, but because you can just make this move, like Dan said: what if we could just massively increase the kind of human-equivalent researcher workforce with machines? You know, maybe they’re working as tools, right — they’re saving researchers time, and therefore kind of giving them more time on the important stuff. Maybe they are in some ways replacing them, substituting for them. Again, I actually don’t think at this kind of high-level place it really matters to get into those details. But if in some way they could speed up research radically quickly, that could then speed up tech progress.

Why does that matter? I think so many of the really hairy social, political questions that we’ve faced as a civilisation over the last century or two have been originated by just finding ourselves with new technological possibilities and asking ourselves: how do we deal with this? How do we kind of structure society around it? So maybe the most salient example is nuclear weapons, right? Fifteen years ago, no one had even thought of the nuclear chain reaction; now we have a working nuclear bomb. How the hell do we deal with this? That’s a really rough question. If that kind of process happens 10x as fast, maybe even faster, that starts to feel like a really hairy thing to deal with.

Dan: So there’s this model here, which is: increasing the amount of sort of machine-based cognitive labour, in the way that we seem to be on track to increase it in the coming years and decades — there’s just a very sort of strong case for thinking that that is going to dramatically speed up the rate of technological progress. I think that in and of itself raises lots of interesting questions, like: is that really the correct model of what drives technological progress? But I think it is worth just, sort of, firstly spending a little bit more time on, like, this question of what do we mean when we say improving the capabilities of AI, and kind of scaling up the amount of AI-generated cognitive labour that we have — that that’s going to lead to, like, a massive expansion in the amount of research that we can do. And I think one thing you’re saying is you get that result even if you don’t make a whole set of, like, really controversial assumptions.

So, for example, you needn’t assume that intelligence is, like, a simple matter, where, like, one thing is just as intelligent as another, or more intelligent than another. You might even get it if AI systems are just kind of augmenting human capabilities, not fully substituting for them. I guess my question there is: it’s sort of difficult for me to see how you’re going to get the really rapid increase in the rate of — let’s say research progress, or in the expansion of the amount of R&D that we can do — if AI systems are merely augmenting and extending human capabilities. Like, I totally understand the argument that says once AI systems can just flexibly substitute for human beings — so that you can get AI systems in some sense creating their own kind of autonomous domain of research — then just the numbers, in terms of how many AI systems we can get and how we might improve their capabilities and so on, that could result in this huge increase in the amount of research that we’re doing. That I buy. But the idea that you’re gonna get that merely by their augmenting human capabilities — that seems a bit more controversial to me.

Fin: Yeah, I think I do somewhat agree with that. So the example that comes to mind is maths, since it’s been on the, you know, on the Twitter feeds recently. You know, at some point around, I guess, the mid-twentieth century, it became effectively free to calculate. And as far as I understand, a reasonable fraction of a mathematician’s job before that point was to do numerical calculations by hand. You might naively think: now it turns out that I can calculate something like 10,000 times faster than I ever could before — that’s going to accelerate mathematics in this kind of, you know, radical singularitarian way. And it didn’t, right? Basically not at all. Similarly, you think about, you know, just the spreadsheet as an innovation across knowledge work: it was useful, but it didn’t meaningfully accelerate basically any kind of field of knowledge work. And also symbolic algebra as well, which kind of came later in maths — didn’t have any kind of radical effects.

So if the affordances that we get from AI are always complementary with other human capabilities, and we don’t get more of them, I basically agree: we could get a big one-time speed-up, or a one-time kind of, you know, boost, when we kind of pick all the fruit that were hard to pick before, and then we’re just bottlenecked again. And this is just in some way the story of any, you know, new technology — this is what normal technologies do. One way that AI could be different is if it could be always complementary in some sense — so always augmenting — but the set of tasks that it’s able to augment is always expanding, right? So, like, in some sense human researchers always have some kind of absolute advantage at something, but the set of those things is just kind of halving constantly. That would be one way that things could go where I think you do get a sustained speed-up. But I also just largely agree that for the more kind of dramatic speed-ups, you need — at least in lots of domains — it needs to look more like the AIs are basically capable of flexibly substituting for human researchers.

Henry: Just quickly on this sort of broader topic, since you mentioned — or referred to — the amazing, very cool maths discoveries we’ve seen really in the last six months, last year. There’s this question Dwarkesh Patel asked Dario Amodei a few years ago — I’m sure you know what I’m gonna say. You know, he said — I’ve got it here — what do you make of the fact that LLMs (this was 2023) have basically the entire corpus of human knowledge memorised, and they haven’t been able to make a single new connection that has led to a discovery? Whereas if even a moderately intelligent person had this much stuff memorised, they would notice: this thing causes this symptom, this other thing also causes this symptom — there’s a medical cure right there. So, like, I’m curious what your take is on why we haven’t already seen some amazing scientific outputs from LLMs — only in the domain of maths so far.

Fin: Yeah, it’s a great question, and that’s totally true. So I don’t know the answer at all. One thing that goes through my mind is — yeah, it seems like even the smartest AIs, they seem to really lack kind of research taste, right? So they’re very, very good at kind of knowing relevant results, at making kind of pedantic, you know, points. But when you think about what makes a very successful kind of superstar researcher in their field so successful, often they have this hard-to-pin-down sense for what research agendas are likely to pan out, right? Some of this, I think, comes from — you know, in the language of AI — being trained on non-public data, right? So, you know, imagine you’re a PI: you supervise a bunch of students, they’ll try a bunch of stuff, some of the ideas won’t pan out, some will. The ideas that don’t pan out don’t get put on the internet, don’t get written up as, you know, here’s our report of a failure. But you still get to learn from them.

I expect there’s much more going on there which has to do with, you know, creativity — and, you know, I think the fact that it’s actually quite hard to pin down what exactly is a missing ingredient suggests that we haven’t operationalised it well enough to make the AIs good at it. It may be that there’s no way around collecting very expensive-to-collect data from human researchers who have research taste, in order to teach the AIs that taste, and that could take some time. So the short answer is: I don’t know, but I expect it has something to do with some components of research skill which are basically independent of, you know, kind of crystallised knowledge.

Henry: Interesting. Not to put you on the spot, Dan, but I’m curious — you know, I think we’ve talked around this topic a few times. I’m just curious if you have a take on this — like, why we haven’t seen—

Dan: Yeah. So I think part of it is just to do with the capabilities of the systems. But there is this other aspect to this — and this is, like, a little bit of a digression, but I think it is interesting. I first came across this view from Hugo Mercier; I don’t think he necessarily kind of originated it. But it’s the idea that advancing the frontier of knowledge, in the human case, relies on individual-level irrationality. Because if you imagine you’ve got a community of people, and you’ve got an individual who just wants to figure out what’s true, almost always the most rational thing for you to do is just to defer to majority opinion or something — because it’s much more likely that other people would have figured out the truth collectively than that your own ideas are correct. But if everyone does that, then you just never advance the frontier of knowledge.

And so the way in which we get around this is: human beings just care a lot more about, like, status and glory and esteem than we do about getting things right. So from, like, an individual perspective, you often find people just placing so much weight on their own ideas such that they’re willing to just pursue them to an irrational extent. And one consequence of that is you just get loads of people backing themselves when their ideas are just completely false and ridiculous — which I think we do in fact see all over the place, including in science. But very occasionally you get people who, because they back themselves and they really care about getting the glory of being the one that makes the discovery, they actually kind of advance the frontier of knowledge. And I think if you do look at the history of science, you find a lot of the people that make really significant discoveries are kind of egomaniacs, in a way that that basic model suggests. But anyway — if you think that kind of view is correct, then you might think optimising AI systems to be really good at truth-seeking might in fact not unlock the source of knowledge creation in the human case.

Fin: Yeah — I mean, it’s super interesting. I can tell you kind of what goes through my head. So this seems really plausible for domains like social sciences, humanities and politics, where ideas succeed, you know, partly on their merits, but partly on how they’re communicated and how kind of fit they are to, you know, succeed in certain ways. You could imagine that just an AI system alone can have, in some sense, a good idea in politics or philosophy, but it’s bottlenecked by, you know, just its ability to persuade or, you know, kind of form coalitions. That’s something that — I don’t know — you’d, like, have thought about a lot, Dan, and I haven’t. My sense is, in hard sciences — you know, thinking about drug discovery or, like, solid-state physics or whatever — it’s kind of rarer that new ideas are kind of, you know... you’ve just gotta be really brave and really kind of push for them. They’re often, just to an outside audience, basically inscrutable: very technical, very complicated, often require kind of large teams of people to uncover. Once they’re uncovered, they’re often easy to demonstrate, because they work. And it seems like the AI systems are not generating ideas even in domains where you can easily kind of validate that they’re good ideas. So it seems like there’s at least some more to it than just that.

Dan: I guess the question is — let’s not get derailed by this too much — but in terms of validating ideas, it’s quite complicated. Because the minute you start coming up with a nascent idea or research approach, one rational way of validating it is by thinking: how does this relate to the existing consensus body of knowledge in the field? And I think actually AIs as they exist today are, like, really bad for this. Like, I write a lot, and I’ll have, like, an opinion, and I’ll give it to Claude or ChatGPT for feedback. And they’re like: well, actually, if you look at the existing literature, then what you’ll find is... and they’ll just, like, bog it down in all of the nuance of what people have already said. And it’s like: I know that people have already said that, but I’m trying to argue for something new.

And I think this iterates into even the hard sciences. And this is why you get these, like, ferocious priority disputes. So if you think about Newton and Leibniz and calculus, right — it’s, like, this incredibly ferocious argument that ends up entangling, like, entire nations, over who gets priority, who gets the glory for being the one that made this discovery or invention, depending on which way you look at calculus. And I would argue it’s precisely that motivation — to get the glory for coming up with something new — that drives people to really back themselves, back their own ideas, where if you were just a disinterested, rational, Bayesian, truth-seeking agent, you wouldn’t. You would just defer much more to what other people have already said. But anyway — this takes us a little bit too far afield. Back to the things that we were talking about.

Fin: Actually, maybe I’ll interject very quickly there. So there’s this kind of round of random people on Twitter, like, getting, you know, their AI to generate maths results. I had a spare hour, so I kind of thought, I might as well, you know — I’ll ask my AI: just go and do a, you know, do a maths breakthrough, right? Set up a server — you know, here’s some compute, run your experiments, whatever. And it actually worked. It came up with — you know, as a total non-mathematician and dilettante — it came up with a result that seemed correct and seemed, you know, like, somewhat non-trivial. And just before — you know, I was kind of looking forward to tweeting about it and getting my, like, you know, five seconds of fame — and I, you know, asked my AI to just double-check that I haven’t been scooped. And, you know, lo and behold, it goes and checks — and I totally have been scooped, actually quite recently, by presumably some other guy, you know, using Claude Code or whatever. And I kind of wonder if you get the, like, Newton–Leibniz dynamic a lot, if you have this period where, if only you know what to prompt for — if only you know where the, you know, easy results are — then it’s very cheap to do that. And it’s basically just a kind of race for who can find the, you know, find the prompts, basically, and unlock the kind of credit they can get for it.

Dan: Yeah. I mean, definitely the way in which the use of AI is going to interact with these, like, status games that intellectuals and academics and pundits and so on play — yeah, that’s going to be interesting. Okay — this model you’ve got, Fin, of, like, what drives technological progress. As you mentioned, a lot of this is, according to your way of viewing things — if I’ve understood it correctly — fuelled by just the sheer amount of cognitive labour dedicated to research. So that’s a controversial view, as I understand it. There are some people who think — as you mentioned, there’s a vast literature on this — but they’ll think: look, like, yeah, you need this cognitive labour to get technological progress, but you also need all of this other stuff. Like people in the real world learning by doing. You need, like, fortuitous mistakes. You need particular kinds of institutions and, like, the incentives that they create. Like, merely ramping up the amount of, like, sort of pure disembodied intelligence in and of itself won’t get you huge technological gains. And my sense is you disagree with that — you think it will. So: have I fairly summarised your view? And if I have, why do you disagree with that alternative view?

Fin: Yeah. So, I mean, I should zoom out a bit and try to kind of characterise what my view is. So, I think I’m not trying to claim that, you know, the main important upshot of AI progress — or, you know, the overwhelmingly most important upshot — will be something like across-the-board technological progress driven by AI. I think more what we’re trying to do is just say: look, among the other kinds of transformative effects AI could have, here’s one that seems like, you know, we can kind of reason about, and if it were to happen, it would be a fairly big deal, kind of socially and politically. Also, I don’t think I’m trying to claim that the kind of tech progress we’ll see will be kind of general and across the board, for reasons I’ll get on to. And then finally — although it’s kind of useful to talk in these terms just to kind of give clean arguments — in practice I actually basically agree that, both historically and probably going forwards to the future, most kind of tech progress, or at least a great deal of tech or productivity kind of growth, doesn’t come from, you know, explicit research and development efforts with, you know, people in lab coats in scientific institutions. A lot of it is much messier: it’s just people in the economy in general just tinkering and experimenting and kind of sharing their ideas without labelling it as R&D. So those are some — you know I love some caveats — so those are some caveats.

Still, the question remains, right? Like: doesn’t tech progress just get bottlenecked by these kind of other, you know, complements to just raw thought — like the need to do experiments in the real world? So, one thing to say is that, you know, I think in some ways scale can, you know, eventually have qualitative effects, once you have an insane amount of kind of cognition. For example, in some cases you could begin to run simulations, whereas before they would have been so expensive or so crude as to not be worth doing. I think you could also find kind of alternative approaches, you know, in some domain, which route around the need to run experiments — where, again, they’re kind of wildly less efficient given today’s kind of factor prices, right? But in the future could be very fast.

But then the last thing to say is that I think this is just probably going to be true in at least a bunch of, you know, important tech domains, right? So, for example, if I’m trying to do drug discovery for kind of longevity or something — and, you know, I’m interested in these kind of long-term effects, and I really need to just run these, like, longitudinal studies with the same cohort over, you know, like, a decade or so — I would imagine that it is just much harder to get any kind of confident results in a domain like that, compared to a domain like maths, for example. Software is the kind of huge and obvious domain where you should expect to see major speed-ups. Certain kinds of, yeah, like, kind of predictive coding in biology, where you’ll kind of want to understand the particular mechanism — like how a particular protein binds or whatever. Yeah — presumably you’ll just see much more progress in those domains which are relatively less bottlenecked. I think it’s an open question exactly how many domains there are like that, and how well you can kind of substitute in the other ones. But I think there’s at least a kind of presumptive argument that this isn’t, like, a kind of deal-breaker on the overall thought.

Henry: So we’ve been talking in quite abstract terms so far. I think it might just be helpful to get some examples of what you see as sort of scientific projects that might sound outrageously unfeasible, but actually you might think are in nearer reach than maybe listeners, or us, would assume. So, I mean, like — what are your kind of near, medium and long-term horizons for what kind of progress this might unlock?

Fin: Yep. So the things that seem like potentially the biggest deal to me — maybe the one I would mention first is robotics. Where — you know, what dimensions matter here? One is just getting robots really cheap; another is getting them just very flexible and dexterous; and then also getting the software, the control, which I think is the most important kind of bottleneck right now. That then unlocks, I think, this kind of very general-purpose potential feedback loop, where if you can automate a very wide range of manufacturing tasks, then you start to kind of get something like, you know, self-replicating capital, which can cash out at the end on basically arbitrary physical goods, right? So the kind of cartoon way of putting this is factories building factories. Because this would be so valuable for hard power, or just economically, I would imagine that it could just attract a huge amount of, yeah, engineering effort, basically, and investment.

Then in terms of the kind of outputs — so, things like... yeah, you know, not exactly jazzed about this, but I expect military applications. For example, drones: it just seems like there’s an enormous amount of headroom for, you know, really quite powerful and quite scary capabilities there. And again, if you just look at just, you know, dollar spend on R&D, a lot of this is directed towards kind of hard-power applications. Let’s see — other things that come to mind... Okay, so drug discovery, I think, probably will be huge — I mean, again, despite the need for some physical experimentation.

Henry: I think maybe one interesting question that comes up when we start thinking about really radical technological change that could come down the pipeline is how quickly you run into some quite messy philosophical questions about it. So, like, just for example: mastery of the human body, right? Not just sort of beating diseases, but allowing us to arbitrarily alter ourselves. You know — and again, I think back to things like the Culture series, where they can do glanding, they can artificially put themselves in all sorts of psychological or physiological states. You know, maybe we could set our pleasure centres to maximum, so, you know, we can all be in these hedonic hazes. And, like, quickly you run into questions — it’s like: is this a good future or a bad future? So, I mean, that’s part of the reason I asked the question: because I think there is this kind of, like, initial shelf — low-hanging sort of fruit — of, like, things everyone wants, right? Like: yeah, cure Alzheimer’s, cure cancer, beat climate change. But then the next shelf up quickly starts to get kind of spicy, and I think there are gonna be legitimate value disagreements about, you know, which items from that shelf we wanna reach for.

Fin: Yeah. I think also one thing that’s been on my mind a bit is: there’s a kind of standard menu of, you know, somewhat science-fictional, speculative technologies that we might imagine getting, you know, by the year 2100. And, you know, on this menu there are things, yeah, like flying cars, space travel, you know, Dyson swarms, uploading, and, you know, kind of nanotech — maybe kind of medical uses of nanotech, like kind of healing kind of our brains or whatever. I think, you know, it is often possible to kind of try to dig into what would it actually take to realise these technologies. And in some cases I think it’s possible to come to a view.

So, for example — you know, citation needed — the brain is incredibly complicated. And I don’t think that we have anything near, like, even small-scale proofs of concept for, you know, nanoscale, you know, medical applications that could kind of help reverse neurodegenerative diseases, for example. That’s something people talk about. That’s just a case where — I mean, I don’t know what I’m talking about, right? I’m a total non-expert in this domain — I just don’t see any kind of shot on goal, effectively. Whereas other stories, like spreading to space and just building a ton of compute or whatever, you know, in space — in some ways I think that kind of feels similarly science-fictional. But that’s an example where — again, as a non-expert, right, I’m not an engineer, I don’t have a space background — but talking to people who do, and also just kind of using some amount of just kind of judgement... it just seems like, yeah, we just totally have all the requisite knowledge, and it’s basically an engineering challenge. Like, maybe a very tough one, but we roughly know what it would take to do it; we basically know all the physics involved.

And so it’s possible, I think, to me, that we get a world where, you know, we don’t just get the full menu at once — we get some very wild kinds of tech basically long before it’s possible to get others. Especially around — yeah, I think, you know, kind of engineering around the brain and medicine seems especially tricky to me. Life extension as well — I’m like: seems extremely hard.

Dan: I just want to have a go at summarising where we are in the conversation, and the debates that are relevant, before we move on to another really interesting part of your paper, Fin, which is about grand challenges. And just to make it concrete: so, Dario Amodei has this term that he uses — a country of geniuses in a data centre. And as I understand the argument that you’re making with Will, it’s something like: look, at some point we are gonna get something like a country of geniuses in a data centre. They might start off as just a country of smart researchers in a data centre. But because of the way in which we’re scaling compute, and increasing the efficiency of AI systems for any given level of pre-training and inference compute, it won’t stop at just a country of geniuses. There’ll be this, like, massive population explosion — at least if you’re thinking about exponential growth over the course of a decade — where you’ve got, like, many planets’ worth of geniuses in a data centre.

And then there’s a question about: if you had that, how much technological progress would that generate? And I think you and Will make a good case in the paper that it would generate a lot, even though there are gonna be bottlenecks when it comes to things like real-world experimentation and learning from doing, et cetera, et cetera. And then there are people who take the other side of that debate, where they say: look, even if you massively increased the amount of, like, disembodied cognitive labour of that kind, it’s going to be so severely bottlenecked by these other things that actually you shouldn’t expect such rapid technological progress. So that’s the debate that’s happening at one level.

But then there’s this other thing that you alluded to, which is: you might think, well, okay, once you start getting these populations of geniuses in data centres, they are going to be able to accelerate one specific kind of technological progress — which is technological progress that pertains to AI, both in terms of the software and the robotics. And then once you’ve got that, and you start getting robotics, et cetera, then all of these other bottlenecks will kind of dissolve along with them. So then you’ll get explosive technological progress. Maybe that’s an extra step — so it sort of relies on more assumptions, at least in the short term — but it’s a plausible step. And so this could all be insanely crazy, in a way that lots of people aren’t factoring in. Is that a fair summary of what’s going on, Fin, and the different positions in the debate?

Fin: Yeah, that seems right. Like, I’m trying to think about what actually is my view about what really matters most. So there’s one piece, which is: look, it seems very plausible that AI will accelerate progress in at least a bunch of just concretely important technological domains, right? So epistemic technologies; technologies involved in hard power — things like, you know, drones, and surveillance as well; technologies for spreading to space; technologies for, you know, digital technologies in some sense, for kind of, you know, economies of digital minds; medicine; and so on, right? That in itself seems like it could throw up a bunch of, you know, questions for how we kind of deal with and integrate these technologies in some way.

However, I think there’s also some story where we don’t actually see a great deal of especially novel technologies, but the world still changes dramatically. I think that could be the case if we get, like I alluded to, this kind of combination of robotics and flexible enough AI that can flexibly kind of substitute for humans, in a way that’s sufficient to start kind of autonomously scaling industry — the kind of factories-building-factories idea. So that is a world where, I think, you know, we get something like an industrial explosion, as others have called it. And I think, yeah, like I said, we already know how to build a lot of technologies which would be, at scale, really quite kind of wild — including, for example, going to space, although I’m not sure that’s the, you know, near-term most important example. And what does that take? Well, it takes enough AI progress to get the kind of flexible substitution, and it takes the robotics — but maybe not much more, other than, you know, barring political or social kind of barriers.

And then a final thing that seems really worth paying attention to is: in terms of AI progress, so far it’s mostly driven by people having ideas, and also scaling compute. There is this story that AI could automate so much of the kind of pipeline involved in getting better AI that AI progress itself — you know, however you kind of choose to measure it — inflects upwards, right? Maybe as part of a kind of feedback loop, right, where better AI makes better AI, and so on. So I don’t think that particular scenario is necessary for the others. But it seems like, if it did happen, it would, you know, bring the others forward in time, and be on its own just especially kind of hairy and kind of wild to witness.

Henry: So this is more of just a quick comment rather than a question, because I know we want to move on to other topics. But, you know, speaking as someone who identifies as techno-optimist, even techno-solutionist — I think, you know, that’s often used in a negative sense, but I think we actually have got a good track record of solving a lot of problems with technology — I think there is a failure in a lot of the kind of more optimistic discourse to connect ideas like the benefits of an intelligence explosion to the things that actually matter to everyone who’s not in our own little clique. Right? So, you know, if you say to the average person on the street, you know, here are some of the cool things that AI can do — like: I don’t want to go to space, right? You know, I don’t care about drone weapons, right? I don’t care about breakthroughs in solving the Riemann hypothesis. I care about my chronic pain. I care about the fact that, you know, my bills are high. I care about the fact that I can’t buy a house, right? So I think, you know — without wanting to go down that route too deeply — I think this is a big part of what’s currently missing in the kind of techno-optimist manifesto world: actually explaining how these breakthrough technologies can make a difference to the problems that people have right now.

Fin: Yep, I totally agree with that. I think there’s kind of, in some sense, failures on both sides — or at least both extremes — of debates around AI. You know, on one hand there’s a kind of unwillingness to recognise just how kind of radically great AI progress could be. And again — you know, you can point to specific reasons, but the more compelling reason is just this totally general, historically informed reason that technological progress, writ large, is good. Like, it generates new options and new affordances for people, right? There’s also a kind of, I think, lack of imagination about how further progress — or not just tech progress, but also just wealth — could be good. So I think there’s often this attitude of: look, I can see what I would do with twice as much wealth, right? But this kind of radically abundant world with sci-fi technologies — I’m off the train at that point. You know, surely I would just saturate. And, you know, you think about applying that to any point in history — it kind of turns out that we just fail to imagine the kind of new capabilities, new, you know, products, for example, that opened to us.

On the other hand — and I’m not ascribing this to you, Henry, or anyone in particular — but I think there’s often a kind of... there’s some attitude of swallowing that pill so fully that you just become — you know, almost as a reaction to kind of pessimistic attitudes — you just become, by default, excited about any kind of, you know, acceleration in the pace of change. Now, I think you could very reasonably be sceptical of crazy fast rates of change that, you know, various kind of Bay Area AGI-pilled people like to talk about. But if you do buy that these scenarios are possible — I think, at least, you know, my view is that the kind of appropriate reaction is to be extremely excited about the possibilities that come out of it, but... there’s a kind of missing mood or something. I think that, you know, I would feel like it’s more appropriate to be kind of not frightened, but just really quite apprehensive about those scenarios.

Dan: Okay, so let’s move on to grand challenges. So the paper makes this, I think, really kind of compelling case that even under relatively minimal assumptions — that is, you do at least have to take seriously the possibility of very powerful AI systems — but even under relatively minimal assumptions, we’re likely to get something like an intelligence explosion, which is likely to greatly accelerate the rate of technological progress. And then you can quibble with how fast that acceleration will be, and how it will happen, and so on. But there’s this other part of the paper which is then connected to this, where you reject this sort of view of transformative AI, or superintelligent AI, which I think used to be quite influential in terms of at least the way in which transformative AI gets framed. And it’s the kind of all-or-nothing view on alignment that says: look, if we manage to align advanced — and then superintelligent — AI systems with human values, well, those AI systems are going to be much smarter than we are, going to be much more capable than we are, so we can just then use those systems to solve all of our other problems. Alternatively, if we don’t align these systems, then, because they’re so much more powerful than us, they’re going to disempower us or eliminate us, and things like that.

And what you argue in the paper is: yes, alignment is a really important challenge — and kind of control, and these sorts of issues — but there are actually many grand challenges. And it’s not just as simple as: if you align it, then we can just punt all of our other problems to future superintelligence. So could you walk through that part of the essay, Fin, and then flag what some of these — I think you call them grand challenges — are?

Fin: Yeah, totally — I can try, at least. I think you put it really well. So the point is not that worries about alignment are totally ill-founded. So I think it’s clearly kind of a coherent possibility that we build AIs, and then we mess up, and then in some way they kind of, you know, cause a catastrophe, take everything from humans, it’s game over for the human race — fine. I think, you know — I kind of want to be careful about putting words in people’s mouths — but you could kind of point to a somewhat caricatured view, which says: so there is, let’s say, an alignment problem, you know, conceived of as a kind of well-described, kind of single, mostly conceptual problem, to be solved or not solved. And if it’s solved, we’ll have AIs which are kind of responsive to our, you know, intentions or instructions, in a kind of robust way. And if we get that world, then what we can do is — you know, they’re gonna be much smarter than us, and so we can ask them to solve all the rest of our problems, right? We can say: what’s the solution to ethics, or to politics, right? What should we be doing? What should we want? They can solve those conceptual questions, and they can also, you know, get to work on the kind of engineering and technological side of kind of giving us the future that we didn’t realise we needed.

Right. So, you know — the way I put it, I think I’m, you know, kinda loading the dice to make that sound like a, you know, somewhat naive idea, and maybe not many people truly believe this. But it’s, I think, worth really pressing on how this plan does not seem especially robust, even if we do end up with AI systems that are very sophisticated at thinking about the kinds of questions we’re interested in solving. One way of putting this is: just think about politics, like, historically and today, and kind of other social problems. To some extent these are problems because we’re confused — we’re conceptually confused, right? About what’s the kind of solution; how exactly, you know, is this kind of problem arising; are we missing concepts here? But for the most part, problems of politics are not problems which get solved with, you know, a very smart person coming up with a solution, right? I think, you know, there’s no special reason why politics as such will just end when we get superintelligence.

And then, more to the point, I think that you can imagine scenarios, for example, where we have AIs which are, you know, in some sense intent-aligned — they do what they’re told. But, for example, there’s a small number of actors in the world who, you know, get to tell the large majority of AIs what to do, right? They just hold a great deal of power, and that just presents all the problems of power imbalances that we’re familiar with through history — in some ways, they could be much worse. So that’s one problem: this kind of extreme concentration-of-power worry, which I can kind of elaborate on. But there are other problems too — or, you know, at least challenges. So one is: let’s say that there are ways to get really great outcomes in some kind of impartial sense, and maybe the AIs can tell us about them and help us reason through — but they’re just unpopular, right? Or people are not especially motivated to act on them. That would be one thing. And we can look again to history, where, at least from our vantage point, it’s pretty obvious that societies from the past have kind of engaged in these, you know, morally embarrassing practices, right? In some ways, that wasn’t for lack of a certain kind of awareness.

Yeah — and I think, you know, you can kind of list more specific challenges. So you could also imagine, in the extreme, just one very powerful actor chooses to lock in what they care about at a time. So they use kind of technological means to just seal themselves off from any further kind of, you know, change or evolution — which are, you know, processes we’re just used to today. You could imagine just kind of bad equilibria — kind of just failures of coordination. You know, again: even today we know what coordination failures are. We can write them down on whiteboards, right? Describe them very well. We still get them. So, you know — I realise this is quite kind of abstract, and I could kind of try to be more concrete if you want — but this is the kind of general story where just getting AIs that do what we tell them to do — even very wise AIs, or very smart AIs — could just easily not be enough to get the just truly great futures that we could get.

Dan: I mean, you list a number of grand challenges in the paper — you just alluded to some of them. Like, for example: AI takeover; highly destructive technologies; as you mentioned, value lock-in. Obviously we can’t cover all of these now. There are two that I’m especially interested in discussing — I think Henry is as well. One, which most people don’t immediately think of in this area, has to do with space and space governance. And the other, which touches on these sort of incredibly difficult ethical questions that are going to emerge even once we’ve got potentially superintelligent AI systems that are aligned — which is just this explosion in the number of digital minds, and more broadly kind of AI-based intelligences, and how we should think about those, and what an ethical, just society where we’ve got such systems — what that might look like. Let’s start with space governance — something I know nothing about, but I think Henry’s interested. What’s the connection here at all? Like, what’s the connection between — okay, we’re gonna build superintelligent systems — and space?

Fin: Yeah. So I think the general story would be: there is this kind of, you know, kind of worldview where you can imagine we get this kind of inflection in AI capabilities potentially quite soon, right? As a result, we get this kind of period of rapid technological change and progress, and just across-the-board growth — including industrial growth. That’s the kind of background, right? You might not buy that possibility, but you might think it’s plausible. Then you can ask: okay, on that possibility, just what else pops out as looking like, you know, a more near-term issue than you might previously have thought? Well, space comes to mind, right? So space is very big, and so you can do lots of stuff in it. Eventually there could be — in some handwavey sense — more stuff happening in space than is happening on Earth. We have a rough sense of what it would take, as a matter of engineering, to do a bunch of stuff in space — like, we’re already doing stuff, mostly in orbit. So again, on this kind of world of tech progress and industrial scale-up, potentially just vastly more stuff starts happening in space.

And then finally, it’s kind of amazingly under— how would I put this? — it’s very thin on institutions and laws and regulation, for better or worse. Partly because until fairly recently there hasn’t been a great deal of interest in figuring out, you know: how do we govern space? Partly because mostly it’s just been a matter of states, and it’s quite hard to get agreements between, you know, great powers. But if you just buy that a bunch of, like, stuff could happen in space — I can say more about what that means — you know, in, like, a matter of decades, rather than by the end of the century or so, it could be that we’re in this kind of period now where there’s a bunch of plasticity and kind of openness about what kind of institutions we kind of negotiate — which could then be really quite influential for kind of determining how things go.

Henry: So — just throwing my thoughts in here, because I’d love your feedback on this. I know, Fin, you’ve recently written a report on data centres in space. So I think two things that possibly are worth bearing in mind, when thinking about the potential for space to take off really, really fast in the next decade, are: firstly, there are economies of scale, when it comes to launch and construction in space, that we’re only really beginning to exploit. We’ve really seen this sort of— with SpaceX’s Falcon rocket, you know, launch costs have come down really far already. And we’re still talking about two or three launches a month of the Falcon rocket — and already that’s reduced costs massively. You know, if we move up to 30 launches a month, 100 launches a month — again, I think we’re gonna see bigger and bigger economies of scale. Launch costs could plummet. And equally, of course, data centres in space — this creates a pretty compelling use case, a pretty compelling justification, for putting things in space.

And equally, you know, one of the big obstacles to doing things in space, of course, has been that humans are spectacularly ill-suited to being in any environment outside of Earth. And if you try and do stuff even moderately far from Earth, right, you run into real latency issues in, like, controlling things remotely, for example. So once we have fully autonomous intelligences — you know, we could just send off an intelligence... I mean, this is a crude example, right? But let’s just say you wanted to build a factory on the Moon. Once you have reliable autonomous robots: dump them on the Moon and say, you know, build some solar panels, you know, build outwards, build some factories to start making either data centres or... beginning the Dyson swarm, right? This stuff could actually take off pretty fast, because, A, of the economies of scale, and, B, because some of the constraints don’t apply. Right — so this is how I get my spacepunk future that I’ve been looking forward to since I was a kid. I’m just curious, like: what do you see as the main obstacles to that? How realistic is it?

Fin: Yeah — so, I mean, I totally agree with what you said. One thing to, yeah, really press on is: I think futures in space which look like Star Trek, right — where it’s a bunch of humans in, like, the kind of big O’Neill cylinders, or space habitats, and they’re, like, orbiting around the Sun, or maybe they’re going on, like, you know, space missions to other stars, and literally humans are kind of spreading across space—

Henry: Seeking out new life and new civilisations.

Fin: Quite right. That does not strike me as especially plausible, or easy to imagine. Well — I mean, it could happen, right? But other things could happen so much faster, and in some sense just overtake those processes, right? It’s just wildly kind of inefficient to bring along this kind of, you know, sack of biological human, where you can do it kind of so much more efficiently. So — not imagining Star Wars. But, yeah — what seems, like, actually plausible? So one is, like you said, data centres in space — at least in the kind of medium to long run — could be quite appealing. The simple reason is there’s loads of solar energy in space, and if launch is very cheap, then you might as well throw them up there. There are, yeah, other uses of orbit — so there’s kind of telecoms, and there’s remote sensing.

Stuff on the Moon — so I’m, like, a bit sceptical of a lot of the things people say. So, for example, you know, mining the Moon for, like, nuclear fusion — like, helium-3 — as far as I can tell, doesn’t kind of pan out at all. Similarly, mining asteroids and then, like, throwing rocks back to Earth: there’s just so much, you know, material on Earth — I don’t really see the case for that. The thing that does seem, you know, just very technologically feasible to do is mining and manufacturing in situ with planetary bodies around the Sun. Where — I mean, look, this is, like, really quite far out at this point, at least in terms of the, you know, the kind of engineering required — so I’m not claiming we should kind of expect this to happen in a few years. But you can imagine, if you can kind of build, as it were, like, a seed — so it has some kind of initial robotics, and instructions to set up mining infrastructure, and the kind of infrastructure required to produce the next few seeds — then you have just this kind of self-growing process, and then you can kind of, again, convert that into whatever you want. You know, that seems like a pretty big deal once it becomes possible. Again, I think it’s just—

Henry: Yeah — the von Neumann, sort of Factorio, exponential growth of manufacturing capability in situ.

Dan: I have two thoughts that conflict with each other when I think about this. The first is: this seems hugely important, and we need more people thinking about this in an informed, rational, forward-looking way. Because at the moment, my sense is very, very few people are thinking about this — especially if you’re thinking about, you know, institutions and regulations and laws and so on. Very few relative to the scale of the challenges and the opportunities. Then the other thought I have is: this stuff is so weird, and it’s so, like, out of sample relative to what we’ve experienced in the past, that, like... can we think about this in a sensible way that would improve on just muddling through?

Like, if we go back 500 years: the average human being is living in this, like, micro-world. They’ve got, like, a really tiny community. There is this whole planet out there of, like, different peoples, but they know basically, like, very little of that. And there’s certainly no, like, global governance regimes and so on. And if you really try to think about it, you know — 500 years ago, go back 3,000 years ago — what would a world where you’ve got these huge things called nation states, embedded in the international regulatory apparatus — what would that look like? What should that look like? It’s just not obvious they’d make any progress on that. And now, when we’re thinking about space exploration, and people making use of, like, superintelligent machines and autonomous robots and so on to start exploring — what confidence should we have that really anything we can think of now is going to make much of a positive difference to how that’s likely to unfold?

Fin: Yeah — I think that’s, I mean, just, like, an extremely reasonable question. And, I mean, roughly for the reasons you give... So — you don’t think that this should be kind of top of the list of priorities for people to, you know, kind of bash their heads against and kind of work on, when there are so many just, you know, already quite imminent and real problems in the world. That said — I don’t know exactly how to articulate this, but, you know, I feel like we — as in, you know, humanity writ large — have accumulated some lessons about how to kind of arrange ourselves and build institutions that kind of broadly work. And I think a lot of the lessons that we’ve learned are quite robust — that is, they don’t depend on details, right? So you think about the kind of innovations of kind of liberal political thought: about the idea of rule of law, and, like, independent judiciary, and democratic mechanisms, and market mechanisms, and kind of economic concepts around, you know, efficiency and how markets work. You know, those don’t depend on, or mention, technological details. And in fact, you know, the past century has seen kind of, in some sense, wild amounts of technological change, but a lot of those ideas have, you know, really just proven quite robust.

And so I think, you know, at least we have this kind of null hypothesis, right? Which is: if it’s really hard to reason about how specific technological changes kind of change the game board, at least we kind of know there are certain arrangements which are just better than others. And that seems quite robust. So, for example, you know: arrangements where there is just no rule of law, and there’s some new resource, and there’s nothing better than just racing to grab the resource because no one’s gonna stop you — that just generally seems a bit hairier, and a bit worse, than, you know, building some institutions which kind of allow people to, you know, peacefully kind of negotiate ownership of these resources, right? And I would say at least we can do those kinds of things about space — which is, we can punt all of the details, a lot of the implementation details, to when we know more. In fact, we should. But at a high level, like, maybe we can start advocating for, you know: let’s just have some mechanisms to, like, you know, have these kind of peaceful alternatives to conflict, or to racing, for example.

It’s very tempting, I think, to try to, from scratch, like, devise new kind of mechanisms or institutions that are, like, perfectly fit to, you know, sci-fi technologies — and I think often we do need to at least extend the kinds of arrangements we have. But I think at least we can do that, right? I don’t think we need to — indeed, I think we often can’t — guess about how the details pan out. But those are things that we can punt.

Henry: So, as Dan mentioned earlier, another really big issue it’s worth drilling in on — perennial topic on the show — AI consciousness, AI welfare. So: you have some quite spicy views on this, right? You have quite a deflationary view of consciousness. Do you wanna tell us about it?

Dan: Actually, Fin, just before you do — I just want to frame it in terms of what we’ve already discussed. In terms of: okay, we get rapid progress in AI; this is going to trigger lots of technological progress; this is all throwing up grand challenges, and many of these challenges can’t just be reduced to — in fact, most of them can’t just be reduced to — aligning AI systems. And one set of challenges here is just: we are going to be building huge numbers of artificial intelligences — both kind of disembodied digital minds and, very plausibly as well, lots of embodied, autonomous, semi-autonomous robots. And that then raises huge ethical and political questions, like: how should we treat these systems? How should we live together with them? And I think many people think the first question you should ask, for any of those other questions, is just: are these systems actually conscious — in the sense of: are there lights on inside? Do they have phenomenal consciousness, as many philosophers would put it? And so you flagged this sort of question of how to treat and interact with digital minds as a grand challenge in your paper. You’ve also got this really interesting — and I think very persuasive — essay where you kind of give a deflationary answer to that question concerning AI consciousness. So: what is that deflationary answer? And do you think it’s relevant, then, to how we should think about these big questions about digital welfare and AI rights and so on?

Fin: Yeah. So I think my view is relevant, but it makes me just feel very confused, and not sure how to approach those downstream questions. So I sometimes think about analogies to biological life, right? So, you know, I think sometime around, let’s say, kind of the middle of the eighteenth century, the kind of prevailing view on what distinguished living things from non-living things — that is, you know, plants and mice from rocks and chairs — is some kind of distinctive, you know, life force that was shared by all and only living things, right? And had certain kinds of properties that meant that there were kind of discoverable facts of the matter about which things were living. Maybe it was imbued by God; maybe, you know, there’s something else going on. Now, it turns out — so we did a bunch of biology; we learned a huge amount about how living systems work. And in the course of doing that, one thing we learned is that there is no such life force, right? There’s just a collection of mechanisms that we associate with life. And there are also fuzzy boundaries to the kind of life concept — well, there’s an ambiguity about which kind of systems you might treat as living. Viruses, for example, right? And so, you know, you can ask this question: is a virus alive? If you were a kind of, you know, scientist in 1760, you would say, you know: we just need to go and discover that — you know, or it depends, kind of, what God chose. Now, what we say is: well, just tell me what you mean by alive. Nothing much else hinges on that, other than your choice of definition, right?

As a side note, by the way, I’ll say: I actually looked into, like, this life force thing and how real it was. My sense is it’s actually much more complicated — like, the history of science is just much more all over the place. But it’s a very convenient myth, so I’ll stick with it.

So, by analogy, I think that something similar might be going on with this kind of concept of consciousness — or, you know, phenomenal consciousness, you might say, right, in particular. Where it seems that people share an intuition — at least people who have decided to kind of, you know, read philosophy about consciousness — that there is a kind of discoverable and non-controversial, or let’s say non-ambiguous, fact of the matter about which systems are conscious. That is: you take some system; it might superficially appear conscious; but you can still ask, are the lights actually on inside, right? Is there something it’s like to be this chatbot? Or to be this mouse, right? The analogy to life would suggest that things are more complicated, and that really there are not facts of the matter — at least in this kind of sharp and non-ambiguous way.

Why might this be true? Well, there’s this kind of view on consciousness that’s often called illusionism — or strong illusionism — about consciousness, which says: you know, this thing that we have the impression exists, this phenomenal consciousness thing that has this kind of bundle of properties — like, it’s irreducible, it’s in some ways incommunicable, it’s kind of strictly always present or not present — this thing just is not real. So we only have impressions of it, but those are mistaken impressions in some way — in the same way that this life force thing that we think exists only appears to exist but doesn’t in fact exist. That doesn’t mean that consciousness in general — tout court is a phrase that I picked up from yesterday — is totally meaningless, right? So clearly we can ask interesting scientific questions about various phenomena we associate with phenomenal consciousness. We can also ask questions about which systems, you know, have beliefs about whether they are in fact conscious, right? So I think it’s clearly true that both of you and myself believe that we are kind of phenomenally conscious in some way. But we just can’t ask well-posed scientific questions about which systems are in fact phenomenally conscious — or, rather, we can, and the answer is: no systems.

And so that’s kind of where I think — at least where I’ve ended up, having tried to think about this a little bit. I think the arguments are fairly compelling. I think the result is just a lot of confusion about how to answer the downstream questions about: okay, what do we actually do? You know, we have all these ethical questions — how do we answer them now? But that’s where I’m at.

Henry: So I do want to take this to questions about the implications of this view for ethics. But just to sort of go through some of the kind of standard dialectical moves when illusionism comes up. So, you know, this is obviously an increasingly mainstream view, I think, in philosophy of mind and consciousness studies: that consciousness in some sense doesn’t exist — or at least doesn’t have the kind of properties we associate with it. This idea of the élan vital as a historical analogy — the life force. So the standard response to that from the anti-illusionist camp is: look, the reason we posited a life force to begin with was to explain certain kinds of structures and behaviours, right? So, you know: why is there this class of things that can reproduce, grow, and have all these properties, and then all these other things, like rocks and stones, that don’t have these properties? We need to come up with a mechanism to explain this difference in behaviour. Consciousness, according to the anti-illusionist camp, is fundamentally different, right? Because we’re not invoking it to explain behaviours, right? We’re invoking it to explain a manifest phenomenon that we are all immediately acquainted with, right? Now, the illusionist can say: well, you know, that’s sort of begging the question, right? If my view is correct, then there is no manifest phenomenon. But I think that’s a very hard pill for people to swallow. You know, there’s this kind of standard interior pointing move that says: but come on, this — this, inside my head — as if you can point inwardly. That can’t be an illusion. There’s something real there. So I do think there is an interesting disanalogy there with the kind of élan vital move — that, you know, we’re not trying to explain behaviour. So, like, you know, you’re welcome to give your reflections on that.

But I also wanted to say that, you know, one reason I think it’s quite hard to at least get rid of consciousness, without just smuggling it back in under a new name, is that it seems to have this kind of direct connection to value. You know — so if I’m gonna go for an invasive piece of surgery, and, you know, I’m about to be given a new type of anaesthetic, and I say to the doctor, you know: will this render me unconscious during the operation? And the doctor says: well, you know, it’ll reduce your attentional faculties, it’ll knock out your working memory, it’ll reduce your social cognition... It’s like: no, no, no, doc — I wanna know, am I gonna be conscious, right? This is the property that matters to me, right? And in the same way — if I’m talking about shrimp, right? If I wanna say: do I wanna allocate money to shrimp welfare, right? You know, you can say: well, you know, here are some interesting facts about shrimp behaviour and shrimp social complexity and shrimp intelligence... It’s like: that’s, you know, maybe interesting, but that’s ultimately not what matters morally. What matters morally is: are there lights on inside, and do those lights hurt or feel good? If so, they are welfare subjects; if not, they’re not. So, like, every illusionist attempt I’ve seen to try and sort of grapple with this problem says: well, okay, you know, maybe consciousness isn’t, you know, what matters, but this quasi-consciousness, or something like that — you end up reinventing this notion of consciousness. So, I don’t know — do you have any thoughts on this? Am I being harsh on the illusionist response?

Fin: Yeah, totally. So, I mean, yeah, I think two points there. I guess the first, you’re saying, is: unlike this analogy to biological life, the thing that phenomenal consciousness is invoked as a concept to explain — it’s not this kind of appearance, you know, of just various behaviours; it’s this kind of manifest, obvious, first-personal reality that I am conscious right now, right? And so one response is: well, I can explain why you say that without ever invoking phenomenal consciousness, right? Because I can talk about how your brain works. I think that is not enough, because you might just say: look, I’m not trying to persuade you of anything — I’m just saying I know, right? I appreciate I can’t tell you; but, you know, anyone, you know, can just kind of introspect and appreciate that they are conscious, right? Then the response has to be — I often think this is just where the argument ends, because there’s a kind of just appeal to brute fact, right? For better or worse.

One potential response is: so, the illusionist is not trying to kind of eliminate beliefs, or judgements, or, you know, kind of epistemic concepts. So — I have, kind of, again, reflecting internally — so, not trying to persuade you, but I’m thinking — so, I certainly have a belief that I’m phenomenally conscious, and that’s non-mysterious, so I can help myself to that, even if I’m an illusionist. What more is there to explain? I think that at least starts to feel very confusing.

Henry: But I mean — just very quickly — you know, a base-model LLM will believe it’s conscious, right? Yep, presumably. And so you might say: look, you know, you believe you’re conscious; so does LaMDA; so does GPT-3 in its untrained form. We can just give explanations for why both of you think that — problem solved. But clearly there is a disanalogy in why we believe we’re conscious. You know, GPT-3 or whatever believes that it’s conscious for some really bad reasons. At the very least, my reasons for believing I’m conscious are very different.

Fin: Yeah — I mean, it’s very hard to know what to say. So, I mean, obviously I share your intuitions, because I’m human, right? So I have this intuition that there is something extra. What I’m saying is that if my kind of criterion for, like, submitting this as an argument is that I can articulate it in any more detail than just “this seems wrong” — that is, without just begging the question — I really struggle to know what to say, right? So, you know, it’s conceivable that I actually just can’t articulate to you what is the difference between me and some AI system that, you know, merely has the kind of, you know, epistemic stances associated with believing that you’re conscious. So, look — there’s not much more to say on this, or at least I don’t know what more to say on this. I’m sure other people do have more to say.

Henry: I mean, I think there is, like, several shows’ worth more to say on this. But with that in mind, maybe we should move on to the ethics side, which I’m also really curious to hear your thoughts on.

Dan: Well — actually, before we move on to the ethics side, could I just sort of chip in to try to summarise, especially for people who aren’t in the weeds in the kind of philosophy of consciousness literature. So, Fin, I take it, as someone with your kind of illusionist view, they’re gonna say: look, we can give functional, behavioural analyses of different psychological states and processes. So we can talk about: this system can perceive, it can imagine, it can deliberate, it can feel pain even — in the following sense — and then you would cash that out in terms of the capabilities, the propensities, of an information-processing mechanism. So there’s no issue with saying that a system has psychological states, including those that we typically think of as being connected to ethics, in that sense. Now, there are some people who then want to say: okay, it can perceive and imagine and deliberate and plan and act and feel, in those functionally, behaviourally defined ways — but I’ve got this extra question, which is: are the lights on inside? And your view, as I understand it, is just to say: well, there’s nothing else to explain. There’s no extra property of the universe. We can still care about whether a system can perceive and feel and act and so on. But once we’ve cashed that out functionally and behaviourally, in kind of usual, standard scientific terms, there’s no extra deep metaphysical question about whether the lights are on inside.

But there is this question of why animals like us are disposed to say that there’s an extra question — why we’re disposed to say: are the lights on inside? And if we can offer a convincing explanation of why we’re disposed to say that, that doesn’t itself appeal to there actually being some extra property, then we’ve done everything we need to do from the perspective of a kind of satisfying scientific worldview. Is that a summary of your position?

Fin: Yeah — I think that’s... you said it better than I could have said it, again. What can I add to that? It’s useful to think about other cases of debunking. So if someone comes to me and says, “You’ll never believe it — I saw a UFO last night,” then I can ask: what grounds do I have to believe that there was actually a UFO, right? What I can do is — let’s say I, you know, have access to their brain state — I could explain why they said, in the moment, that they saw a UFO, without ever kind of using the word UFO, right? So, you know: neurons fired in their brain caused them to make those sounds. That is actually not enough to debunk whether or not there was a UFO, right? Just commonsensically. However, suppose that I, you know, checked the news: there was a weather balloon, right, and it had a light on that looked like a UFO. Then I can debunk whether it was really a UFO, because I just never have to invoke anything like a UFO — even if I’m trying to conceal it in, you know, kind of other language. So that would be the claim with illusionism.

And then I think there is, like you said, something that the illusionist still has to explain — which is, I think, so far not explained, or at least there’s no consensus on what the story is — which is: if consciousness is an illusion, it’s super weird that basically everyone who thinks about consciousness ends up with the same illusion, the same impressions. And basically no one seems to be able to convince themselves out of it — unlike virtually every other, you know, perceptual illusion or kind of epistemic mistake that people can, you know, tend to find themselves in, but then persuade themselves out of. Like, you know, imagine a visual illusion where you can just kind of look at it another way and realise you’re wrong. So that remains to be explained. So illusionism is not, I think, to date, a full explanation of what the hell’s going on with consciousness. But it’s a kind of proviso, or a kind of: look, here’s the direction that seems most likely to pan out — which just denies that consciousness is real at all. Yeah.

Dan: I think it’s worth just saying two things that make me very sympathetic to this view, which have mostly just been implicit in what we’ve been saying, but are worth kind of making more explicit. One is, like: how compelling this view seems partly depends on how you frame it. So if you say “consciousness is an illusion”, that sounds bizarre. But really the view is not that consciousness is an illusion, but that it’s an illusion that there are these intrinsic, ineffable properties that elude functional, behavioural explanation. And if you put it like that, it starts to seem, I think, much more reasonable.

And then, second: in the background here is just the idea that, like, this seems continuous with ordinary scientific investigation, of a kind that has been basically the only source of knowledge about the world we’ve ever produced. And to the extent that the illusionist position seems much more coherent with that kind of scientific epistemology, that’s, like, a really, really strong point in its favour — which should be relevant to this debate you two are having about, like, who’s begging the question against who, and which intuitions are we relying on. If you just step back and ask, like, which position endorsed here is more aligned with a general scientific attitude to these questions — I don’t know, my sense is clearly it’s this illusionist position. But my sense is, Henry, you disagree with that, at least.

Henry: Well, here’s sort of one question I have — a kind of meta-question for the illusionist — which is: is this a view you’ll ever be able to convince people of? And if the answer is no, then it’s almost — okay, maybe that’s not irrelevant, right? But I do think this question of how do we achieve something like consensus in this domain is a question that, you know, the non-illusionist might have an answer to. Like, if you’re just a non-illusionist reductionist, you might say: look, we are gonna figure out what consciousness is. It’s gonna be amazing. It’s gonna be the best theory ever. Maybe we need superintelligence to help us figure it out, but we’ll be able to pin it down, and we’ll be able to see which animals have it, which animals don’t. We’ll be able to salvage a lot of our intuitions. It’ll be nice and clean. Everyone will be convinced. You know, they’ll come up with a new Nobel Prize just for philosophy, just to give to whoever comes up with it. But, like, the illusionist story — I’m not sort of sure what happens. It’s like: you just get progressively more convincing arguments, and, like, fewer and fewer people believe in it over time, and just gradually it falls out of the discourse. I mean, maybe — maybe that would happen. But, yeah — I mean, now, that’s not an argument against the truth of illusionism, but I think in some ways it would be much messier and worse if something like illusionism was true. I also have my other first-order arguments — or first-order disagreements — with illusionism, but maybe I’ll save those for another time.

Fin: Yep — okay. Yeah, I’ll be curious to chat about that at some point. So, I mean, I agree and disagree with what you said, Dan. I think that it just clearly is very hard to internalise. So, you know, I would say that I’m very sympathetic towards illusionism, although not fully confident. I still just can’t kind of wrangle my intuitions to get on board, right? This is kind of the System 2 part of my brain that’s doing all the work. And I think it’s not possible to sympathetically reframe illusionism to make it sound, you know, kind of obvious after all. Like, you really are just denying something that is kind of just pervasively and strongly kind of intuitive. One analogy might be to the sense that people have of a kind of continuous and distinctive self — where I might think that there is always an unambiguous fact of the matter about who is me at a given time, as long as I’m still alive. So that’s an example where I think philosophy — philosophers — have done a good job showing that those intuitions are false. And I find those arguments moving. And, you know, the day after I read those arguments, I’m like — my intuitions haven’t changed, right? Except when I kind of go and kind of read about that stuff again. I think that shows—

Henry: Hold on, hold on — you don’t feel like you’re living in the open air, as Parfit put it, you know, once you get rid— Yeah. Yeah, exactly.

Fin: Right, right — my glass tunnel. Yeah. I haven’t read that passage enough to internalise it. But, you know, it shows, I think, that part of the job of philosophy is to try to kind of challenge pervasive intuitions, even when it doesn’t succeed in kind of actually undermining them in any lasting way. But then the second thing you said, Dan, I think I agree with — which is that illusionism is in some sense the view that consciousness goes the way of every other kind of empirical or philosophical mystery historically, which is that it gets in some way dissolved by just standard, you know, conceptual or empirical methods, right? So it would be exceptional in that sense — uniquely exceptional — if we needed to kind of devise new methods to understand it.

Now — and then, I guess, going back to Henry’s point earlier, which was: well, hang on, here’s a kind of, you know, huge problem, which is: we base so much of our thinking about ethics on consciousness concepts, right? So I care about helping, you know, other people, or non-human animals, you know, in some cases exactly because I think that they’re capable of suffering, or feeling, you know, joy — and I, you know, want to avoid the former and, you know, promote the latter, or whatever. And so — I mean, here’s some reactions, right? One reaction is that — and I read Henry as saying this — you can use this as an argument against illusionism. Which is: I think my kind of ethical intuitions are true, and they imply consciousness — or the kind of phenomenal consciousness required to kind of ground those intuitions, right? Therefore, we have to help ourselves to those consciousness concepts. I do think that’s a good argument.

Henry: Or maybe the kind of Moorean version of this is: I’m more confident that there is an ethically salient difference between animals and rocks, grounded in a specific kind of internal property — I’m more confident in that ethically salient distinction than I am that any clever philosophy-of-mind argument to the contrary is false.

Dan: Moorean in the sense of G. E. Moore, who used that style of argument to dismiss certain kinds of arguments.

Henry: Thank you — thank you, Dan. Sorry, yeah.

Fin: Good stuff — thank you, Dan. So — how to react to this. One thing to say is that this might be a bit cheap, so I’m kind of curious for your reaction. When I think about kind of ethical attitudes, especially historically, they don’t seem to me grounded explicitly in kind of phenomenal-consciousness-related concepts. Or at least they don’t always seem to be. Sometimes they seem to be grounded in concepts which I think are confused, right? So maybe you might care about helping someone to save their soul. If you think there are no souls, the fact that people had kind of soul-saving intuitions is not an especially compelling reason to believe in souls. But also you can kind of ground ethical intuitions in terms of things like, you know, fairness, or objective goods, or preferences. So it’s not clear that you’re totally at sea. But then, finally, even if you were — even if you thought, my god, like, I have no idea what kind of, you know, foothold I have here — I don’t know... maybe I’m, like, a bit sceptical of the kind of Moorean-type arguments in general. I’m like: well, maybe that’s a bullet you have to bite. It seems possible to me.

Henry: So I guess one point you could make on this — that I would expect you to be sort of sympathetic to, Fin — is that human moral psychology is just a really messy grab bag of System 1–type responses that we’ve evolved over the years, mixing everything from disgust and basic survival-related feelings and intuitions to all sorts of more contractarian, or sort of social-dynamical, intuitions. And, like, one of the achievements of twentieth-century politics and ethics is realising that some of these are probably less well-founded than others. And that, you know, the idea that I have a greater moral obligation towards someone just because they’re in my in-group is maybe not a great way to think about things. And, more generally, you know, the broader idea of the expanding moral circle, right — which, again, I expect all of us are to some degree sympathetic to.

And this has been grounded primarily on a basically sentientist framework. Maybe not essentially, right? To be clear, I don’t think you can’t have an expanding moral circle without being a sentientist. But, like, Peter Singer is a sentientist, as are many of the loudest advocates of expanding the moral circle — and particularly, you know, when we think about non-human animals and so forth, right? There is this kind of sentientist idea in the background. So I would say that the move away from, you know, moral obligations grounded in souls and honour and piety, towards a more harm-based framework, has been a real form of moral progress. And illusionism threatens to kick away, you know, the last supporting pillar there.

Fin: Yeah. I mean — so, like, to be clear: my actual view on what the hell this means for ethics just is overwhelmingly a view of being, like, very confused, and, like, slightly worried, for the reasons you say. So, I mean, some things to say here. One is: imagine that you have a faith, and you’re questioning your faith, and you’re worried that, you know, on the other side is this kind of moral void, right? If God is dead, everything is permitted, and so on. I think — I don’t know if either of you are God-fearing men — but if you are an atheist, I think most people have this attitude of: it turns out that, you know, you still just care ethically about other people, right? And if you kind of originally thought that your ethical attitudes were grounded in belief in God in some way, and then you lose your faith, you can kind of either change your motivations, or realise that, after all, your motivations weren’t grounded in that thing you thought they were.

I think that kind of goes on in my head, where I think: okay, what if it turns out that phenomenal consciousness is not real? So I previously did think that the most kind of compelling way to ground my ethical intuitions, you know, in a kind of systematic and fair way, is in terms of facts about how much pleasure and pain there is in the world. Now, you know, if illusionism is true, I can’t help myself to those facts. I don’t feel like I therefore have to give up on the intuitions I have — that it’s good to help other beings, and, you know, to prevent needless suffering, for example. Yeah — I don’t feel like I need to kind of revise those attitudes. It feels way more kind of plausible to me that those attitudes just don’t really necessarily depend on other metaphysical beliefs. Now — you know, exactly what do I do? Like, how do you kind of cash out ethical views in, like, a systematic way? I have no idea. But my sense is that there’s no need to panic, if that makes sense.

Henry: So — two very quick reflections on this, and then, you know, I’ll pass back to you, Dan — or Fin, if you’d like to respond. So I think there is this kind of viable project, that you’ve seen some people start to sketch out, where, you know, we can take our intuitions around phenomenal consciousness and maybe salvage some of them, maybe lose others. You know: maybe rather than saying this pain matters because it’s conscious, we say this pain matters because it’s globally accessible, or this pain matters because it’s the target of a higher-order thought. And even as I say that, right, I think the shape of the worry starts to come through — which is: it’s not clear that any other properties in the brain can carry that same sort of foundational role as phenomenality. There’s just an obvious wrongness — one that’s very hard to vocalise — about the badness I experience when I am in conscious pain. And if I say, look: negative reward signals that are globally broadcast are intrinsically bad... that suddenly starts to look really mysterious to me, and much weaker. Go on, Dan, yeah.

Dan: Can I just double-click on that? So presumably the view would not be, you know: pain is bad because it involves global workspace architecture under such-and-such conditions. Presumably the view is: pain is bad. What is pain? Well, pain is the following. You’re not giving a kind of justification-based argument, where you’re saying pain is bad because of this other thing. You’re just giving an account of what pain is. And then there’s a question, which is: well, why is pain bad? But that’s going to be a question for anyone. I don’t get why that’s a distinctive question for the illusionist. What am I missing with that?

Henry: Well, I mean, I think part of the answer that the non-illusionist can give is that pain is bad because of its phenomenal character, right? The felt unpleasantness of pain is what makes pain bad. And if you come along and say, well, that phenomenal character doesn’t exist, then it’s just not clear what you’re gonna ground that badness in.

Dan: Yeah — I worry that there’s just sort of talking in circles there, though. Why is pain bad? Because it involves suffering. Why is suffering bad? Well, it’s, like, a negative experience. Ultimately, I think you’re just going to be able to cash these things out functionally and behaviourally. Then there is a question about, like: why is anything bad? And that’s just, I think, not going to have a satisfying answer — because I’m a nihilist, for independent reasons. So I don’t think anything is going to ground our moral judgements. So, like, in some sense it is all just a kind of make-believe that we are motivated to play, as a certain kind of pro-social ape that’s been encultured in a particular way. So, like, there are complicated, like, meta-ethical questions. But I just don’t get why I should be that bothered by the implications of illusionism specifically. Like, it’s just giving an analysis of what pain is. It’s not, in and of itself, trying to give you an answer to the question of why pain is bad.

Henry: So I would agree that on a non-illusionist position there is still this question of, well, what is it about the phenomenal character of pain that’s so bad, right? But the non-illusionist is not claiming to have explained this whole situation, right? Whereas the illusionist, I think, is already more confident in an answer than the non-illusionist — namely, that this phenomenal property, this kind of thing, doesn’t exist. And just to offer a historical analogy here again, right? So, you know, the illusionist can appeal to the élan vital. Here’s another analogy you could appeal to, you know: if you read ancient Greek philosophy, it’s full of people like Parmenides and Zeno, who confidently assert that, like, motion is impossible, or that there’s only one object. And they come to all these really radical, well-argued logical conclusions, right? And the answer, usually, to them is often things like: well, they didn’t have convergent series back then, right? They didn’t have the maths to understand this; or, like, physics was just at such a primitive level that they weren’t in a position to realise that these things that they thought somehow were logically impossible were actually broadly resolvable within a broader theory of physics.

And I think we know so little about the brain, right, that the illusionist’s call that, like, there can never be anything like this — this clearly just doesn’t exist — is premature, right? Our knowledge of neuroscience — neuroscience is such an infant field, right? In the sense that, like, we still don’t have any really good models of how just basic stuff, like representation, works in most of the brain. And so I think to, you know, to give up and say, look, there can never be anything that actually satisfies most of our intuitions about this — might well be premature. And as we start to understand more about how coding works in the brain, how we create this kind of unified psychology, right — something might well emerge that actually is a target for some kind of reasonable identification with at least enough of our pre-theoretical ideas about consciousness, for us to say we found consciousness, rather than consciousness doesn’t exist.

Fin: I’m not sure how persuasive I find that. So — maybe I’m misreading you. So suppose that we found some kind of brain process that just perfectly correlates with our intuitions about kind of whether and when phenomenal consciousness is present. I agree that maybe that deserves just the name “consciousness”. I think it’d also be, in some sense, good news, right? Because we don’t have to confront these ambiguous or edge cases, where, you know, some kind of associated properties are present but others aren’t. I think, nonetheless, you could still maintain an illusionist view — which is: we remain metaphysically confused about consciousness, phenomenal consciousness. That is, we are confused, for example, that it’s conceivable or possible you could have these particular brain processes without phenomenal consciousness coming along for the ride — or you could kind of wonder whether, you know, there’s a possibility that it doesn’t, and so on. So it seems to me that illusionism is not kind of trying to make a claim about what we’ll discover in the course of doing neuroscience — although there are still quite interesting and important questions about what we’ll discover, which no one knows, including illusionists.

Henry: Then it sounds like — so, just very quickly, I’d say that the danger here, then, is that there’s a kind of motte-and-bailey thing going on: where the illusionist starts out with a very specific and very strong set of claims, within a much broader set of broadly reductionist, physicalist positions, and ends up saying, well, look: as long as there’s no supernatural stuff, right — or no fundamentally different substance, no non-physical property — then illusionism just wins, right? We claim all of physicalism as our territory, right? Whereas I think there are viable physicalist views, that we are just not in a position to evaluate yet, that are not illusionist in the strict sense of the term. And we just don’t know which of these views yet are correct.

Fin: Yeah — go ahead, then.

Dan: I think it’s worth wrapping things up with the following. So: you make a really good case in this paper, Fin, with Will, that, look — rapid AI progress, of the sort that is happening and is very likely to continue in the coming years and decades, is going to produce lots of artificial intelligences that are very sophisticated, have agency. Some of them, as I’ve said, are going to be kind of disembodied digital agents, in some sense. Some of them will be semi-autonomous, if not completely autonomous, robots. And whatever we think about these questions concerning the philosophy and science of consciousness, this is going to create a huge set of, like, ethical and political questions. How should we interact with these systems? How should we treat them? What kinds of societies do we want to build when we’re living in the context of non-human intelligences that are as cognitively sophisticated — maybe much more cognitively sophisticated — than we are?

So this is, like — to put it mildly — a grand challenge. And also one that gets basically no attention among people who are currently academics and researchers. Obviously there is a small community of devoted researchers who are thinking about this in a really serious way. But, like, relative to the scale of these challenges, it seems like it’s dramatically under-researched. So my view is, like: all of those challenges are real, and, in a sense, we can start thinking about those challenges — both intellectually and in terms of policies — without settling these really difficult questions about the nature of consciousness. So that’s how I would sort of want to wrap up that conversation. And also the message I would want to send to people who are listening to this or watching this: that this is a hugely important set of questions that is currently massively under-researched. So — do you two agree with that? Do you have anything to add?

Fin: I can try pitching in. I, unsurprisingly, do agree with that. One way of putting this is to say: settling questions on AI consciousness, and the metaphysics of consciousness — that strikes me as neither necessary nor sufficient to figure out, you know, how the hell to kind of arrange, you know, our institutions and society at large around, you know, this potential influx of digital minds, right? So: it’s not sufficient, because — let’s just say we have answers on, you know, which AI systems are conscious in some sense. Well, there’s another kind of, you know, kind of being where we have pretty confident answers — which is humans, right? And it turns out we still face questions about how to, you know, live together with each other, and associate with each other, and what kinds of, you know, legal, political institutions to build. And so just figuring out, you know, which systems are conscious does not tell us how to integrate with them. Should we ascribe them rights, or obligations, and so on? Those are just, as far as I can tell, largely independent — and extremely important, tricky, interesting — questions to answer.

And then I’d say it’s also not necessary to answer these questions — at least plausibly not totally necessary — in the sense that I can imagine, as we kind of approach a world where we do just, you know, live with some of these AI systems — we actually get to interact with them — it’s kind of plausible to me that some of these questions around consciousness will start to feel a little less kind of relevant, and will sort of kind of go from a kind of thought-experiment land into: I just, like, have these attitudes by default to these systems. In the same way where, if we watch a, you know, sci-fi movie where there’s an alien race, and they’re just very sympathetic — they seem very understanding and wise and empathetic — we just don’t tend to find ourselves speculating about whether they’re truly conscious or not. And we could imagine some of that happening as well. So that would be my kind of... my claim — which is not to say that these questions are not important, around consciousness. It just seems to me that there are so many other questions as well that we should get right.

Henry: So I agree with half of that. I certainly agree it wouldn’t be sufficient — insofar as I think our moral landscape is complex and pluralist, and we can acquire obligations that are not grounded purely in consciousness. You know, I think you might think we have obligations based on things like keeping promises and so forth, as well as, you know, maybe obligations grounded in relations to communities and so forth. And I expect, as sort of worlds of agents emerge — of ubiquitous AI agents that we’re forming close relationships with — yeah, there are gonna be a whole lot of moral questions that, you know, aren’t immediately settled by figuring out which ones are conscious and which ones aren’t.

That said, I do think it’s necessary, at least for a full moral theory, to have a sense of which beings are conscious — or — even if the answer to that is, like, just demonstrated conclusively that consciousness is the wrong property, right? That would be one way, right? But until we get an answer to that question, I think there’s gonna be this sort of yawning gap at the heart of our moral theories. Because pre-theoretically — and, frankly, theoretically, for maybe the plurality, if not the majority, of sort of ethicists and philosophers of mind — consciousness is a very, very, very, very special moral property. And to your example, Fin — I think, again, I sort of agree with you, right, about sci-fi and other shows. I think the reason that, you know, we don’t hesitate to ascribe moral status to Data, or, you know, the Klingons, or R2-D2 and so forth, is because we ascribe them interiority — we ascribe them an interior mental life, right? And I think that’s gonna happen: I absolutely expect that we’ll start to see widespread attributions of consciousness to AI systems by the general public. But that’s a different story from: we’ll just stop caring about consciousness. I think, instead, we will be convinced by the kind of thick relational properties that form between humans and AI systems — you know, that’ll drive our consciousness ascriptions.

Fin: Yeah — do you know what, I think that you’re right, in the sense that I think I do actually agree that in some important ways it is necessary to get clearer on metaphysical questions around consciousness. So maybe I should, yeah, reduce my claim down to just this point: that, very plausibly, once we enter this world where we get to interact with sophisticated AIs, it’ll feel potentially just quite natural and automatic how we should kind of relate to them and empathise with them. And we won’t be constantly asking ourselves, you know, which kind of metaphysics of consciousness is most plausible. I think it’ll, you know, potentially come quite organically. But we might be wrong. So—

Dan: Okay. With that, Fin — this has been great. Thanks for coming on.

Fin: Thank you both so much. Great questions.

Henry: Thanks so much, Fin. It’s been great.

Discussion about this video

User's avatar

Ready for more?