There is currently much debate about how impressive frontier AI models really are and what role they can play in research, including in academic philosophy. To get some evidence, I asked GPT-6 Astra, OpenAI’s new model, to review my book, ‘The Social Roots of Delusions’, co-authored with Sam Wilkinson and Kengo Miyazono and recently published by Oxford University Press. (For a brief summary of the book, see here.)
I used this prompt: “You are a world-leading academic. Please write an insightful review of this recent book published by Oxford University Press. The review should be suitable for publication in a top academic journal. It should be no longer than 2,000 words.”
I used the ‘pro’ setting for the amount of “thinking effort” (i.e., how much inference compute) the model uses, and turned off memory so it didn’t know it was talking to me, and so the book isn’t used as training data.
A few quick thoughts:
The review is very good.
If a human wrote it, I would be impressed not just by their detailed reading of the book but also by their intelligence, depth of knowledge, and insight.
There are some annoying AI-isms in the writing style (“not X but Y”).
The review’s objections are reasonable, although I think I could address them. (I also tend to have this feeling whenever humans criticise my work…)
We are now at the point where people are simply kidding themselves if they think frontier AI models lack genuine intelligence, understanding, insight, or academic utility. I’m unsure exactly how research should change now that we have access to smart expert intelligences (that will only become more impressive and useful with time), but I am certain it must.
Delusion and Epistemic Membership
Daniel Williams, Sam Wilkinson and Kengo Miyazono, The Social Roots of Delusions. Oxford: Oxford University Press, 2026.
The deepest provocation of The Social Roots of Delusions is not that social life can make us irrational. It is that remaining in contact with reality is itself a precarious social achievement. Williams, Wilkinson and Miyazono ask readers to reconsider the contrast between the isolated person whose beliefs resist correction and the community whose members collectively sustain an extraordinary conviction. These are not manifestations of a single underlying pathology. Rather, they reveal different vulnerabilities of epistemic interdependence: the loss of access to support that normally constrains individual error, and the mobilisation of that support to protect collective error. The book’s considerable achievement is to make these possibilities illuminate one another without requiring a uniform explanation of delusion.
Its three parts address classification, popular delusions and clinical delusions. The first advances the most philosophically revisionary claim. Drawing on Edward Craig’s function-first approach to knowledge, the authors ask what attributing delusion does within social life. Their answer is that it expresses “folk-epistemic bafflement” and identifies someone as a “rational lost cause” concerning a particular belief. Attribution does not merely identify an unreliable informant; it recommends abandoning ordinary attempts at rational correction. This genealogy loosens the conceptual connection between delusion and pathology and accommodates substantial heterogeneity among the beliefs attracting the designation. Delusionality becomes a social-epistemic status rather than an intrinsic property awaiting discovery (pp. 54–70).
Part 2 explains how otherwise capable individuals sustain “popular delusions”. The target is not religion, ideology or conspiracy belief indiscriminately, but a cluster comprising pronounced yet selective epistemic irrationality, compartmentalisation, emotional attachment and distinctive sustaining practices. The proposed explanation centres on socially motivated cognition. People can internalise claims they are motivated to propagate, while convictions can signal allegiance and cooperativeness. Sincerity is therefore compatible with strategic social function; the account is not a theory of universal cynicism. Crucially, beliefs need not be comforting to be advantageous. A terrifying narrative can justify hostility, secure standing or demonstrate loyalty even while making its adherent miserable. The distinction between wanting a proposition to be true and wanting the benefits of believing it is illuminating.
The account becomes most compelling in Chapter 5. Motivated cognition faces a rationalisation constraint: people cannot simply decide to believe whatever serves their interests. They need apparent grounds. The authors show how communities can collectively satisfy this constraint through “epistemic reward systems”. Dissent is discouraged, favourable evidence amplified and intellectual labour rewarded for producing increasingly persuasive defences of preferred conclusions. These practices also alter higher-order evidence: the apparent conviction of trusted others becomes further reason to believe. The discussions of white supremacy and QAnon illustrate how radically different formations can sustain such processes (pp. 142–148). Collective irrationality here is not merely individual irrationality multiplied. It involves the social production of an evidential environment within which otherwise implausible beliefs acquire credibility. This moves the analysis beyond familiar appeals to gullibility or an absence of critical thinking: considerable critical ingenuity may be invested in keeping a conviction intact.
Part 3 resists an equally familiar reduction in the opposite direction. Acknowledging delusions’ social character does not require locating their cause in a defective social-cognitive module. Against the “suspicion system” approach discussed in Chapter 7, the authors argue that detecting concealed intentions demands access to too much contextual information to be plausibly assigned to a highly encapsulated mechanism. Their alternative builds on multifactorial approaches, especially Daniel Freeman’s work, to explain three phenomena: the importance of social adversity, the recurrent social themes of delusions and resistance to social evidence. Perceptual disturbance, cognition, affect and behaviour remain indispensable; sociality specifies the circumstances in which these factors acquire their explanatory significance.
The treatment of persecution is particularly effective. Suspicion can generate anxiety and withdrawal; withdrawal can diminish corrective contact and increase vulnerability; these consequences can intensify the original suspicion. Protective responses thereby become constituents of a damaging feedback process (pp. 212–215). The distinction between “misfunction” and “malfunction” is important here: a mechanism can operate unsuccessfully because its inputs or conditions are abnormal, without being intrinsically damaged (pp. 247–248). This is not an argument against neurobiological explanation. It is an argument against assuming that every explanatory contribution must identify a further defect inside the person.
Chapter 9 sharpens the account by distinguishing testimonial isolation from testimonial discount. Evidence may fail to reach someone, or it may arrive but be assigned little weight. The distinction prevents apparent resistance to correction from being treated automatically as an individual incapacity. Moreover, discounting testimony can reflect ordinary epistemic vigilance under extraordinary conditions: apparently compelling personal experience may outweigh statements from people whose competence or sincerity is distrusted. The authors’ attention to these asymmetries makes “insensitivity to evidence” look too coarse a description. The relevant questions concern which evidence, delivered by whom, under what conditions of trust.
The resulting synthesis is substantial. Its originality lies less in announcing that delusions are social than in connecting social motives, distributed evidence and classificatory practices. Nevertheless, explanatory inclusiveness has costs. The authors candidly acknowledge that their framework remains largely qualitative, that mechanisms of compartmentalisation require clarification and that several central hypotheses await direct testing (pp. 275–277). The important next step is not merely to add further contributing factors, but to specify contrasts that discriminate among explanations. Under otherwise comparable conditions, when should changing the source of testimony alter conviction more than changing its content? Which features of a community’s reward structure make it responsive to correction rather than simply cohesive? Such questions would turn the framework’s breadth into testable predictions.
There are also conceptual difficulties that empirical elaboration alone will not resolve. First, the genealogy of attribution does not establish the authors’ stronger non-descriptivism. They maintain that delusion attributions are appropriate or inappropriate rather than accurate or inaccurate (p. 65). Yet an utterance can regulate social conduct while also describing a condition. A warning can direct attention and make a truth-apt claim about danger; similarly, an attribution might recommend a response while making defeasible claims about someone’s evidential responsiveness. Nor does the relational character of a status prevent its objective description. The absence of a single natural kind likewise need not eliminate descriptive constraints. The argument establishes that delusion attribution is not merely descriptive more securely than it establishes that it is non-descriptive. A hybrid account remains available.
Bafflement raises a related issue. The authors distinguish sincere bafflement from appropriate bafflement, recognising that inadequate empathy can lead us to overestimate incorrigibility (pp. 66–67). This qualification is essential, but it returns explanatory weight to judgements about the person and the interaction. Consider a clinician who recognises a familiar delusion without being baffled and responds by seeking a different basis for engagement. Such a case need not defeat a genealogy of ordinary usage. It does, however, suggest that identifying one important social function cannot exhaust the concept’s clinical and epistemic possibilities. The authors’ acknowledgment that not all uses fit their proposal makes this boundary especially consequential.
Second, the analysis of popular delusions leaves the distribution of irrationality insufficiently settled. Part 2 explicitly targets genuinely irrational belief systems, not merely beliefs outsiders dislike. Yet Chapter 5’s success in explaining how communities manufacture apparent warrant strengthens the possibility that some participants respond rationally to misleading evidence. The authors recognise this, describing insiders’ beliefs as less irrational than they appear externally (pp. 149–150). But the concession potentially goes further. Imagine a newcomer who receives convergent testimony from apparently credible sources, lacks access to the processes selecting that testimony and has no adequate reason to suspect manipulation. The community’s belief-forming arrangements may be epistemically corrupt without this participant’s acceptance being irrational.
An adequate account must therefore distinguish those who produce rationalisations, those who enforce conformity and those who inherit the resulting informational environment. These roles can overlap, but their epistemic defects need not. It must also distinguish a belief’s motivated acquisition from its subsequent maintenance on apparently good grounds. None of this vindicates popular delusions collectively. It does mean that identifying socially motivated processes somewhere in a network does not establish the irrationality of every believer downstream. The authors could restrict their target to participants who independently exhibit the stipulated epistemic defects. But identifying those defects would then be a separate empirical task, not something established by the social explanation itself. The book’s own social epistemology calls for a more differentiated reconciliation with the rational-social-learning explanations it contests.
Third, reading the account of attribution alongside the account of maintenance exposes a normative problem. The authors themselves identify the crucial feedback: labelling someone delusional can discourage testimonial interaction, deepen isolation and thereby help sustain the beliefs that occasioned the label (p. 243). They carefully qualify this claim; attribution is neither a sufficient nor an unmediated cause. Nevertheless, a practice whose function is to recommend disengagement may partly produce the conditions that make disengagement seem warranted. The apparent incorrigibility of a person’s belief cannot then be evaluated independently of how others have approached that person. This is not an objection to describing an exclusionary practice. It is a demand that the account’s own norms of appropriate attribution incorporate the practice’s causal effects.
The distinction required is between withholding assent and withdrawing epistemic recognition. Rejecting a proposition need not mean treating its speaker as incapable of contributing to understanding the situation. Likewise, declining repetitive contradiction need not amount to abandoning rational engagement. The book’s perspectival question—unpersuadable by whom and according to which norms?—should therefore extend to the quality, duration and institutional conditions of the relationship. Its postponement of sustained ethical analysis to future work (pp. 277–278) leaves unfinished something internal to the theory: how to distinguish warranted limits on argument from an exclusion that intensifies the problem.
The discussion of folie à deux supplies a complementary qualification. Here intense trust within a relationship can coexist with isolation from, or discounting of, outsiders (pp. 254–258). Thus neither increased connection nor increased trust is intrinsically corrective. What matters is their organisation: who can be heard, how credibility is allocated and whether a relationship permits revision. Read together, the book’s accounts of popular and clinical delusions suggest that epistemic health cannot be measured simply by conformity to a community. Communities themselves must remain answerable to evidence and challenge.
These reservations qualify rather than diminish the book’s importance. Williams, Wilkinson and Miyazono offer a powerful framework for relating processes too often considered separately, while making clear why neither pathology nor irrationality alone settles the explanatory task. Their strongest contribution is not a final account of what delusions are, but a demanding account of what explaining them requires. The unresolved tension between epistemic correction and epistemic exclusion is itself productive. A theory that understands knowledge as socially sustained must explain not only how communities protect themselves from error, but how they can do so without abandoning the people whose beliefs they find hardest to understand.



Nice idea, and unsurprising. Also good practice to see what models produce on topics where you're an expert. (I still have this very strong aversion to read LLM content not produced in my own chats!)