This post is part of a series on the new questions raised by AI.1 In this post, I look at how AI is set to change how we argue with each other.
The philosopher Leibniz, in his 1685 text The Art of Discovery, proposed a symbolic language that would allow disagreements to be resolved through calculation. His ambition was to make everyday arguments as clearly resolvable as mathematical questions. He described this idea with the words “Let us calculate”.
Unfortunately, Leibniz’s ambition was never realised. Arguments are messy, and many, if not most, end up without a clear outcome, with people moving only slightly away from their initial positions. In the nineteenth century, Schopenhauer proposed a radically different vision of debates: as jousts where winning is what matters and where tricks and bad-faith moves are all possible strategies to come out on top.
Modern psychology offers some support for Schopenhauer’s view. In their argumentative theory of reasoning, Mercier and Sperber argued that our reasoning abilities developed partly to persuade others and assess their arguments.2 More broadly, convincing others to be our friends, allies and partners may have exerted selective pressure on our reasoning abilities. This helps explain why we often reason like lawyers defending our own causes rather than scientists, with systematic motivated reasoning: selective use of evidence, logical flaws and cognitive biases that conveniently help us land on our preferred conclusions.
Social media are full of examples. In fact, even in everyday life, people sometimes admit that they continued arguing after they had realised they were probably wrong.3
In this post, I look at how AI could change the way our arguments work. I think it will have two effects. First, AI makes it much easier to verify arguments: factual assertions and interpretations of the evidence can be checked right away. Second, AI makes it much easier to craft good arguments: people can use AI agents as personal lawyers, helping them find supporting evidence, formulate their position and respond to counter-arguments. Taken together, I think these two effects are likely to improve the quality of the arguments we use in our discussions.
AI and third-party checks
An arbiter of truth available in any of our arguments
I was recently discussing international news with two colleagues, one American and one from Pakistan. Let’s call them John and Amir. At some point, the discussion drifted towards Pakistan’s domestic politics. Amir started explaining that the former Pakistani prime minister, Imran Khan, had been ousted by a US-organised military coup. I had never heard of this, and my default reaction when hearing about large geopolitical conspiracies I have never heard of is some level of scepticism. Not because I think conspiracies never happen—in the case of the US, from the overthrow of Iranian PM Mohammad Mossadegh to covert efforts against Salvador Allende in Chile, they clearly have—but because our psychology may be calibrated to overdetect conspiracies and, as a consequence, there are likely far fewer actual conspiracies than stories about alleged conspiracies.4
John was somewhat sceptical too. He quickly asked a reasonable question, likely to become fairly common in such discussions: “Is this role of the US clearly established? Is it something that ChatGPT would back up?” As he finished speaking, he was already taking out his phone and repeating Amir’s story in a fairly neutral way, with something like “Is this true?” at the end.
ChatGPT provided a lengthy answer which basically presented the fall of the PM as something triggered by domestic events, while pointing to evidence that the US had tried to influence the turn of events in favour of his removal.5 There was therefore more than a grain of truth in the claim of US meddling, though this did not establish that the US had organised his removal.
Amir and John seemed willing to accept that interpretation, at least provisionally, and moved on.
This type of exchange, likely to be repeated over and over all over the world, is quite novel. We now have an inexpensive third party that can be called into an argument within seconds and provide an assessment both informative and relatively independent of either side.
Hence, in an argument within a family, at work or between friends, one can increasingly turn to AI to settle factual issues that may be critical to the disagreement at hand. Without ChatGPT, John and Amir might have exited the discussion with beliefs markedly further apart than they did that evening, when it was used to adjudicate that factual issue in contention. This is, I think, one of the reasons why AI models may reduce political polarisation.6
Verifiability and information unravelling
There is another reason why third-party checks might improve the quality of arguments. By making many statements more easily verifiable, they change the strategic dynamic of discussions.
To the extent that arguments are agonistic, people may have an interest not to be entirely upfront about everything they know. They may retain pieces of evidence or information that are inconvenient for the stance they are defending. They may justify this selection to themselves by thinking that these elements are “not useful” and would hamper the process of persuading others towards the point of view they “know” to be right.
The use of AI models in arguments might change this. One of the most interesting results from the game theory of communication is that when information can be verified, disclosure can unravel: people progressively acquire an incentive to reveal exactly what they know.
Consider Alice, Bob and Candice arguing for different positions. Suppose, for simplicity, that an AI can assess the quality of the evidence supporting an argument using three categories: excellent, good and unclear.
Alice has evidence which she believes an independent AI assessment would rate as excellent.
Bob believes his evidence would be rated as good.
Candice believes her evidence would be rated as unclear.
In an “old-style” argument, all may have an interest to inflate the strength of their evidence. But suppose that anyone can ask for the evidence to be checked by an AI model. Alice can now credibly say that her position is supported by excellent evidence, since this claim can be checked.
Once Alice has made such a statement, silence becomes informative. If Bob and Candice do not also claim to have excellent evidence, listeners can infer that they do not. Bob now has an incentive to reveal that his evidence is good. By admitting that his evidence is not excellent, he can at least reassure the audience that it is not unclear. Once Bob has made this statement, Candice’s silence reveals the only remaining possibility: that her evidence is unclear.
At the end of this process, listeners have much better information about what the speakers themselves know about the quality of their arguments. The possibility of verification makes exaggeration harder, and silence itself becomes informative.
Notably, this effect would arise even when nobody actually calls the AI. Once people know that claims can easily be checked, they have stronger incentives to make claims that will survive scrutiny.
AI and argumentative contests
Verification is only one side of the transformation. AI models also give each side much better tools for advocacy.
AI as personal lawyers
Because arguments have an agonistic dimension. People often try to “win” when debating an issue. To do so, they need to find good arguments to make their case and counterarguments to repel the arguments of the other side. Such types of conflict are obvious in political arguments between people defending different positions, but a bit of introspection reveals that even in discussions in the household or the office, we are often engaged in this kind of verbal jousting.
Given this conflictual aspect, AI models can be used as wingmen to help us find better arguments and articulate them. With their tremendous research ability, AI models can act as expert lawyers, crafting precise and well-documented arguments. Each side can advise its AI like a client advises a lawyer, and then use it to help build the case to present to the audience.
A consequence may be an increase in the quality of arguments, both because AI agents can build better arguments quickly and because poor-quality arguments become easy targets for the opponent’s AI.
This kind of debate is already taking place over social media. There is little doubt that some people involved in lengthy back-and-forth arguments online use AI models to help craft their arguments and find counterarguments against points made by the other side. If anything, I suspect this is likely to be good for public debate, as it imposes greater discipline on the evidence and reasoning put forward on both sides.
In fact, as AI models continue to improve in speed, they could play a role in live debates. One problem with political interviews, for instance, is that politicians make factual statements across a range of domains for which it is practically impossible for journalists to have the expertise to assess immediately whether the statement is warranted. Often, it is not even the statement itself that is factually incorrect, but its implication or interpretation, which would require proper contextualisation for the audience to draw the right conclusion.
It is easy to imagine journalists soon having at their disposal an AI fact-checker listening live to assertions and offering corrections or suggesting clarifying questions. Again, the main effect might come before any correction is actually made: if speakers anticipate that dubious claims will immediately be challenged, they have less incentive to make them in the first place.
AI as judges of AI arguments?
A natural question is who would judge such arguments as they become better and possibly harder for a lay audience to follow. In practice, an AI-enhanced debate might become to a debate what a Matrix-like fight scene would be to a boxing match.
It might then become hard for an audience to follow the arguments and judge them. This leads to a more speculative possibility: AI models themselves could be used as third-party adjudicators, deciding which side has the better arguments. Debaters could agree on a codified process whereby they submit their arguments to an independent AI application tasked with providing an assessment of their respective merits.
I suspect that such formal resolution will never become common in ordinary social arguments. This is, in part, due to a strategic reason. Someone whose argumentative position is weak may be able to discover this beforehand by asking their own AI. If they expect to lose badly under an agreed adjudication procedure, they may prefer not to enter the debate in the first place. Potential debates might therefore be settled before they reach the AI courtroom.
Science as a limiting case
There is, however, one domain where AI seems to have the potential to take on such an adjudicating role: science. As scientists increasingly use AI models to help craft technical contributions, editors and reviewers will increasingly use AI models to assess research. AI may, in that sense, become an arbiter, even without being formally delegated the decision. It can influence human judgement by providing highly technical assessments that would otherwise be costly to obtain.
There are potentially very good consequences. Research output may improve in quality and explanations may become clearer. Artificially technical language will provide less protection for weak arguments if an AI can rapidly unpack the jargon and assess the underlying reasoning for an inquisitive reader.
But science also points to the potential challenges associated with an increasing role of AI in technical arguments. Suppose AI eventually becomes better than human scientists at evaluating some forms of research. Imagine that a researcher, working with an AI model, proposes a revolutionary theory unifying quantum theory and general relativity. Suppose also that the theory is sufficiently complex that no human can fully assess it, while other AI models judge it to be correct. Should humans accept their verdict?
Hopefully, through repeated queries and investigations, humans would eventually be able to make sense of the theory. But there is no guarantee that the human brain will always be able to grasp every aspect of the universe that a superior intelligence might understand. As Brian Greene has pointed out, his dog will never understand the laws of the universe.7 Likewise, there is no guarantee that the human brain, which evolved to solve very earthly problems, is capable of grasping the deepest laws of the world.
At that point, the role of AI as an arbiter of arguments would raise a very different question: what does it mean to rely on an argument whose validity we can no longer independently establish? Things have changed so fast here that we have more questions about the future than clear answers.
Human debates are messy and filled with errors in reasoning and tricks to “win”, even when the ideas defended would not withstand careful scrutiny. Recognising this, Schopenhauer wrote a cynical guide to winning arguments at almost any cost.
Many of these limitations stem from our cognitive constraints. It is costly to cross-check every point in an assertion: its factual accuracy, its suggested implications and interpretation, possible contradictions, and whether relevant evidence has been omitted.
By making it easier to craft high-quality arguments and verify their quality, AI is bound to change how we argue, possibly reducing many of the biases and poor practices that characterise human argumentation.
Leibniz’s ideal of a perfect language that would determine what is right and wrong is likely outside the realm of what is feasible. But the tremendous power of AI may increase scrutiny of arguments to the point where we move somewhat closer to Leibniz’s dream and away from Schopenhauer’s cynical reality: many bad arguments that are common today may simply become much harder to get away with.
References
Fetterman, A.K., Curtis, S., Carre, J. and Sassenberg, K. (2019), ‘On the willingness to admit wrongness: Validation of a new measure and an exploration of its correlates’, Personality and Individual Differences, 138, pp. 193–202.
Hruschka, T.M.J. and Appel, M. (2026) ‘Reducing political polarization through conversations with artificial intelligence’, Journal of Computer-Mediated Communication, 31(2), zmag003.
Leibniz, G.W. (1685/1951) ‘The Art of Discovery’, in Wiener, P.P. (ed.) Leibniz: Selections. New York: Charles Scribner’s Sons.
Mercier, H. and Sperber, D. (2011) ‘Why do humans reason? Arguments for an argumentative theory’, Behavioral and Brain Sciences, 34(2), pp. 57–74.
Schopenhauer, A. (1831) Eristische Dialektik: Die Kunst, Recht zu behalten [Eristic Dialectic: The Art of Being Right].
van Prooijen, J.-W. and van Vugt, M. (2018), “Conspiracy Theories: Evolved Functions and Psychological Mechanisms,” Perspectives on Psychological Science, 13(6), 770–788.
My previous posts discussed AI risks, the moral status of AI agents, and the impact of AI on peer-reviewed research, the impact of AI on the ability of laypeople to assess experts’ takes and the possible impact of AI on the higher education sector.
Mercier and Sperber (2011).
See Fetterman et al. (2019).
See van Prooijen and van Vugt (2018).
Khan had lost coalition support after falling out with the military establishment and was ultimately removed through a parliamentary vote of no confidence. During that process, a leaked diplomatic cable suggested that Washington favoured that outcome and had communicated this preference to Pakistani officials. According to the cable, US Assistant Secretary of State Donald Lu said that “if the no-confidence vote against the Prime Minister succeeds, all will be forgiven in Washington” and that otherwise “it will be tough going ahead”.
Hruschka and Appel (2026)
Part of The Elegant Universe
No matter how hard you try, you can’t teach physics to a dog. Their brain is not wired to grasp it. But what about us? — Greene









I think it will allow us to significantly improve the quality of discussion and debate. An example is the issue of women’s wages relative to men’s. A simple AI query now explains the issue as both real and easily explained primarily by differences in choices, goals, values and careers.
A few other points.
I see no reason for someone to honestly and fairly only ask their AI to argue one side of the debate. Both sides should ask their AI to lay out the pros and cons, strengths and weaknesses before the discussion escalates. Doing so puts the respectable burden on the debaters to first convince their own AI before trying to persuade anyone else.
Second, I find that when I push AI to lay out a position it believes are best, that it lays out extremely sound and fairly well balanced positions on every conceivable topic from abortion, to homelessness, to law enforcement Iran, to health care. Does everyone else not ask their AI these things, or debate them when they disagree (for their own sake)? Assuming they do, are we all getting similar answers?
Finally, I am gobsmacked that nobody has yet taken the initiative to set up substack dedicated to debates using AI. This could be debates between models Claude vs Chat, or between people and models. I honestly think this would be the most interesting site possible. If I had the time, that is what I would build.