AI and Judging: Confronting Misconceptions

Discussing AI and decision aids with lawyers can be frustrating. Limited technological expertise is paired with an often weirdly romanticized picture of human decision making. You’re told AI systems amplify biases. You get the inevitable reference to COMPAS. The rule of law. Yes, human judges make mistakes. But it doesn’t matter. Because they are human mistakes. Fairness and biases are constantly bandied out but assumed rather than defined. Often the discussion seems to discard, wholesale, the use of questionnaires and structured risk assessment in judicial fora. Implying that we should return to clinical evaluations (the world of not ten, but fifty years ago). Or not — put the risk assessment questionnaire on a sheet of paper, not on a screen, and of all a sudden it is perfectly fine. I have even seen instances where old, controversial authors like Jerome Frank (famous for the “what the judge ate for breakfast” quote) get bizarrely praised for their celebration of arbitrary decision making — what seems to be great then is that humans, unlike machines, can decide like cases not alike, but in all sorts of surprising ways. That this itself flies in the face of longstanding normative conceptions of judging and more than two centuries of legal and political scholarship, since Montesquieu and Beccarria, seemingly, need not bother anyone.

Sure, a flawed argument doesn’t mean its conclusion is necessarily wrong. Plus, I am also a lawyer (of sort). Yet the repeated frustration supplied essential motivation to come up with a more rigorous assessment in the form of a synthetic review of the empirical literature: Man and Machine: Artificial Intelligence and Judicial Decision Making (coauthored with Ahmad Shahvaroughi; manuscript available as pre-print on ArXiv: https://arxiv.org/abs/2603.19042, and SSRN: https://papers.ssrn.com/sol3/papers.cfm?abstract_id=7548322). In the revised manuscript, we begin by introducing an analytical framework to make the discussion more precise. We then apply it to survey existing empirical work on three questions: How accurately do algorithms predict outcomes? How do their predictions compare with human judgment? And what happens when judges receive algorithmic advice? What comes out of the review is more complex and more nuanced than much of what we hear in legal circles. Which goes for the strengths and weakness of real human judges, fears of automation bias or algorithms and racial discriminations.

Leave a Comment

Filed under Arthur Dyevre

Leave a Reply

Your email address will not be published. Required fields are marked *