A person stands before a court while a file is opened. Some pages contain facts: the charge, the evidence, the person’s record, the circumstances allowed by law. Then another page adds a coloured band or a number said to represent the risk of future offending. Nothing on that page records something the person has done. It is an estimate made by comparing this person with other people.

That difference is the boundary. Statistical analysis can help institutions study populations, evaluate programmes and detect where systems fail. It can sometimes help professionals organise relevant information. But an individual forecast of offending or dangerousness should not anchor guilt, sentence, detention or any other decision about liberty. A model can describe patterns in a group; it cannot turn a possible future into an individual fact.

A defendant sits at a courtroom table as abstract lines from a crowd of silhouettes converge behind him.
A statistical pattern drawn from groups cannot become a fact about one person.

What a risk score actually predicts

A criminal-risk model begins by defining an outcome. That may be rearrest, reconviction, breach of supervision or another recorded event within a chosen period. Those outcomes are not interchangeable. Arrest depends partly on police presence and enforcement practice. Conviction depends on evidence, procedure, legal representation and prosecutorial choices. A breach of supervision may reflect a missed appointment rather than a new offence. Calling all of these outcomes “recidivism” can make an administrative label sound more objective than it is.

The model then uses selected variables to find associations with that outcome in historical data. Depending on the system, these may include age, previous convictions, employment or housing information, substance-use history, answers to questionnaires, or features derived from administrative records. A statistical or machine-learning procedure estimates how combinations of those variables relate to the recorded outcome. The result may be a probability, a rank or a category such as low, medium or high.

Every stage contains a decision. Someone chooses the outcome, the observation period, the population, the variables, the treatment of missing data and the threshold between categories. A threshold is especially consequential. It converts a continuous estimate into an instruction-like label. Moving it changes how many people are classified as high risk and therefore changes the balance between false positives and false negatives. The threshold is a policy choice about which error to tolerate, not a fact discovered by the model.

Even the best-calibrated score remains a statement about groups. In its 2016 State v. Loomis judgment, the Wisconsin Supreme Court described the COMPAS score before it as comparing information about an individual with a similar data group. The court noted that the assessment did not predict the specific likelihood that the individual defendant would reoffend. That distinction is easily lost once a group-derived score is displayed beside a person’s name.

The base-rate trap

A simple hypothetical example shows why apparently respectable performance can mislead. Imagine 1,000 people in a setting where 100 will meet the chosen recidivism definition during the observation period. Suppose a tool correctly flags 80 of those 100 and correctly clears 80 per cent of the 900 who will not meet the definition. It therefore finds 80 true positives, but it also produces 180 false positives. Among the 260 people labelled positive, fewer than one in three will in fact meet the outcome.

The arithmetic is not an unusual defect. It follows from the base rate. When the event being predicted is relatively uncommon, false positives can outnumber true positives even with apparently good sensitivity and specificity. The person carrying the false-positive label experiences the consequence as an individual—perhaps closer supervision, a harsher bail condition or a loss of liberty—while the institution reports aggregate accuracy.

Calibration does not solve this problem. A score is calibrated if, among people assigned a stated probability, roughly that proportion experiences the defined outcome. If 30 out of every 100 people receiving a 30 per cent score meet the outcome, the score can be calibrated. It still cannot say which 30. For the remaining 70, the group rate was never an individual destiny. The uncertainty is not a small technical footnote; it is the central fact that a coercive decision must confront.

Fairness cannot be selected from a settings menu

Risk tools are often defended or criticised through a single fairness measure, but different measures ask different questions. One may require people with the same score to have similar outcome rates across groups. Another may require false-positive rates to be equal. A third may require false-negative rates to be equal. These conditions can conflict when the recorded outcome rates differ between groups.

Alexandra Chouldechova’s peer-reviewed analysis of recidivism instruments showed that commonly invoked fairness criteria cannot all be satisfied simultaneously when prevalence differs between groups, except in restricted circumstances. The point is not that fairness is impossible. It is that a technical team cannot silently optimise one definition and claim to have solved the institutional question. The choice determines who bears which errors and must therefore be exposed to legal and democratic scrutiny. The study also explains how disparate impact can arise even when a score satisfies one recognised fairness condition.

Historical data deepen the difficulty. Police records do not contain a neutral census of all offending. They contain events that institutions detected, recorded and processed. Areas subjected to more enforcement generate more observations. People with fewer resources may be less able to avoid administrative breaches or secure alternatives to detention. A model trained on those records can learn the footprint of institutional attention and return it as a forecast about individuals.

If the score then influences where police, supervision or services concentrate, the system creates a feedback loop. More attention produces more recorded events, those records strengthen the apparent association, and the next model treats the result as confirmation. Auditability therefore requires more than inspecting code. It requires tracing how labels were produced, which populations were absent, how institutional behaviour shaped the target, and whether deployment changes the data later used to validate the system.

The strongest case for structured assessment

The serious counterargument is not that algorithms are infallible. It is that unaided human judgment is also inconsistent, opaque and vulnerable to bias. A tired decision-maker may overweight a vivid fact, treat similar cases differently or rely on intuition that is never tested. Structured instruments can force relevant questions to be asked consistently. They can identify needs, support rehabilitation planning and help a system examine whether comparable cases receive comparable treatment.

That case deserves weight. The choice is not between a flawed model and a perfectly wise human. Research by Julia Dressel and Hany Farid found, in one well-known experiment using the COMPAS data, that the commercial system was not more accurate or fair than predictions made by people with little criminal-justice expertise, and that a simple model using two features performed similarly. The published study does not establish that every risk instrument is useless. It shows why complexity and proprietary sophistication cannot substitute for demonstrated value.

Structured assessment is most defensible when it changes support rather than punishment. Information about housing instability, treatment needs or practical barriers may help professionals offer services, provided people can correct the data and refusing assistance does not become evidence of dangerousness. Aggregate analysis may help compare programmes, estimate staffing needs or discover that a policy produces unequal burdens. Those uses ask what an institution should improve. They do not claim to know what a named person will do.

The danger arises when an instrument built for one purpose migrates into another. In Loomis, COMPAS had been designed to support correctional placement, management and treatment planning, yet its risk assessment appeared in a sentencing file. The Wisconsin Supreme Court allowed consideration under specified limitations, while stressing that risk scores must not determine whether someone is incarcerated, the severity of the sentence or, as the determinative factor, whether community supervision is possible. It also required written cautions when the assessment appeared in a presentence report. That operational history demonstrates how a nominally supplementary tool can move towards the centre of a liberty decision.

The European legal boundary

The EU AI Act draws an important but carefully qualified line. Article 5(1)(d) prohibits placing on the market, putting into service or using an AI system to assess or predict a natural person’s risk of committing a criminal offence when the assessment is based solely on profiling or on personality traits and characteristics. The exception concerns systems supporting a human assessment of a person’s involvement in criminal activity where that assessment is already based on objective and verifiable facts directly linked to criminal activity. The prohibition has applied since 2 February 2025.

The Act does not treat every other criminal-risk application as acceptable. Its Annex III classifies permitted law-enforcement systems for assessing offending or reoffending risk in specified circumstances as high-risk. The regulation’s text connects this classification to the power imbalance of law enforcement and to possible effects on liberty, defence rights, effective remedy and the presumption of innocence. As of 22 August 2026, the Commission states that enforcement powers began applying on 2 August 2026, while the detailed rules for Annex III high-risk systems are scheduled to apply from 2 December 2027 following the AI Omnibus changes. The current dates and scope are set out in the Commission’s enforcement guidance.

The AI Act sits alongside data-protection law. Article 11 of the EU Law Enforcement Directive requires Member States to prohibit a decision based solely on automated processing, including profiling, when it produces an adverse legal effect or significantly affects a person, unless Union or national law authorises it and provides appropriate safeguards, at least the right to obtain human intervention. The directive also requires, as far as possible, a distinction between data based on facts and data based on personal assessments.

These rules matter, but “human intervention” can become ceremonial. If the score is presented first, the reviewer is under time pressure, the underlying data cannot be challenged and disagreement creates extra work, the human may merely ratify the machine. A lawful process therefore needs more than a person positioned at the end of a workflow. It needs someone with the evidence, competence, time and authority to disregard the recommendation and give reasons that stand without it.

Auditability begins before the model

An audit should be able to reconstruct the entire chain from legal purpose to human consequence. The institution must identify who authorised the use, which decision it may influence, what outcome the model predicts, how the training population differs from the current population, which variables are used, how missing or disputed data are handled, what threshold is applied, and how performance changes across relevant groups and over time.

It must also record how people encounter the system. Can the defendant know that a score was used? Can counsel inspect the input data and the validation evidence? Is the model version preserved for appeal? Can an independent expert reproduce the result? Does the written decision explain the evidence and law, or merely translate the score into judicial language? A proprietary product that prevents meaningful challenge creates a problem even if its average performance is respectable.

The Council of Europe’s CEPEJ Charter frames judicial AI around fundamental rights, non-discrimination, quality and security, transparency, impartiality and user control. Its later Resource Centre also distinguishes operational systems from pilots and conceptual projects rather than treating every announced tool as an established deployment. That distinction is essential in criminal justice, where a vendor demonstration cannot establish reliability in a different population, under different procedures and with liberty at stake. The Charter’s principles place institutional control, not model novelty, at the centre.

What technology should do instead

Technology can improve criminal proceedings without forecasting the defendant. It can organise authenticated records, build a traceable chronology, identify missing documents, compare submissions, retrieve current authorities and show where factual claims conflict. It can make disclosure searchable, preserve version history and help counsel find the evidence needed to challenge the state’s case. These functions reduce cognitive and administrative burden while leaving the legal relevance and weight of each fact open to argument.

At system level, statistical analysis can reveal delays, unequal access to diversion, patterns in pretrial detention, programme outcomes and differences in correction rates. Such work still requires careful definitions and safeguards, but its object is the institution rather than a forecast attached to a person. The question changes from “What will this defendant do?” to “What is our system doing, for whom, and with what evidence?”

Where individual decision support is used, the interface should lead with verified facts, provenance, missing information and lawful options—not a recommendation score. A judge or other authorised decision-maker should construct reasons from the record and accept responsibility for the outcome. If removing the forecast would make the reasons collapse, the forecast was not merely supportive.

The decision that remains human

A court must decide what has been proved, what the law permits and why a coercive measure is necessary and proportionate in this case. Those are not predictions to be optimised. They are exercises of public authority owed to a person who must be able to hear, understand and challenge the reasons.

If you want to follow Alkemata’s continuing work on human-centred justice, the latest articles are collected on the site as the series develops.

Continue exploring

The practical difference between nominal and meaningful oversight is examined in When Human Review Becomes a Rubber Stamp. The same institutional test applies here: does the human reviewer possess the time, evidence and authority needed to disagree?

The unresolved decision is therefore not which model should predict the defendant most accurately. It is whether a justice institution will keep group statistics out of individual coercive judgment and invest instead in verified records, contestable evidence, transparent reasons and the human capacity to decide without borrowing certainty from a score.

Author