A caseworker opens the next file. The system recommends refusal. A coloured warning is already visible, the supporting paragraph is pre-written and dozens of cases remain in the tray. The worker can technically disagree, but doing so means opening several records, writing a reason and asking a supervisor. Approval takes one click.
This is called human review, yet the decision has already acquired direction, language and institutional momentum. Oversight is meaningful only when the reviewer has enough time, evidence and competence to make an independent assessment—and enough authority to change the outcome. Without those conditions, the human is not a safeguard but a signature surface.

The recommendation arrives first
Decision-support systems do not merely add information. They shape the order in which information is encountered. When a recommendation appears before the reviewer has formed a view, it becomes an anchor. The later evidence is examined from a starting point supplied by the system: why is this recommendation right or wrong? That is a narrower question than: what decision follows from this case?
Automation bias adds another pressure. People can treat an automated output as a reliable shortcut, especially when the task is repetitive or verification is costly. A 2025 systematic review of 35 peer-reviewed studies found that over-reliance is shaped by interacting factors including expertise, AI literacy, task demands and explanation complexity. It also cautioned that explanations alone do not necessarily improve accuracy: a persuasive explanation may increase acceptance without producing genuine verification.
This is why adding a “reason” beside a recommendation is not sufficient. The explanation may itself frame the case, omit alternative interpretations or give confidence without exposing the underlying evidence. A useful interface should let the worker reach the source record, see uncertainty and identify what the model did not consider.
Workload decides whether review happens
Even a well-designed screen cannot create time. If an agency uses automation to increase the number of files assigned to each reviewer, the nominal review stage can survive while the practical one disappears. Staff learn that opening the evidence behind every recommendation is incompatible with throughput targets. The safest action for the worker becomes accepting the default.
Alert fatigue is the corresponding problem for exceptions. When a system raises many low-value warnings, staff must repeatedly spend attention discovering that nothing requires action. The signal gradually becomes part of the background. A rare consequential warning then arrives through the same channel and can be dismissed by habit. More alerts do not necessarily create more oversight; they can consume the attention on which oversight depends.
The UK Information Commissioner’s Office makes the organisational requirement explicit in its AI audit framework for human review. It says reviewers should have appropriate knowledge, experience, authority and independence, as well as manageable caseloads and sufficient time. It also expects documented testing, override logs and a fallback route when the system’s performance falls below an acceptable level.
Authority is more than an override control
A reviewer may be allowed to click “override” while lacking the power to make disagreement stick. The exception may require managerial approval, harm performance statistics or trigger an investigation into the worker rather than the model. In that environment, the formal control exists but the organisation has made it costly to use.
Meaningful authority includes the ability to pause a case, request missing evidence, consult a specialist, change the decision and escalate a suspected systemic problem. It also requires a route to suspend automated processing when reviewers find a pattern of error. Otherwise each worker repeatedly repairs individual cases while the same fault continues upstream.
The ICO’s wider guidance on individual rights in AI systems says human intervention cannot be a token gesture: the reviewer must have the authority and capability to change the decision and must assess all relevant data, including information supplied by the affected person. That turns review from internal quality control into a point where a person’s evidence can alter the administrative outcome.
Design for an independent view
Where the risk justifies it, the reviewer should assess the material issues before seeing the model’s conclusion. The system can first present verified facts, sources, missing information and the applicable rule. After the worker records a provisional view, the recommendation can appear as a comparison or challenge. This makes disagreement visible and reduces the pull of a default answer.
Not every routine case needs this sequence. A risk-based workflow can reserve independent assessment for consequential, ambiguous, unusual or randomly sampled cases. But the organisation should decide the boundary explicitly. It should not quietly classify almost everything as routine because independent review is expensive.
For high-risk AI systems, the consolidated EU Artificial Intelligence Act, current to 27 July 2026, requires interfaces and oversight measures that enable designated people to understand relevant limitations, detect anomalies, remain aware of automation bias, interpret outputs and disregard or reverse them. That is a useful test of the whole workflow. If staff cannot realistically exercise those capacities during an ordinary shift, oversight is ineffective regardless of what the policy manual promises.
The consistency counterargument
The strongest case for recommendation-first review is consistency. Public bodies want similar cases treated similarly, and independent human judgments can vary. Requiring staff to reconstruct every routine decision may duplicate work, slow services and reintroduce personal bias. A standard recommendation, followed by review, appears to combine machine consistency with human discretion.
That combination can work, but consistency should not mean agreement with the machine. It should mean consistent attention to the same legally relevant factors, evidence standards and reasons. Technology can structure those elements without displaying a conclusion first. It can also identify comparable cases and contradictions while preserving the worker’s responsibility to decide.
The UK government’s AI Playbook states the principle as meaningful human control at the right stage. The timing matters. Human involvement after a recommendation has become the default is not equivalent to human judgment supported by organised evidence.
Measure disagreement without turning it into a quota
An agency cannot evaluate oversight by reporting that every case was seen by a person. It should examine how often reviewers seek underlying evidence, request additional information, disagree, escalate and reverse decisions. It should measure the time available for review, the error rate in independent samples, complaint and appeal outcomes, and whether disagreement is concentrated among particular teams or types of case.
Neither a high nor a low disagreement rate is inherently good. High disagreement may reveal a poor model, or reviewers who misunderstand a good one. Near-perfect agreement may reflect excellent performance, or a workforce that lacks time and authority. The useful question is whether disagreements are justified when checked against reliable evidence and whether the organisation learns from them.
Independent sampling is crucial because accepted recommendations otherwise escape scrutiny. A separate team can examine a stratified sample without seeing the original recommendation, including ordinary approvals and refusals as well as overrides. This estimates errors that frontline agreement statistics cannot reveal. The NIST AI Risk Management Framework playbook recommends documenting human oversight, tracking overrides, complaints and response times, and using independent assessment and affected-community feedback to evaluate system performance.
If you want to follow this series on the institutional design of human-centred public technology, you can subscribe to Alkemata for the next article.
The remaining decision is organisational, not cosmetic. Will the agency reserve enough capacity and power for reviewers to test the system, or will every efficiency gain be converted into a larger caseload? Human review becomes a safeguard only when disagreement is practical, consequential and studied. Otherwise the person at the end of the process merely lends human legitimacy to a decision already made.