A caseworker opens a file before speaking to the person across the desk. One record shows a recent address. Another still carries the old one. An uploaded payslip is current, but a medical document is missing and a deadline is close. The software could compress all this into a red risk score. Or it could show what is known, where it came from, what conflicts and which lawful routes remain open.

The second design is more useful and more honest. Decision support in public services should give caseworkers a map, not a score: a structured view of evidence, provenance, gaps, uncertainty and lawful options. Technology can make the case easier to understand. It should not turn a person into a recommendation that staff are expected to accept.

A caseworker and a citizen sit across a desk beside branching pathways connecting housing, health, calendar and document symbols, while a closed machine stands unused
A useful decision-support system exposes the evidence and possible routes while leaving the caseworker responsible for the decision.

What disappears inside a score

A score looks precise because it has one dimension. Higher may mean more urgent, more likely to qualify, more likely to default or more likely to require investigation. But the case from which it was calculated has many dimensions. Some data are verified; some are self-reported. Some are recent; some may be stale. Absence can mean that a condition is not met, that a document has not arrived, or that the administration never asked the right question.

Compression is sometimes necessary. Public services have queues, finite staff and rules that require comparable treatment. The problem begins when the compressed output becomes the main object that the worker sees. A recommendation displayed first supplies an initial frame for everything that follows. Evidence can then be read as support for the score rather than as material from which an independent decision should be built.

This is not solved by writing “human review required” under the number. The human needs an interface that makes review possible. The current consolidated text of the EU Artificial Intelligence Act requires effective human oversight for high-risk AI systems. Article 14 specifically addresses automation bias and says overseers must be able to understand capabilities and limitations, interpret outputs, disregard or reverse them and stop the system. Those capacities depend on what the interface reveals and what the organisation permits, not merely on the presence of a person at the end.

A map preserves the route to each fact

A case map begins with evidence rather than a predicted outcome. Each material fact should carry its source, date and status. The caseworker should be able to distinguish a current register entry from an applicant’s statement, a confirmed document from an inferred relationship, and a contradiction from a simple absence. Selecting a fact should lead back to the underlying record rather than to a model-generated explanation of it.

The map should also preserve missingness. If eligibility depends on residence during a particular period, the interface should show which dates are covered and which are not. It should not silently treat an empty field as “no”. If two official registers disagree, both entries and their provenance should remain visible until the discrepancy is resolved. Technology is valuable here because it can assemble records, detect conflicts, extract dates and keep a chronology current faster than a person working across several systems.

Uncertainty should be attached to the claim it qualifies. A document classifier may be unsure what kind of form was uploaded. An extraction tool may be unsure whether a date refers to issue, treatment or expiry. A predictive model may perform differently for cases unlike its training data. One overall confidence score hides these distinct problems. Local uncertainty tells the caseworker what to verify.

Finally, the map should connect facts to lawful options without selecting one as the natural answer. It can show that one route requires additional evidence, another requires a supervisor’s authority, and a third triggers a statutory notice or human specialist. The decision remains a reasoned act by the caseworker, with the applicable rule and supporting facts recorded.

Consistency is the serious counterargument

The strongest case for a score is not convenience. It is equal treatment. Two workers may read the same ambiguous file differently. Workload, experience and unconscious bias can influence who receives extra time, scrutiny or help. A common model appears to offer a disciplined reference point, and large differences from its recommendation can reveal uneven practice.

That concern is real. Simply returning to unstructured discretion would protect neither citizens nor staff. But consistency should be built into the evidence and reasoning process rather than imposed through a concealed conclusion. The system can require the same eligibility questions, apply explicit rules to verified facts, flag the same missing items, show comparable precedents where legally appropriate and record why an exception or escalation was used.

This makes disagreement informative. If workers repeatedly choose a different lawful option when a particular kind of evidence is present, the organisation can examine whether the policy, training, interface or model is wrong. If only agreement with a score is measured, justified departures look like errors and staff learn to comply. A supposedly advisory system then becomes an unofficial policy without legislative authority.

The institution must make independence practical

A good interface cannot compensate for impossible workloads. A caseworker who has three minutes, cannot obtain missing evidence and must explain every override to a manager is not exercising meaningful discretion. Nor is a worker independent if the system’s output affects performance ratings or if only agreement is quick enough to meet targets.

The UK Algorithmic Transparency Recording Standard asks public bodies to describe how an algorithmic tool is integrated into decision-making, what information it gives the decision-maker, how human review works, what training is required and what appeal routes exist. This is a useful shift from asking whether a human appears somewhere in the process to asking what that person actually sees and can do.

Public transparency matters because the caseworker is not the only human involved. The person affected needs an intelligible explanation of the decision, the important evidence used, the source of that evidence and the route for correction or appeal. A map designed only for internal efficiency still fails if the citizen receives a conclusion they cannot contest. The same provenance that helps staff investigate a discrepancy should help the person point to the wrong record.

The standard also recognises that accountability belongs to an organisation, not to a lone frontline employee. Its guidance, updated in May 2025, makes records mandatory for certain UK central-government bodies when algorithmic tools significantly influence decisions with public effect or interact directly with the public. A named senior responsible role, documented deployment context and published limits make it harder to blame an individual worker for a system that management procured and configured.

Measure whether the map improves judgment

Accuracy alone is too blunt a test. A service should measure whether the system helps workers find missing evidence, detects stale or contradictory records, shortens the time to a correct decision and improves the quality of reasons. It should examine correction and appeal outcomes, differences across groups, cases escalated, justified overrides and the time available for review.

The NIST AI Risk Management Framework playbook recommends maintaining statistics on overrides, reported errors, complaints, response times and adjudication, while testing explanations for decision-makers and people affected by decisions. Those measures treat human oversight as an observable capability. A low override rate is not automatically success; it may indicate an excellent tool, an anchored workforce or a workplace in which disagreement is punished.

There is also a simpler test. Remove the model’s recommendation and ask whether the interface still helps the worker reach a better supported decision. If the answer is yes because records are assembled, provenance is visible, gaps are explicit and lawful routes are clear, the technology has created durable capability. If the value disappears with the score, the system may be replacing judgment rather than supporting it.

This distinction fits the broader evidence. The OECD’s 2025 review of AI in public-service design and delivery finds substantial scope for automating recording, preparation, sorting, classification and verification so that public servants can spend more time on work requiring judgment and discretion. That benefit will materialise only if saved time is actually returned to the human part of the service.

If you want to follow this series on human-centred legal technology and government, you can subscribe to Alkemata for the next article.

Continue exploring

The same boundary between useful coordination and hidden decision-making appears in When Government Knocks First, which examines how event-driven public services can act earlier without removing citizen control.

The decision now belongs to service owners and procurement teams: will they buy a system that optimises agreement with its recommendation, or one that makes evidence easier to inspect and human reasons easier to defend? The unresolved evidence need is equally concrete. Before scaling, agencies should show that caseworkers with the map catch more consequential errors and produce better reasons under real workloads—not merely that they reach the model’s answer faster.

By rdi

I am the vice-boss here; in charge of online activities and the technical stuff. I have a background as engineer and scientist in fields as different as aerospace, plasma physics, biosensing, I am currently here to find people motivated to build stuff together and to share adventures together

Leave a Reply

Your email address will not be published. Required fields are marked *