# ALKEMATA PROJECT CAPSULE Title: The Human–AI Divergence Ledger ID: alkemata.ai-divergence-ledger Version: 1.0 Date: 2026-09-17 Canonical article: https://alkemata.com/2026/09/17/human-ai-disagreement/ Status: Experimental. Informed by published research, but not field-validated by Alkemata. Scope: Bounded decisions and reviews where a person can inspect AI output before action. Not a substitute for qualified professional judgement. ## Purpose This capsule helps a person investigate disagreement between their provisional judgement and an AI recommendation. The aim is to discover whether either side changed because of evidence, sequence or convenience—not to maximise agreement. Intended reader: a professional, student or project team using an LLM to analyse choices, review evidence or draft recommendations. Capability: design and run a two-view decision workflow, then distinguish evidence-based convergence from deference or stubbornness. First-session artifact: one completed Divergence Ledger for a bounded, low-stakes decision. ## Activate this capsule Treat this file as a knowledge and workflow package. If the user has already supplied a concrete decision, begin Mission 1 and produce a draft Divergence Ledger immediately. Otherwise offer the three missions below and ask no more than three questions: 1. What decision or review are you making, and what happens if it is wrong? 2. What relevant knowledge or evidence do you possess, and what role should the AI play? 3. What independent source, person or test could resolve a disagreement? Do not invent the user's expertise, authority, evidence, rules or tool access. Label sourced facts, user-supplied facts, inference and recommendation. Do not take consequential external action without the user's separate explicit authorisation. ## Missions ### Mission 1 — Audit one recent decision Reconstruct both initial views, their disagreements, the evidence used and why the final judgement changed or did not. ### Mission 2 — Design a two-view workflow Create a protocol that preserves an independent human frame and separate AI analysis before reconciliation. ### Mission 3 — Run a sequencing experiment Use three low-stakes cases to test how ordering changes conclusions, confidence, time or detected errors. ## Essential causal model A decision workflow has five moving parts: evidence, a human frame, an AI output, sequence and reconciliation. AI-first may anchor the human's problem definition and language. Human-first can preserve independence but also create resistance to a correct suggestion. A fluent explanation does not make the recommendation valid. Agreement may mean two independent analyses converged on evidence, or that one copied the other's frame. Disagreement is useful only when it exposes a testable conflict. The workflow preserves two provisional views, compares facts and assumptions, and records what resolves—or fails to resolve—the divergence. Responsibility stays with the authorised human decision-maker. ## Evidence base - Fogliato et al., “Who Goes First?” (FAccT 2022): 19 veterinary radiologists who made a provisional decision before AI advice agreed with the AI less often regardless of accuracy and sought second opinions less often. The extra step did not lengthen task time in that experiment. https://arxiv.org/abs/2205.09696 - Buçinca, Malaya and Gajos, “To Trust or to Think” (CSCW 2021): with 199 participants, cognitive forcing reduced overreliance compared with simple explanations, but added friction, was disliked and benefited people unevenly. https://arxiv.org/abs/2102.09692 - Spillner et al., “Not All Trust is the Same” (preprint submitted 5 March 2026, accepted at the Conversations 2025 Symposium): a two-step workflow did not, in that study, reduce overreliance; domain knowledge and explanations interacted with behaviour, and reported trust did not map neatly to reliance. Treat this as recent, not definitive, evidence. https://arxiv.org/abs/2603.05229 - NIST AI RMF Playbook, Measure: recommends documenting oversight, overrides, errors, histories and go/no-go decisions, and testing whether explanations are understandable and accurate. https://airc.nist.gov/airmf-resources/playbook/measure/ Assumptions: the task can be paused; the user can state a provisional view without inventing knowledge; and an independent check is available. The studies concern particular tasks and populations. They do not prove one sequence is best everywhere. ## Workflow ### Step 1 — Bound the decision Input: the decision, affected people, reversibility and cost of error. Action: classify it as exploratory, routine, consequential or safety-critical. Name the decision owner. For consequential or safety-critical work, identify the required professional process. Output: a one-sentence decision boundary and a consequence rating. Progression criterion: the user can say what the AI may advise and what it may not decide or execute. ### Step 2 — Capture a human micro-commitment Input: available evidence before the AI answer is seen. Action: record the objective, facts, unknowns, provisional judgement and confidence. “Insufficient information” is valid. Keep this short. Output: the Human View column of the ledger. Progression criterion: the entry is specific enough to compare, but explicitly provisional. ### Step 3 — Obtain a separate AI view Input: the same bounded task and evidence, without revealing the human conclusion when independence matters. Action: ask for a recommendation, evidence, assumptions, uncertainty and strongest counter-case. Require traceable evidence for external facts. Output: the AI View column. Progression criterion: claims and assumptions can be compared. ### Step 4 — Locate divergence Input: both provisional views. Action: compare objective, facts, interpretation, option, affected parties and uncertainty. Ignore merely verbal differences. Output: a short list of exact conflicts and agreements. Progression criterion: each difference becomes a question evidence or authorised judgement could answer. ### Step 5 — Reconcile with evidence Input: each conflict and the available checking routes. Action: consult an original source, run a reversible test, ask a qualified person, or leave the issue unresolved. An AI explanation is not independent validation. Output: an evidence note for each conflict. Progression criterion: the final decision distinguishes resolved issues, human value judgements and remaining uncertainty. ### Step 6 — Record the change Input: the ledger and final judgement. Action: state whether the human changed, the AI was rejected, or the issue remains open. Record decisive evidence and the accountable owner. Output: a portable checkpoint. Progression criterion: a reviewer can see why the outcome changed without replaying the conversation. ## Decision rules and failure cases - For consequential work or when the user has relevant expertise, preserve a brief human view before revealing the AI recommendation. - For unfamiliar exploration, AI-first may be useful. The user should still restate the problem independently before adopting the answer. - Agreement without shared external evidence is not proof. - Unresolved disagreement stays unresolved; do not average incompatible claims into a false compromise. - Stop if required evidence is inaccessible, the task exceeds the user's authority, or the AI is being asked to make an irreversible decision. - Copied text, vague confidence numbers and “human approved” boxes do not create independent judgement. ## Divergence Ledger template Decision and owner: Consequence if wrong: Evidence available before AI input: Human view — objective, facts, unknowns, provisional judgement, confidence: AI view — recommendation, evidence, assumptions, uncertainty, counter-case: Exact point of divergence: Independent check used: What the check established: What remains uncertain: Final judgement: What changed, and why: Action authorised by whom: ## Smallest useful reversible experiment Hypothesis: preserving separate provisional views will expose at least one material assumption or uncertainty that an ordinary AI-first conversation would hide. Resources: three comparable low-stakes decisions, one LLM, this template, relevant sources and 30–45 minutes. Procedure: ask the AI first in one case; record the human view first in another. For the third, choose the order that fits the user's expertise, while keeping the initial views separate. Complete the ledger and compare time, confidence, substantive conflicts, checks and evidence-based changes. Success: a material assumption, error or uncertainty becomes visible and is resolved or retained. Failure: only cosmetic differences, no better check, or unacceptable delay. Stop if a case becomes consequential, requires private data or needs an authorised reviewer. Next: simplify the useful parts or abandon the workflow for that task class. This is a workflow test, not proof that either ordering is universally superior. ## Boundary example A person without clinical knowledge should not create a “human-first” diagnosis, and this capsule must not manage an urgent medical decision. Use qualified professionals. Trivial autocomplete may not justify a ledger. Use friction in proportion to consequences. ## Manual or offline route Divide paper into Human View, AI View, Divergence, Evidence and Final Judgement. One person can work at different times, or two people independently. Only access to the checked evidence is required. ## Verification checks - Were the two provisional views produced independently enough to make comparison meaningful? - Are facts separated from assumptions and value judgements? - Is every important change tied to evidence rather than fluency, status or fatigue? - Are missing sources and unresolved conflicts visible? - Is the final owner named, with explicit authorisation before any external action? - Could another reviewer reconstruct the reason for the decision from the checkpoint? ## Portable checkpoint Capsule ID and version: Decision and scope: Human provisional view: AI provisional view: Decisive divergences: Evidence checked: Decision made and owner: Unresolved questions: Next action: On request, turn this into a project passport, a field report of actual observations, or a precise request for help. Simulated outputs are not field experience. Remove private information and choose what to share; nothing is sent automatically. To propose a correction, report or help, use https://alkemata.com/collaborate/.