A lawyer preparing an appeal searches a hearing transcript for the moment a witness qualified an answer. The words are there, neatly punctuated and instantly searchable. But the crucial “not” has disappeared, or the sentence has been attributed to counsel rather than the witness. The audio contains two people speaking at once. The page does not look uncertain.
This is the central danger of automated court transcription: a probabilistic output can acquire the visual authority of a legal record before anyone has decided whether it deserves that authority. Automatic speech recognition can make proceedings easier to search, caption and review. It can give a trained reporter a useful first draft. It should not, by itself, decide what the court officially heard.
The workable boundary is therefore not “machines or humans”. It is between a machine-generated working record and a certified court record. The first can be fast, provisional and richly searchable. The second needs an accountable correction process, a traceable history and a named human or judicial authority willing to attest that it accurately represents the proceeding.

A transcript is not simply speech turned into text
Courts already use more than one method to preserve proceedings. In the United States federal judiciary, proceedings may be recorded by shorthand, stenotype, stenomask or electronic sound recording, with the method chosen by the judge. When electronic sound recording is used, designated transcription services may produce the written transcript. Yet the reporter or transcriber must file a certified electronic copy with the clerk. That division between recording method and official transcript is explicit in the Federal Court Reporting Program.
The underlying statute makes the boundary sharper. Under 28 U.S.C. § 753, court sessions are recorded verbatim by an authorised method, but a transcript is not official unless it is made from records certified by the reporter or another designated person. Technology can change how the raw record is captured. It does not erase the institutional act that gives a transcript legal standing.
That matters because a transcript performs several jobs at once. It helps a party understand what happened, lets a judge revisit evidence, supports an appeal, enables quotation and can make a hearing accessible to people who cannot use the audio. A search index may tolerate some errors if it reliably brings a reviewer to the relevant audio. An authoritative transcript cannot be judged by the same standard when a missing negation, a mistaken number or the wrong speaker name can change the meaning of evidence.
The pipeline has more than one place to fail
Automatic speech recognition is often discussed as though it were one conversion: sound enters, words emerge. A courtroom system is a chain. Microphones first capture the acoustic scene. Software decides where speech begins and ends. A recognition model maps audio to likely word sequences. Another component may add punctuation, capitalisation and paragraph breaks. A diarisation system answers a different question: who spoke when? Formatting logic then attaches roles such as judge, witness or counsel. Names, citations and technical terms may be normalised afterwards.
Each stage can be right while another is wrong. A recogniser may identify the words correctly but assign them to the wrong speaker. A diarisation system may split one person into two apparent speakers after they move away from a microphone. Two simultaneous voices may be reduced to whichever is louder. A language model may prefer a fluent, common phrase to an unusual but legally important name. Punctuation can transform a hesitant exchange into a confident statement.
Diarisation deserves particular attention because attribution is part of meaning in court. “I object” has one procedural effect when spoken by counsel and another when attached to a witness. “That was not my decision” cannot safely float free of its speaker. Research systems usually measure diarisation separately from recognition. A NIST multi-party speech challenge, for example, evaluated word error rate alongside diarisation error and Jaccard error. The separate measures acknowledge that recognising words and allocating speaking time are distinct technical problems.
Confidence scores do not solve this on their own. They are model estimates under particular training and test conditions, not probabilities that a sentence is legally safe to rely on. A low score can help route a segment to review, but a high score may still conceal the wrong speaker, an omitted short word or an unfamiliar proper noun. Court-specific risk is uneven: “fifteen” versus “fifty”, “can” versus “cannot”, and one surname substituted for another may matter more than several harmless punctuation errors.
Conditions in a hearing are not laboratory conditions
Courts contain the acoustic conditions that make the pipeline difficult. People interrupt. Counsel speak while turning away from a microphone. A witness answers quietly. Paper moves, ventilation hums and remote participants arrive through compressed audio. Names, addresses, statutes and specialist evidence are often rare in general training data. A single hearing may include dialect shifts, code-switching and several languages.
Performance can also vary across speakers. A 2020 study in the Proceedings of the National Academy of Sciences tested five commercial speech-recognition systems on conversational speech and found substantially higher average word error rates for Black speakers than for white speakers in its US datasets. The study does not establish how every current system will perform in every court. It does establish why a court cannot infer equal reliability from a single overall accuracy figure.
Interpreted proceedings add another layer. The source-language testimony, the interpreter’s rendering and any correction by the interpreter are related but not interchangeable events. If the transcript records only the interpreted speech, it should make that status clear. If both channels are preserved, the system must not merge them into one apparently seamless answer. The interpreter may ask for repetition or correct a term after hearing more context; version history must preserve that sequence rather than silently replacing the earlier words.
Non-verbal events create a further limit. A transcript can record that a document was indicated, that speech was inaudible or that several people spoke together, but it cannot reconstruct an event the microphones did not capture. The correct output is sometimes not a guessed sentence but a visible marker of uncertainty linked to the audio. “Inaudible” is less satisfying than fluent prose, yet it is more honest and more reviewable.
What automation is genuinely good for
The strongest case for automation is practical. Transcripts can be costly and slow to produce. A searchable draft helps a reporter navigate hours of audio, locate names and compare repeated terms. Live captions can improve immediate participation for some users. Judges and lawyers can search a provisional record to find a passage, then listen to the source recording. Structured timestamps can connect a page to the exact audio segment that supports it.
This is not a trivial gain. In England and Wales, the updated HM Courts & Tribunals Service transcript guidance describes a process in which applicants request particular portions of recorded proceedings, authorised companies produce transcripts, and some outputs require judicial approval before delivery. Private proceedings have additional permission and secure-transcription requirements. Better indexing could reduce the labour of finding the relevant portion without collapsing those legal controls.
A well-designed system therefore uses automation to reduce navigation and drafting work, not to make certification disappear. It can highlight low-confidence segments, unusual terms, speaker changes and overlaps. It can compare the draft against a case-specific glossary. It can expose, rather than hide, competing possibilities. Most importantly, it can keep the audio one click away from every disputed passage.
The counterargument is that requiring human review of every line preserves the very bottleneck automation is meant to remove. That objection is serious. A system that merely adds machine output to an unchanged manual workflow may cost more without improving access. The answer is not ceremonial review. It is risk-based work allocation within a clear legal boundary: automation handles first-pass transcription, timing, search and consistency checks; trained people concentrate on attribution, uncertainty and consequential passages; certification remains an accountable act.
The record needs a history, not just a final file
A trustworthy transcript system should preserve distinct artefacts. The original authorised recording is the evidential source. The machine output is a draft. Human edits should produce a reviewable revision. Certification creates the official version. Later corrections should create a new version linked to the order, stipulation or decision that authorised the change.
That sequence should be technically visible. Each edit needs an actor, time and reason. The system should show whether a change corrected recognition, speaker attribution, punctuation, a proper name or a redaction. File integrity checks can demonstrate that the underlying audio has not changed. Access controls can protect sealed or sensitive material without destroying the audit trail. Retention rules should distinguish the source recording, provisional drafts, the certified transcript and logs of subsequent correction.
Without that separation, a vendor may improve its model, rerun old audio and silently produce different words. A well-intentioned editor may clean up grammar and erase a hesitation that mattered. A redaction may be applied to one copy but not another. Version provenance turns these from invisible risks into governable events.
Correction is not an embarrassing exception; it is part of what makes a record legitimate. The Federal Rules of Appellate Procedure, Rule 10(e), as reproduced by the Fourth Circuit, provide that a dispute over whether the record truly discloses what happened is settled by the trial court. Material omissions or misstatements can be corrected, and a supplemental record can be certified. An automated system should make that institutional process easier to use, not overwrite it with an opaque “latest version”.
Accuracy must be measured at the point of consequence
Word error rate is useful, but it averages substitutions, deletions and insertions across a body of speech. Courts need additional measures. Speaker-attribution error should be reported separately. Evaluators should count mistakes in names, dates, amounts, citations and negations. They should test overlapping speech, quiet voices, remote connections, interpreters and the language varieties heard in the actual jurisdiction.
Operational measures matter as much as benchmark accuracy. How often do reporters have to return to the audio? Which errors survive into the certified version? How long does correction take? Do parties receive notice of revisions? Can a self-represented person identify and challenge a disputed passage without specialist software? Are errors concentrated among particular speakers or hearing types? A court should know not only whether automation saves transcription time, but whether it reduces the time between a challenge and a reliable correction.
Review interfaces should also avoid anchoring the certifier to a plausible machine sentence. Showing the proposed text before the audio encourages people to hear what the screen suggests. For flagged or high-consequence segments, the reviewer may need to hear the audio first, compare alternative renderings and confirm the speaker independently. Review must have enough time and authority to change the draft; otherwise certification becomes a decorative signature.
Who, then, owns the transcript?
The useful answer is not primarily about copyright or possession. The court must own the rules that determine authority. It must decide which recording is the source, who may transcribe it, who certifies the result, how parties can inspect and challenge it, which version is filed and how corrections propagate to the appellate record. A contractor may supply software or labour, but it should not determine these legal facts through product defaults.
The certifier owns a professional responsibility: to compare the draft with the source and refuse certainty where the recording does not support it. The judge or court owns the responsibility for resolving formal disputes about the record. Parties need effective access to the relevant version and a usable route to correction. The technology’s responsibility is narrower but demanding: preserve provenance, expose uncertainty, link text to audio and prevent silent revision.
If you value reporting that examines where technology should assist and where human authority must remain visible, you can subscribe to Alkemata for the next article in this series.
The unresolved decision is concrete. Before a court buys an automated transcription system, it should specify which output has legal force, who can certify it, and what evidence will show that correction works across the speakers and hearing conditions the court actually encounters. If those answers live only in a vendor interface, the court has automated not transcription but authority.