Apollo 13 radio communications transcription benchmark Published by InstantTranscriber on 2026-07-10 from an evaluation run dated 2026-05-19. Files - apollo-13-radio-audio.ogg: 178.66-second NASA mission audio sample. - reference-transcript.txt: human-reviewed reference transcript and roles. - instanttranscriber-output.txt: plain model output from the tested configuration. - production-output.json: sanitized production-style segments, words, timestamps, and speakers. - direct-no-vad-two-speakers-output.json: sanitized direct no-VAD, two-speaker output. - forced-three-speakers-output.json: sanitized direct no-VAD output forced to three speakers. - no-word-timestamps-output.json: sanitized direct no-VAD output without word timestamps. - *-speaker-output.txt: human-readable versions of the four scored outputs. - results.json: machine-readable configuration, metrics, findings, hashes, and limitations. - score.py: dependency-free scorer for the published WER and role-accuracy results. - speaker-uncertainty-evaluation.md: uncertainty-flag evaluation details. Tested configuration - ASR: Whisper large-v3, English, word timestamps enabled. - Speaker labeling: Fast ECAPA with two mapped roles. - VAD: requested, with the sparse-output fallback used on this sample. Headline result - Text WER: 29.7%. - Speaker-attributed WER: 30.9%. - Matched-word speaker-role accuracy: 98.2%. - GPU worker wall time: 24.7 seconds. Reproduce the displayed quality metrics Run from this directory: python3 score.py reference-transcript.txt \ production-output.json \ direct-no-vad-two-speakers-output.json \ forced-three-speakers-output.json \ no-word-timestamps-output.json The scorer lowercases text, removes punctuation, and tokenizes English words and numbers. Text WER is Levenshtein word error rate. Speaker-attributed WER applies the same distance to (role, word) pairs. Role accuracy is the share of exactly aligned words assigned to the correct role. Reference SPEAKER_01 is HOUSTON; the other reference speakers are SPACECRAFT. For each output, the scorer tests all mappings from anonymous speaker IDs to those two roles and reports the mapping with the lowest speaker-attributed WER. This mapping step uses the reference and should not be interpreted as automatic speaker identification. This is deliberately difficult radio audio. It is one sample, not a general accuracy claim. Results depend on language, recording quality, vocabulary, number of speakers, overlap, and evaluation rules. Source and usage NASA is the source of the mission audio. Its media content is generally available for informational use subject to NASA's media usage guidelines and without implying NASA endorsement. Mission context: https://www.nasa.gov/history/houston-weve-had-a-problem/ NASA media usage guidelines: https://www.nasa.gov/nasa-brand-center/images-and-media/