Transcription accuracy benchmark with open test files
This first-party benchmark publishes the difficult audio sample, human-reviewed reference, model output, scoring definitions, configuration, results, and limitations. It is evidence from one noisy radio recording, not a universal accuracy claim.
Representative difficulty
The 178.66-second Apollo 13 sample contains radio noise, technical vocabulary, short backchannels, and two mapped roles: HOUSTON and SPACECRAFT.
Repeatable scoring
Text WER measures normalized word errors. Speaker-attributed WER also requires the mapped speaker role to match. Matched-word role accuracy isolates speaker assignment on exactly aligned words.
Configuration disclosure
The tested path used Whisper large-v3, English, word timestamps, Fast ECAPA speaker labeling, and the sparse-VAD fallback.
Honest scope
A single English radio sample cannot predict performance on meetings, accents, languages, music, crosstalk, or different microphones. Use the downloads to repeat the test and add your own representative files.
Test sample and configuration
Whisper large-v3, English, word timestamps, Fast ECAPA, two mapped roles, sparse-VAD fallback
Source: NASA Apollo 13 mission communications . Used for factual evaluation under the NASA media usage guidelines ; no endorsement is implied.
Measured results
Lower WER is better. Higher matched-word role accuracy is better.
| Configuration | Text WER | Speaker WER | Role accuracy | GPU worker time |
|---|---|---|---|---|
| Production-style sparse-VAD fallback | 29.7% | 30.9% | 98.2% | 24.7 s |
| Direct no-VAD, two speakers | 29.7% | 30.9% | 98.2% | 15.4 s |
| Forced three speakers | 29.7% | 33.2% | 95.0% | 16.7 s |
| Word timestamps disabled | 26.9% | 47.6% | 72.3% | 15.3 s |
What this test shows
- The production-style path reached 29.7% text WER and 98.2% matched-word role accuracy on this noisy sample.
- Direct no-VAD matched quality with lower runtime here, but the result does not justify disabling VAD for arbitrary user audio.
- Forcing extra speakers reduced role accuracy, and disabling word timestamps severely damaged speaker attribution even though plain text WER improved.
- A separate uncertainty flag caught six of eight wrong-speaker words, with two false positives. It is useful for review, not automatic correction.
Download and reproduce the benchmark
Inspect the same files, verify their hashes, or score another transcription system.
Limitations
- One English, two-role radio recording is not representative of every transcription workflow.
- The reference transcript and speaker-role mapping include human editorial judgment.
- Anonymous output speakers are mapped to roles using the reference; this evaluation step is not automatic speaker identification.
- Runtime is GPU worker wall time, not end-to-end latency or a billing measurement.
- NASA is the source of the mission audio and does not endorse InstantTranscriber or this benchmark.
FAQ
What result did InstantTranscriber get on this file?
The production-style configuration measured 29.7% text WER, 30.9% speaker-attributed WER, and 98.2% matched-word role accuracy on this difficult Apollo 13 radio sample.
Does this benchmark prove general transcription accuracy?
No. It is one repeatable, difficult sample. Accuracy depends on audio quality, language, vocabulary, speakers, overlap, and evaluation rules.
Can I download the benchmark files?
Yes. The audio, reference transcript, model output, machine-readable results, hashes, configuration, and limitations are published on this page.
Transcribe with clear data handling
Upload audio or video, generate a transcript, and export only the formats your workflow needs.
Instant access. No credit card required.
Sign up is required before uploading or transcribing.
Free Plan
$0
No credit card required
- 3 transcriptions per day
- Max 35 minutes per file
- Max 50 MB per file
- First transcript summary included
- Export to TXT, DOCX, PDF, SRT, VTT
Pro Plan
$9.99
per month - or $5.99/mo billed annually (save 40%)
- Unlimited transcriptions
- Parallel transcription jobs
- Max 10 hours per file
- Max 3 GB per file
- Higher-quality speaker labels
- Transcript summaries
- Priority support
No credit card needed for trial