Trust

Transcription accuracy benchmark with open test files

This first-party benchmark publishes the difficult audio sample, human-reviewed reference, model output, scoring definitions, configuration, results, and limitations. It is evidence from one noisy radio recording, not a universal accuracy claim.

Representative difficulty

The 178.66-second Apollo 13 sample contains radio noise, technical vocabulary, short backchannels, and two mapped roles: HOUSTON and SPACECRAFT.

Repeatable scoring

Text WER measures normalized word errors. Speaker-attributed WER also requires the mapped speaker role to match. Matched-word role accuracy isolates speaker assignment on exactly aligned words.

Configuration disclosure

The tested path used Whisper large-v3, English, word timestamps, Fast ECAPA speaker labeling, and the sparse-VAD fallback.

Honest scope

A single English radio sample cannot predict performance on meetings, accents, languages, music, crosstalk, or different microphones. Use the downloads to repeat the test and add your own representative files.

Test sample and configuration

Whisper large-v3, English, word timestamps, Fast ECAPA, two mapped roles, sparse-VAD fallback

Source: NASA Apollo 13 mission communications . Used for factual evaluation under the NASA media usage guidelines ; no endorsement is implied.

Audio duration178.66 seconds
Reference words391
Predicted words350
Matched words scored for role281

Measured results

Lower WER is better. Higher matched-word role accuracy is better.

ConfigurationText WERSpeaker WERRole accuracyGPU worker time
Production-style sparse-VAD fallback29.7%30.9%98.2%24.7 s
Direct no-VAD, two speakers29.7%30.9%98.2%15.4 s
Forced three speakers29.7%33.2%95.0%16.7 s
Word timestamps disabled26.9%47.6%72.3%15.3 s

What this test shows

  • The production-style path reached 29.7% text WER and 98.2% matched-word role accuracy on this noisy sample.
  • Direct no-VAD matched quality with lower runtime here, but the result does not justify disabling VAD for arbitrary user audio.
  • Forcing extra speakers reduced role accuracy, and disabling word timestamps severely damaged speaker attribution even though plain text WER improved.
  • A separate uncertainty flag caught six of eight wrong-speaker words, with two false positives. It is useful for review, not automatic correction.

Download and reproduce the benchmark

Inspect the same files, verify their hashes, or score another transcription system.

Limitations

  • One English, two-role radio recording is not representative of every transcription workflow.
  • The reference transcript and speaker-role mapping include human editorial judgment.
  • Anonymous output speakers are mapped to roles using the reference; this evaluation step is not automatic speaker identification.
  • Runtime is GPU worker wall time, not end-to-end latency or a billing measurement.
  • NASA is the source of the mission audio and does not endorse InstantTranscriber or this benchmark.

FAQ

What result did InstantTranscriber get on this file?

The production-style configuration measured 29.7% text WER, 30.9% speaker-attributed WER, and 98.2% matched-word role accuracy on this difficult Apollo 13 radio sample.

Does this benchmark prove general transcription accuracy?

No. It is one repeatable, difficult sample. Accuracy depends on audio quality, language, vocabulary, speakers, overlap, and evaluation rules.

Can I download the benchmark files?

Yes. The audio, reference transcript, model output, machine-readable results, hashes, configuration, and limitations are published on this page.

Transcribe with clear data handling

Upload audio or video, generate a transcript, and export only the formats your workflow needs.

Instant access. No credit card required.

Sign up is required before uploading or transcribing.

Free Forever

Free Plan

$0

No credit card required

  • 3 transcriptions per day
  • Max 35 minutes per file
  • Max 50 MB per file
  • First transcript summary included
  • Export to TXT, DOCX, PDF, SRT, VTT