Benchmark run

2026 09 v4 english tail

English clips 951 to corpus end, 13 multilingual models, matching the Flow depth

Sep 22, 2026

Models
13
Samples per dataset
1986
Warmup runs
3
Hardware
Apple M4 Max / 36 GB / macOS 26.6.2

Full results

Accuracy is the share of words transcribed correctly per language condition, so higher is better; speed is milliseconds per second of audio, so lower is better. The best value in each column is emphasised, and a cell reading a dash says on hover why there is no figure in it.

Accuracy (%) per language, and speed in ms per second of audio, so lower is better. A cell reading a dash was never measured, and hovering any cell says what is behind it.

Whisper Large V3 Turbo Q8_01667834 MB1.1 GB88 ms96.0%Every word from every language counted together, not an average of per-language scores.96.0%Every word from every language counted together, not an average of per-language scores.Not scored on any benchmark language.97.1%94.9%
Whisper Large V3 Turbo Q5_01667574 MB783 MB80 ms95.9%Every word from every language counted together, not an average of per-language scores.95.9%Every word from every language counted together, not an average of per-language scores.Not scored on any benchmark language.97.1%94.9%
Whisper Large V3 Q5_016671.1 GB1.6 GB123 ms95.9%Every word from every language counted together, not an average of per-language scores.95.9%Every word from every language counted together, not an average of per-language scores.Not scored on any benchmark language.97.1%94.8%
Whisper Large V316673.0 GB3.5 GB192 ms95.9%Every word from every language counted together, not an average of per-language scores.95.9%Every word from every language counted together, not an average of per-language scores.Not scored on any benchmark language.97.1%94.8%
Whisper Large V3 Turbo16671.6 GB1.9 GB111 ms95.9%Every word from every language counted together, not an average of per-language scores.95.9%Every word from every language counted together, not an average of per-language scores.Not scored on any benchmark language.97.1%94.7%
Whisper Large V216673.0 GB3.5 GB193 ms95.4%Every word from every language counted together, not an average of per-language scores.95.4%Every word from every language counted together, not an average of per-language scores.Not scored on any benchmark language.96.8%94.1%
Whisper Large V2 Q8_016671.6 GB2.1 GB142 ms95.4%Every word from every language counted together, not an average of per-language scores.95.4%Every word from every language counted together, not an average of per-language scores.Not scored on any benchmark language.96.8%94.1%
Whisper Large V2 Q5_016671.1 GB1.5 GB124 ms95.4%Every word from every language counted together, not an average of per-language scores.95.4%Every word from every language counted together, not an average of per-language scores.Not scored on any benchmark language.96.8%94.0%
Whisper Large V116673.0 GB3.5 GB194 ms95.2%Every word from every language counted together, not an average of per-language scores.95.2%Every word from every language counted together, not an average of per-language scores.Not scored on any benchmark language.96.7%93.7%
Parakeet TDT 0.6B v31667500 MB79 MB9 ms95.2%Every word from every language counted together, not an average of per-language scores.95.2%Every word from every language counted together, not an average of per-language scores.Not scored on any benchmark language.96.2%94.2%
Whisper Medium16671.5 GB1.9 GB108 ms94.8%Every word from every language counted together, not an average of per-language scores.94.8%Every word from every language counted together, not an average of per-language scores.Not scored on any benchmark language.96.4%93.4%
Whisper Medium Q8_01667766 MB1.2 GB82 ms94.8%Every word from every language counted together, not an average of per-language scores.94.8%Every word from every language counted together, not an average of per-language scores.Not scored on any benchmark language.96.4%93.4%
Whisper Medium Q5_01667514 MB881 MB74 ms94.8%Every word from every language counted together, not an average of per-language scores.94.8%Every word from every language counted together, not an average of per-language scores.Not scored on any benchmark language.96.4%93.3%

At a glance

Ratings computed from benchmark data, scaled 1 to 10, and written with their scale in every cell. Accuracy is based on Word Error Rate (WER) and does not include punctuation yet.

Ratings run 1 to 10, higher is better. A cell reading a dash was never measured, and hovering any cell says what is behind it.

Whisper Large V3 Turbo Q5_0All languagesNo8/1010/10Every word from every language counted together, not an average of per-language scores.
Whisper Large V3 Q5_0All languagesYes7/1010/10Every word from every language counted together, not an average of per-language scores.
Parakeet TDT 0.6B v325 languagesNo10/1010/10Every word from every language counted together, not an average of per-language scores.
Whisper Large V1All languagesYes5/1010/10Every word from every language counted together, not an average of per-language scores.
Whisper Large V2All languagesYes5/1010/10Every word from every language counted together, not an average of per-language scores.
Whisper Large V2 Q5_0All languagesYes7/1010/10Every word from every language counted together, not an average of per-language scores.
Whisper Large V2 Q8_0All languagesYes6/1010/10Every word from every language counted together, not an average of per-language scores.
Whisper Large V3All languagesYes5/1010/10Every word from every language counted together, not an average of per-language scores.
Whisper Large V3 TurboAll languagesNo7/1010/10Every word from every language counted together, not an average of per-language scores.
Whisper Large V3 Turbo Q8_0All languagesNo8/1010/10Every word from every language counted together, not an average of per-language scores.
Whisper MediumAll languagesYes7/1010/10Every word from every language counted together, not an average of per-language scores.
Whisper Medium Q5_0All languagesYes8/1010/10Every word from every language counted together, not an average of per-language scores.
Whisper Medium Q8_0All languagesYes8/1010/10Every word from every language counted together, not an average of per-language scores.

Charts

The plots generated alongside this run. Select a chart to open it full size.

Bar chart comparing model accuracy across English, Spanish, Danish, and Hungarian benchmark conditions.
Accuracy by model and test condition
Bar chart comparing transcription speed across models and test conditions.
Speed comparison across conditions
Bar chart showing average English, multilingual, and overall accuracy per model.
Average accuracy by language group