Benchmark run

Crispasr vs whisper

to test which asr harnes is best.

Aug 17, 2026

Models
6
Samples per dataset
20
Warmup runs
3
Hardware
Apple M4 Max / 36 GB / macOS 26.5.1

Full results

Accuracy is the share of words transcribed correctly per language condition, so higher is better; speed is milliseconds per second of audio, so lower is better. The best value in each column is emphasised, and a cell reading a dash says on hover why there is no figure in it.

Accuracy (%) per language, and speed in ms per second of audio, so lower is better. A cell reading a dash was never measured, and hovering any cell says what is behind it.

Whisper Large V3 Turbo Q5_0 (whisper-cli)17574 MB801 MB108 ms96.0%Every word from every language counted together, not an average of per-language scores.95.9%Every word from every language counted together, not an average of per-language scores.96.2%Every word from every language counted together, not an average of per-language scores.96.9%94.6%96.2%
Whisper Large V3 Q5_0 (whisper-cli)171.1 GB2.0 GB154 ms95.7%Every word from every language counted together, not an average of per-language scores.95.7%Every word from every language counted together, not an average of per-language scores.95.7%Every word from every language counted together, not an average of per-language scores.96.9%94.3%95.7%
Whisper Large V3 Q5_0171.1 GB1.5 GB104 ms95.7%Every word from every language counted together, not an average of per-language scores.95.9%Every word from every language counted together, not an average of per-language scores.95.5%Every word from every language counted together, not an average of per-language scores.96.9%94.6%95.5%
Whisper Large V3 Turbo Q5_017574 MB738 MB69 ms95.5%Every word from every language counted together, not an average of per-language scores.95.0%Every word from every language counted together, not an average of per-language scores.96.2%Every word from every language counted together, not an average of per-language scores.95.6%94.3%96.2%
Whisper Medium Q5_0 en (whisper-cli)17514 MB1.1 GB99 ms61.7%Every word from every language counted together, not an average of per-language scores.94.1%Every word from every language counted together, not an average of per-language scores.14.5%Every word from every language counted together, not an average of per-language scores.95.0%93.0%14.5%
Whisper Medium Q5_0 en17514 MB832 MB60 ms61.6%Every word from every language counted together, not an average of per-language scores.94.4%Every word from every language counted together, not an average of per-language scores.13.7%Every word from every language counted together, not an average of per-language scores.95.0%93.6%13.7%

At a glance

Ratings computed from benchmark data, scaled 1 to 10, and written with their scale in every cell. Accuracy is based on Word Error Rate (WER) and does not include punctuation yet.

Ratings run 1 to 10, higher is better. A cell reading a dash was never measured, and hovering any cell says what is behind it.

Whisper Large V3 Turbo Q5_0 (whisper-cli)All languagesNo7/1010/10Every word from every language counted together, not an average of per-language scores.
Whisper Large V3 Turbo Q5_0All languagesNo8/1010/10Every word from every language counted together, not an average of per-language scores.
Whisper Large V3 Q5_0 (whisper-cli)All languagesYes6/1010/10Every word from every language counted together, not an average of per-language scores.
Whisper Large V3 Q5_0All languagesYes7/1010/10Every word from every language counted together, not an average of per-language scores.
Whisper Medium Q5_0 enEnglish onlyNo8/103/10Every word from every language counted together, not an average of per-language scores.
Whisper Medium Q5_0 en (whisper-cli)English onlyNo7/103/10Every word from every language counted together, not an average of per-language scores.

Charts

The plots generated alongside this run. Select a chart to open it full size.

Bar chart comparing model accuracy across English, Spanish, Danish, and Hungarian benchmark conditions.
Accuracy by model and test condition
Bar chart comparing transcription speed across models and test conditions.
Speed comparison across conditions
Bar chart showing average English, multilingual, and overall accuracy per model.
Average accuracy by language group