Benchmark run

Hviske vs main models

This benchmark compares the hviske model to the already main models recommended in Codictate

Aug 18, 2026

Models
7
Samples per dataset
200
Warmup runs
3
Hardware
Apple M4 Max / 36 GB / macOS 26.5.1

Full results

Accuracy is the share of words transcribed correctly per language condition, so higher is better; speed is milliseconds per second of audio, so lower is better. The best value in each column is emphasised, and a cell reading a dash says on hover why there is no figure in it.

Accuracy (%) per language, and speed in ms per second of audio, so lower is better. A cell reading a dash was never measured, and hovering any cell says what is behind it.

Hviske v5 Tiny Q5_0197181 MB282 MB16 ms88.7%Every word from every language counted together, not an average of per-language scores.Not scored on any benchmark language.88.7%Every word from every language counted together, not an average of per-language scores.88.7%
Hviske v5 Tiny Q4_K197153 MB253 MB16 ms88.6%Every word from every language counted together, not an average of per-language scores.Not scored on any benchmark language.88.6%Every word from every language counted together, not an average of per-language scores.88.6%
Hviske v5 Tiny F16197503 MB601 MB21 ms88.5%Every word from every language counted together, not an average of per-language scores.Not scored on any benchmark language.88.5%Every word from every language counted together, not an average of per-language scores.88.5%
Hviske v5 Tiny Q8_0197268 MB368 MB18 ms88.5%Every word from every language counted together, not an average of per-language scores.Not scored on any benchmark language.88.5%Every word from every language counted together, not an average of per-language scores.88.5%
Hviske v5 Tiny Q6_K197232 MB332 MB17 ms88.3%Every word from every language counted together, not an average of per-language scores.Not scored on any benchmark language.88.3%Every word from every language counted together, not an average of per-language scores.88.3%
Whisper Large V3 Turbo Q5_0197574 MB742 MB52 ms84.8%Every word from every language counted together, not an average of per-language scores.Not scored on any benchmark language.84.8%Every word from every language counted together, not an average of per-language scores.84.8%
Parakeet TDT 0.6B v3197500 MB79 MB17 ms80.6%Every word from every language counted together, not an average of per-language scores.Not scored on any benchmark language.80.6%Every word from every language counted together, not an average of per-language scores.80.6%

At a glance

Ratings computed from benchmark data, scaled 1 to 10, and written with their scale in every cell. Accuracy is based on Word Error Rate (WER) and does not include punctuation yet.

Ratings run 1 to 10, higher is better. A cell reading a dash was never measured, and hovering any cell says what is behind it.

Hviske v5 Tiny F16Danish onlyNo9/109/10Every word from every language counted together, not an average of per-language scores.
Hviske v5 Tiny Q8_0Danish onlyNo10/109/10Every word from every language counted together, not an average of per-language scores.
Hviske v5 Tiny Q6_KDanish onlyNo10/109/10Every word from every language counted together, not an average of per-language scores.
Hviske v5 Tiny Q5_0Danish onlyNo10/109/10Every word from every language counted together, not an average of per-language scores.
Hviske v5 Tiny Q4_KDanish onlyNo10/109/10Every word from every language counted together, not an average of per-language scores.
Whisper Large V3 Turbo Q5_0All languagesNo9/108/10Every word from every language counted together, not an average of per-language scores.
Parakeet TDT 0.6B v325 languagesNo10/107/10Every word from every language counted together, not an average of per-language scores.

Charts

The plots generated alongside this run. Select a chart to open it full size.

Bar chart comparing model accuracy across English, Spanish, Danish, and Hungarian benchmark conditions.
Accuracy by model and test condition
Bar chart comparing transcription speed across models and test conditions.
Speed comparison across conditions
Bar chart showing average English, multilingual, and overall accuracy per model.
Average accuracy by language group