Benchmarks

Measured, including where we lose.

19 models across 34 runs

Every model transcribed the same recordings in each of 5 languages, 902 to 2,936 samples per language and 8,287 in all. Every transcript was scored the same way. How we measure.

Head to head with Wispr Flow

Codictate leads four of five languages. Wispr Flow leads Danish by 5.1 points.

English (clean)

CodictateWhisper Large V3 Q5_0
97.1% word accuracy.
97.1%
Wispr Flow
96.0% word accuracy.
96.0%

English (noisy)

CodictateWhisper Large V3
94.9% word accuracy.
94.9%
Wispr Flow
93.3% word accuracy.
93.3%

Spanish

CodictateWhisper Large V3 Q5_0
97.1% word accuracy.
97.1%
Wispr Flow
95.8% word accuracy.
95.8%

Danish

Wispr Flow
93.8% word accuracy.
93.8%
CodictateHviske v5 Tiny Q5_0
88.7% word accuracy.
88.7%

Hungarian

CodictateWhisper Large V3 Q5_0
85.6% word accuracy.
85.6%
Wispr Flow
77.9% word accuracy.
77.9%
50100

Bar length is word accuracy on a shared 50 to 100 scale, so a longer bar is a better result.

Response time

Milliseconds of waiting for every second of audio you dictate, so lower is better. Point at a cell for the actual wait per clip.

ModelEnglish (clean)English (noisy)SpanishDanishHungarian
Wispr Flow 1.6.89760.2 ms/sWait per clip, after you stop speaking: median 401 ms, p90 653 ms.65.1 ms/sWait per clip, after you stop speaking: median 384 ms, p90 607 ms.40.8 ms/sWait per clip, after you stop speaking: median 471 ms, p90 655 ms.158.7 ms/sWait per clip, after you stop speaking: median 1.8 s, p90 2.3 s.155.5 ms/sWait per clip, after you stop speaking: median 1.8 s, p90 2.4 s.
Whisper Large V3 Turbo Q5_075.2 ms/sWait per clip, after you stop speaking: median 553 ms, p90 591 ms.85.0 ms/sWait per clip, after you stop speaking: median 550 ms, p90 585 ms.47.9 ms/sWait per clip, after you stop speaking: median 584 ms, p90 615 ms.52.3 ms/sWait per clip, after you stop speaking: median 592 ms, p90 627 ms.50.5 ms/sWait per clip, after you stop speaking: median 611 ms, p90 655 ms.
Whisper Large V3 Q5_0116.7 ms/sWait per clip, after you stop speaking: median 842 ms, p90 1.0 s.129.7 ms/sWait per clip, after you stop speaking: median 824 ms, p90 971 ms.80.0 ms/sWait per clip, after you stop speaking: median 965 ms, p90 1.1 s.89.3 ms/sWait per clip, after you stop speaking: median 1.0 s, p90 1.2 s.90.3 ms/sWait per clip, after you stop speaking: median 1.1 s, p90 1.3 s.
Whisper Large V3181.7 ms/sWait per clip, after you stop speaking: median 1.3 s, p90 1.5 s.202.7 ms/sWait per clip, after you stop speaking: median 1.3 s, p90 1.5 s.121.9 ms/sWait per clip, after you stop speaking: median 1.5 s, p90 1.7 s.135.0 ms/sWait per clip, after you stop speaking: median 1.5 s, p90 1.7 s.135.5 ms/sWait per clip, after you stop speaking: median 1.6 s, p90 1.9 s.
Whisper Large V3 Turbo104.8 ms/sWait per clip, after you stop speaking: median 771 ms, p90 820 ms.118.8 ms/sWait per clip, after you stop speaking: median 770 ms, p90 820 ms.67.2 ms/sWait per clip, after you stop speaking: median 819 ms, p90 864 ms.72.9 ms/sWait per clip, after you stop speaking: median 826 ms, p90 876 ms.70.6 ms/sWait per clip, after you stop speaking: median 856 ms, p90 917 ms.
Parakeet TDT 0.6B v311.4 ms/sWait per clip, after you stop speaking: median 61 ms, p90 174 ms.12.0 ms/sWait per clip, after you stop speaking: median 59 ms, p90 166 ms.10.4 ms/sWait per clip, after you stop speaking: median 127 ms, p90 194 ms.10.9 ms/sWait per clip, after you stop speaking: median 123 ms, p90 184 ms.10.7 ms/sWait per clip, after you stop speaking: median 130 ms, p90 197 ms.
Hviske v5 Tiny Q5_0Targets Danish only.Targets Danish only.Targets Danish only.18.5 ms/sWait per clip, after you stop speaking: median 206 ms, p90 248 ms.Targets Danish only.

Wispr Flow 1.6.897 vs Codictate, September 2026

Response times are not measured the same way for both products: Codictate is timed at the direct adapter call boundary, Wispr Flow is timed from the UI-observed paste.

Wispr Flow streams while you speak, and is measured over the network.

Speed figures from runs before 4 September 2026 are not shown. Full method.

All Codictate models

The 19 models worth running day to day. Ratings for a quick overview, exact percentages if you want the detail.

Ratings run 1 to 10 across every benchmark language. Point at any cell for the figure behind it.

Whisper Large V3 Turbo Q5_0574 MB778 MBAll languagesNo8/1010/1093.5% word accuracy, pooled across every language it measured.
Whisper Large V3 Q5_01.1 GB1.5 GBAll languagesYes7/1010/1094.0% word accuracy, pooled across every language it measured.
Whisper Large V23.0 GB3.5 GBAll languagesYes6/1010/1092.9% word accuracy, pooled across every language it measured.
Whisper Large V2 Q5_01.1 GB1.5 GBAll languagesYes7/1010/1092.8% word accuracy, pooled across every language it measured.
Whisper Large V2 Q8_01.6 GB2.1 GBAll languagesYes7/1010/1092.8% word accuracy, pooled across every language it measured.
Whisper Large V33.0 GB3.5 GBAll languagesYes6/1010/1094.0% word accuracy, pooled across every language it measured.
Whisper Large V3 Turbo1.6 GB1.9 GBAll languagesNo8/1010/1093.6% word accuracy, pooled across every language it measured.
Whisper Large V3 Turbo Q8_0834 MB1.1 GBAll languagesNo8/1010/1093.6% word accuracy, pooled across every language it measured.
Parakeet TDT 0.6B v3500 MB80 MB25 languagesNo10/109/1092.0% word accuracy, pooled across every language it measured.
Hviske v5 Tiny F16503 MB647 MBDanish onlyNo9/109/10da88.6% word accuracy on Danish, the only language this model targets.
Hviske v5 Tiny Q8_0268 MB413 MBDanish onlyNo10/109/10da88.6% word accuracy on Danish, the only language this model targets.
Hviske v5 Tiny Q6_K232 MB377 MBDanish onlyNo10/109/10da88.6% word accuracy on Danish, the only language this model targets.
Hviske v5 Tiny Q5_0181 MB327 MBDanish onlyNo10/109/10da88.7% word accuracy on Danish, the only language this model targets.
Hviske v5 Tiny Q4_K153 MB298 MBDanish onlyNo10/109/10da88.6% word accuracy on Danish, the only language this model targets.
Whisper Large V13.0 GB3.5 GBAll languagesYes6/109/1091.9% word accuracy, pooled across every language it measured.
Whisper Medium1.5 GB1.9 GBAll languagesYes8/109/1090.8% word accuracy, pooled across every language it measured.
Whisper Medium Q5_0514 MB878 MBAll languagesYes8/109/1090.8% word accuracy, pooled across every language it measured.
Whisper Medium Q8_0766 MB1.2 GBAll languagesYes8/109/1090.8% word accuracy, pooled across every language it measured.
Whisper Small Q5_1181 MB383 MBAll languagesYes9/107/1080.7% word accuracy, pooled across every language it measured.

Which model should you pick

Five ways to read the table above, depending on what you dictate and what your machine can spare.

Best accuracy across languages

Whisper Large V3 Q5_0 is the most accurate multilingual model we ship, at 94.0% word accuracy counted across every language we test. The full-size Whisper Large V3 scores 94.0% and needs 3.0 GB on disk. The compressed build needs 1.1 GB for the same result, so it is the easier one on your computer.

Fastest

Parakeet TDT 0.6B v3 is the quickest thing we ship and rates 10 out of 10 for speed, while holding 92.0% word accuracy. It covers 25 languages rather than the whole Whisper list and does not translate, so it is the pick when you dictate in a language it supports and want the text back immediately.

Daily driver

If you switch languages or work in noise, Whisper Large V3 Turbo Q5_0 is the safer default: 93.5% word accuracy at 574 MB on disk, and 8 out of 10 for speed. If you dictate English in a quiet room, Parakeet is faster and covers less ground.

Translation

Whisper Large V3 Q5_0 is the accuracy pick above and it also translates. The Turbo builds do not, which the translation column above records, so if you need translation at this accuracy this is the model.

Lightweight

Whisper Small Q5_1 gives the best accuracy for its size: 80.7% word accuracy at 181 MB on disk and 383 MB of peak memory, across every language, translation included. If your machine is tight on memory, this is the one.

All benchmark runs

Each run records the dataset, the sample count and the hardware with its results.