Benchmarks
Measured, including where we lose.
19 models across 34 runs
Every model transcribed the same recordings in each of 5 languages, 902 to 2,936 samples per language and 8,287 in all. Every transcript was scored the same way. How we measure.
Head to head with Wispr Flow
Codictate leads four of five languages. Wispr Flow leads Danish by 5.1 points.
English (clean)
English (noisy)
Spanish
Danish
Hungarian
Bar length is word accuracy on a shared 50 to 100 scale, so a longer bar is a better result.
Response time
Milliseconds of waiting for every second of audio you dictate, so lower is better. Point at a cell for the actual wait per clip.
| Model | English (clean) | English (noisy) | Spanish | Danish | Hungarian |
|---|---|---|---|---|---|
| Wispr Flow 1.6.897 | 60.2 ms/sWait per clip, after you stop speaking: median 401 ms, p90 653 ms. | 65.1 ms/sWait per clip, after you stop speaking: median 384 ms, p90 607 ms. | 40.8 ms/sWait per clip, after you stop speaking: median 471 ms, p90 655 ms. | 158.7 ms/sWait per clip, after you stop speaking: median 1.8 s, p90 2.3 s. | 155.5 ms/sWait per clip, after you stop speaking: median 1.8 s, p90 2.4 s. |
| Whisper Large V3 Turbo Q5_0 | 75.2 ms/sWait per clip, after you stop speaking: median 553 ms, p90 591 ms. | 85.0 ms/sWait per clip, after you stop speaking: median 550 ms, p90 585 ms. | 47.9 ms/sWait per clip, after you stop speaking: median 584 ms, p90 615 ms. | 52.3 ms/sWait per clip, after you stop speaking: median 592 ms, p90 627 ms. | 50.5 ms/sWait per clip, after you stop speaking: median 611 ms, p90 655 ms. |
| Whisper Large V3 Q5_0 | 116.7 ms/sWait per clip, after you stop speaking: median 842 ms, p90 1.0 s. | 129.7 ms/sWait per clip, after you stop speaking: median 824 ms, p90 971 ms. | 80.0 ms/sWait per clip, after you stop speaking: median 965 ms, p90 1.1 s. | 89.3 ms/sWait per clip, after you stop speaking: median 1.0 s, p90 1.2 s. | 90.3 ms/sWait per clip, after you stop speaking: median 1.1 s, p90 1.3 s. |
| Whisper Large V3 | 181.7 ms/sWait per clip, after you stop speaking: median 1.3 s, p90 1.5 s. | 202.7 ms/sWait per clip, after you stop speaking: median 1.3 s, p90 1.5 s. | 121.9 ms/sWait per clip, after you stop speaking: median 1.5 s, p90 1.7 s. | 135.0 ms/sWait per clip, after you stop speaking: median 1.5 s, p90 1.7 s. | 135.5 ms/sWait per clip, after you stop speaking: median 1.6 s, p90 1.9 s. |
| Whisper Large V3 Turbo | 104.8 ms/sWait per clip, after you stop speaking: median 771 ms, p90 820 ms. | 118.8 ms/sWait per clip, after you stop speaking: median 770 ms, p90 820 ms. | 67.2 ms/sWait per clip, after you stop speaking: median 819 ms, p90 864 ms. | 72.9 ms/sWait per clip, after you stop speaking: median 826 ms, p90 876 ms. | 70.6 ms/sWait per clip, after you stop speaking: median 856 ms, p90 917 ms. |
| Parakeet TDT 0.6B v3 | 11.4 ms/sWait per clip, after you stop speaking: median 61 ms, p90 174 ms. | 12.0 ms/sWait per clip, after you stop speaking: median 59 ms, p90 166 ms. | 10.4 ms/sWait per clip, after you stop speaking: median 127 ms, p90 194 ms. | 10.9 ms/sWait per clip, after you stop speaking: median 123 ms, p90 184 ms. | 10.7 ms/sWait per clip, after you stop speaking: median 130 ms, p90 197 ms. |
| Hviske v5 Tiny Q5_0 | Targets Danish only. | Targets Danish only. | Targets Danish only. | 18.5 ms/sWait per clip, after you stop speaking: median 206 ms, p90 248 ms. | Targets Danish only. |
Wispr Flow 1.6.897 vs Codictate, September 2026
Response times are not measured the same way for both products: Codictate is timed at the direct adapter call boundary, Wispr Flow is timed from the UI-observed paste.
Wispr Flow streams while you speak, and is measured over the network.
Speed figures from runs before 4 September 2026 are not shown. Full method.
All Codictate models
The 19 models worth running day to day. Ratings for a quick overview, exact percentages if you want the detail.
Ratings run 1 to 10 across every benchmark language. Point at any cell for the figure behind it.
| Whisper Large V3 Turbo Q5_0 | 574 MB | 778 MB | All languages | No | 8/10 | 10/1093.5% word accuracy, pooled across every language it measured. |
|---|---|---|---|---|---|---|
| Whisper Large V3 Q5_0 | 1.1 GB | 1.5 GB | All languages | Yes | 7/10 | 10/1094.0% word accuracy, pooled across every language it measured. |
| Whisper Large V2 | 3.0 GB | 3.5 GB | All languages | Yes | 6/10 | 10/1092.9% word accuracy, pooled across every language it measured. |
| Whisper Large V2 Q5_0 | 1.1 GB | 1.5 GB | All languages | Yes | 7/10 | 10/1092.8% word accuracy, pooled across every language it measured. |
| Whisper Large V2 Q8_0 | 1.6 GB | 2.1 GB | All languages | Yes | 7/10 | 10/1092.8% word accuracy, pooled across every language it measured. |
| Whisper Large V3 | 3.0 GB | 3.5 GB | All languages | Yes | 6/10 | 10/1094.0% word accuracy, pooled across every language it measured. |
| Whisper Large V3 Turbo | 1.6 GB | 1.9 GB | All languages | No | 8/10 | 10/1093.6% word accuracy, pooled across every language it measured. |
| Whisper Large V3 Turbo Q8_0 | 834 MB | 1.1 GB | All languages | No | 8/10 | 10/1093.6% word accuracy, pooled across every language it measured. |
| Parakeet TDT 0.6B v3 | 500 MB | 80 MB | 25 languages | No | 10/10 | 9/1092.0% word accuracy, pooled across every language it measured. |
| Hviske v5 Tiny F16 | 503 MB | 647 MB | Danish only | No | 9/10 | 9/10da88.6% word accuracy on Danish, the only language this model targets. |
| Hviske v5 Tiny Q8_0 | 268 MB | 413 MB | Danish only | No | 10/10 | 9/10da88.6% word accuracy on Danish, the only language this model targets. |
| Hviske v5 Tiny Q6_K | 232 MB | 377 MB | Danish only | No | 10/10 | 9/10da88.6% word accuracy on Danish, the only language this model targets. |
| Hviske v5 Tiny Q5_0 | 181 MB | 327 MB | Danish only | No | 10/10 | 9/10da88.7% word accuracy on Danish, the only language this model targets. |
| Hviske v5 Tiny Q4_K | 153 MB | 298 MB | Danish only | No | 10/10 | 9/10da88.6% word accuracy on Danish, the only language this model targets. |
| Whisper Large V1 | 3.0 GB | 3.5 GB | All languages | Yes | 6/10 | 9/1091.9% word accuracy, pooled across every language it measured. |
| Whisper Medium | 1.5 GB | 1.9 GB | All languages | Yes | 8/10 | 9/1090.8% word accuracy, pooled across every language it measured. |
| Whisper Medium Q5_0 | 514 MB | 878 MB | All languages | Yes | 8/10 | 9/1090.8% word accuracy, pooled across every language it measured. |
| Whisper Medium Q8_0 | 766 MB | 1.2 GB | All languages | Yes | 8/10 | 9/1090.8% word accuracy, pooled across every language it measured. |
| Whisper Small Q5_1 | 181 MB | 383 MB | All languages | Yes | 9/10 | 7/1080.7% word accuracy, pooled across every language it measured. |
| The model Codictate runs, or the cloud product it is compared against. | What the model weighs once downloaded. Smaller is better. | Peak memory while the model ran, averaged over its languages. Smaller is better. | Milliseconds of waiting for every second of audio you dictate. Lower is better. Response times are not measured the same way for both products: Codictate is timed at the direct adapter call boundary, Wispr Flow is timed from the UI-observed paste. | English (clean): the share of words transcribed correctly. | English (noisy): the share of words transcribed correctly. | Spanish: the share of words transcribed correctly. | Danish: the share of words transcribed correctly. | Hungarian: the share of words transcribed correctly. | Word accuracy over every language in the row, pooled by reference words. |
|---|---|---|---|---|---|---|---|---|---|
| Whisper Large V3 | 3.0 GB | 3.5 GB | 164 ms/s | 97.0% | 94.9% | 97.1% | 87.6% | 85.5% | 94.0%Every word from every language counted together, not an average of per-language scores. |
| Whisper Large V3 Q5_0 | 1.1 GB | 1.5 GB | 106 ms/s | 97.1% | 94.9% | 97.1% | 87.4% | 85.6% | 94.0%Every word from every language counted together, not an average of per-language scores. |
| Whisper Large V3 Turbo Q8_0 | 834 MB | 1.1 GB | 73 ms/s | 97.0% | 94.8% | 96.7% | 86.5% | 83.4% | 93.6%Every word from every language counted together, not an average of per-language scores. |
| Whisper Large V3 Turbo | 1.6 GB | 1.9 GB | 93 ms/s | 97.0% | 94.7% | 96.7% | 86.5% | 83.4% | 93.6%Every word from every language counted together, not an average of per-language scores. |
| Whisper Large V3 Turbo Q5_0 | 574 MB | 778 MB | 66 ms/s | 97.0% | 94.8% | 96.7% | 86.1% | 83.0% | 93.5%Every word from every language counted together, not an average of per-language scores. |
| Wispr Flow 1.6.897cloud product | No disk file: it is a cloud product. | Memory of a cloud product cannot be observed. | 88 ms/s | 96.0% | 93.3% | 95.8% | 93.8% | 77.9% | 93.0%Every word from every language counted together, not an average of per-language scores. |
| Whisper Large V2 | 3.0 GB | 3.5 GB | 164 ms/s | 96.7% | 94.0% | 96.7% | 85.3% | 81.0% | 92.9%Every word from every language counted together, not an average of per-language scores. |
| Whisper Large V2 Q8_0 | 1.6 GB | 2.1 GB | 123 ms/s | 96.7% | 94.0% | 96.7% | 85.3% | 81.0% | 92.8%Every word from every language counted together, not an average of per-language scores. |
| Whisper Large V2 Q5_0 | 1.1 GB | 1.5 GB | 106 ms/s | 96.7% | 94.0% | 96.7% | 85.0% | 80.9% | 92.8%Every word from every language counted together, not an average of per-language scores. |
| Parakeet TDT 0.6B v3 | 500 MB | 80 MB | 11 ms/s | 96.2% | 94.1% | 95.1% | 80.4% | 81.6% | 92.0%Every word from every language counted together, not an average of per-language scores. |
| Whisper Large V1 | 3.0 GB | 3.5 GB | 165 ms/s | 96.5% | 93.6% | 96.3% | 82.1% | 77.3% | 91.9%Every word from every language counted together, not an average of per-language scores. |
| Whisper Medium | 1.5 GB | 1.9 GB | 93 ms/s | 96.4% | 93.2% | 96.1% | 78.8% | 73.2% | 90.8%Every word from every language counted together, not an average of per-language scores. |
| Whisper Medium Q8_0 | 766 MB | 1.2 GB | 72 ms/s | 96.4% | 93.2% | 96.1% | 78.7% | 73.1% | 90.8%Every word from every language counted together, not an average of per-language scores. |
| Whisper Medium Q5_0 | 514 MB | 878 MB | 65 ms/s | 96.4% | 93.2% | 96.1% | 78.5% | 72.8% | 90.8%Every word from every language counted together, not an average of per-language scores. |
| Whisper Small Q5_1 | 181 MB | 383 MB | This row's runs never recorded how long the wait was. | 95.4% | 91.1% | 93.8% | 62.5% | 56.3% | 80.7%Every word from every language counted together, not an average of per-language scores. |
| Hviske v5 Tiny Q5_0 | 181 MB | 327 MB | 18 ms/s | Targets Danish only. | Targets Danish only. | Targets Danish only. | 88.7% | Targets Danish only. | 88.7%Pooled over Danish only. |
| Hviske v5 Tiny F16 | 503 MB | 647 MB | 22 ms/s | Targets Danish only. | Targets Danish only. | Targets Danish only. | 88.6% | Targets Danish only. | 88.6%Pooled over Danish only. |
| Hviske v5 Tiny Q4_K | 153 MB | 298 MB | 18 ms/s | Targets Danish only. | Targets Danish only. | Targets Danish only. | 88.6% | Targets Danish only. | 88.6%Pooled over Danish only. |
| Hviske v5 Tiny Q8_0 | 268 MB | 413 MB | 19 ms/s | Targets Danish only. | Targets Danish only. | Targets Danish only. | 88.6% | Targets Danish only. | 88.6%Pooled over Danish only. |
| Hviske v5 Tiny Q6_K | 232 MB | 377 MB | 19 ms/s | Targets Danish only. | Targets Danish only. | Targets Danish only. | 88.6% | Targets Danish only. | 88.6%Pooled over Danish only. |
Wispr Flow 1.6.897 vs Codictate, September 2026
Which model should you pick
Five ways to read the table above, depending on what you dictate and what your machine can spare.
Best accuracy across languages
Whisper Large V3 Q5_0 is the most accurate multilingual model we ship, at 94.0% word accuracy counted across every language we test. The full-size Whisper Large V3 scores 94.0% and needs 3.0 GB on disk. The compressed build needs 1.1 GB for the same result, so it is the easier one on your computer.
Fastest
Parakeet TDT 0.6B v3 is the quickest thing we ship and rates 10 out of 10 for speed, while holding 92.0% word accuracy. It covers 25 languages rather than the whole Whisper list and does not translate, so it is the pick when you dictate in a language it supports and want the text back immediately.
Daily driver
If you switch languages or work in noise, Whisper Large V3 Turbo Q5_0 is the safer default: 93.5% word accuracy at 574 MB on disk, and 8 out of 10 for speed. If you dictate English in a quiet room, Parakeet is faster and covers less ground.
Translation
Whisper Large V3 Q5_0 is the accuracy pick above and it also translates. The Turbo builds do not, which the translation column above records, so if you need translation at this accuracy this is the model.
Lightweight
Whisper Small Q5_1 gives the best accuracy for its size: 80.7% word accuracy at 181 MB on disk and 383 MB of peak memory, across every language, translation included. If your machine is tight on memory, this is the one.
All benchmark runs
Each run records the dataset, the sample count and the hardware with its results.
- Sep 22, 2026
2026 09 v4 english tail
English clips 951 to corpus end, 13 multilingual models, matching the Flow depth
- Models
- 13
- Recordings per language
- 1986
- Hardware
- Apple M4 Max
- Sep 21, 2026
2026 09 v5 english tail
Wispr Flow 1.6.897 external-product benchmark; Flow 1.6.897; multilingual (en, es, da, hu selected); Context Awareness off; clean dictionary; BlackHole 2ch
- Models
- 1
- Recordings per language
- 1986
- Hardware
- Apple M4 Max
- Sep 20, 2026
2026 09 v5 full 1 6 897
Wispr Flow 1.6.897 external-product benchmark; Flow 1.6.897; multilingual (en, es, da, hu selected); Context Awareness off; clean dictionary; BlackHole 2ch
- Models
- 1
- Recordings per language
- 950
- Hardware
- Apple M4 Max
- Sep 18, 2026
2026 09 v4 hviske catchup
Clips 401 to end of da_dk, 5 Danish-pinned hviske models
- Models
- 5
- Recordings per language
- 527
- Hardware
- Apple M4 Max
- Sep 18, 2026
2026 09 v4 multilingual catchup
Clips 401-950, 13 multilingual models, matching the Flow comparison depth
- Models
- 13
- Recordings per language
- 550
- Hardware
- Apple M4 Max
- Sep 17, 2026
2026 09 v4 english tail
Wispr Flow 1.6.886 external-product benchmark; Wispr Flow, English clips 2201 to end of both corpora
- Models
- 1
- Recordings per language
- 736
- Hardware
- Apple M4 Max
- Sep 17, 2026
2026 09 v4 english
Wispr Flow 1.6.872 external-product benchmark; Wispr Flow, English clips 951-2200
- Models
- 1
- Recordings per language
- 1250
- Hardware
- Apple M4 Max
- Sep 16, 2026
Da recheck 1 6 827
Wispr Flow 1.6.872 external-product benchmark; da_dk clips 401-927 again on Flow 1.6.827
- Models
- 1
- Recordings per language
- 527
- Hardware
- Apple M4 Max
- Sep 5, 2026
2026 09 v2 hviske v5 tiny q4 k
benchmark-v2 publication batch 2026-09-v2; codictate hviske-v5-tiny-q4_k; clips [0, 400) of the consumable range; hotkey option+z
- Models
- 1
- Recordings per language
- 400
- Hardware
- Apple M4 Max
- Sep 5, 2026
2026 09 v2 hviske v5 tiny q5 0
benchmark-v2 publication batch 2026-09-v2; codictate hviske-v5-tiny-q5_0; clips [0, 400) of the consumable range; hotkey option+z
- Models
- 1
- Recordings per language
- 400
- Hardware
- Apple M4 Max
- Sep 5, 2026
2026 09 v2 hviske v5 tiny q6 k
benchmark-v2 publication batch 2026-09-v2; codictate hviske-v5-tiny-q6_k; clips [0, 400) of the consumable range; hotkey option+z
- Models
- 1
- Recordings per language
- 400
- Hardware
- Apple M4 Max
- Sep 5, 2026
2026 09 v2 hviske v5 tiny q8 0
benchmark-v2 publication batch 2026-09-v2; codictate hviske-v5-tiny-q8_0; clips [0, 400) of the consumable range; hotkey option+z
- Models
- 1
- Recordings per language
- 400
- Hardware
- Apple M4 Max
- Sep 5, 2026
2026 09 v2 hviske v5 tiny f16
benchmark-v2 publication batch 2026-09-v2; codictate hviske-v5-tiny-f16; clips [0, 400) of the consumable range; hotkey option+z
- Models
- 1
- Recordings per language
- 400
- Hardware
- Apple M4 Max
- Sep 5, 2026
2026 09 v2 medium q8 0
benchmark-v2 publication batch 2026-09-v2; codictate medium-q8_0; clips [0, 400) of the consumable range; hotkey option+z
- Models
- 1
- Recordings per language
- 400
- Hardware
- Apple M4 Max
- Sep 5, 2026
2026 09 v2 medium q5 0
benchmark-v2 publication batch 2026-09-v2; codictate medium-q5_0; clips [0, 400) of the consumable range; hotkey option+z
- Models
- 1
- Recordings per language
- 400
- Hardware
- Apple M4 Max
- Sep 5, 2026
2026 09 v2 medium
benchmark-v2 publication batch 2026-09-v2; codictate medium; clips [0, 400) of the consumable range; hotkey option+z
- Models
- 1
- Recordings per language
- 400
- Hardware
- Apple M4 Max
- Sep 5, 2026
2026 09 v2 large v2 q8 0
benchmark-v2 publication batch 2026-09-v2; codictate large-v2-q8_0; clips [0, 400) of the consumable range; hotkey option+z
- Models
- 1
- Recordings per language
- 400
- Hardware
- Apple M4 Max
- Sep 5, 2026
2026 09 v2 large v2 q5 0
benchmark-v2 publication batch 2026-09-v2; codictate large-v2-q5_0; clips [0, 400) of the consumable range; hotkey option+z
- Models
- 1
- Recordings per language
- 400
- Hardware
- Apple M4 Max
- Sep 5, 2026
2026 09 v2 large v2
benchmark-v2 publication batch 2026-09-v2; codictate large-v2; clips [0, 400) of the consumable range; hotkey option+z
- Models
- 1
- Recordings per language
- 400
- Hardware
- Apple M4 Max
- Sep 5, 2026
2026 09 v2 large v1
benchmark-v2 publication batch 2026-09-v2; codictate large-v1; clips [0, 400) of the consumable range; hotkey option+z
- Models
- 1
- Recordings per language
- 400
- Hardware
- Apple M4 Max
- Sep 5, 2026
2026 09 v2 parakeet tdt 0 6b v3
benchmark-v2 publication batch 2026-09-v2; codictate parakeet-tdt-0.6b-v3; clips [0, 400) of the consumable range; hotkey option+z
- Models
- 1
- Recordings per language
- 400
- Hardware
- Apple M4 Max
- Sep 5, 2026
2026 09 v2 large v3 turbo q8 0
benchmark-v2 publication batch 2026-09-v2; codictate large-v3-turbo-q8_0; clips [0, 400) of the consumable range; hotkey option+z
- Models
- 1
- Recordings per language
- 400
- Hardware
- Apple M4 Max
- Sep 5, 2026
2026 09 v2 large v3 turbo q5 0
benchmark-v2 publication batch 2026-09-v2; codictate large-v3-turbo-q5_0; clips [0, 400) of the consumable range; hotkey option+z
- Models
- 1
- Recordings per language
- 400
- Hardware
- Apple M4 Max
- Sep 5, 2026
2026 09 v2 large v3 turbo
benchmark-v2 publication batch 2026-09-v2; codictate large-v3-turbo; clips [0, 400) of the consumable range; hotkey option+z
- Models
- 1
- Recordings per language
- 400
- Hardware
- Apple M4 Max
- Sep 5, 2026
2026 09 v2 large v3
benchmark-v2 publication batch 2026-09-v2; codictate large-v3; clips [0, 400) of the consumable range; hotkey option+z
- Models
- 1
- Recordings per language
- 400
- Hardware
- Apple M4 Max
- Sep 5, 2026
2026 09 v2 large v3 q5 0
benchmark-v2 publication batch 2026-09-v2; codictate large-v3-q5_0; clips [0, 400) of the consumable range; hotkey option+z
- Models
- 1
- Recordings per language
- 400
- Hardware
- Apple M4 Max
- Sep 5, 2026
2026 09 v2 flow
Wispr Flow 1.6.774 external-product benchmark; benchmark-v2 publication batch 2026-09-v2; wispr-flow wispr-flow; clips [0, 400) of the consumable range; hotkey option+z
- Models
- 1
- Recordings per language
- 400
- Hardware
- Apple M4 Max
- Sep 4, 2026
Hviske danish 400
hviske-v5-tiny-q5_0 on FLEURS Danish at 400 samples to match the Wispr Flow 1.6.765 depth for the published comparison
- Models
- 1
- Recordings per language
- 400
- Hardware
- Apple M4 Max
- Sep 4, 2026
Curated 400 wispr comparison
Curated models at 400 samples per dataset to match the Wispr Flow 1.6.765 external-product run at equal depth for the published comparison
- Models
- 4
- Recordings per language
- 400
- Hardware
- Apple M4 Max
- Aug 18, 2026
Hviske vs main models
This benchmark compares the hviske model to the already main models recommended in Codictate
- Models
- 7
- Recordings per language
- 200
- Hardware
- Apple M4 Max
- Aug 17, 2026
Crispasr vs whisper
to test which asr harnes is best.
- Models
- 6
- Recordings per language
- 20
- Hardware
- Apple M4 Max
- May 9, 2026
Full run except tiny base
Initial model benchmarks to see if there are some models that we should leave out for a more extensive benchmark that will include way more samples. This does not include the Tiny/Base models which has been tested initially already
- Models
- 34
- Recordings per language
- 50
- Hardware
- Apple M4 Max
- May 9, 2026
Tiny base triage
Triage of Tiny and Base model families (all quantization and language variants) to evaluate whether these smaller models are worth further benchmarking.
- Models
- 12
- Recordings per language
- 50
- Hardware
- Apple M4 Max
- May 8, 2026Featured
Main model comparison
Comparison of Codictate's curated speech models (Small q5_1, Large V3 Turbo q5_0, Large V3 q5_0, Parakeet) to determine the best-performing default model.
- Models
- 4
- Recordings per language
- 200
- Hardware
- Apple M4 Max