Blog
We ran Wispr Flow on every recording we have. 8,287 of them, 19.8 hours of audio.
7 min read
Emil Lykke Grann
Three weeks ago I tested Wispr Flow on 2,000 recordings, 400 per language. That was enough to rank it. It was not enough to trust the decimals. So I kept going until there was nothing left to run.
This time Flow heard every recording in our test set. That is 8,287 recordings in total, or 19.8 hours of audio. The recordings come from two public collections:
- LibriSpeech is English audiobooks. It has a clean set (2,617 recordings) and a noisier set (2,936).
- FLEURS is people reading sentences aloud. We use Spanish (905), Danish (927) and Hungarian (902).
Every recording ran on the same Flow version, 1.6.897.
TL;DR
- Flow got 93.02% right across all five languages together.
- The best Codictate model scores 94.04% accuracy (Whisper Large V3 Q5_0).
- Flow beats Codictate by 5.12 points in Danish (Flow: 93.81%, Codictate: 88.69%).
- Flow is fast in English and Spanish, but slow in Danish and Hungarian.
Flow's results on every recording
Word accuracy is the share of spoken words Flow typed correctly. Higher is better.
| Collection | Language | Recordings | Words spoken | Word accuracy | Character accuracy |
|---|---|---|---|---|---|
| LibriSpeech clean | English | 2,617 | 52,543 | 96.00% | not scored |
| LibriSpeech noisy | English | 2,936 | 52,311 | 93.35% | not scored |
| FLEURS | Spanish | 905 | 23,175 | 95.81% | 97.08% |
| FLEURS | Danish | 927 | 19,846 | 93.81% | 95.20% |
| FLEURS | Hungarian | 902 | 16,906 | 77.94% | 86.42% |
| All five | 8,287 | 164,781 | 93.02% |
In total, Flow got 11,509 words wrong out of 164,781. Clean English at 96% is good. Hungarian is a different story: more than one word in five comes out wrong, which is a lot of fixing if that's your language.
The ranking from the first post did not change. Two numbers moved by about half a point. Clean English went up from 95.52%, and Danish went down from 94.55%. Both were on an older Flow version back then. So the small test got the big picture right. It just couldn't be trusted to the decimal.
Eight recordings came back empty. Flow did not paste anything within our 45 second limit. Four were Danish and four were Hungarian. We count them as every word wrong. That's rare, but notice it only happened in the two languages where Flow is also slowest.
On the same recordings, Codictate wins four languages and loses Danish
We ran Codictate's models on the exact same 8,287 recordings. Higher is better.
| English | English noisy | Spanish | Danish | Hungarian | All five | |
|---|---|---|---|---|---|---|
| Wispr Flow 1.6.897 | 96.00% | 93.35% | 95.81% | 93.81% | 77.94% | 93.02% |
| Whisper Large V3 | 97.03% | 94.92% | 97.09% | 87.61% | 85.47% | 94.05% |
| Whisper Large V3 Q5_0 | 97.08% | 94.88% | 97.09% | 87.43% | 85.56% | 94.04% |
| Whisper Large V3 Turbo | 96.98% | 94.74% | 96.71% | 86.47% | 83.43% | 93.58% |
| Whisper Large V3 Turbo Q5_0 | 96.96% | 94.76% | 96.69% | 86.11% | 82.96% | 93.48% |
| Parakeet TDT 0.6B v3 | 96.21% | 94.08% | 95.11% | 80.44% | 81.63% | 91.98% |
| Hviske v5 Tiny Q5_0 | 88.69% |
So what does that mean if you're choosing? In English, Spanish and Hungarian, a free model on your own Mac makes fewer mistakes than Flow. The best one is Whisper Large V3 Q5_0, a 1.5 GB download. In English and Spanish the gap is about a point, so you'd barely notice it day to day. Hungarian is not close: 7.62 points, or about one extra wrong word in every 13.
Danish goes the other way, and clearly. Flow beats Hviske v5 Tiny, the best Danish model Codictate ships, by 5.12 points. If you mostly dictate in Danish, Flow is the more accurate choice today.
The one that surprised me most is Parakeet. It's a small model that answers in about 11 ms per second of audio, and it still beats Flow in both English sets. So in English you don't have to pick between fast and accurate.
Speed: Flow got faster, but is still slowest in Danish and Hungarian
The table shows how long you wait for the text, per second of audio you spoke. We time it from when you stop talking to the last change in the pasted text. Lower is faster.
| Wait in ms per second of audio | English | English noisy | Spanish | Danish | Hungarian |
|---|---|---|---|---|---|
| Wispr Flow 1.6.897 | 60.2 | 65.1 | 40.8 | 158.7 | 155.5 |
| Whisper Large V3 Q5_0 | 116.7 | 129.7 | 80.0 | 89.3 | 90.3 |
| Whisper Large V3 Turbo Q5_0 | 75.2 | 85.0 | 47.9 | 52.3 | 50.5 |
| Parakeet TDT 0.6B v3 | 11.4 | 12.0 | 10.4 | 10.9 | 10.7 |
Good news if you use Flow: it got faster since September. September's test only had 400 recordings per language, so for a fair comparison we looked at just those 400 again. That's why these numbers differ a little from the table, which covers every recording.
One caveat. We can't time the two products at the same point. Codictate is timed when the speech model is called. Flow is timed from the keypress to the text appearing. So treat any difference smaller than two times as noise.
How 29 hours of testing happened
Flow has no API I can use. It only listens to a microphone. So we play each recording into BlackHole, a virtual microphone, and point Flow at it. A simulated Option+Z keypress starts and stops Flow. Flow pastes its text into a text box. Once the text has not changed for 750 ms, we take it as final and score it with the same code we use for Codictate.
Everything plays in real time. So 19.8 hours of audio takes at least 19.8 hours to run, plus the wait for Flow after each recording. The test ran in two sessions:
- The first session covered 4,634 recordings and took 18 hours 27 minutes.
- The second covered the remaining 3,653 English recordings and took 10 hours 51 minutes.
The Mac can't be used in the meantime, because the text box has to stay focused. So you can imagine how long that felt for me...
Two things made runs this long possible. First, every result is saved the moment it is scored, so a crash only loses one recording. Second, we track each recording by its audio file, not by its sentence. In FLEURS, several speakers read the same sentence. The first version of our setup thought it had tested 400 Danish recordings. It had only heard 264 different files.
Flow also updates itself. That matters, because Danish swings between versions. On the same recordings, version 1.6.827 got 10.08% of Danish words wrong, and 1.6.872 got 6.40% wrong. This whole test stayed on 1.6.897. We checked the version before and after.
The test machine was an Apple M4 Max with 36 GB of RAM, on macOS 26.6.2. In Flow, Context Awareness was off and the dictionary was empty and the languages we were benchmarking were selected in the settings.
What this benchmark does not tell you
We tested five languages. Flow's homepage says it supports "100+". All of the audio we tested is people reading, not dictating. LibriSpeech is audiobooks, and FLEURS is people reading sentences from Wikipedia. Nobody changes their mind halfway through a Slack message.
The audio goes straight into a virtual microphone. There is no room, no fan and no earbuds. That makes things easier for both products equally. It is fine for comparing them, but it won't predict what you get at your desk.
Flow had Context Awareness off and an empty dictionary. A Flow that has learned your words may do better. For English we only score words, not characters. And the speed numbers are timed at different points for the two products.
FAQ
How accurate is Wispr Flow? Flow 1.6.897 got 93.02% of words right across all 8,287 recordings in five languages. Clean English was 96.00%. Noisy English was 93.35%, Spanish 95.81%, Danish 93.81% and Hungarian 77.94%.
Is Wispr Flow more accurate than Whisper and other local models? Not in our test. On the same 2,617 clean English recordings, Whisper Large V3 Q5_0 in Codictate got 97.08% right. Flow got 96.00%. In Hungarian it was 85.56% for Codictate against 77.94% for Flow. Flow did beat every local model in Danish, 93.81% against 88.69%.
Does the 400-recording test still hold? The ranking does. Flow still wins Danish and loses the other four. Clean English and Danish each moved about half a point, on a newer Flow version.
Reproduce it or argue with it
The test setup, every run and every transcript are at github.com/EmilLykke/dictation-benchmark. You need a Mac, BlackHole 2ch and a Wispr Flow account. You also need about 30 hours of not touching the computer.
The Codictate results are on the benchmarks page. The first post is here.