Blog
Wispr shipped a new speech model the day I finished testing the old one. So I ran all 8,287 recordings again.
7 min read
Emil Lykke Grann
On 17 September my last English test run finished. By then Flow had heard every English recording in our test set. That is 2,617 clean and 2,936 noisy recordings. They all come from LibriSpeech, a collection of English audiobooks.
The same day, an email from Wispr arrived. Flow now runs on Canto, "the first speech model we've built ourselves". It was "rolled out for everyone, with nothing to switch on".
So every number I had was for a model Wispr had just replaced. Great timing. I updated Flow to version 1.6.897, the first one after the switch. Then I ran every recording again. That is 8,287 recordings in total, or 19.8 hours of audio. It took two nights.
TL;DR
- Canto barely changed accuracy. Clean English went from 4.10% to 4.00% of words wrong, and noisy English from 6.60% to 6.65%. Both changes are within the noise.
- Flow got 17 to 26 percent faster in every language, including the three Canto doesn't cover.
- On clean English, Whisper Large V3 Q5_0 running on a Mac still makes about a quarter fewer mistakes than Flow (2.92% against 4.00% wrong).
- Canto is English only. Flow still beats Codictate in Danish (Flow: 93.81%, Codictate: 88.69%).
Before and after, on the same recordings
Every recording below ran twice, once before Canto and once after. The table shows the share of words Flow got wrong, so lower is better. The range shows how far the change could move with a different set of recordings.
| LibriSpeech | Recordings | Words spoken | Before Canto | After Canto | Change | Range |
|---|---|---|---|---|---|---|
| Clean | 2,617 | 52,543 | 4.10% | 4.00% | -0.10 points | -0.23 to +0.02 |
| Noisy | 2,936 | 52,311 | 6.60% | 6.65% | +0.06 points | -0.11 to +0.21 |
Both ranges include zero, so we can't tell before and after apart. If you use Flow in English, you won't notice a difference in accuracy.
The "before" numbers come from four Flow versions, because Flow updated itself while I was testing. Each of those four pieces moved by less than 0.3 points. That is about how much Flow moves between two ordinary updates anyway.
Canto only handles English. Wispr says so on the Canto page. So Spanish, Danish and Hungarian work as a check: if they moved too, something other than Canto changed. Here is the share of words wrong on the first 400 recordings per language, from September to now:
- Spanish: 3.95% to 3.94%
- Danish: 5.45% to 6.30%
- Hungarian: 23.30% to 21.74%
Danish and Hungarian jump around between Flow versions, with or without Canto. On one version Danish got 10.08% of words wrong. On the next it was back to 6.40%. So if Flow suddenly feels worse in your language one week, it might not be you.
What did change: Flow is faster in every language
The table shows how long you wait for the text, per second of audio you spoke. It uses the same 400 recordings per language. Lower is faster.
| Wait in ms per second of audio | English | English noisy | Spanish | Danish | Hungarian |
|---|---|---|---|---|---|
| September, version 1.6.774 | 73.3 | 74.5 | 50.3 | 213.3 | 212.0 |
| Now, version 1.6.897 | 56.5 | 62.1 | 40.3 | 166.1 | 157.1 |
Flow is 17 to 26 percent faster in every language. That includes Spanish, Danish and Hungarian, which Canto doesn't handle. So the speedup isn't Canto. Something else in Flow got quicker at the same time. So for Flow users, less waiting is the real change in this update, not better accuracy.
Wispr now tests on the same public sets we use
A while ago I made a video calling Wispr out for not publishing any benchmarks. In it I show Wispr saying that standard speech recognition benchmarks wouldn't make sense for their models. With Canto, they've used them anyway.
Luckily, the two collections on the Canto page, LibriSpeech and FLEURS, are exactly the ones this whole test uses. LibriSpeech is our English, and FLEURS is our Spanish, Danish and Hungarian. So we don't have to take Wispr's LibriSpeech claim on trust. We can check it on the same collection they named.
What Wispr claims, and what we can check
Wispr's Canto page says Canto "tied for the lowest WER on LibriSpeech". It also says Canto "remained competitive on FLEURS and Common Voice, although it did not lead either evaluation". The headline result is this chart:

Source: wisprflow.ai/canto
That 3.4% is on ten hours of real Flow dictations from over 2,300 speakers. Only Wispr has those recordings, so nobody else can check it. Notice also that every model in the chart is a cloud service. There's no open model you can run yourself, like Whisper.
Here's what we can check. We can't test Canto on its own, only Flow, the app: audio goes in, and text comes out in a window. The first row below is Wispr's test and the other two are ours, on different audio. So only compare the last two with each other.
| Audio | Words wrong | |
|---|---|---|
| Canto, Wispr's test | 10 hours of Flow dictations, only Wispr has them | 3.4% |
| Flow on Canto, our test | LibriSpeech clean, 2,617 recordings | 4.00% |
| Whisper Large V3 Q5_0 in Codictate, our test | LibriSpeech clean, 2,617 recordings | 2.92% |
On the same audio, the free model on your Mac makes 1,536 mistakes against Flow's 2,102. So one of two things is true. Either Flow adds mistakes after the model, or Canto's LibriSpeech "tie" is at a level a free 1.5 GB model already beats. Either way, you don't use the model. You use the app, and that's what we measured.
Wispr's older claims are still online
A Wispr blog post from December 2025 says Flow works "across every platform at 95%+ accuracy". The homepage says "100+ Languages".
On Canto, Flow gets 96.00% of words right in clean English. So in English, the 95% claim holds. In Hungarian it gets 77.94%, which is nowhere near it. And Canto, the model built to improve accuracy, only covers English. That's one language out of the hundred.
Where Flow is still ahead: Danish
We have 927 Danish recordings, from FLEURS, a collection of people reading sentences aloud. Flow gets 93.81% of the words right. The best Codictate model for Danish, Hviske v5 Tiny, gets 88.69%.
That has nothing to do with Canto. Flow was already ahead in Danish in September. It's the one language of the five where the cloud beats the laptop, by five points. If you dictate in Danish, that's a real reason to pick Flow today, and the gap we'd most like to close.
What this benchmark does not tell you
Wispr didn't say which Flow version got Canto. So "before" and "after" here are dates and version numbers. We couldn't see the switch itself. My last "before" run was on the evening of 17 September, the day Canto launched. Some of its 1,153 recordings may already have used Canto. They changed no more than the rest, so the result doesn't depend on them.
Everything from the full test post applies too. The audio is people reading, not dictating. It goes into a virtual microphone, not a room. Flow had Context Awareness off and an empty dictionary.
Canto was built for noisy, real-world speech. Noisy LibriSpeech is the closest we have to that, and it got 0.06 points worse. If Canto is better on real dictation, this test can't show it. Wispr hasn't published a test anyone else can run.
FAQ
Did Canto make Wispr Flow more accurate? Not on read English. Clean English went from 4.10% to 4.00% of words wrong. Noisy English went from 6.60% to 6.65%. That is over the same 5,553 recordings, and both changes are within the noise.
Is Wispr Flow faster after Canto? Yes, 17 to 26 percent faster in every language. That includes three languages Canto doesn't cover. Clean English went from 73.3 to 56.5 ms of waiting per second of audio.
Is Canto better than Whisper? We can only test Flow, not Canto on its own. On 2,617 clean English recordings, Flow got 96.00% of words right. Whisper Large V3 Q5_0, running locally, got 97.08%.
Reproduce it or argue with it
Every run, before and after, is at github.com/EmilLykke/dictation-benchmark. Each run records its Flow version.
The full results are in the full test post. The comparison with Codictate is on the benchmarks page. The first post, from September, is here.