Blog

Edda v0.2 is now Codictate's most accurate Danish model

3 min readEmil Lykke Grann

Codictate has a new Danish model, Edda v0.2. It's the most accurate Danish model we've measured in the app.

It makes about a third fewer mistakes than hviske, our previous best for Danish. The price is a bigger download and a slower answer.

Edda is a Danish version of Whisper Large V3 Turbo

Edda v0.2 started as OpenAI's Whisper Large V3 Turbo. The Alexandra Institute trained it further on Danish as part of the CoRal project. Danish Foundation Models released it on Hugging Face under the Apache 2.0 license.

It only does Danish. That's the same bet hviske made: give up every other language to get better at one.

Edda gets fewer Danish words wrong than any model we've measured in the app

Every model below heard the same 927 Danish recordings from FLEURS, a public set of people reading sentences aloud. "Words wrong" is the share of words a model missed, swapped or invented. "Wait" is the time a model takes for each second of speech. Lower is better in every column.

ModelDownloadMemory while runningWords wrongWait per second of audio
Edda v0.2 full1.5 GB1.9 GB7.51%74 ms
Edda v0.2 Q8834 MB1.1 GB7.52%57 ms
Edda v0.2 Q5547 MB787 MB7.53%53 ms
Hviske v5 Tiny Q5181 MB327 MB11.31%18 ms
Whisper Large V3 Turbo Q5574 MB776 MB13.89%52 ms

Edda Q5 is almost exactly Large V3 Turbo's size and speed. That makes sense, since Turbo is what it was built from. The difference is all in the Danish.

Against hviske it's a real trade. Edda Q5 is three times the download and about three times slower. In return, it gets one word in 13 wrong instead of one in 9.

All three Edda sizes score the same, so take Q5

The smaller versions store the same model with less precision, which saves space. Over 19,846 spoken words, the full version and Q5 are four mistakes apart. That's 0.02 percentage points.

For those four mistakes, Q5 is about a third of the download and needs less than half the memory. It's also faster, at 53 ms per second of speech instead of 74. So Q5 is the one I'd pick: it's the leanest version and gives up practically no accuracy.

Getting Edda into the app took a conversion and two workarounds

Edda ships in a file format that Codictate's speech engine can't read. So I converted it to one the engine can read, and we host that copy ourselves.

Two snags came up on the way. The converter wanted two tokenizer files that Edda doesn't include. Those hold the list of word pieces the model writes with. I rebuilt them from Edda's own tokenizer, and they came out identical to Whisper Turbo's. The ready-made tool for making smaller versions only reads a newer format, so I built the older tool from source.

Every download is checked against a fingerprint before it's installed. So the file you install is the one we measured. Someone had already made a community conversion for Fennec, a Linux dictation app. Their notes on feeding Edda the previous sentence for better punctuation were a useful read. The full steps are in the repo.

Turn it on in Settings, under Browse more models

Open Settings and click Browse more models. Download Edda V0.2 Q5, select it, and dictate in Danish.

Because Edda only does Danish, the app locks the language to Danish and turns Translate to English off while it's selected.

The numbers come from our benchmark on an Apple M4 Max. FLEURS is read speech, so treat them as a ranking between models rather than a promise for your own voice. Every result, and how we measure it, is on the benchmarks page. Our other Danish model has its own post.