One style guide, checked by scripts
A frontier language model wrote the example outputs Typd learns from, all under one written style guide. Scripts check them against what was said, so the words are not paraphrased.
Voice dictation for Windows · Private beta
Tap Right Alt, speak, then tap again. Typd pastes the finished text into whatever app you are using.
You say
hi sam um i was going to can you send me the file from tuesday actually wednesday thanks
Typd writes
Hi Sam,
Can you send me the file from Wednesday?
Thanks
In any app: email, chat, documents, code editors or an AI chat.
Talk the way you would to a person. Pause, correct yourself, start a sentence over.
One model on your computer writes the finished text and pastes it where your cursor is.
What the model takes care of
AI dictation apps usually chain two models. A speech recognizer writes a raw transcript. Then a large language model, usually in the cloud, cleans it up. Typd does both steps in one small model that runs on your computer.
The model
I compare Typd with a cloud pipeline I built for reference. It uses Whisper for speech, then a frontier language model for cleanup. The cloud reference is still ahead on every test, by 0.2 to 1.8 percentage points. Typd runs on a laptop with a consumer 4 GB GPU.
| Test | Typdone local model | Cloud referenceWhisper + frontier LLM |
|---|---|---|
| Dictation320 scripted clips, synthetic voices | 1.5 % | 1.1 % |
| Conversational speechReal speech, DisfluencySpeech | 5.4 % | 4.8 % |
| MeetingsReal meetings, AMI | 8.0 % | 7.1 % |
| Earnings callsReal calls, Earnings-22 | 9.5 % | 7.7 % |
The cloud reference is a two-model pipeline I built for this comparison. It is not a commercial product, and I have not tested Typd head to head against commercial dictation apps.
Scores compare the output with cleaned-up reference text written to one style guide, not with word-for-word transcripts. They cannot be compared with public speech-recognition benchmarks. The references were written by the same family of model that does the cloud reference’s cleanup, which favours the cloud reference.
Speed and memory, on that laptop
Formatting, on test sets not used in training
Typd starts from the open Qwen3-ASR-0.6B model, fine-tuned for dictation.
A frontier language model wrote the example outputs Typd learns from, all under one written style guide. Scripts check them against what was said, so the words are not paraphrased.
About 2,000 passages of real, spontaneous speech were relabelled with paragraphs. They are CC-BY YouTube speech from the YODAS-Granary dataset. The rest of the data is synthetic speech for lists, emails, commands and restarts, real meetings and conversations, and my own dictations.
The model added paragraphs and lists to scripted speech, but rarely to natural speech. Converting voices, adding noise and changing the pacing ruled out the voice and the microphone. So later training rounds changed the data and kept the model the same size.
People often start a sentence, stop, and say it another way. An October 2026 training run targeted this. On an internal test set, its error rate on these restarts fell from 18.7 % to 2.6 %. That run is not yet the model in the app.
Averaging the weights of separate training runs, a method known as a model soup, improved accuracy without making the model bigger.
Long dictations are split at pauses. In a 163-second test dictation, the old version lost 42 % of the text. The new one keeps all of it.
Private beta builds on request — via mehmetdedeler.com.
Typd is built by one person, Mehmet Dedeler. That includes the model, the training data and the Windows app.