Typd

Voice dictation for Windows · Private beta

One small model on your computer hears you and writes what you meant.

Tap Right Alt, speak, then tap again. Typd pastes the finished text into whatever app you are using.

Fully local
Your voice never leaves the computer. No account. No internet needed to dictate.
One model
0.78 billion parameters. It hears and writes in one step.
Fast
Text about 0.2 s after you stop, for an 8-second dictation on a laptop GPU.

You say

hi sam um i was going to can you send me the file from tuesday actually wednesday thanks

Typd writes

Hi Sam,

Can you send me the file from Wednesday?

Thanks

An illustration, not a recorded sample. Struck words are what the model leaves out. Measured accuracy is further down the page.

How it works

  1. Tap Right Alt

    In any app: email, chat, documents, code editors or an AI chat.

  2. Speak, then tap again

    Talk the way you would to a person. Pause, correct yourself, start a sentence over.

  3. The text is pasted

    One model on your computer writes the finished text and pastes it where your cursor is.

What the model takes care of

Fillers
Removes “um”, “uh” and the like.
Corrections
Resolves self-corrections and restarts to what you meant.
Punctuation
Adds punctuation and formats numbers.
Structure
Lists, paragraphs, and the layout of emails and messages.
Spoken commands
“New paragraph”, “open parenthesis” and similar.
Names
Reads the screen for names, to help spell them.

One model instead of two

AI dictation apps usually chain two models. A speech recognizer writes a raw transcript. Then a large language model, usually in the cloud, cleans it up. Typd does both steps in one small model that runs on your computer.

The usual two-model pipeline compared with Typd's single local model Top row, the usual pipeline: microphone, then a speech recognizer, which passes a transcript to a large language model that is usually in the cloud, then text. Bottom row, Typd, all on your computer: microphone, then one model of 0.78 billion parameters that hears and writes in one step, then finished text. The usual pipeline · two models Microphone Speech recognizer transcript Large language model usually in the cloud Text Typd · one model, all on your computer Microphone One model · 0.78B parameters hears and writes in one step Finished text The usual two-model pipeline compared with Typd's single local model Left column, the usual pipeline: microphone, then a speech recognizer, which passes a transcript to a large language model that is usually in the cloud, then text. Right column, Typd, all on your computer: microphone, then one model of 0.78 billion parameters that hears and writes in one step, then finished text. Usual · two models Typd · one model Microphone Speech recognizer transcript Large language model usually in the cloud Text Microphone One model 0.78B parameters hears and writes Finished text All on your computer
Dashed: the step that usually runs in the cloud. In Typd nothing leaves the computer. There is no account, and dictation needs no internet.

The model

Base
Qwen3-ASR-0.6B, open
Training
Fine-tuned for dictation
Parts
Audio encoder ~0.18B + 0.6B language model
Runtime
llama.cpp, open source
Weights
397 MB, 4-bit
Audio projector
214 MB

Measured results

I compare Typd with a cloud pipeline I built for reference. It uses Whisper for speech, then a frontier language model for cleanup. The cloud reference is still ahead on every test, by 0.2 to 1.8 percentage points. Typd runs on a laptop with a consumer 4 GB GPU.

Word error rate: the share of words that come out wrong. Lower is better.
Test Typdone local model Cloud referenceWhisper + frontier LLM
Dictation320 scripted clips, synthetic voices1.5 %1.1 %
Conversational speechReal speech, DisfluencySpeech5.4 %4.8 %
MeetingsReal meetings, AMI8.0 %7.1 %
Earnings callsReal calls, Earnings-229.5 %7.7 %

The cloud reference is a two-model pipeline I built for this comparison. It is not a commercial product, and I have not tested Typd head to head against commercial dictation apps.

Scores compare the output with cleaned-up reference text written to one style guide, not with word-for-word transcripts. They cannot be compared with public speech-recognition benchmarks. The references were written by the same family of model that does the cloud reference’s cleanup, which favours the cloud reference.

Speed and memory, on that laptop

8-second dictationGPU, time after you stop
~0.2 s
60-second dictationGPU
~1.5 s
8-second dictationCPU only
~0.9 s
GPU memoryWhole app: idle, and at most during long dictations
0.84–1.0 GB

Formatting, on test sets not used in training

Spoken commandsFully right
96 %
Emails and messagesLaid out right
75 %
Lists and paragraphsLaid out right
82 %

How it was built

Typd starts from the open Qwen3-ASR-0.6B model, fine-tuned for dictation.

One style guide, checked by scripts

A frontier language model wrote the example outputs Typd learns from, all under one written style guide. Scripts check them against what was said, so the words are not paraphrased.

Real speech, relabelled

About 2,000 passages of real, spontaneous speech were relabelled with paragraphs. They are CC-BY YouTube speech from the YODAS-Granary dataset. The rest of the data is synthetic speech for lists, emails, commands and restarts, real meetings and conversations, and my own dictations.

Natural speech needed new data

The model added paragraphs and lists to scripted speech, but rarely to natural speech. Converting voices, adding noise and changing the pacing ruled out the voice and the microphone. So later training rounds changed the data and kept the model the same size.

Sentence restarts

People often start a sentence, stop, and say it another way. An October 2026 training run targeted this. On an internal test set, its error rate on these restarts fell from 18.7 % to 2.6 %. That run is not yet the model in the app.

Averaging training runs

Averaging the weights of separate training runs, a method known as a model soup, improved accuracy without making the model bigger.

Long dictations kept whole

Long dictations are split at pauses. In a 163-second test dictation, the old version lost 42 % of the text. The new one keeps all of it.

Where it is going

Now

  • A working Windows app. Beta builds on request.
  • One local model, under 1 GB of GPU memory for the whole app.

Next

  • More languages, starting with Turkish.
  • Better rare names.
  • A Mac version.

Goal

  • Match, then beat, the leading cloud dictation app on accuracy and formatting, running locally.
  • Good on laptops without a GPU.

Private beta

Private beta builds on request — via mehmetdedeler.com.

Typd is built by one person, Mehmet Dedeler. That includes the model, the training data and the Windows app.