Voice input

Dictation on Windows – offline, straight into any program

You think faster than you type. Press the hotkey, speak, press it again — the text lands exactly where your cursor is: in Outlook, Word, the browser, your chat window. Speech recognition runs with Whisper or Parakeet on your own machine. Neither the recording nor the text is uploaded.

No account required · Free in your system language

Windows 10 and later · No admin rights needed · No account

Typing is the bottleneck

You know exactly what the email should say. You just need twelve minutes to type it.

Dictation apps exist. Most of them come with three catches: they send your voice to a server, they charge monthly, and they are a separate application — you dictate into it, copy the text, and paste it somewhere else.

Celerando has no window to speak into. The text appears directly in whatever program you are already working in.

Three steps, and no fourth

01
Press the hotkey

Anywhere in Windows, no matter which program is in front.

02
Speak

Normally. No commands, no pauses in particular places.

03
Press it again

The text is where your cursor was. No copying, no pasting, no switching windows. Prefer hold-and-release? Starting with Business, push-to-talk is available too.

Two engines, 99 languages — all on your PC

Celerando recognises speech with two open-source engines. Both run locally, both are downloaded by the app straight from Hugging Face:

Whisper (OpenAI)

99 languages, five model sizes from Tiny (75 MB) to Large v3 Turbo (1.6 GB). Runs on the CPU, via CUDA on NVIDIA cards, via Vulkan on any other graphics card. The engine of every edition.

Parakeet TDT 0.6B v3 (NVIDIA NeMo)

25 European languages with automatic language detection, very low latency, streaming-capable. Runs on any DirectX 12 graphics card — AMD, Intel or NVIDIA. From Business.

What each edition recognises

Free: your system language with Whisper Base. From Pro: all 99 Whisper languages, free choice of model, GPU acceleration. From Business: automatic language detection and Parakeet.

  • German
  • English
  • Spanish
  • French
  • Italian
  • Portuguese
  • Dutch
  • Polish
  • Czech
  • Danish
  • Swedish
  • Norwegian
  • Finnish
  • Hungarian
  • Romanian
  • Greek
  • Turkish
  • Russian
  • Ukrainian
  • Arabic
  • Hebrew
  • Hindi
  • Chinese
  • Japanese
  • Korean
  • Vietnamese
  • Thai
  • Indonesian

… and 71 more. The language is set per dictation or — from Business — detected automatically.

The models: download once, then offline

The models are downloaded once from inside the app and then live in your user profile — no admin rights. Whisper Base is the default and is fetched when you enable voice input; you pick every other model in the settings.

After that, speech recognition needs no connection at all. Train, plane, basement, isolated company network — dictation keeps working.

ModelDownloadBest for
Whisper Tiny 75 MBVery quick notes, lowest accuracy
Whisper Base (default) 141 MBSolid all-rounder, runs on any PC without a graphics card
Whisper Small 466 MBGood accuracy, still fine on CPU
Whisper Medium / Medium q5 1.5 GB / 539 MBHigh accuracy, GPU recommended — q5 is the variant for machines with 8 GB RAM
Whisper Large v3 Turbo / Turbo q5 1.6 GB / 574 MBBest accuracy, GPU strongly recommended
Parakeet TDT 0.6B v3 2.5 GBLowest latency, 25 languages detected automatically — from Business

The optional CUDA runtime for NVIDIA cards (about 70 MB) comes from downloads.menterium.net once you select it in the settings. Without it Celerando uses Vulkan — not an error, just a different route to the same graphics card.

Will it run on your machine?

Three tiers, from office laptop to workstation — and none of them strictly needs a graphics card:

Minimum

Windows 10 (build 19041) or Windows 11, 4 GB RAM, any x86-64 CPU. Whisper Tiny and Base run purely on the CPU — even on an older laptop.

Recommended

CPU with AVX2 (Intel from 2013, AMD from 2015), 8 GB RAM, DirectX 12 graphics from Intel, AMD or NVIDIA. Everyday dictation with Small or Medium q5; Parakeet already runs here.

Performance

16 GB RAM, NVIDIA card with 4 GB or more of graphics memory (CUDA) or a modern GPU with 6 GB or more. Large v3 Turbo, long dictations, Smart Modes.

The Voice input – requirements page has a self-check: three clicks and you know which engine and model fit your machine.

“Menterium” is not a word. To Celerando it is.

Every speech recogniser trips over the same things: personal names, company names, product names, the vocabulary of your profession. You dictate your own surname and get back something that merely sounds similar.

In the Personal Dictionary you enter those words once — together with the variants the recogniser wrongly produces. From then on it sticks.

It works in both directions: the terms are handed to the recogniser before transcription, and wrong spellings are corrected afterwards.

Included from Pro.

From dictation to finished text

After recognition the text passes through three stages, all local, all optional:

Rules

Spoken symbols become characters (“euro” → €), abbreviations are expanded or shortened, your own replacements via regular expressions. From Pro.

Smart Modes

A local language model tidies up: filler words out, punctuation in, soften the tone or phrase it as an email. Usually needs a DirectX 12 graphics card with enough memory. From Business.

Streaming

The text appears while you are still speaking — piece by piece, directly in the target program. From Business, as is push-to-talk; the two are separate settings.

Your voice never leaves your machine

Cloud dictation means your recording is transferred to somebody else's server, processed there, and the text is sent back.

Anyone dictating client conversations, patient data, personnel files or draft contracts is dealing with professional confidentiality and the GDPR. A cloud service often cannot be approved for that.

Celerando processes recording and text on your PC. What the app contacts at all: the download of each model you select, the update check, a single installation ping on the very first launch, and — only if you select it in the settings — the optional CUDA runtime. It carries a random installation ID, a truncated hash of the device fingerprint, and the app and Windows versions — no account, no name, no content. The server does not store the IP address in plain text: a hash of it is deleted after 48 hours, while the approximate location derived from it (country, region, city with an accuracy radius) stays with the record for 24 months. Audio and text are never part of it. Any network monitor will confirm it.

The optional transcription history (from Pro) stores text and recording only in your user profile — off by default.

Celerando and Windows voice typing compared

Windows ships its own voice typing behind Win + H. The difference is not whether it works — it is where.

CelerandoWindows voice typing (Win + H)Cloud dictation apps
Where recognition happens On your PCOn Microsoft's servers (online speech recognition)On the provider's servers
Internet required Never while dictating — only downloads, updates and the first-launch pingFor every dictationFor every dictation
Engine Whisper or Parakeet, model of your choice from ProMicrosoft, not selectableProprietary
Languages 99 (Whisper, all from Pro), 25 auto-detected (Parakeet, from Business)Depends on installed language packDepends on the service
Your own terms Personal Dictionary, rules, Smart Modes—Depends on the service
Price Free (system language), Pro €29.99 onceIncluded in WindowsUsually €7–15 per month

As of September 2026. Windows voice typing is perfectly fine for occasional notes — the only question is whether your dictation is allowed on someone else's server.

When typing is not an option

When your hands stop cooperating

Tenosynovitis, RSI, tendinitis. Dictation is then not a convenience but the condition for continuing to work at all.

When you write in a foreign language

Speaking is almost always easier than writing.

When you produce text all day

Emails, minutes, documentation, notes.

Pay once. Not every month.

Comparable dictation tools cost €7–15 per month. After a year you are past €100, and the bill keeps coming.

Celerando Pro costs €29.99 once. Updates within the version are included.

Frequently asked questions about offline dictation

Does dictation work without internet?

Yes, fully. The speech model is downloaded once from inside the app — Whisper Base at 141 MB is the default, the largest Whisper model is 1.6 GB and Parakeet is 2.5 GB — and then lives in your user profile. Recognition itself opens no connection; try it in flight mode.

Which languages are recognised?

With Whisper, 99 languages — including German, English, Spanish, French, Italian, Portuguese, Dutch, Polish, Turkish, Russian, Ukrainian, Arabic, Chinese, Japanese and Korean. Your system language is free, all 99 from Pro. From Business, Celerando detects the language automatically, and Parakeet adds 25 European languages with very low latency.

Whisper or Parakeet — which engine should I use?

Whisper is the default: the most languages, the widest choice of models, accelerated via CUDA on NVIDIA cards. Parakeet pays off if you have an AMD or Intel graphics card or want the lowest delay — 25 European languages with automatic detection, streaming-capable. Parakeet is included from Business.

How big is the download, and where are the models stored?

Whisper ranges from 75 MB (Tiny) to 1.6 GB (Large v3 Turbo), Parakeet takes 2.5 GB, and the optional CUDA runtime adds about 70 MB. The models come straight from Hugging Face and live in your Windows user profile under %LOCALAPPDATA%\Menterium\Celerando — no admin rights, and they can be deleted from inside the app.

Do I need a graphics card?

No. Whisper Tiny, Base and Small run on the CPU, even on an older office PC. A graphics card makes the larger models practical: NVIDIA via CUDA, everything else via Vulkan; Parakeet uses any DirectX 12 card. GPU acceleration is included from Pro. If you select CUDA without the runtime installed, Celerando falls back to Vulkan — no error, no abort.

Is my voice stored anywhere?

Not outside your machine. The optional transcription history (from Pro) keeps text and recording in your user profile so a dictation is not lost when it lands in the wrong window — it is off by default and can be cleared at any time.

In which programs does dictation work?

In all that accept text input — Outlook, Word, browsers, Teams, Slack, editors, forms. Celerando inserts the text via the clipboard and Ctrl+V by default; for remote desktop and terminals there is a mode that types character by character.

How long can a recording be?

In the free version 30 seconds per recording. From Pro, unlimited.

Does the text appear while I am still speaking?

That is streaming mode, included from Business — the text appears piece by piece as you speak. In Free and Pro the text is inserted as soon as you stop the recording.

What are Smart Modes?

A local language model that post-processes the recognised text — removing filler words, adding punctuation, softening the tone, phrasing it as an email. The models (such as Llama 3.2 3B or Mistral 7B) run via DirectML on a DirectX 12 graphics card with enough memory; for machines without one there is a small CPU model that is enough for punctuation and light smoothing. From Business.

Is it suitable for law firms, medical practices and businesses bound to confidentiality?

The starting point differs from a cloud service: there is no third party processing recording or text, and so no data processor that would have to be bound by contract. Whether that is sufficient for your case is for your data-protection review to decide — this page provides the technical facts for it.

Do I need admin rights?

No. Celerando installs into your user profile, and so do the models.