-
refactor the recorder so it streams samples instead of buffering the whole take
-
Hi Sarah, the beta build is ready, I'll send the download link this afternoon
-
Standup: beta ships Thursday, QA gets one extra day, Léa owns the changelog
-
Running ten minutes late, start without me and I'll demo Inkvox when I get there
-
Meeting notes: scope the on-device history search, ship the privacy page first
On-device dictation · v0.1 beta
Speak. It types. Nothing leaves your machine.
Inkvox turns your voice into text in any app, transcribed by Whisper running on your GPU. Your audio never reaches a server, and there's no account to create.
sound in · words out · nothing leaves your machine
Works where you type
Your favorite apps, now voice-powered.
On-device · measured
Faster than your fingers.
Sub-second transcription, and a private log of everything you've dictated. You can search it whenever you want, and it never leaves your machine.
0
words dictated
this week
0h
saved vs typing1
0s
avg transcription2
1 Illustrative activity, based on ~40 wpm typed vs ~150 wpm spoken. 2 11 s of speech on a mid-range GPU (RTX 3070), Whisper large-v3-turbo.
How it works
You tap, you talk, it's typed.
Press whichever key you like: the pill opens, you talk, and the text appears straight into the app you're already working in.
model: claude-fable-5 · cwd: ~/dev/inkvox
✓ Found 3 files to refactor. Ready when you are.
Léa
online
You're still good for 3pm? 14:02
On my way 14:03 ✓✓
Stylized illustrations. Product names and logos belong to their respective owners.
100+ languages
Speak whichever language you like.
Whisper detects the language as you speak. You can switch mid-sentence, and there's nothing to set up.
Privacy
The cloud never hears you.
Cloud dictation streams your voice to someone else's servers: every meeting note, every half-formed idea, everything you mutter near your microphone.
Inkvox runs the open Whisper model on your own GPU. Audio goes from your microphone to your screen and nowhere else: it's never written to disk, and it's gone the moment it's transcribed. There's no account to create, and it works just as well on a plane.
Features
A small app that does one thing well.
The whole product fits in one loop: you speak, it transcribes, it's typed.
You speak
100+ languages. Whisper detects them for you, and you can switch mid-sentence.
Inkvox transcribes
Whisper on your GPU: Vulkan on NVIDIA, AMD and Intel, with a CPU fallback if needed. Offline after one ~800 MB model download.
audio uploaded → 0 bytes
The text appears in the focused app
If you can type there, you can dictate there. Your clipboard is put back exactly as it was.
Inkvox Pro · in the works
Say it messy.
Pro cleans it up.
The Pro rewrite layer turns raw speech into clean, structured text. It's a small language model running on the same GPU, not in the cloud.
you said
okay so um what I want is like a python script that uh goes through a folder, you know, and finds all the images, the duplicates I mean, and uh deletes them, well actually no, moves them somewhere, like a trash folder or something like that
Write a Python script that scans a folder for duplicate images and moves them into a trash subfolder.
Every "um" is a token you pay for.
Paste raw dictation into Claude, ChatGPT or Cursor and the filler words bill you twice: they cost tokens, and they blur what you're asking. Pro strips them before they ever reach the model.
−0%tokens for the same request2
2 This very example, with counts approximated at ~1.3 tokens per word.
Better answers, first try
A clear prompt saves you two round trips. Pro reshapes your ramble into the question you actually meant to ask, so the model stops guessing.
Your AI bill, trimmed
Filler words are tokens. Strip them locally, for free, and every paid model call you make after that gets cheaper.
Dictation stays free
Raw dictation will never move behind the paywall. Pro adds the cleanup, unlocked with a license key, and it runs on your GPU too.
The difference
The same job, without the subscription.
Cloud dictation apps charge you monthly to run your voice through their servers. Inkvox does the same job on hardware you already own. You pay nothing, and there's nothing to leak.
Cloud dictation apps
Your voice is processed
✕on their servers, every word
Works offline
✕no internet, no dictation
Account
✕email, login, sync required
Who can replay your audio
✕whoever their policy allows
A year of dictation
$144–$180
Inkvox
Your voice is processed
✓on your GPU, it never leaves
Works offline
✓plane, train, dead zone, same speed
Account
✓none, install and talk
Who can replay your audio
✓no one, never stored, never sent
A year of dictation
$0free while in beta, Pro will be a one-time license, not a subscription
Typical cloud dictation pricing, 2026: $12–15 / month.
FAQ
Questions we get asked.
Is it free?
Dictation is free and will stay free. The upcoming AI rewrite will be the paid Pro feature, unlocked with a license key. There's no subscription wall in front of the basics.
What do I need to run it?
Windows 10/11 or macOS. Any Vulkan-capable GPU (NVIDIA, AMD or Intel, which covers most machines from the last decade) gets you sub-second dictation. Without a GPU, Inkvox falls back to CPU with a lighter model.
Where does my audio go?
From your microphone to your GPU, then to your screen. It's processed in memory, never written to disk, and discarded the moment the text is inserted.
Which model does it use?
Whisper large-v3-turbo (quantized, ~800 MB) by default, downloaded on first launch with a progress bar. On a more modest machine, Inkvox picks a lighter model automatically.
What about macOS?
Yes, macOS is part of the beta. It's the same codebase, with a Metal backend. Windows and macOS run the exact same Inkvox.
Is the demo on this page real?
It's a faithful recreation of the real pill, rebuilt in HTML: same look, same states. The actual app is native, and faster than the animation.
Stop typing.
Keep your voice to yourself.
Free beta · one email when it opens · no spam, ever