Windows
Windows 11
Download for WindowsVoice input, an agent and autocomplete for Windows. On top of any program.
Works on top of what you already use
Speed
Speech runs at 150 words a minute. Typing, at 40.
The left-hand clock runs on everything: the speech, and then recognition, formatting and insertion - the same three phases the first screen shows. The program shows no draft while you speak either. On the right, typing at forty words a minute with one typo.
The second step
It sees where you corrected yourself and where the thought ended.
so um I looked at that report uh on sales for the third no for the fourth quarter .
“The third - no, the fourth” is a self-correction. Only what you meant stays in the text.
Talking aloud
It hears you while you speak and answers aloud. Interrupt it.
YouOpen Spotify and put on something calm
AssistantOpening Spotify
YouAnd remind me about the call in half an hour
AssistantGot it. I will tell you at seven, while this conversation is open
1.4 seconds from the end of your sentence to the first sound back
An assistant with hands
You say what you need. Up to three workers get on it.
Midri VoiceWhat the assistant is allowed to do on this computer
This is your computer and your files. Turn on only what you intend to use: a permission can be taken back at any time, but what has been done will not undo itself.
Yours
Four shapes, seven colour sets. Pick one, or watch it try them on.
Tap a shape or a colour
When you type anyway
It finishes your sentence in your window. Agree with Tab.
Code and prompts
“Use state” becomes useState. Names come out as code.
project
Code names restored from pronunciation and from what is open on screen.
Translation as you speak
Translation is not a step. It is the same key press.
Insert in: ES
Someone else's text
Select someone else’s text and the bar appears above it.
Written chat
A hotkey over any window. Same memory as the voice.
You are working. No browser tab open, nothing to switch to.
Memory
Not rows in a file but a web of links: neighbours come too.
Midri VoiceConversational speaking pace
From your last word to a spoken answer
Translation languages on insert
What it can do
Twenty-three things, in three groups, in one program.
What it does when you speak to it.
10It hears you while you are still speaking and answers in its own voice, with no transcription in between. About a second and a half to the first sound back.
It sees everything installed and launches by name. Ask for one already running and it brings that window forward instead of opening a second copy.
The assistant has a separate Chromium with its own profile: it reads pages, clicks and fills fields only in there. You sign in to the sites you need once, inside it. The browser is downloaded separately: 149 MB to fetch, 344 MB on disk.
Low and calm, lively and quick, soft and unhurried - the voice changes from the orb menu and from settings. Two of the thirty work on the free plan.
Share your screen and ask about what is on it: read an error, walk you through steps, make sense of an unfamiliar window.
“Remind me in half an hour” - and in half an hour it says so out loud, as long as the conversation is open. Not a notification, just a sentence in the middle of the talk.
What happens between your voice and the text.
9Recognition runs on your computer. Lose the connection and dictation keeps going, cleaned by rules. Your speech is never sent anywhere.
Forty-five well-known names such as GitHub, macOS and Kubernetes are restored automatically. Add the rest as a list with no length limit - matches are found by sound.
Names from the active window join the dictionary for the current phrase. “Slack” becomes Slack exactly where you are talking about it.
Right Alt, Caps Lock, Ctrl+Shift - anything, a single key included. Hold it while you talk, like a walkie-talkie.
“Business tone”, “no exclamation marks”, “short sentences” - rules are written in words and applied to every phrase.
Everything you dictated stays at hand and can be pasted again in one click. Stored only on your computer.
How it sits there when you are not using it.
4Dictation, text processing, chat, Live AI, connectors, history, appearance, system, and the subscription tile. The assistant's permissions and its workers live inside “Live AI”, where you switch the conversation on. Theme, twelve accent colours, window material, and the edge the panel sits on.
Until you touch it the panel is a strip at the edge of the screen. Hover and it unfolds into the microphone, conversation, assistant and settings - then tucks away again.
If you would rather not pay through us, paste your own key from OpenAI, Anthropic, Google or another provider and its models line up beside ours for the workers. The key sits in the Windows Credential Manager and never travels to us.
The panel does not take focus, does not pop over your work, and does not move the cursor. The assistant opens a window of its own in one case only - when you have granted it the browser and it is using it.
Install
Recognition runs on the local model - no internet, nothing sent anywhere. An account is needed to use the program, and it opens what runs on our side: smart processing, translation, talking aloud, and the assistant.
The program is in a closed beta; there is no public build yet. Create an account and we will write to it as soon as the build can be handed over. macOS and Linux versions are planned - work on them has not started.
Pricing
Recognition runs on your machine and costs us nothing. We charge for what runs on our side: processing, translation, talking aloud and the agent.
I write all day
I write and I delegate
There are several of us
Allowances reset on the first of the month, together with your billing.
Questions
Midri Voice is a voice assistant for Windows. You hold a key and speak - the text arrives finished in whatever window you were typing in. Or you talk to it out loud and it answers in a voice, opens apps, and sorts out your files.
It is free while in preview. There are no payments on the site yet and no card is asked for anywhere.
Yes. Sign-in is through Google and takes one click - there is no password to invent or forget. The account is what the paid plans and the usage allowance are attached to.
Dictation is recognised on your own machine and the audio does not go anywhere. A spoken conversation is different: there your voice travels through the service, because otherwise nothing could answer you.
Windows 10 and Windows 11 today. Other systems are coming - the assistant itself is not tied to Windows.
Any window with a text field: Cursor, VS Code, JetBrains, the terminal, Telegram, Slack, Gmail, Notion, a browser. It remembers the field you clicked before you started speaking, so the text lands there even if the window moved.
Built-in dictation writes down sounds. This one understands what you meant: it punctuates, drops the ums, keeps code and names intact, translates as it goes, and remembers what you told it yesterday.
Your speech is recognised on your own computer. The account is for what is computed on our side: smart processing, translation, talking aloud, and the assistant.
Windows 11 · free while in preview