Emacs Carnival Oct 2026: Exploring keys and voice input in Emacs
Making more use of the keyboard
October's Emacs Carnival theme "Are keyboards and M-x here to stay?" is a good nudge to think about the voice interface I've been experimenting with and where I want to take it.
Like Chris Maiorana, I love the way Emacs's M-x (execute-extended-command) allows me to call any command by name.
It means that I can call commands or come up with my own commands without worrying about assigning them to keyboard shortcuts or making sense of where to place them in a hierarchical menu.
With flexible completion via the vertico and orderless packages, I don’t even need to type in the full name of the command or remember the words in the right order.
If I still can’t remember the command, it’s easy enough to use defalias to make my own nickname for it.
I tend to limit keyboard shortcuts to only my most common commands.
Everything else is hard to practise enough so that it sticks in my muscle memory.
Even with marginalia showing me keybindings when I use M-x and which-key showing me hints when I pause, it doesn’t sink in.
Maybe I could use key-quiz to drill a small number of keyboard shortcuts I’m actively working on learning, or modify help-quick-sections to show me reminders when I use C-h C-q (help-quick-toggle).
Getting one keyboard shortcut to do the right thing in different
contexts is great because it means fewer shortcuts to memorize. Sometimes I can do that by binding the right functions to different mode maps with bind-key, like the way my sacha-org-insert-link-dwim uses the right link syntax in Markdown or OddMuse too.
I also like the context-sensitive keyboard shortcuts I get from Embark.
I find myself using a handful of those regularly, and C-.
(embark-act) starts a bunch of shortcuts across different types of
things: i to insert, w to copy, and so on.
For example, I know C-x C-f (find-file) will let me select a file to open.
Instead of opening it, I can use C-. i to insert the filename into the buffer.
Simpler keyboard shortcuts help. I rarely use the numpad on my laptop keyboard, so those keys can be mapped to functions that don't need to be followed by lots of typing. For example, 0 is a large button. I use it to start and stop speech recognition. Actually, I bind f9 in Emacs to sacha-whisper-run-at-point, and I bind numpad-0 as a keyboard shortcut in KDE (my window manager) instead of just in Emacs so that it's available globally. I'm still on X11 instead of Wayland. The shell script xdotool-emacs focuses Emacs briefly and sends the keys, and the shortcut is set to /home/sacha/bin/xdotool-emacs key --clearmodifiers F9. I do that instead of using emacsclient because emacsclient starts a different buffer, and I might want to keep the context. My sacha-whisper-maybe-type function detects if Emacs is not the focused application and types the text if it isn't, which means I can use it as keyboard input for oher applications while still taking advantage of the corrections and commands I've defined in Emacs. An easy-to-press global shortcut means I can use it from anywhere, even when my foreground app is the browser or another application.
An additional modifier opens up more possibilities for keyboard shortcuts near the regular typing position. My Windows key acts like a Super modifier. I've set it up as a one-shot modifier with Kanata so that I can either hold it down to access more shortcuts or tap it to treat it like a leader key.
There are quite a few commands for which neither M-x nor keyboard shortcuts feel like the right fit: things I want to do faster or more fluidly than M-x, but which I don't do often enough to memorize a good shortcut for. I've been experimenting a bit with voice input, and I really like it.
Some reasons why I'm curious about speech recognition
- I could use more practice speaking, since I don't talk to a lot of people in real life.
- English: Videos and livestreams can be a way to share interesting things and start conversations. I want to organize my thoughts, improve my breathing, and reduce my stutters and false starts. That will make the audio more pleasant to listen to. Dictating to my computer gives me more practice. It also makes things easier to livestream if I can reuse my speech as input.
French: I like challenging my brain and I want to help the kiddo learn it too. Improving my speaking will also make it easier for me to listen to myself and close the feedback loop. Maybe I can even use speech recognition to give me feedback and increase motivation while speaking, like keeping track of words or phrases I want to use. When I use them in conversation, the buffer can cross them off the list, which lets me focus on the words remaining.
Figure 3: Screenshot showing balance and target words
- Dictation can capture ideas quickly, with more natural wording, especially since typing sometimes gets messed up on my computer.
- Voice interfaces might let me reduce some kinds of task friction or work hands-free. For example, I like setting timers by voice on my phone while cooking. It might be interesting to be able to dictate a draft while knitting or hand-sewing.
- Brain fog is something I will likely need to deal with. I got a preview of it thanks to the sleep deprivation of early parenting. I know that brain fog is probably going to happen around menopause, it will probably also be part of aging, and it's a common symptom of long COVID which I might still end up getting despite my precautions. I could get into an accident and hurt my hands. I like Emacs and I want to be able to keep using it even if cognitive or physical changes make it harder to remember or use lots of keyboard shortcuts. If I start playing around with this now, the capabilities and the habits will be there when I need it.
- Other people are already dealing with those kinds of physical or cognitive constraints, so things I figure out might be useful for other people.
- A global shortcut for voice notes and commands let me quickly dash off a reminder before I take off. I don't have to find Emacs and go through my Org capture templates.
- It's fun. I like saying "Prepare for take-off" and having Emacs get ready for a spontaneous stream by updating YouTube using my current Org Mode subtree title and description. I like seeing if my dictionary of substitutions can fix the commonly misrecognized words. Saying "I command you to…" to get Emacs to do a M-x command is totally inefficient but also playful, and silly things stick in my brain better thanks to the mnemonic advantage of humor. My brain likes that kind of novelty.
For context, I use Linux on a Lenovo P52 laptop (manufactured 2018, 64 GB RAM but hardly any GPU), and I have an Android phone. I like writing regular text, Emacs Lisp, and some Python and JavaScript. I'm slowly starting to have more focused time now that the kiddo is in virtual school, although it's still broken up by recess and lunch. I will probably have an increasing amount of focused time over the next few years.
Quick benefits from voice commands
I love the way voice reduces context shifts. For example, some of my most-used voice commands are:
- remind me [today/tomorrow/next week/next month/someday] to…: capture and schedule the task; include an Org link and screenshot for context, and link to the audio recording in case the text was misrecognized
- switch to…: capture the task and immediately clock into it
- where was I?: jump to the current task
- what was I doing?: list recently-clocked tasks
I don't have to think about which keys I've configured in org-capture-templates. I just use a natural phrase.
I've defined "open …" to open things from my bookmarked links stored in an Org file. Because my numpad-0 keybinding works from anywhere, I can even use it to open sites or capture notes while looking at my browser.
Because Emacs can interact with lots of other applications, I can even experiment with controlling those applications by voice. For example, I use obs-websocket.el to control OBS for streaming or recording a video. Using "Let's go live" and "Over and out" to start or stop streaming or "Start recording" and "Stop recording" to manage the recording means I can do that without needing to find the right window. I stream with a 10-second delay so that there's a little bit of time to catch any private information that's been accidentally shared on the screen. My new "panic" command lets me discard that delay and interrupt the stream by using an OBS keyboard shortcut to kill the stream immediately, even if I'm probably not going to remember that exact keyboard shortcut when I'm stressed. It should be a little easier to remember that than a keyboard shortcut I hardly ever use or a menu item that is a really small target (and my Emacs might not even be my foreground app).
For regular dictation, I'm experimenting with having it automatically place each sentence on a new line or list item. This makes it easy to reorganize in Org. I can refill it into a paragraph when I'm done.
I still have to deal with the limits of my memory, of course.
I have a "what can I say" command that opens up an Org file with my notes. I keep this as a manual note instead of just dumping the contents of my sacha-whisper-commands variable so that I can prioritize the things I'm trying to build as habits.
Here's the part of my config that sets up some speech commands.
I might want more complex grammar someday for things that regular expressions can't easily handle. That's something other projects have done a good job of figuring out.
Prior work on voice control of Emacs: Voicemacs, Emacslisten
There's been a lot of good work on voice input over the years, especially since a number of Emacs users have a hard time with or would like to avoid repetitive stress injury. In addition to the innumerable projects that let you use Whisper or similar speech recognition engines to dictate prose, there are also a few projects that use a constrained vocabulary to more reliably control Emacs. Two projects in this area are based on the Talon voice control system: Voicemacs and Emacslisten.
- The Talon community hub defines some Emacs commands. I think it simulates keyboard input and assumes that you have flx or some other fuzzy completion matching so that input like
M-x t-d-on-erunstoggle-debug-on-error, as defined in community/apps/emacs/emacs_commands.csv. I seem to be having a hard time with getting orderless to do that withorderless-flex, so I'm just going to remove that column of the CSV from my emacs_commands.csv. - Emacslisten uses Emacs Lisp macros to define the command vocabulary. Cursorfree defines a whole bunch of them, using hatty for voice navigation within a buffer.
- Voicemacs defines its commands in Python in the user's Talon config (see example). It feels fairly complex, like it assumes it's being cloned as your
~/.talon/userdirectory. It lets Emacs and Talon communicate via RPC to send additional context. It also extends some things like dired and company to make them more voice-friendly.
I'm new to both projects, so I'm not sure if I'm reading this right, but I think I might be able to use Emacslisten to define voice commands with Emacs Lisp, and then maybe someday use Voicemacs to provide richer context for some of those commands. For example, Voicemacs can list snippets, but I'm not quite sure how to use that yet.
It looks like I can use Talon for commands and my whisper.el setup for regular dictation at the same time. The Pipewire virtual sinks help with that, I think, since the programs don't have to have an exclusive lock on the audio device. This means that I can explore the kinds of vocabularies that people have set up while also taking advantage of greater accuracy and ease of correction in my whisper dictation workflow. I think I can use the Talon process socket to turn off recognition when I want to dictate using my whisper.el setup, and then I can turn it back on afterwards.
A note about the underlying speech recognition: Talon is cross-platform and free, but not open source. Subscribing to Talon (CAD 36/month) lets you use a Conformer + Whisper hybrid engine that uses whisper-large for more accuracy, but that's a bit on the expensive side for me. Dragonfly is a free open source engine that can use Kaldi or CMU Pocket Sphinx for speech recognition, but I'm not sure what the performance is like.
Interestingly, because Talon simulates keystrokes, using the commands from the community set means keys sometimes arrive out of order on my computer. "split solo" is supposed to run M-x delete-other-windows, but it sometimes ends up sending M-x delete-other-wnidows instead. This is the same sort of annoyance that gets in the way of typing quickly on my computer. Keystrokes get processed out of order when the CPU gets busy, which is part of the reason why I find dictation to be more reliable than typing sometimes. I should fix that someday, but my attempts so far aren't working. (I wonder if it has something to do with the GPU under Linux; maybe I should drop back down to CPU only since it turns out I can't use the GPU for much anyway.) Quick fix, let's try adding a brief delay in my settings.talon:
settings():
# Adds a small pause (in ms) between each character typed
insert_wait = 5
# Ensures modifiers (Shift/Ctrl) don't un-latch too fast
key_wait = 1.0
I think Voicemacs uses a sevver and and Emacslisten uses emacsclient to communicate with Emacs directly, so they probably won't run into this keystroke problem.
It might be interesting to see if I can port some of my simpler commands over to Emacslisten or straight Talon so that they're always active. Then I can slowly work on getting used to the other commands people find useful, or incorporating them into my setup.
Imagining a hybrid approach: keyboard, voice, even mouse/stylus?
I love shifting fluidly from one form of input to another, depending on what makes sense at that moment.
For example, I often start my thoughts by drawing on a light dotted grid using Noteful on my iPad. That lets me quickly cluster ideas spatially, drawing connections and highlighting words as I figure out what I want to say.
Then I use dictation and typing to sketch the outline, jumping around as thoughts occur to me. If I'm livestreaming at the same time, that helps me demonstrate things along the way and lets me benefit from people's questions. Dictation lets me explain things to people and capture the text at the same time.
I can move around with the keyboard, but it would be really cool to use voice for quick navigation between sections of my outline using either line numbers or words. I've implemented a "line" command that lets me jump to a line using the last two digits of its line number, since I probably won't have more than 100 lines visible on the screen at a time. For example, this line is currently 119, so referring to it as "nineteen" would work. I'd like to add more operations for editing the text. I think it would be great to be able to say "kill line 33" to delete that line or "move lines 16 to 20 after 27" to have it move those lines there.
I now have a sacha-whisper-one-sentence-per-line variable that splits up dictation if non-nil, so we'll see where those take me. Orukeet does a reasonable job of punctuating my sentences if I don't pause too much between phrases. My M-q re-wraps paragraphs between one sentence per line, one paragraph per line, and wrapping to fill-column, and it's probably the sort of thing that more voice commands would make even more convenient.
Maybe I can even use that spatial sense of a layout by touching buttons in an Emacs-served webpage on my tablet to direct my dictation to different places. Org-draw uses a similar Emacs web server approach to let you draw on another device and embed the results into an Org file, and there are also simpler HTTP servers (web-server, simple-httpd), so it's probably quite doable. I'm imagining something that takes the top-level items from my current nested list and makes them nice big buttons, so I could have something like "Introduction", "Keyboard", "Reasons", "Experiment", "Hybrid", and "Conclusion", and clicking on those could trigger a POST request to the server inside Emacs which moves my point accordingly. I might even have buttons for various commands. It sounds pretty straightforward to implement.
When my brain comes up with lots of interesting ideas, my org-capture voice commands let me choose which rabbit holes to postpone and which ones I want to dive into.
If I get interrupted by life, I can write on my phone while waiting for the kiddo. Orgzly Revived and Syncthing let me pick up where I left off.
As I get thoughts down, I can use keyboard shortcuts to rearrange things and then convert the list of sentences into paragraphs. I can easily turn text into links using my favourites.
I can imagine combining the audio clips from the text properties with the videos and screenshots to create a quick narrated clip that I can edit (using Emacs, of course), or publishing the narration so that people can listen to it if they like.
This combination of drawing + audio + text works better for me than any of those components alone. Visuals give me anchors for my thoughts. Audio helps with the braindump. Text permits precise, searchable details we can copy. Like George Jones, I find they all have their uses and their strengths.
When I try to dictate longer sections without that kind of visual support, I often lose my train of thought. I want to see where I've been and where I want to go. I don't trust LLMs to clean up that kind of rough transcript; the output doesn't quite resonate with me. I like keeping speech recognition constrained to sentences I can immediately check or reorganize, and I have the recordings as text properties so I can go back to that segment if I need to (at least during the current editing section). That way, I'm never left scratching my head and wondering what I meant.
I love that local speech recognition is getting better and better. I'm curious about whether small language models can do better turn detection to help compensate for my intra-phrase delays when I'm thinking of a word, and whether decision models can automatically classify a sentence into its place in my outline. I think it would be cool if embeddings or other language models could suggest related words or resources as I go along. Lots of interesting possibilities even with the technology that's already available.
Thanks to George Jones for hosting this month's Emacs Carnival on input alternatives. Check out the page for more posts!







