// the find
yetone/voice-input-dist
A macOS menu-bar app that turns a held hotkey into text typed into whatever field has focus, using Apple's Speech framework instead of a bundled model. It suits people who dictate into many different apps and don't want to run a Whisper stack for it. This repo is the packaged build; the README points to a separate repo for the canonical source.
- Uses Apple's Speech framework rather than shipping a model, so there is nothing to download or manage and the install stays small. The tradeoff is recognition quality on technical vocabulary like code identifiers, which is the first thing to test against a Whisper-class model.
- The codebase is eight Swift files split by responsibility: KeyMonitor, SpeechEngine, TextInjector, OverlayPanel, LLMRefiner, SettingsWindow, AppDelegate and main. That split keeps hotkey handling separate from recognition and injection, so the code is easy to read in one sitting and easy to fork.
- Post-processing sits in its own LLMRefiner stage, so cleanup is a discrete step between raw transcript and injected text rather than tangled into the recognizer.
- The repo ships a prebuilt VoiceInput.app binary, and the reproducibility claim depends on a separate source repo and an asciinema recording. Nothing here lets you check the binary against the source, so if you intend to run it, build from source yourself and compare the output.
- The README defers the license to the source repo, so this repo alone grants no reuse or redistribution rights. For a project people are expected to fork or install, that is a real gap.
- Injecting text into other apps needs Accessibility permission on macOS, and a global key monitor typically needs Input Monitoring. The README mentions neither, so first-run setup will be confusing.
- LLMRefiner implies transcripts may be sent to a language model, but the README says nothing about what gets sent, to which endpoint, or how to turn it off. Anyone dictating sensitive material needs that answered before use.