Vibeleaderboard
Index / app
Visit handy.computer
Category
Productivity
Rank
Pricing
Open Source
Type
APP
Use case
Productivity & Collaboration
Interfaces
Desktop
Builder
@cj_pais
Latest release
v0.9.7
Date

About

Handy is a free, open-source speech-to-text desktop app that lets you dictate into any text field using a keyboard shortcut. Press and hold to speak, release to transcribe — no cloud, no subscription, just local voice-to-text wherever your cursor is. It works across platforms and keeps your voice data private on your own machine.

What it does

Handy sits in the system tray and converts a held hotkey into transcribed text wherever the cursor is focused. While the key is down it records audio, filters out silence with voice-activity detection, then runs the clip through a locally loaded model, either a Whisper-family model or a lighter CPU-tuned alternative with automatic language detection, and pastes the result into whatever field was active. The whole pipeline runs on the machine itself, so no audio or transcript is sent anywhere during that process.

Why it's ranked here

Handy backs its pitch with real engineering. The Rust backend separates audio capture, voice detection, transcription, and system integration into distinct managers instead of one tangled process, and the project documents its own rough edges, a Bluetooth microphone quirk, a hardware-restricted key combo, Wayland paste failures, rather than hiding them. It ships through official installers plus community packages on every supported platform, and a command-line interface lets it run headless for scripting. That mix of real architecture and honesty about limits is what separates a working tool from a demo.

What's good

The model choice is genuinely useful rather than cosmetic: heavier Whisper variants run faster with a GPU, while the CPU-optimized alternative detects the spoken language automatically and stays usable on machines without one. Voice-activity detection trims silence before it reaches the model, so short holds don't waste a transcription pass. The command line goes beyond a toggle, it can batch-transcribe an audio file completely offline and emit the result as JSON, which makes the same engine usable from a script instead of only from the hotkey. A community-built Raycast extension shows the interface is open enough for others to build on.

Tradeoffs

The command-line surface is labeled a beta feature, so scripting against it today means accepting it may still change. On Linux the reliable path requires installing a separate text-input tool for your display server; skip that step and Handy falls back to a library with weaker Wayland support. A keyboard shortcut built around the fn or Globe key only fires on Apple's own keyboards, never on third-party ones, because of how that key is reported at the hardware level. And a Bluetooth headset microphone can audibly drop in quality mid-recording on macOS, a tradeoff of how Bluetooth handles simultaneous audio in and out.

How to use it well

This fits someone who wants private dictation into whatever app has focus, without paying for a cloud transcription subscription or worrying about audio leaving the machine. Pick the heavier model if a GPU is available for the best accuracy, or the lighter CPU-optimized one on a laptop without one. On Linux, install the display-server-specific text-input helper the project recommends rather than relying on the fallback. Wayland users who need a truly system-level shortcut should bind it through their desktop environment or window manager to the command-line toggle rather than the app's own hotkey handling. It is not the choice if you need shared or cloud-synced transcripts across devices.

Technical notes+

The frontend (package.json) is a Tauri 2 app built with React, TypeScript, Vite, Zustand for state, and i18next for translations, with Playwright end-to-end tests, a dedicated keyboard-handling test run through Bun, and separate lint/format scripts for the frontend and for the Rust backend via cargo fmt. On the Rust side, main.rs applies platform-specific startup fixups (disabling the Linux WebKit DMABUF renderer and, on Windows, disabling Vulkan implicit layers unless overridden) before handing off to lib.rs's run(). lib.rs wires AudioRecordingManager, ModelManager, TranscriptionManager, and HistoryManager into Tauri's managed state and deliberately defers Enigo (keyboard/mouse simulation) and shortcut initialization until the frontend calls dedicated commands after onboarding, to avoid triggering macOS permission dialogs early; src-tauri/src/commands/mod.rs exposes those as initialize_enigo and initialize_shortcuts. cli.rs defines the full CLI surface, including a headless --transcribe-file mode with device selection, model override, repeat-for-timing, and JSON output. src-tauri/src/audio_toolkit/mod.rs re-exports the recording, Silero/Earshot VAD, language-id, and text post-processing (custom word substitution, filler-word removal) primitives; src-tauri/src/managers/mod.rs lists six coordinating managers including gguf_meta and model_capabilities. App.tsx and main.tsx show the same deferred-initialization pattern on the frontend: theme is applied synchronously before first render to avoid a flash, and Enigo/shortcut commands are only invoked after onboarding completes. BUILD.md documents several resolved rough edges: an Intel-Mac ONNX Runtime dynamic-link workaround, a Windows path-length workaround for the Vulkan shader build via a short NTFS junction, and an AppImage bundling failure on rolling-release Linux distros traced to an outdated bundled strip binary.

Observed

Architecture
Desktop application built with Tauri, combining a Rust backend with a React and TypeScript frontend.
Packaging
Ships official installers for Windows, macOS, and Linux, alongside a Homebrew cask for macOS and a winget package for Windows.
Interfaces
Includes a command-line interface for remote-controlling a running instance (toggle recording, toggle post-processing, cancel) and for headless batch transcription of a WAV file, with JSON output.
Model support
Supports two local transcription model families: multiple Whisper size variants with GPU acceleration, and a CPU-optimized alternative with automatic language detection.
Voice detection
Uses Silero for voice activity detection to filter silence before transcription runs.
Linux support
On Linux, reliable text insertion depends on installing a display-server-specific helper (a tool for X11, a different one for Wayland); without one the app falls back to a library with reduced Wayland compatibility.
Testing
Repository includes Playwright end-to-end tests, a dedicated keyboard-handling test, and separate lint/format tooling for the frontend and the Rust backend.
Known limitations
README documents platform-specific limitations directly, including a Bluetooth microphone quality issue on macOS during recording, a hardware restriction on certain keyboard shortcuts, and Wayland focus/paste issues.

Read from README.md, package.json, src-tauri/src/main.rs, src-tauri/src/lib.rs, src-tauri/src/cli.rs, src/App.tsx, src/main.tsx, src-tauri/src/commands/mod.rs, src-tauri/src/audio_toolkit/mod.rs, src-tauri/src/managers/mod.rs, BUILD.md.

What it can do

  • Transcribe spoken words into text

    Voice audio captured via microphone → Transcribed text pasted into the active text field

  • Trigger voice recording via push-to-talk

    Held keyboard shortcut → Voice recording session that transcribes on release

  • Trigger voice recording via toggle mode

    Single keyboard shortcut press to start and stop → Voice recording session that transcribes when toggled off

  • Paste transcribed text into any active text field

    Completed voice transcription → Text inserted at the current cursor position in any application

  • Perform speech-to-text processing locally on device

    Microphone audio → Transcribed text processed entirely on the local machine without cloud calls

  • Customize keyboard shortcuts for recording activation

    User-defined key combination → Configured hotkey that triggers voice dictation

Tags

speech-to-textvoice inputaccessibilityopen-sourcelocal aidictationcross-platformprivacy

Tech Stack

Node.jsTypeScriptVite

Media

Handy

Comments (0)

No comments yet

Editorially curated, with community endorsements as a secondary signal. Corrections welcome.