
FluidVoice
github.com/altic-dev/fluidvoice- Category
- AI Tools
- Rank
- No. 772Tools index
- Listed in
- #5 Dictate instead of type
- Pricing
- Open Source
- Type
- APP
- Use case
- Productivity & Collaboration
- Interfaces
- Desktop
- Builder
- altic-dev
- GitHub
- 11.8k stars
- Latest release
- v1.6.10-beta.7
- Date
About
An open-source macOS dictation app that runs speech-to-text fully on-device using models like Nemotron Speech, Parakeet, and Whisper, with an optional local AI enhancement layer ('Fluid Intelligence') for smart formatting and context-aware cleanup without sending data to the cloud. It also supports voice-driven Command Mode for controlling the Mac and Write Mode for rewriting text in any app.
What it does
FluidVoice listens through a Mac's microphone and turns the recording into typed text wherever the cursor is focused, running the conversion on the machine itself rather than in a remote service. A hotkey starts and stops capture, and a person can switch among several bundled speech engines depending on accuracy, latency and language needs. Beyond plain typing, the same trigger can issue a spoken instruction that fires a Mac action, or select existing text and restyle it in place. A separate, closed source enhancement engine can clean up punctuation and phrasing before anything lands on the page, or a person can swap in a paid cloud provider instead.
Why it's ranked here
The core app is real open source, copyleft under GPLv3, but the feature the project promotes hardest, its private on-device enhancement layer, is closed and maintained outside the public repository. So privacy without inspection is the more accurate claim than privacy through openness. What is verifiable holds up well: several distinct speech engines spanning a fast light model to a slower, high accuracy multilingual one, cloud API keys handled through the system credential store instead of a config file, and a realtime audio pipeline built for one job. That combination, a transparent shell wrapped around a private add-on, is an unusual shape for this category, not a broken promise.
What's good
Cloud provider keys go into the macOS system credential vault as a single entry, with old per-provider entries swept up and removed automatically rather than left behind. Requests aimed at a local address skip attaching a cloud key entirely, so a locally hosted OpenAI-compatible server never receives a credential meant for someone else's API. The realtime audio capture path is a fixed-size, allocation-free ring buffer feeding a dedicated audio callback, the standard shape for code that cannot afford to stall while a person is mid-sentence. Several speech engines ship together, covering a real range from a fast small model to slower, more accurate multilingual options.
Tradeoffs
The most differentiated feature, the private local enhancement model, sits entirely outside the GPLv3 codebase and is described as staying that way for now, so the part doing the most interesting engineering cannot be read, audited or forked the way the surrounding app can. GPLv3 itself is a strong copyleft that limits how the app's own code can be reused elsewhere. None of the fetched files include a test suite; a plan for automated UI test coverage exists only as notes for a separate, unmerged branch. Every bundled speech model outside the Whisper family also needs Apple Silicon, so Intel Mac owners get a narrower set of choices.
How to use it well
It fits someone who dictates constantly across many different Mac apps and wants the audio and text handled locally by default, with a hotkey driven workflow instead of a separate recording step. Picking a smaller, faster speech engine suits quick notes; a larger multilingual one suits longer or non-English dictation. Anyone who wants the cleanup step to stay fully auditable should leave it off and either accept the raw transcript or route enhancement through their own cloud key, since the private local model cannot be inspected. It is a macOS-only tool today: iOS and Windows support is not shipped, only a waitlist signup.
Technical notes+
The package manifest (Package.swift) pulls in FluidAudio, PromiseKit, DynamicNotchKit and transcribe-cpp-swift as Swift dependencies, plus AppUpdater for update checks, and links sqlite3 into the main executable target. Realtime input capture lives in a separate C target, Sources/CoreAudioCaptureSupport/CoreAudioCaptureSupport.c, which reads Core Audio callbacks into a fixed 64-slot ring buffer (FV_RING_CAPACITY, up to 8192 frames per packet) using only atomics and no heap allocation in the callback. Sources/Fluid/Persistence/KeychainService.swift stores every provider API key together as one JSON blob under a single Keychain item (service com.fluidvoice.provider-api-keys), and on save it migrates and deletes any older per-provider legacy Keychain entries. Sources/Fluid/Networking/AIProvider.swift implements an OpenAI-compatible chat client: it detects localhost and private IP ranges to skip the Authorization header, drops the temperature field for o1, o3 and gpt-5 prefixed models, and adds a reasoning_effort parameter for gpt-oss or openai prefixed models. Sources/Fluid/AppDelegate.swift starts a LocalAPIServer on launch and, on quit, shuts down both the speech engine and a separate PrivateAIIntegrationService before finishing termination, consistent with the private enhancement layer being an out-of-process runtime the app talks to rather than inline code. docs/MACOS_UI_AUTOMATION_BRANCH_PLAN.md lays out planned XCUITest coverage as work for an unmerged branch. No test target appears among the files read.
Observed
- License
- GPLv3 for the application code
- AI enhancement layer
- The on-device enhancement model ('Fluid Intelligence') is a separate, privately maintained runtime distributed outside the GPLv3 codebase
- Packaging and install surface
- Homebrew cask, direct release download, or build from source with Swift Package Manager and Xcode
- Platform support
- macOS 15.0 or later; Apple Silicon required for the bundled speech and enhancement models, Intel Macs limited to the Whisper models
- Interfaces
- Desktop menu bar application; also starts a local API server on launch
- Credential storage
- Cloud provider API keys are stored in the macOS Keychain rather than in a plain settings file
- Structural observation
- No test directory or test target is present among the fetched repository files
Read from README.md, Package.swift, Sources/Fluid/fluidApp.swift, Sources/Fluid/AppDelegate.swift, Sources/Fluid/Networking/AIProvider.swift, Sources/Fluid/Persistence/SettingsStore.swift, Sources/Fluid/Persistence/KeychainService.swift, Sources/Fluid/Persistence/TranscriptionHistoryStore.swift, Sources/CoreAudioCaptureSupport/CoreAudioCaptureSupport.c, docs/MACOS_UI_AUTOMATION_BRANCH_PLAN.md.
What it can do
Transcribe speech to text on-device
Voice audio → Text
Enhance and format transcribed text using local AI
Raw transcribed text → Cleaned/formatted text
Control the Mac via voice commands
Voice command → System action
Rewrite text in any app via voice
Voice instruction and existing text → Rewritten text
Tags
Tech Stack
Comments (0)
No comments yet
Editorially curated, with community endorsements as a secondary signal. Corrections welcome.