Abhishek Barali has released SpeakoFlow, a free, open-source voice dictation and AI assistant that runs locally on Windows, macOS, and Linux. Built by a solo developer who wanted to escape the limitations of paid dictation tools, the app lets users dictate text, ask questions, and hold conversations using keyboard shortcuts. It stands out by defaulting to local processing, ensuring voice data stays on your machine unless you explicitly choose a cloud provider.

Core Features and Workflow

The tool operates through three primary modes: Dictate, Ask, and Converse. In Dictate mode, holding a specific key combination (like Left Ctrl + Left Win on Windows) allows you to speak, and the text appears at your cursor in any application. The Ask mode lets you select text and query the assistant, which can translate, explain, or rewrite content. For deeper interactions, Converse mode initiates a hands-free dialogue where the assistant answers out loud. The app also includes a beta Meeting Notes feature that records both microphone and system audio to generate summaries and action items.

Local-First Architecture and Model Support

SpeakoFlow prioritizes privacy and performance by running transcription and AI tasks locally whenever possible. It supports a wide range of speech-to-text models, including Parakeet, Nemotron, and Whisper, as well as 65 other speech models in its catalog. For the AI assistant, users can choose between a built-in offline engine powered by llama.cpp, or connect to local servers like Ollama and LM Studio. If cloud services are preferred, it supports API keys for OpenAI, Anthropic, Google Gemini, and others, storing these credentials securely in the system keychain.

Installation and Platform-Specific Notes

The application is built using Tauri 2 with a Rust backend and a React/TypeScript frontend. Windows users receive a standard .exe installer, though it may trigger a SmartScreen warning due to lack of code signing. macOS users must download a .dmg and may need to run a specific terminal command to bypass Gatekeeper restrictions, as the app is not yet signed by an Apple Developer account. Linux support includes packages for Arch (AUR), Debian, and Ubuntu, along with AppImages for other distributions, though some advanced features like system-wide audio recording on macOS require additional setup with virtual audio devices.

Key Takeaways

  • SpeakoFlow is free and open-source (MIT license), offering a viable alternative to paid services like Wispr Flow and Superwhisper.
  • It defaults to local processing for transcription and AI, enhancing privacy and reducing latency.
  • The tool supports cross-platform use on Windows, macOS, and Linux with customizable keyboard shortcuts.
  • Features include real-time dictation, contextual AI queries, voice conversations, and meeting transcription.
  • Users can integrate their own local models (Ollama, LM Studio) or cloud APIs, providing flexibility in performance and cost.

The Bottom Line

SpeakoFlow is a powerful, privacy-respecting tool for developers who want to ditch subscription fees without sacrificing the convenience of voice-first workflows. Its local-first architecture and extensive model support make it a standout choice for the open-source community.

Technical Details for Builders

For those interested in contributing or building from source, SpeakoFlow requires Rust and Bun. The project structure leverages Tauri for the desktop shell, with speech processing handled by transcribe.cpp and whisper.cpp. The AI cleanup feature, known as SpeakoFlow Mini, is a specialized 795 MB model trained to remove filler words and correct grammar locally. This granular control over the stack allows developers to tailor the experience to their specific hardware capabilities, whether they are running on high-end GPUs or modest processors.