Dictata offers a 100% local voice dictation solution for Windows, leveraging whisper.cpp for accurate transcription. With features like global hotkeys, continuous dictation, and GPU support, users can efficiently convert speech to text without compromising privacy. Experience seamless integration into any application with this innovative tool.
Dictata is a fully local, system-wide voice dictation tool designed specifically for Windows, developed using native Rust. This application leverages the powerful whisper.cpp framework for accurate voice transcription without transmitting any data off the machine. Simply press a global hotkey, speak, and then press again to place the transcribed text directly into the active application.
Key Features
- Hotkey Dictation: Choose between toggle or push-to-talk modes, with automatic pasting of text into your current application.
- Continuous Mode: Experience real-time transcription as text is inserted during pauses in speech.
- Silence Skipping: Utilize voice-activity detection (VAD) to skip over silence, optimizing processing time and reducing erroneous inputs.
- GPU and CPU Support: High-efficiency transcription via Vulkan for supported AMD, Intel, or NVIDIA graphics cards, as well as CPU-only options.
- Flexible Audio Input: Transcribe from a primary microphone, system audio, or a combination for meetings.
- Custom Output Modes: Processed outputs through a local OpenAI-compatible language model, allowing for custom prompts and options for refined text (e.g., polished emails or lists).
- File Transcription: Easily transcribe any audio or video file format supported by ffmpeg through the History page of the application.
- Custom Vocabulary: Inject personalized vocabulary and replacements as needed.
- Model Library: Access a built-in ggml model library for recommendations, downloads, and management based on detected hardware configuration.
- Multilingual Interface: Support for multiple languages, including French, English, and Spanish.
System Requirements
- Optimized for Windows 10/11. Note: Linux is currently not supported. Please refer to the
LINUX.md file for more details.
- Ensure ffmpeg is available in your PATH for file transcription functions.
- For GPU support, the Vulkan SDK and Visual Studio Build Tools are necessary.
How to Use
- Launch the application; it runs in the system tray.
- Access the Settings to choose a model, set a hotkey, select the audio source, and configure your preferences.
- In the desired text input area, press the defined hotkey (default:
Ctrl+Alt+Space), start speaking, and press again to finish the transcription. Use Esc to cancel if needed.
Architecture Overview
- The application is structured into several modules, each handling specific functionality, from audio capture and model interaction to settings management and user interface. This design ensures a robust and responsive user experience.
Testing
Complete with comprehensive unit tests, the application allows developers to validate functionality effectively and ensure continued reliability as updates are implemented.
Comments
0Start the conversation
Share the first comment.