QuickWhisper
QuickWhisper is a macOS application for transcription, dictation, and AI summarization using OpenAI's Whisper model. It runs entirely on-device with no cloud dependency required.
The application transcribes audio from local files, YouTube videos, online meetings, and system audio. QuickWhisper can record meetings with calendar integration while keeping the recording interface hidden during screen sharing.
System-wide dictation works across all macOS applications, replacing keyboard input with voice. All transcription runs on your Mac.
AI summarization is available through cloud providers (OpenAI, Anthropic, Google, xAI, Mistral, Groq) or on-device via Ollama and LM Studio.
QuickWhisper also includes batch transcription, Watch Folders for automatic background transcription, speaker diarization, Apple Shortcuts integration, and webhooks for third-party service integration.
Learn more
Scribe
ElevenLabs has introduced Scribe, an advanced Automatic Speech Recognition (ASR) model designed to deliver highly accurate transcriptions across 99 languages. Scribe is engineered to handle diverse real-world audio scenarios, providing features such as word-level timestamps, speaker diarization, and audio-event tagging. Benchmark tests, including FLEURS and Common Voice, demonstrate Scribe's superior performance over leading models like Gemini 2.0 Flash, Whisper Large V3, and Deepgram Nova-3, achieving the lowest word error rates in languages such as Italian (98.7%) and English (96.7%). Notably, Scribe also significantly reduces errors in languages that have been traditionally underserved, including Serbian, Cantonese, and Malayalam, where other models often exhibit error rates exceeding 40%. Developers can integrate Scribe through ElevenLabs' speech-to-text API, receiving structured JSON transcripts that include detailed annotations.
Learn more
FastScribe
AI transcription tool that converts audio and video to text with timestamps and automatic speaker labels. Identifies who is speaking and separates the transcript into labelled turns, which you can rename.
Supports MP3, M4A, WAV, AAC, FLAC, OGG, Opus, WMA, AMR, MP4, MOV, WEBM, AVI, MKV and more. Exports TXT, SRT, VTT and DOCX subtitles with speaker names included.
Free tier with no signup required for the first file, speaker labels included.
Speech recognition and speaker diarization both run on private self-hosted GPU hardware, and audio is deleted immediately after transcription.
Supports Spanish, French, German, Portuguese, Italian, Japanese, Hindi, Korean and more.
Learn more
Ecango
Ecango is an AI-powered audio and video transcription tool that converts spoken content into accurate, searchable text in seconds. Users can upload or drag and drop audio or video files, let Ecango generate the transcript, then edit it directly in the browser and export it in popular formats including DOCX, ODT, PDF, SRT, and TXT. It supports transcription, subtitles, and translation across more than 90 languages, dialects, and accents, using advanced speech recognition to deliver up to 99.8% accuracy. Speaker identification and diarization detect different people speaking within the same recording and organize their dialogue into an easy-to-read transcript. Ecango supports popular audio and video formats and automatically handles video files without requiring users to separate the audio first. Its AI can also filter background noise to improve transcription and translation results when recordings are less than ideal.
Learn more