Audio foundation model excelling in audio understanding
Repo of Qwen2-Audio chat & pretrained large audio language model
Large Audio Language Model built for natural interactions
Multilingual speech recognition and audio understanding model
Speech recognition module for Python
Fast and accurate automatic speech recognition (ASR) for edge devices
Robust Speech Recognition via Large-Scale Weak Supervision
Multi-modal large language model designed for audio understanding
Captcha solver extension for humans
Uses Qwen3-ASR, local LLM, Whisper, TEN-VAD
Speech-to-text, text-to-speech, and speaker recognition
Buzz transcribes and translates audio offline
HTML5 js recording mp3 wav ogg webm amr format
Automatic Speech Recognition with Word-level Timestamps
Fast multimodal LLM for real-time voice interaction and AI apps
Speech recognition for your site
AsrTools: Smart Voice-to-Text Tool
Framework for building real-time voice and multimodal AI agents
A free, open source, and extensible speech-to-text application
Open speech-to-speech models and pipelines by Hugging Face toolkit AI
State-of-the-art Machine Learning for Pytorch, TensorFlow, and JAX
Voice Recognition to Text Tool
Capable of understanding text, audio, vision, video
VoiceStudio is the open-source, fully-local ElevenLabs alternative
Python Audio Analysis Library: Feature Extraction, Classification