Translate the video from one language to another and embed dubbing
Capable of understanding text, audio, vision, video
Buzz transcribes and translates audio offline
VoiceStudio is the open-source, fully-local ElevenLabs alternative
Automagically synchronize subtitles with video
Open-source Video Translation Skill
Framework for building real-time voice and multimodal AI agents
Uses Qwen3-ASR, local LLM, Whisper, TEN-VAD
AI-powered tool for generating, optimizing, and translating subtitles
Voice Recognition to Text Tool
Qwen3-omni is a natively end-to-end, omni-modal LLM
A GPT-4o Level MLLM for Vision, Speech and Multimodal Live Streaming
Official MiniMax Model Context Protocol (MCP) server
Use Microsoft Edge's online text-to-speech service from Python
Official Python inference and LoRA trainer package
AI-powered video clipping and highlight generation
Open Vision Agents by Stream. Build voice and vision agents quickly
Build Vision Agents quickly with any model or video provider
Automatically translates the text of a video based on a subtitle file
A python tool that uses GPT-4, FFmpeg, and OpenCV
AI framework for automated short video creation and editing tools
"VideoRAG: Chat with Your Videos
AI generative media user experience highlighting use of APIs
A Web UI for easy subtitle using whisper model
Generate blog articles from video or audio