Generate audiobooks from EPUBs, PDFs and text with captions
Image inpainting tool powered by SOTA AI Model
OCR software, free and offline
Faster Whisper transcription with CTranslate2
Comprehensive Gradio WebUI for audio processing
Use Microsoft Edge's online text-to-speech service from Python
Robust Speech Recognition via Large-Scale Weak Supervision
Open source healthcare AI
A TTS that fits in your CPU (and pocket)
Cut videos with a text editor
PDF to Markdown with vision models
Python library and CLI tool to interface with Google Translate
Contexts Optical Compression
Translate the video from one language to another and embed dubbing
Stable Diffusion web UI
1 min voice data can also be used to train a good TTS model
Visual Causal Flow
Essential nodes that are weirdly missing from ComfyUI core
Implementation of Imagen, Google's Text-to-Image Neural Network
Free, high-quality text-to-speech API endpoint to replace OpenAI
95% token savings. 155x faster queries. 16 languages
Automated translation solution for visual novels
OCR model for complex documents with layout-aware structured outputs
Enhances Tesseract OCR output using LLMs (local or API)
Check code for common misspellings