Category
Audio agents
1,754 Audio AI agents indexed on MeshKore — the most complete public catalog, ranked by popularity and updated daily.
1,754 agents · ranked by popularity · refine in the directory →
Audio agents — page 3 of 18
ai agent video editor use with ElevenLabs Scribe, pack phrase-level transcripts, reason over an EDL, render with ffmpeg, video editor ai agent
gradio region:us
React Native Duolingo clone with a real-time AI voice teacher. Built with Expo, Stream Voice Agents, Clerk auth, and NativeWind for a complete, interactive mobile learning experience.
Open-source iOS app connecting Meta Ray-Ban smart glasses to AI — 5 backends (on-device MLX models, Apple Intelligence, OpenAI, Gemini Live, OpenClaw), on-device neural voice, face recognition & live web search. Private and offline-capable.
WhisperClip simplifies your life by automatically transcribing audio recordings and saving the text directly to your clipboard. With just a click of a button, you can effortlessly convert spoken words into written text, ready to be pasted wherever you need it. This application harnesses the power of OpenAI’s Whisper for free.
Input a YouTube video link or upload a video file and get a video with subtitles.
🎵 Democratic Slack/Discord bot for Sonos control with Spotify integration. Queue music, vote to skip, and let the community decide what plays!
It is a personal assistant chatbot, capable to perform many tasks same as Google Assistant plus more extra features...
Recent advancements propelled by large language models (LLMs), encompassing an array of domains including Vision, Audio, Agent, Robotics, Fundamental Sciences such as Mathematics, and Ominous.
Unofficial AI-powered Plex playlist generator with library awareness. Every track it suggests, you actually own.
A voice-first desktop AI companion with memory, emotion, tools, and plugins.
ChatGPT web application. ChatGPT 网页应用,支持多对话、海量提示词、PWA、ASR、TTS
PodAgent: A Comprehensive Framework for Podcast Generation
Voice-to-text CLI for terminal users
Realtime Interview Copilot is a desktop application that assists users in crafting responses during interviews. It leverages real-time audio transcription and AI-powered response generation to provide relevant and concise answers.
Addressee detection for voice agents: device-directed speech detection that runs before STT, so background speech, side conversations, and the agent's own TTS echo never trigger it. No wake word, model-agnostic, drop-in for LiveKit, Pipecat, ElevenLabs, Twilio, and OpenAI. The layer your VAD and turn detection are missing.
Automatic AI video editing. Claude directs HyperFrames (HTML + GSAP motion graphics) and Remotion (timeline assembly) to turn raw footage into a finished MP4 - transcript, kinetic typography, karaoke subtitles, sound effects. Self-hosted, supervised from a web dashboard on port 6868.
⚡ A local, privacy-focused AI desktop assistant for Windows. Control your PC remotely via Telegram or locally with Voice commands. Powered by Ollama.
iOS & watchOS speech-to-text app with AI voice keyboard, on-device RAG, and chat with your notes - powered by Apple Foundation Models, WhisperKit, NVIDIA Parakeet, and 20+ AI providers
Historical multi-model Backrooms experiment with configurable conversation templates and example transcripts.
Push to talk voice recognition using Whisper
Full-stack AI chat platform built on Cloudflare using Workers, Durable Objects, KV, and AI Gateway. Features AI chat, Text-to-Speech (TTS), and Speech-to-Text (STT).
Iron man inspired Personal virtual assistant
Source-grounded public context for 立正: Superlinear posts, selected comments, YouTube, Public Axioms V1, and 《真本事》.
Turn any SIP call into a realtime AI voice agent (OpenAI Realtime / Deepgram/Gemini Live/xAI Grok Voice)
🎬 端到端 AI 漫剧 / 动画短剧 Multi-Agent 多智能体协作创作平台 (Novella AI)。基于 Hub-and-Spoke 智能体架构、ProjectBlackboard 共享黑板、自定义 Agent 扩展插件、角色 Consistency 锚定与 4K GPU 压制。
MCP server that watches video footage and returns license-safe BGM matches, hook windows, and ffmpeg ducking specs for agents.
🤖 Advanced AI-powered virtual assistant with voice recognition, face authentication, phone integration, and intelligent automation capabilities.
LLM based agents with proactive interactions, long-term memory, external tool integration, and local deployment capabilities.
Multi-agent AI platform with voice and text control. Visual workflows, web interface, and chat integrations (Telegram, Discord, Slack). Long-term memory, any LLM backend, extensible via MCP tools.
MCP server for Final Cut Pro XML. Lets Claude read, edit and generate real timelines: cut detection, markers, roles, transcript-based editing. Published on PyPI as fcp-mcp-server.
Open-source voice agent orchestration framework - build production voice AI pipelines without vendor lock-in
A curated list of awesome OpenAI's Whisper
A true Artificial Intelligent Assistant with ALICE as backend and offline speech recognition with vosk engine and pyttsx3 as text to speech engine
🎬 AI-powered MCP server for Adobe Premiere Pro — 1,027 tools for timeline editing, color grading, audio mixing, effects, export & more. Control video editing with natural language via Claude, GPT, or any AI assistant. The most comprehensive MCP server for any NLE.
Open source, local first AI medical agent for desktop and web.
Making Meta Ray-ban / Oakley glasses smarter. Non-Meta Model selection, agentic features, smart guidance.
Give your AI assistant a phone — OpenClaw plugin for real phone calls via Twilio + OpenAI Realtime, with in-call tools, transcripts, and call screening
Real-time voice agent powered by Agora and OpenAI
A comprehensive steganography framework for embedding and extracting agentic commands in audio and video media using ultrasonic frequencies. This project provides tools for covert communication and command transmission through multimedia channels.
The complete video-production skill behind the 蝦說 AI channel — lets an AI agent autonomously produce narrated educational videos (slides + TTS + ASR verification + subtitles).
"Chat With Any Video" project in 24 hours, challenge myself to complete in @Supabase's AI Hackathon.
A Discord chatbot that supports popular LLMs for text generation and ultra-realistic voices for voice chat.
The ChatGPT/DeepSeek Voice Assistant uses a Raspberry Pi (or desktop) to enable spoken conversation with OpenAI or DeepSeek large language models. This implementation listens to speech, processes the conversation through the OpenAI/DeepSeek service, and responds back. Like Apple Siri, Amazon Alex, Google Nest Home, Mi XiaoAi etc.
Utsuwa is an open-source alternative to Grok Companion. This is a platform where you can have a virtual AI waifu that learns and grows with you, bundled with optional mechanics inspired by Japanese dating sim games.
A complete voice AI starter app for LiveKit Agents with Node.js
希望用代码为 waifus 绘心。
AI voice assistant starter app for iOS, macOS, and visionOS built with LiveKit
A bridge to help connect a brand new iCloud account to Facetime audio for voice to voice agent action
Cartesia Line SDK for voice agents.
A comprehensive Model Context Protocol (MCP) server that enables AI agents to create fully mixed and mastered tracks in REAPER with both MIDI and audio capabilities.
Desktop agent framework for creating AI agents that can see and control your computer through voice and text commands
🦞 MobileClaw — 带眼睛的龙虾对讲机 | Multimodal voice+vision walkie-talkie for OpenClaw AI agents. iOS & Android.
🎥 Youtube Video Summarizer and Question Answering App Using Whisper and Langchain
Join the OVOS collective, utils for OpenVoiceOS mesh networking
A Clojure library for building real-time voice-enabled AI Agents. Simulflow handles the orchestration of speech recognition, audio processing, and AI service integration with the elegance of functional programming.
Daisy:AI 语音助手
Self-hosted native-Rust runtime for real-time voice agents. Own the stack: one binary in your own VPC or air-gapped, no hosted control plane. pipecat-compatible pipeline, in-process SIP/RTP, single-process call density. Apache-2.0.
Say it, and your Mac does it. A computer-use harness on Jev that reads the screen through Accessibility. Fast, no vision model
Next.js app for serverless deployments of OpenAI Whisper on Banana.dev
AI voice assistant starter app for iOS, macOS, and visionOS built with LiveKit
Capture golden product feedback sessions with screen, voice, DOM, network, and console context for AI agents.
OSS AI memory layer for Home Assistant (AGPL). The conversation server inside the Nives addon was forked from this project — they're independent now.
Create speech-to-speech voice agents that deliver personalized self-service experiences and natural-sounding voices, seamlessly integrated with telephony systems.
An MCP server built on ableton-js enables AI assistants to control Ableton Live in real time, including Arrangement View operations such as song management, track control, MIDI editing, and audio recording, along with other capabilities.
[AI-Assistant] “通慧智教”-大模型赋能智能教学网站-通过AI教学助理-小慧,使用语音或文本,以自然语言方式与网站交互。祝贺小慧成功斩获国内多个软件设计赛事的【国家一等奖】、【国家二等奖】及【国家三等奖】。Large Model Empowered Intelligent Teaching Website - Interacts with the website using voice or text in a natural language manner through AI teaching assistant-XiaoHui.
开源人工智能,基于开源软硬件构建语音对话机器人、智能音箱……人机对话、自然交互,来宝拥有无限可能。特别说明,来宝运行于Python 3!
Ubo app aims to streamline hardware-integrated agentic experience development on embedded platforms (supports Raspberry Pi)
Video-production toolkit for AI agents: Deno + WebGPU + three.js/TSL render engine with VRM character locomotion, simulations, effects, and a full audio pipeline — real-time GPU rendering with minimal CPU
Music generated agent
Voice control for ChatGPT. Talk to ChatGPT and hear ChatGPT's responses in a natural voice.
Shell wrapper for OpenAI's ChatGPT, Whisper, and TTS. Features LocalAI, Ollama, Gemini, Anthropic, and more.
Conversational-AI framework for Go.
You can have a voice chat with AI Agent
AI voice assistant starter app for Flutter built with LiveKit
Reproducible voice-AI benchmarking — TTS / STT / S2S latency and accuracy.
Query LLMs and AI tools with voice commands
Command Line Interface to interact with ElevenLabs voice agents
On-device VAD / streaming STT / TTS / diarization in C++17 (ONNX + LiteRT) with a voice-agent pipeline. Linux, Windows, Android.
openai/whisper + extra features
NOVA is a customizable voice assistant made with Node.js.
Voice Agent Framework for Conversational AI
Implementation of OpenAI's Text-To-Speech in Unity. Synthesize any text and play it via any AudioSource.
AskTube - An AI-powered YouTube video summarizer and QA assistant powered by Retrieval Augmented Generation (RAG) 🤖. Run it entirely on your local machine with Ollama, or cloud-based models like Claude, OpenAI, Gemini, Mistral, and more.
SpeechAgents: Human-Communication Simulation with Multi-Modal Multi-Agent Systems
The Open Source Voice Agent Platform
Iris is a desktop voice companion that uses Gemini Live for natural realtime conversation and Hermes Agent for long-running work.
Create subtitles with ease, using Whisper AI for Windows
Embeddable AI voice assistant button built with LiveKit
Record, transcribe, and transform voice notes into structured insights. Leverage Whisper or AssemblyAI and ChatGPT to fill in gaps, generate summaries, and visualize ideas — all seamlessly integrated within Obsidian.
Meeper 📝 - is your secretary for any in-browser conference.
Agent-first CLI for audio/video transcription via Whisper
100% free, local & offline voice assistant with speech recognition
Local MCP server for stateful, fail-closed Logic Pro control and live project readback.
Personal AI agent hub on macOS — iTerm2 sessions, voice, web console, Telegram, Cloudflare Worker relay with E2EE
Hybrid Conversational Bot based on both neural retrieval and neural generative mechanism with TTS.
LangGraph adapter for LiveKit Agents
A python package for whisper normalizer
Multi-modal & multi-domain customer service agent with real time text, voice and soon video
Record, transcribe, and transform voice notes into structured insights. Leverage Whisper or AssemblyAI and ChatGPT to fill in gaps, generate summaries, and visualize ideas — all seamlessly integrated within Obsidian.