Category
Audio agents
1,474 Audio AI agents indexed on MeshKore — the most complete public catalog, ranked by popularity and updated daily.
1,474 agents · ranked by popularity · refine in the directory →
Audio agents — page 3 of 15
Push to talk voice recognition using Whisper
Addressee detection for voice agents: device-directed speech detection that runs before STT, so background speech, side conversations, and the agent's own TTS echo never trigger it. No wake word, model-agnostic, drop-in for LiveKit, Pipecat, ElevenLabs, Twilio, and OpenAI. The layer your VAD and turn detection are missing.
Unofficial AI-powered Plex playlist generator with library awareness. Every track it suggests, you actually own.
用 AI 大模型复刻聊天对象的本地对话 Agent:导入真实聊天记录,LLM 学习 TA 的语气、表情和回复节奏并以人物身份延续对话,支持语音、主动联系与长期记忆,数据全在本地。 | Clone anyone's texting style from real chat history: a local-first LLM agent that learns their tone, stickers and reply rhythm, then chats as them — voice, proactive messages, long-term memory, fully private.
Free, open-source voice dictation and AI assistant for Windows, macOS, and Linux. Offline Whisper speech-to-text: everything Wispr Flow and Superwhisper do, plus screen vision. Press a key and speak into any app. Say "Hey Flow" and it writes the reply from what's on your screen.
Full-stack AI chat platform built on Cloudflare using Workers, Durable Objects, KV, and AI Gateway. Features AI chat, Text-to-Speech (TTS), and Speech-to-Text (STT).
LLM based agents with proactive interactions, long-term memory, external tool integration, and local deployment capabilities.
Realtime Interview Copilot is a desktop application that assists users in crafting responses during interviews. It leverages real-time audio transcription and AI-powered response generation to provide relevant and concise answers.
A curated list of awesome OpenAI's Whisper
AI Phone Agent: A starter kit to build AI agents that answer real phone calls and talk to customers in real time (OpenAI Realtime). Node.js backend for Twilio & Amazon Connect - ship phone-first voice agents faster. Dial +1 (855) 522-2348 to try the demo.
Open-source voice agent orchestration framework - build production voice AI pipelines without vendor lock-in
A true Artificial Intelligent Assistant with ALICE as backend and offline speech recognition with vosk engine and pyttsx3 as text to speech engine
Real-time voice agent powered by Agora and OpenAI
Multi-agent AI platform with voice and text control. Visual workflows, web interface, and chat integrations (Telegram, Discord, Slack). Long-term memory, any LLM backend, extensible via MCP tools.
Open source, local first AI medical agent for desktop and web.
iOS & watchOS speech-to-text app with AI voice keyboard, on-device RAG, and chat with your notes - powered by Apple Foundation Models, WhisperKit, NVIDIA Parakeet, and 20+ AI providers
Open-source iOS app connecting Meta Ray-Ban smart glasses to AI — 5 backends (on-device MLX models, Apple Intelligence, OpenAI, Gemini Live, OpenClaw), on-device neural voice, face recognition & live web search. Private and offline-capable.
"Chat With Any Video" project in 24 hours, challenge myself to complete in @Supabase's AI Hackathon.
A Discord chatbot that supports popular LLMs for text generation and ultra-realistic voices for voice chat.
Iron man inspired Personal virtual assistant
The ChatGPT/DeepSeek Voice Assistant uses a Raspberry Pi (or desktop) to enable spoken conversation with OpenAI or DeepSeek large language models. This implementation listens to speech, processes the conversation through the OpenAI/DeepSeek service, and responds back. Like Apple Siri, Amazon Alex, Google Nest Home, Mi XiaoAi etc.
A complete voice AI starter app for LiveKit Agents with Node.js
希望用代码为 waifus 绘心。
Give your AI assistant a phone — OpenClaw plugin for real phone calls via Twilio + OpenAI Realtime, with in-call tools, transcripts, and call screening
Cartesia Line SDK for voice agents.
A comprehensive steganography framework for embedding and extracting agentic commands in audio and video media using ultrasonic frequencies. This project provides tools for covert communication and command transmission through multimedia channels.
Turn any SIP call into a realtime AI voice agent (OpenAI Realtime / Deepgram/Gemini Live/xAI Grok Voice)
A comprehensive Model Context Protocol (MCP) server that enables AI agents to create fully mixed and mastered tracks in REAPER with both MIDI and audio capabilities.
Desktop agent framework for creating AI agents that can see and control your computer through voice and text commands
🦞 MobileClaw — 带眼睛的龙虾对讲机 | Multimodal voice+vision walkie-talkie for OpenClaw AI agents. iOS & Android.
The complete video-production skill behind the 蝦說 AI channel — lets an AI agent autonomously produce narrated educational videos (slides + TTS + ASR verification + subtitles).
🎥 Youtube Video Summarizer and Question Answering App Using Whisper and Langchain
Join the OVOS collective, utils for OpenVoiceOS mesh networking
AI voice assistant starter app for iOS, macOS, and visionOS built with LiveKit
Next.js app for serverless deployments of OpenAI Whisper on Banana.dev
A Clojure library for building real-time voice-enabled AI Agents. Simulflow handles the orchestration of speech recognition, audio processing, and AI service integration with the elegance of functional programming.
开源人工智能,基于开源软硬件构建语音对话机器人、智能音箱……人机对话、自然交互,来宝拥有无限可能。特别说明,来宝运行于Python 3!
AI voice assistant starter app for iOS, macOS, and visionOS built with LiveKit
Self-hosted native-Rust runtime for real-time voice agents. Own the stack: one binary in your own VPC or air-gapped, no hosted control plane. pipecat-compatible pipeline, in-process SIP/RTP, single-process call density. Apache-2.0.
Voice control for ChatGPT. Talk to ChatGPT and hear ChatGPT's responses in a natural voice.
Shell wrapper for OpenAI's ChatGPT, Whisper, and TTS. Features LocalAI, Ollama, Gemini, Anthropic, and more.
[AI-Assistant] “通慧智教”-大模型赋能智能教学网站-通过AI教学助理-小慧,使用语音或文本,以自然语言方式与网站交互。祝贺小慧成功斩获国内多个软件设计赛事的【国家一等奖】、【国家二等奖】及【国家三等奖】。Large Model Empowered Intelligent Teaching Website - Interacts with the website using voice or text in a natural language manner through AI teaching assistant-XiaoHui.
You can have a voice chat with AI Agent
An MCP server built on ableton-js enables AI assistants to control Ableton Live in real time, including Arrangement View operations such as song management, track control, MIDI editing, and audio recording, along with other capabilities.
AI voice assistant starter app for Flutter built with LiveKit
openai/whisper + extra features
Query LLMs and AI tools with voice commands
NOVA is a customizable voice assistant made with Node.js.
Capture golden product feedback sessions with screen, voice, DOM, network, and console context for AI agents.
Create speech-to-speech voice agents that deliver personalized self-service experiences and natural-sounding voices, seamlessly integrated with telephony systems.
Voice Agent Framework for Conversational AI
SpeechAgents: Human-Communication Simulation with Multi-Modal Multi-Agent Systems
Implementation of OpenAI's Text-To-Speech in Unity. Synthesize any text and play it via any AudioSource.
Ubo app aims to streamline hardware-integrated agentic experience development on embedded platforms (supports Raspberry Pi)
Daisy:AI 语音助手
AskTube - An AI-powered YouTube video summarizer and QA assistant powered by Retrieval Augmented Generation (RAG) 🤖. Run it entirely on your local machine with Ollama, or cloud-based models like Claude, OpenAI, Gemini, Mistral, and more.
Iris is a desktop voice companion that uses Gemini Live for natural realtime conversation and Hermes Agent for long-running work.
Meeper 📝 - is your secretary for any in-browser conference.
Create subtitles with ease, using Whisper AI for Windows
Agent-first CLI for audio/video transcription via Whisper
100% free, local & offline voice assistant with speech recognition
Making Meta Ray-ban / Oakley glasses smarter. Non-Meta Model selection, agentic features, smart guidance.
Personal AI agent hub on macOS — iTerm2 sessions, voice, web console, Telegram, Cloudflare Worker relay with E2EE
Automatic AI video editing. Claude directs HyperFrames (HTML + GSAP motion graphics) and Remotion (timeline assembly) to turn raw footage into a finished MP4 - transcript, kinetic typography, karaoke subtitles, sound effects. Self-hosted, supervised from a web dashboard on port 6868.
Hybrid Conversational Bot based on both neural retrieval and neural generative mechanism with TTS.
LangGraph adapter for LiveKit Agents
Embeddable AI voice assistant button built with LiveKit
🎬 AI-powered MCP server for Adobe Premiere Pro — 1,027 tools for timeline editing, color grading, audio mixing, effects, export & more. Control video editing with natural language via Claude, GPT, or any AI assistant. The most comprehensive MCP server for any NLE.
A python package for whisper normalizer
Record, transcribe, and transform voice notes into structured insights. Leverage Whisper or AssemblyAI and ChatGPT to fill in gaps, generate summaries, and visualize ideas — all seamlessly integrated within Obsidian.
OSS AI memory layer for Home Assistant (AGPL). The conversation server inside the Nives addon was forked from this project — they're independent now.
Multi-modal & multi-domain customer service agent with real time text, voice and soon video
Simple GUI around whisper.cpp for voice-to-text on Linux
Voice Prompts, GPT-4o prompts, Voice Agent Prompts, ChatGPT Prompts, HumeAI Prompts
A comprehensive Model Context Protocol (MCP) server that enables AI agents to create fully mixed and mastered tracks in REAPER with both MIDI and audio capabilities.
Whisper is an automatic speech recognition (ASR) system Gradio Web UI Implementation
AI Voice Assistant: Talk to an AI agent that helps you with event scheduling, contact management, accessing your knowledge base, and web searches using simple voice commands
🎬 端到端 AI 漫剧 / 动画短剧 Multi-Agent 多智能体协作创作平台 (Novella AI)。基于 Hub-and-Spoke 智能体架构、ProjectBlackboard 共享黑板、自定义 Agent 扩展插件、角色 Consistency 锚定与 4K GPU 压制。
Twitch livestream bot that can control colors for overlays from Stream Elements, play sound effects, handle custom rewards (like text-to-speech) and more!
Deploy generative AI agents in your contact center for voice and chat using Amazon Connect, Amazon Lex, and Amazon Bedrock Knowledge Bases
Omnigram is a Flutter-based file reader and audiobook . It accommodates EPUB and PDF and offers audiobook functionality, supporting TTS model and other AI chat technologies for enhanced reading experiences
Leopard Chat UI - A Teneo Chat Client based on Vue and Vuetify
An Android ChatBot powered by IBM Watson Services (Assistant V1, Text-to-Speech, and Speech-to-Text with Speaker Recognition) on IBM Cloud.
Voice AI agent starter kit with Groq, Llama 4, and (optionally) Twilio
A voice assistant application built with the LiveKit Agents framework, capable of using Model Context Protocol (MCP) tools to interact with external services
Production-ready audio and video transcription app that can run on your laptop or in the cloud.
Talk to your second brain personal assistant using speech 🧠
Utsuwa is an open-source alternative to Grok Companion. This is a platform where you can have a virtual AI waifu that learns and grows with you, bundled with optional mechanics inspired by Japanese dating sim games.
A minimalistic automatic speech recognition streamlit based webapp powered by OpenAI's Whisper "State of the Art" models
Voice-powered AI assistant platform — connect any LLM, any TTS, with a live web canvas, music generation, and agent orchestration using openclaw. Install: npx openvoiceui setup
A modern, serverless web application that connects users to a Microsoft Copilot Studio agent for booking and managing appointments via chat and voice input.
A voice activated interface for your custom AI Agent.
Pipecat framework based orchestrator for building real-time, voice-enabled, and multimodal conversational AI agents
Audio-Oscar is a multi-agent framework for generating long-form, controllable audio from complex audio scene descriptions.
An autonomous, context-aware AI desktop companion. Built with Python, featuring real-time screen vision, custom ONNX voice synthesis, active window tracking, and a dynamic floating UI with reactive facial expressions.
gradio region:us
A WebRTC-native, audio-first conversational-AI framework for Go.
A personal AI OS for macOS, powered by Claude.
OmniAgent (ICML 2026): the first native omni-modal agent for active video perception — a 7B agent that beats Qwen2.5-VL-72B with 73% fewer frames on LVBench.