Category

Audio agents

1,754 Audio AI agents indexed on MeshKore — the most complete public catalog, ranked by popularity and updated daily.

1,754 agents · ranked by popularity · refine in the directory →

Audio agents — page 3 of 18

ai-agent-video-editor★ 151

ai agent video editor use with ElevenLabs Scribe, pack phrase-level transcripts, reason over an EDL, render with ffmpeg, video editor ai agent

Multi-voice-TTS-GPT-SoVITS★ 150

gradio region:us

react-native-lingua★ 147

React Native Duolingo clone with a real-time AI voice teacher. Built with Expo, Stream Voice Agents, Clerk auth, and NativeWind for a complete, interactive mobile learning experience.

OpenVision★ 147

Open-source iOS app connecting Meta Ray-Ban smart glasses to AI — 5 backends (on-device MLX models, Apple Intelligence, OpenAI, Gemini Live, OpenClaw), on-device neural voice, face recognition & live web search. Private and offline-capable.

whisper-clip★ 139

WhisperClip simplifies your life by automatically transcribing audio recordings and saving the text directly to your clipboard. With just a click of a button, you can effortlessly convert spoken words into written text, ready to be pasted wherever you need it. This application harnesses the power of OpenAI’s Whisper for free.

Auto-Subtitled-Video-Generator★ 136

Input a YouTube video link or upload a video file and get a video with subtitles.

SlackONOS★ 136

🎵 Democratic Slack/Discord bot for Sonos control with Spotify integration. Queue music, vote to skip, and let the community decide what plays!

PersonalAssistantChatbot★ 133

It is a personal assistant chatbot, capable to perform many tasks same as Google Assistant plus more extra features...

Awesome-Colorful-LLM★ 130

Recent advancements propelled by large language models (LLMs), encompassing an array of domains including Vision, Audio, Agent, Robotics, Fundamental Sciences such as Mathematics, and Ominous.

mediasage★ 129

Unofficial AI-powered Plex playlist generator with library awareness. Every track it suggests, you actually own.

Noema★ 129

A voice-first desktop AI companion with memory, emotion, tools, and plugins.

chatgpt-web★ 127

ChatGPT web application. ChatGPT 网页应用,支持多对话、海量提示词、PWA、ASR、TTS

PodAgent★ 127

PodAgent: A Comprehensive Framework for Podcast Generation

whis★ 126

Voice-to-text CLI for terminal users

realtime-interview-copilot★ 125

Realtime Interview Copilot is a desktop application that assists users in crafting responses during interviews. It leverages real-time audio transcription and AI-powered response generation to provide relevant and concise answers.

saa-sdk★ 122

Addressee detection for voice agents: device-directed speech detection that runs before STT, so background speech, side conversations, and the agent's own TTS echo never trigger it. No wake word, model-agnostic, drop-in for LiveKit, Pipecat, ElevenLabs, Twilio, and OpenAI. The layer your VAD and turn detection are missing.

AIEV★ 122

Automatic AI video editing. Claude directs HyperFrames (HTML + GSAP motion graphics) and Remotion (timeline assembly) to turn raw footage into a finished MP4 - transcript, kinetic typography, karaoke subtitles, sound effects. Self-hosted, supervised from a web dashboard on port 6868.

zyron-assistant★ 120

⚡ A local, privacy-focused AI desktop assistant for Windows. Control your PC remotely via Telegram or locally with Voice commands. Powered by Ollama.

VivaDicta★ 119

iOS & watchOS speech-to-text app with AI voice keyboard, on-device RAG, and chat with your notes - powered by Apple Foundation Models, WhisperKit, NVIDIA Parakeet, and 20+ AI providers

UniversalBackrooms★ 118

Historical multi-model Backrooms experiment with configurable conversation templates and example transcripts.

MisterWhisper★ 115

Push to talk voice recognition using Whisper

workersai★ 115

Full-stack AI chat platform built on Cloudflare using Workers, Durable Objects, KV, and AI Gateway. Features AI chat, Text-to-Speech (TTS), and Speech-to-Text (STT).

J.A.R.V.I.S★ 115

Iron man inspired Personal virtual assistant

lizheng-open-context★ 115

Source-grounded public context for 立正: Superlinear posts, selected comments, YouTube, Public Axioms V1, and 《真本事》.

sip-to-ai★ 114

Turn any SIP call into a realtime AI voice agent (OpenAI Realtime / Deepgram/Gemini Live/xAI Grok Voice)

novella★ 114

🎬 端到端 AI 漫剧 / 动画短剧 Multi-Agent 多智能体协作创作平台 (Novella AI)。基于 Hub-and-Spoke 智能体架构、ProjectBlackboard 共享黑板、自定义 Agent 扩展插件、角色 Consistency 锚定与 4K GPU 压制。

sonic-match-mcp★ 114

MCP server that watches video footage and returns license-safe BGM matches, hook windows, and ffmpeg ducking specs for agents.

jarvis-ai-assistant★ 112

🤖 Advanced AI-powered virtual assistant with voice recognition, face authentication, phone integration, and intelligent automation capabilities.

speechless★ 109

LLM based agents with proactive interactions, long-term memory, external tool integration, and local deployment capabilities.

magec★ 109

Multi-agent AI platform with voice and text control. Visual workflows, web interface, and chat integrations (Telegram, Discord, Slack). Long-term memory, any LLM backend, extensible via MCP tools.

fcp-mcp-server★ 109

MCP server for Final Cut Pro XML. Lets Claude read, edit and generate real timelines: cut detection, markers, roles, transcript-based editing. Published on PyPI as fcp-mcp-server.

modelguide★ 108

Open-source voice agent orchestration framework - build production voice AI pipelines without vendor lock-in

awesome-openai-whisper★ 107

A curated list of awesome OpenAI's Whisper

JARVIS-AI-ASSISTANT★ 106

A true Artificial Intelligent Assistant with ALICE as backend and offline speech recognition with vosk engine and pyttsx3 as text to speech engine

AdobePremiereProMCP★ 106

🎬 AI-powered MCP server for Adobe Premiere Pro — 1,027 tools for timeline editing, color grading, audio mixing, effects, export & more. Control video editing with natural language via Claude, GPT, or any AI assistant. The most comprehensive MCP server for any NLE.

phlox★ 105

Open source, local first AI medical agent for desktop and web.

OpenGlasses★ 105

Making Meta Ray-ban / Oakley glasses smarter. Non-Meta Model selection, agentic features, smart guidance.

openclaw-voice-call-realtime★ 105

Give your AI assistant a phone — OpenClaw plugin for real phone calls via Twilio + OpenAI Realtime, with in-call tools, transcripts, and call screening

openai-realtime-python★ 104

Real-time voice agent powered by Agora and OpenAI

ultrasonic★ 103

A comprehensive steganography framework for embedding and extracting agentic commands in audio and video media using ultrasonic frequencies. This project provides tools for covert communication and command transmission through multimedia channels.

video-production-skill★ 103

The complete video-production skill behind the 蝦說 AI channel — lets an AI agent autonomously produce narrated educational videos (slides + TTS + ASR verification + subtitles).

ChatVox★ 102

"Chat With Any Video" project in 24 hours, challenge myself to complete in @Supabase's AI Hackathon.

LLMChat★ 101

A Discord chatbot that supports popular LLMs for text generation and ultra-realistic voices for voice chat.

gptspeaker★ 101

The ChatGPT/DeepSeek Voice Assistant uses a Raspberry Pi (or desktop) to enable spoken conversation with OpenAI or DeepSeek large language models. This implementation listens to speech, processes the conversation through the OpenAI/DeepSeek service, and responds back. Like Apple Siri, Amazon Alex, Google Nest Home, Mi XiaoAi etc.

utsuwa★ 101

Utsuwa is an open-source alternative to Grok Companion. This is a platform where you can have a virtual AI waifu that learns and grows with you, bundled with optional mechanics inspired by Japanese dating sim games.

agent-starter-node★ 100

A complete voice AI starter app for LiveKit Agents with Node.js

XnneHangLab★ 100

希望用代码为 waifus 绘心。

agent-starter-android★ 100

AI voice assistant starter app for iOS, macOS, and visionOS built with LiveKit

facetime-bridge★ 100

A bridge to help connect a brand new iCloud account to Facetime audio for voice to voice agent action

line★ 99

Cartesia Line SDK for voice agents.

reaper-mcp★ 99

A comprehensive Model Context Protocol (MCP) server that enables AI agents to create fully mixed and mastered tracks in REAPER with both MIDI and audio capabilities.

tankwork★ 98

Desktop agent framework for creating AI agents that can see and control your computer through voice and text commands

mobileclaw★ 98

🦞 MobileClaw — 带眼睛的龙虾对讲机 | Multimodal voice+vision walkie-talkie for OpenClaw AI agents. iOS & Android.

GPTube★ 97

🎥 Youtube Video Summarizer and Question Answering App Using Whisper and Langchain

ZZZ-HiveMind-core★ 97

Join the OVOS collective, utils for OpenVoiceOS mesh networking

simulflow★ 97

A Clojure library for building real-time voice-enabled AI Agents. Simulflow handles the orchestration of speech recognition, audio processing, and AI service integration with the elegance of functional programming.

Daisy-Voice-Agent★ 97

Daisy:AI 语音助手

flowcat★ 97

Self-hosted native-Rust runtime for real-time voice agents. Own the stack: one binary in your own VPC or air-gapped, no hosted control plane. pipecat-compatible pipeline, in-process SIP/RTP, single-process call density. Apache-2.0.

jev-use★ 97

Say it, and your Mac does it. A computer-use harness on Jev that reads the screen through Accessibility. Fast, no vision model

whisper-nextjs★ 96

Next.js app for serverless deployments of OpenAI Whisper on Banana.dev

agent-starter-swift★ 96

AI voice assistant starter app for iOS, macOS, and visionOS built with LiveKit

riffrec★ 96

Capture golden product feedback sessions with screen, voice, DOM, network, and console context for AI agents.

home-mind★ 95

OSS AI memory layer for Home Assistant (AGPL). The conversation server inside the Nives addon was forked from this project — they're independent now.

call-center-voice-agent-accelerator★ 95

Create speech-to-speech voice agents that deliver personalized self-service experiences and natural-sounding voices, seamlessly integrated with telephony systems.

ableton-copilot-mcp★ 94

An MCP server built on ableton-js enables AI assistants to control Ableton Live in real time, including Arrangement View operations such as song management, track control, MIDI editing, and audio recording, along with other capabilities.

ai-assistant-teaching-website★ 94

[AI-Assistant] “通慧智教”-大模型赋能智能教学网站-通过AI教学助理-小慧,使用语音或文本,以自然语言方式与网站交互。祝贺小慧成功斩获国内多个软件设计赛事的【国家一等奖】、【国家二等奖】及【国家三等奖】。Large Model Empowered Intelligent Teaching Website - Interacts with the website using voice or text in a natural language manner through AI teaching assistant-XiaoHui.

laibot-client★ 93

开源人工智能,基于开源软硬件构建语音对话机器人、智能音箱……人机对话、自然交互,来宝拥有无限可能。特别说明,来宝运行于Python 3!

ubo_app★ 93

Ubo app aims to streamline hardware-integrated agentic experience development on embedded platforms (supports Raspberry Pi)

eidoverse-video★ 93

Video-production toolkit for AI agents: Deno + WebGPU + three.js/TSL render engine with VRM character locomotion, simulations, effects, and a full audio pipeline — real-time GPU rendering with minimal CPU

MusicGen★ 93

Music generated agent

ChatGPT-voice-control★ 92

Voice control for ChatGPT. Talk to ChatGPT and hear ChatGPT's responses in a natural voice.

shellChatGPT★ 92

Shell wrapper for OpenAI's ChatGPT, Whisper, and TTS. Features LocalAI, Ollama, Gemini, Anthropic, and more.

jargo★ 92

Conversational-AI framework for Go.

conversational-voice-ai-agent★ 91

You can have a voice chat with AI Agent

agent-starter-flutter★ 91

AI voice assistant starter app for Flutter built with LiveKit

benchmarks★ 91

Reproducible voice-AI benchmarking — TTS / STT / S2S latency and accuracy.

talon-ai-tools★ 90

Query LLMs and AI tools with voice commands

cli★ 90

Command Line Interface to interact with ElevenLabs voice agents

speech-core★ 90

On-device VAD / streaming STT / TTS / diarization in C++17 (ONNX + LiteRT) with a voice-agent pipeline. Linux, Windows, Android.

pywhisper★ 89

openai/whisper + extra features

NOVA-NodeJS★ 88

NOVA is a customizable voice assistant made with Node.js.

nimble-pipecat★ 88

Voice Agent Framework for Conversational AI

OpenAI-Text-To-Speech-for-Unity★ 88

Implementation of OpenAI's Text-To-Speech in Unity. Synthesize any text and play it via any AudioSource.

asktube★ 87

AskTube - An AI-powered YouTube video summarizer and QA assistant powered by Retrieval Augmented Generation (RAG) 🤖. Run it entirely on your local machine with Ollama, or cloud-based models like Claude, OpenAI, Gemini, Mistral, and more.

SpeechAgents★ 87

SpeechAgents: Human-Communication Simulation with Multi-Modal Multi-Agent Systems

phoneflow★ 87

The Open Source Voice Agent Platform

iris★ 86

Iris is a desktop voice companion that uses Gemini Live for natural realtime conversation and Hermes Agent for long-running work.

WinWhisper★ 85

Create subtitles with ease, using Whisper AI for Windows

agent-starter-embed★ 85

Embeddable AI voice assistant button built with LiveKit

obsidian-scribe★ 85

Record, transcribe, and transform voice notes into structured insights. Leverage Whisper or AssemblyAI and ChatGPT to fill in gaps, generate summaries, and visualize ideas — all seamlessly integrated within Obsidian.

meeper★ 84

Meeper 📝 - is your secretary for any in-browser conference.

trx★ 84

Agent-first CLI for audio/video transcription via Whisper

alts★ 82

100% free, local & offline voice assistant with speech recognition

logic-pro-mcp★ 82

Local MCP server for stateful, fail-closed Logic Pro control and live project readback.

roboot★ 80

Personal AI agent hub on macOS — iTerm2 sessions, voice, web console, Telegram, Cloudflare Worker relay with E2EE

Chatbot★ 80

Hybrid Conversational Bot based on both neural retrieval and neural generative mechanism with TTS.

langgraph-livekit-agents★ 80

LangGraph adapter for LiveKit Agents

whisper_normalizer★ 80

A python package for whisper normalizer

multi-modal-customer-service-agent★ 79

Multi-modal & multi-domain customer service agent with real time text, voice and soon video

obsidian-scribe★ 79

Record, transcribe, and transform voice notes into structured insights. Leverage Whisper or AssemblyAI and ChatGPT to fill in gaps, generate summaries, and visualize ideas — all seamlessly integrated within Obsidian.

Browse other category pages