Category
Audio agents
1,754 Audio AI agents indexed on MeshKore — the most complete public catalog, ranked by popularity and updated daily.
1,754 agents · ranked by popularity · refine in the directory →
Audio agents — page 10 of 18
OpenClaw WhoBot Skill — WhoBot (呼波特) AI电话数字员工知识技能,为 OpenClaw AI 代理注入 WhoBot 全量知识 | OpenClaw knowledge skill for AI Phone Digital Employee platform
Transformative Evolving Neural Kinetic Agent - a local-first Python voice agent for Windows.
Open-source AI memory engine for agents: multimodal memory, temporal knowledge graphs, graph RAG, and evidence-backed long-term context.
Wisp - A hotkey-driven AI overlay for your desktop. Press a key, pick an intent, and Wisp reads the right context, then streams an answer without making you leave what you're doing. Local-first, voice in/out, bring your own model provider.
用 AI 大模型复刻聊天对象的本地对话 Agent:导入真实聊天记录,LLM 学习 TA 的语气、表情和回复节奏并以人物身份延续对话,支持语音、主动联系与长期记忆,数据全在本地。 | Clone anyone's texting style from real chat history: a local-first LLM agent that learns their tone, stickers and reply rhythm, then chats as them — voice, proactive messages, long-term memory, fully private.
Conversational AI Agent & Multi-Agent Workflow Orchestration Platform
Voice AI agent with basic capabilities including telephony
Livekit STT plugin for openAI whisper models with faster-whisper backend
Gesture + voice control for macOS. Pinch to move cursor. Swipe for arrows. Circle to screenshot. Speak to run commands.
Self-hosted AI accountant. Reads receipts / invoices, reconciles ledgers, checks its own numbers, and asks before anything irreversible. Open source, agentic, runs on your infrastructure.
A private, offline AI assistant running entirely on your local machine.
A sophisticated AI-powered personal assistant inspired by Iron Man's JARVIS, featuring voice recognition, AI interaction, system control, and personalized information services.
Hardware voice recorder audio processor & task dispatcher (SecondBrain style)
Voice-first dictation & multimodal note synthesizer (Wispr style)
Voice-to-action semantic intent parser & webhook dispatcher (Wispr Action style)
Conversational candidate screening & voice interview agent (Talvo style)
Your personal AI agent. 72 built-in tools, 21 messaging channels, voice, vision, and persistent memory. Runs locally.
Ultra-low-latency real-time voice dictation & meeting memo synthesizer (Wispr style)
Real Time Interview Copilot — live dual-channel transcription (Deepgram) + LLM answers (DeepSeek/Gemini/OpenAI/Ollama), grounded in your résumé/JD. Open-source Electron app for interview practice.
A highly customizable AI companion for Telegram. Create digital clones with unique personalities, voice, and vision directly from Google Colab.
🔮 Lucy — a screen-aware AI voice agent for Android. She sees your screen, understands it with Gemini, and taps and types for you.
Bud WaaV is high performant, Scalable Audio AI Gateway written in Rust.
The Raspberry Pi desk dashboard you can talk to: widgets, touch, and a voice button wired to any AI including one running entirely on the Pi. No accounts, no cloud.
State-space Mamba architecture voice synthesizer with sub-95ms TTFT (Cartesia style)
Production-grade high-volume voice agent orchestrator with warm human transfer & CRM telephony sync (Retell style)
Multimodal meeting audio scribe & agenda task extractor (Granola AI style)
Duplex telephony voice agent turn orchestrator & SIP dialer (AI Phone Call)
Duplex low-latency conversational voice agent & WebRTC stream manager (ElevenLabs style)
Acoustic latent neural speech synthesizer & audio tokenizer (WhisperSpeech style)
Context-aware voice dictation and real-time intent formatter (Wispr Flow style)
Voice command transcription & structured autonomous action dispatcher (Wispr Flow style)
Conversational podcast dialogue audio synthesizer & multi-speaker banter (NotebookLM)
0xNuller — unified Control, Agent, Voice, Chat, Playground and Market for Web and Android
Talk to your Roland TR-8S and get a track back: a live studio, reverse-engineered SysEx, and an AI assistant on your own Claude subscription
OpenClaw skill: genpark localmusicai
OpenClaw skill: genpark voice shop
Multimodal video semantic search engine indexing video frames and audio to extract timestamped clips (Clipto style)
Source-grounded notebook and audio overview podcast synthesizer (NotebookLM style)
static region:us
Neural voiceprint speaker diarization & overlap resolver (PyAnnote style)
Ultra-lightweight 82M edge neural speech synthesizer & 20x RTF TTS (Kokoro style)
Sub-100ms conversational turn-taking & barge-in handler (Cartesia Sonic / ElevenLabs)
Sub-100ms conversational turn-taking & barge-in handler (Cartesia Sonic / ElevenLabs)
Adaptive token bucket concurrency rate limiter and backoff scheduler preventing HTTP 429 exceptions
Real-time acoustic noise floor estimator and spectral gate suppressing ambient reverberation
Semantic prompt cache deduplicator calculating Jaccard token overlap to prevent redundant LLM cost
Voice-Agents is a production-ready Python library for building enterprise-grade voice-enabled AI applications. Built by Swarms Corporation, it provides seamless integration with multiple TTS/STT providers including OpenAI, ElevenLabs, and Groq, with real-time streaming capabilities optimized for agent-based architectures.
A voice keyboard for Android powered by your own AI — dictate, and polished text appears. Bring your own OpenAI-compatible STT + LLM.
The Tiledesk Automation Engine: Your open-source solution for multichannel workflow automation. An alternative to Voiceflow, Botpress, Stack AI, Flowise, Gumloop, VectorShift, Lyzr AI
Fully automated AI presentation system — pre-generates narration, fires live n8n demo workflows, and answers audience questions via voice Q&A. Zero manual input during the talk.
Voice-to-voice LLM conversation, Full-local
Play Music in Discord.
This project uses OpenAI's GPT-3 model to create a simple assistant that can interact with you via speech or text
Onfire Games's Love Delivery Heroine Latte, An Unofficial Implementation of ChatGPT and MB-iSTFT-VITS
Your personal AI agent. 72 built-in tools, 21 messaging channels, voice, vision, and persistent memory. Runs locally.
Transform your Dialogflow NLP model to a NLP.js model
A Pokedex Discord Bot with Rich Embed and Voice Playback on Pokemon Look Up
通过Rokid AR眼镜和OpenAI Whisper实现现实生活中的字幕
AI Agent Playground 🚀
MCP Server wrapper for TTS engines (Kokoro TTS and OpenAI TTS)
AI-based YouTube summarizer with chat, notes, and chapter generation using LangChain + MERN.
Voice chatbot with voice+screen output to show that "not everything needs to be spoken"
Your personal AI assistant that runs entirely offline — with voice control, memory, web search and full system interaction. No cloud. No tracking. Just you and your AI.
An intelligent WhatsApp bot that leverages Google's Gemini AI to provide automated, context-aware responses through a modern web dashboard.
Free AI tool for mock interviews using your resume + job description. Get better with feedback after each session.
An AI-powered backseat coach to fix your skill issue and/or ruin your day :). Supports popular models from OpenAI, Anthropic and Google and self-hosted. Customizable prompting and voice cloning thanks to ElevenLabs and Coqui TTS.
Voice-first J.A.R.V.I.S. desk companion for the Waveshare ESP32-S3 Touch AMOLED 1.75C — Gemini Live, reactive round UI, physical privacy, motion, OTA, and policy-gated MCP tools.
Hotkey-first, local-first voice and text workflows for macOS & Linux. Build reusable recipes with speech-to-text, LLM prompts, text-to-speech, and desktop actions, then run them in any app using local or cloud models.
Give your Hermes agent a body — an animated face, live RGB presence, and mood, all driven by the agent's real state. Minnie is the flagship example.
Documented case of Claude Opus 4.6 entering an infinite generation loop in Cursor. 3,400 lines, 294 self-termination attempts, full transcript and recordings
Just A Rather Very Intelligent System
A turnkey, fully-local AI workstation engineered for the AMD Ryzen AI Max+ 395. LLM inference, voice, document parsing, browser automation, agents — all on-device.
Give your AI agent a mouth, ears, and a phone number. Framework-agnostic VoIP/IVR service with Twilio — plug in any agent backend (Hermes, OpenAI, Ollama, etc).
gradio Agents-MCP-Hackathon mcp-server-track region:us
gradio region:us
gradio region:us
Swift-first toolkit for shipping on-device AI on iOS, macOS, visionOS, tvOS and watchOS — Core ML LLM runtime with SwiftUI prefabs, structured output, tool calling, RAG and speech.
DeskRoute is AI receptionist for appointment-based local businesses - real phone calls via LiveKit, Google Calendar bookings by voice, and an admin dashboard for escalations and call history.
My Raze - AI 虚拟女友 Web + PWA 应用 | 智能对话 · 场景自拍 · 语音互动 · 亲密度系统
Jarvis AI OS - one brain, many shells. Landing, console, agent marketplace, plugin SDK, and desktop shell (Tauri).
Local prompt-cache economics for LLM agents: split cache reads/writes from input spend, reconcile dollars to invoices, and surface concrete TTL/marker fixes.
AI-assisted D&D 5e session management for Roll20 + D&D Beyond, built on MCP — live combat, map prep with Claude-vision walls, and a voice HUD.
Voice chatbot to use your personal AI agents.
A powerful starting point for creating a feature-rich Discord bot that utilizes modern slash commands.
AI-powered Windows desktop assistant with voice mode, Electron UI, local memory indexing, and system automation.
Official TypeScript/JavaScript SDK for Cognipeer Console — OpenAI-compatible chat, batch, realtime voice, embeddings, RAG, MCP, agent tracing, and guardrails for multi-tenant AI products.
Privacy-first AI meeting assistant. Captures mic + system audio, transcribes live, and writes the summary — entirely on your machine. Tauri/Rust core, transcribe.cpp for STT, bundled llama.cpp. No account, no cloud, no telemetry.
Multimodal video dossiers for agents: transcripts, frames, OCR, evidence, and RAG-ready knowledge.
A chat bot to emulate Donald Trump's speech - built at (and winner of) Local Hack Day (Durham) 2016
AI 面试助手: 素材整理 / 全程记录 / 自动复盘 / 模拟练习。让每一场面试都变成下一场的准备。
An AI agentic operating system for mixed reality on the Meta Quest 3 — your own J.A.R.V.I.S. Multi-agent orchestration, multimodal perception (sight/hearing/gaze), 42 holographic widgets, 20 LLM providers.
Audio dynamic range compressor and peak limiter with logarithmic decibel thresholding, ratio attenuation, and linear makeup gain.
Time-domain Pitch-Synchronous Overlap-Add (TD-PSOLA) engine stretching or compressing speech timescales without altering pitch.
A futuristic AI assistant inspired by JARVIS — featuring a stunning holographic Three.js interface, ultra-realistic voice cloning & TTS synthesis, intelligent conversations, and automated macOS system actions. Built to feel like a real sci-fi AI companion with immersive visuals, smooth interactions, and powerful desktop control capabilities. 🚀
TypeScript-фреймворк для разработки голосовых навыков и чат-ботов. Он даёт единую бизнес-логику для всех платформ — но одинаково эффективен, даже если вы работаете только с одной. Поддерживаются: `Яндекс.Алиса`, `Маруся`, `Сбер Салют`, а также `Telegram`, `VK`, `MAX` и `Viber` из коробки.
Semantic prompt cache deduplicator calculating Jaccard token overlap to prevent redundant LLM cost
A voice assistant that can recognize and respond to user’s voice commands using Porcupine, Whisper, pyttsx4, OpenAI GPT, and LangChain.
AI-powered voice task management application with natural language commands. Built with Streamlit and OpenAI.
Open-source web dashboard for Vexa – manage meeting transcriptions, view real-time transcripts, and chat with your meetings using AI.
Just A Rather Very Intelligent System