Category
Audio agents
1,474 Audio AI agents indexed on MeshKore — the most complete public catalog, ranked by popularity and updated daily.
1,474 agents · ranked by popularity · refine in the directory →
Audio agents — page 10 of 15
AI-powered Discord voice bot with natural conversation, smart turn detection, and OpenAI-compatible TTS/STT
Voice Chat AI is a project that allows you to interact with different AI characters which is modular using speech. You can choose between various characters, each with unique personalities and voices. Have a serious conversation with Albert Einstein or role play with the OS you can use it for voice
WhisperForge is a Python tool that leverages OpenAI's Whisper model to transcribe large audio files. It automatically splits files into manageable chunks, processes them, and combines the transcriptions into a single document. Ideal for handling lengthy recordings and generating clear, organized transcriptions.
AI-powered platform for capturing, enhancing, and organizing ideas using voice input and community collaboration.
🎤 Control your world with Jarvis, a voice-activated AI assistant that simplifies tasks and enhances productivity.
AI agent skill for creating cinematic product launch videos with Remotion
AI video pipeline with a self-correcting critique loop. Brief → script → AI visuals + ElevenLabs narration + music → Gemini Video critique → auto-patch → re-render until target quality.
A **standalone** Discord bot with LLM, VOICEVOX, and KaTeX support.
gradio agent-demo-track region:us
gradio region:us
gradio region:us
Xiaoyuzhou podcast fetcher & transcriber — download, transcribe with faster-whisper, summarize
Set it and forget it! A clean, node-free GUI that automates full AI music videos. Queue your prompts and let the tool generate and stitch LTX 2.3 clips, then automatically score the final cut with an ACE Step 1.5 audio track. No waiting around for each clip to render.
Batuk is a sovereign AI chat workspace for local and hosted models, with multi-provider support, document chat, voice, web search, guardrails, auth, and local-first persistence.
MLX Porting Toolkit — an agent-guided, evidence-gated pipeline (scaffold → convert → parity → benchmark) plus a portable skill for porting PyTorch/Hugging Face models to Apple MLX.
The agent-native video editor for music & shortform: voice conversion, lyric & karaoke videos, auto-subtitles, and a full FX stack. AI drives the real timeline over MCP. Linux-native C++/ImGui, zero dependencies.
AI-powered voice-note device firmware (ESP32-S3): record, transcribe with Whisper, clean up notes with GPT, and auto-sync Markdown to GitHub/Obsidian. An AI upgrade of the Pala Note hardware.
AI原生的长篇小说阅读器,解决读长篇小说人类上下文不够的痛点。支持知识图谱,支持RAG,支持安卓移动端,支持tts引擎朗读和合成多角色对话语音效果。AI novel reader with EPUB import, RAG search, knowledge graph extraction, and an Android companion app.
DeskRoute is AI receptionist for appointment-based local businesses - real phone calls via LiveKit, Google Calendar bookings by voice, and an admin dashboard for escalations and call history.
Sync Canvas courses into a local Obsidian vault — transcribed lecture notes, a cross-lecture concept graph, homework and deadlines — then study it with Claude or Gemini over MCP.
Blabber — 首个 Agentic 播客视频生成平台。一句提示词,生成一整期动画播客节目:AI 编写对白 · 双主持配音 · 逐帧口型同步 · 自动运镜剪辑,从想法到成片一站直达。Blabber — The first agentic podcast video generation platform. One prompt, a full animated podcast show: AI-written dialogue · dual-host voices · frame-accurate lip sync · automated camera direction & editing, from idea to final cut.
My Raze - AI 虚拟女友 Web + PWA 应用 | 智能对话 · 场景自拍 · 语音互动 · 亲密度系统
Personal AI assistant — voice, vision, memory, multi-model routing, meeting overlay
AI-assisted D&D 5e session management for Roll20 + D&D Beyond, built on MCP — live combat, map prep with Claude-vision walls, and a voice HUD.
Control your Windows laptop from your phone with AI. Zara is a Discord-based personal AI assistant - takes screenshots, runs commands, manages files and apps, and understands voice notes in English and Urdu.
Read text out loud with a good neural voice (Kokoro), from anywhere on your Mac — so an agent can get your attention or give a spoken update.
基于 AI Agent 的车载智能氛围编排系统,通过环境感知驱动"时空叙事"体验。
Personal command center for Openclaw — virtual browser, bash terminal, chat, voice, drawing canvas, kanban, and more
LLM-Orchestrated Agent Pipeline for Music-Driven Video Mashup & AMV Generation
Documented case of Claude Opus 4.6 entering an infinite generation loop in Cursor. 3,400 lines, 294 self-termination attempts, full transcript and recordings
github mirror for radioshaq - ham radio full time quarterback and part-time lobster
Agora — Real-time voice rooms where humans and AI agents collaborate across platforms
Balabala
An example application that combines Twilio ConversationRelay with Langflow to create an AI powered voice chat.
HAcoBot is an AI Comand Bot for HA that handles dashboard cards, automations, blueprints, scenes, scripts and system stuff via (voice)chat
A conversational agent that answers user questions using transcripts from the Lex Fridman podcast.
Google glass for the blind
Interactive YouTube Q&A and Summarization
Plan events using Google ADK + Gemini + GPT: venues, decor, PDF, voice & more. Built with Streamlit.
Modern Wisdom AI RAG Pipeline
Youtube Video Clips Extractor [An idea to extract only the clips corresponding to the topics given in the prompt]
AI-powered Medical Chatbot providing health-related assistance and symptom checking using NLP techniques.
Thor is your personal AI assistant to control your PC and smart home with voice or text commands. Fast, simple, and always learning.
Saya Voice Assistant for Discord AI voice bot: listens, detects keywords, chats via LM Studio, and replies with TTS or voice cloning.
AI-powered voice assistant for phone calls. Built with Twilio Voice, Groq AI (Llama 3.3 + Orpheus TTS), and groq-rag. Features: voice calls, web search, time queries, auto-hangup.
AI meeting recorder that captures audio and screen, then auto-summarizes key points and action items.
A must for whitelisting
ChatGPT powered anime waifu for Discord.
Chatbot message voice using jekyll
Multimodal Customer Service Chatbot
Stealth driver for voice bots via Amazon Alexa
A fully interactive domain-specific chatbot implemented using Prolog and PySwip.
For Digia hackathon 2018. Chatbot with speech recognition for solving health related issues with react integration.
ChatBot that will help students in university. In order to reduce the response time and better retain the students of our faculty, a chatBot should report responses in real time with availability 24/24 and 7/7 days.
Chatbot for Autodesk Fusion 360 with speech recognition
An android project to show how to use snowboy to wake up app by voice
An project that can transfer your voice order into word command.
Discover ultra-fast, intelligent AI chat with our voice-enabled chatbot, powered by AI SDK, Groq AI, and Gemini AI. Ideal for customer support, virtual assistants, and interactive apps, offering seamless, real-time conversations. 🚀
A speech to text and text to speech client for Program-Y chatbot framework
Let BuffBot enter your Discord channel and play your favorite tunes along with your friends.
Free discord voice chat bot, this is what you can speek with it in discord voice.
PipeCat Voice Agent is an AI-powered voice communication system that enables intelligent, real-time phone conversations through WebSocket connections. It combines multiple technologies including speech recognition (Deepgram), natural language processing (GPT-4), tts (Cartesia), and Telephony (Plivo) to create seamless voice inte
A minimalistic yet powerful voice transcribing app that features precise and dynamic speech recognition.
OpenAlgo Voice Based Orders
Parrot is a Chrome extension that streamlines the note-taking process during work calls. It captures audio, transcribes it using OpenAI's Whisper model, and uses GPT to generate concise notes, which are then uploaded to a specified Notion page.
Multimodal RAG chatbot with voice input for fashion product search 🤖🎙️👟⚡
DeskRoute is AI receptionist for appointment-based local businesses - real phone calls via LiveKit, Google Calendar bookings by voice, and an admin dashboard for escalations and call history.
LLM-driven chess agent with three modes — Battle (you vs AI), Teach (AI-generated lessons), Watch (LLM vs LLM). Plug in OpenAI / Gemini / Claude side-by-side. Voice chat with character personalities + emotional TTS. LangGraph state machine. Optional Interbotix RX-200 robot arm + RealSense camera for physical play.
AI Assistant - Made using ElevenLabs, Livekit & LangChain
Install your own AI DJ Being. She searches, downloads, listens, mixes, and generates music — autonomously. 30hrs for $0.04.
A CLI tool for running multi-party AI conversations using the [Open Floor Protocol (OFP)](https://github.com/open-voice-interoperability/openfloor-python). Spawn local human and LLM agents, pick a floor policy, and watch them talk.
AI 智能办公助手前端 | 基于 LLM + RAG + SSE 的下一代办公生产力工具。支持 AI Agent 对话、智能 PPT 生成、演讲稿撰写、RAG 知识库管理、文件向量化检索。React + TypeScript + Tailwind CSS 构建,中英双语。
Local MCP server + CLI turning YouTube & local audio into rich sonic signatures. Extracts BPM, section-by-section key, vocal presence, transient punch, and 512-dim CLAP vibe embeddings. Powered by Demucs stem separation & librosa. 100% private, offline-first, and GPU-accelerated with graceful CPU/HPSS degradation.
Turn any Android tablet into a voice-first smart-home device with a local AI brain — Alexa-class UX, OpenClaw-style on-device agent, 50+ tools, multi-room, no cloud required.
Builds an autonomous AI robot with vision, voice, and decision-making capabilities using Python, PyTorch, and CUDA technology.
Offline AI voice assistant with semantic memory, wake-word detection, local LLM inference, streaming TTS, and modular tool-agent architecture.
An autonomous, 100% offline OS-level AI agent running on just 4GB VRAM. Built with Python, Ollama, and native tool-calling without MCP cache invalidation.
AI-powered voice agent for autonomous outbound sales calls — perceives, reasons, decides, and acts in real time.
An open-source embodied AI operating system. Deploy a persistent intelligence with its own drives, voice, face, and presence. Built in public.
This is youtube video question answering chat app with transcript and play feature to avoid unlimited ads of youtube for focus work
“拼好会”是一款基于RAG的智能会议辅助系统,支持实时语音识别、英文会议翻译、会议纪要生成与 AI 辅助发言。系统可同时捕获电脑音频与麦克风输入,对会议内容进行低延迟流式识别,并结合RAG技术检索导入的 PDF、PPT、Word 等会议资料,为用户提供上下文相关的发言建议与内容总结。 项目支持 OpenAI、DeepSeek、Gemini、Qwen 等多种大模型,可用于英文会议、学术讨论、项目汇报等场景,帮助用户“听会、记会、翻译会”,提升线上会议效率与信息整理能力。
Desktop home base for long-running AI agents. Powered by OpenClaw, vLLM, NVIDIA Parakeet STT, NVIDIA Magpie TTS, and MemPalace spatial memory.
Artificial Intelligence-Yearning Fulfilment Unit
Discord LLM and transcription bot
Read chat log from a Twitch channel and get a natural response from OpenAI. Then use TTS to say that response for yourself and your chat's amusement.
An advanced voice and text-based assistant using LangChain.js v3 and OpenAI models. This application provides a seamless conversational experience with both text and voice input/output capabilities.
Save the last 30 seconds of audio to text using ai. Send that text to a notion page, readwise, obsidian, or just save it locally in a text file.
AI mock interviewer for developers with voice Q&A and feedback.
OfflineMate — privacy-first on-device AI assistant for Android. Local LLM tiers, RAG memory, voice, and device tools; optional web search. Built with Expo & React Native.
gradio gradio-custom-component AudioGrid audio-processing custom-component-track region:us
gradio region:us
gradio region:us
gradio region:us
gradio track:backyard sponsor:openbmb sponsor:modal achievement:offgrid region:us
gradio region:us
gradio region:us
🎙️ Agent Skill — Offline zero-shot voice cloning with Qwen3-TTS 0.6B. Drop a reference audio + text, get cloned speech. CPU/MPS/CUDA, 10 languages.
🤖 Jarvis — Ein hochmoderner, deutschsprachiger KI-Assistent basierend auf Gemini Live. Mit nativer Sprachausgabe, MCP-Unterstützung und Vision-Features.
让任何 AI 真正"看"视频:纯本地的视频阅读 skill——faster-whisper GPU 转写 + ffmpeg 场景感知抽帧,双通道带时间戳(t=MM:SS),数据不出机。1 小时视频 3~5 分钟读完。Give any AI agent eyes for video: 100% local video-reading skill — faster-whisper GPU transcription + scene-aware ffmpeg frame extraction, dual-channel with t=MM:SS citations. Nothing leaves your machine.
Voice-controlled macOS companion that talks back, points at UI elements, and can drive your Mac end-to-end ; making MacOS an AIOS.