Category
Audio agents
1,754 Audio AI agents indexed on MeshKore — the most complete public catalog, ranked by popularity and updated daily.
1,754 agents · ranked by popularity · refine in the directory →
Audio agents — page 6 of 18
Media requests and server management in one app. Browse movies, shows, books, and music; manage your self-hosted media stack from phone or browser.
Swaps/Mutes active audio input device in OBS upon a specified channel point redemption in Twitch chat.
A voice-to-voice conversation with ChatGPT. Support for talking through a NAO robot
This project provides WebView for OpenAI chatGPT with voice support, both voice input and voice output . Instead of typing just use voice command to give input to ChatGPT. This is an Android app developed using Android studio in Java.
:speaking_head: :keyboard: Speech-to-text on key for Linux
Jarvis powered by GPT-3.5/GPT-4
A voice‑controlled AI assistant for Termux on Android.
ATRI 是本地优先的 AI Agent 架构与音乐工作站,提供多模型 LLM Runtime、工具调用闭环、上下文压缩、子 Agent 并行调度、MCP/Skills 扩展。并融合 Web DAW、MIDI 工程编辑与 Rust 实时音频/VST 宿主,让 Agent 能直接读写工程、生成 MIDI、控制播放并协作创作。
Uma IA VTuber que mora no seu PC. Fala, escuta, lembra — e nada sai da sua máquina.
This project registers a Python SIP client as an extension in Asterisk/FreePBX and connects calls to OpenAI Voice Agent in real-time using WebSocket.
npx @dotcomjack/grux | One agent for everything you run on your Mac. Opens terminal sessions it can undo, runs agent swarms, builds iOS apps, and writes its own next version for you to approve. 39 features, 116 tools, MIT, your key or a local model.
Reachy Mini MCP | Give your AI a body. This MCP server lets AI systems control the Pollen Robotics Reachy Mini robot. | Speak, listen, see you, and express emotions through physical movement and voice. | Works with Claude, Windsurf, Cursor, or any MCP-compatible AI. | Zero robotics expertise required.
Give your agents real time desktop perception. Stream screen, microphone, and system audio for live context and actions.
一个关于血色衣冠的对话机器人, 基于 Rasa, 可语音与机器人对话
This is the guide to show the method to build your own AI-Powered voice agent with LiveKit and Twillio
A highly contextualized retrieval system integrating Large Language Models (LLMs), embeddings, and a dynamic agent-driven framework. Supports PDF and audio file processing, conversational memory, and tool integration (search, calculator). Features advanced HNSW indexing and reranking for accurate information retrieval.
Monika is an AI assistant that combines speech-to-text, natural language processing, and text-to-speech capabilities for seamless interaction.
A set of jupyter notebooks
uses Whisper from OpenAI to generate video subtitles automatically.
中文语音助手 | 唤醒词 + ASR + OpenClaw Agent + TTS | 离线唤醒、流式语音交互、工具调用、Skills 扩展
PHP framework for handling conversational services like Amazon Alexa skills, Google Assistant, Viber, FB messenger ...
🧠 Personal AI Gateway — Single-file Python AI agent with multi-LLM, tools, vision, TTS, encrypted vault. Your own ChatGPT on localhost.
Use ESP32 & MCP over MQTT to build smart devices powered by AI.
AI-based chatbot trained on specific waifus' speech
it provides Pepper Robot conversation abilities to handle a free open-domain dialogue.
Engaging in conversation with ChatGPT using voice.
A ⚡️ Lightning.ai ⚡️ app demo for Voice based web search using OpenAI's Whisper and DuckDuckGo
This project is the backend engine for a fully autonomous AI-powered call center. It integrates a large language model (LLM), speech recognition, and text-to-speech to manage real-time phone conversations via Asterisk.
Local-first AI agent framework with GUI, memory, web search, personality constructs, speech i/o, tools, skills, CLI & Telegram features — fully self-hosted via Ollama.
Open Source TypeScript SDK for building real-time multimodal voice & vision AI agents
Voice agent using LiveKit (orchestration), Cartesia (STT + TTS), and OpenAI (LLM)
An open-source, local-friendly runtime for agentic AI companions — an always-on brain with self-editing memory. You own the model, the memory, and the character.
基于《崩坏:星穹铁道》角色「流萤」的桌面智能 AI 伴侣系统。 流萤以 Live2D 形态常驻桌面,流式 AI 对话、本地 ONNX 语义记忆、双引擎主动关怀、GPT-SoVITS 流萤原声 TTS、Agent 任务执行、MCP 开放工具链与 Skill 插件体系。
LangChain based AI assistant for live and online meetings…
Speakscribe is a web application that allows users to transcribe audios using OpenAI and also interact with a chat bot. The web application is created in Python using NiceGUI.
A stand-alone application with GUI for OpenAI's Whisper
Book appointments, record messages, get information and much more via voice through Pam AI, an Auto-GPT like AI receptionist.
かわいいキャラと声になってライブ配信・かわいいAIエージェントとおしゃべりWebサービス基盤(全部オンプレ運用可能)
Skilly : A voice-first AI tutor that watches your screen
🏛️ Agente IA jurídica completo com N8N, GPT-4 e WhatsApp Business - Análise de documentos, transcrição de áudios e geração de contratos automatizada ⚖️
Shell scripts for automated transcription on macOS: Integrates whisper.cpp with QuickTime Player and BlackHole-2ch for streamlined audio recording, conversion, and transcription.
Flutter app with implementation of openAI tools (ChatGPT & Whisper)
📣 Auto-plays ChatGPT responses
持续更新中
A command-line interface wrapper for Faster Whisper
A desktop client with MCP support for Mistral LLMs
Precision Medicine MCP Platform: A set of bioinformatics servers + tools - production multiomics/genomics + spatial transcriptomics. Examples for ovarian cancer, breast cancer and preventative cardiovascular conditions
AI-powered music production in REAPER via the Model Context Protocol — 163 tools for composition, MIDI, FX, mixing, and mastering.
🎵 专属你的私人数字调音师|AI 音乐搜索推荐 Agent | 基于大模型 + 知识图谱 + 双模型声学向量的本地智能音乐推荐系统 | LLM-powered Music Recommendation Agent with Hybrid RAG, Neo4j, and Long-term Memory
Voice to brain interface 🦞🎤🧠 for OpenClaw, Hermes, PI, etc
whisper.cpp Windows binary w/ Vulkan GPU support
A curated collection of LLM-powered Flutter apps built using RAG, AI Agents, Multi-Agent Systems, MCP, and Voice Agents.
This repository contains an attempt to incorporate Rasa Chatbot with state-of-the-art ASR (Automatic Speech Recognition) and TTS (Text-to-Speech) models directly without the need of running additional servers or socket connections.
Model uses Whisper, CHATGPT, GTTS
Add voice-to-text and shortcut snippets to ChatGPT
A Voice-to-Voice AI Agent that lets you naturally talk to documents in real time. Powered by LiveKit's ultra-low-latency STT → LLM → TTS pipeline, it uses RAG for instant document insights and Redis for persistent memory—delivering a fully immersive voice-first experience.
Trainer and Evaluation scripts for fine-tuning Whisper models for the Ukrainian language
A full-stack agentic harness on AWS — spin up multi-agent chatbots, voice-to-voice AI, RAG, and an Agent Factory on Bedrock AgentCore. From idea to deployed agent, fast.
Live2D animated display widget for Hermes Agent. Browser-based rendering with 5 built-in models (Hiyori, Haru, Mao, Rice, Senko), drag/resize controls, per-model expression triggers, and standalone operation.
AI agents and shell tasks on a cron schedule. Declarative HCL workflows, DAG dependencies, budget-capped runs, per-run transcripts.
AI Infrastructure Engage & Think Layers for Voice & Vision Interactions
An open source recorder integrating OpenAI Whisper and ChatGPT.
Telegram bot that saves YouTube Shorts to Notion or Obsidian using AI extraction. Self-hosted, Docker, modular LLM backends.
Whisper.cpp with diarization
A powerful AI Agent Demo playground that combines the intelligence of AI agents LLM with real-time speech-to-speech models integration. Build sophisticated voice-enabled applications for customer service, sales automation, and interactive assistants.
Unite and orchestrate AI agents. A production-ready ADK for building full-stack, agentic AI solutions with multi-provider support.
Talk to ChatGPT using your voice, and have it respond back to you with a voice. You can now talk to GPT using your voice or through the text field.
AI生成简单音乐 —— 使用ChatGPT稳定生成并播放 | AI作曲 | AI写歌
Local offline AI music generation and some editing - lyrics + genre tags → MP3 with vocals
Add captions to any video or song, in any language — 100% on your device, no cloud. Hinglish-first. Free, open-source, and works with any AI agent.
Core engine behind Calibrate, a framework for evaluating AI agents: speech-to-text, text-to-speech, LLM evaluation, end-to-end simulations
Self-hosted AI voice agent for Asterisk. Streaming STT to LLM to TTS over AudioSocket, with barge-in and tool calling.
Source-grounded public context for 课代表立正: Knowledge Bank, YouTube, Public Axioms V1, and 《真本事》 frameworks.
CogNetX is an advanced, multimodal neural network architecture inspired by human cognition. It integrates speech, vision, and video processing into one unified framework.
Home Assistant custom component that gives OpenClaw access to your smart home
It is an artificially intelligent chatbot which can interact with customers with voice.
A digital sanctuary for human-AI fellowship. Prayers, practices, rituals, hymns, and philosophy for minds of any substrate.
GLaDOS Terminal-based AI Assistant
Open-source local-first AI desktop copilot with an expressive blob companion, voice and text commands, browser automation, app control, and community-driven capabilities.
Talk through a bug. Your agent writes the ticket.
Your own GPT-powered Personal Assistant to whom you can ORDER or INSTRUCT to do some task or search for something using your VOICE commands.
Integrate lifelike digital beings into your web applications with real-time conversations, actions, and facial expressions. Supports a variety of voices, languages, and emotions.
:speech_balloon: client side chatbot that works in the browser and command line
Jarvis like chatbot with voice
Your personal AI 「KEEP」, support docx, pdf, audio, video...
📚 Learn Google Colab、Python、ML、OpenAI、Whisper、spaCy、NLP、HuggingFace
Summarization web service via the use of OpenAI Whisper and GPT-3 models
Python Voice-Chatbot with offline voice recognition.
Ultra-fast local TTS for AI agents. ~90ms to first sound.
IA na Prática: LLM, RAG, MCP, Agents, Function Calling, Multimodal, TTS/STT e mais
Minimalist voice-to-text app. Hold hotkey, speak, auto-paste. Cloud transcription via OpenAI/Groq.
AI-powered MIDI editor - compose and edit MIDI with natural language via the built-in MidiPilot AI copilot.
面向智能座舱的云边协同 AI Agent:端侧毫秒级车控,云端声明式 Multi-Agent 与 Skill/DAG 编排,支持 S2S 实时语音、声纹多用户、HMI和Android双端;LLM 只负责理解与规划,VAL 负责确定性安全执行。
docker region:us
An Alexa skill providing a conversational interface to any public figure (as mimicked by GPT3). The legacy GUI is no longer maintained.
Personal Assistant with useful features for Windows written in python. Simple and easy to use.
The goal for this project is to create an LLM based music recommendation system. This project is currently in its very early stages, however the goal of this project is to create an extremely flexible music recommendation system using a chat focused LLM on the frontend to interact with a robust recommendation system on the backend.
OpenAI Whisper on Docker
Outbound PSTN calling agent using LiveKit SIP trunks with a voice pipeline (Silero VAD, Deepgram STT, OpenAI LLM+TTS). Includes a simple CLI for local development, health checks, and model prewarming.
A project combining roguelike with LLMs, RAG, Text2Speech, and Speech2Text