Category
Audio agents
1,754 Audio AI agents indexed on MeshKore — the most complete public catalog, ranked by popularity and updated daily.
1,754 agents · ranked by popularity · refine in the directory →
Audio agents — page 4 of 18
AI Voice Assistant: Talk to an AI agent that helps you with event scheduling, contact management, accessing your knowledge base, and web searches using simple voice commands
Voice Prompts, GPT-4o prompts, Voice Agent Prompts, ChatGPT Prompts, HumeAI Prompts
Utsuwa is an open-source alternative to Grok Companion. This is a platform where you can have a virtual AI waifu that learns and grows with you, bundled with optional mechanics inspired by Japanese dating sim games.
Simple GUI around whisper.cpp for voice-to-text on Linux
Voice-powered AI assistant platform — connect any LLM, any TTS, with a live web canvas, music generation, and agent orchestration using openclaw. Install: npx openvoiceui setup
A comprehensive Model Context Protocol (MCP) server that enables AI agents to create fully mixed and mastered tracks in REAPER with both MIDI and audio capabilities.
星⭐收藏家:Bilibili 收藏视频 Markdown 多 Agent 本地编排工作台
Whisper is an automatic speech recognition (ASR) system Gradio Web UI Implementation
Deploy generative AI agents in your contact center for voice and chat using Amazon Connect, Amazon Lex, and Amazon Bedrock Knowledge Bases
Home Assistant voice stack integration for Hermes Agent — on-device voice control, wake word → STT → LLM → TTS → HA media_player. No cloud. No latency. No subscription.
OmniAgent (ICML 2026): the first native omni-modal agent for active video perception — a 7B agent that beats Qwen2.5-VL-72B with 73% fewer frames on LVBench.
Android-first Flutter client for self-hosted Hermes Agent — chat, Bots, Voice and private remote control.
Twitch livestream bot that can control colors for overlays from Stream Elements, play sound effects, handle custom rewards (like text-to-speech) and more!
Omnigram is a Flutter-based file reader and audiobook . It accommodates EPUB and PDF and offers audiobook functionality, supporting TTS model and other AI chat technologies for enhanced reading experiences
Zorv AI — 安卓 Android 端开源 AI Agent 智能体助手。多模型对话、人格系统、记忆、语音全双工 TTS/STT、定时任务、可扩展工具链、飞书/QQ/微信接入。Kotlin + Jetpack Compose 开发。
Leopard Chat UI - A Teneo Chat Client based on Vue and Vuetify
Voice AI agent starter kit with Groq, Llama 4, and (optionally) Twilio
Shadow AI: stealth AI assistant for restricted/locked-down environments, enabling cross-device interaction over LAN, screen & audio capture, real-time AI inference, and low-friction automation. Supports OpenAI, Claude, Gemini, Kimi, Antigravity.
An Android ChatBot powered by IBM Watson Services (Assistant V1, Text-to-Speech, and Speech-to-Text with Speaker Recognition) on IBM Cloud.
A voice assistant application built with the LiveKit Agents framework, capable of using Model Context Protocol (MCP) tools to interact with external services
Production-ready audio and video transcription app that can run on your laptop or in the cloud.
Audio-Oscar is a multi-agent framework for generating long-form, controllable audio from complex audio scene descriptions.
Your feedback layer for AI collaboration. Point at anything on your Mac - text, screenshots, web elements, or voice - and your agent reads and resolves your comments over MCP.
Talk to your second brain personal assistant using speech 🧠
gradio region:us
gradio region:us
A minimalistic automatic speech recognition streamlit based webapp powered by OpenAI's Whisper "State of the Art" models
Local-first voice memory assistant — capture speech, transcribe live, and ask a private RAG chatbot grounded in what you've actually said. No cloud required.
Pipecat framework based orchestrator for building real-time, voice-enabled, and multimodal conversational AI agents
A modern, serverless web application that connects users to a Microsoft Copilot Studio agent for booking and managing appointments via chat and voice input.
A voice activated interface for your custom AI Agent.
Open-source, local-first desktop AI agent for voice, gesture, gaze, browser and system automation with permission gates and verification.
Turn Hermes into an autonomous WhatsApp manager. Take full control of your messaging. Forget typing: schedule messages, send voice notes, share PC files, and search your contacts instantly. Your WhatsApp, on total autopilot.
🐬 A pink qFlipper fork with LOTEI — a 100% local AI dolphin (Ollama) that chats, talks, watches your Flipper's screen, and one-click-installs custom firmware.
A personal AI OS for macOS, powered by Claude.
Effect building blocks for agentic ai
An AI buddy that lives on your Windows PC, she sees your screen, points, and actually does things. Open-source Clicky for Windows, built on Claude.
Audio to summary with openAI Whisper & GPT 3.5/4 using streamlit
The first intelligent AI agent running natively on a smartwatch. NullClaw + Vosk offline STT + Claude on Galaxy Watch
A voice assistant powered by OpenAI's ChatGPT language model, currently available in six languages.
Modular Discord bot using JavaScript and optionally Python and Lua.
Modern Desktop Application offering a suite of tools for audio/video text recognition and a variety of other useful utilities.
JARVIS AI Assistant 🤖 A virtual assistant project inspired by Tony Stark's JARVIS, powered by speech recognition, AI chat, web browsing, and more. Features: 🎙️ Voice-based interaction using speech recognition. 🧠 AI-powered chat with OpenAI's language model. 🌐 Web browsing capabilities to open websites. 🎵 Music playback. ⏰Current time display
A multi engine TTS & LLM edge computing playground with audio book features and more!
Official Repo for "AgenticASR: Refining Speech Recognition in Real-World Scenarios via an Agentic Approach"
gradio region:us
A distributed multi-modal agent orchestration framework implementing advanced natural language processing, computer vision, and audio processing capabilities through a microservices architecture.
Practice your job interview skills with AI-powered voice interviews, featuring real-time feedback and dynamic questions
AIOPE — 70-tool on-device AI agent for Android. Linux terminal, realtime voice, browser automation, SSH, on-device RAG, dynamic UI, MCP. BYOK, any model. By XNet Inc.
It's like ChatGPT for videos.
🤖 AI Dialer ☎️ – Autonomous Voice Agent for Appointment Scheduling 🗓️
A distributed multi-modal agent orchestration framework implementing advanced natural language processing, computer vision, and audio processing capabilities through a microservices architecture.
Free, open-source Opus Clip alternative that runs entirely on your own PC. Turn long streams and videos into ready-to-post vertical Shorts: multimodal clip detection, speaker-aware face tracking, editable captions, AI edit chat. No uploads, no subscription, no watermark.
Self-hosted realtime voice gateway for Hermes Agent. Keep talking while Hermes runs background tasks through local, Gemini Live, and OpenAI Realtime adapters.
Automatically generate subtitles from an input audio or video file using OpenAI Whisper
A painted music video for "Functional Emotions", made by Claude Opus 5.5: custom GPU brushstroke renderer, storyboard, and 7 parallel chapter agents.
A curated collection of tools to aid transcriptionists and subtitlers.
Simple RAG tutorials that can be run locally or using Google Colab (only Pro version).
A curated list of voice AI agent frameworks, tools, resources, and best practices
Multi-agent debate system for Hermes Agent — spawn custom-composed expert agents to debate any question in structured rounds, producing visible transcripts and synthesized recommendations. Pure SKILL.md, zero new infrastructure.
Run your Mac with your voice — gpt-realtime-2 + agent-desktop. Clone it and teach your agent any app.
Download music and play on Telegram for free
This project demonstrates a multi-agent system using Google's Agent Development Kit (ADK), Agent2Agent (A2A) and Model Context Protocol (MCP). that integrates Notion for information retrieval and ElevenLabs for text-to-speech conversion.
A distributed multi-modal agent orchestration framework implementing advanced natural language processing, computer vision, and audio processing capabilities through a microservices architecture.
AI Voice Assistant for Solana
Just an .exe that can be used for those unable to build whisper.cpp in Windows.
OpenAI realtime audio with WebRTC
重生之我是 AI 打工人。前世,我的身份默默无闻,来去匆匆,不知道自己将在何地出生。然而,命运给予了我难得的机会,让我重生为一名 AI 打工人。
Home Assistant custom integration — Mistral AI as conversation agent and Voxtral as speech-to-text engine. ⚠️ Please note this is not an officially supported integration and is not affiliated with Mistral AI in any way.
A WebView2 UI framework for foobar2000 - build the entire player in web tech (Vue/React) with native Windows 11 Mica/Acrylic effects. Ships a TypeScript SDK and an MCP server for AI agents.
Open source Android AI agent. Control your device with voice & AI. Lightweight, any LLM, no root.
A basic voice agent built with Python agents framework
Build safe on-brand voice and chat agents with structured flows, guardrails, and compliance
Open source CLI agent that generates complete lesson bundles in your teaching voice. 48+ tools, zero-touch 9-file output, local-first privacy.
Streamlit Audio Transcription with OPENAI's Whisper Ai: An interactive Streamlit app demonstrating real-time audio transcription using OPENAI's Whisper Ai.
🤖 Home Assistant custom integration for using webhook-based systems (e.g. n8n) as conversation agents and voice assistants.
The Open Source Voice Agent Platform. Orchestrate ultra-low latency AI pipelines for real-time conversations over WebRTC.
LLM integration into Crusader Kings 3
A beautiful, native macOS desktop application for transcribing audio and video files using whisper.cpp
Blazor Server playground for OpenAI using Cledev.OpenAI .NET library
This Guidance provides a sample foundation for building real-time voice AI agents on AWS. It demonstrates how to build a voice assistant that handles phone calls and voice interactions in web and mobile applications via WebRTC.
Reusable skills, scripts, and production workflows for agent-led video creation with HyperFrames, FFmpeg, captions, audio sync, and render QA. Community project from real HyperFrames video work; not an official HyperFrames or HeyGen repository.
livekit agent plugins
AI Boardroom: A fully configured multi-agent corporate simulation platform featuring the real-time Jarvis Voice Widget.
MusicAgent is a MAS (Multi Agent System) that programs songs in Sonic Pi. It uses generative AI to generate song structures, arrangements, lyrics, ... based on user preferences.
This Guidance provides a sample foundation for building real-time voice AI agents on AWS. It demonstrates how to build a voice assistant that handles phone calls and voice interactions in web and mobile applications via WebRTC.
贾维斯语音助手(JARVIS-Voice Assistant),妈妈我再也不用羡慕钢铁侠了😭。
From-scratch voice agents in Python: end-to-end speech pipelines, runnable chapters, and a small shared library. Local models, explicit streaming behavior.
Voice Agent Framework
Voice-enabled JARVIS mission-control dashboard for Hermes Agent
Agent Skill — WeChat 4.x chat decrypt & query (macOS + Windows): MCP read/search, export, voice transcription.
久远:一个开发中的大模型语音助手,当前关注易用性,简单上手,支持对话选择性记忆和Model Context Protocol (MCP)服务。 KUON:A large language model-based voice assistant under development, currently focused on ease of use and simple onboarding. It supports selective memory in conversations and the Model Context Protocol (MCP) service.
ALICE and its prior work, Voice2Action: Language Models as Agent for Efficient Real-Time Interaction in Virtual Reality
Your private AI companion that lives on your wrist. Complete local AI assistant with emotional intelligence.
Local AI Music Discovery for Lidarr - Brainarr is a privacy-focused AI-powered import list plugin for Lidarr that generates intelligent music recommendations using local AI models. It exclusively supports local providers (Ollama, LM Studio) to ensure your music preferences never leave your network.
The self-evolving agentic framework for bioinformatics
Menu-bar app that turns the Teenage Engineering EP-2350 Ting mic into a voice and button controller for AI agent harnesses
An OpenAI's Whisper-based full-stack project to transcribe audio and video files using React & Django.