Category
Audio agents
1,474 Audio AI agents indexed on MeshKore — the most complete public catalog, ranked by popularity and updated daily.
1,474 agents · ranked by popularity · refine in the directory →
Audio agents — page 4 of 15
A voice assistant powered by OpenAI's ChatGPT language model, currently available in six languages.
Audio to summary with openAI Whisper & GPT 3.5/4 using streamlit
An open-source, privacy-first AI System Control Agent (JARVIS-like) using voice and hand gestures
Shadow AI: stealth AI assistant for restricted/locked-down environments, enabling cross-device interaction over LAN, screen & audio capture, real-time AI inference, and low-friction automation. Supports OpenAI, Claude, Gemini, Kimi, Antigravity.
Turn Hermes into an autonomous WhatsApp manager. Take full control of your messaging. Forget typing: schedule messages, send voice notes, share PC files, and search your contacts instantly. Your WhatsApp, on total autopilot.
Modular Discord bot using JavaScript and optionally Python and Lua.
Modern Desktop Application offering a suite of tools for audio/video text recognition and a variety of other useful utilities.
Home Assistant voice stack integration for Hermes Agent — on-device voice control, wake word → STT → LLM → TTS → HA media_player. No cloud. No latency. No subscription.
gradio region:us
An AI buddy that lives on your Windows PC, she sees your screen, points, and actually does things. Open-source Clicky for Windows, built on Claude.
A distributed multi-modal agent orchestration framework implementing advanced natural language processing, computer vision, and audio processing capabilities through a microservices architecture.
Practice your job interview skills with AI-powered voice interviews, featuring real-time feedback and dynamic questions
JARVIS AI Assistant 🤖 A virtual assistant project inspired by Tony Stark's JARVIS, powered by speech recognition, AI chat, web browsing, and more. Features: 🎙️ Voice-based interaction using speech recognition. 🧠 AI-powered chat with OpenAI's language model. 🌐 Web browsing capabilities to open websites. 🎵 Music playback. ⏰Current time display
🐬 A pink qFlipper fork with LOTEI — a 100% local AI dolphin (Ollama) that chats, talks, watches your Flipper's screen, and one-click-installs custom firmware.
It's like ChatGPT for videos.
🤖 AI Dialer ☎️ – Autonomous Voice Agent for Appointment Scheduling 🗓️
A distributed multi-modal agent orchestration framework implementing advanced natural language processing, computer vision, and audio processing capabilities through a microservices architecture.
A distributed multi-modal agent orchestration framework implementing advanced natural language processing, computer vision, and audio processing capabilities through a microservices architecture.
A multi engine TTS & LLM edge computing playground with audio book features and more!
🤖 Advanced AI-powered virtual assistant with voice recognition, face authentication, phone integration, and intelligent automation capabilities.
Automatically generate subtitles from an input audio or video file using OpenAI Whisper
A curated collection of tools to aid transcriptionists and subtitlers.
Download music and play on Telegram for free
This project demonstrates a multi-agent system using Google's Agent Development Kit (ADK), Agent2Agent (A2A) and Model Context Protocol (MCP). that integrates Notion for information retrieval and ElevenLabs for text-to-speech conversion.
Local MCP server for stateful, fail-closed Logic Pro control and live project readback.
AI Voice Assistant for Solana
Simple RAG tutorials that can be run locally or using Google Colab (only Pro version).
Multi-agent debate system for Hermes Agent — spawn custom-composed expert agents to debate any question in structured rounds, producing visible transcripts and synthesized recommendations. Pure SKILL.md, zero new infrastructure.
OpenAI realtime audio with WebRTC
重生之我是 AI 打工人。前世,我的身份默默无闻,来去匆匆,不知道自己将在何地出生。然而,命运给予了我难得的机会,让我重生为一名 AI 打工人。
Just an .exe that can be used for those unable to build whisper.cpp in Windows.
A basic voice agent built with Python agents framework
Home Assistant custom integration — Mistral AI as conversation agent and Voxtral as speech-to-text engine. ⚠️ Please note this is not an officially supported integration and is not affiliated with Mistral AI in any way.
Blazor Server playground for OpenAI using Cledev.OpenAI .NET library
Build safe on-brand voice and chat agents with structured flows, guardrails, and compliance
Open source CLI agent that generates complete lesson bundles in your teaching voice. 48+ tools, zero-touch 9-file output, local-first privacy.
Streamlit Audio Transcription with OPENAI's Whisper Ai: An interactive Streamlit app demonstrating real-time audio transcription using OPENAI's Whisper Ai.
🤖 Home Assistant custom integration for using webhook-based systems (e.g. n8n) as conversation agents and voice assistants.
The Open Source Voice Agent Platform. Orchestrate ultra-low latency AI pipelines for real-time conversations over WebRTC.
The first intelligent AI agent running natively on a smartwatch. NullClaw + Vosk offline STT + Claude on Galaxy Watch
Official Repo for "AgenticASR: Refining Speech Recognition in Real-World Scenarios via an Agentic Approach"
Run your Mac with your voice — gpt-realtime-2 + agent-desktop. Clone it and teach your agent any app.
星⭐收藏家:Bilibili 收藏视频 Markdown 多 Agent 本地编排工作台
久远:一个开发中的大模型语音助手,当前关注易用性,简单上手,支持对话选择性记忆和Model Context Protocol (MCP)服务。 KUON:A large language model-based voice assistant under development, currently focused on ease of use and simple onboarding. It supports selective memory in conversations and the Model Context Protocol (MCP) service.
livekit agent plugins
LLM integration into Crusader Kings 3
MusicAgent is a MAS (Multi Agent System) that programs songs in Sonic Pi. It uses generative AI to generate song structures, arrangements, lyrics, ... based on user preferences.
This Guidance provides a sample foundation for building real-time voice AI agents on AWS. It demonstrates how to build a voice assistant that handles phone calls and voice interactions in web and mobile applications via WebRTC.
ALICE and its prior work, Voice2Action: Language Models as Agent for Efficient Real-Time Interaction in Virtual Reality
A beautiful, native macOS desktop application for transcribing audio and video files using whisper.cpp
An OpenAI's Whisper-based full-stack project to transcribe audio and video files using React & Django.
AI Boardroom: A fully configured multi-agent corporate simulation platform featuring the real-time Jarvis Voice Widget.
👻 kwami.io | A 3D Interactive AI Companion Library for creating engaging AI companions with visual (blob), audio, and AI speech capabilities.
This WebMCP Music Composer project is a functional demonstration of the WebMCP Protocol, illustrating how AI agents can interact with local browser contexts (tools) to achieve complex workflows autonomously.
Efficient AI English Learning: Read & Speak via Web | 英语学习,英语单词,通过 AI 学英语朗读,对话的高效 Web 应用
Reference implementation of an end-to-end voice agent built using the NVIDIA Nemotron models
A Whisper + ChatGPT MagicMirror Module.
Multi-agent TTS production harness: Fish TTS + WhisperX + Claude, with cross-episode memory and auto-fix loop
The self-evolving agentic framework for bioinformatics
Private Open General Assistant Platform
A Python based Voice Assistant like Siri
An AI assistant building SDK in python
Your private AI companion that lives on your wrist. Complete local AI assistant with emotional intelligence.
Multi-Agent Indonesian Fintech Fraud Detection Platform powered by Xiaomi MiMo V2.5 — QRIS, phishing, invoice fraud, WhatsApp scams
streamlit region:us
Generate an engaging podcast based on your document using Azure OpenAI and Azure Speech.
A production-ready voice agent implementation using LiveKit and Python, featuring advanced conversational AI capabilities and optional telephony integration. It provides intelligent turn detection, function calling, comprehensive logging, telephony integration, and audio enhancement.
Spotify Web AI DJ - client side agentic smarts using Gemma 2, two billion parameter LLM, to play what a user wants via natural speech input
A fully local, open-source voice-to-text tool that acts as a system-wide AI dictation layer, converting speech into clean, formatted text.
A modified version of SalesGPT with the addition of TTS, STT, and Twilio to make calls. A Context-aware AI Sales Agent to automate sales outreach
Complete AI Agent System for OpenClaw — memory, self-healing, self-improvement, voice, automation
A fully local, open-source voice-to-text tool that acts as a system-wide AI dictation layer, converting speech into clean, formatted text.
Local AI Music Discovery for Lidarr - Brainarr is a privacy-focused AI-powered import list plugin for Lidarr that generates intelligent music recommendations using local AI models. It exclusively supports local providers (Ollama, LM Studio) to ensure your music preferences never leave your network.
A curated list of voice AI agent frameworks, tools, resources, and best practices
A minimalistic automatic speech recognition streamlit based webapp powered by OpenAI's Whisper
AI Agent Skills toolkit for automated product introduction video generation with Remotion, Playwright, and edge-tts
Video-production toolkit for AI agents: Deno + WebGPU + three.js/TSL render engine with VRM character locomotion, simulations, effects, and a full audio pipeline — real-time GPU rendering with minimal CPU
A Chat Client for LLMs, written in Compose Multiplatform.
Automatic subtitles for DaVinci Resolve with OpenAI Whisper
a tribute to multi-agent pathfinding composed of simple mechanics, borrowed art, and questionable music taste
This is a fun Python project that allows you to chat with a chatbot about the PDF you uploaded. and generate a PDF transcript of the conversation. The project is built using Python and Streamlit framework.
AI native 的跨平台离线语音输入法
This repository shows you how to build real-time, voice-enabled AI agents with Pipecat and Amazon Bedrock
Sky LiveKit Agent Perplexica is a local, free solution integrating LiveKit with advanced internet search. It uses a local Perplexica instance with function calling to retrieve and summarise search results in natural language. Powered by Faster Whisper, Ollama (Qwen 2.5) and Kokoro-82M TTS.
🐍 Personal AI agent in pure Python — OpenClaw reimagined. Memory, RAG, skills marketplace, cron, voice. Telegram/Discord/WhatsApp/Web. Works with DeepSeek, Claude, Gemini, Kimi, GLM, Ollama.
Generate subtitles for all the videos in a folder with OpenAI's Whisper privately in your computer.
fine-tune Whipser model for Taiwanese speech recognition
Python platform for working with LLMs
This is a Next js project that implements a conversational AI Agents using ElevenLabs' SDK. The application features a voice assistant interface that allows users to interact with the AI through voice commands.
Cross-platform Electron app for simultaneously streaming & recording microphone and speaker audio
V.I.S.O.R., my in-development AI-powered voice assistant with integrated memory!
Simple Python audio transcriber using OpenAI's Whisper speech recognition model
MOM AI transcribes audio into meeting summary and generate minutes of meeting. Built using Langchain, OpenAI GPT-3, Open Whisper.
Native iOS app for talking to your OpenClaw agents by voice or text. On-device speech recognition, streaming responses, multi-agent channels.
Tater is a local-first AI platform with built-in LLMs, vision, voice, memory, automation, integrations, and voice satellite support.
Full Apple Music integration for Claude via MCP — search catalog, manage library & playlists
Learning chatbot that can automatically fetch lecture transcript
A WebView2 UI framework for foobar2000 - build the entire player in web tech (Vue/React) with native Windows 11 Mica/Acrylic effects. Ships a TypeScript SDK and an MCP server for AI agents.
A framework for creating voice based agents. Integrations LLMs with speech recognition and text-to-speech