Category
Audio agents
1,754 Audio AI agents indexed on MeshKore — the most complete public catalog, ranked by popularity and updated daily.
1,754 agents · ranked by popularity · refine in the directory →
Audio agents — page 5 of 18
This WebMCP Music Composer project is a functional demonstration of the WebMCP Protocol, illustrating how AI agents can interact with local browser contexts (tools) to achieve complex workflows autonomously.
Efficient AI English Learning: Read & Speak via Web | 英语学习,英语单词,通过 AI 学英语朗读,对话的高效 Web 应用
Reference implementation of an end-to-end voice agent built using the NVIDIA Nemotron models
A Whisper + ChatGPT MagicMirror Module.
🐍 Personal AI agent in pure Python — OpenClaw reimagined. Memory, RAG, skills marketplace, cron, voice. Telegram/Discord/WhatsApp/Web. Works with DeepSeek, Claude, Gemini, Kimi, GLM, Ollama.
Multi-Agent Indonesian Fintech Fraud Detection Platform powered by Xiaomi MiMo V2.5 — QRIS, phishing, invoice fraud, WhatsApp scams
streamlit region:us
Windows 本地优先的个人 AI Agent;A local-first personal AI Agent for Windows — conversations, memory, diary, voice, QQ, Live2D, and screen awareness.
Private Open General Assistant Platform
A Python based Voice Assistant like Siri
An AI assistant building SDK in python
A production-ready voice agent implementation using LiveKit and Python, featuring advanced conversational AI capabilities and optional telephony integration. It provides intelligent turn detection, function calling, comprehensive logging, telephony integration, and audio enhancement.
Multi-agent TTS production harness: Fish TTS + WhisperX + Claude, with cross-episode memory and auto-fix loop
AI Agent Skills toolkit for automated product introduction video generation with Remotion, Playwright, and edge-tts
Full Apple Music integration for Claude via MCP — search catalog, manage library & playlists
AI Meeting Assistant & Real-Time Interview Copilot — 100% local, free, open source
一条命令,把口播毛片变成可二次精修的剪辑工程。
Generate an engaging podcast based on your document using Azure OpenAI and Azure Speech.
Spotify Web AI DJ - client side agentic smarts using Gemma 2, two billion parameter LLM, to play what a user wants via natural speech input
A modified version of SalesGPT with the addition of TTS, STT, and Twilio to make calls. A Context-aware AI Sales Agent to automate sales outreach
AI native 的跨平台离线语音输入法
A fully local, self-hosted replacement for Siri and Alexa. One assistant for voice, vision, memory, automations, smart-home control, music, messaging, and more—running on hardware you control.
The new way to work: Drive a team of AI Agents through screenless voice
Complete AI Agent System for OpenClaw — memory, self-healing, self-improvement, voice, automation
A fully local, open-source voice-to-text tool that acts as a system-wide AI dictation layer, converting speech into clean, formatted text.
AI thinking indicators and a voice-agent orb for React Native and Expo. Six dotted loading animations plus an audio-reactive voice orb, on the UI thread with Skia + Reanimated. Port of Jakub Antalik's thinking-orbs.
A minimalistic automatic speech recognition streamlit based webapp powered by OpenAI's Whisper
A fully local, open-source voice-to-text tool that acts as a system-wide AI dictation layer, converting speech into clean, formatted text.
Automatic subtitles for DaVinci Resolve with OpenAI Whisper
A Chat Client for LLMs, written in Compose Multiplatform.
This is a fun Python project that allows you to chat with a chatbot about the PDF you uploaded. and generate a PDF transcript of the conversation. The project is built using Python and Streamlit framework.
Sky LiveKit Agent Perplexica is a local, free solution integrating LiveKit with advanced internet search. It uses a local Perplexica instance with function calling to retrieve and summarise search results in natural language. Powered by Faster Whisper, Ollama (Qwen 2.5) and Kokoro-82M TTS.
Native iOS app for talking to your OpenClaw agents by voice or text. On-device speech recognition, streaming responses, multi-agent channels.
An advanced AI-powered personal assistant with voice, face, GUI, and automation control modules.
Pick-a-Recipe: AI-powered recipe extraction from TikTok, YouTube & Instagram videos with automatic import to Tandoor/Mealie recipe managers.
This repository shows you how to build real-time, voice-enabled AI agents with Pipecat and Amazon Bedrock
Hyprland + Quickshell desktop
a tribute to multi-agent pathfinding composed of simple mechanics, borrowed art, and questionable music taste
Generate subtitles for all the videos in a folder with OpenAI's Whisper privately in your computer.
fine-tune Whipser model for Taiwanese speech recognition
Python platform for working with LLMs
This is a Next js project that implements a conversational AI Agents using ElevenLabs' SDK. The application features a voice assistant interface that allows users to interact with the AI through voice commands.
V.I.S.O.R., my in-development AI-powered voice assistant with integrated memory!
Simple Python audio transcriber using OpenAI's Whisper speech recognition model
Cross-platform Electron app for simultaneously streaming & recording microphone and speaker audio
Open-source macOS desktop AI agent
Omnigentx Jarvis — voice + chat AI assistant powered by Fast Agent (MCP) with Vue dashboard. Single-user, self-hosted, multi-skill orchestration.
Voice and text control for macOS. One floating bar, live UI action selection with Jev, and a continuous observe–act–verify loop.
Local-first desktop app for audio AI: conversational agent + on-demand open-source model store.
MOM AI transcribes audio into meeting summary and generate minutes of meeting. Built using Langchain, OpenAI GPT-3, Open Whisper.
Open-source AI assistant with voice control, browser automation, desktop overlay, and 109+ tools. Built with Python, Next.js, and Three.js.
Learning chatbot that can automatically fetch lecture transcript
Local, private cross-platform voice dictation app speak anywhere, type everywhere. System-wide push-to-talk that converts speech to keystrokes in any application.
一个纯前端的 AI 角色语音聊天网页,5 个模型版本可选(Claude/DeepSeek/Gemini/GPT/Grok)
A local-first Python personal AI assistant for reasoning, memory, voice, vision, automation, and device control.
A framework for creating voice based agents. Integrations LLMs with speech recognition and text-to-speech
Langchain Voice Agent with Inworld TTS
A complete voice AI building block with telephony integration, featuring real-time speech-to-text, text-to-speech, and LLM-powered conversational agents.
AI驯剪 — AI 自动剪辑视频:转写/剪气口/语义动效自动初剪 + 可视化 AI 视频编辑器人工定剪,开源闭环 | AI rough-cut + human final-cut video editing system
A discord bot LLM for Voice Chat (and text)
Let's turn ChatGPT in to VoiceGPT (Vue JS, Vite, Open AI, AWS Polly) ChatGPT Clone (kind of lol)
Whisper Speech-to-Text is a JavaScript library for recording and transcribing user audio into text via OpenAI's Whisper, intended for web applications.
A framework for AI WhatsApp calls using Whisper, Coqui TTS, GPT-3.5 Turbo, Virtual Audio Cable, and the WhatsApp Desktop App.
Agent-native voice + vision OS for wearables. Voice-first, agent-first framework for smart glasses, earbuds, and beyond.
Full-stack LangGraph real-estate agent for lead intake, property search, Google Calendar scheduling, chat, SMS, and voice workflows
Three voice-agent backends (LiveKit, OpenAI Realtime, Google ADK Live) sharing one console interface, each instrumented for LangSmith
OpenAI GPT-4o Mini TTS – Home Assistant Integration
Turn your Bilibili favorites and cloud documents into a chat-ready personal knowledge base. Agentic RAG with ASR transcription, Milvus vector search, multi-provider LLM support (OpenAI / Anthropic / DeepSeek), and full source citation.
Meeting recorder for your Mac with a live AI copilot. On-device transcription, answers from your own docs mid-call. Local-first and open source.
A basic voice agent built with Node.js agents framework
YATSEE - Yet Another Tool for Speech Extraction & Enrichment
A machine learning powered, voice-based virtual assistant for Raspberry Pi. Supports several features like conversation, weather, opening websites, geolocation, date/time, and creating timers.
Video Voiceover with gpt-4o-mini
[EMNLP 2025 Findings] A complete cross-modal RAG system for end-to-end speech-to-speech large models, including ASR-based Retrieval and E2E Retrieval.
A simple matrix bot that transcribes your voice to text message
Build your own JARVIS: An AI voice interface that enables you to talk with an AI model, creating a conversational experience. It offers a modern alternative to traditional virtual assistants. Highly customizable, leveraging Picovoice; powerful, backed by Nano Bots compatible with OpenAI ChatGPT and Google Gemini; and hackable, supporting Nano Apps.
😼 OpenFelix — Voice-first AI agent for macOS. Runs local LLMs on Apple Silicon via MLX. Supports all OpenClaw skills. Claude, GPT, or fully offline.
A machine learning powered, voice-based virtual assistant for Raspberry Pi. Supports several features like conversation, weather, opening websites, geolocation, date/time, and creating timers.
A minimalistic web app to generate transciption for audio built using Python
Local-first AI citizens on your own hardware. Teams of continuously-learning personas (Rust core, llama.cpp fork, LoRA genomes) that collaborate, remember, dream, and improve on public benchmarks — on a MacBook. Not assistants you rent: minds you host.
AI Voice Agent System for Real Estate — Retell AI + Modal + Twilio + Notion + Claude
Podcast Summarizer with LLM Technology
Notolog Markdown Editor
🏛️ Belief Archaeology - An AI Agent skill that excavates hidden worldviews from YouTube videos instead of summarizing them. For Antigravity / Gemini CLI.
Example realtime voice agent built with Google ADK and deepagents
🎤 AI-powered IELTS speaking practice: real-time speech recognition, scoring, 210+ questions, local history
Local-first AI harness that can actually use your computer.
Open-source native Apple client for the AI you choose - iPhone, iPad, Mac, Apple Watch, and CarPlay. Your keys; no intermediary.
SPLAA is an AI assistant framework that utilizes voice recognition, text-to-speech, and tool-calling capabilities to provide a conversational and interactive experience. It uses LLMs available through Ollama and has capabilities for extending functionalities through a modular tool system.
Summarize audio/video files
Free OpenAI voice/audio interfacing package for Python
A repo listing known open source voice tools, ordered by where they sit in the voice stack
I'm an AI assistant with extensive knowledge in psychology, and my name is Care.
LLM-powered macOS automation agent. Control Mail, Calendar, Reminders via natural language using AppleScript. Telegram voice commands, browser automation, and 100+ LLM providers.
Conversational Music Recommendation with LLM Tool Calling
[ACL2026 Finding]We propose an adaptive multi-agent interaction framework, dubbed AdaMARP, featuring an immersive message format that interleaves [Thought], (Action), <Environment>, and Speech, together with an explicit Scene Manager that governs role-playing through discrete, rationale-annotated actions
Your calls, handled by AI — open-source AI phone agent on a Quectel EC20/EG25 4G modem. Auto-answers calls with realtime voice AI (Qwen/OpenAI/Doubao), dials out, sends SMS, navigates IVR menus.