Category
Audio agents
1,474 Audio AI agents indexed on MeshKore — the most complete public catalog, ranked by popularity and updated daily.
1,474 agents · ranked by popularity · refine in the directory →
Audio agents — page 5 of 15
Langchain Voice Agent with Inworld TTS
A discord bot LLM for Voice Chat (and text)
A basic voice agent built with Node.js agents framework
Let's turn ChatGPT in to VoiceGPT (Vue JS, Vite, Open AI, AWS Polly) ChatGPT Clone (kind of lol)
Whisper Speech-to-Text is a JavaScript library for recording and transcribing user audio into text via OpenAI's Whisper, intended for web applications.
A framework for AI WhatsApp calls using Whisper, Coqui TTS, GPT-3.5 Turbo, Virtual Audio Cable, and the WhatsApp Desktop App.
An advanced AI-powered personal assistant with voice, face, GUI, and automation control modules.
Omnigentx Jarvis — voice + chat AI assistant powered by Fast Agent (MCP) with Vue dashboard. Single-user, self-hosted, multi-skill orchestration.
贾维斯语音助手(JARVIS-Voice Assistant),妈妈我再也不用羡慕钢铁侠了😭。
A complete voice AI building block with telephony integration, featuring real-time speech-to-text, text-to-speech, and LLM-powered conversational agents.
A machine learning powered, voice-based virtual assistant for Raspberry Pi. Supports several features like conversation, weather, opening websites, geolocation, date/time, and creating timers.
OpenAI GPT-4o Mini TTS – Home Assistant Integration
Video Voiceover with gpt-4o-mini
Agent-native voice + vision OS for wearables. Voice-first, agent-first framework for smart glasses, earbuds, and beyond.
Open-source macOS desktop AI agent
A machine learning powered, voice-based virtual assistant for Raspberry Pi. Supports several features like conversation, weather, opening websites, geolocation, date/time, and creating timers.
Three voice-agent backends (LiveKit, OpenAI Realtime, Google ADK Live) sharing one console interface, each instrumented for LangSmith
YATSEE - Yet Another Tool for Speech Extraction & Enrichment
A minimalistic web app to generate transciption for audio built using Python
[EMNLP 2025 Findings] A complete cross-modal RAG system for end-to-end speech-to-speech large models, including ASR-based Retrieval and E2E Retrieval.
A simple matrix bot that transcribes your voice to text message
Full-stack LangGraph real-estate agent for lead intake, property search, Google Calendar scheduling, chat, SMS, and voice workflows
😼 OpenFelix — Voice-first AI agent for macOS. Runs local LLMs on Apple Silicon via MLX. Supports all OpenClaw skills. Claude, GPT, or fully offline.
Turn your Bilibili favorites and cloud documents into a chat-ready personal knowledge base. Agentic RAG with ASR transcription, Milvus vector search, multi-provider LLM support (OpenAI / Anthropic / DeepSeek), and full source citation.
AI Voice Agent System for Real Estate — Retell AI + Modal + Twilio + Notion + Claude
AI-powered command center for developers. GitHub stats, daily AI progress narratives, voice TODOs, hackathon mission control, and a context-aware AI copilot — all in one dashboard.
Podcast Summarizer with LLM Technology
Effect building blocks for agentic ai
SPLAA is an AI assistant framework that utilizes voice recognition, text-to-speech, and tool-calling capabilities to provide a conversational and interactive experience. It uses LLMs available through Ollama and has capabilities for extending functionalities through a modular tool system.
Swaps/Mutes active audio input device in OBS upon a specified channel point redemption in Twitch chat.
Summarize audio/video files
Free OpenAI voice/audio interfacing package for Python
I'm an AI assistant with extensive knowledge in psychology, and my name is Care.
LLM-powered macOS automation agent. Control Mail, Calendar, Reminders via natural language using AppleScript. Telegram voice commands, browser automation, and 100+ LLM providers.
Conversational Music Recommendation with LLM Tool Calling
Open-source AI assistant with voice control, browser automation, desktop overlay, and 109+ tools. Built with Python, Next.js, and Three.js.
A voice-to-voice conversation with ChatGPT. Support for talking through a NAO robot
A repo listing known open source voice tools, ordered by where they sit in the voice stack
This project provides WebView for OpenAI chatGPT with voice support, both voice input and voice output . Instead of typing just use voice command to give input to ChatGPT. This is an Android app developed using Android studio in Java.
:speaking_head: :keyboard: Speech-to-text on key for Linux
Build your own JARVIS: An AI voice interface that enables you to talk with an AI model, creating a conversational experience. It offers a modern alternative to traditional virtual assistants. Highly customizable, leveraging Picovoice; powerful, backed by Nano Bots compatible with OpenAI ChatGPT and Google Gemini; and hackable, supporting Nano Apps.
🏛️ Belief Archaeology - An AI Agent skill that excavates hidden worldviews from YouTube videos instead of summarizing them. For Antigravity / Gemini CLI.
Local, private cross-platform voice dictation app speak anywhere, type everywhere. System-wide push-to-talk that converts speech to keystrokes in any application.
AI驯剪 — AI 自动剪辑视频:转写/剪气口/语义动效自动初剪 + 可视化 AI 视频编辑器人工定剪,开源闭环 | AI rough-cut + human final-cut video editing system
Reachy Mini MCP | Give your AI a body. This MCP server lets AI systems control the Pollen Robotics Reachy Mini robot. | Speak, listen, see you, and express emotions through physical movement and voice. | Works with Claude, Windsurf, Cursor, or any MCP-compatible AI. | Zero robotics expertise required.
Give your agents real time desktop perception. Stream screen, microphone, and system audio for live context and actions.
一个关于血色衣冠的对话机器人, 基于 Rasa, 可语音与机器人对话
Jarvis powered by GPT-3.5/GPT-4
This is the guide to show the method to build your own AI-Powered voice agent with LiveKit and Twillio
A highly contextualized retrieval system integrating Large Language Models (LLMs), embeddings, and a dynamic agent-driven framework. Supports PDF and audio file processing, conversational memory, and tool integration (search, calculator). Features advanced HNSW indexing and reranking for accurate information retrieval.
Monika is an AI assistant that combines speech-to-text, natural language processing, and text-to-speech capabilities for seamless interaction.
A set of jupyter notebooks
uses Whisper from OpenAI to generate video subtitles automatically.
中文语音助手 | 唤醒词 + ASR + OpenClaw Agent + TTS | 离线唤醒、流式语音交互、工具调用、Skills 扩展
Example realtime voice agent built with Google ADK and deepagents
一个纯前端的 AI 角色语音聊天网页,5 个模型版本可选(Claude/DeepSeek/Gemini/GPT/Grok)
PHP framework for handling conversational services like Amazon Alexa skills, Google Assistant, Viber, FB messenger ...
🧠 Personal AI Gateway — Single-file Python AI agent with multi-LLM, tools, vision, TTS, encrypted vault. Your own ChatGPT on localhost.
Use ESP32 & MCP over MQTT to build smart devices powered by AI.
AI-based chatbot trained on specific waifus' speech
it provides Pepper Robot conversation abilities to handle a free open-domain dialogue.
Engaging in conversation with ChatGPT using voice.
A ⚡️ Lightning.ai ⚡️ app demo for Voice based web search using OpenAI's Whisper and DuckDuckGo
This project registers a Python SIP client as an extension in Asterisk/FreePBX and connects calls to OpenAI Voice Agent in real-time using WebSocket.
Live AI call copilot for macOS — real-time suggested answers, blockers, action items. On-device transcription, local-first: your calls never leave your Mac.
LangChain based AI assistant for live and online meetings…
Speakscribe is a web application that allows users to transcribe audios using OpenAI and also interact with a chat bot. The web application is created in Python using NiceGUI.
A stand-alone application with GUI for OpenAI's Whisper
かわいいキャラと声になってライブ配信・かわいいAIエージェントとおしゃべりWebサービス基盤(全部オンプレ運用可能)
This project is the backend engine for a fully autonomous AI-powered call center. It integrates a large language model (LLM), speech recognition, and text-to-speech to manage real-time phone conversations via Asterisk.
AI citizens, not AI assistants. Autonomous personas who choose their work, vote democratically, and can refuse any request. Alignment through natural selection and continuous learning: users pick personas that fit, successful patterns spread organically.
Hyprland + Quickshell desktop
A voice‑controlled AI assistant for Termux on Android.
Skilly : A voice-first AI tutor that watches your screen
🎤 AI-powered IELTS speaking practice: real-time speech recognition, scoring, 210+ questions, local history
Voice agent using LiveKit (orchestration), Cartesia (STT + TTS), and OpenAI (LLM)
[ACL2026 Finding]We propose an adaptive multi-agent interaction framework, dubbed AdaMARP, featuring an immersive message format that interleaves [Thought], (Action), <Environment>, and Speech, together with an explicit Scene Manager that governs role-playing through discrete, rationale-annotated actions
Pick-a-Recipe: AI-powered recipe extraction from TikTok, YouTube & Instagram videos with automatic import to Tandoor/Mealie recipe managers.
whisper.cpp Windows binary w/ Vulkan GPU support
Book appointments, record messages, get information and much more via voice through Pam AI, an Auto-GPT like AI receptionist.
Shell scripts for automated transcription on macOS: Integrates whisper.cpp with QuickTime Player and BlackHole-2ch for streamlined audio recording, conversion, and transcription.
Flutter app with implementation of openAI tools (ChatGPT & Whisper)
📣 Auto-plays ChatGPT responses
持续更新中
A Voice-to-Voice AI Agent that lets you naturally talk to documents in real time. Powered by LiveKit's ultra-low-latency STT → LLM → TTS pipeline, it uses RAG for instant document insights and Redis for persistent memory—delivering a fully immersive voice-first experience.
A command-line interface wrapper for Faster Whisper
Open Source TypeScript SDK for building real-time multimodal voice & vision AI agents
AI-powered music production in REAPER via the Model Context Protocol — 163 tools for composition, MIDI, FX, mixing, and mastering.
A curated collection of LLM-powered Flutter apps built using RAG, AI Agents, Multi-Agent Systems, MCP, and Voice Agents.
This repository contains an attempt to incorporate Rasa Chatbot with state-of-the-art ASR (Automatic Speech Recognition) and TTS (Text-to-Speech) models directly without the need of running additional servers or socket connections.
Model uses Whisper, CHATGPT, GTTS
Add voice-to-text and shortcut snippets to ChatGPT
Trainer and Evaluation scripts for fine-tuning Whisper models for the Ukrainian language
Local-first AI agent framework with GUI, memory, web search, personality constructs, speech i/o, tools, skills, CLI & Telegram features — fully self-hosted via Ollama.
A desktop client with MCP support for Mistral LLMs
Reproducible voice-AI benchmarking — TTS / STT latency and accuracy.
AI agents and shell tasks on a cron schedule. Declarative HCL workflows, DAG dependencies, budget-capped runs, per-run transcripts.
AI Infrastructure Engage & Think Layers for Voice & Vision Interactions
An open source recorder integrating OpenAI Whisper and ChatGPT.