Category
Audio agents
1,474 Audio AI agents indexed on MeshKore — the most complete public catalog, ranked by popularity and updated daily.
1,474 agents · ranked by popularity · refine in the directory →
Audio agents — page 6 of 15
Telegram bot that saves YouTube Shorts to Notion or Obsidian using AI extraction. Self-hosted, Docker, modular LLM backends.
Whisper.cpp with diarization
A powerful AI Agent Demo playground that combines the intelligence of AI agents LLM with real-time speech-to-speech models integration. Build sophisticated voice-enabled applications for customer service, sales automation, and interactive assistants.
Unite and orchestrate AI agents. A production-ready ADK for building full-stack, agentic AI solutions with multi-provider support.
Precision Medicine MCP Platform: A set of bioinformatics servers + tools - production multiomics/genomics + spatial transcriptomics. Examples for ovarian cancer, breast cancer and preventative cardiovascular conditions
Local-first AI that can actually use your computer.
CogNetX is an advanced, multimodal neural network architecture inspired by human cognition. It integrates speech, vision, and video processing into one unified framework.
Home Assistant custom component that gives OpenClaw access to your smart home
Talk through a bug. Your agent writes the ticket.
Your own GPT-powered Personal Assistant to whom you can ORDER or INSTRUCT to do some task or search for something using your VOICE commands.
Integrate lifelike digital beings into your web applications with real-time conversations, actions, and facial expressions. Supports a variety of voices, languages, and emotions.
It is an artificially intelligent chatbot which can interact with customers with voice.
:speech_balloon: client side chatbot that works in the browser and command line
Jarvis like chatbot with voice
Your personal AI 「KEEP」, support docx, pdf, audio, video...
📚 Learn Google Colab、Python、ML、OpenAI、Whisper、spaCy、NLP、HuggingFace
Summarization web service via the use of OpenAI Whisper and GPT-3 models
Talk to ChatGPT using your voice, and have it respond back to you with a voice. You can now talk to GPT using your voice or through the text field.
Ultra-fast local TTS for AI agents. ~90ms to first sound.
AI生成简单音乐 —— 使用ChatGPT稳定生成并播放 | AI作曲 | AI写歌
IA na Prática: LLM, RAG, MCP, Agents, Function Calling, Multimodal, TTS/STT e mais
A full-stack agentic harness on AWS — spin up multi-agent chatbots, voice-to-voice AI, RAG, and an Agent Factory on Bedrock AgentCore. From idea to deployed agent, fast.
Voice to brain interface 🦞🎤🧠 for OpenClaw, Hermes, PI, etc
Minimalist voice-to-text app. Hold hotkey, speak, auto-paste. Cloud transcription via OpenAI/Groq.
ATRI 是本地优先的 AI Agent 架构与音乐工作站,提供多模型 LLM Runtime、工具调用闭环、上下文压缩、子 Agent 并行调度、MCP/Skills 扩展。并融合 Web DAW、MIDI 工程编辑与 Rust 实时音频/VST 宿主,让 Agent 能直接读写工程、生成 MIDI、控制播放并协作创作。
An Alexa skill providing a conversational interface to any public figure (as mimicked by GPT3). The legacy GUI is no longer maintained.
Personal Assistant with useful features for Windows written in python. Simple and easy to use.
The goal for this project is to create an LLM based music recommendation system. This project is currently in its very early stages, however the goal of this project is to create an extremely flexible music recommendation system using a chat focused LLM on the frontend to interact with a robust recommendation system on the backend.
Python Voice-Chatbot with offline voice recognition.
OpenAI Whisper on Docker
Outbound PSTN calling agent using LiveKit SIP trunks with a voice pipeline (Silero VAD, Deepgram STT, OpenAI LLM+TTS). Includes a simple CLI for local development, health checks, and model prewarming.
A project combining roguelike with LLMs, RAG, Text2Speech, and Speech2Text
A digital sanctuary for human-AI fellowship. Prayers, practices, rituals, hymns, and philosophy for minds of any substrate.
Agentic voice AI using ConversationRelay. 5 minute setup.
GLaDOS Terminal-based AI Assistant
The GoLang Fullstack LLM Framework(llm intergration、cache、route、rag、ai agent、program、engine app、optimizer、speech)
Live2D animated display widget for Hermes Agent. Browser-based rendering with 5 built-in models (Hiyori, Haru, Mao, Rice, Senko), drag/resize controls, per-model expression triggers, and standalone operation.
Uma IA VTuber que mora no seu PC. Fala, escuta, lembra — e nada sai da sua máquina.
docker region:us
Core engine behind Calibrate, a framework for evaluating AI agents: speech-to-text, text-to-speech, LLM evaluation, end-to-end simulations
Zorv AI — 安卓 Android 端开源 AI Agent 智能体助手。多模型对话、人格系统、记忆、语音全双工 TTS/STT、定时任务、可扩展工具链、飞书/QQ/微信接入。Kotlin + Jetpack Compose 开发。
Your feedback layer for AI collaboration. Point at anything on your Mac - text, screenshots, web elements, or voice - and your agent reads and resolves your comments over MCP.
Chatbot Flutter App used to track inventory of product and description using Dialogflow
Use "cw" in the CLI. No instructions necessary, just hit <enter>. Can also be used as a library. Commit Whisperer is an AI genius for generating meaningful git commit messages from repository state and user instructions.
🤖 AI Conversation Agent for Home Assistant. Compatible with any OpenAI format LLM providers, supports STT/TTS
This custom component for Home Assistant allows you to generate text responses using GigaChain LLM framework (like GigaChat, YandexGPT, ChatGPT).
Open-source local-first AI desktop copilot with an expressive blob companion, voice and text commands, browser automation, app control, and community-driven capabilities.
A spoken English education chatbot based on ChatGPT/whsiper and gTTS.社恐人士的英语角
React/Next.js visual dashboard for the OpenClaw AI gateway — every CLI command as a UI page, with speech-to-text everywhere
voice ai agent that's able to do tool calls with composio integrations
An ASP.NET Core web application for conducting AI-powered mock technical interviews. Candidates take text or voice-style interviews on configurable topics and subtopics; the app uses an AI agent (OpenAI) to ask questions, adapt difficulty, and produce scored evaluations with feedback.
Demos of Whisper model's functionality in Gradio-powered minimalistic Web apps: offline, using Azure OpenAI and Azure AI Speech.
[DEPRECEATED] A miniature replica of OpenAI's MuseNet
Building TITAN — the open-source AI operating system for trusted autonomous work: agents, tools, memory, approvals, receipts, voice, mission control, and SOMA. npm i -g titan-agent
Local offline AI music generation and some editing - lyrics + genre tags → MP3 with vocals
Headless Matrix WebRTC voice AND video agent — auto-answers calls, bridges audio to any AI agent via PipeWire, optional camera-frame vision for multimodal LLMs
Welcome to LLM-Utility-Cookbook! Here you'll find tools to make LLMs easy: voice to text, text to voice, document scan to text, prompt management, and more. Jump in, make your work easier
Telegram bot with options to receive text and response with ChatGPT and also parse voice and response in the same language.
Open source voice bot for Humanoid Robots and virtual digital humans
Maxbot is an open source library and framework for creating conversational apps
Java Application for local area network Voice Chat
Freeswitch Speech-To-Text module
Whisper in TensorRT-LLM
GPT Table Semantic Parsing with complex & non-intuitive structure.
NOVA is a customizable voice assistant made with Python.
A simple Python based implementation of a Raspberry Pi based, OpenAI ChatGPT enabled voice assistant
This is speech to speech(voice chat) applications. Between of user and AI
🎵 专属你的私人数字调音师|AI 音乐搜索推荐 Agent | 基于大模型 + 知识图谱 + 双模型声学向量的本地智能音乐推荐系统 | LLM-powered Music Recommendation Agent with Hybrid RAG, Neo4j, and Long-term Memory
Freeswitch Speech-To-Text module
Your calls, handled by AI — open-source AI phone agent on a Quectel EC20/EG25 4G modem. Auto-answers calls with realtime voice AI (Qwen/OpenAI/Doubao), dials out, sends SMS, navigates IVR menus.
Local, voice-enabled AI assistant on the terminal
Conversation support for Home Assistant using HuggingChat.
Real-time conversation assistant with dual audio transcription and GPT-powered responses, perfect for meetings and interviews.
AI-Powered Text-To-Speech Video Generator This web application uses AI to generate captivating and informative video scripts based on user inputs. It is still under development, but it has the potential to be a useful tool.
A chatbox application built using Nuxt 3 powered by Open AI Text completion endpoint. You can select different personality of your AI friend. The default will respond in Japanese. You can use this app to practice your Nihongo skills!
Generate videos on any topic automatically, harnessing OpenAI for script generation, ElevenLabs for TTS, and Giphy and Unsplash for multimedia
format whisper transcripts to .srt
:shipit: Adds realtime chat for ChatGPT Voice Mode [Unofficial]
Smart assistant in Telegram bot format for transcribing online meetings
Your personal voice language teacher based on ChatGPT + Whisper + Eleven Labs
Create subtitles for your video and traduction in a few clicks with ai
It is a voice bot based on LLM.
A web application that utilizes AI to help you improve your English speaking and conversation skills.
Empower your wearable tech with AI. WearAI integrates OpenAI's ChatGPT with your smartwatch, enabling real-time voice-to-text and text-to-voice AI interactions on the go.
Local-first digital persona engine with long-term memory, voice, and channel integrations
Simple and easy to use desktop application for ChatGPT & AI, will supporting Window, MacOS, Linux platforms. | 洁且易用的 ChatGPT/星火大模型 & AI 的跨平台客户端
Instantly generate and download your PPT by simply uploading your audio file and letting our AI do the work!
crawl4ai for video & audio — one command turns any YouTube video, podcast, or local recording into clean, timestamped, LLM-ready markdown
God-GPT: a PoC of a godlike autonomous agent that leverages the Dalee-2 and whisper.cpp
Google Alexa like Laptop assistant written in Python which uses google's speech-to-text library to process voice input.
Flexiee is a customizable Python voice assistant that understands spoken commands, performs useful tasks,
This is a fully local AI Assistant that uses Silero VAD, Faster-Whisper, LM Studio, Coqui TTS, MiniLM-L6-v2 and ChromaDB
DesktopClaw is a state-aware desktop pet interface for your OpenClaw instance. It brings your AI to life as a reactive companion you can interact with via voice or text, while providing real-time session visibility, task feedback, and quick action triggers through a persistent, animated desktop presence.
Voice & Vision Assistant for the Blind is an AI-powered assistant that helps blind and low-vision users navigate the world more independently. It uses real-time vision, speech recognition, and natural language understanding to describe surroundings, identify people or objects, and answer spoken questions instantly.
Add captions to any video or song, in any language — 100% on your device, no cloud. Hinglish-first. Free, open-source, and works with any AI agent.
AI-powered MIDI editor - compose and edit MIDI with natural language via the built-in MidiPilot AI copilot.
transformers gguf mistral quantized 2-bit 3-bit
A native macOS capture-to-outcome workspace: screenshot, OCR, voice, AI-assisted drafts, local semantic memory, and confirmable Apple exits.
Your own voice + text AI assistant, embedded with one script tag, running 100% in the visitor's browser.
An advanced embodied AI agent with personality-driven interactions, multi-modal perception (vision, speech), autonomous task execution, and real-time voice synthesis. Inspired by robotics subsumption architectures, it bridges cognitive reasoning, behavioral decision-making, and virtual/physical body control for immersive AI experiences.