Category
Audio agents
1,474 Audio AI agents indexed on MeshKore — the most complete public catalog, ranked by popularity and updated daily.
1,474 agents · ranked by popularity · refine in the directory →
Audio agents — page 2 of 15
Set of 📝 with 🔗 to help those building Voice AI agents 🎙️🤖
Discord AI Chatbot using DialoGPT, trained on the game transcript of The World Ends With You
Node.js bindings for OpenAI's Whisper. (C++ CPU version by ggerganov)
OpenClaw voice assistant app for Android - Wake word activation & system assistant integration
Cross-platform, real-time, offline speech recognition plugin for Unreal Engine. Based on Whisper OpenAI technology, whisper.cpp.
Allows you to have an engaging and safely emotive spoken / CLI conversation with the AI ChatGPT / GPT-4 while giving you the option to let it remember things discussed.
Voice-activated AI assistant with speech recognition and NLP. Automate tasks effortlessly with this Jarvis-like personal assistant.
AI Talks - ChatGPT Assistant via Streamlit
TranscriberBot for Telegram
AI Device Template Featuring Whisper, TTS, Groq, Llama3, OpenAI and more
Firefox Voice is an experiment in a voice-controlled web user agent
Safestclaw is the alternative to openclaw.. You can naturally chat with it via text and voice, and you can choose not to use a language model., By default it picks up on intent and semantics.. No prompt injection while you get over ninety percent of what openclaw does plus tts and voice to text
A unified interface for multiple Text-to-Speech (TTS) providers.
She's the AI agent you come home to.
AIxplora is a open-source tool which let's you query all kind of files not limited to any length or format.
AI-WEBUI: A universal web interface for AI creation, 一款好用的图像、音频、视频AI处理工具
DB-GPT WebUI,LLM to vision.
Sayna is a unified Voice Layer for AI Agents with a seemless integration to an existing agentic frameworks
The main repo for Stage Whisper — a free, secure, and easy-to-use transcription app for journalists, powered by OpenAI's Whisper automatic speech recognition (ASR) machine learning models.
llama.cpp (GGUF LLMs) and llava.cpp (GGUF VLMs) for ROS 2
🤖 A Multipurpose Discord Bot with a Music System & Utility commands used by 200K+ users!
(Spring Boot 3. X Microservices framework) 基于Spring Boot 3.X 的 Spring Cloud Alibaba / Spring Cloud Tencent + React的微服务框架。🔝 🔝 点个starrred 关注更新。Chat GPT(RAG、TTS、STT、LLM)
Sample application to add voice capabilities to the Agents SDK
Your voice-controlled Mac assistant
gpt_server是一个用于生产级部署LLMs、Embedding、Reranker、ASR、TTS、文生图、图片编辑和文生视频的开源框架。
:speech_balloon: Easy way to create conversation chats
Documentation and Wiki for SEPIA. Please post your questions and bug-reports here in the issues section! Thank you :-)
A complete voice AI starter for LiveKit Agents with Python.
QQ bot platform with cognitive memory architecture and multi-agent Skills, via OneBot V11.
Daily Bots Web Demo showcasing how to build real-time voice AI agents
Kotlin framework for conversational voice assistants and chatbots development
M.I.L.E.S, a GPT-4-Turbo voice assistant, self-adapts its prompts and AI model, can play any Spotify song, adjusts system and Spotify volume, performs calculations, browses the web and internet, searches global weather, delivers date and time, autonomously chooses and retains long-term memories. Available for macOS and Windows.
The definitive, open-source Swift framework for interfacing with generative AI.
The AI Podcast Studio: generate podcasts scripts and their audio version with a team of AI workers in a Podcast Studio 🎙️📜
🎙️ AI generated subtitles and segmented chapters for podcasts
Vue 3 terminal UI toolkit for browser DOM and CLI stdout: components, ANSI rendering, markdown transcripts, log views, and agent consoles.
NodeJS Bindings for Whisper - the CPU version of OpenAI's Whisper, as initially crafted in C++ by ggerganov.
[DEPRECATED] Superseded by systempromptio/systemprompt-template and systempromptio/systemprompt-core. Multi-modal MCP client for voice-powered agentic workflows.
Amazon Sumerian Hosts (Hosts) is an experimental open source project that aims to make it easy to create interactive animated 3D characters for Babylon.js, three.js, and other web 3D frameworks. It leverages AWS services including Amazon Polly (text-to-speech) and Amazon Lex (chatbot).
Hermes Agent made portable desktop for Windows — 100 tools, GUI, local models via LM Studio, TTS, Music, ComfyUI, workflows, tool maker. No install. No Docker. No admin rights.
Custom TTS component for Home Assistant. Utilizes the OpenAI speech engine or any compatible endpoint to deliver high-quality speech. Optionally offers chime and audio normalization features.
OpenAI-Compatible Proxy Middleware for the Wyoming Protocol
一个基于Indextts和Qwen3TTS的 AI 有声书制作工具。利用 LLM 自动拆解剧本与识别情绪,集成多角色 TTS 语音合成(可智能分析音色并使用Qwen3TTS语音设计模型从音色描述文本生成音色),支持音效(SFX)、背景音乐(BGM)混音及实时台词音频滤波器的自动插入和匹配,可直接在浏览器导出 wav 成品,本工具本体无需配置环境即可跨平台在浏览器使用。现已支持背景图片提示词生成功能,可一键导出带情节背景图片和故事音频的mp4视频。
Full stack voice chatbot
基于Maibot核心修改而成的多功能笨蛋机器人
Experimental open-source macOS menubar app for speech-to-text workflows
An Android ChatBot powered by Watson Services - Assistant, Speech-to-Text and Text-to-Speech on IBM Cloud.
⌨️ Command-line interface (CLI) for a better use of Leon, your open-source personal assistant. GNU/Linux, macOS and Windows supported.
Open-source AI meeting notes for Mac. Records mic + system audio with no bot, transcribes on-device or via OpenAI / Deepgram / Groq, identifies speakers offline, and writes summaries that fuse your notes with the transcript. Ask your notes and get cited answers. Tauri 2 + Rust + Swift.
A DAW built for automation, transformation, and fast musical iteration
A voice-enabled chatbot application built using of 🦜️🔗 LangChain, text-to-speech, and speech-to-text models from 🤗 Hugging Face, and 🍱 BentoML.
pdf reader app with note taking, annotations, collaboration, ai features (chat, flashcards generation w. ai-feedbacks), tts and ocr.
A New End-to-end Framework for Evaluating Voice Agents
Quickly deploy Open-AutoGLM agent on Android phone using Termux. Support AI voice recognition and enable automated operation of your phone without Root or PC!
Open-source AI interview platform for voice, chat & video
A quick experiment to achieve almost realtime transcription using Whisper.
Reference architecture for agentic AI chatbots with Strands Agents and Amazon Bedrock AgentCore
transformers safetensors funaudiochat text-generation audio-language-model speech-to-speech
A Multi-Agent AI Tool that creates beautiful presentations with voice-overs 🎦🔥
A real-time Agent framework for audio and video.
Exposes internet search tools for use by LLM-backed Assist in Home Assistant
Input text from speech in any Linux window, the lean, fast and accurate way, using whisper.cpp OFFLINE. Speak with local LLMs via llama.cpp.
Flutter App That Can Transcribe Audio Offline/On Device with Whisper C++ Bindings via Rust
🦞 一个可爱的桌面龙虾AI助手 - Desktop lobster pet with OpenClaw AI, Edge TTS voice, and emotion animations
Cute voice assistant built on ESP32 to help users with reminders, productivity, and daily conversations.
🛡 Установщик разблокировщика зарубежных AI-сервисов (и не только) для России 🌍
Like ChatGPT's voice conversations with an AI, but entirely offline/private/trade-secret-friendly, using local AI models such as LLama 2 and Whisper
💻 Desktop AI assistant built for real productivity. Voice, automation, memory, vision, web search, and workflow tools in one experience.
7 production n8n workflows from Jacobo, a multi-agent AI system (WhatsApp + Voice). Open source by default.
This repository contains a Python script that allows users to download the audio from a YouTube video, transcribe it into text, detect the language and save the transcription in txt file automatically.
transformers pytorch safetensors llama text-generation en
OpenAI GPT based informational audiobook/podcast mp3 generator
A voice chatbot based on GPT4All and talkGPT, running on your local pc!
Automatically generate engaging AI podcasts from nothing but an episode title.
gradio region:us
AAHL's Agent Skills. 汇集了多种实用的智能体技能,涵盖Home Assistant智能家居控制、微软Edge TTS和智谱GLM-TTS文本转语音、DuckDuckGo搜索、DeepWiki文档检索、加密货币行情、天气预报、Lark/飞书、影视搜索、商品比价等功能
Open-source realtime voice agent server in Go with WebRTC (WHIP), barge-in, streaming STT/LLM/TTS pipelines, plugin system, multi-language SDKs, SIP telephony, ESP32 support & fully local mode.
AI VTuber Waifu and voice assistant
Secure AI conversations with documents, video, audio, and more. Personal workspaces for focused context, group spaces for shared insight. Classify docs, reuse prompts, and extend with modular features.
The most cost-effective, highest performance AI voice agent possible today
Make Local AI Toys, Robots, Devices that work with a MacBook and an Arduino ESP32
AI assistant in Telegram that remembers everything and helps you run your life. Self-hosted in one command.
ai agent video editor use with ElevenLabs Scribe, pack phrase-level transcripts, reason over an EDL, render with ffmpeg, video editor ai agent
WhisperClip simplifies your life by automatically transcribing audio recordings and saving the text directly to your clipboard. With just a click of a button, you can effortlessly convert spoken words into written text, ready to be pasted wherever you need it. This application harnesses the power of OpenAI’s Whisper for free.
303 份 AI/LLM 中文讲义,支持在线阅读、PDF 下载和 LaTeX 源码查看 | Stanford CS336/CS224R/CS25 | Berkeley LLM Agents | Agent 工程实践
🎵 Democratic Slack/Discord bot for Sonos control with Spotify integration. Queue music, vote to skip, and let the community decide what plays!
Input a YouTube video link or upload a video file and get a video with subtitles.
It is a personal assistant chatbot, capable to perform many tasks same as Google Assistant plus more extra features...
Deploy a complete self-hosted AI stack with Docker Compose: Ollama, LiteLLM, AnythingLLM, Whisper, WhisperLive, Kokoro, Embeddings, Docling and MCP Gateway. Local-first, private by default, with lightweight stacks, optional HTTPS and NVIDIA CUDA acceleration. Multi-arch: amd64, arm64.
Recent advancements propelled by large language models (LLMs), encompassing an array of domains including Vision, Audio, Agent, Robotics, Fundamental Sciences such as Mathematics, and Ominous.
A voice-first desktop AI companion with memory, emotion, tools, and plugins.
ChatGPT web application. ChatGPT 网页应用,支持多对话、海量提示词、PWA、ASR、TTS
PodAgent: A Comprehensive Framework for Podcast Generation
Fulloch - The Fully Local Home Voice Assistant
Voice-to-text CLI for terminal users
Iron-Man-style voice assistant + holographic HUD for Hermes Agent. Local Whisper STT, ElevenLabs voice, agent-summoned media panels, runs on your own hardware.
Hermes-Relay — Your Hermes AI agent, in your pocket — chat, voice, and control.
React Native Duolingo clone with a real-time AI voice teacher. Built with Expo, Stream Voice Agents, Clerk auth, and NativeWind for a complete, interactive mobile learning experience.
⚡ A local, privacy-focused AI desktop assistant for Windows. Control your PC remotely via Telegram or locally with Voice commands. Powered by Ollama.