Category

Audio agents

1,754 Audio AI agents indexed on MeshKore — the most complete public catalog, ranked by popularity and updated daily.

1,754 agents · ranked by popularity · refine in the directory →

Audio agents — page 2 of 18

maxheadbox★ 354

Tiny truly local voice-activated LLM Agent that runs on a Raspberry Pi

say★ 353

say - command line tool for voice and video calling

macos-local-voice-agents★ 341

Pipecat voice AI agents running locally on macOS

jarvis★ 333

Jarvis is a voice-activated, conversational AI assistant powered by a local LLM (Qwen via Ollama). It listens for a wake word, processes spoken commands using a local language model with LangChain, and responds out loud via TTS. It supports tool-calling for dynamic functions like checking the current time.

Skill-Anything★ 333

Any source (PDF, video, web, audio, text) to interactive learning package with quizzes, flashcards and spaced repetition. One command, 12-section study guide.

hack-interview★ 328

AI-powered tool for real-time interview question transcription and response generation.

voiceai★ 328

Set of 📝 with 🔗 to help those building Voice AI agents 🎙️🤖

BYDMate★ 325

BYD DiLink app (3.0/5.0/5.1, UI7): split screen 1/3+2/3, navigation on instrument cluster, Yandex Navigator on HUD, blind-spot cameras on turn signal, Russian voice AI agent, real BMS consumption, trips/charges journal, automation, ABRP + webhook telemetry. Leopard 3 / Bao 3, Sea Lion 07, Song, Atto 3.

tiledesk★ 322

Install Tiledesk on your server using Helm for Kubernetes orchestration and Docker Compose for running multi-container Docker applications. Tiledesk provides an open-source solution comparable to Voiceflow, empowering you to create sophisticated LLM-enabled chatbots that seamlessly transition interactions to human agents when needed.

openclaw-assistant★ 320

OpenClaw voice assistant app for Android - Wake word activation & system assistant integration

twewy-discord-chatbot★ 310

Discord AI Chatbot using DialoGPT, trained on the game transcript of The World Ends With You

whisper-node★ 306

Node.js bindings for OpenAI's Whisper. (C++ CPU version by ggerganov)

RuntimeSpeechRecognizer★ 302

Cross-platform, real-time, offline speech recognition plugin for Unreal Engine. Based on Whisper OpenAI technology, whisper.cpp.

gpt-voice-conversation-chatbot★ 302

Allows you to have an engaging and safely emotive spoken / CLI conversation with the AI ChatGPT / GPT-4 while giving you the option to let it remember things discussed.

AI-Talks★ 295

AI Talks - ChatGPT Assistant via Streamlit

TranscriberBot★ 294

TranscriberBot for Telegram

firefox-voice★ 292

Firefox Voice is an experiment in a voice-controlled web user agent

ai-devices★ 291

AI Device Template Featuring Whisper, TTS, Groq, Llama3, OpenAI and more

humla★ 290

Open-source AI meeting notes for Mac. Records mic + system audio with no bot, transcribes on-device or via OpenAI / Deepgram / Groq, identifies speakers offline, and writes summaries that fuse your notes with the transcript. Ask your notes and get cited answers. Tauri 2 + Rust + Swift.

sapphire★ 288

She's the AI agent you come home to.

safestclaw★ 280

Safestclaw is the alternative to openclaw.. You can naturally chat with it via text and voice, and you can choose not to use a language model., By default it picks up on intent and semantics.. No prompt injection while you get over ninety percent of what openclaw does plus tts and voice to text

tetos★ 277

A unified interface for multiple Text-to-Speech (TTS) providers.

aixplora★ 272

AIxplora is a open-source tool which let's you query all kind of files not limited to any length or format.

ai_webui★ 271

AI-WEBUI: A universal web interface for AI creation, 一款好用的图像、音频、视频AI处理工具

Amadeus★ 268

Real-time multimodal desktop agent evolving toward a persistent AI OS interface (0.15 α).

DB-GPT-Web★ 266

DB-GPT WebUI,LLM to vision.

llama_ros★ 264

llama.cpp (GGUF LLMs) and llava.cpp (GGUF VLMs) for ROS 2

agent-starter-python★ 263

A complete voice AI starter for LiveKit Agents with Python.

Stage-Whisper★ 261

The main repo for Stage Whisper — a free, secure, and easy-to-use transcription app for journalists, powered by OpenAI's Whisper automatic speech recognition (ASR) machine learning models.

Undefined★ 259

QQ bot platform with cognitive memory architecture and multi-agent Skills, via OneBot V11.

openai-voice-agent-sdk-sample★ 256

Sample application to add voice capabilities to the Agents SDK

twelvet★ 255

(Spring Boot 3. X Microservices framework) 基于Spring Boot 3.X 的 Spring Cloud Alibaba / Spring Cloud Tencent + React的微服务框架。🔝 🔝 点个starrred 关注更新。Chat GPT(RAG、TTS、STT、LLM)

vue-tui★ 255

Vue 3 terminal UI toolkit for browser DOM and CLI stdout: components, ANSI rendering, markdown transcripts, log views, and agent consoles.

GPT-Automator★ 254

Your voice-controlled Mac assistant

KarmaBot★ 254

🤖 A Multipurpose Discord Bot with a Music System & Utility commands used by 200K+ users!

react-native-chatbot★ 253

:speech_balloon: Easy way to create conversation chats

gpt_server★ 253

gpt_server是一个用于生产级部署LLMs、Embedding、Reranker、ASR、TTS、文生图、图片编辑和文生视频的开源框架。

sepia-docs★ 253

Documentation and Wiki for SEPIA. Please post your questions and bug-reports here in the issues section! Thank you :-)

SpeakoFlow★ 249

Free, open-source offline voice dictation for Windows, macOS, and Linux. A Wispr Flow alternative with an AI assistant that can read your screen on request and answer questions.

jaicf-kotlin★ 248

Kotlin framework for conversational voice assistants and chatbots development

daily-bots-web-demo★ 247

Daily Bots Web Demo showcasing how to build real-time voice AI agents 

hermes-relay★ 247

Hermes-Relay — Your Hermes AI agent, in your pocket — chat, voice, and control.

M.I.L.E.S★ 236

M.I.L.E.S, a GPT-4-Turbo voice assistant, self-adapts its prompts and AI model, can play any Spotify song, adjusts system and Spotify volume, performs calculations, browses the web and internet, searches global weather, delivers date and time, autonomously chooses and retains long-term memories. Available for macOS and Windows.

AI★ 230

The definitive, open-source Swift framework for interfacing with generative AI.

neuralnoise★ 227

The AI Podcast Studio: generate podcasts scripts and their audio version with a team of AI workers in a Podcast Studio 🎙️📜

portable-hermes-agent★ 226

Hermes Agent made portable desktop for Windows — 100 tools, GUI, local models via LM Studio, TTS, Music, ComfyUI, workflows, tool maker. No install. No Docker. No admin rights.

streamcore-server★ 225

Open-source realtime voice agent server in Go with WebRTC (WHIP), barge-in, streaming STT/LLM/TTS pipelines, plugin system, multi-language SDKs, SIP telephony, ESP32 support & fully local mode.

Unitale★ 224

一个基于Indextts和Qwen3TTS的 AI 有声书制作工具。利用 LLM 自动拆解剧本与识别情绪,集成多角色 TTS 语音合成(可智能分析音色并使用Qwen3TTS语音设计模型从音色描述文本生成音色),支持音效(SFX)、背景音乐(BGM)混音及实时台词音频滤波器的自动插入和匹配,可直接在浏览器导出 wav 成品,本工具本体无需配置环境即可跨平台在浏览器使用。现已支持背景图片提示词生成功能,可一键导出带情节背景图片和故事音频的mp4视频。

magda-core★ 221

An open ecosystem for music production.

SpotifyTranscripts★ 220

🎙️ AI generated subtitles and segmented chapters for podcasts

eva★ 220

A New End-to-end Framework for Evaluating Voice Agents

iva-agent★ 219

AI assistant in Telegram that remembers everything and helps you run your life. Self-hosted in one command.

openai_tts★ 217

Custom TTS component for Home Assistant. Utilizes the OpenAI speech engine or any compatible endpoint to deliver high-quality speech. Optionally offers chime and audio normalization features.

wyoming_openai★ 216

OpenAI-Compatible Proxy Middleware for the Wyoming Protocol

nodejs-whisper★ 211

NodeJS Bindings for Whisper - the CPU version of OpenAI's Whisper, as initially crafted in C++ by ggerganov.

amazon-sumerian-hosts★ 210

Amazon Sumerian Hosts (Hosts) is an experimental open source project that aims to make it easy to create interactive animated 3D characters for Babylon.js, three.js, and other web 3D frameworks. It leverages AWS services including Amazon Polly (text-to-speech) and Amazon Lex (chatbot).

multimodal-mcp-client★ 210

[DEPRECATED] Superseded by systempromptio/systemprompt-template and systempromptio/systemprompt-core. Multi-modal MCP client for voice-powered agentic workflows.

NachoBot★ 210

基于Maibot核心修改而成的多功能笨蛋机器人

llm_intents★ 208

Exposes internet search tools for use by LLM-backed Assist in Home Assistant

blitztext-app★ 208

Experimental open-source macOS menubar app for speech-to-text workflows

aural-oss★ 207

Open-source AI interview platform for voice, chat & video

IRIS-AI★ 205

💻 Desktop AI assistant built for real productivity. Voice, automation, memory, vision, web search, and workflow tools in one experience.

gdansk-ai★ 199

Full stack voice chatbot

chatbot-watson-android★ 198

An Android ChatBot powered by Watson Services - Assistant, Speech-to-Text and Text-to-Speech on IBM Cloud.

leon-cli★ 197

⌨️ Command-line interface (CLI) for a better use of Leon, your open-source personal assistant. GNU/Linux, macOS and Windows supported.

uxie★ 196

pdf reader app with note taking, annotations, collaboration, ai features (chat, flashcards generation w. ai-feedbacks), tts and ocr.

BentoChain★ 194

A voice-enabled chatbot application built using of 🦜️🔗 LangChain, text-to-speech, and speech-to-text models from 🤗 Hugging Face, and 🍱 BentoML.

sample-strands-agent-with-agentcore★ 191

Reference architecture for agentic AI chatbots with Strands Agents and Amazon Bedrock AgentCore

AutoGLM-TERMUX★ 190

Quickly deploy Open-AutoGLM agent on Android phone using Termux. Support AI voice recognition and enable automated operation of your phone without Root or PC!

ryza-ai-revive★ 190

Ryza AI companion. Bring your own LLM/TTS. 莱莎 AI 陪伴,自填大模型与语音。

Fun-Audio-Chat-8B★ 188

transformers safetensors funaudiochat text-generation audio-language-model speech-to-speech

youtube-agent-skill★ 187

Eleven Claude skills that run a YouTube channel: scripts off 21 hook formulas with a scored hook, title and thumbnail linted as one pairing, an edit decision list from your transcript, a retention reader that finds where viewers actually leave, a Shorts cutter, and a virality engine. Free, MIT.

openai-whisper-realtime★ 186

A quick experiment to achieve almost realtime transcription using Whisper.

presenter★ 186

A Multi-Agent AI Tool that creates beautiful presentations with voice-overs 🎦🔥

realtime-ai★ 184

A real-time Agent framework for audio and video.

Goida-AI-Unlocker★ 180

🛡 Установщик разблокировщика зарубежных AI-сервисов (и не только) для России 🌍

BlahST★ 177

Input text from speech in any Linux window, the lean, fast and accurate way, using whisper.cpp OFFLINE. Speak with local LLMs via llama.cpp.

kkclaw★ 174

🦞 一个可爱的桌面龙虾AI助手 - Desktop lobster pet with OpenClaw AI, Edge TTS voice, and emotion animations

iva★ 172

AI assistant in Telegram that remembers everything and helps you run your life. Self-hosted in one command.

flutter_whisper.cpp★ 171

Flutter App That Can Transcribe Audio Offline/On Device with Whisper C++ Bindings via Rust

Wally★ 170

Cute voice assistant built on ESP32 to help users with reminders, productivity, and daily conversations.

jarvis_ai★ 170

Iron-Man-style voice assistant + holographic HUD for Hermes Agent. Local Whisper STT, ElevenLabs voice, agent-summoned media panels, runs on your own hardware.

ai-course-notes★ 169

303 份 AI/LLM 中文讲义,支持在线阅读、PDF 下载和 LaTeX 源码查看 | Stanford CS336/CS224R/CS25 | Berkeley LLM Agents | Agent 工程实践

Logue★ 168

A local-first workspace for documents, tasks and meetings on macOS — nested spaces, typed properties, relations and saved views, with transcription, Smart Minutes and an agent running entirely on-device via MLX on Apple Silicon

jacobo-workflows★ 166

7 production n8n workflows from Jacobo, a multi-agent AI system (WhatsApp + Voice). Open source by default.

kobold_assistant★ 164

Like ChatGPT's voice conversations with an AI, but entirely offline/private/trade-secret-friendly, using local AI models such as LLama 2 and Whisper

OpenToys★ 164

Make Local AI Toys, Robots, Devices that work with a MacBook and an Arduino ESP32

audio-to-text-transcription★ 160

This repository contains a Python script that allows users to download the audio from a YouTube video, transcribe it into text, detect the language and save the transcription in txt file automatically.

skills★ 160

AAHL's Agent Skills. 汇集了多种实用的智能体技能,涵盖Home Assistant智能家居控制、微软Edge TTS和智谱GLM-TTS文本转语音、DuckDuckGo搜索、DeepWiki文档检索、加密货币行情、天气预报、Lark/飞书、影视搜索、商品比价等功能

B★ 160

An autonomous, context-aware AI desktop companion. Built with Python, featuring real-time screen vision, custom ONNX voice synthesis, active window tracking, and a dynamic floating UI with reactive facial expressions.

self-hosted-ai-stack★ 158

Deploy a complete self-hosted AI stack with Docker Compose: Ollama, LiteLLM, AnythingLLM, Whisper, WhisperLive, Kokoro, Embeddings, Docling and MCP Gateway. Local-first, private by default, with lightweight stacks, optional HTTPS and NVIDIA CUDA acceleration. Multi-arch: amd64, arm64.

ChatMusician★ 158

transformers pytorch safetensors llama text-generation en

podgenai★ 157

OpenAI GPT based informational audiobook/podcast mp3 generator

VoiceAgentRAG★ 157
fulloch★ 156

Fulloch - The Fully Local Home Voice Assistant

hypercheap-voiceAI★ 154

The most cost-effective, highest performance AI voice agent possible today

ai-waifu★ 153

AI VTuber Waifu and voice assistant

simplechat★ 153

Secure AI conversations with documents, video, audio, and more. Personal workspaces for focused context, group spaces for shared insight. Classify docs, reuse prompts, and extend with modular features.

talkGPT4All★ 152

A voice chatbot based on GPT4All and talkGPT, running on your local pc!

podcast-llm★ 152

Automatically generate engaging AI podcasts from nothing but an episode title.

Browse other category pages