Category
Audio agents
1,754 Audio AI agents indexed on MeshKore — the most complete public catalog, ranked by popularity and updated daily.
1,754 agents · ranked by popularity · refine in the directory →
Audio agents — page 2 of 18
Tiny truly local voice-activated LLM Agent that runs on a Raspberry Pi
say - command line tool for voice and video calling
Pipecat voice AI agents running locally on macOS
Jarvis is a voice-activated, conversational AI assistant powered by a local LLM (Qwen via Ollama). It listens for a wake word, processes spoken commands using a local language model with LangChain, and responds out loud via TTS. It supports tool-calling for dynamic functions like checking the current time.
Any source (PDF, video, web, audio, text) to interactive learning package with quizzes, flashcards and spaced repetition. One command, 12-section study guide.
AI-powered tool for real-time interview question transcription and response generation.
Set of 📝 with 🔗 to help those building Voice AI agents 🎙️🤖
BYD DiLink app (3.0/5.0/5.1, UI7): split screen 1/3+2/3, navigation on instrument cluster, Yandex Navigator on HUD, blind-spot cameras on turn signal, Russian voice AI agent, real BMS consumption, trips/charges journal, automation, ABRP + webhook telemetry. Leopard 3 / Bao 3, Sea Lion 07, Song, Atto 3.
Install Tiledesk on your server using Helm for Kubernetes orchestration and Docker Compose for running multi-container Docker applications. Tiledesk provides an open-source solution comparable to Voiceflow, empowering you to create sophisticated LLM-enabled chatbots that seamlessly transition interactions to human agents when needed.
OpenClaw voice assistant app for Android - Wake word activation & system assistant integration
Discord AI Chatbot using DialoGPT, trained on the game transcript of The World Ends With You
Node.js bindings for OpenAI's Whisper. (C++ CPU version by ggerganov)
Cross-platform, real-time, offline speech recognition plugin for Unreal Engine. Based on Whisper OpenAI technology, whisper.cpp.
Allows you to have an engaging and safely emotive spoken / CLI conversation with the AI ChatGPT / GPT-4 while giving you the option to let it remember things discussed.
AI Talks - ChatGPT Assistant via Streamlit
TranscriberBot for Telegram
Firefox Voice is an experiment in a voice-controlled web user agent
AI Device Template Featuring Whisper, TTS, Groq, Llama3, OpenAI and more
Open-source AI meeting notes for Mac. Records mic + system audio with no bot, transcribes on-device or via OpenAI / Deepgram / Groq, identifies speakers offline, and writes summaries that fuse your notes with the transcript. Ask your notes and get cited answers. Tauri 2 + Rust + Swift.
She's the AI agent you come home to.
Safestclaw is the alternative to openclaw.. You can naturally chat with it via text and voice, and you can choose not to use a language model., By default it picks up on intent and semantics.. No prompt injection while you get over ninety percent of what openclaw does plus tts and voice to text
A unified interface for multiple Text-to-Speech (TTS) providers.
AIxplora is a open-source tool which let's you query all kind of files not limited to any length or format.
AI-WEBUI: A universal web interface for AI creation, 一款好用的图像、音频、视频AI处理工具
Real-time multimodal desktop agent evolving toward a persistent AI OS interface (0.15 α).
DB-GPT WebUI,LLM to vision.
llama.cpp (GGUF LLMs) and llava.cpp (GGUF VLMs) for ROS 2
A complete voice AI starter for LiveKit Agents with Python.
The main repo for Stage Whisper — a free, secure, and easy-to-use transcription app for journalists, powered by OpenAI's Whisper automatic speech recognition (ASR) machine learning models.
QQ bot platform with cognitive memory architecture and multi-agent Skills, via OneBot V11.
Sample application to add voice capabilities to the Agents SDK
(Spring Boot 3. X Microservices framework) 基于Spring Boot 3.X 的 Spring Cloud Alibaba / Spring Cloud Tencent + React的微服务框架。🔝 🔝 点个starrred 关注更新。Chat GPT(RAG、TTS、STT、LLM)
Vue 3 terminal UI toolkit for browser DOM and CLI stdout: components, ANSI rendering, markdown transcripts, log views, and agent consoles.
Your voice-controlled Mac assistant
🤖 A Multipurpose Discord Bot with a Music System & Utility commands used by 200K+ users!
:speech_balloon: Easy way to create conversation chats
gpt_server是一个用于生产级部署LLMs、Embedding、Reranker、ASR、TTS、文生图、图片编辑和文生视频的开源框架。
Documentation and Wiki for SEPIA. Please post your questions and bug-reports here in the issues section! Thank you :-)
Free, open-source offline voice dictation for Windows, macOS, and Linux. A Wispr Flow alternative with an AI assistant that can read your screen on request and answer questions.
Kotlin framework for conversational voice assistants and chatbots development
Daily Bots Web Demo showcasing how to build real-time voice AI agents
Hermes-Relay — Your Hermes AI agent, in your pocket — chat, voice, and control.
M.I.L.E.S, a GPT-4-Turbo voice assistant, self-adapts its prompts and AI model, can play any Spotify song, adjusts system and Spotify volume, performs calculations, browses the web and internet, searches global weather, delivers date and time, autonomously chooses and retains long-term memories. Available for macOS and Windows.
The definitive, open-source Swift framework for interfacing with generative AI.
The AI Podcast Studio: generate podcasts scripts and their audio version with a team of AI workers in a Podcast Studio 🎙️📜
Hermes Agent made portable desktop for Windows — 100 tools, GUI, local models via LM Studio, TTS, Music, ComfyUI, workflows, tool maker. No install. No Docker. No admin rights.
Open-source realtime voice agent server in Go with WebRTC (WHIP), barge-in, streaming STT/LLM/TTS pipelines, plugin system, multi-language SDKs, SIP telephony, ESP32 support & fully local mode.
一个基于Indextts和Qwen3TTS的 AI 有声书制作工具。利用 LLM 自动拆解剧本与识别情绪,集成多角色 TTS 语音合成(可智能分析音色并使用Qwen3TTS语音设计模型从音色描述文本生成音色),支持音效(SFX)、背景音乐(BGM)混音及实时台词音频滤波器的自动插入和匹配,可直接在浏览器导出 wav 成品,本工具本体无需配置环境即可跨平台在浏览器使用。现已支持背景图片提示词生成功能,可一键导出带情节背景图片和故事音频的mp4视频。
An open ecosystem for music production.
🎙️ AI generated subtitles and segmented chapters for podcasts
A New End-to-end Framework for Evaluating Voice Agents
AI assistant in Telegram that remembers everything and helps you run your life. Self-hosted in one command.
Custom TTS component for Home Assistant. Utilizes the OpenAI speech engine or any compatible endpoint to deliver high-quality speech. Optionally offers chime and audio normalization features.
OpenAI-Compatible Proxy Middleware for the Wyoming Protocol
NodeJS Bindings for Whisper - the CPU version of OpenAI's Whisper, as initially crafted in C++ by ggerganov.
Amazon Sumerian Hosts (Hosts) is an experimental open source project that aims to make it easy to create interactive animated 3D characters for Babylon.js, three.js, and other web 3D frameworks. It leverages AWS services including Amazon Polly (text-to-speech) and Amazon Lex (chatbot).
[DEPRECATED] Superseded by systempromptio/systemprompt-template and systempromptio/systemprompt-core. Multi-modal MCP client for voice-powered agentic workflows.
基于Maibot核心修改而成的多功能笨蛋机器人
Exposes internet search tools for use by LLM-backed Assist in Home Assistant
Experimental open-source macOS menubar app for speech-to-text workflows
Open-source AI interview platform for voice, chat & video
💻 Desktop AI assistant built for real productivity. Voice, automation, memory, vision, web search, and workflow tools in one experience.
Full stack voice chatbot
An Android ChatBot powered by Watson Services - Assistant, Speech-to-Text and Text-to-Speech on IBM Cloud.
⌨️ Command-line interface (CLI) for a better use of Leon, your open-source personal assistant. GNU/Linux, macOS and Windows supported.
pdf reader app with note taking, annotations, collaboration, ai features (chat, flashcards generation w. ai-feedbacks), tts and ocr.
A voice-enabled chatbot application built using of 🦜️🔗 LangChain, text-to-speech, and speech-to-text models from 🤗 Hugging Face, and 🍱 BentoML.
Reference architecture for agentic AI chatbots with Strands Agents and Amazon Bedrock AgentCore
Quickly deploy Open-AutoGLM agent on Android phone using Termux. Support AI voice recognition and enable automated operation of your phone without Root or PC!
Ryza AI companion. Bring your own LLM/TTS. 莱莎 AI 陪伴,自填大模型与语音。
transformers safetensors funaudiochat text-generation audio-language-model speech-to-speech
Eleven Claude skills that run a YouTube channel: scripts off 21 hook formulas with a scored hook, title and thumbnail linted as one pairing, an edit decision list from your transcript, a retention reader that finds where viewers actually leave, a Shorts cutter, and a virality engine. Free, MIT.
A quick experiment to achieve almost realtime transcription using Whisper.
A Multi-Agent AI Tool that creates beautiful presentations with voice-overs 🎦🔥
A real-time Agent framework for audio and video.
🛡 Установщик разблокировщика зарубежных AI-сервисов (и не только) для России 🌍
Input text from speech in any Linux window, the lean, fast and accurate way, using whisper.cpp OFFLINE. Speak with local LLMs via llama.cpp.
🦞 一个可爱的桌面龙虾AI助手 - Desktop lobster pet with OpenClaw AI, Edge TTS voice, and emotion animations
AI assistant in Telegram that remembers everything and helps you run your life. Self-hosted in one command.
Flutter App That Can Transcribe Audio Offline/On Device with Whisper C++ Bindings via Rust
Cute voice assistant built on ESP32 to help users with reminders, productivity, and daily conversations.
Iron-Man-style voice assistant + holographic HUD for Hermes Agent. Local Whisper STT, ElevenLabs voice, agent-summoned media panels, runs on your own hardware.
303 份 AI/LLM 中文讲义,支持在线阅读、PDF 下载和 LaTeX 源码查看 | Stanford CS336/CS224R/CS25 | Berkeley LLM Agents | Agent 工程实践
A local-first workspace for documents, tasks and meetings on macOS — nested spaces, typed properties, relations and saved views, with transcription, Smart Minutes and an agent running entirely on-device via MLX on Apple Silicon
7 production n8n workflows from Jacobo, a multi-agent AI system (WhatsApp + Voice). Open source by default.
Like ChatGPT's voice conversations with an AI, but entirely offline/private/trade-secret-friendly, using local AI models such as LLama 2 and Whisper
Make Local AI Toys, Robots, Devices that work with a MacBook and an Arduino ESP32
This repository contains a Python script that allows users to download the audio from a YouTube video, transcribe it into text, detect the language and save the transcription in txt file automatically.
AAHL's Agent Skills. 汇集了多种实用的智能体技能,涵盖Home Assistant智能家居控制、微软Edge TTS和智谱GLM-TTS文本转语音、DuckDuckGo搜索、DeepWiki文档检索、加密货币行情、天气预报、Lark/飞书、影视搜索、商品比价等功能
An autonomous, context-aware AI desktop companion. Built with Python, featuring real-time screen vision, custom ONNX voice synthesis, active window tracking, and a dynamic floating UI with reactive facial expressions.
Deploy a complete self-hosted AI stack with Docker Compose: Ollama, LiteLLM, AnythingLLM, Whisper, WhisperLive, Kokoro, Embeddings, Docling and MCP Gateway. Local-first, private by default, with lightweight stacks, optional HTTPS and NVIDIA CUDA acceleration. Multi-arch: amd64, arm64.
transformers pytorch safetensors llama text-generation en
OpenAI GPT based informational audiobook/podcast mp3 generator
Fulloch - The Fully Local Home Voice Assistant
The most cost-effective, highest performance AI voice agent possible today
AI VTuber Waifu and voice assistant
Secure AI conversations with documents, video, audio, and more. Personal workspaces for focused context, group spaces for shared insight. Classify docs, reuse prompts, and extend with modular features.
A voice chatbot based on GPT4All and talkGPT, running on your local pc!
Automatically generate engaging AI podcasts from nothing but an episode title.