Category
Audio agents
1,754 Audio AI agents indexed on MeshKore — the most complete public catalog, ranked by popularity and updated daily.
1,754 agents · ranked by popularity · refine in the directory →
Audio agents — page 8 of 18
:sunglasses: Simple ChatBot, which replies to anything you ask ! Now with voice recgnition ! Say "Hey" to wake the bot ! :D
SpeakingAI is a demo of privately deployable 'GPT-4o like AI + RAG', a fully functional web AI server with audio query/answer in streaming, using LLM and RAG for backend knowledge.
Patterns and insights about yourself, picked up your Wispr Flow chats
ACE Step 1.5 XL music generation with dynamics-preserving mastering. Part of the AEON Media Production family.
Suno AI Premium — text-to-music generator with vocal synthesis, genre control and commercial licensing for creators producing original songs.
The open-source alternative to Suno and ElevenLabs Music. Natural language music composition, run locally, own everything.
Building AI products solo while traveling the world · Core contributor @openimsdk · Voice AI & Agents · cubxxw.com
Fully local AI meeting assistant: speaker diarization + self-updating knowledge graph. Nothing leaves your Mac.
A universal git-native AI agent framework — written in Rust. Single binary. Zero dependencies. 383 models. Blazing fast.
A full-stack application that allows practitioners to record voice notes and also export them to Google Docs/Sheets.
Voice Assistant AI This project is a high-performance, action-oriented voice assistant that uses ElevenLabs Conversational AI to provide human-like vocal interaction. It enables users to perform digital tasks—like building websites or creating art—entirely through natural speech.
Have a conversation with chatGPT. speech-to-text and text-to-speech is here using chatGPT
Node.js app where you can ask questions to ChatGPT using voice prompts, see the ChatGPT-like word-by-word answer, and then listen to the responses with voice
An AI-enabled music curation software suite based on Python and PydanticAI, utilizing Spotify and Brave Search as well as YouTube.
基于多AI协同的智能对话模块,让AI像真人一样自然参与聊天
J.A.R.V.I.S is a smart, lightweight, and extensible voice assistant that runs directly on your phone. Built with React Native and integrated with OpenAI's powerful language model!
OpenClaw skill: genpark voice shop
Open-source AI assistent with voice interaction, automation and enginering tools
Make any web app voice-controllable in 37 languages. Drop-in widget + open-source agent server — visitors speak, the agent talks back and drives their browser in real time.
Webagent, Turn your website Into an AI site. Press Cmd/Ctrl+K to do anything — command palette, DOM-grounded agent, voice, inline AI, and Dwell long-press in one SDK. Ships with your product, not as a chatbot widget.
OpenAI Sora 2 — full premium build for Windows with pro features unlocked. Creative Suite Pro Pack
Open-source real-time meeting copilot — live transcription + AI answers in a floating overlay. Anthropic, OpenAI, Gemini, Ollama, or any OpenAI-compatible endpoint.
本地优先的中文语音 AI 管家——说一句话,它替你操作 Windows(GLM-5.2 Agent + 本地 Whisper + GUI 自动化)
Generate original AI music tracks via Claude Desktop — MCP server powered by ACE-Step 1.5, with vocal/instrumental stem splitting and reference-style matching. No music software or AI/ML knowledge needed.
Local AI voice assistant with wake word detection, screen vision, 31 tools, and Iron Man web UI. Fully offline — Whisper STT, Kokoro TTS, Ollama LLM.
Hackable voice AI assistant for Windows: wake it with your voice, bring your own LLM (Claude/Groq/Ollama/OpenAI/Gemini), and it actually does things on your PC. Full source, free for personal use. by KloomStudio.com.ar
Aoi VR Agent - AI voice assistant that lives in your SteamVR hand: holographic hand panel (OpenVR overlay) driven by a native C++ agent (LLM, streaming TTS, ASR, simultaneous interpretation, VR screenshots). Unity IL2CPP + C++20, MIT licensed.
Self-hosted Rust LLM gateway for the Rabbit R1. WebSocket streaming, sandboxed Python agents, OpenAI / Anthropic / OpenAI-compatible providers.
ATRI 是本地优先的 AI Agent 架构与音乐工作站,提供多模型 LLM Runtime、工具调用闭环、上下文压缩、子 Agent 并行调度、MCP/Skills 扩展。并融合 Web DAW、MIDI 工程编辑与 Rust 实时音频/VST 宿主,让 Agent 能直接读写工程、生成 MIDI、控制播放并协作创作。
🔮 quami.io | Kwami's web2 platform
Inquiry videos with ChatGPT
🤖🎙️ Explore Lex Fridman Podcast Transcripts with a smart chatbot!
Infrastructure IA distribuée Linux — cluster GPU on-premise, 900+ agents autonomes, RGPD natif
An application that converts PowerPoint slides into videos and adds voice or text explanations automatically.
Eva the AI assistant
Edison - The Voice Controlled AI Assistant and Clock Radio
CursedGPT leverages the Hugging Face Transformers library to interact with a pre-trained GPT-2 model. It employs TensorFlow for model management and AutoTokenizer for efficient tokenization. The script enables users to input prompts interactively, generating text responses from the GPT-2 model.
Emilia - Desktop Character.AI Client
Integrate ChatGPT, Dall-E, Whisper and other AI models in Replicate into Messenger and Telegram bot
This is a personal health care tools.
🔒 A bot to log and prevent hate speech across many servers on Discord
Discord bot for Google and Polly Text-to-Speech
A turnbase mmoprg chat based game with 3 wifus.
A sample Nuxt 3 application that listens to chatter in the background and transcribes it using the powerful OpenAI Whisper, an automatic speech recognition (ASR) system.
Faster-whisper一键启动整合包带GUI界面
A web app to explore the life, works and world of famous composers through the years
Enhanced version of sora-extend with Docker support and CLI arguments. Original by @mattshumer_
MediBot AI; Your not-so-human, never-on-vacation, slightly overqualified AI medical assistant. Built with Ollama + RAG for brainpower, Whisper to hear your complaints, and React to look good while doing it. Finds hospitals, books appointments, tracks your cholesterol like a judgmental auntie, and soon might even talk back in your language.
Perl Framework for AI - Langertha - the viking of AI
Set of abstraction libraries to easily build Text and Audio based bots
This repository will guide you to create automatically generate YouTube Transcription using Using OpenAI's Whisper
Bot de conversão de áudio no whatsapp para texto
A platform to enhance medical e-Shadowing.
Freeswitch Speech-To-Text module
Speech-enabled Retrieval-Augmented Generation solution
Open-source AI operating system for computer use agents
Fluid dialogue manager plugin for Godot 4.x built on RiveScript
Classical music exploration chatbot built using Embabel and Neo4j
Talk to Rawan voice-to-voice using speech recognition or text-to-speech, with elevenlabs technology and chatgpt on the web.
CLI educacional para transcrição com OpenAI Whisper
电视语音换台源码参考,动态对接夏杰语音,直接支持语音换台。
An OpenClaw skill that uses faster-whisper (a faster implementation of the Whisper transcription model) to transcribe audio, with additional features such as speaker diarization
An open-source voice agent built on the PamirAI Distiller device, combining speech recognition, and text-to-speech to create a conversational AI assistant with OpenClaw you can talk to.
Discord Bot by GDjkhp
A powerful Video RAG system that enables users to upload videos, automatically builds searchable indexes from transcripts, and answers questions using Google Gemini. Built with a focus on local processing and only free tools.
Self-hosted multi-user AI assistant with agents, memory, scheduler, ESP32 voice satellites, and a built-in dashboard. Works with Claude, OpenAI, Ollama, LM Studio, and more.
NVIDIA DGX Spark ressources
Home Assistant add-on that replaces Xiaozhi cloud for StackChan ESP32-S3 robot — switchable OpenAI Realtime / Google Gemini Live voice assistant with HA device control. No Xiaozhi account needed.
A collection of real-world use cases built with agentcall.dev — each folder is a self-contained, working example. By Pattern AI Labs.
29 controlled experiments on Nemotron 3 Ultra: what prompting can and cannot do to a frontier model. Voice 4/8 to 7/8, hidden bugs 1/5 to 5/5 — blind dual-judge graded, no fine-tuning.
Minimalistic dictation app for Mac. (OpenAI Whisper wrapper)
臺語教材 AI Agent:依 108 課綱自動生成台語講義、測驗、互動網站、教學影片與語音教材(意傳媠聲 TTS+教育部官方音檔+臺羅標音檢核)
A runtime-agnostic local agent workflow that turns video links or transcripts into neutral, readable Obsidian notes — captions or local Whisper, free and fully local.
An AI teacher that lives next to your cursor on Windows. Sees your screen, talks back, points at where to click. Drop a Markdown doc into the knowledge folder and Clicky becomes an expert on any software, even niche or company-internal stuff.
鳳問 Ask Audrey — cited AI answers from Audrey Tang's 30-year transcript archive (archive.tw). Cloudflare Workers + Hono; web + LINE bot. 以 AI 檢索唐鳳逐字稿、附出處作答
Персональный ИИ-ассистент нового поколения для Windows — голос, зрение, управление ПК, навыки, память.
Control your live Kimi CLI sessions from your phone — self-hosted PWA over tmux: live chat, swarm progress, push notifications, voice input. Zero dependencies.
Create an AI by chatting - it gets a URL, a memory you can see, a body across your devices, a voice, and a wallet.
Simulación interactiva del HAL 9000 de 2001: Una Odisea del Espacio, con una interfaz inspirada en la nave Discovery One. Permite conversar con la icónica IA, recreando su personalidad serena pero inquietante para una experiencia inmersiva.
Voice-first language practice with Flutter, Go, OpenAI Realtime, Firebase, and Cloud Run.
real-time ai voice assistant with animated orb ui powered by anthropic claude
Local-first Video-to-Markdown for Bilibili, YouTube, Douyin and audio. CLI, Web UI, Docker and AI Skill for searchable knowledge bases.
This project is made for a hackathon and is named **AgentOps**. It is a platform where business owners can integrate automation into their websites to automate repetitive tasks like proposal generation and invoice generation, ensuring they never miss a client.
A command deck for your own agent. Voice in, voice out, tools, memory and cron — on your own model, or as the face for a self-hosted Hermes Agent.
The complete AI chat toolkit for Flutter — streaming, tools, generative UI, voice, and a batteries-included UI kit. Provider-agnostic (OpenAI, Anthropic, Gemini), state-manager-agnostic. The Flutter answer to Vercel's AI SDK + AI Elements.
AI-powered voice-note device firmware (ESP32-S3): record, transcribe with Whisper, clean up notes with GPT, and auto-sync Markdown to GitHub/Obsidian. An AI upgrade of the Pala Note hardware.
Open-source JARVIS-class desktop AI. Zen mode orb, computer control, plugin marketplace, custom tools, local RAG, knowledge graph, voice personality, automations, vision, 3D Holodeck, blackout privacy — all local, no cloud.
Discord AI bot capable of chatting and moderating, trained on conversation transcripts of Elon Musk
AI-powered meeting recorder on XIAO ESP32-S3 — live transcription, summaries, and AI chat
💫 𝗣ᴜʀᴠɪ 𝗖ʜᴀᴛ (𝗩2) 💢 𝗟ᴀᴛᴇsᴛ ᴜᴘᴅᴀᴛᴇs : ᴄʟᴇᴀɴ ᴄᴏᴅᴇ, ʀᴇᴍᴏᴠᴇ ʙᴜɢs, sʜɪғᴛᴇᴅ ᴛᴏ ᴄʜᴀᴛ ᴀᴘɪ 🔥 ɪғ ʏᴏᴜ ʜᴀᴠᴇ ᴀɴʏ ǫᴜᴇʀʏ : @KingxAra 🧿
Open-source realtime voice AI with adaptive VAD, latency telemetry, interruptible TTS, an inspectable TypeScript WebSocket gateway, and experimental Ollama/local AI.
ThorOS - a voice-controlled, AI-first Linux distribution built on Debian 13. Say what you want; on-device AI plans it and permissioned agents carry it out. Built so people who cannot use a keyboard or mouse can work independently. 100% local, no cloud required.
Taste-first club-music production that automates everything but the judgment. You direct in plain language; the agent writes each track's source, drives the synths, renders, and inspects. You never operate a DAW or plug-in; you make every musical call. Open-source, keyboard-first.
Let AI agents watch videos — cloud-hosted MCP server with transcript + frame extraction for TikTok, YouTube, and 1000+ platforms
Give your AI an eye — local screen → vision-model captions → memory store your AI reads on wake-up.
KI Operating System — Local-first AI infrastructure. Multi-Agent Orchestration, Ghost Control, One Voice Engine.
AI Voice Agents: Exploring the Next Generation of Human-Machine Interaction! 🎙️🤖🎧
end-to-end voicebot that answers open domain questions.
AI Vtuber for Streaming on Youtube/Twitch
An Interactive Personalized LLM Creating Realtime Musical Agents With Free RTOS