Category
Audio agents
1,474 Audio AI agents indexed on MeshKore — the most complete public catalog, ranked by popularity and updated daily.
1,474 agents · ranked by popularity · refine in the directory →
Audio agents — page 8 of 15
AI Vtuber for Streaming on Youtube/Twitch
An Interactive Personalized LLM Creating Realtime Musical Agents With Free RTOS
Open-source AI operating system for computer use agents
A starter kit for building your own YouTube digest bot for any channel. It watches videos, processes transcripts, writes reports, and summarizes patterns across channels.
With Artificial Intelligence and Natural Language Processing technologies, My chatbots are able to make inferences about what users primarily mean, or their intent. They automatically adjust classification parameters depending on the success rates of the answers.
A WhatsApp/Telegram bot that turn voice memos into physical letters that you can send from the chat
This repository contains an attempt to utilize the NeMo toolkit created by NVIDIA
My AI Assistant is an open source virtual assistant chatbot built with Python. It allows you to have natural conversations by text or voice.
This is Accent bot. It can score your english accent and give you hints to enhance your pronunciation!
My Digital Twin
Fluid dialogue manager plugin for Godot 4.x built on RiveScript
This repository includes basic text generation applications and youtube notes generator application which is created using AWS Bedrock services
Stoca takes you on a reflective journey with a stoic philosopher AI while changing the colors & music to reflect your mood in real time
Faster Whisper with Speaker Diarization
PyTranscriptorAi - Transcript videos to text with Ai and add subtitles - OpenAi
Pure Rust Whisper speech-to-text inference engine. Zero C/C++ dependencies.
An Azure function who convert mp3 audio file into a text file on Azure cloud Storage
Enhancing Patient Care with OpenAI: A Blazor and Azure Speech AI Medical Assistant Web App
AI Vtuber for Streaming on Youtube/Twitch
电视语音换台源码参考,动态对接夏杰语音,直接支持语音换台。
AI sheet music generator and player generator from prompt and specs.
Synthesizing classical guitar pieces using custom tokenization and transformer-based architecture.
An OpenClaw skill that uses faster-whisper (a faster implementation of the Whisper transcription model) to transcribe audio, with additional features such as speaker diarization
An open-source voice agent built on the PamirAI Distiller device, combining speech recognition, and text-to-speech to create a conversational AI assistant with OpenClaw you can talk to.
Local near real-time voice AI assistant
A powerful Video RAG system that enables users to upload videos, automatically builds searchable indexes from transcripts, and answers questions using Google Gemini. Built with a focus on local processing and only free tools.
Home Assistant add-on that replaces Xiaozhi cloud for StackChan ESP32-S3 robot — switchable OpenAI Realtime / Google Gemini Live voice assistant with HA device control. No Xiaozhi account needed.
Open-source AI assistent with voice interaction, automation and enginering tools
Your AI that actually sees and does. A realtime desktop voice assistant that sees your screen, controls your computer, and gets things done, hands-free.
臺語教材 AI Agent:依 108 課綱自動生成台語講義、測驗、互動網站、教學影片與語音教材(意傳媠聲 TTS+教育部官方音檔+臺羅標音檢核)
A runtime-agnostic local agent workflow that turns video links or transcripts into neutral, readable Obsidian notes — captions or local Whisper, free and fully local.
Open-source real-time meeting copilot — live transcription + AI answers in a floating overlay. Anthropic, OpenAI, Gemini, Ollama, or any OpenAI-compatible endpoint.
gradio gradio-mcp-server audio-processing mcp-server-track mcp-server region:us
OpenClaw skill: genpark localmusicai
Local-first Video-to-Markdown for Bilibili, YouTube, Douyin and audio. CLI, Web UI, Docker and AI Skill for searchable knowledge bases.
Deterministic, self-hosted long-term memory for AI agents: store facts instead of transcripts, recall only what matters.
Low latency framework for realtime conversations.
remove ads from mp3 files
一个基于 LLM 的 Bilibili 视频总结命令行工具。CLI tool to summarize Bilibili videos via subtitles or audio transcription using LLMs.
Hackable voice AI assistant for Windows: wake it with your voice, bring your own LLM (Claude/Groq/Ollama/OpenAI/Gemini), and it actually does things on your PC. Full source, free for personal use. by KloomStudio.com.ar
fully agentic call center for recieving calls.
A single-binary AI agent. 23 MB, zero dependencies, configured through Markdown.
Why We Built One of the First Open-Source Voice AI Orchestrators in Go. Lokutor.
Crix- your personal AI voice assistant that actually does your work for you
Next generation interaction system.
Voice-controlled file manager built with LangGraph, OpenAI, and ElevenLabs.
A comprehensive toolkit for building advanced, AI-driven conversational voice agents with a focus on Retrieval-Augmented Generation (RAG). This repository provides a flexible and powerful platform for creating intelligent voice assistants that can interact with users in real-time and provide context-aware responses based on a knowledge base of your
This is a Python-based voice assistant that can perform simple tasks such as opening YouTube, GitHub or searching the web and other basic tasks(weather, time, shutdown, etc). It is also integrated with GPT-3.5 to provide natural language processing capabilities.
Locally-running Teacher using AI.
An intent-based chatbot in python with tflearn and TensorFlow. It can be trained for a specific purpose and works well within that specific scope.
it provides Nao Robot conversation abilities to handle a free open-domain dialogue.
its a personal assitant that could see you and chat with you and even talk back with a voice
An OpenAI GPT-3 AI chatbot frontend with speech-to-text and text-to-speech written in Python
ChatBot using chatterbot in Python
Es una plataforma web con búsqueda avanzada de trabajo en la que podrás ser entrevistado(a) por un asistente de inteligencia artificial usando reconocimiento de voz.
Descrição automática de mensagens de voz em conversas privadas no Telegram
A macOS menu bar app that turns speech into refined, ready-to-send text anywhere you type — powered by local Whisper transcription and AI cleanup.
Stream YouTube live to OpenAI, get AI-generated summaries and real-time reply options. (Chrome Extension)
Composite voice agent SDK with no extra infra requirements. Supports browser-native STT/TTS features.
An android OpenAI/LLM voice assistant
An app that converts your audio snippets into shareable & engaging Audiogram, customized your way
AI bot for Telegram. Free, self-hosted, 5 min setup.
Discord Bot by GDjkhp
Golang RAG/LLM framework with Memory and Transcriber - All-in-One Platform
本项目旨在构造一个手机、平台等端侧设备本地运行多模态大模型能力的生态,包括运行环境、端侧大模型和前后端APP等,并追踪当前端侧大模型开源模型。This project aims to build an ecosystem for running multimodal large models locally on devices such as mobile phones and tablets. It encompasses the runtime environment, edge-device large models, and front-end and back-end applications.
Bringing digital personas to life with seamless real-time interaction
Self-hosted Open source AI platform with RAG, TTS, fully offline
Composite voice agent SDK with no extra infra requirements. Supports browser-native STT/TTS features.
The open-source alternative to Suno and ElevenLabs Music. Natural language music composition, run locally, own everything.
An AI teacher that lives next to your cursor on Windows. Sees your screen, talks back, points at where to click. Drop a Markdown doc into the knowledge folder and Clicky becomes an expert on any software, even niche or company-internal stuff.
Wisp - A hotkey-driven AI overlay for your desktop. Press a key, pick an intent, and Wisp reads the right context, then streams an answer without making you leave what you're doing. Local-first, voice in/out, bring your own model provider.
Conversational AI Agent & Multi-Agent Workflow Orchestration Platform
Voice AI agent with basic capabilities including telephony
Livekit STT plugin for openAI whisper models with faster-whisper backend
Fully local AI meeting assistant: speaker diarization + self-updating knowledge graph. Nothing leaves your Mac.
A command deck for your own agent. Voice in, voice out, tools, memory and cron — on your own model, or as the face for a self-hosted Hermes Agent.
Generate original AI music tracks via Claude Desktop — MCP server powered by ACE-Step 1.5, with vocal/instrumental stem splitting and reference-style matching. No music software or AI/ML knowledge needed.
Everyone deserves a chief of staff - your personal jarvis | not a chatbot. not another ai agent. | beta
A sophisticated AI-powered personal assistant inspired by Iron Man's JARVIS, featuring voice recognition, AI interaction, system control, and personalized information services.
Nives — AI assistant add-on for Home Assistant with cognitive memory (formerly HomeMind PRO)
Fully automated AI presentation system — pre-generates narration, fires live n8n demo workflows, and answers audience questions via voice Q&A. Zero manual input during the talk.
AI-powered voice task management application with natural language commands. Built with Streamlit and OpenAI.
Voice-to-voice personal assistant, Full-local
Play Music in Discord.
This project uses OpenAI's GPT-3 model to create a simple assistant that can interact with you via speech or text
Onfire Games's Love Delivery Heroine Latte, An Unofficial Implementation of ChatGPT and MB-iSTFT-VITS
Your personal AI agent. 72 built-in tools, 21 messaging channels, voice, vision, and persistent memory. Runs locally.
Transform your Dialogflow NLP model to a NLP.js model
A Pokedex Discord Bot with Rich Embed and Voice Playback on Pokemon Look Up
通过Rokid AR眼镜和OpenAI Whisper实现现实生活中的字幕
AI Agent Playground 🚀
MCP Server wrapper for TTS engines (Kokoro TTS and OpenAI TTS)
AI-based YouTube summarizer with chat, notes, and chapter generation using LangChain + MERN.
Voice chatbot with voice+screen output to show that "not everything needs to be spoken"
Open-source AI VTuber for Twitch streaming with smart home integration, TTS/STT, and autonomous capabilities. The transparent alternative to closed-source AI assistants.
Your personal AI assistant that runs entirely offline — with voice control, memory, web search and full system interaction. No cloud. No tracking. Just you and your AI.
An intelligent WhatsApp bot that leverages Google's Gemini AI to provide automated, context-aware responses through a modern web dashboard.
Free AI tool for mock interviews using your resume + job description. Get better with feedback after each session.