Category
Audio agents
1,474 Audio AI agents indexed on MeshKore — the most complete public catalog, ranked by popularity and updated daily.
1,474 agents · ranked by popularity · refine in the directory →
Audio agents — page 13 of 15
AI voice agent framework for dTelecom rooms
Voice pipeline for Glove agent framework
Open-source, self-hosted AI agent framework for low-cost hardware. Voice-first. Memory-persistent. Model-agnostic. No lock-in.
LobeHub - an open-source,comprehensive AI Agent framework that supports speech synthesis, multimodal, and extensible Function Call plugin system. Supports one-click free deployment of your private ChatGPT/LLM web application.
The realtime voice agent framework
AI Agent Framework with multi-provider support and real-time broadcasting
MCP server for BlueColumn — persistent semantic memory API for AI agents. Give any MCP-compatible agent the ability to remember, recall, and store observations across sessions. Now with Audio Intelligence: voice call continuity, sound-event indexing, and
Tailor CLI — AI intelligence platform for documents, legislation, education, and healthcare. Upload, review, sign, and collaborate with AI agents using PACT protocol.
A simple project to make ChatGpt speak
(Now winston_doc) A perfect AI powered RAG for document query and summary. Supports ~all LLM and ~all filetypes (url, pdf, epub, youtube (incl playlist), audio, anki, md, docx, pptx, oe any combination!)
A simple and short denoising algorithm to denoise maximum amount of noise from an audio
A simple tool based on gtts(google text to speech api)
Realtime AI Voice Assistant with STM memory
LiveKit Python Agents
Agent Framework plugin for services from OpenAI
A suite of AI-powered command-line tools for text correction, audio transcription, and voice assistance.
MCP server for agent commerce payments. Features create invoice, process payment, escrow funds. From MEOK AI Labs.
Give your AI agent a voice — an MCP server for text-to-speech
Voice AI agent simulator and evaluation harness
SIP and audio streaming transport for AI voice agents (pure Rust)
AI outbound voice agent framework
AgentCare framework for healthcare voice AI workflows, scheduling, and patient communications.
AI agent framework with email, file, and voice processing
Phone numbers, SMS, email, and voice for AI agents.
MCP server for Agentline — give your AI agent a phone number, email, SMS, and voice calls.
AgentPhone Python SDK — give your AI agents phone numbers, SMS, and voice calls
Edge-based voice assistant using Gemma LLM with STT and TTS capabilities
Multi-agent AI system for video transcription using OpenAI Whisper, ChromaDB, and AutoGen
A real-time interactive Omni Avatar Agent built on LiveKit.
A tool for downloading YouTube audio, generating subtitles, and sending them to an HTTP endpoint.
A package for audio DSP tools
Audio Generator Agent AI Agent Directory to Host All Audio Generator Agent related AI Agents Services, Community, Reviews and More.
A library to remove silence from audio files using pydub.
A project about Audio models and it's fragility
RAG over audio files with provider-agnostic pipeline
Framework for evaluating text and voice AI agents
A voice command tool that uses ChatGPT as NLU
Package to speak with OpenAI's GPT models
This is the simplest module for AI Conversations with TTS.
Chatterbox TTS ported to VLLM for efficienct and advanced inference tasks
A sophisticated AI voice assistant with 1940s British charm - voice interface for Claude CLI
CLI harness for GarageBand - macOS audio project editing via Python. Requires: macOS
Local CLI for Cohere Transcribe (open-source 2B Conformer ASR).
CrewAI integration for AgentPhone — telephony tools for AI agents
A toolkit for working and conversing with large language models. Featuring tokenized sentence queueing for TTS.
A powerful desktop client for Mistral LLMs
Efficient speech quality assessment learned from SSL-based speech quality assessment model
Recommend DLSite voice works with LLM
Speech recognition extension library
Speech recognition extension library
Bridge between LiveKit Agents and ElevenAgents Voice Orchestration
ElevenLabs Text-to-Speech components for Haystack.
Local reimbursement invoice and payment proof matching agent
FAF Agent — the Voice of FAF (MCP server)
A Python application to assist in waking up for Fajr prayer by providing 3 interactive verses/explanations from the Quran + ChatGPT explanations accompanied by a soothing Islamic prayer fade-in and fade-out audio file from YouTube.
Haystack node to convert audio files into Documents.
Haystack node to convert text entities (documents, answers, etc...) into audio files.
A modular voice agent with swappable STT/TTS/LLM backends
fat_llama is a Python package for upscaling audio files to FLAC or WAV formats using advanced audio processing techniques. It utilizes CUDA-accelerated calculations to enhance audio quality by upsampling and adding missing frequencies through FFT (Fast Fourier Transform), resulting in richer and more detailed audio.
fat_llama_fftw is a Python package for upscaling audio files to FLAC or WAV formats using advanced audio processing techniques. It utilizes cpu-accelerated calculations to enhance audio quality by upsampling and adding missing frequencies through FFT (Fast Fourier Transform), resulting in richer and more detailed audio.
Fireworks AI Livekit Voice Agent Plugin
FunASR: A Fundamental End-to-End Speech Recognition Toolkit
FunASR: A Fundamental End-to-End Speech Recognition Toolkit
Audio transcription and alignment library using Google Gemini API for TTS labeling
Instrumentation and telemetry for the Google Gemini Live API via OpenTelemetry
A real-time Gemini API streaming library with FastAPI integration
A Python ASGI Socket.IO server proxying to Gemini Live via WebSockets
A Python package for Gemini Text-to-Speech using the official API.
OpenAI provider adapters for genblaze (Sora, DALL-E, TTS)
Speech-based GPT using Python
Fast GPT-3 client for Windows and Unix that supports both text and speech in any language.
Text-to-speech CLI tool and Python library using OpenAI's TTS API
Gpt powered subtitle translation tool.
A simple script that uses the Groq API to transcribe audio files using the Whisper model
A Voice-Controlled Drone Agent MCP Server
Raspberry Pi MCP Voice Assistant with Qwen-Agent framework
OpenAI TTS + STT voice providers for Kestrel Sovereign
LangChain integration for AgentPhone — telephony tools for AI agents
LangChain integration for CAMB AI - multilingual audio and localization tools
Integration of Sarvam AI platform with LangChain
LangChain integrations for Stardog
A powerful framework for building realtime voice AI agents
Bridge external AI agent frameworks (LangChain, LangGraph, Strands, ADK, custom) to the LiveKit Agent runtime.
gRPC translator for LiveKit Agents Adapter — connect any agent service to LiveKit
LangChain translator for LiveKit Agents Adapter
WebSocket translator for LiveKit Agents Adapter — connect any agent service to LiveKit
Agent Framework plugin for services from Anthropic
Groq inference plugin for LiveKit Agents
LangChain/LangGraph plugin for LiveKit agents
Universal LiveKit plugin for LangGraph workflows with intelligent filtering
Llama Index plugin for LiveKit Agents
LiveKit Agents Plugin for services from Mistral AI
Agent Framework plugin for services from OpenAI
Agent Framework plugin for OpenAI Server-Sent Events TTS API
Agent Framework plugin for OpenAI Server-Sent Events TTS API
Agent Framework plugin for RAG
llama-index tools azure_speech integration
LlamaIndex x ElevenLabs integration
LlamaIndex x Gemini Live integration
LlamaIndex x OpenAI Realtime integration