Category
Audio agents
1,754 Audio AI agents indexed on MeshKore — the most complete public catalog, ranked by popularity and updated daily.
1,754 agents · ranked by popularity · refine in the directory →
Audio agents — page 15 of 18
AIMakeSong review: compare features, pricing, use cases, access model, and alternatives for this Music agent in 2026. Transform your ideas into music with just a few clicks.
SongMaker-AI Music Generator review: compare features, pricing, use cases, access model, and alternatives for this Music agent in 2026. An AI-powered platform for creating and customizing music tracks.
Docs Made Easy review, features, use cases and alternatives (5/5 from 6 reviews): Document generation tool for invoices, reports, and quotes in Salesforce
Atmoscapia review: compare features, pricing, use cases, access model, and alternatives for this Music agent in 2026. Ambient Music Generator Studio for Creators
Dialora.ai review: compare features, pricing, use cases, access model, and alternatives for this Voice AI Agents agent in 2026. Rated 5/5 from 2 verified reviews. Smart Conversations, Real Results 24/7 AI Voice Agents for Growing Businesses
Instant AI voices that laugh and show real emotion.
Create smart AI voice agents using real-time speech and natural text.
Create ultra-fast, natural AI voice agents with ease.
Create AI voices that sound real and understand emotions.
TalkBud is an AI voice companion for natural, real-time conversations.
Build human-like voice AI agents and scale your phone operations.
Automated AI voice calls for customer service and outreach.
Rime gives your AI human-like voices for better customer talks.
Build AI voice agents that talk like humans and automate tasks.
Build and deploy AI voice agents to handle calls, customer support, and more.
Transcribe and understand speech accurately with advanced AI models.
Turn speech into text, text into speech, and power AI voice agents.
Turn text into lifelike speech with AI voices. Easy and free!
AI voice agents to automate your call center and boost satisfaction.
Self-hosted AI agent platform — autonomous, multi-channel, multi-model.
Voice pipeline for Cloudflare Agents — STT, TTS, VAD, streaming, and SFU utilities
The OpenAI Agents SDK is a lightweight yet powerful framework for building multi-agent workflows. This package contains the logic for building realtime voice agents on the server or in the browser.
TypeScript client OpenAI's realtime voice API.
OpenAI implementation for @micdrop/server
Multi-provider voice library for Next.js: ElevenLabs, Fish Audio, OpenAI TTS, Deepgram STT, and ElevenLabs ConvAI — with React hooks and drop-in route handlers
Embeddable AI assistant SDK for FieldStone — booking, chat, and voice on any website.
React hooks for TanStack AI streaming chat, realtime voice, structured outputs, and media generation.
Official TypeScript SDK for the Amigo Platform API
Official TypeScript SDK for sipgate flow — the GDPR-compliant, EU-hosted real-time voice and telephony layer for AI phone agents. Connect phone calls to your own AI agents (speech-to-text, text-to-speech, instant barge-in).
JavaScript/TypeScript SDK for Lokutor Real-time Voice AI
Rev AI makes speech applications easy to build!
Official Telnyx AI node for n8n
Browser SDK for Retell voice calls: start a web call with an agent, or monitor an ongoing call.
React Native component for enhancing app with voice interaction using Alan.AI SDK
Soniox transcription provider for the Vercel AI SDK.
AI voice agent framework for dTelecom rooms
Voice pipeline for Glove agent framework
Open-source, self-hosted AI agent framework for low-cost hardware. Voice-first. Memory-persistent. Model-agnostic. No lock-in.
LobeHub - an open-source,comprehensive AI Agent framework that supports speech synthesis, multimodal, and extensible Function Call plugin system. Supports one-click free deployment of your private ChatGPT/LLM web application.
The realtime voice agent framework
AI Agent Framework with multi-provider support and real-time broadcasting
Unified billing ledger for voice AI — one key for every voice tool and model, budget-capped.
MCP server for BlueColumn — persistent semantic memory API for AI agents. Remember, recall, and store observations across sessions. Audio Intelligence (calls, sound events, music) plus Music Memory: store practice recordings and lessons with instrument/ke
Tailor CLI — AI intelligence platform for documents, legislation, education, and healthcare. Upload, review, sign, and collaborate with AI agents using PACT protocol.
A simple project to make ChatGpt speak
(Now winston_doc) A perfect AI powered RAG for document query and summary. Supports ~all LLM and ~all filetypes (url, pdf, epub, youtube (incl playlist), audio, anki, md, docx, pptx, oe any combination!)
A simple and short denoising algorithm to denoise maximum amount of noise from an audio
A simple tool based on gtts(google text to speech api)
Realtime AI Voice Assistant with STM memory
LiveKit Python Agents
Agent Framework plugin for services from OpenAI
A suite of AI-powered command-line tools for text correction, audio transcription, and voice assistance.
MCP server for agent commerce payments. Features create invoice, process payment, escrow funds. From MEOK AI Labs.
Give your AI agent a voice — an MCP server for text-to-speech
Voice AI agent simulator and evaluation harness
SIP and audio streaming transport for AI voice agents (pure Rust)
AI outbound voice agent framework
AgentCare framework for healthcare voice AI workflows, scheduling, and patient communications.
AI agent framework with email, file, and voice processing
Phone numbers, SMS, email, and voice for AI agents.
MCP server for Agentline — give your AI agent a phone number, email, SMS, and voice calls.
AgentPhone Python SDK — give your AI agents phone numbers, SMS, and voice calls
Edge-based voice assistant using Gemma LLM with STT and TTS capabilities
Multi-agent AI system for video transcription using OpenAI Whisper, ChromaDB, and AutoGen
A real-time interactive Omni Avatar Agent built on LiveKit.
A tool for downloading YouTube audio, generating subtitles, and sending them to an HTTP endpoint.
A package for audio DSP tools
Audio Generator Agent AI Agent Directory to Host All Audio Generator Agent related AI Agents Services, Community, Reviews and More.
A library to remove silence from audio files using pydub.
A project about Audio models and it's fragility
RAG over audio files with provider-agnostic pipeline
Framework for evaluating text and voice AI agents
A voice command tool that uses ChatGPT as NLU
Package to speak with OpenAI's GPT models
This is the simplest module for AI Conversations with TTS.
Chatterbox TTS ported to VLLM for efficienct and advanced inference tasks
A sophisticated AI voice assistant with 1940s British charm - voice interface for Claude CLI
CLI harness for GarageBand - macOS audio project editing via Python. Requires: macOS
Local CLI for Cohere Transcribe (open-source 2B Conformer ASR).
CrewAI integration for AgentPhone — telephony tools for AI agents
A toolkit for working and conversing with large language models. Featuring tokenized sentence queueing for TTS.
A powerful desktop client for Mistral LLMs
Efficient speech quality assessment learned from SSL-based speech quality assessment model
Recommend DLSite voice works with LLM
Speech recognition extension library
Speech recognition extension library
Bridge between LiveKit Agents and ElevenAgents Voice Orchestration
ElevenLabs Text-to-Speech components for Haystack.
Local reimbursement invoice and payment proof matching agent
FAF Agent — the Voice of FAF (MCP server)
A Python application to assist in waking up for Fajr prayer by providing 3 interactive verses/explanations from the Quran + ChatGPT explanations accompanied by a soothing Islamic prayer fade-in and fade-out audio file from YouTube.
Haystack node to convert audio files into Documents.
Haystack node to convert text entities (documents, answers, etc...) into audio files.
A modular voice agent with swappable STT/TTS/LLM backends
fat_llama is a Python package for upscaling audio files to FLAC or WAV formats using advanced audio processing techniques. It utilizes CUDA-accelerated calculations to enhance audio quality by upsampling and adding missing frequencies through FFT (Fast Fourier Transform), resulting in richer and more detailed audio.
fat_llama_fftw is a Python package for upscaling audio files to FLAC or WAV formats using advanced audio processing techniques. It utilizes cpu-accelerated calculations to enhance audio quality by upsampling and adding missing frequencies through FFT (Fast Fourier Transform), resulting in richer and more detailed audio.
Fireworks AI Livekit Voice Agent Plugin
FunASR: A Fundamental End-to-End Speech Recognition Toolkit
FunASR: A Fundamental End-to-End Speech Recognition Toolkit