Category
Audio agents
1,754 Audio AI agents indexed on MeshKore — the most complete public catalog, ranked by popularity and updated daily.
1,754 agents · ranked by popularity · refine in the directory →
Audio agents — page 18 of 18
Bingeable is an AI agent that builds product tutorial videos for you. It clicks through your app, records the process, and adds voiceovers without you ever need
React voice agent UI components with audio-reactive adapters for Vapi, ElevenLabs, LiveKit, Pipecat, OpenAI Live, OpenAI Realtime, Gemini Live, and custom voice AI apps.
BDB Invoice & Receipt Scraper Suite for headless automation, cross-platform Uber & AliExpress tax invoice fetcher, receipt PNG-to-PDF normalizer, and PDF accounting analyzer
Self-hosted AI voice agent for Asterisk. Streaming speech-to-text to LLM to text-to-speech over AudioSocket, with barge-in and correct 20 ms frame pacing.
Gandr text to speech component for Haystack
LangChain tools that return finished files: PDF, docx, charts, subtitles, invoices, screenshots — one hosted endpoint, free quota, no API key to start.
RunAPI OpenAI Transcription SDK for speech-to-text integration
Provider-neutral realtime multimodal agent for Swarmauri live models.
Stateless Swarmauri agent for converting independent text prompts into audio through any TTSBase provider.
Standalone PlayHT text-to-speech provider, stateless agent integration, and CLI for Swarmauri.
Speech Generation review: compare features, pricing, use cases, access model, and alternatives for this Voice AI Agents agent in 2026. Free online text-to-speech AI that converts text to natural human-like audio.
LiveKit voice integration for Mastra agents — realtime voice with semantic turn detection and barge-in
Speech-to-speech (realtime voice model) adapters for the Glove agent framework
Live avatar adapters for the Glove agent framework — a face over the speech-to-speech voice stack
CrewAI text to speech tool for the Gandr TTS API. First audio byte in 146 ms over the open internet, 116 ms p50 first audio, server side warm. 23 languages, every render watermarked.
LangChain text to speech tool for the Gandr TTS API. First audio byte in 146 ms over the open internet, 116 ms p50 first audio, server side warm. 23 languages, every render watermarked.
AgentDuet transport for Pipecat: phone/WhatsApp calls in any Pipecat pipeline
Bring the brain, we bring the voice: write a plain Python Brain, Voqalize runs the voice.
RapGen review: compare features, pricing, use cases, access model, and alternatives for this Music agent in 2026. Create rap lyrics, beats, vocals, and complete AI rap tracks in your browser.
AI Clean Audio review: compare features, pricing, use cases, access model, and alternatives for this Voice AI Agents agent in 2026. AI-powered background noise removal for audio and video files
SelfAgent — Cross-platform Autonomous AI & WhatsApp Gateway Agent. Install once, chat from your terminal or WhatsApp. Secure key storage via OS keychain (Windows DPAPI, macOS Keychain, Linux libsecret).
Multi-domain MCP Server for Voice Agent (Finance, Retail, Telecom)
Clean Voice App review: compare features, pricing, use cases, access model, and alternatives for this Voice AI Agents agent in 2026. AI tool that removes background noise from audio and video recordings.
AI Recruiting Agent by CogniAgent review: compare features, pricing, use cases, access model, and alternatives for this Recruiting agent in 2026. Screens, qualifies, and books the interview across text, chat, and voice.
Python SDK for the DhivehiGPT API
Velvet AI review: compare features, pricing, use cases, access model, and alternatives for this NSFW agent in 2026. Private AI companionship for adults. Emotionally intelligent AI partners with lifelike chat, voice a
MusicBento review: compare features, pricing, use cases, access model, and alternatives for this Music agent in 2026. Create Original Songs with AI Music Generator
CLI and Python API to transcribe videos longer than 30 minutes with gemini-3.5-transcribe while mitigating free-tier 429 errors
Connect a voice agent to the TeleCMI telephony platform
Speak pydantic-ai streamed output aloud, as it streams
Sori, an audio language model agent. This name is reserved; the package is released together with the model.
Sori, an audio language model agent. This name is reserved; the package is released together with the model.
Bring your own AI agent into Google Meet, Zoom & Discord voice channels, as a real voice participant.
Terminal AI Agent with desktop pet Arona — eye-tracking pupils, voice cloning, Computer Use, TTS/STT, MCP.
Open-source voice-and-video chat SDK for AI applications.
Node.js package of ai-coustics SDK
Ojin AI voice agent widget — drop-in web component for any website
Speechmatics Agent STT Python client for agent transcription
Online Music Visualizer review: compare features, pricing, use cases, access model, and alternatives for this Music agent in 2026. Create audio-reactive music visualizer videos online
Convert Audio to Text review: compare features, pricing, use cases, access model, and alternatives for this Productivity agent in 2026. Rated 5/5 from 5 verified reviews. Turn audio, video, and YouTube links into accurate transcripts with timestamps and AI summaries.
AI agent interface for WeChat Web — send/receive messages, manage contacts, upload media, transcribe voice
Fast, interruptible browser voice agent: VAD → STT → LLM → TTS with real barge-in, echo filtering, streaming first-sentence TTS and tap-to-seek. Zero runtime dependencies.
A voice agent that runs on your own machine — talk to her, she hands the slow work to bots with a shell, a browser and skills.
Node.js package of ai-coustics SDK
Voice agent framework powered by JEV (Joint Embedding Vectors) for dynamic conversation flow
Local real-time speech-to-speech translation pipeline (VAD -> streaming ASR -> translation -> TTS) for Intel AI PCs.
Ultralight CLI for local text-to-speech, powered by Kokoro by default
A powerful framework for building realtime voice AI agents
Agent Framework plugin for services from OpenAI
Dialog transformer for OVOS, designed to return responses from Ollama
Palabra Realtime STT/TTS models for the OpenAI Agents SDK VoicePipeline
Local-first terminal companion agent (Ollama-powered)
Reusable voice agent sessions with local speech and OpenRouter adapters
An open-source, provider-agnostic realtime voice agent framework in Python