Category
Audio agents
1,474 Audio AI agents indexed on MeshKore — the most complete public catalog, ranked by popularity and updated daily.
1,474 agents · ranked by popularity · refine in the directory →
Audio agents — page 15 of 15
The AssemblyAI JavaScript SDK provides an easy-to-use interface for interacting with the AssemblyAI API, which supports async and real-time transcription, as well as the latest LeMUR models.
Evaluation judges for AI voice agents (hallucination, response accuracy, intent, tool correctness, etc.)
FunASR speech-to-text transcriber for Haystack
sora2 video free review: compare features, pricing, use cases, access model, and alternatives for this AI Video Agents agent in 2026. A tool for creating realistic AI videos with synchronized audio instantly
MCP server for processing audio and video files using Google's Gemini multimodal models
Make A Song AI review: compare features, pricing, use cases, access model, and alternatives for this Music agent in 2026. AI music generator for songs, lyrics-to-song tracks, jingles, and creator audio.
Voice-Interactive Reasoning Agent — a voice-controlled AI agent that accesses local folder/markdown context during collaborative sessions.
Real-time AI-powered noise suppression solution, designed to work seamlessly across all browsers. For Video Conferencing, Live Sreaming and Recording solutions.
Official Dial LangChain tools — phone numbers, SMS, WhatsApp, and voice calls for AI agents
Build AI agents for phone and WhatsApp calls with real-time audio and messaging
Hexa - Voice Controlled AI Desktop Assistant
Production-grade CLI multi-agent AI assistant, real-time voice controller, and automated background job scheduler for local and cloud environments.
An embeddable AI-powered customer engagement and interactive product walkthrough SDK. With SalesGenie, you can dynamically guide users through your application using virtual pointers, page highlighting, voice/audio interfaces, and an AI chat assistant.
Multi-agent simulation harness that stress-tests conversational AI agents with adversarial synthetic users, then grades every transcript with auditable decision-tree judges.
moxxy command-line binary. Subcommand dispatcher consuming the moxxy SDK.
LangChain integration for FunASR (SenseVoice / Paraformer / Fun-ASR-Nano) speech-to-text
llama-index readers FunASR (SenseVoice / Paraformer / Fun-ASR-Nano) integration
Cloudflare-native LLM agent runtime for the Factory platform. One hardened orchestration engine — tool registry, reasoning loop, memory, guardrails — that powers the vertical SaaS products (Voice, Video, Astrology).
Local-first Goal Loop OS for long-term AI work across terminal, web, desktop, mobile app, and messengers.
RunAPI Gemini Omni SDK for audio, character, and video workflows in JavaScript, Python, Ruby, Go, Java, and PHP
Haystack integration for whisper
Soulmate IO review: compare features, pricing, use cases, access model, and alternatives for this AI Avatar agent in 2026. Private AI companions with voice chat and persistent memory.
Chime Labs review: compare features, pricing, use cases, access model, and alternatives for this Voice AI Agents agent in 2026. AI voice receptionist for Australian tradies — answers every call 24/7 and books the job.
Official Dial tools for the Microsoft Agent Framework — phone numbers, SMS, OTP, and voice calls
Official Dial AutoGen tools — phone numbers, SMS, OTP, and voice calls for Microsoft AutoGen agents
Official Dial CrewAI tools — phone numbers, SMS, OTP, and voice calls for multi-agent crews
VAT number validator for AI agents. EU VIES, AU ABR. Fraud risk scoring and name cross-check. PROCEED/HOLD verdict before any invoice payment.
Fine Voice review: compare features, pricing, use cases, access model, and alternatives for this Voice AI Agents agent in 2026. Free online AI voice generator and text-to-speech studio with voice cloning and API
MelodySeek review: compare features, pricing, use cases, access model, and alternatives for this Music agent in 2026. Identify music from video or audio links instantly
own your computer-use — your LLM on your own mouse and keyboard. Part of Own Your Stack.
Composable, agentic pipelines for MCP — fact-checking (DSPy), RAG (hybrid search), document intelligence (DeepPipe), and clinical voice (Groq/ElevenLabs). 31 AI tools, one command, zero manual setup.
AgentSea - Unite and orchestrate AI agents. A production-ready ADK for building agentic AI applications with multi-provider support.
Pipecat community TTS integration for the XTTSv2-vLLM streaming server
Stripe-backed Swarmauri billing provider for checkout, payments, subscriptions, invoices, refunds, disputes, Connect, and webhooks.
ARCOX DEX MCP server and terminal agent for retail swap, bridge, send, ARCOX Pay invoices, balances, history, retry bridge, and agentic jobs.
Neuralix AI SDK - Universal AI Agent Infrastructure for Modern Applications
Provider-agnostic realtime voice orchestration for GG tools and agents
Full-Stack Carnatic Music Notation Engine — Lyrics to Playable Audio
LumiMusic review: compare features, pricing, use cases, access model, and alternatives for this Music agent in 2026. AI music creation platform for generating and refining complete songs
UniMusic AI review: compare features, pricing, use cases, access model, and alternatives for this Music agent in 2026. AI music generator and song maker.
Paigy MCP server — a voice inbox for your AI agents. Lets an agent notify a user and await their reply.
Lobe Chat - an open-source, high-performance chatbot framework that supports speech synthesis, multimodal, and extensible Function Call plugin system. Supports one-click free deployment of your private ChatGPT/LLM web application.
The Supafone agent framework: create hosted voice/web/campaign agents with managed numbers, stages, tools, artifacts, and Supafone Pro watcher built in.
Official TypeScript SDK for the DocImprint document intelligence API
Voice synthesis endpoint for agent_core — per-agent Qwen3-TTS via ICL voice cloning
Agent-ready command-line interface for Lexware Office — contacts, invoices, and the receivables/AR-aging the API refuses to total for you.
Local, free voice artifacts for AI agents
Model adapters for the AgentDuet VoiceAgent layer: Gemini Live, xAI Grok Voice, Alibaba Qwen-Omni, Amazon Nova Sonic
Unofficial high-throughput batch inference for Cohere Arabic/English ASR
Unofficial high-throughput batch inference for Cohere Arabic/English ASR
Time-Accurate Automatic Speech Recognition using Cohere Transcribe, with word-level timestamps and speaker diarization.
Configurable LangChain agent core for voice applications with pluggable domains, prompts, tools, and RAG documents.
LangChain tools for FlowSpeech text-to-speech generation.
An integration package connecting Plivo and LangChain
LangChain integration for the Tunova music generation API (Suno-powered songs for AI agents)
Average purchase and sale prices computed from invoices
RunAPI Gemini TTS SDK for multi-speaker speech generation in JavaScript, Python, Ruby, Go, Java, and PHP
RunAPI OpenAI TTS SDK for speech generation in JavaScript, Python, Ruby, Go, Java, and PHP
Thin client SDK for the Multilingual Voice Agent service
SuperBryn SDK — sync your voice-agent configuration to SuperBryn for review, versioning, and monitoring
A friendly Python wrapper (sync + async) around the Vapi voice AI API.
BeatBun - dynamic ai music review: compare features, pricing, use cases, access model, and alternatives for this Music agent in 2026. Rated 5/5 from 1 verified review. AI music generator that creates royalty-free tracks from text
Lacuna review: compare features, pricing, use cases, access model, and alternatives for this Music agent in 2026. Chat-based AI Song Generator and AI Lyrics Generator that turns text into full songs with vocals
AIMusicFlow review: compare features, pricing, use cases, access model, and alternatives for this AI Agents Platform agent in 2026. AI music platform for creators, turning lyrics into songs with vocals, stems, mastering, and video.
Bingeable is an AI agent that builds product tutorial videos for you. It clicks through your app, records the process, and adds voiceovers without you ever need
React voice agent UI components with audio-reactive adapters for Vapi, ElevenLabs, LiveKit, Pipecat, OpenAI Realtime, Gemini Live, and custom voice AI apps.
BDB Invoice & Receipt Scraper Suite for headless automation, cross-platform Uber & AliExpress tax invoice fetcher, receipt PNG-to-PDF normalizer, and PDF accounting analyzer
Self-hosted AI voice agent for Asterisk. Streaming speech-to-text to LLM to text-to-speech over AudioSocket, with barge-in and correct 20 ms frame pacing.
Gandr text to speech component for Haystack
LangChain tools that return finished files: PDF, docx, charts, subtitles, invoices, screenshots — one hosted endpoint, free quota, no API key to start.
RunAPI OpenAI Transcription SDK for synchronous speech-to-text integration
Provider-neutral realtime multimodal agent for Swarmauri live models.
Stateless Swarmauri agent for converting independent text prompts into audio through any TTSBase provider.
Standalone PlayHT text-to-speech provider, stateless agent integration, and CLI for Swarmauri.