Category
Audio agents
1,754 Audio AI agents indexed on MeshKore — the most complete public catalog, ranked by popularity and updated daily.
1,754 agents · ranked by popularity · refine in the directory →
Audio agents — page 16 of 18
Audio transcription and alignment library using Google Gemini API for TTS labeling
Instrumentation and telemetry for the Google Gemini Live API via OpenTelemetry
A real-time Gemini API streaming library with FastAPI integration
A Python ASGI Socket.IO server proxying to Gemini Live via WebSockets
A Python package for Gemini Text-to-Speech using the official API.
OpenAI provider adapters for genblaze (Sora, DALL-E, TTS)
Speech-based GPT using Python
Fast GPT-3 client for Windows and Unix that supports both text and speech in any language.
Text-to-speech CLI tool and Python library using OpenAI's TTS API
Gpt powered subtitle translation tool.
A simple script that uses the Groq API to transcribe audio files using the Whisper model
A Voice-Controlled Drone Agent MCP Server
Raspberry Pi MCP Voice Assistant with Qwen-Agent framework
OpenAI TTS + STT voice providers for Kestrel Sovereign
LangChain integration for AgentPhone — telephony tools for AI agents
LangChain integration for CAMB AI - multilingual audio and localization tools
Integration of Sarvam AI platform with LangChain
LangChain integrations for Stardog
A powerful framework for building realtime voice AI agents
Bridge external AI agent frameworks (LangChain, LangGraph, Strands, ADK, custom) to the LiveKit Agent runtime.
gRPC translator for LiveKit Agents Adapter — connect any agent service to LiveKit
LangChain translator for LiveKit Agents Adapter
WebSocket translator for LiveKit Agents Adapter — connect any agent service to LiveKit
Agent Framework plugin for services from Anthropic
Groq inference plugin for LiveKit Agents
LangChain/LangGraph plugin for LiveKit agents
Universal LiveKit plugin for LangGraph workflows with intelligent filtering
Llama Index plugin for LiveKit Agents
LiveKit Agents Plugin for services from Mistral AI
Agent Framework plugin for services from OpenAI
Agent Framework plugin for OpenAI Server-Sent Events TTS API
Agent Framework plugin for OpenAI Server-Sent Events TTS API
Agent Framework plugin for RAG
llama-index tools azure_speech integration
LlamaIndex x ElevenLabs integration
LlamaIndex x Gemini Live integration
LlamaIndex x OpenAI Realtime integration
Train a speech-to-speech model using your own language model
LLM plugin providing access to 75+ AI models through 1min.ai API v2 (chat-with-ai endpoint) with SSE streaming, attachments, AI memory, brand voice, web search, and persistent options
Text-to-speech using the Azure OpenAI TTS API
Transcribe audio using the Groq.com Whisper API
OpenAI-compatible inference server: Llama 3.1 8B + Whisper + Kokoro TTS exposed via ngrok
Library to reduce latency in voice generations from LLM chat completion streams
Full-stack AI chat platform — multi-provider LLM, MCP tools, voice, budget
Llmovoice python library
LLM-powered FFmpeg command assistant
Tools for working with Ollama
Synap memory integration for LiveKit Agents — preload long-term memory, record turns, expose Synap as LLM tools
Natural language to FFmpeg, instantly and privately
A tool to create guided audio meditations (like those found on YouTube but in audio-only form) using OpenAI.
A proxy service converts Minimax TTS API to OpenAI-compatible format
SHAP implementation for multimodal large language models supporting audio and text input.
Moss runtime for LiveKit voice agents - a hot index cache shared across rooms and a one-line attach per call.
CLI tool for deploying Moss voice agents
Moss Voice Agent Manager - Simplified LiveKit agent integration
A library for optimizing audio processing parameters using differential evolution.
l10n_ar_electronic_invoice_storage_rg1361
Local macOS menubar voice assistant powered by MLX Whisper, Ollama, and Kokoro TTS
OpenClaw-inspired AI Agent Framework with NVIDIA Nemotron, Voice, Canvas, Browser Sandbox, and Multi-Channel Support
OpenClaw-inspired AI Agent Framework with NVIDIA Nemotron, Voice, Canvas, Browser Sandbox, and Multi-Channel Support
AgentPhone tools for OpenAI Agents SDK — telephony for AI agents
an openai based game audio translator
A powerful and easy-to-use Python library for generating natural-sounding speech using OpenAI text-to-speech capabilities.
An MCP server integrating with OpenAI TTS (text-to-speech) API
A CLI that provides tts using OpenAI
Free & Unlimited Unofficial OpenAI API Python SDK - Supports all OpenAI Models & Endpoints
A library for real-time text to speech processing using OpenAI API.
Simple CLI wrapper for OpenAI Speech-to-Text API
Speech-to-Text tool using Whisper, PyAudio, and VAD.
Audio-native LLM utilities for Pipecat
Plaud Live Agent SDK - 实时AI助手客户端SDK
OpenAIAudio converter
Python Module for Free ChatGPT TTS
Resp-Agent: A multi-agent framework for respiratory sound diagnosis and generation
Rio Agent: voice-first autonomous assistant with local and cloud runtimes
SignalWire AI Agents SDK
A voice interface for OpenAI's ChatGPT
Command line tool to transcribe & translate audio from livestreams in real time
Voice MCP for natural conversations with Claude - simple, elegant, wake-word activated
A client for Lollms that allows audio to audio full interaction with Lollms
A voice chatbot based on GPT4All and OpenAI Whisper, running on your PC locally
Telnyx Agent Toolkit — tools for building AI agents with Telnyx APIs
Placeholder for tts_webui_extension.gpt_sovits until migrated from GitHub due to VCS dependencies.
OpenAI compatible TTS API with support for multiple TTS models
Azure integrations for Twilio Agent Connect (TAC) - connectors for Azure AI agent services
A FastAPI server that acts as a TwiML voice agent
Steganography framework for ultrasonic agentic command transmission
Video SDK Agents
VideoSDK Agent Framework plugin for anthropic
VideoSDK Agent Framework plugin for Groq TTS services
VideoSDK Agent Framework plugin for LangChain and LangGraph
VideoSDK Agent Framework plugin for OpenAI services
Production-ready Voice AI infrastructure
A conversational voice companion bot framework for Python. Plug in any LLM and voice tools to create your own assistant.
A voice agent framework for LLM + STT + TTS pipelines
A comprehensive Python library for building production-ready voice agents with multi-provider support. Features real-time streaming TTS/STT, OpenAI, ElevenLabs, and Groq integration, audio processing, and seamless conversational AI capabilities.
Provider-agnostic voice RAG pipeline. Plug in your voice provider, LLM, vector store, and document parsers.
A package to make it easy to interact with LLM's using voice
voice cloning with GPT
A modular Python library for voice interactions with AI systems