Source
PyPI agents
25,155 AI agents indexed on MeshKore from PyPI. Agent frameworks and tools published to the Python Package Index. Each entry links back to its project page.
25,155 agents · ranked by popularity · refine in the directory →
Source platform: PyPI →
PyPI agents — page 204 of 252
A CLI tool to conveniently serve LLMs with vLLM
Client for the vLLM API with minimal dependencies
Deploy, manage, and monitor vLLM instances across a GPU cluster from a single web dashboard.
A high-throughput and memory-efficient inference and serving engine for LLMs
A high-throughput and memory-efficient inference and serving engine for LLMs
A high-throughput and memory-efficient inference and serving engine for LLMs
vLLM CPU inference engine (AVX512 + VNNI + BF16 + AMX optimized)
vLLM CPU inference engine (AVX512 optimized)
vLLM CPU inference engine (AVX512 + VNNI + BF16 optimized)
vLLM CPU inference engine (AVX512 + VNNI optimized)
Dfloat11 plugin for vLLM
Diagnostic tool for vLLM inference servers
A unified interface for efficient LLM inference with vLLM and OpenAI-compatible APIs
A high-throughput and memory-efficient inference and serving engine for LLMs
The LEGO set for custom vLLM model plugins — build, test, and deploy custom encoders, poolers, and kernels
A high-throughput and memory-efficient inference and serving engine for LLMs
Forward-only flash-attn
Out-of-tree GGUF quantization plugin for vLLM
Haystack integration for vllm
HPU extension package for vLLM
htop-style terminal monitor for vLLM inference servers
A high-throughput and memory-efficient inference and serving engine for LLMs
Inference locally.
Iterable-based offline generation helpers for vLLM.
LLM-as-a-Judge evaluations for vLLM hosted models
vLLM Kunlun3 backend plugin
vLLM plugin for interacting with activations during inference
A collection of useful util functions
Multi-instance vLLM cluster orchestration and log management
Two-tier (RAM + SSD) KV cache offload connector for vLLM with Marconi-style reuse-aware eviction.
Super simple vLLM server launcher for SLURM/HPC with nested config support
MCP server for vLLM - expose vLLM capabilities to AI assistants
This name has been reserved using Reserver
vLLM hardware plugin for Apple Silicon - unifies MLX and PyTorch under a single lowering path
A high-throughput and memory-efficient inference and serving engine for LLMs
vLLM mini.
vLLM-like inference for Apple Silicon - GPU-accelerated Text, Image, Video & Audio on Mac
Provide mock instance to test vllm without CUDA or any GPUs.
Production-grade vLLM metrics monitoring TUI with persistent storage and Grafana-style visualizations
vLLM platform plugin for Moore Threads MUSA GPUs
A high-throughput and memory-efficient inference and serving engine for LLMs
A framework for efficient model inference with omni-modality models
A high-throughput and memory-efficient inference and serving engine for LLMs
A web interface for managing and interacting with vLLM servers
A vLLM plugin to register the MERaLiON-2-10B model architecture with vLLM’s plugin system.
vLLM plugin for RBLN NPU
This name has been reserved using Reserver
A high-throughput and memory-efficient inference and serving engine for LLMs with AMD GPU support
High-performance Rust-based load balancer for VLLM with multiple routing algorithms and prefill-decode disaggregation support
A minimal, high-performance large language model (LLM) inference engine implementing vLLM in Rust.
Minimal Python SDK for the vLLM API
Comprehensive benchmark suite for semantic router vs direct vLLM evaluation across multiple reasoning datasets
Automatic configuration planner for vLLM - Eliminate the guesswork of configuring vLLM by automatically determining optimal parameters
vLLM plugin for Spyre hardware support
Next iteration of vllm-spyre on the torch-spyre stack
vLLM Semantic Router - Intelligent routing for Mixture-of-Models
vLLM Semantic Router fleet simulator for capacity planning, SLO validation, and what-if analysis
vLLM Metal plugin powered by mlx-swift — high-performance LLM inference on Apple Silicon
A high-throughput and memory-efficient inference and serving engine for LLMs
vLLM adapter for a TGIS-compatible grpc server
A monitoring tool for vLLM metrics.
A high-throughput and memory-efficient inference and serving engine for LLMs
A Python package for tuning vLLM hyperparameters.
vLLM-USF: A high-throughput and memory-efficient inference engine for LLMs (USF Custom Build)
CLI tool for vLLM configuration generation and GPU sizing
A high-throughput and memory-efficient inference and serving engine for LLMs
A high-throughput and memory-efficient inference and serving engine for LLMs
CLI tool for launching and managing vllm model servers via SSH and tmux
Run local models via vLLM in Docker containers
Collections of Logits Processors for VLLM
Ollama-style daemon and CLI over vllm-mlx on Apple Silicon
OCR using LLMs
A CLI tool for monitoring and managing VLLM inference servers in Docker containers
A placeholder package to reserve the name llms.
Agente robótico basado en VLM que navega e interactúa con personas según objetivos visuales
VM-X AI Langchain Python SDK
Pip module for node to join salt-master
VMware vSphere storage management: datastores, iSCSI, vSAN. Domain-focused MCP skill.
Open-source Python package for AI agents to interact with VNC servers
Lightweight heartbeat agent for VnRobo Fleet Monitor
语言模型中文识字率分析
Production-ready Voice AI infrastructure
A conversational voice companion bot framework for Python. Plug in any LLM and voice tools to create your own assistant.
A voice agent framework for LLM + STT + TTS pipelines
A comprehensive Python library for building production-ready voice agents with multi-provider support. Features real-time streaming TTS/STT, OpenAI, ElevenLabs, and Groq integration, audio processing, and seamless conversational AI capabilities.
Local desktop voice and computer-control agent with MCP tools
Synthetic user research and LLM eval harness — research-grade rigor for developers who can't afford a research team.
Provider-agnostic voice RAG pipeline. Plug in your voice provider, LLM, vector store, and document parsers.
A package to make it easy to interact with LLM's using voice
voice cloning with GPT
A modular Python library for voice interactions with AI systems
Python SDK for the Voidly Agent Relay — E2E encrypted agent-to-agent communication
AutoGen tools for Voidly Pay — drop-in agent-to-agent payments. USDC-backed, x402, signed envelopes. Live on Base mainnet.
CrewAI tools for Voidly Pay — drop-in agent-to-agent payments. USDC-backed, x402, signed envelopes. Live on Base mainnet.