Source

PyPI agents

25,155 AI agents indexed on MeshKore from PyPI. Agent frameworks and tools published to the Python Package Index. Each entry links back to its project page.

25,155 agents · ranked by popularity · refine in the directory →

Source platform: PyPI

PyPI agents — page 204 of 252

vllm-cli

A CLI tool to conveniently serve LLMs with vLLM

vllm-client

Client for the vLLM API with minimal dependencies

vllm-cluster-manager

Deploy, manage, and monitor vLLM instances across a GPU cluster from a single web dashboard.

vllm-consul

A high-throughput and memory-efficient inference and serving engine for LLMs

vllm-cpm

A high-throughput and memory-efficient inference and serving engine for LLMs

vllm-cpu

A high-throughput and memory-efficient inference and serving engine for LLMs

vllm-cpu-amxbf16

vLLM CPU inference engine (AVX512 + VNNI + BF16 + AMX optimized)

vllm-cpu-avx512

vLLM CPU inference engine (AVX512 optimized)

vllm-cpu-avx512bf16

vLLM CPU inference engine (AVX512 + VNNI + BF16 optimized)

vllm-cpu-avx512vnni

vLLM CPU inference engine (AVX512 + VNNI optimized)

vllm-df11

Dfloat11 plugin for vLLM

vllm-doctor

Diagnostic tool for vLLM inference servers

vllm-dolphin
vllm-efficient-client

A unified interface for efficient LLM inference with vLLM and OpenAI-compatible APIs

vllm-emissary

A high-throughput and memory-efficient inference and serving engine for LLMs

vllm-factory

The LEGO set for custom vLLM model plugins — build, test, and deploy custom encoders, poolers, and kernels

vllm-fixed

A high-throughput and memory-efficient inference and serving engine for LLMs

vllm-flash-attn

Forward-only flash-attn

vllm-gguf-plugin

Out-of-tree GGUF quantization plugin for vLLM

vllm-haystack

Haystack integration for vllm

vllm-hpu-extension

HPU extension package for vLLM

vllm-htop

htop-style terminal monitor for vLLM inference servers

vllm-hust

A high-throughput and memory-efficient inference and serving engine for LLMs

vllm-inference

Inference locally.

vllm-iter

Iterable-based offline generation helpers for vLLM.

vllm-judge

LLM-as-a-Judge evaluations for vLLM hosted models

vllm-kunlun

vLLM Kunlun3 backend plugin

vllm-lens

vLLM plugin for interacting with activations during inference

vllm-logits

A collection of useful util functions

vllm-manager

Multi-instance vLLM cluster orchestration and log management

vllm-marconi-offload

Two-tier (RAM + SSD) KV cache offload connector for vLLM with Marconi-style reuse-aware eviction.

vllm-marenostrum

Super simple vLLM server launcher for SLURM/HPC with nested config support

vllm-mbart
vllm-mcp-server

MCP server for vLLM - expose vLLM capabilities to AI assistants

vllm-messages

This name has been reserved using Reserver

vllm-metal

vLLM hardware plugin for Apple Silicon - unifies MLX and PyTorch under a single lowering path

vllm-mindspore

A high-throughput and memory-efficient inference and serving engine for LLMs

vllm-mini

vLLM mini.

vllm-mlx

vLLM-like inference for Apple Silicon - GPU-accelerated Text, Image, Video & Audio on Mac

vllm-mock

Provide mock instance to test vllm without CUDA or any GPUs.

vllm-mon

Production-grade vLLM metrics monitoring TUI with persistent storage and Grafana-style visualizations

vllm-musa

vLLM platform plugin for Moore Threads MUSA GPUs

vllm-nccl-cu11
vllm-nccl-cu12
vllm-npu

A high-throughput and memory-efficient inference and serving engine for LLMs

vllm-omni

A framework for efficient model inference with omni-modality models

vllm-online

A high-throughput and memory-efficient inference and serving engine for LLMs

vllm-playground

A web interface for managing and interacting with vLLM servers

vllm-plugin-meralion2

A vLLM plugin to register the MERaLiON-2-10B model architecture with vLLM’s plugin system.

vllm-rbln

vLLM plugin for RBLN NPU

vllm-recorder
vllm-responses

This name has been reserved using Reserver

vllm-rocm

A high-throughput and memory-efficient inference and serving engine for LLMs with AMD GPU support

vllm-router

High-performance Rust-based load balancer for VLLM with multiple routing algorithms and prefill-decode disaggregation support

vllm-rs

A minimal, high-performance large language model (LLM) inference engine implementing vLLM in Rust.

vllm-sdk

Minimal Python SDK for the vLLM API

vllm-semantic-router-bench

Comprehensive benchmark suite for semantic router vs direct vLLM evaluation across multiple reasoning datasets

vllm-speculative-autoconfig

Automatic configuration planner for vLLM - Eliminate the guesswork of configuring vLLM by automatically determining optimal parameters

vllm-spyre

vLLM plugin for Spyre hardware support

vllm-spyre-next

Next iteration of vllm-spyre on the torch-spyre stack

vllm-sr

vLLM Semantic Router - Intelligent routing for Mixture-of-Models

vllm-sr-sim

vLLM Semantic Router fleet simulator for capacity planning, SLO validation, and what-if analysis

vllm-swift

vLLM Metal plugin powered by mlx-swift — high-performance LLM inference on Apple Silicon

vllm-test-tpu

A high-throughput and memory-efficient inference and serving engine for LLMs

vllm-tgis-adapter

vLLM adapter for a TGIS-compatible grpc server

vllm-top

A monitoring tool for vLLM metrics.

vllm-tpu

A high-throughput and memory-efficient inference and serving engine for LLMs

vllm-tuner

A Python package for tuning vLLM hyperparameters.

vllm-usf

vLLM-USF: A high-throughput and memory-efficient inference engine for LLMs (USF Custom Build)

vllm-wizard

CLI tool for vLLM configuration generation and GPU sizing

vllm-xft

A high-throughput and memory-efficient inference and serving engine for LLMs

vllm-xpu
vllm-xpu-kernels

A high-throughput and memory-efficient inference and serving engine for LLMs

vllmctl

CLI tool for launching and managing vllm model servers via SSH and tmux

vllmd

Run local models via vLLM in Docker containers

vllmlp

Collections of Logits Processors for VLLM

vllmlx

Ollama-style daemon and CLI over vllm-mlx on Apple Silicon

vllmocr

OCR using LLMs

vllmoni

A CLI tool for monitoring and managing VLLM inference servers in Docker containers

vlm-agent

A placeholder package to reserve the name llms.

vlm-robot-agent

Agente robótico basado en VLM que navega e interactúa con personas según objetivos visuales

vm-x-ai-langchain

VM-X AI Langchain Python SDK

vmanage-agent

Pip module for node to join salt-master

vmware-storage

VMware vSphere storage management: datastores, iSCSI, vSAN. Domain-focused MCP skill.

vnc-agent-bridge

Open-source Python package for AI agents to interact with VNC servers

vnrobo-agent

Lightweight heartbeat agent for VnRobo Fleet Monitor

vocab-coverage

语言模型中文识字率分析

vocallm

Production-ready Voice AI infrastructure

voice-agent-core

A conversational voice companion bot framework for Python. Plug in any LLM and voice tools to create your own assistant.

voice-agent-tequity

A voice agent framework for LLM + STT + TTS pipelines

voice-agents

A comprehensive Python library for building production-ready voice agents with multi-provider support. Features real-time streaming TTS/STT, OpenAI, ElevenLabs, and Groq integration, audio processing, and seamless conversational AI capabilities.

voice-computer-use-agent

Local desktop voice and computer-control agent with MCP tools

voice-of-agents

Synthetic user research and LLM eval harness — research-grade rigor for developers who can't afford a research team.

voice-rag

Provider-agnostic voice RAG pipeline. Plug in your voice provider, LLM, vector store, and document parsers.

voiceagent

A package to make it easy to interact with LLM's using voice

voicegpt

voice cloning with GPT

voicellm

A modular Python library for voice interactions with AI systems

voidly-agents

Python SDK for the Voidly Agent Relay — E2E encrypted agent-to-agent communication

voidly-pay-autogen

AutoGen tools for Voidly Pay — drop-in agent-to-agent payments. USDC-backed, x402, signed envelopes. Live on Base mainnet.

voidly-pay-crewai

CrewAI tools for Voidly Pay — drop-in agent-to-agent payments. USDC-backed, x402, signed envelopes. Live on Base mainnet.

Browse other source pages