Source
PyPI agents
24,677 AI agents indexed on MeshKore from PyPI. Agent frameworks and tools published to the Python Package Index. Each entry links back to its project page.
24,677 agents · ranked by popularity · refine in the directory →
Source platform: PyPI →
PyPI agents — page 123 of 247
Benchmark autonomous AI agents on task completion, tool use, goal adherence, and safety. Works with any agent — just provide a callable.
A short description of your library
plugin for simonw's llm to add a langchain agent
Polymorphic Prompt Assembler to protect LLM agents from prompt injection and prompt leak
A scaffolding tool to quickly generate a boilerplate for a custom LLM Agent Framework.
LLM Agent Toolkit provides minimal, modular interfaces for core components in LLM-based applications.
Detect silent failures in LLM agents
llm_agents
Build an LLM agent, equipped with MCP, from scratch.
pytest-style behavioral contracts for LLM agent tool-call sequences
Tree-structured multi-agent runtime with a FastAPI core, LangChain agents, and event-driven executors.
A model aggregator service for multiple LLM backends.
A package for generating diagrams using various models
LLM plugin to save prompt options alongside an alias
A small example package
Make all LLM sync models available as async
Latency and Memory Analysis of Transformer Models for Training and Inference
An LLM analysis assistant to help you understand and implement PMF
An easy-to-extend LLM annotator for robust, resumable data annotation.
CLI tool to anonymize code using local LLM before sending to Claude Code
LLM access to models by Anthropic, including the Claude series
Antibodies for LLM hallucinations
A flexible gateway for connecting and managing multiple LLM providers
LLM plugin for models hosted by Anyscale Endpoints
A package to query popular LLMs
Lightweight, pluggable adapter for multiple LLM APIs (OpenAI, Anthropic, Google)
Runs a throughput benchmark for LLM APIs, measuring generation throughput, prompt throughput, and Time To First Token (TTFT) under various concurrency levels.
A unified interface for querying multiple LLM providers via REST APIs
A client for interacting with LLM completion APIs and tracking usage.
Unified LLM API provider and engine library
The official Python library for the llm_api API
Unofficial open APIs for popular LLMs with self-hosted redirect capability
A python lib to call LLM API models.
A unified API routing library for Large Language Models
read and cache structured documents from remote for LLM agents
LLM plugin to expose a FastAPI server with compatible APIs for popular LLM clients
Debug your AI programs with ease through a web-based interface. Modify inputs and outputs, and leverage the power of custom Python filters.
LLM Application Performance Monitoring - Real-time monitoring for LLM-powered applications
LLM-App is a library for creating responsive AI applications leveraging OpenAI/Hugging Face APIs to provide responses to user queries based on live data sources. Build your own LLM application in 30 lines of code, no vector database required.
LLM plugin for Apple Foundation Models (Apple Intelligence)
A/B testing for LLMs. Statistical proof, not vibes.
LLM plugin for loading arXiv papers
AI browser pilot — test your app 60x cheaper. MCP server binary wrapper.
AI browser pilot — test your app 60x cheaper. MCP server binary wrapper.
Embed your LLM into a python function
DEPRECATED: llm-assert has been renamed to callspec. Install with: pip install callspec
Collection of simple LLM applications
A flexible framework for building AI assistants using various LLMs
LLM-A*: Large Language Model Enhanced Incremental Heuristic Search on Path Planning.
Multi-LLM Provider Library
LLM Attestation → This package is now `aiir`. Install: pip install aiir
A Python package to run LLMAttributor in your computational notebooks.
Static security analyzer for LLM applications — eslint for LLM security
CLI tool for creating code snapshots with configurable settings
An intelligent automated exploratory data analysis tool powered by Large Language Models, providing in-depth insights and visualizations for your datasets
Automatic micro-batching for HTTP LLM calls and local PyTorch inference, backed by a Rust core.
llm-autodocs 🧞: Automatically generate docstrings for your codebase using LLMs.
A placeholder package for llm_autoeval
39% faster TTFT, 67% less KV cache, zero config — autotune optimises local LLMs on Ollama, LM Studio, and MLX
A toolkit for quickly implementing llm powered functionalities.
Azure AI Foundry plugin for LLM - access Anthropic, Llama, Mistral and more via Azure
Add your description here
LLM plugin to access model deployments on Azure AI Foundry and Foundry Local
Create embeddings using the Azure API
LLM plugin to access Azure OpenAI models
Text-to-speech using the Azure OpenAI TTS API
LLM Base
Batch CLI tool for running batch inference with local LLMs
Batch-oriented LLM annotation workflows for tabular datasets with OpenAI Batch support.
A package to process CSV text data in batches using OpenAI API
Cron-friendly batch LLM processing for Polars.
Run prompts against models hosted on AWS Bedrock
LLM plugin for Amazon Titan on AWS Bedrock
LLM plugin for Anthropic's Claude on AWS Bedrock
LLM plugin for AWS Bedrock Converse API with tool calling and embeddings support
LLM plugin for Meta Llama2 on AWS Bedrock
LLM plugin for Mistral on AWS Bedrock
LLM plugin for Amazon's Nova on AWS Bedrock
An LLM plugin for supporting Amazon Titan image generators on Amazon Bedrock.
Behavioral testing for LLM applications. pytest plugin with semantic assertions, multi-turn conversation testing, and drift detection. No LLM judge needed.
Behavioral regression testing tool for LLM model upgrades. Compare model versions and detect behavioral changes.
Evaluate large-language models for undesirable behaviors such as bias.
LLM Benchmarking tool for OLLAMA
A developer-centric CLI tool to systematically evaluate and compare Large Language Models (LLMs)
CLI tool for comparing LLM API pricing, ranked by cost-effectiveness against LMSYS Arena scores
LLM evaluation harness with custom metrics, LLM-as-judge, and regression tracking
LLM Benchmark
MCP Server for LLM comparison, benchmarks, and pricing — find the best model for any task
LLM Inference Benchmark CLI - measure TTFT, TPS, ITL, E2E latency for any OpenAI-compatible API
Benchmark LLMs with 10 benchmarks & 132K+ questions. 8 providers: OpenAI, Anthropic, Groq, Together, Fireworks, DeepSeek, Ollama, HuggingFace. Unified CLI + Web dashboard.
Benchmark any LLM provider against your actual prompts — latency, cost, quality
LLM trainer
Unified Python library for LLM APIs (OpenAI, Anthropic, Gemini, xAI, Groq, custom)
LLM-Blender, an innovative ensembling framework to attain consistently superior performance by leveraging the diverse strengths and weaknesses of multiple open-source large language models (LLMs). LLM-Blender cut the weaknesses through ranking and integrate the strengths through fusing generation to enhance the capability of LLMs.
Simple interface for creating and managing LLM chains
Python library for developing LLM bots
Pre-flight LLM cost estimation and budget enforcement
Prevent LLM cache poisoning with Confidence Gap Analysis — intent-aware thresholds, zero false serves
Lite library for local caching of LLM API requests