Category
Data agents
6,796 Data AI agents indexed on MeshKore — the most complete public catalog, ranked by popularity and updated daily.
6,796 agents · ranked by popularity · refine in the directory →
Data agents — page 59 of 68
ZeroDB vector store for LlamaIndex - AI-native vector database with free embeddings, semantic search, and RAG support. Pinecone alternative.
LlamaIndex integration for ZeusDB vector database. Enterprise-grade RAG with high-performance vector search.
Easy deployment of quantized llama models on cpu
Data processing pipeline using MLX (scraper, chunker, extractor).
Data processing pipeline using MLX (scraper, chunker, extractor).
A comprehensive simulation framework for AI research and testing scenarios
Web scraping tool potentially using Llama models.
Blockchain intelligence and analytics platform
Dataset management and processing library for LlamaSearch.ai applications
Next-Gen Hybrid Python/Rust Data Platform with MLX
Database management and query optimization library for Python
The soul ecosystem for LlamaIndex: persistent memory, identity, database schema intelligence, and SoulMate API integration.
A powerful library for AI-powered search and data processing
A comprehensive PDF processing toolkit for document workflows
LLAMASS is a Loader for the AMASS dataset
An easy-to-extend LLM annotator for robust, resumable data annotation.
LLM-A*: Large Language Model Enhanced Incremental Heuristic Search on Path Planning.
An intelligent automated exploratory data analysis tool powered by Large Language Models, providing in-depth insights and visualizations for your datasets
Batch-oriented LLM annotation workflows for tabular datasets with OpenAI Batch support.
A package to process CSV text data in batches using OpenAI API
Evaluate large-language models for undesirable behaviors such as bias.
Benchmark LLMs with 10 benchmarks & 132K+ questions. 8 providers: OpenAI, Anthropic, Groq, Together, Fireworks, DeepSeek, Ollama, HuggingFace. Unified CLI + Web dashboard.
Semantic-aware chunking with provenance tracking for production RAG and LLM data pipelines
Real-time cost tracking, budget enforcement, and usage analytics for LLM applications
Best open-source document to markdown converter for LLM training data. Convert PDF, Word, PowerPoint, Excel, images, URLs to clean markdown, JSON, HTML locally. Alternative to Unstructured, Docling, Marker, MarkItDown, MinerU, PaddleOCR, Tesseract
A Python package for calculating key metrics to assess LLM performance in various tasks, including extracting structured dataa.
A processor for LLM tasks
Generate synthetic evaluation datasets from your own documents.
LLM access to Databricks model serving
A dataclass interface for llms
极简高性能流式数据加工库
Python3 library for converting between various LLM dataset formats.
Meta-library that combines all llm-dataset-converter libraries.
A collection of datasets for language model training including scripts for downloading, preprocesssing, and sampling.
RAG llm datatech
LLM plugin of various deep research implementations
Intelligent LLM dispatching with performance-based routing, multimodal support, streaming, monitoring, and comprehensive analytics
Advanced Knowledge Graph Engine with Document Processing, Semantic Search and Multi-LLM Integration
Tools that use LLM to explain datasets
A Lakehouse LLM Explorer. Wrapper for spark, databricks and langchain processes
A framework that enables efficient extraction of structured data from unstructured text using large language models (LLMs).
Extract structured, validated JSON from any LLM  OpenAI, Anthropic, Gemini  with batch extraction, caching, per-field confidence scoring, schema evolution, multi-schema extraction, output transforms, partial extraction, extraction diff, pipeline extraction, and smart auto-retry.
Automated feature engineering using Large Language Models (LLMs) for tabular data
Generate interpretable feature schemas and tabular exports from text, images, tabular data, and video with LLMs
Token-efficient data format converter for LLM contexts
LLM fragments plugin for PyPI packages metadata
Minimal metadata and discovery helpers for typed Python tools used by llm_function runtimes.
LLM-Guard is a comprehensive tool designed to fortify the security of Large Language Models (LLMs). By offering sanitization, detection of harmful language, prevention of data leakage, and resistance against prompt injection attacks, LLM-Guard ensures that your interactions with LLMs remain safe and secure.
LLM-optimized HTML cleaning: hydration extraction, token budgets, multiple output formats
a data preprocessing toolkit that makes it easy to create common LLM-related data structures; from training data to chain payloads!
Async multi-provider LLM client with typed tool calling, lazy output parsing, and Azure/Databricks/OpenRouter/LiteLLM/DeepSeek/local backends
A proxy server to intercept and store LLM API calls for fine-tuning dataset collection
Fuzzy join pandas DataFrames using LLM scoring and embedding retrieval — record linkage, entity resolution, approximate join
A minimal, notebook-first framework for running reproducible LLM experiments and comparing multiple models over JSONL/CSV datasets.
LLM Labeling UI is an open source project for large language model data labeling
Mask sensitive data in documents using a local OpenAI-compatible LLM
A tool for harvesting metadata from dataset landing pages using Large Language Models.
Generates LLM context by scraping and summarizing documentation for Python libraries listed in a requirements.txt file.
Access Nous Research models via API
Multi-agent debate for better AI decisions. Research-backed, local-first.
Parse data from documents optimised for downstream llm tasks.
LLM Patch Driver is a framework for patching data objects using LLMs. It can generate and apply a single patch, or start the patching loop to fix complex validation issues.
Pricing + release metadata and cost estimation for LLMs
A library for scraping and managing LLM pricing information
A comprehensive framework for systematic A/B testing, optimization, performance analytics, security, and monitoring of LLM prompts across multiple providers with enterprise-ready API
Pydantic data models and adapters for the LLM Protocol Suite.
Privacy-first text redaction using local LLM models with rule generation capabilities
A minimum Python package built on top of the LangChain framework to interact with LLM.
LLM plugin for comprehensive research using Jina AI's Search Foundation APIs
Salvage structured data from LLM responses that didn't follow instructions.
Turn any webpage into structured data using LLMs
Generate Data from LLM easily using LLMKit
This package would process text input, such as a research paper title or abstract snippet, and generate a structured summary of the core idea or problem addressed. It uses an LLM to interpret the inpu
A lightweight, rule-based text splitter for LLM context window management, handles multiple file formats and enriches chunks with metadata.
Expose Datasette instances to LLM as a tool
LLM data collection and synthetic fine-tuning dataset pipeline
A precision-focused LLM-powered web research tool that prioritizes accuracy over quantity
A comprehensive Python wrapper for Large Language Models with database integration and usage tracking
A comprehensive Python wrapper for Large Language Models with database integration and usage tracking
Plugin for LLM that adds support for fetching transcripts from YouTube videos using Supadata API
Enable large language models to output structured data.
LLM extraction from documents
Directly Connecting Python to LLMs - Dataclasses & Interfaces <-> LLMs
Talk to your CSV data with your huggingface llm models
🪄 Dataset augmentation using LLMs
A comprehensive toolkit for building, training, and deploying language models
Calculate LLM token costs from litellm pricing data
Research library for black-box experiments on language models.
A library for compressing large language models utilizing the latest techniques and research in the field for both training aware and post training techniques. The library is designed to be flexible and easy to use on top of PyTorch and HuggingFace Transformers, allowing for quick experimentation.
A library for compressing large language models utilizing the latest techniques and research in the field for both training aware and post training techniques. The library is designed to be flexible and easy to use on top of PyTorch and HuggingFace Transformers, allowing for quick experimentation.
rubrics
llm dataset
Agent-first documentation platform (CLI and server)
llove — a cute, terminal-first Artifact for inspecting LLMesh data with llove
A library to extract structured information from unstructured text using LLMs, powered by LangChain.
Financial Data Assistant - A PocketFlow-based LLM library for financial data chat with tool calling
A filesystem-metaphor memory layer for LLMs and AI agents
llmfsd: LLM Fake Structured Data, faking Structured Data from any LLM
Protect OpenAI and Anthropic API calls from prompt injection, jailbreaks, and data-extraction attacks.
Embed signed semantic-metadata layers into images, PDFs, and audio files