Category

Data agents

6,796 Data AI agents indexed on MeshKore — the most complete public catalog, ranked by popularity and updated daily.

6,796 agents · ranked by popularity · refine in the directory →

Data agents — page 59 of 68

llama-index-vector-stores-zerodb

ZeroDB vector store for LlamaIndex - AI-native vector database with free embeddings, semantic search, and RAG support. Pinecone alternative.

llama-index-vector-stores-zeusdb

LlamaIndex integration for ZeusDB vector database. Enterprise-grade RAG with high-performance vector search.

llama-memory

Easy deployment of quantized llama models on cpu

llama-mlx-pipeline

Data processing pipeline using MLX (scraper, chunker, extractor).

llama-mlx-pipeline-llamasearch

Data processing pipeline using MLX (scraper, chunker, extractor).

llama-simulation

A comprehensive simulation framework for AI research and testing scenarios

llama-web-scraper

Web scraping tool potentially using Llama models.

llamachain

Blockchain intelligence and analytics platform

llamadatasets-llamasearch

Dataset management and processing library for LlamaSearch.ai applications

llamadb-llamasearch

Next-Gen Hybrid Python/Rust Data Platform with MLX

llamadb3-llamasearch

Database management and query optimization library for Python

llamaindex-soul

The soul ecosystem for LlamaIndex: persistent memory, identity, database schema intelligence, and SoulMate API integration.

llamapublisher

A powerful library for AI-powered search and data processing

llamasearch-pdf-llamasearch

A comprehensive PDF processing toolkit for document workflows

llamass

LLAMASS is a Loader for the AMASS dataset

llm-annotator

An easy-to-extend LLM annotator for robust, resumable data annotation.

llm-astar

LLM-A*: Large Language Model Enhanced Incremental Heuristic Search on Path Planning.

llm-auto-eda

An intelligent automated exploratory data analysis tool powered by Large Language Models, providing in-depth insights and visualizations for your datasets

llm-batch-annotate

Batch-oriented LLM annotation workflows for tabular datasets with OpenAI Batch support.

llm-batch-processor

A package to process CSV text data in batches using OpenAI API

llm-behavior-eval

Evaluate large-language models for undesirable behaviors such as bias.

llm-benchmark-toolkit

Benchmark LLMs with 10 benchmarks & 132K+ questions. 8 providers: OpenAI, Anthropic, Groq, Together, Fireworks, DeepSeek, Ollama, HuggingFace. Unified CLI + Web dashboard.

llm-chunk-optimizer

Semantic-aware chunking with provenance tracking for production RAG and LLM data pipelines

llm-cost-guard

Real-time cost tracking, budget enforcement, and usage analytics for LLM applications

llm-data-converter

Best open-source document to markdown converter for LLM training data. Convert PDF, Word, PowerPoint, Excel, images, URLs to clean markdown, JSON, HTML locally. Alternative to Unstructured, Docling, Marker, MarkItDown, MinerU, PaddleOCR, Tesseract

llm-data-lens

A Python package for calculating key metrics to assess LLM performance in various tasks, including extracting structured dataa.

llm-data-processor

A processor for LLM tasks

llm-data-simulator

Generate synthetic evaluation datasets from your own documents.

llm-databricks

LLM access to Databricks model serving

llm-dataclass

A dataclass interface for llms

llm-datagen

极简高性能流式数据加工库

llm-dataset-converter

Python3 library for converting between various LLM dataset formats.

llm-dataset-converter-all

Meta-library that combines all llm-dataset-converter libraries.

llm-datasets

A collection of datasets for language model training including scripts for downloading, preprocesssing, and sampling.

llm-datatech

RAG llm datatech

llm-deep-research

LLM plugin of various deep research implementations

llm-dispatcher

Intelligent LLM dispatching with performance-based routing, multimodal support, streaming, monitoring, and comprehensive analytics

llm-exo-graph

Advanced Knowledge Graph Engine with Document Processing, Semantic Search and Multi-LLM Integration

llm-explain

Tools that use LLM to explain datasets

llm-explorer

A Lakehouse LLM Explorer. Wrapper for spark, databricks and langchain processes

llm-extractinator

A framework that enables efficient extraction of structured data from unstructured text using large language models (LLMs).

llm-extractor

Extract structured, validated JSON from any LLM — OpenAI, Anthropic, Gemini — with batch extraction, caching, per-field confidence scoring, schema evolution, multi-schema extraction, output transforms, partial extraction, extraction diff, pipeline extraction, and smart auto-retry.

llm-feat

Automated feature engineering using Large Language Models (LLMs) for tabular data

llm-feature-gen

Generate interpretable feature schemas and tabular exports from text, images, tabular data, and video with LLMs

llm-file-format

Token-efficient data format converter for LLM contexts

llm-fragments-pypi

LLM fragments plugin for PyPI packages metadata

llm-function-tools

Minimal metadata and discovery helpers for typed Python tools used by llm_function runtimes.

llm-guard

LLM-Guard is a comprehensive tool designed to fortify the security of Large Language Models (LLMs). By offering sanitization, detection of harmful language, prevention of data leakage, and resistance against prompt injection attacks, LLM-Guard ensures that your interactions with LLMs remain safe and secure.

llm-html

LLM-optimized HTML cleaning: hydration extraction, token budgets, multiple output formats

llm-hygiene

a data preprocessing toolkit that makes it easy to create common LLM-related data structures; from training data to chain payloads!

llm-interaction

Async multi-provider LLM client with typed tool calling, lazy output parsing, and Azure/Databricks/OpenRouter/LiteLLM/DeepSeek/local backends

llm-intercept

A proxy server to intercept and store LLM API calls for fine-tuning dataset collection

llm-join

Fuzzy join pandas DataFrames using LLM scoring and embedding retrieval — record linkage, entity resolution, approximate join

llm-lab

A minimal, notebook-first framework for running reproducible LLM experiments and comparing multiple models over JSONL/CSV datasets.

llm-labeling-ui

LLM Labeling UI is an open source project for large language model data labeling

llm-mask

Mask sensitive data in documents using a local OpenAI-compatible LLM

llm-metadata-harvester

A tool for harvesting metadata from dataset landing pages using Large Language Models.

llm-min

Generates LLM context by scraping and summarizing documentation for Python libraries listed in a requirements.txt file.

llm-nous

Access Nous Research models via API

llm-parliament

Multi-agent debate for better AI decisions. Research-backed, local-first.

llm-parse

Parse data from documents optimised for downstream llm tasks.

llm-patch-driver

LLM Patch Driver is a framework for patching data objects using LLMs. It can generate and apply a single patch, or start the patching loop to fix complex validation issues.

llm-price

Pricing + release metadata and cost estimation for LLMs

llm-pricing

A library for scraping and managing LLM pricing information

llm-prompt-optimizer

A comprehensive framework for systematic A/B testing, optimization, performance analytics, security, and monitoring of LLM prompts across multiple providers with enterprise-ready API

llm-protocol-suite

Pydantic data models and adapters for the LLM Protocol Suite.

llm-redact

Privacy-first text redaction using local LLM models with rule generation capabilities

llm-research

A minimum Python package built on top of the LangChain framework to interact with LLM.

llm-researcher

LLM plugin for comprehensive research using Jina AI's Search Foundation APIs

llm-salvage

Salvage structured data from LLM responses that didn't follow instructions.

llm-scraper-py

Turn any webpage into structured data using LLMs

llm-stack-kit

Generate Data from LLM easily using LLMKit

llm-structured-summary

This package would process text input, such as a research paper title or abstract snippet, and generate a structured summary of the core idea or problem addressed. It uses an LLM to interpret the inpu

llm-text-splitter

A lightweight, rule-based text splitter for LLM context window management, handles multiple file formats and enriches chunks with metadata.

llm-tools-datasette

Expose Datasette instances to LLM as a tool

llm-web-crawler

LLM data collection and synthetic fine-tuning dataset pipeline

llm-web-research

A precision-focused LLM-powered web research tool that prioritizes accuracy over quantity

llm-wrapper-test1

A comprehensive Python wrapper for Large Language Models with database integration and usage tracking

llm-wrapper-testing

A comprehensive Python wrapper for Large Language Models with database integration and usage tracking

llm-youtube-transcript

Plugin for LLM that adds support for fetching transcripts from YouTube videos using Supadata API

llm2dict

Enable large language models to output structured data.

llm_etl_pipeline

LLM extraction from documents

llm_strategy

Directly Connecting Python to LLMs - Dataclasses & Interfaces <-> LLMs

llmanalyst

Talk to your CSV data with your huggingface llm models

llmaug

🪄 Dataset augmentation using LLMs

llmbuilder

A comprehensive toolkit for building, training, and deploying language models

llmcalc

Calculate LLM token costs from litellm pricing data

llmcomp

Research library for black-box experiments on language models.

llmcompressor

A library for compressing large language models utilizing the latest techniques and research in the field for both training aware and post training techniques. The library is designed to be flexible and easy to use on top of PyTorch and HuggingFace Transformers, allowing for quick experimentation.

llmcompressor-nightly

A library for compressing large language models utilizing the latest techniques and research in the field for both training aware and post training techniques. The library is designed to be flexible and easy to use on top of PyTorch and HuggingFace Transformers, allowing for quick experimentation.

llmdata

rubrics

llmdataset

llm dataset

llmdocs-mcp

Agent-first documentation platform (CLI and server)

llmesh-llove

llove — a cute, terminal-first Artifact for inspecting LLMesh data with llove

llmextract

A library to extract structured information from unstructured text using LLMs, powered by LangChain.

llmfa-agent

Financial Data Assistant - A PocketFlow-based LLM library for financial data chat with tool calling

llmfs

A filesystem-metaphor memory layer for LLMs and AI agents

llmfsd

llmfsd: LLM Fake Structured Data, faking Structured Data from any LLM

llmgateways

Protect OpenAI and Anthropic API calls from prompt injection, jailbreaks, and data-extraction attacks.

llmind-cli

Embed signed semantic-metadata layers into images, PDFs, and audio files

Browse other category pages