Category
Data agents
6,796 Data AI agents indexed on MeshKore — the most complete public catalog, ranked by popularity and updated daily.
6,796 agents · ranked by popularity · refine in the directory →
Data agents — page 62 of 68
MCP server for reading LlamaIndex documents stored in Qdrant vector database
An integration of Qdrant ANN vector database backend with txtai
AI-powered quant research knowledge base & brainstorm agent
A collection of storage modules like for file management, metadata and more
A Python package for preprocessing and augmenting data for large language models by quantum Neural networks.
Synchronize users and data between radosgw clusters
A project to show good CLI practices with a fully fledged RAG system.
The rag-core-api contains the API layer for the RAG template for document retrieval, question answering, knowledge base management in the vector database and more.
Agentic RAG pipeline failure diagnosis — no database, no LLM API, no cloud
Generate QA pairs from JSON documents to evaluate RAG pipelines
Enterprise-grade data poisoning detection & alerting for RAG systems
This repository contains a project that implements a Retrieval-Augmented Generation (RAG) system using the LLaMA3 model. The project focuses on creating embeddings for instructions of a professional bioinformatic software to help users conduct biology research.
PostgreSQL pgvector-based RAG memory system with MCP server
AI-powered document analysis and query generation tool with RAG capabilities
Scan RAG documents for hidden prompt injections, invisible text attacks, and data exfiltration payloads before they enter your vector database.
Generate startup ideas grounded in real YC data using Retrieval-Augmented Generation (RAG).
A Python SDK for sending RAG trace node data
RAGCAR: Retrieval-Augmented Generative Companion for Advanced Research
RagChat transforms unstructured data for LLM interaction.
Datagen & RagEval for various LLM (Large Language Model) for RAG Apps
Build knowledge bases for RAG
A patent-pending, embedded, multimodal RAG database that performs automated ingestion, hybrid vector+keyword search, and offline retrieval entirely inside a portable single-file SQLite container.
Useful Tools for Database, RAG and LLM
A library for generating dataset and evaluating these datasets on RAG based solutions
Efficient RaggedBuffer datatype that implements 3D arrays with variable-length 2nd dimension.
RAG dataset generator
Simple vector database operations with Qdrant
Version control for your RAG pipeline — compare embedding models on your own data
Permission-aware retrieval for RAG applications
scraping stuff
A simple, clean Python library for Retrieval-Augmented Generation (RAG)
MCP Server for RAG documentation search with Qdrant and Ollama
ragl: retrieval-augmented generation (RAG) for text.
Local-first RAG toolkit backed by a single SQLite database
RAGoon : High level library for batched embeddings generation, blazingly-fast web-based RAG and quantized indexes processing ⚡
A RAG system for creating knowledge bases from different document formats
RAG in 3 functions. Ingest any data source into vector databases.
A modular RAG SDK for ingesting web, document, and API sources, chunking them, and storing embeddings in pluggable vector databases.
Unified data extraction and preprocessing toolkit for Retrieval-Augmented Generation (RAG) pipelines.
The Fastest Way to Audit Your RAG - Generate QA datasets & evaluate RAG systems in Colab, Jupyter, or CLI. Privacy-first, any LLM, visual reports.
ragsearch is a Python library designed for building a Retrieval-Augmented Generation (RAG) application that enables natural language querying over both structured and unstructured data. This tool leverages embedding models and a vector database (FAISS or ChromaDB) to provide an efficient and scalable search engine.
DataStax RAGStack
DataStax RAGStack Colbert implementation
DataStax RAGStack Knowledge Graph
DataStax RAGStack Graph Store
DataStax RAGStack Langchain
DataStax RAGStack Llama Index
A package that abstracts most of the utilities used in RAG applications
Production-grade RAG toolkit for document ingestion and retrieval with hybrid search support
A simple RAG (Retrieval-Augmented Generation) framework for Python.
An up-to-date simple random user-agent with real world database.
Client for the Research Data Storage Registry, Research Computing Services, University of Melbourne.
reshape data type in py language
reeln-cli plugin for OpenAI-powered LLM integration (metadata, translation, zoom)
Research Agent AI Agent Directory to Host All Research Agent related AI Agents Services, Community, Reviews and More.
Research API and SDK
Intelligent research paper analysis pipeline with LLM-driven categorization
A Python package for creating a research brief agent
Automated Research. Powered by LoopGPT.
Aids users in publishing data to a Record Evolution Datapod
A Python library that meters LiteLLM usage to Revenium with context-based metadata injection and framework integrations.
Fast, efficient, minimal, extendible and elegant RAG system
An EOS agent to collect and process RPKI VRP data
RPX — wrap any robot training command with full end-to-end analytics.
Run a Python script with coverage tracking and allow the user to specify the coverage data file.
Shadow-Sandbox DB Layer -- let AI agents modify your database safely with tenant isolation, Pydantic validation, and atomic sync.
A document-ingesting agent that monitors specified directories, keeping stored documents up to date in a vector database for Retrieval-Augmented Generation (RAG) queries.
Sandlake Storage SDK - 基于 S3 存储后端的模型和数据集下载/上传 SDK,提供类似 ModelScope 风格的 API
LLM components for the Sayou Data Platform
A tool for comparing populations in single-cell RNA-seq data with average overlap of marker gene lists
Supply chain compromise scanner — detects known PyPI and npm attacks via data-driven threat profiles
Secure Cloud Data Migration Agent (Linux local agent)
LLM-driven agent for describing data tables based on domain schemas
Multi-agent scientific literature research system with persistent memory
An open-source agentic harness for scientific literature analysis — multi-model data extraction, systematic reviews, and provenance tracking
Python library for scraping ChatGPT
Use a random User-Agent provided by fake-useragent for every request
Use a random User-Agent provided by fake-useragent for every request
A Scrapy extension for data extraction using LangChain
Automatically pick an User-Agent for every request
Purpose-scoped ADK agents for SDC4 data operations
TOON-Native Auto-Embedding & RAG Toolkit for MariaDB — VECTOR(N), HNSW, VEC_DISTANCE_COSINE, with TOON tabular output that saves 10-55% LLM tokens vs JSON.
A taskflow agent for the SecLab project, enabling secure and automated workflow execution.
Tools for extract sensitive configuration out of your project
AI-powered dataset segmentation agent for ML workflows
Add your description here
Semantic Kernel plugins for Built-Simple research APIs (PubMed, ArXiv, Wikipedia)
SGR Agent Core - Schema-Guided Reasoning for building agent
Local package containing the ShareGPT V3 unfiltered cleaned split dataset.
Simple HTTP client for SheetBase - Use Google Sheets as a database
SIA: Self-Improving AI framework
A small, easy to use module that helps you store data easily.
This package will help you talk to your data using Retrieval Augmented Generation (RAG).
Package that provides support for Langchain community data loaders.
SincroX Agent — lightweight SQL bridge that runs next to your database
Orchestrator, Generic Agent, and Research Agent components of the Sirji AI agentic framework.
LLM-driven self-healing API discovery for undocumented SaaS portals via CDP
Bayesian Transfer Learning for Small-Data Predictive Analytics
smappdragon is a set of tools for working with twitter data