Category
Image agents
2,343 Image AI agents indexed on MeshKore — the most complete public catalog, ranked by popularity and updated daily.
2,343 agents · ranked by popularity · refine in the directory →
Image agents — page 2 of 24
High-fidelity HTML design and prototype guidance skill for AI agents
End-to-end RAG system design, evaluation, and optimization. 极客时间RAG训练营,RAG 10大组件全面拆解,4个实操项目吃透 RAG 全流程。RAG的落地,往往是面向业务做RAG,而不是反过来面向RAG做业务。这就是为什么我们需要针对不同场景、不同问题做针对性的调整、优化和定制化。魔鬼全在细节中,我们深入进去探究。
Train Models Contrastively in Pytorch
🤖 Beautifully designed chatbot components based on shadcn/ui
AI-First Album: Chat with your gallery using plain language! LLM Vision + RAG + Album/Gallery.
🎨 Image collector, support for custom acquisition source, compatible with Windows and MacOS!| 图像采集器,支持自定义采集源,兼容Windows和MacOS!
Building LLM-Enabled Multi Agent Applications from Scratch
AI Agnostic (Multi-user and Multi-bot) Chat with Fictional Characters. Designed with scale in mind.
Microsoft Foundry (demos, documentation, accelerators).
OpenCluely is a free, open source Cluely (alternative), built for technical interviews like DSA, OAs, and CP. It offers an invisible overlay, real-time AI help, Smart Image Processing for question capture, and multi-language support : 100% customizable and private.
Autonomous self-evolving agents. Vision-grounded layered memory and self-written skills for LLM agents that operate your computer.
Multi-agent framework for design, simulation, and auditing.
JTokkit is a Java tokenizer library designed for use with OpenAI models.
公开的 Java 后端 / AI Agent / 系统设计 / 算法面试复习资料库
🦀An agentic AI assistant that lives in your chats, inspired by nanoclaw and incorporating some of its design ideas. Built with Rust 🦀
ComfyUI-IF_AI_tools is a set of custom nodes for ComfyUI that allows you to generate prompts using a local Large Language Model (LLM) via Ollama. This tool enables you to enhance your image generation workflow by leveraging the power of language models.
✨ The agentic motion layer — an open-source, chat-native motion engine. Describe the feeling; your AI ships the animation.
The llama-cpp-agent framework is a tool designed for easy interaction with Large Language Models (LLMs). Allowing users to chat with LLM models, execute structured function calls and get structured output. Works also with models not fine-tuned to JSON output and function calls.
InnoShop is an AI-powered open source e-commerce system built on Laravel 13, designed for global commerce. It supports multiple AI models for intelligent automation, with native multi-language and multi-currency capabilities. The modular architecture and flexible plugin/theme system make it highly adaptable to diverse business needs.
基于Stable Diffusion优化的AI绘画模型。支持输入中英文文本,可生成多种现代艺术风格的高质量图像。| An optimized text-to-image model based on Stable Diffusion. Both Chinese and English text inputs are available to generate images. The model can generate high-quality images in several modern art styles.
Claude skill: design any process, idea, or problem into an ICM workspace (folder structure as agent architecture), or restructure an existing folder into one
This repository hosts a suite of specialized agents designed to power your brainstorming sessions. Each agent brings a unique perspective and tailored approach to generate creative, targeted solutions for diverse challenges. Choose the right agent for your specific brainstorming needs and unlock innovative ideas with precision.
transformers safetensors qwen3_5_moe image-text-to-text moe vlm
Bridge between LLM-Agent and Cadence Virtuoso. A new infrastructure for Agentic Analog and Mixed-Signal Design.
Chain together LLMs for reasoning & orchestrate multiple large models for accomplishing complex tasks
A wechat robot based on ChatGPT with no risk, very stable! 🚀
VMAS is a vectorized differentiable simulator designed for efficient Multi-Agent Reinforcement Learning benchmarking. It is comprised of a vectorized 2D physics engine written in PyTorch and a set of challenging multi-robot scenarios. Additional scenarios can be implemented through a simple and modular interface.
MCP-Universe is a comprehensive framework designed for RL training, benchmarking, and developing AI agents for general tool-use.
Control Figma from the command line. Full read/write access for AI agents — create shapes, text, components, set styles, export images. 100+ commands.
Official implementation for "Blended Diffusion for Text-driven Editing of Natural Images" [CVPR 2022]
Unofficial Linux packages for Claude Desktop AI assistant with automated updates.
Emergence World: A world designed to reveal what no benchmark can: emergent intelligence.
Use agent to learn agent - A skeleton course on how to design, build, and operate production AI agents
Chatbot Arena meets multi-modality! Multi-Modality Arena allows you to benchmark vision-language models side-by-side while providing images as inputs. Supports MiniGPT-4, LLaMA-Adapter V2, LLaVA, BLIP-2, and many more!
Image to text, fast.
Create a private chatgpt website via vercel
CLI toolkit for AI agents — converts images into structured JSON evidence (OCR, layout, semantics), bridging the vision gap for text-only LLMs.
Open Source Project Management with Conversational AI Task Execution. Built for teams who want conversational workflow management alongside traditional PM features. Self-hostable with modular architecture.
Local AI Assistant
Unofficial Linux packages for Claude Desktop AI assistant with automated updates.
首个中医大语言模型——“仲景”。受古代中医学巨匠张仲景深邃智慧启迪,专为传统中医领域打造的预训练大语言模型。 The first-ever Traditional Chinese Medicine large language model - "CMLM-ZhongJing". Inspired by the profound wisdom of the ancient Chinese medical master Zhang Zhongjing, it is a pre-trained large language model designed specifically for the field of Traditional Chinese Medicine.
🔍 Search local images with natural language on Android, powered by OpenAI's CLIP model. / 在 Android 上用自然语言搜索本地图片 (基于 OpenAI 的 CLIP 模型)
Open-source agent skills for generating editorial-style information cards from natural-language input.
Editable, part-aware 3D generation from text or reference images. Open-source client for nova3d.xyz.
Universal AI Agent using Amazon Bedrock, capable of customize to create/edit files, execute commands, search the web, use knowledge base, use multi-agents, generative images and more.
Contrastive Language-Image Forensic Search allows free text searching through videos using OpenAI's machine learning model CLIP
An open SDK for agentic payments. Let AI agents make payments, hold funds, and move money across chains with policy enforcement and human approval built in.
ChatGPT-Pro is an advanced application that combines the power of ChatGPT and DALL.E.
[Deprecated & ingrated in docker-agent] Docker image for a Jenkins agent which can connect to Jenkins using TCP or Websocket protocols
Official DESIGN.md files published by companies themselves. First-party design systems that AI agents read to generate consistent UI.
🤖️ 基于 Golang + Vue3 + NaiveUI 的全新的个人、团队、企业私有化AIGC平台 | 支持AI对话、AI绘画、AI视频、AI音乐、AI无限画布工作流、AI PPT、思维导图、应用广场、同时支持多用户多模型聚合平台管理使用等等,具有非常强大的运营能力以及企业内部使用能力,欢迎体验~
A filesystem designed for agents, with SOTA retrieval, automatic memory profiles, sync engine. Drop any file type (pdf, images, videos), and grep through them.
ThunderID is a high-performance, open-source identity stack designed for developers to secure and manage access for humans, AI agents, and workloads through fully composable identity flows.
Official repo of VLABench, a large scale benchmark designed for fairly evaluating VLA, Embodied Agent, and VLMs.
A CLI tool/python module for generating images from text using guided diffusion and CLIP from OpenAI.
ACP is the Agent Control Plane - a distributed agent scheduler optimized for simplicity, clarity, and control. It is designed for outer-loop agents that run without supervision, and make asynchronous tool calls like requesting human feedback on key operations. Full MCP support.
Clipboard Conqueror is a novel copy and paste copilot alternative designed to bring your very own LLM AI assistant to any text field.
AgentScope Spark Design - UI Component Library for Alibaba Cloud Apsara Lab
AI-friendly semantic HTML architecture for better human-agent collaboration.Replacing long Markdown with stable, interactive artifacts.
A highly efficient, secure, and stable application development platform with excellent performance, easy scalability, and deep integration of AI capabilities such as LLM, AI Chat, RAG, and Agents.高效、安全、稳定的服务研发平台,具备良好性能,同时易扩展,深度集成LLM、AIChat、RAG、Agent等AI能力
[CVPR' 2026] JarvisEvo: Towards a Self-Evolving Photo Editing Agent with Synergistic Editor-Evaluator Optimization
Practical system design, tools, and hands-on resources for building Gen-AI agents & agentic AI systems.
❤开箱即用❤an unofficial implement of ChatGPT in QQ/Wechat. 一个非官方的ChatGPT腾讯qq/微信(非公众号)实现版,快来把你的qq或微信变成chatgpt吧
A curated archive of breakthroughs in Agents, Architecture, Training, RAG, and On-Device AI.
Bring back Clippy on Windows 10/11!
AI Assistant that reduces the size of your application's Docker Image
Multi-agent AI pipeline that turns one line of text into a finished short-form drama: script, cinematic storyboards, character-consistent video. Provider-agnostic (OpenAI/Claude, MJ, Minimax, Veo/Sora, fal, ComfyUI). MIT.
Design Your AI Agents
🧠 世界上覆盖最全的优秀Qwen提示语大全,欢迎贡献你的提示词。🧠 The most comprehensive collection of excellent Qwen prompts in the world. Feel free to contribute your own prompts!
A versatile tool designed to help prototype intelligent assistants, agents and multi-agentic systems
Universal CPU profiler designed for humans and AI agents
CLI client for podwise.ai — turn any podcast episode into AI-powered insights, designed for use in AI agents and skills workflows.
clawdcursor compiles whatever's on screen into one UI map — accessibility tree and OCR fused into stable, addressable elements, with a screenshot only when needed — then drives apps through reusable scripts, verifying every action and routing it through a single safety gate.
Gen-Searcher: Reinforcing Agentic Search for Image Generation
Open-WebUI-Functions is a collection of custom pipelines, filters, and integrations designed to enhance Open WebUI. These functions enable seamless interactions with Azure AI, N8N, Google Gemini and other AI models, providing dynamic request handling, preprocessing, and automation.
A next-generation LLM4AD platform focused on intuitive UI interactions and seamless collaboration with AI agents, making automated algorithm design more accessible and easier to use
Alfred workflow using ChatGPT, DALL·E 2 and other models for chatting, image generation and more.
55个精选网站设计系统 DESIGN.md,可给AI Agent使用生成匹配UI
A curated registry of reusable Mercury Agent, Open Claw or Hermes Agent skills designed for real developer workflows, persistent memory, and token-efficient execution.
High-performance framework for building interactive workflow systems in Rust. Designed for complex workflows and multi-agent systems
Generate a picture book from a single prompt using OpenAI function calling, replicate, and Deep Lake
A lightweight, modular Java application framework for web and CLI development, designed for AI integration and plugin-based architecture. Enabling developers to create robust solutions with ease for building efficient and scalable applications.
智能体设计模式
This repository is a hub for AI Agent projects, including GitHub Sentinel, LanguageMentor, and ChatPPT, designed to enhance enterprise workflows, language learning, and multimodal interaction. Explore a growing family of agents geared towards revolutionizing various industries with cutting-edge AI solutions.
Jenkins agent (base image) and inbound agent Docker images
The world's most powerful open-source bio AI assistant - Access academic literature, clinical trials, drug labels, and more, all through natural conversation.
[IJCV 2024] InterGen: Diffusion-based Multi-human Motion Generation under Complex Interactions
A C++/Python implementation of the StreetLearn environment based on images from Street View, as well as a TensorFlow implementation of goal-driven navigation agents solving the task published in “Learning to Navigate in Cities Without a Map”, NeurIPS 2018
Design notes for MyInvestPilot — an agent-friendly investment workflow system built with strategy primitives, DAG engines, validation loops, and observable decision infrastructure.
Human-taught Computer-use Agent Designed for Real Windows and MacOS Desktops.
[ICML 2025 Oral] Official repo of EmbodiedBench, a comprehensive benchmark designed to evaluate MLLMs as embodied agents.
The RAG Experiment Accelerator is a versatile tool designed to expedite and facilitate the process of conducting experiments and evaluations using Azure Cognitive Search and RAG pattern.
An intelligent Wiki overlay tool designed specifically for gamers, featuring GEMINI-powered RAG (Retrieval-Augmented Generation) capabilities for smart Q&A services.
Mirror of Stable-Baselines: a fork of OpenAI Baselines, implementations of reinforcement learning algorithms
Harnessclaw is a powerful, Electron-based desktop application designed to manage, chat with, and operate AI agents and skills seamlessly.
Official Model Studio CLI(阿里云百炼 CLI)built for AI Agent frameworks, exposing models, search, multimodal, and workflow capabilities as structured tool calls.
[ICLR-2025] POGEMA stands for Partially-Observable Grid Environment for Multiple Agents. This is a grid-based environment that was specifically designed to be flexible, tunable and scalable. It can be tailored to a variety of PO-MAPF / MAPF settings.
An enterprise-grade AI retriever designed to streamline AI integration into your applications, ensuring cutting-edge accuracy.
META‑AGENTIC α‑AGI 👁️✨ — Mission 🎯 End‑to‑end: Identify 🔍 → Out‑Learn 📚 → Out‑Think 🧠 → Out‑Design 🎨 → Out‑Strategise ♟️ → Out‑Execute ⚡
[NeurIPS 2024] ReEvo: Large Language Models as Hyper-Heuristics with Reflective Evolution