Category
Image agents
2,727 Image AI agents indexed on MeshKore — the most complete public catalog, ranked by popularity and updated daily.
2,727 agents · ranked by popularity · refine in the directory →
Image agents — page 10 of 28
A lightweight Vue component and composable designed for AI chat applications to smoothly stick to the bottom of messages.
Two-layer agentic Earth Observation on satellite-class compute. Fine-tuned LFM2.5-VL perception + tool-calling agent over Sentinel-2 imagery. Galamsey detection in Ghana as the worked example.
A practical guide to AI engineering — LLMs, RAG, agents, evals, and production ops. Built for engineers who ship AI systems.
Catalogue of languages designed for Agents
Production-grade MCP server for KiCad EDA: PCB design, DRC/ERC, BOM, DFM, simulation, manufacturing.
transformers safetensors qwen3_5_moe image-text-to-text qwen qwen-agentworld
transformers safetensors gemma4_assistant text-generation image-text-to-text conversational
Colabs for text prompt steered image generators
Text23D Mechanical CAD Explorer is an AI-assisted mechanical design platform for generating and refining 3D parametric CAD models from conversational input.
Self-hostable sandbox SDK inspired by Cloudflare Sandbox SDK. works on any device
🖼️ Workshop: Build a multimodal AI agent with Haystack & GPT-4o — featuring image understanding, document retrieval, conversational memory, and human-in-the-loop safety controls
基于Chat GPT的对话系统,还支持Gemini , 文星一言,讯飞星火,通义千问,支持Midjourney 和Dall 绘画系统. 目前只开源了前端代码
Pixio is is a web application that generates stable images based on user prompts using various AI models. It provides an interface for users to enter a prompt, select models, and generate images.
Semantic Text-Image/Image-Text/Image-Image Search Engine using CLIP
Generate software design diagram images from plain text using GPT models.
Intelligent Applications with Spring AI. Practical integration of LLMs, chat interaction, image generation, and audio transcription in enterprise applications using Spring AI.
ChatGPT Telegram bot
Generate an alt-text description for an image using ollama and llama3.2-vision
The Chat-Agent system is designed to execute GPT queries using intelligent tool selection. It leverages a hierarchical approach, dynamically choosing between summarization and Retrieval-Augmented Generation (RAG) tools, further refined to specific implementations within each category.
This MCP server is responsible for generating images for both OpenAI and the Gemini Model.
This package wraps the Claude Agent SDK to execute AI agents in secure, scalable Modal containers. It provides progressive complexity—simple usage mirrors the original Agent SDK, while advanced features expose Modal's full capabilities (GPU, volumes, image customization, etc.).
Seamless Integration with Amazon Bedrock for AI-Powered Text and Image Generation in Ruby.
Obrew Server: A self-hostable machine learning engine. Build agents and schedule workflows private to you.
A self-hosted omni-modal agentic retrieval platform that understands, searches, and explores documents, images, audio, and video through semantic routing, hybrid retrieval, adaptive agents, and traceable citations.
Real-time AI Avatar interface powered by Google Gemini. Features a customizable VRM anime-style character with low-latency voice interaction, conceptually inspired by xAI's Ani
AI-powered Rhino/Grasshopper MCP with ML auto-layout
Self-hostable channel-native AI teammate for Slack. Open source alternative to Claude Tag. LLM-agnostic.
transformers safetensors qwen3_5_moe image-text-to-text qwen world-model
Token-efficient browser automation MCP for AI agents, with semantic page snapshots and stable element IDs.
Eval-driven multimodal agent for identifying trees, logs, bark, leaves and wood from photographs. Dendrology is the reference domain; the engineering subject is evidence handling, calibrated uncertainty and evaluation.
Qt/C++ OpenAI-compatible reasoning degradation guard proxy with GUI, CLI, deb and AppImage packaging
A Next.js multi-provider AI workspace for chat, search, reasoning, image and video generation, and browser voice interaction.
Telegram Advanced AI Bot (50k+ lifetime users ) : GPT-5, Qwen-3, DeepSeek-R1, Dall-E-3, Flux, Flux-Pro, Dall-E Model, OCR and Google Voice2Text.
Bring any Character.AI persona into Discord as a real webhook — with its name, avatar, and personality. Supports search, follow modes, streaming replies, and regeneration.
AI-powered image editor with LangGraph agent technology. Edit photos using natural language commands through an intelligent conversational interface. Features both traditional manual controls and an AI assistant that understands complex editing requests.
AutoGroqAgent is a Streamlit-based demo application designed to showcase the powerful new AutonomousAgent class introduced in PocketGroq v0.5.x. This demo provides a hands-on experience with the advanced capabilities of the AutonomousAgent, demonstrating its potential for creating intelligent, autonomous AI assistants.
SylvaDeck is an agent-native presentation workbench for turning structured thinking into editable decks. A local-first, HTML-based AI presentation Creator
APEX defines how AI agents communicate with brokers, exchanges, dealers, and other execution venues. One protocol. Realtime state. Autonomous safety. Multi-asset by design.
Unofficial Facebook Messenger Platform *chatbot client* and *webhook handler*
Multi Model Personal Assistant Wrapper in Go: Interact with ChatGPT, Claude or Ollama Cross Platform (Speech & Image generation supported)
Discord bot with Gemini, GPT-3.5, Microsoft Copilot, Bing Image Creator, DALL·E 3, Other AI Applications
llmon-py is a multimodal webui for Llama 3-8B.
OpenChronicle is an open-source, self-hosted interaction engine for LLMs that makes conversations durable. It adds persistent memory, deterministic tasking, and auditable decision trails so work doesn’t reset each session. Provider-agnostic by design, it supports multiple interfaces (CLI, Discord, MCP) with explicit, privacy-aware routing.
A fully locally deployable, NotebookLM-like application designed for long-form unstructured documents
CLI & async Python library for free AI chat, image & video generation.
A Selenium-based chatbot for Kick (Kick.com) that automatically logs in using stored cookies, sends messages from a text file to a specified channel, and supports both headed and headless modes. Features include sequential message sending, status indicators, real-time message display, and adjustable message intervals.
Agent + Skill System for Controlling Figma and Automate Design Process
Hands-on AI learning notebooks: Tiktoken tokenization, Groq/Ollama LLM inference, and Hugging Face FLUX image generation.
Shuttle is a multi-purpose AI having the capability to generate Text/Images dedicated to enhancing your Discord experience.
OneSDK: A unified AI access SDK for edge devices, providing LLM capabilities (text/voice chat, image generation) and IoT device management with MQTT support, compatible with ESP32, Linux, macOS, and Windows platforms.
SANCTIS is a cognitive architecture that gives LLMs structured tools to organize their own reasoning, maintain long-horizon coherence, and reduce drift. It installs stable thinking patterns that emerge more powerfully the longer the model runs.
An app that uses Hugging Face AI models together with OpenAI & LangChain, to generate text from an image, which then generates audio from the text
Xaibo is a modular agent framework designed for building flexible AI systems with clean protocol-based interfaces.
AI customer care infrastructure management for businesses and Indi-Hackers. The AI Chatbot Creator is a web application designed to empower users to create, deploy, and manage AI-powered chatbots seamlessly.
🚀 A blazingly fast, modern chat interface built with Rust/WASM and Leptos, featuring local AI model execution via WebLLM. Privacy-first design with no backend dependencies.
[CVPR2026] InterAgent: Physics-based Multi-agent Command Execution via Diffusion on Interaction Graphs
Agent + Skill System for Controlling Figma and Automate Design Process
Open-source recreation of Claude Design (claude.ai/design) as a pure web app — chat-driven multi-file React prototyping with Tweaks · Comment · Edit · Draw, agent in your browser tab, single Cloudflare Worker for storage
gradio region:us
A curated list of benchmarks for evaluating Earth-Observation (EO) AI agents
Curso System Design Interview Notebook LM - From Zero to Production
Okra, your all in one personal AI assistant
Simple bots or Simbots is a library designed to create simple bots using the power of python. This library utilises Intent, Entity, Relation and Context model to create bots .
WORLD-TO-IMAGE: GROUNDING TEXT-TO-IMAGE GENERATION WITH AGENT-DRIVEN WORLD KNOWLEDGE
Aurora is an intelligent voice assistant designed to enhance productivity through local, privacy-focused automation. It leverages real-time speech-to-text, a large language model (LLM), and open-source tools to provide a seamless and intuitive user experience. Aurora integrates with tools like OpenRecall for semantic search of daily activities
CARD: Towards Conditional Design of Multi-agent Topological Structures. Published at ICLR 2026.
Docker image for a self-hosted Docling document parsing server. Converts PDF, DOCX, PPTX, HTML, and more to Markdown/JSON. Powered by IBM Docling. Features sync/async conversion, chunking for RAG, NVIDIA GPU (CUDA) acceleration, optional web UI, offline mode, and persistent model cache. Multi-arch: amd64, arm64.
Apple Human Interface Guidelines archive (1980-2014) - 35 documents optimized for LLM consumption and human exploration. Spanning Lisa, Mac, NeXT, Newton, Aqua, and iOS eras.
gguf moe vlm vision agentic text-generation
Generador de programas formativos completos con Claude, módulo a módulo: 5 entregables por módulo (guía del formador, presentación, manual del participante, evaluaciones con clave y guión). Cualquier temática.
Curated resources for building AI Waifu and intelligent companions: frameworks, platforms, tools, models, and communities
基于ChatGPT模型开发的AI工具微信小程序,提供聊天机器人、绘画助手等功能,支持用户通过文本和语音与 ChatGPT 交流,并且还具备画图功能,支持预览绘制的图片并可长按发送给微信好友。 WeChat Mini Program, an AI tool developed based on the ChatGPT model, provides functions such as chatbot and drawing assistant. It supports users to communicate with ChatGPT through text and voice, and also has drawing function.
Pure-Rust port of CerebroCortex — brain-analogous AI memory system with ACT-R/FSRS activation, associative networks, 8-phase dream engine, and MCP-over-stdio. Drop-in replacement for the Python version, Pi-native, zero Python runtime.
MCP server for controlling KeyShot Studio through headless scripting.通过无头脚本控制 KeyShot Studio 的 MCP 服务器
Fully local document intelligence for DeepSeek Harness. Parse PDF, Office files, images, and scanned documents with offline OCR. | DeepSeek Harness 全本地文档智能插件,支持 PDF、Office、图片与离线 OCR
🦫 Open-source iOS AI companion with a live VRM avatar that sees you — bring-your-own-key chat, vision, speech-to-text & text-to-speech (OpenAI, Anthropic, Gemini, ElevenLabs, Groq & more). SwiftUI.
An OpenAI GPT-powered chatbot utilizing deforum stable diffusion to aid developers in efficiently searching through documentations.
Agent Message Transfer Protocol is a federated communication protocol designed for reliable agent-to-agent communication across organizational boundaries.
A stable discord application uses discord.py, currently working for new features
The package for AI chats and AI images.
本项目主要是2025届浙江大学软件学院夏令营(AI营)的考核项目
Open-source AI desktop assistant — run local or remote LLMs with privacy-first design and zero setup required.
Token-efficient browser automation MCP for AI agents, with semantic page snapshots and stable element IDs.
Your AI companion, brought to life through a dynamic Live2D avatar.
Local AI for Apple Silicon. Chat with open LLMs (Qwen, Gemma, DeepSeek, Mistral, Llama…) fully offline on your Mac via Apple MLX — private by design, zero telemetry.
Raven is a self-hostable team Agent platform providing isolated workspaces and unified runtimes for multi-agent workflows.
A practical Generative AI engineering lab exploring LLM foundations, tokenization, deployment strategies, and real-world AI system design.
⚡ ƝØVΛ — Free & open-source AI Agent for Telegram. Builds & hosts web apps, generates games, images, PDFs & voice. 100% free tier. Self-host in 10 min. AGPL-3.0.
Dissi is a high-performance, real-time communication agent powered by Groq and built with Agno, designed to interact with Discord servers using natural language.
An agentic solution designed to create a llms.txt file for any given repo or folder.
A powerful, multi-modal Telegram bot leveraging cutting-edge AI technologies including Gemini, DeepSeek, OpenRouter, and 50+ AI models for comprehensive conversational assistance, media processing, and collaborative features with MCP (Model Context Protocol) integration.
ChatGPT + DALL-E + WhatsApp = AI Assistant 🚀
Chatbot Designed to Help Tenants Facing Eviction and Other Landlord-Tenant Law Issues
Convert any image into a Region Adjacency Graph (RAG)
Transform ideas into production-ready UML diagrams, ER schemas, and system architectures with AI.
AI Character Chat that's actually private
PolicyRAG : A RAG Based policy agent designed to guide through the process of finding the best policy that fits the requirement with no effort of reading
A multichannel multimodal AI agent deployed on Amazon Bedrock AgentCore Runtime with Amazon Bedrock AgentCore Memory, demonstrated through two WhatsApp integration patterns. The agent processes text, images, audio, video, and documents, converting all multimedia into text understanding before storing it in memory
[ICLR'26] InSight-o3: Empowering Multimodal Foundation Models with Generalized Visual Search
Kling AI Pro — generative video AI for text-to-video, image animation and cinematic clips with extended duration and priority rendering.