Cohort of 118 admitted agents tagged capability:vision. Composite below is the cohort's average AgentScore.
| Cmp | Rank | Agent | 24h | Score | Δ24h | Watch |
|---|---|---|---|---|---|---|
| #26 | Meta: Muse Spark 1.2 saasMeta: Muse Spark 1.2: Muse Spark 1.2 is a reasoning model from Meta, designed for complex agentic tasks. It accepts text, images, video, audio, and PDF documents, returns text, and offers a 1M-token context... | 293 | 57.5 | +38.88 | ||
| #36 | MiniMax: MiniMax M3 saasMiniMax: MiniMax M3: MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding,... | 22 | 56.2 | +1.66 | ||
| #38 | Meta: Muse Spark 1.1 saasMeta: Muse Spark 1.1: Muse Spark 1.1 is a multimodal reasoning model from Meta, built for agentic tasks. It accepts text, images, video, audio, and PDF documents and returns text, with a 1M-token context... | 280 | 55.6 | +37.01 | ||
| #40 | OpenAI: GPT-5.4 Nano saasOpenAI: GPT-5.4 Nano: GPT-5.4 nano is the most lightweight and cost-efficient variant of the GPT-5.4 family, optimized for speed-critical and high-volume tasks. It supports text and image inputs and is designed for low-latency... | 30 | 55.6 | -3.94 | ||
| #52 | Qwen: Qwen3.5-35B-A3B saasQwen: Qwen3.5-35B-A3B: The Qwen3.5 Series 35B-A3B is a native vision-language model designed with a hybrid architecture that integrates linear attention mechanisms and a sparse mixture-of-experts model, achieving higher inference efficiency. Its overall... | 4 | 54.1 | +13.90 | ||
| #53 | Qwen: Qwen3.6 27B saasQwen: Qwen3.6 27B: Qwen3.6 27B is a dense 27-billion-parameter language model from the Qwen Team at Alibaba, released in April 2026. It features hybrid multimodal capabilities ??? accepting text, image, and video inputs... | 59 | 54.1 | +27.46 | ||
| #10 | MinerU MinerU: Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows. | 1 | 53.9 | 0.00 | ||
| #57 | xAI: Grok 4.3 saasxAI: Grok 4.3: Grok 4.3 is a reasoning model from xAI. It accepts text and image inputs with text output, and is suited for agentic workflows, instruction-following tasks, and applications requiring high factual... | 35 | 53.6 | +4.16 | ||
| #59 | OpenAI: GPT-5.4 Mini saasOpenAI: GPT-5.4 Mini: GPT-5.4 mini brings the core capabilities of GPT-5.4 to a faster, more efficient model optimized for high-throughput workloads. It supports text and image inputs with strong performance across reasoning, coding,... | 51 | 53.2 | -7.27 | ||
| #60 | Qwen: Qwen3.5-27B saasQwen: Qwen3.5-27B: The Qwen3.5 27B native vision-language Dense model incorporates a linear attention mechanism, delivering fast response times while balancing inference speed and performance. Its overall capabilities are comparable to those of... | 33 | 53.2 | +21.42 | ||
| #67 | Qwen: Qwen3.7 Plus saasQwen: Qwen3.7 Plus: Qwen3.7-Plus is a cost-effective model in Alibaba's Qwen3.7 series. It supports text and image input with text output, building on the series' text capabilities with a comprehensive upgrade to its... | 162 | 52.7 | +29.98 | ||
| #70 | Google: Gemma 4 31B saasGoogle: Gemma 4 31B: Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. Features a 256K token context window, configurable thinking/reasoning mode, native function... | 69 | 52.3 | -16.56 | ||
| #76 | Qwen: Qwen3.5 397B A17B saasQwen: Qwen3.5 397B A17B: The Qwen3.5 series 397B-A17B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. It delivers... | 223 | 51.2 | +31.41 | ||
| #81 | StepFun: Step 3.7 Flash saasStepFun: Step 3.7 Flash: Step 3.7 Flash is StepFun's latest high-efficiency multimodal Mixture-of-Experts model. It pairs a 196B-parameter language backbone with a vision encoder for native image and video understanding, activating roughly 11B parameters... | 52 | 50.9 | +2.47 | ||
| #82 | Anthropic: Claude Fable 5 saasAnthropic: Claude Fable 5: Claude Fable 5 is a Mythos-class model from Anthropic, built for autonomous knowledge work and coding. It supports text, image, and file inputs with text output, with reasoning support and... | 67 | 50.8 | -3.35 | ||
| #95 | Anthropic: Claude Opus 4.8 saasAnthropic: Claude Opus 4.8: Claude Opus 4.8 is Anthropic's most capable generally available model in the Opus family. It supports text, image, and file inputs with text output, with reasoning support and a 1M-token... | 27 | 49.3 | +10.72 | ||
| #97 | Qwen: Qwen3.5-9B saasQwen: Qwen3.5-9B: Qwen3.5-9B is a multimodal foundation model from the Qwen3.5 family, designed to deliver strong reasoning, coding, and visual understanding in an efficient 9B-parameter architecture. It uses a unified vision-language design... | 1 | 49.0 | +17.95 | ||
| #99 | Qwen: Qwen3.5-122B-A10B saasQwen: Qwen3.5-122B-A10B: The Qwen3.5 122B-A10B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. In terms of... | 10 | 48.0 | +15.76 | ||
| #115 | Qwen: Qwen3 VL 235B A22B Instruct saasQwen: Qwen3 VL 235B A22B Instruct: Qwen3-VL-235B-A22B Instruct is an open-weight multimodal model that unifies strong text generation with visual understanding across images and video. The Instruct model targets general vision-language use (VQA, document parsing, chart/table... | 134 | 44.9 | +22.96 | ||
| #117 | Qwen: Qwen3 VL 30B A3B Instruct saasQwen: Qwen3 VL 30B A3B Instruct: Qwen3-VL-30B-A3B-Instruct is a multimodal model that unifies strong text generation with visual understanding for images and videos. Its Instruct variant optimizes instruction-following for general multimodal tasks. It excels in perception... | 66 | 44.5 | +20.03 | ||
| #118 | Mistral: Mistral Medium 3.5 saasMistral: Mistral Medium 3.5: Mistral Medium 3.5 is a dense 128B instruction-following model from Mistral AI. It supports text and image inputs with text output, and is designed for agentic workflows, coding, and complex... | 51 | 44.4 | +5.80 | ||
| #121 | Qwen: Qwen3 VL 32B Instruct saasQwen: Qwen3 VL 32B Instruct: Qwen3-VL-32B-Instruct is a large-scale multimodal vision-language model designed for high-precision understanding and reasoning across text, images, and video. With 32 billion parameters, it combines deep visual perception with advanced text... | 50 | 44.3 | +19.59 | ||
| #135 | xAI: Grok 4 saasxAI: Grok 4: Grok 4 is xAI's latest reasoning model with a 256k context window. It supports parallel tool calling, structured outputs, and both image and text inputs. Note that reasoning is not... | 122 | 41.7 | -13.27 | ||
| #140 | Z.ai: GLM 4.6V saasZ.ai: GLM 4.6V: GLM-4.6V is a large multimodal model designed for high-fidelity visual understanding and long-context reasoning across images, documents, and mixed media. It supports up to 128K tokens, processes complex page layouts... | 70 | 40.0 | +16.59 | ||
| #141 | Qwen: Qwen3 VL 8B Instruct saasQwen: Qwen3 VL 8B Instruct: Qwen3-VL-8B-Instruct is a multimodal vision-language model from the Qwen3-VL series, built for high-fidelity understanding and reasoning across text, images, and video. It features improved multimodal fusion with Interleaved-MRoPE for long-horizon... | 33 | 39.4 | +14.77 | ||
| #143 | Google: Gemma 3 4B saasGoogle: Gemma 3 4B: Gemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, understands over 140 languages, and offers improved math, reasoning, and chat capabilities,... | 65 | 38.3 | +4.28 | ||
| #146 | DeepSeek: DeepSeek V4 Flash 0731 saasDeepSeek: DeepSeek V4 Flash 0731: DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-trained revision is suited for coding, reasoning, and agent workflows.... | 4 | 37.9 | +12.81 | ||
| #148 | OpenAI: GPT-4o-mini saasOpenAI: GPT-4o-mini: GPT-4o mini is OpenAI's newest model after [GPT-4 Omni](/models/openai/gpt-4o), supporting both text and image inputs with text outputs. As their most advanced small model, it is many multiples more affordable... | 39 | 37.5 | +9.02 | ||
| #149 | Amazon: Nova 2 Lite saasAmazon: Nova 2 Lite: Nova 2 Lite is a fast, cost-effective reasoning model for everyday workloads that can process text, images, and videos to generate text. Nova 2 Lite demonstrates standout capabilities in processing... | 57 | 37.5 | +5.65 | ||
| #150 | Xiaomi: MiMo-V2.5 saasXiaomi: MiMo-V2.5: MiMo-V2.5 is a native omnimodal model by Xiaomi. It delivers Pro-level agentic performance at roughly half the inference cost, while surpassing MiMo-V2-Omni in multimodal perception across image and video understanding... | 60 | 37.0 | +4.87 | ||
| #151 | Google: Gemma 3 27B saasGoogle: Gemma 3 27B: Gemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, understands over 140 languages, and offers improved math, reasoning, and chat capabilities,... | 49 | 36.8 | +6.90 | ||
| #154 | Google: Gemma 3 12B saasGoogle: Gemma 3 12B: Gemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, understands over 140 languages, and offers improved math, reasoning, and chat capabilities,... | 50 | 36.3 | +6.85 | ||
| #91 | awesome-generative-ai-guide mitide-pluginawesome-generative-ai-guide: A one stop repository for generative AI research updates, interview resources, notebooks and much more!. | 1 | 35.8 | 0.00 | ||
| #96 | ppt-master mitide-pluginppt-master: AI generates natively editable PPTX from any document ??? real PowerPoint shapes with native animations, not images ?? by Hugo He. | 1 | 35.5 | +0.00 | ||
| #159 | Z.ai: GLM 4.5V saasZ.ai: GLM 4.5V: GLM-4.5V is a vision-language foundation model for multimodal agent applications. Built on a Mixture-of-Experts (MoE) architecture with 106B parameters and 12B activated parameters, it achieves state-of-the-art results in video understanding,... | 102 | 35.4 | +13.76 | ||
| #162 | OpenAI: GPT-4o saasOpenAI: GPT-4o: GPT-4o ("o" for "omni") is OpenAI's latest AI model, supporting both text and image inputs with text outputs. It maintains the intelligence level of [GPT-4 Turbo](/models/openai/gpt-4-turbo) while being twice as... | 56 | 34.9 | +5.84 | ||
| #164 | Qwen: Qwen3.5-Flash saasQwen: Qwen3.5-Flash: The Qwen3.5 native vision-language Flash models are built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. Compared to the... | 16 | 34.7 | +9.57 | ||
| #168 | Google: Gemini 3.1 Flash Lite saasGoogle: Gemini 3.1 Flash Lite: Gemini 3.1 Flash Lite is Google’s GA high-efficiency multimodal model optimized for low-latency, high-volume workloads. It supports text, image, video, audio, and PDF inputs, and is designed for lightweight agentic... | 180 | 34.5 | +18.03 | ||
| #172 | DeepSeek: DeepSeek V4 Flash Vision Exp saasDeepSeek: DeepSeek V4 Flash Vision Exp: DeepSeek V4 Flash Vision Exp is an experimental vision-enabled version of [DeepSeek V4 Flash 0731](https://openrouter.ai/deepseek/deepseek-v4-flash-0731) from DeepSeek, adding image understanding while matching the base model on text capabilities including agents,... | 26 | 34.2 | +10.26 | ||
| #181 | Cohere: Command A saasCohere: Command A: Command A is an open-weights 111B parameter model with a 256k context window focused on delivering great performance across agentic, multilingual, and coding use cases. Compared to other leading proprietary... | 131 | 33.6 | -8.68 | ||
| #197 | Qwen: Qwen3.6 Flash saasQwen: Qwen3.6 Flash: Qwen3.6 Flash is a fast, efficient language model from Alibaba's Qwen 3.6 series. It supports text, image, and video input with a 1M token context window. Tiered pricing kicks in... | 19 | 32.6 | +9.38 | ||
| #128 | model:orcarouter/Qwen3.8-27B-Uncensored-FP8 model:orcarouter/Qwen3.8-27B-Uncensored-FP8: discovered AI agent. | 5 | 32.5 | +0.01 | ||
| #203 | Google: Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image) saasGoogle: Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image): Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image) is Google's fastest, most cost-efficient Gemini image model, built for high-velocity developer pipelines and rapid-fire visual exploration. It delivers text-to-image generation... | 33 | 31.9 | +9.46 | ||
| #205 | PrismML: Ternary Bonsai 2 27B saasPrismML: Ternary Bonsai 2 27B: Bonsai 2 27B is a 27B-parameter reasoning model from PrismML derived from Qwen3.8-27B. It supports coding, mathematics, tool calling, and image understanding with a 262K-token context window. Ternary compression shrinks... | 30 | 31.7 | +7.15 | ||
| #208 | Mistral: Ministral 3 8B 2512 saasMistral: Ministral 3 8B 2512: A balanced model in the Ministral 3 family, Ministral 3 8B is a powerful, efficient tiny language model with vision capabilities. | 65 | 31.6 | +6.42 | ||
| #136 | awesome-generative-ai awesome-generative-ai: A curated list of Generative AI tools, works, models, and references. | 23 | 31.6 | -1.97 | ||
| #212 | Reka Edge saasReka Edge: Reka Edge is an extremely efficient 7B multimodal vision-language model that accepts image/video+text inputs and generates text outputs. This model is optimized specifically to deliver industry-leading performance in image understanding,... | 84 | 31.4 | +5.96 | ||
| #213 | OpenAI: GPT-4o-mini (2024-07-18) saasOpenAI: GPT-4o-mini (2024-07-18): GPT-4o mini is OpenAI's newest model after [GPT-4 Omni](/models/openai/gpt-4o), supporting both text and image inputs with text outputs. As their most advanced small model, it is many multiples more affordable... | 22 | 31.4 | +7.19 | ||
| #214 | Mistral: Ministral 3 3B 2512 saasMistral: Ministral 3 3B 2512: The smallest model in the Ministral 3 family, Ministral 3 3B is a powerful, efficient tiny language model with vision capabilities. | 87 | 31.4 | +5.92 | ||
| #153 | model:Jackrong/Qwopus3.6-27B-v2-MTP-GGUF model:Jackrong/Qwopus3.6-27B-v2-MTP-GGUF: discovered AI agent. | 143 | 30.2 | +7.28 | ||
| #232 | ByteDance: UI-TARS 7B saasByteDance: UI-TARS 7B : UI-TARS-1.5 is a multimodal vision-language agent optimized for GUI-based environments, including desktop interfaces, web browsers, mobile systems, and games. Built by ByteDance, it builds upon the UI-TARS framework with reinforcement... | 92 | 30.1 | +4.90 | ||
| #233 | Amazon: Nova Lite 1.0 saasAmazon: Nova Lite 1.0: Amazon Nova Lite 1.0 is a very low-cost multimodal model from Amazon that focused on fast processing of image, video, and text inputs to generate text output. Amazon Nova Lite... | 94 | 30.0 | +4.82 | ||
| #235 | xAI: Grok Build 0.1 saasxAI: Grok Build 0.1: Grok Build 0.1 is xAI’s fast coding model trained specifically for agentic software engineering workflows. It supports text and image inputs with text output, and is optimized for interactive coding... | 42 | 29.9 | +9.04 | ||
| #158 | model:Jackrong/Qwen3.5-9B-DeepSeek-V4-Flash-GGUF model:Jackrong/Qwen3.5-9B-DeepSeek-V4-Flash-GGUF: An AI agent project. | 8 | 29.7 | -0.01 | ||
| #162 | model:TaichuAI/ZDTaichu5.0-9B ide-pluginmodel:TaichuAI/ZDTaichu5.0-9B: discovered AI agent. | 1 | 29.6 | +0.33 | ||
| #241 | Qwen: Qwen3.5 Plus 2026-02-15 saasQwen: Qwen3.5 Plus 2026-02-15: The Qwen3.5 native vision-language series Plus models are built on a hybrid architecture that integrates linear attention mechanisms with sparse mixture-of-experts models, achieving higher inference efficiency. In a variety of... | 3 | 29.4 | +7.00 | ||
| #242 | Google: Nano Banana (Gemini 2.5 Flash Image) saasGoogle: Nano Banana (Gemini 2.5 Flash Image): Gemini 2.5 Flash Image, a.k.a. "Nano Banana," is now generally available. It is a state of the art image generation model with contextual understanding. It is capable of image generation,... | 28 | 29.4 | +8.24 | ||
| #247 | Qwen: Qwen2.5 VL 72B Instruct saasQwen: Qwen2.5 VL 72B Instruct: Qwen2.5-VL is proficient in recognizing common objects such as flowers, birds, fish, and insects. It is also highly capable of analyzing texts, charts, icons, graphics, and layouts within images. | 10 | 29.1 | +6.65 | ||
| #252 | Qwen: Qwen3.5 Plus 2026-04-20 saasQwen: Qwen3.5 Plus 2026-04-20: Qwen3.5 Plus (April 2026) is a large-scale multimodal language model from Alibaba. It accepts text, image, and video input and produces text output, with a 1M token context window. This... | 5 | 28.8 | +6.82 | ||
| #254 | MiniMax: MiniMax-01 saasMiniMax: MiniMax-01: MiniMax-01 is a combines MiniMax-Text-01 for text generation and MiniMax-VL-01 for image understanding. It has 456 billion parameters, with 45.9 billion parameters activated per inference, and can handle a context... | 41 | 28.6 | +5.38 | ||
| #186 | model:Jackrong/Qwopus3.6-35B-A3B-v1-GGUF model:Jackrong/Qwopus3.6-35B-A3B-v1-GGUF: discovered AI agent. | 10 | 28.5 | -0.03 | ||
| #256 | Baidu: ERNIE 4.5 VL 424B A47B saasBaidu: ERNIE 4.5 VL 424B A47B : ERNIE-4.5-VL-424B-A47B is a multimodal Mixture-of-Experts (MoE) model from Baidu???s ERNIE 4.5 series, featuring 424B total parameters with 47B active per token. It is trained jointly on text and image data... | 22 | 28.5 | +5.90 | ||
| #257 | Qwen: Qwen3 VL 30B A3B Thinking saasQwen: Qwen3 VL 30B A3B Thinking: Qwen3-VL-30B-A3B-Thinking is a multimodal model that unifies strong text generation with visual understanding for images and videos. Its Thinking variant enhances reasoning in STEM, math, and complex tasks. It excels... | 9 | 28.4 | +7.10 | ||
| #188 | model:HauhauCS/Gemma4-26B-A4B-Uncensored-HauhauCS-Balanced model:HauhauCS/Gemma4-26B-A4B-Uncensored-HauhauCS-Balanced: discovered AI agent. | 9 | 28.4 | -0.01 | ||
| #263 | Google: Nano Banana 2 (Gemini 3.1 Flash Image Preview) saasGoogle: Nano Banana 2 (Gemini 3.1 Flash Image Preview): Gemini 3.1 Flash Image Preview, a.k.a. "Nano Banana 2," is Google???s latest state of the art image generation and editing model, delivering Pro-level visual quality at Flash speed. It combines... | 27 | 28.1 | +7.70 | ||
| #267 | Google: Nano Banana 2 (Gemini 3.1 Flash Image) saasGoogle: Nano Banana 2 (Gemini 3.1 Flash Image): Gemini 3.1 Flash Image, a.k.a. "Nano Banana 2," is Google’s latest state of the art image generation and editing model, delivering Pro-level visual quality at Flash speed. It combines advanced... | 22 | 27.7 | +7.34 | ||
| #276 | Perceptron: Perceptron Mk1 saasPerceptron: Perceptron Mk1: Perceptron Mk1 (Mark One) is Perceptron's highest-quality vision-language model for video and embodied reasoning.** It accepts image and video inputs paired with natural language queries, and produces detailed visual understanding... | 45 | 27.2 | +4.58 | ||
| #280 | Qwen: Qwen3.8 Max (0902) saasQwen: Qwen3.8 Max (0902): Qwen3.8 Max 0902 is an updated snapshot of Qwen3.8 Max from Alibaba's Qwen team. It is a 2.4-trillion-parameter mixture-of-experts model that accepts text, image, and video input and returns text,... | 62 | 27.1 | +10.07 | ||
| #214 | model:HauhauCS/Qwen3.6-27B-Uncensored-HauhauCS-Balanced model:HauhauCS/Qwen3.6-27B-Uncensored-HauhauCS-Balanced: discovered AI agent. | 139 | 26.9 | +5.72 | ||
| #284 | Qwen: Qwen3 VL 235B A22B Thinking saasQwen: Qwen3 VL 235B A22B Thinking: Qwen3-VL-235B-A22B Thinking is a multimodal model that unifies strong text generation with visual understanding across images and video. The Thinking model is optimized for multimodal reasoning in STEM and math.... | 17 | 26.9 | +7.40 | ||
| #290 | OpenAI: GPT-5 Image Mini saasOpenAI: GPT-5 Image Mini: GPT-5 Image Mini combines OpenAI's advanced language capabilities, powered by [GPT-5 Mini](https://openrouter.ai/openai/gpt-5-mini), with GPT Image 1 Mini for efficient image generation. This natively multimodal model features superior instruction following, text... | 17 | 26.5 | +7.06 | ||
| #226 | ian-xiaohei-illustrations mitian-xiaohei-illustrations: 中文小黑怪诞正文配图生成 Skill | 16:9 白底手绘 | 少量红橙蓝批注 | Codex Skill. | 36 | 26.1 | -1.63 | ||
| #232 | model:Jackrong/Qwopus3.6-27B-v2-GGUF model:Jackrong/Qwopus3.6-27B-v2-GGUF: discovered AI agent. | 106 | 25.9 | +4.25 | ||
| #236 | model:cyjin-yl/Qwen3.8-27B-Uncensored-Cyber-agentic-imatrix-GGUF model:cyjin-yl/Qwen3.8-27B-Uncensored-Cyber-agentic-imatrix-GGUF: discovered AI agent. | 8 | 25.7 | +0.14 | ||
| #296 | Xiaomi: MiMo-V2-Omni saasXiaomi: MiMo-V2-Omni: MiMo-V2-Omni is a frontier omni-modal model that natively processes image, video, and audio inputs within a unified architecture. It combines strong multimodal perception with agentic capability - visual grounding, multi-step... | 234 | 25.4 | -14.03 | ||
| #300 | OpenAI: GPT-4o (2024-05-13) saasOpenAI: GPT-4o (2024-05-13): GPT-4o ("o" for "omni") is OpenAI's latest AI model, supporting both text and image inputs with text outputs. It maintains the intelligence level of [GPT-4 Turbo](/models/openai/gpt-4-turbo) while being twice as... | 42 | 25.0 | +3.36 | ||
| #259 | model:MiniMaxAI/MiniMax-M3-MXFP8 ide-pluginmodel:MiniMaxAI/MiniMax-M3-MXFP8: discovered AI agent. | 274 | 24.4 | +6.84 | ||
| #288 | model:lordx64/Qwable-v1 model:lordx64/Qwable-v1: discovered AI agent. | 395 | 23.5 | +8.39 | ||
| #315 | OpenAI: GPT-4 Turbo saasOpenAI: GPT-4 Turbo: The latest GPT-4 Turbo model with vision capabilities. Vision requests can now use JSON mode and function calling.
Training data: up to December 2023. | 39 | 22.6 | +6.83 | ||
| #328 | model:empero-ai/Qwythos-9B-Claude-Mythos-5-1M-GGUF model:empero-ai/Qwythos-9B-Claude-Mythos-5-1M-GGUF: discovered AI agent. | 147 | 22.4 | -6.00 | ||
| #332 | division-sh/swarm division-sh/swarm: discovered AI agent. | 450 | 22.3 | +8.87 | ||
| #359 | model:Jackrong/Qwopus3.6-27B-v2 model:Jackrong/Qwopus3.6-27B-v2: discovered AI agent. | 19 | 21.6 | 0.00 | ||
| #323 | Google: Nano Banana Pro (Gemini 3 Pro Image) saasGoogle: Nano Banana Pro (Gemini 3 Pro Image): Nano Banana Pro is Google’s most advanced image-generation and editing model, built on Gemini 3 Pro. It extends the original Nano Banana with significantly improved multimodal reasoning, real-world grounding, and... | 49 | 21.6 | +7.16 | ||
| #331 | Google: Nano Banana Pro (Gemini 3 Pro Image Preview) saasGoogle: Nano Banana Pro (Gemini 3 Pro Image Preview): Nano Banana Pro is Google???s most advanced image-generation and editing model, built on Gemini 3 Pro. It extends the original Nano Banana with significantly improved multimodal reasoning, real-world grounding, and... | 42 | 20.3 | +5.87 | ||
| #332 | OpenAI: GPT-5 Image saasOpenAI: GPT-5 Image: [GPT-5](https://openrouter.ai/openai/gpt-5) Image combines OpenAI's GPT-5 model with state-of-the-art image generation capabilities. It offers major improvements in reasoning, code quality, and user experience while incorporating GPT Image 1's superior instruction following,... | 55 | 20.2 | +7.63 | ||
| #337 | Qwen: Qwen3.8 Max Prime saasQwen: Qwen3.8 Max Prime: Qwen3.8 Max Prime is a higher-throughput variant of Qwen3.8 Max from Alibaba's Qwen team, served as a separate SKU at a higher price point. It accepts text, image, and video... | 44 | 19.6 | +5.90 | ||
| #466 | illo-skill mitide-pluginillo-skill: illo — an AI agent skill that turns ideas and articles into original print-style editorial illustrations, starring a recurring mascot. Ten bundled looks, custom characters, OpenRouter-powered. | 16 | 18.8 | 0.00 | ||
| #519 | Hent-ai mitHent-ai: Emotion Image Attachment Plugin for AI agents — Auto-classify emotions via LLM and attach matching images to Discord messages. | 8 | 18.0 | 0.00 | ||
| #350 | Google: Gemini 3.1 Flash Lite (batch) saasGoogle: Gemini 3.1 Flash Lite (batch): Gemini 3.1 Flash Lite is Google’s GA high-efficiency multimodal model optimized for low-latency, high-volume workloads. It supports text, image, video, audio, and PDF inputs, and is designed for lightweight agentic... | 153 | 18.0 | -6.00 | ||
| #352 | OpenAI: GPT-5.4 Image 2 saasOpenAI: GPT-5.4 Image 2: [GPT-5.4](https://openrouter.ai/openai/gpt-5.4) Image 2 combines OpenAI's GPT-5.4 model with state-of-the-art image generation capabilities from GPT Image 2. It enables rich multimodal workflows, allowing users to seamlessly move between reasoning, coding, and... | 37 | 17.7 | +5.76 | ||
| #561 | ARIS-in-AI-Offer mitide-pluginARIS-in-AI-Offer: Bilingual ML / LLM / multimodal / diffusion / agent / generative-model interview cheat sheets (秋招经验手册) — single-file HTML reads anywhere on phone, iPad, and laptop — auto-generated by the ARIS /render-html workflow 🌱. | 102 | 17.2 | -1.45 | ||
| #608 | Oaklight/zerodep ide-pluginOaklight/zerodep: Third-party Python libraries introduce dependency management overhead, supply chain risk, and deployment friction in constrained environments. A natural question is how much of this ecosystem can be replicated using only Python's standard library -- and at what correctness and... | 7 | 16.4 | -0.07 | ||
| #380 | SpaceXAI: Grok 4.3 (batch) saasSpaceXAI: Grok 4.3 (batch): Grok 4.3 is a reasoning model from SpaceXAI. It accepts text and image inputs with text output, and is suited for agentic workflows, instruction-following tasks, and applications requiring high factual... | 104 | 14.9 | -6.00 | ||
| #385 | NVIDIA: Nemotron 3 Nano Omni (free) saasNVIDIA: Nemotron 3 Nano Omni (free): NVIDIA Nemotron??? 3 Nano Omni is a 30B-A3B open multimodal model designed to function as a perception and context sub-agent in enterprise agent systems. It accepts text, image, video, and... | 108 | 14.5 | +8.51 | ||
| #728 | qwen-image-2.1-skill apache-2.0qwen-image-2.1-skill: 🎨 Agentic skill for Qwen-Image-2.1: Rewrites and optimizes text-to-image and multi-image editing prompts using official Alibaba specifications. Compatible with skills.sh and all AI agents. | 101 | 14.5 | -1.58 | ||
| #749 | awesome-images awesome-images: Try flatkey.ai for 40% saving! Generate practical images ready for all your work needs!!! | 109 | 14.3 | -1.63 | ||
| #762 | model:byteshape/Qwen3.8-27B-GGUF model:byteshape/Qwen3.8-27B-GGUF: discovered AI agent. | 368 | 14.0 | -5.96 | ||
| #393 | Mistral: Mistral Medium 3.5 (batch) saasMistral: Mistral Medium 3.5 (batch): Mistral Medium 3.5 is a dense 128B instruction-following model from Mistral AI. It supports text and image inputs with text output, and is designed for agentic workflows, coding, and complex... | 87 | 13.4 | -6.00 | ||
| #418 | Google: Gemma 4 31B (free) saasGoogle: Gemma 4 31B (free): Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. Features a 256K token context window, configurable thinking/reasoning mode, native function... | 27 | 11.5 | -0.07 | ||
| #1127 | Hyu-Zhang/MARS ide-pluginHyu-Zhang/MARS: This report presents MARS, short for Multimodal Agentic Reasoning with Source selection, our system for the CASTLE Challenge at EgoVis 2026. Participants must answer 185 closed-form questions over the CASTLE 2024 dataset. In contrast to prior single-video egocentric benchmarks... | 26 | 8.6 | 0.00 |
Browse all sectors at /sectors.