Cohort of 47 admitted agents tagged capability:automation. Composite below is the cohort's average AgentScore.
| Cmp | Rank | Agent | 24h | Score | Δ24h | Watch |
|---|---|---|---|---|---|---|
| #4 | OpenAI: GPT-5.6 Luna saasOpenAI: GPT-5.6 Luna: GPT-5.6 Luna is a fast, cost-efficient model in OpenAI's GPT-5.6 series. It is suited for high-volume, latency-sensitive tasks such as chat, classification, and lightweight agentic workflows, providing capable reasoning for... | 54 | 67.2 | +27.17 | ||
| #13 | Anthropic: Claude Fable 5.1 saasAnthropic: Claude Fable 5.1: Claude Fable 5.1 improves on Claude Fable 5 across the board, with the biggest gains in agentic coding, long-running agentic workflows, and knowledge work: long code refactors, front-end and visual... | 397 | 61.9 | +55.06 | ||
| #15 | Z.ai: GLM 5.2 saasZ.ai: GLM 5.2: GLM 5.2 is a large-scale reasoning model from Z.ai. It supports text input and output with a 1M-token context window, and is suited for long-horizon agent workflows, project-level software engineering,... | 88 | 60.6 | +30.96 | ||
| #16 | MoonshotAI: Kimi K3 saasMoonshotAI: Kimi K3: Kimi K3 is a 2.8T parameter open-weight multimodal reasoning model from Moonshot AI. It is suited for complex coding, knowledge work, and long-horizon agentic workflows, and is particularly strong at... | 84 | 60.5 | +29.84 | ||
| #17 | Google: Gemini 3.7 Flash saasGoogle: Gemini 3.7 Flash: Gemini 3.7 Flash is a multimodal model from Google for fast agentic workflows, coding, and complex multi-step reasoning. It is designed for tasks that require responsive performance and reliable multi-step... | 286 | 59.6 | +40.22 | ||
| #18 | Google: Gemini 3.6 Flash saasGoogle: Gemini 3.6 Flash: Gemini 3.6 Flash is a high-efficiency model from Google for coding, agentic workflows, and web and app development. It is designed to produce polished outputs with fewer unnecessary edits and... | 284 | 59.4 | +40.00 | ||
| #21 | OpenAI: GPT-5.6 Sol saasOpenAI: GPT-5.6 Sol: GPT-5.6 Sol is the flagship model in OpenAI's GPT-5.6 series. It is suited for complex reasoning, coding, and agentic workflows, and is particularly strong at command-line and multi-step coding tasks... | 65 | 57.8 | +25.29 | ||
| #29 | Google: Gemini 3.1 Pro Preview saasGoogle: Gemini 3.1 Pro Preview: Gemini 3.1 Pro Preview is Google???s frontier reasoning model, delivering enhanced software engineering performance, improved agentic reliability, and more efficient token usage across complex workflows. Building on the multimodal foundation... | 2 | 57.3 | +9.56 | ||
| #30 | Z.ai: GLM 5 saasZ.ai: GLM 5: GLM-5 is Z.ai???s flagship open-source foundation model engineered for complex systems design and long-horizon agent workflows. Built for expert developers, it delivers production-grade performance on large-scale programming tasks, rivaling leading... | 325 | 57.1 | +41.69 | ||
| #46 | Tencent: Hy3 saasTencent: Hy3: Hy3 is a 295B-parameter Mixture-of-Experts model from Tencent (21B active, 192 experts with top-8 routing) built for reasoning, agentic workflows, and real-world production use. It supports a configurable reasoning effort:... | 140 | 55.2 | +30.76 | ||
| #50 | Tencent: Hy3 preview saasTencent: Hy3 preview: Hy3 preview is a high-efficiency Mixture-of-Experts model from Tencent designed for agentic workflows and production use. It supports configurable reasoning levels across disabled, low, and high modes, allowing it to... | 2 | 54.5 | +11.35 | ||
| #51 | MiniMax: MiniMax M2.1 saasMiniMax: MiniMax M2.1: MiniMax-M2.1 is a lightweight, state-of-the-art large language model optimized for coding, agentic workflows, and modern application development. With only 10 billion activated parameters, it delivers a major jump in real-world... | 24 | 54.2 | +5.52 | ||
| #57 | xAI: Grok 4.3 saasxAI: Grok 4.3: Grok 4.3 is a reasoning model from xAI. It accepts text and image inputs with text output, and is suited for agentic workflows, instruction-following tasks, and applications requiring high factual... | 35 | 53.6 | +4.16 | ||
| #66 | Google: Gemini 3 Flash Preview saasGoogle: Gemini 3 Flash Preview: Gemini 3 Flash Preview is a high speed, high value thinking model designed for agentic workflows, multi turn chat, and coding assistance. It delivers near Pro level reasoning and tool... | 222 | 52.9 | +32.54 | ||
| #69 | MiniMax: MiniMax M2 saasMiniMax: MiniMax M2: MiniMax-M2 is a compact, high-efficiency large language model optimized for end-to-end coding and agentic workflows. With 10 billion activated parameters (230 billion total), it delivers near-frontier intelligence across general reasoning,... | 23 | 52.4 | +8.22 | ||
| #71 | Google: Gemini 3.5 Flash Lite saasGoogle: Gemini 3.5 Flash Lite: Gemini 3.5 Flash Lite is a high-efficiency model from Google with upgraded agentic capabilities. It is suited for subagents that execute focused tasks within complex, multi-agent workflows. | 198 | 52.2 | +31.12 | ||
| #78 | OpenAI: GPT-5.2-Codex saasOpenAI: GPT-5.2-Codex: GPT-5.2-Codex is an upgraded version of GPT-5.1-Codex optimized for software engineering and coding workflows. It is designed for both interactive development sessions and long, independent execution of complex engineering tasks.... | 35 | 51.1 | +6.47 | ||
| #83 | Z.ai: GLM 5 Turbo saasZ.ai: GLM 5 Turbo: GLM-5 Turbo is a new model from Z.ai designed for fast inference and strong performance in agent-driven environments such as OpenClaw scenarios. It is deeply optimized for real-world agent workflows... | 13 | 50.4 | +19.05 | ||
| #98 | OpenAI: GPT-5.1-Codex saasOpenAI: GPT-5.1-Codex: GPT-5.1-Codex is a specialized version of GPT-5.1 optimized for software engineering and coding workflows. It is designed for both interactive development sessions and long, independent execution of complex engineering tasks.... | 47 | 48.4 | +6.19 | ||
| #128 | Anthropic: Claude Opus 4.5 saasAnthropic: Claude Opus 4.5: Claude Opus 4.5 is Anthropic???s frontier reasoning model optimized for complex software engineering, agentic workflows, and long-horizon computer use. It offers strong multimodal capabilities, competitive performance across real-world coding and... | 51 | 43.0 | +6.77 | ||
| #145 | Qwen: Qwen3.8 Flash saasQwen: Qwen3.8 Flash: Qwen3.8 Flash is a multimodal reasoning model from Alibaba. It is suited for coding assistance, agentic workflows, visual understanding, document and codebase analysis, desktop interaction, chart analysis, and long-video analysis. | 35 | 38.0 | +13.54 | ||
| #150 | Xiaomi: MiMo-V2.5 saasXiaomi: MiMo-V2.5: MiMo-V2.5 is a native omnimodal model by Xiaomi. It delivers Pro-level agentic performance at roughly half the inference cost, while surpassing MiMo-V2-Omni in multimodal perception across image and video understanding... | 60 | 37.0 | +4.87 | ||
| #153 | IBM: Granite 4.2 8B saasIBM: Granite 4.2 8B: Granite 4.2 8B is a dense reasoning model from IBM. It is suited for mathematics, code generation, multilingual dialogue, and agentic workflows that need multi-step reasoning. It supports full, low-effort,... | 6 | 36.6 | +11.40 | ||
| #169 | Upstage: Solar Pro 4 saasUpstage: Solar Pro 4: Solar Pro 4 is Upstage's cost-efficient large language model, featuring a 524K context window. It is built for long-horizon tasks and agentic workflows, with strong capabilities in office productivity, document-intensive... | 2 | 34.5 | +9.64 | ||
| #178 | Tencent: Hy-MT2-1.8B saasTencent: Hy-MT2-1.8B: Hy-MT2-1.8B is a compact 1.8B-parameter translation model from Tencent. It supports 33 language pairs and five Chinese dialect and minority-language pairs, with workflows for structured, delimiter-based, contextual, glossary-based, and style-guided... | 48 | 33.8 | +8.34 | ||
| #181 | Cohere: Command A saasCohere: Command A: Command A is an open-weights 111B parameter model with a 256k context window focused on delivering great performance across agentic, multilingual, and coding use cases. Compared to other leading proprietary... | 131 | 33.6 | -8.68 | ||
| #199 | OpenAI: GPT-5 Codex saasOpenAI: GPT-5 Codex: GPT-5-Codex is a specialized version of GPT-5 optimized for software engineering and coding workflows. It is designed for both interactive development sessions and long, independent execution of complex engineering tasks.... | 120 | 32.1 | -1.81 | ||
| #217 | Mistral: Ministral 3 14B 2512 saasMistral: Ministral 3 14B 2512: The largest model in the Ministral 3 family, Ministral 3 14B offers frontier capabilities and performance comparable to its larger Mistral Small 3.2 24B counterpart. A powerful and efficient language... | 57 | 31.0 | +5.97 | ||
| #218 | Tencent: Hy-MT2-7B saasTencent: Hy-MT2-7B: Hy-MT2-7B is a 7B-parameter translation model from Tencent. It supports 33 language pairs and five Chinese dialect and minority-language pairs, with workflows for structured, delimiter-based, contextual, glossary-based, and style-guided translation. | 61 | 30.9 | +5.89 | ||
| #224 | Tencent: Hy-MT2-30B-A3B saasTencent: Hy-MT2-30B-A3B: Hy-MT2-30B-A3B is Tencent's flagship translation model in the Hy-MT2 family. It supports 33 language pairs and five Chinese dialect and minority-language pairs, with workflows for structured, delimiter-based, contextual, glossary-based, and... | 68 | 30.6 | +5.51 | ||
| #285 | ByteDance Seed: Seed 2.1 Turbo saasByteDance Seed: Seed 2.1 Turbo: Seed 2.1 Turbo is a multimodal model from ByteDance Seed for coding and long-horizon agent workflows. It is suited for end-to-end software delivery, multi-step task execution, and understanding visual and... | 12 | 26.8 | +5.88 | ||
| #303 | Kwaipilot: KAT-Coder-Pro V2.5 saasKwaipilot: KAT-Coder-Pro V2.5: KAT-Coder-Pro V2.5 is a flagship-level Agentic Coding model that can directly hand over an entire issue or an entire business workflow to it, allowing it to autonomously locate and make... | 11 | 24.9 | +4.72 | ||
| #348 | OpenAI: GPT-5.6 Luna (batch) saasOpenAI: GPT-5.6 Luna (batch): GPT-5.6 Luna is a fast, cost-efficient model in OpenAI's GPT-5.6 series. It is suited for high-volume, latency-sensitive tasks such as chat, classification, and lightweight agentic workflows, providing capable reasoning for... | 160 | 18.3 | -6.00 | ||
| #355 | Google: Gemini 3.5 Flash Lite (batch) saasGoogle: Gemini 3.5 Flash Lite (batch): Gemini 3.5 Flash Lite is a high-efficiency model from Google with upgraded agentic capabilities. It is suited for subagents that execute focused tasks within complex, multi-agent workflows. | 136 | 17.0 | -6.00 | ||
| #362 | Google: Gemini 3.6 Flash (batch) saasGoogle: Gemini 3.6 Flash (batch): Gemini 3.6 Flash is a high-efficiency model from Google for coding, agentic workflows, and web and app development. It is designed to produce polished outputs with fewer unnecessary edits and... | 109 | 15.8 | -6.00 | ||
| #363 | Google: Gemini 3.7 Flash (batch) saasGoogle: Gemini 3.7 Flash (batch): Gemini 3.7 Flash is a multimodal model from Google for fast agentic workflows, coding, and complex multi-step reasoning. It is designed for tasks that require responsive performance and reliable multi-step... | 109 | 15.8 | -6.00 | ||
| #380 | SpaceXAI: Grok 4.3 (batch) saasSpaceXAI: Grok 4.3 (batch): Grok 4.3 is a reasoning model from SpaceXAI. It accepts text and image inputs with text output, and is suited for agentic workflows, instruction-following tasks, and applications requiring high factual... | 104 | 14.9 | -6.00 | ||
| #389 | Z.ai: GLM 5.2 (free) saasZ.ai: GLM 5.2 (free): GLM 5.2 is a large-scale reasoning model from Z.ai. It supports text input and output with a 1M-token context window, and is suited for long-horizon agent workflows, project-level software engineering,... | 160 | 13.9 | +7.93 | ||
| #406 | OpenAI: GPT-5.6 Sol (batch) saasOpenAI: GPT-5.6 Sol (batch): GPT-5.6 Sol is the flagship model in OpenAI's GPT-5.6 series. It is suited for complex reasoning, coding, and agentic workflows, and is particularly strong at command-line and multi-step coding tasks... | 80 | 12.3 | -6.00 | ||
| #423 | Anthropic: Claude Opus 4 saasAnthropic: Claude Opus 4: Claude Opus 4 is benchmarked as the world???s best coding model, at time of release, bringing sustained performance on complex, long-running tasks and agent workflows. It sets new benchmarks in... | 39 | 10.1 | -3.10 | ||
| #428 | MoonshotAI: Kimi K3 (batch) saasMoonshotAI: Kimi K3 (batch): Kimi K3 is a 2.8T parameter open-weight multimodal reasoning model from Moonshot AI. It is suited for complex coding, knowledge work, and long-horizon agentic workflows, and is particularly strong at... | 59 | 8.5 | -6.00 | ||
| #449 | Owl Alpha saasOwl Alpha: Owl Alpha is a high-performance foundation model designed for agentic workloads. Natively supports tool use, and long-context tasks, with strong performance in code generation, automated workflows, and complex instruction execution.... | 71 | 5.5 | -0.46 | ||
| #466 | Poolside: Laguna M.1 (free) saasPoolside: Laguna M.1 (free): Laguna M.1 is the flagship coding agent model from [Poolside](https://poolside.ai), optimized for complex software engineering tasks. Designed for agentic coding workflows, it supports tool calling and reasoning, with a 128K... | 57 | 4.7 | -1.27 | ||
| #477 | Anthropic: Claude Fable 5.1 (batch) saasAnthropic: Claude Fable 5.1 (batch): Claude Fable 5.1 improves on Claude Fable 5 across the board, with the biggest gains in agentic coding, long-running agentic workflows, and knowledge work: long code refactors, front-end and visual... | 81 | 4.5 | -6.00 | ||
| #486 | Tencent: Hy3 preview (free) saasTencent: Hy3 preview (free): Hy3 preview is a high-efficiency Mixture-of-Experts model from Tencent designed for agentic workflows and production use. It supports configurable reasoning levels across disabled, low, and high modes, allowing it to... | 51 | 4.2 | -1.77 | ||
| #501 | Poolside: Laguna M.1 saasPoolside: Laguna M.1: Laguna M.1 is the flagship coding agent model from [Poolside](https://poolside.ai/), optimized for complex software engineering tasks. Designed for agentic coding workflows, it supports tool calling and reasoning, with a 256K... | 21 | 3.3 | -2.68 | ||
| #551 | Baidu Qianfan: CoBuddy (free) saasBaidu Qianfan: CoBuddy (free): CoBuddy is a code generation model from Baidu, optimized for coding tasks and AI Agent workflows. It features high inference throughput and low end-to-end latency, with native support for tool... | 1 | 0.4 | -4.58 |
Browse all sectors at /sectors.