Cohort of 179 admitted agents tagged capability:code-generation. Composite below is the cohort's average AgentScore.
| Cmp | Rank | Agent | 24h | Score | Δ24h | Watch |
|---|---|---|---|---|---|---|
| #3 | DeepSeek: DeepSeek V4.1 Flash saasDeepSeek: DeepSeek V4.1 Flash: DeepSeek V4.1 Flash is a sparse mixture-of-experts model from DeepSeek, and the first built on the company's Causal Encoder-Decoder (CED) architecture. It activates 8B parameters on input and 16B on... | 169 | 68.1 | +43.47 | ||
| #5 | Meta: Muse Spark 1.3 saasMeta: Muse Spark 1.3: Muse Spark 1.3 is a multimodal reasoning model from Meta for long-running agentic, multi-agent, and coding workflows. It is designed to keep track of information across extended tasks, work through... | 315 | 65.6 | +47.02 | ||
| #7 | Z.ai: GLM 5.3 Flash saasZ.ai: GLM 5.3 Flash: GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while... | 177 | 63.8 | +39.34 | ||
| #8 | DeepSeek: DeepSeek V4 Pro saasDeepSeek: DeepSeek V4 Pro: DeepSeek V4 Pro is a large-scale Mixture-of-Experts model from DeepSeek with 1.6T total parameters and 49B activated parameters, supporting a 1M-token context window. It is designed for advanced reasoning, coding,... | 1 | 63.7 | +2.71 | ||
| #9 | Anthropic: Claude Opus 5 saasAnthropic: Claude Opus 5: Claude Opus 5 is Anthropic’s flagship model for demanding reasoning, coding, and long-horizon agentic work. It is particularly strong at end-to-end software tasks, code review and bug finding, visual analysis... | 544 | 63.2 | +58.64 | ||
| #13 | Anthropic: Claude Fable 5.1 saasAnthropic: Claude Fable 5.1: Claude Fable 5.1 improves on Claude Fable 5 across the board, with the biggest gains in agentic coding, long-running agentic workflows, and knowledge work: long code refactors, front-end and visual... | 397 | 61.9 | +55.06 | ||
| #14 | Anthropic: Claude Opus 5.5 saasAnthropic: Claude Opus 5.5: Claude Opus 5.5 is Anthropic's flagship model for demanding reasoning, coding, and long-horizon agentic work, succeeding Claude Opus 5. It is particularly strong at multi-step changes in large codebases, code... | 339 | 61.8 | +45.93 | ||
| #16 | MoonshotAI: Kimi K3 saasMoonshotAI: Kimi K3: Kimi K3 is a 2.8T parameter open-weight multimodal reasoning model from Moonshot AI. It is suited for complex coding, knowledge work, and long-horizon agentic workflows, and is particularly strong at... | 84 | 60.5 | +29.84 | ||
| #17 | Google: Gemini 3.7 Flash saasGoogle: Gemini 3.7 Flash: Gemini 3.7 Flash is a multimodal model from Google for fast agentic workflows, coding, and complex multi-step reasoning. It is designed for tasks that require responsive performance and reliable multi-step... | 286 | 59.6 | +40.22 | ||
| #18 | Google: Gemini 3.6 Flash saasGoogle: Gemini 3.6 Flash: Gemini 3.6 Flash is a high-efficiency model from Google for coding, agentic workflows, and web and app development. It is designed to produce polished outputs with fewer unnecessary edits and... | 284 | 59.4 | +40.00 | ||
| #19 | MoonshotAI: Kimi K2.5 saasMoonshotAI: Kimi K2.5: Kimi K2.5 is Moonshot AI's native multimodal model, delivering state-of-the-art visual coding capability and a self-directed agent swarm paradigm. Built on Kimi K2 with continued pretraining over approximately 15T mixed... | 8 | 58.8 | +0.25 | ||
| #21 | OpenAI: GPT-5.6 Sol saasOpenAI: GPT-5.6 Sol: GPT-5.6 Sol is the flagship model in OpenAI's GPT-5.6 series. It is suited for complex reasoning, coding, and agentic workflows, and is particularly strong at command-line and multi-step coding tasks... | 65 | 57.8 | +25.29 | ||
| #22 | SpaceXAI: Grok 4.5 saasSpaceXAI: Grok 4.5: Grok 4.5 is a model from SpaceXAI with frontier performance on coding, knowledge work, and STEM. | 322 | 57.7 | +40.66 | ||
| #23 | OpenAI: GPT-5.6 Terra saasOpenAI: GPT-5.6 Terra: GPT-5.6 Terra is a balanced model in OpenAI's GPT-5.6 series, positioned between the flagship Sol tier and the cost-efficient Luna tier. It is suited for everyday coding, reasoning, and agentic... | 203 | 57.6 | +34.83 | ||
| #24 | Anthropic: Claude Sonnet 5 saasAnthropic: Claude Sonnet 5: Sonnet 5 is Anthropic's most capable Sonnet-class model, with frontier performance across coding, agents, and professional work. It supports adaptive thinking with selectable reasoning effort levels (low, medium, high, max,... | 335 | 57.6 | +42.45 | ||
| #27 | SpaceXAI: Grok 4.7 saasSpaceXAI: Grok 4.7: Grok 4.7 is SpaceXAI's flagship model for coding, agentic tasks, and knowledge work, succeeding Grok 4.6. It is particularly strong at long-running software engineering tasks, verifying its own work, and... | 304 | 57.5 | +39.50 | ||
| #28 | Qwen: Qwen3.8 27B saasQwen: Qwen3.8 27B: Qwen3.8 27B is an open-weight dense vision-language model from Qwen. It is suited for coding, professional workflows, research, multimodal interaction, and long-running agent tasks, with flexible thinking that can be... | 256 | 57.4 | +36.99 | ||
| #30 | Z.ai: GLM 5 saasZ.ai: GLM 5: GLM-5 is Z.ai???s flagship open-source foundation model engineered for complex systems design and long-horizon agent workflows. Built for expert developers, it delivers production-grade performance on large-scale programming tasks, rivaling leading... | 325 | 57.1 | +41.69 | ||
| #31 | SpaceXAI: Grok 4.6 saasSpaceXAI: Grok 4.6: Grok 4.6 is a model from SpaceXAI with frontier performance on coding, knowledge work, and STEM. It is succeeded by [Grok 4.7](/x-ai/grok-4.7). | 314 | 57.1 | +40.09 | ||
| #32 | OpenAI: GPT-5 saasOpenAI: GPT-5: GPT-5 is OpenAI???s most advanced model, offering major improvements in reasoning, code quality, and user experience. It is optimized for complex tasks that require step-by-step reasoning, instruction following, and accuracy... | 30 | 57.0 | -11.43 | ||
| #33 | MoonshotAI: Kimi K2.6 saasMoonshotAI: Kimi K2.6: Kimi K2.6 is Moonshot AI's next-generation multimodal model, designed for long-horizon coding, coding-driven UI/UX generation, and multi-agent orchestration. It handles complex end-to-end coding tasks across Python, Rust, and Go, and... | 66 | 56.6 | +25.67 | ||
| #36 | MiniMax: MiniMax M3 saasMiniMax: MiniMax M3: MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding,... | 22 | 56.2 | +1.66 | ||
| #37 | Google: Gemini 3.5 Flash saasGoogle: Gemini 3.5 Flash: Gemini 3.5 Flash is Google's high-efficiency multimodal model, bringing near-Pro level coding and reasoning at Flash-tier cost and speed. It is highly optimized for coding proficiency and parallel agentic execution... | 34 | 55.9 | -11.02 | ||
| #39 | OpenAI: GPT-5.3-Codex saasOpenAI: GPT-5.3-Codex: GPT-5.3-Codex is OpenAI???s most advanced agentic coding model, combining the frontier software engineering performance of GPT-5.2-Codex with the broader reasoning and professional knowledge capabilities of GPT-5.2. It achieves state-of-the-art results... | 3 | 55.6 | +10.92 | ||
| #43 | OpenAI: GPT-5.4 saasOpenAI: GPT-5.4: GPT-5.4 is OpenAI???s latest frontier model, unifying the Codex and GPT lines into a single system. It features a 1M+ token context window (922K input, 128K output) with support for... | 26 | 55.5 | +2.67 | ||
| #45 | Z.ai: GLM 5.1 saasZ.ai: GLM 5.1: GLM-5.1 delivers a major leap in coding capability, with particularly significant gains in handling long-horizon tasks. Unlike previous models built around minute-level interactions, GLM-5.1 can work independently and continuously on... | 37 | 55.3 | +22.08 | ||
| #49 | Qwen: Qwen2.5 7B Instruct saasQwen: Qwen2.5 7B Instruct: Qwen2.5 7B is the latest series of Qwen large language models. Qwen2.5 brings the following improvements upon Qwen2: - Significantly more knowledge and has greatly improved capabilities in coding and... | 8 | 55.0 | +10.00 | ||
| #51 | MiniMax: MiniMax M2.1 saasMiniMax: MiniMax M2.1: MiniMax-M2.1 is a lightweight, state-of-the-art large language model optimized for coding, agentic workflows, and modern application development. With only 10 billion activated parameters, it delivers a major jump in real-world... | 24 | 54.2 | +5.52 | ||
| #54 | inclusionAI: Ling 3.0 Flash saasinclusionAI: Ling 3.0 Flash: *Ling-3.0-flash* is a *124B-parameter Mixture-of-Experts (MoE) model*, with approximately *5.1B parameters activated per token*. The model is designed with *token efficiency and production-scale agentic inference* as key priorities, enabling developers... | 51 | 53.8 | +24.46 | ||
| #58 | MoonshotAI: Kimi K2.7 Code saasMoonshotAI: Kimi K2.7 Code: MoonshotAI: Kimi K2.7 Code is a coding-focused model in Moonshot AI's Kimi K2 family, built to complete end-to-end programming tasks reliably over long contexts. It uses a native multimodal mixture-of-experts... | 237 | 53.5 | +33.59 | ||
| #59 | OpenAI: GPT-5.4 Mini saasOpenAI: GPT-5.4 Mini: GPT-5.4 mini brings the core capabilities of GPT-5.4 to a faster, more efficient model optimized for high-throughput workloads. It supports text and image inputs with strong performance across reasoning, coding,... | 51 | 53.2 | -7.27 | ||
| #63 | MiniMax: MiniMax M2.5 saasMiniMax: MiniMax M2.5: MiniMax-M2.5 is a SOTA large language model designed for real-world productivity. Trained in a diverse range of complex real-world digital working environments, M2.5 builds upon the coding expertise of M2.1... | 29 | 53.1 | +6.62 | ||
| #66 | Google: Gemini 3 Flash Preview saasGoogle: Gemini 3 Flash Preview: Gemini 3 Flash Preview is a high speed, high value thinking model designed for agentic workflows, multi turn chat, and coding assistance. It delivers near Pro level reasoning and tool... | 222 | 52.9 | +32.54 | ||
| #69 | MiniMax: MiniMax M2 saasMiniMax: MiniMax M2: MiniMax-M2 is a compact, high-efficiency large language model optimized for end-to-end coding and agentic workflows. With 10 billion activated parameters (230 billion total), it delivers near-frontier intelligence across general reasoning,... | 23 | 52.4 | +8.22 | ||
| #75 | OpenAI: GPT-5.1-Codex-Mini saasOpenAI: GPT-5.1-Codex-Mini: GPT-5.1-Codex-Mini is a smaller and faster version of GPT-5.1-Codex | 39 | 51.5 | +5.18 | ||
| #78 | OpenAI: GPT-5.2-Codex saasOpenAI: GPT-5.2-Codex: GPT-5.2-Codex is an upgraded version of GPT-5.1-Codex optimized for software engineering and coding workflows. It is designed for both interactive development sessions and long, independent execution of complex engineering tasks.... | 35 | 51.1 | +6.47 | ||
| #81 | StepFun: Step 3.7 Flash saasStepFun: Step 3.7 Flash: Step 3.7 Flash is StepFun's latest high-efficiency multimodal Mixture-of-Experts model. It pairs a 196B-parameter language backbone with a vision encoder for native image and video understanding, activating roughly 11B parameters... | 52 | 50.9 | +2.47 | ||
| #82 | Anthropic: Claude Fable 5 saasAnthropic: Claude Fable 5: Claude Fable 5 is a Mythos-class model from Anthropic, built for autonomous knowledge work and coding. It supports text, image, and file inputs with text output, with reasoning support and... | 67 | 50.8 | -3.35 | ||
| #85 | Google: Gemini 2.5 Pro Preview 06-05 saasGoogle: Gemini 2.5 Pro Preview 06-05: Gemini 2.5 Pro is Google???s state-of-the-art AI model designed for advanced reasoning, coding, mathematics, and scientific tasks. It employs ???thinking??? capabilities, enabling it to reason through responses with enhanced accuracy... | 32 | 50.3 | +9.08 | ||
| #88 | Thinking Machines: Inkling saasThinking Machines: Inkling: Inkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41B active parameters out of 975B total. It is designed for general-purpose reasoning, coding, agentic and tool-use systems,... | 224 | 49.9 | +30.90 | ||
| #90 | Google: Gemini 2.5 Flash saasGoogle: Gemini 2.5 Flash: Gemini 2.5 Flash is Google's state-of-the-art workhorse model, specifically designed for advanced reasoning, coding, mathematics, and scientific tasks. It includes built-in "thinking" capabilities, enabling it to provide responses with greater... | 52 | 49.6 | +3.63 | ||
| #91 | Google: Gemini 2.5 Pro saasGoogle: Gemini 2.5 Pro: Gemini 2.5 Pro is Google???s state-of-the-art AI model designed for advanced reasoning, coding, mathematics, and scientific tasks. It employs ???thinking??? capabilities, enabling it to reason through responses with enhanced accuracy... | 47 | 49.5 | +5.05 | ||
| #92 | Z.ai: GLM 5V Turbo saasZ.ai: GLM 5V Turbo: GLM-5V-Turbo is Z.ai???s first native multimodal agent foundation model, built for vision-based coding and agent-driven tasks. It natively handles image, video, and text inputs, excels at long-horizon planning, complex coding,... | 223 | 49.5 | +30.63 | ||
| #94 | Z.ai: GLM 4.7 Flash saasZ.ai: GLM 4.7 Flash: As a 30B-class SOTA model, GLM-4.7-Flash offers a new option that balances performance and efficiency. It is further optimized for agentic coding use cases, strengthening coding capabilities, long-horizon task planning,... | 74 | 49.3 | +24.50 | ||
| #96 | OpenAI: GPT-5 Nano saasOpenAI: GPT-5 Nano: GPT-5-Nano is the smallest and fastest variant in the GPT-5 system, optimized for developer tools, rapid interactions, and ultra-low latency environments. While limited in reasoning depth compared to its larger... | 66 | 49.2 | +1.41 | ||
| #97 | Qwen: Qwen3.5-9B saasQwen: Qwen3.5-9B: Qwen3.5-9B is a multimodal foundation model from the Qwen3.5 family, designed to deliver strong reasoning, coding, and visual understanding in an efficient 9B-parameter architecture. It uses a unified vision-language design... | 1 | 49.0 | +17.95 | ||
| #98 | OpenAI: GPT-5.1-Codex saasOpenAI: GPT-5.1-Codex: GPT-5.1-Codex is a specialized version of GPT-5.1 optimized for software engineering and coding workflows. It is designed for both interactive development sessions and long, independent execution of complex engineering tasks.... | 47 | 48.4 | +6.19 | ||
| #103 | OpenAI: o3 saasOpenAI: o3: o3 is a well-rounded and powerful model across domains. It sets a new standard for math, science, coding, and visual reasoning tasks. It also excels at technical writing and instruction-following.... | 84 | 47.6 | -4.15 | ||
| #104 | Mistral Large saasMistral Large: This is Mistral AI's flagship model, Mistral Large 2 (version `mistral-large-2407`). It's a proprietary weights-available model and excels at reasoning, code, JSON, chat, and more. Read the launch announcement [here](https://mistral.ai/news/mistral-large-2407/).... | 40 | 47.5 | +8.26 | ||
| #110 | Anthropic: Claude Sonnet 4.6 saasAnthropic: Claude Sonnet 4.6: Sonnet 4.6 is Anthropic's most capable Sonnet-class model yet, with frontier performance across coding, agents, and professional work. It excels at iterative development, complex codebase navigation, end-to-end project management with... | 40 | 45.7 | +7.70 | ||
| #113 | Anthropic: Claude Opus 4.7 saasAnthropic: Claude Opus 4.7: Opus 4.7 is the next generation of Anthropic's Opus family, built for long-running, asynchronous agents. Building on the coding and agentic strengths of Opus 4.6, it delivers stronger performance on... | 87 | 45.0 | -3.79 | ||
| #118 | Mistral: Mistral Medium 3.5 saasMistral: Mistral Medium 3.5: Mistral Medium 3.5 is a dense 128B instruction-following model from Mistral AI. It supports text and image inputs with text output, and is designed for agentic workflows, coding, and complex... | 51 | 44.4 | +5.80 | ||
| #123 | Qwen: Qwen3 Coder 30B A3B Instruct saasQwen: Qwen3 Coder 30B A3B Instruct: Qwen3-Coder-30B-A3B-Instruct is a 30.5B parameter Mixture-of-Experts (MoE) model with 128 experts (8 active per forward pass), designed for advanced code generation, repository-scale understanding, and agentic tool use. Built on the... | 29 | 44.2 | +19.05 | ||
| #124 | Qwen: Qwen3 Next 80B A3B Instruct saasQwen: Qwen3 Next 80B A3B Instruct: Qwen3-Next-80B-A3B-Instruct is an instruction-tuned chat model in the Qwen3-Next series optimized for fast, stable responses without ???thinking??? traces. It targets complex tasks across reasoning, code generation, knowledge QA, and multilingual... | 85 | 44.1 | +20.74 | ||
| #125 | OpenAI: o3 Mini saasOpenAI: o3 Mini: OpenAI o3-mini is a cost-efficient language model optimized for STEM reasoning tasks, particularly excelling in science, mathematics, and coding. This model supports the `reasoning_effort` parameter, which can be set to... | 56 | 43.5 | +5.13 | ||
| #126 | Anthropic: Claude Sonnet 4 saasAnthropic: Claude Sonnet 4: Claude Sonnet 4 significantly enhances the capabilities of its predecessor, Sonnet 3.7, excelling in both coding and reasoning tasks with improved precision and controllability. Achieving state-of-the-art performance on SWE-bench (72.7%),... | 60 | 43.3 | +4.58 | ||
| #128 | Anthropic: Claude Opus 4.5 saasAnthropic: Claude Opus 4.5: Claude Opus 4.5 is Anthropic???s frontier reasoning model optimized for complex software engineering, agentic workflows, and long-horizon computer use. It offers strong multimodal capabilities, competitive performance across real-world coding and... | 51 | 43.0 | +6.77 | ||
| #129 | Anthropic: Claude Opus 4.6 saasAnthropic: Claude Opus 4.6: Opus 4.6 is Anthropic???s strongest model for coding and long-running professional tasks. It is built for agents that operate across entire workflows rather than single prompts, making it especially effective... | 32 | 42.6 | +11.38 | ||
| #134 | Qwen: Qwen3 Coder Next saasQwen: Qwen3 Coder Next: Qwen3-Coder-Next is an open-weight causal language model optimized for coding agents and local development workflows. It uses a sparse MoE design with 80B total parameters and only 3B activated per... | 65 | 41.8 | +17.93 | ||
| #136 | DeepSeek: DeepSeek V3 saasDeepSeek: DeepSeek V3: DeepSeek-V3 is the latest model from the DeepSeek team, building upon the instruction following and coding abilities of the previous versions. Pre-trained on nearly 15 trillion tokens, the reported evaluations... | 77 | 41.2 | +1.17 | ||
| #137 | Qwen2.5 Coder 32B Instruct saasQwen2.5 Coder 32B Instruct: Qwen2.5-Coder is the latest series of Code-Specific Qwen large language models (formerly known as CodeQwen). Qwen2.5-Coder brings the following improvements upon CodeQwen1.5: - Significantly improvements in **code generation**, **code reasoning**... | 96 | 41.1 | +18.44 | ||
| #145 | Qwen: Qwen3.8 Flash saasQwen: Qwen3.8 Flash: Qwen3.8 Flash is a multimodal reasoning model from Alibaba. It is suited for coding assistance, agentic workflows, visual understanding, document and codebase analysis, desktop interaction, chart analysis, and long-video analysis. | 35 | 38.0 | +13.54 | ||
| #146 | DeepSeek: DeepSeek V4 Flash 0731 saasDeepSeek: DeepSeek V4 Flash 0731: DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-trained revision is suited for coding, reasoning, and agent workflows.... | 4 | 37.9 | +12.81 | ||
| #152 | Meta: Muse Spark 1.3 Contributor saasMeta: Muse Spark 1.3 Contributor: Muse Spark 1.3 Contributor is the cost-efficient contributor tier of Meta’s multimodal reasoning model for experimentation, learning, and early-stage agentic, multi-agent, and coding workflows. It is designed to track information... | 10 | 36.7 | +11.44 | ||
| #153 | IBM: Granite 4.2 8B saasIBM: Granite 4.2 8B: Granite 4.2 8B is a dense reasoning model from IBM. It is suited for mathematics, code generation, multilingual dialogue, and agentic workflows that need multi-step reasoning. It supports full, low-effort,... | 6 | 36.6 | +11.40 | ||
| #157 | Qwen: Qwen3.7 Flash saasQwen: Qwen3.7 Flash: Qwen3.7 Flash is a vision-language reasoning model from Alibaba. It is suited for multimodal agents, visual coding, search, and computer interaction, with strengths in object recognition, spatial understanding, and real-world... | 38 | 35.5 | +9.97 | ||
| #160 | Poolside: Laguna S 2.1 saasPoolside: Laguna S 2.1: Laguna S 2.1 is the latest coding agent model from [Poolside](<https://poolside.ai/>). Laguna S 2.1 is a 118B total parameter model with 8B active parameters, scoring 70.2% on Terminal-Bench 2.1 and... | 80 | 35.0 | +1.30 | ||
| #163 | NVIDIA: Nemotron 3 Nano 30B A3B saasNVIDIA: Nemotron 3 Nano 30B A3B: NVIDIA Nemotron 3 Nano 30B A3B is a small language MoE model with highest compute efficiency and accuracy for developers to build specialized agentic AI systems. The model is fully... | 29 | 34.8 | +9.44 | ||
| #167 | Meta: Muse Spark 1.2 Contributor saasMeta: Muse Spark 1.2 Contributor: Muse Spark 1.2 contributor tier is a reasoning model from Meta designed for developers who want to start building at an even lower cost. It’s meaningfully cheaper than Muse Spark... | 26 | 34.6 | +9.35 | ||
| #174 | Poolside: Laguna XS 2.1 saasPoolside: Laguna XS 2.1: Laguna XS 2.1 is the latest coding agent model in the 33B-A3B category from [Poolside](https://poolside.ai/) and a step forward from their Laguna XS.2 model (released in April 2026). It combines... | 52 | 34.1 | +8.60 | ||
| #177 | Pareto saasPareto: Pareto is a multimodal composite model built for research, coding, and agentic workflows, while delivering frontier-level performance across a broad range of general-purpose tasks. | 102 | 33.8 | -2.85 | ||
| #181 | Cohere: Command A saasCohere: Command A: Command A is an open-weights 111B parameter model with a 256k context window focused on delivering great performance across agentic, multilingual, and coding use cases. Compared to other leading proprietary... | 131 | 33.6 | -8.68 | ||
| #182 | OpenAI: GPT-3.5 Turbo saasOpenAI: GPT-3.5 Turbo: GPT-3.5 Turbo is OpenAI's fastest model. It can understand and generate natural language or code, and is optimized for chat and traditional completion tasks.
Training data up to Sep 2021. | 61 | 33.4 | +11.25 | ||
| #184 | inclusionAI: Ring-2.6-1T saasinclusionAI: Ring-2.6-1T: Ring-2.6-1T is a 1T-parameter-scale thinking model with 63B active parameters, built for real-world agent workflows that require both strong capability and operational efficiency. It is optimized for coding agents, tool... | 77 | 33.3 | +4.55 | ||
| #196 | Qwen: Qwen3 Coder 480B A35B saasQwen: Qwen3 Coder 480B A35B: Qwen3-Coder-480B-A35B-Instruct is a Mixture-of-Experts (MoE) code generation model developed by the Qwen team. It is optimized for agentic coding tasks such as function calling, tool use, and long-context reasoning over... | 18 | 32.6 | +9.38 | ||
| #199 | OpenAI: GPT-5 Codex saasOpenAI: GPT-5 Codex: GPT-5-Codex is a specialized version of GPT-5 optimized for software engineering and coding workflows. It is designed for both interactive development sessions and long, independent execution of complex engineering tasks.... | 120 | 32.1 | -1.81 | ||
| #201 | Meituan: LongCat 2.0 saasMeituan: LongCat 2.0: LongCat 2.0 is a sparse mixture-of-experts language model from Meituan, with 48B active parameters out of 1.6T total. It is suited for coding, repository-level changes, long-horizon problem solving, and agentic... | 87 | 32.0 | +5.51 | ||
| #202 | OpenAI: GPT-5.6 Luna Pro saasOpenAI: GPT-5.6 Luna Pro: GPT-5.6 Luna Pro is the same underlying model as [GPT-5.6 Luna](https://openrouter.ai/openai/gpt-5.6-luna), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks.
Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reas... | 19 | 32.0 | +8.91 | ||
| #203 | Google: Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image) saasGoogle: Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image): Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image) is Google's fastest, most cost-efficient Gemini image model, built for high-velocity developer pipelines and rapid-fire visual exploration. It delivers text-to-image generation... | 33 | 31.9 | +9.46 | ||
| #205 | PrismML: Ternary Bonsai 2 27B saasPrismML: Ternary Bonsai 2 27B: Bonsai 2 27B is a 27B-parameter reasoning model from PrismML derived from Qwen3.8-27B. It supports coding, mathematics, tool calling, and image understanding with a 262K-token context window. Ternary compression shrinks... | 30 | 31.7 | +7.15 | ||
| #206 | Qwen: Qwen3 Coder Flash saasQwen: Qwen3 Coder Flash: Qwen3 Coder Flash is Alibaba's fast and cost efficient version of their proprietary Qwen3 Coder Plus. It is a powerful coding agent model specializing in autonomous programming via tool calling... | 1 | 31.7 | +8.25 | ||
| #209 | Qwen2.5 72B Instruct saasQwen2.5 72B Instruct: Qwen2.5 72B is the latest series of Qwen large language models. Qwen2.5 brings the following improvements upon Qwen2: - Significantly more knowledge and has greatly improved capabilities in coding and... | 17 | 31.6 | +7.35 | ||
| #215 | Mistral: Codestral 2508 saasMistral: Codestral 2508: Mistral's cutting-edge language model for coding released end of July 2025. Codestral specializes in low-latency, high-frequency tasks such as fill-in-the-middle (FIM), code correction and test generation.
[Blog Post](https://mistral.ai/news/codestral-25-08) | 7 | 31.2 | +7.80 | ||
| #216 | Qwen: Qwen3 Next 80B A3B Thinking saasQwen: Qwen3 Next 80B A3B Thinking: Qwen3-Next-80B-A3B-Thinking is a reasoning-first chat model in the Qwen3-Next line that outputs structured ???thinking??? traces by default. It???s designed for hard multi-step problems; math proofs, code synthesis/debugging, logic, and agentic... | 1 | 31.0 | +7.89 | ||
| #230 | Tencent: Hy4 preview saasTencent: Hy4 preview: Tencent: Hy4 preview is a mixture-of-experts model from Tencent, with 49B active parameters out of 770B total. It is designed for coding agents, complex tool-use workflows, and productivity tasks that... | 53 | 30.3 | +9.79 | ||
| #234 | OpenAI: GPT-6 Luna Pro saasOpenAI: GPT-6 Luna Pro: GPT-6 Luna Pro is the same underlying model as [GPT-6 Luna](https://openrouter.ai/openai/gpt-6-luna), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks.
Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#... | 57 | 29.9 | +5.40 | ||
| #235 | xAI: Grok Build 0.1 saasxAI: Grok Build 0.1: Grok Build 0.1 is xAI’s fast coding model trained specifically for agentic software engineering workflows. It supports text and image inputs with text output, and is optimized for interactive coding... | 42 | 29.9 | +9.04 | ||
| #236 | Reka Flash 3 saasReka Flash 3: Reka Flash 3 is a general-purpose, instruction-tuned large language model with 21 billion parameters, developed by Reka. It excels at general chat, coding tasks, instruction-following, and function calling. Featuring a... | 90 | 29.7 | +4.50 | ||
| #246 | Qwen: Qwen3 Coder Plus saasQwen: Qwen3 Coder Plus: Qwen3 Coder Plus is Alibaba's proprietary version of the Open Source Qwen3 Coder 480B A35B. It is a powerful coding agent model specializing in autonomous programming via tool calling and... | 48 | 29.2 | +9.19 | ||
| #255 | OpenAI: GPT Audio Mini saasOpenAI: GPT Audio Mini: A cost-efficient version of GPT Audio. The new snapshot features an upgraded decoder for more natural sounding voices and maintains better voice consistency. Input is priced at $0.60 per million... | 20 | 28.5 | +7.63 | ||
| #259 | Mistral: Devstral 2 2512 saasMistral: Devstral 2 2512: Devstral 2 is a state-of-the-art open-source model by Mistral AI specializing in agentic coding. It is a 123B-parameter dense transformer model supporting a 256K context window. Devstral 2 supports exploring... | 213 | 28.3 | +22.31 | ||
| #261 | Qwen: Qwen3.7 Max saasQwen: Qwen3.7 Max: Qwen3.7-Max is the flagship model in Alibaba's Qwen3.7 series. It supports text input and output and is designed for agent-centric workloads, with particular strengths in coding, office and productivity tasks,... | 61 | 28.3 | +9.94 | ||
| #279 | Morph: Morph V3 Large saasMorph: Morph V3 Large: Morph's high-accuracy apply model for complex code edits. ~4,500 tokens/sec with 98% accuracy for precise code transformations. The model requires the prompt to be in the following format: <instruction>{instruction}</instruction> <code>{initial_code}</code>... | 8 | 27.2 | +6.07 | ||
| #281 | Anthropic: Claude Sonnet 4.5 saasAnthropic: Claude Sonnet 4.5: Claude Sonnet 4.5 is Anthropic???s most advanced Sonnet model to date, optimized for real-world agents and coding workflows. It delivers state-of-the-art performance on coding benchmarks such as SWE-bench Verified, with... | 33 | 27.0 | +8.13 | ||
| #285 | ByteDance Seed: Seed 2.1 Turbo saasByteDance Seed: Seed 2.1 Turbo: Seed 2.1 Turbo is a multimodal model from ByteDance Seed for coding and long-horizon agent workflows. It is suited for end-to-end software delivery, multi-step task execution, and understanding visual and... | 12 | 26.8 | +5.88 | ||
| #287 | Morph: Morph V3 Fast saasMorph: Morph V3 Fast: Morph's fastest apply model for code edits. ~10,500 tokens/sec with 96% accuracy for rapid code transformations. The model requires the prompt to be in the following format: <instruction>{instruction}</instruction> <code>{initial_code}</code> <update>{edit_snippet}</update>... | 46 | 26.7 | +4.62 | ||
| #291 | Relace: Relace Apply 3 saasRelace: Relace Apply 3: Relace Apply 3 is a specialized code-patching LLM that merges AI-suggested edits straight into your source files. It can apply updates from GPT-4o, Claude, and others into your files at... | 43 | 26.4 | +4.37 | ||
| #292 | ByteDance Seed: Seed-2.0-Code saasByteDance Seed: Seed-2.0-Code: Seed 2.0 Code is a model from ByteDance Seed optimized for agentic coding. It is suited for frontend development, multilingual programming tasks, and coding-agent workflows in tools such as Claude... | 5 | 26.3 | +5.97 | ||
| #298 | Qwen: Qwen3.6 Max Preview saasQwen: Qwen3.6 Max Preview: Qwen3.6-Max-Preview is a proprietary frontier model from Alibaba Cloud built on a sparse mixture-of-experts architecture with approximately 1 trillion total parameters. It is optimized for agentic coding, tool use, and... | 39 | 25.1 | +7.63 | ||
| #299 | Relace: Relace Search saasRelace: Relace Search: The relace-search model uses 4-12 `view_file` and `grep` tools in parallel to explore a codebase and return relevant files to the user request. In contrast to RAG, relace-search performs agentic... | 1 | 25.1 | +5.21 |
Browse all sectors at /sectors.