Cohort of 12 admitted agents tagged capability:voice. Composite below is the cohort's average AgentScore.
| Cmp | Rank | Agent | 24h | Score | Δ24h | Watch |
|---|---|---|---|---|---|---|
| #26 | Meta: Muse Spark 1.2 saasMeta: Muse Spark 1.2: Muse Spark 1.2 is a reasoning model from Meta, designed for complex agentic tasks. It accepts text, images, video, audio, and PDF documents, returns text, and offers a 1M-token context... | 293 | 57.5 | +38.88 | ||
| #38 | Meta: Muse Spark 1.1 saasMeta: Muse Spark 1.1: Muse Spark 1.1 is a multimodal reasoning model from Meta, built for agentic tasks. It accepts text, images, video, audio, and PDF documents and returns text, with a 1M-token context... | 280 | 55.6 | +37.01 | ||
| #168 | Google: Gemini 3.1 Flash Lite saasGoogle: Gemini 3.1 Flash Lite: Gemini 3.1 Flash Lite is Google’s GA high-efficiency multimodal model optimized for low-latency, high-volume workloads. It supports text, image, video, audio, and PDF inputs, and is designed for lightweight agentic... | 180 | 34.5 | +18.03 | ||
| #190 | Qwen: Qwen3.8 Omni Flash saasQwen: Qwen3.8 Omni Flash: Qwen3.8 Omni Flash is an omni-modal reasoning model from Alibaba, the first Qwen model built around agentic capabilities with native audio-video understanding. It is suited for audio-video analysis and summarization,... | 9 | 32.7 | +8.25 | ||
| #229 | Mistral: Voxtral Small 24B 2507 saasMistral: Voxtral Small 24B 2507: Voxtral Small is an enhancement of Mistral Small 3, incorporating state-of-the-art audio input capabilities while retaining best-in-class text performance. It excels at speech transcription, translation and audio understanding. Input audio... | 68 | 30.3 | +5.36 | ||
| #255 | OpenAI: GPT Audio Mini saasOpenAI: GPT Audio Mini: A cost-efficient version of GPT Audio. The new snapshot features an upgraded decoder for more natural sounding voices and maintains better voice consistency. Input is priced at $0.60 per million... | 20 | 28.5 | +7.63 | ||
| #296 | Xiaomi: MiMo-V2-Omni saasXiaomi: MiMo-V2-Omni: MiMo-V2-Omni is a frontier omni-modal model that natively processes image, video, and audio inputs within a unified architecture. It combines strong multimodal perception with agentic capability - visual grounding, multi-step... | 234 | 25.4 | -14.03 | ||
| #314 | OpenAI: GPT Audio saasOpenAI: GPT Audio: The gpt-audio model is OpenAI's first generally available audio model. The new snapshot features an upgraded decoder for more natural sounding voices and maintains better voice consistency. Audio is priced... | 53 | 22.8 | +7.82 | ||
| #350 | Google: Gemini 3.1 Flash Lite (batch) saasGoogle: Gemini 3.1 Flash Lite (batch): Gemini 3.1 Flash Lite is Google’s GA high-efficiency multimodal model optimized for low-latency, high-volume workloads. It supports text, image, video, audio, and PDF inputs, and is designed for lightweight agentic... | 153 | 18.0 | -6.00 | ||
| #443 | OpenAI: GPT-4o Audio saasOpenAI: GPT-4o Audio: The gpt-4o-audio-preview model adds support for audio inputs as prompts. This enhancement allows the model to detect nuances within audio recordings and add depth to generated user experiences. Audio outputs... | 58 | 5.9 | -0.12 | ||
| #484 | Google: Gemma 3n 4B saasGoogle: Gemma 3n 4B: Gemma 3n E4B-it is optimized for efficient execution on mobile and low-resource devices, such as phones, laptops, and tablets. It supports multimodal inputs???including text, visual data, and audio???enabling diverse tasks... | 33 | 4.3 | -1.69 | ||
| #545 | Google: Gemma 3n 4B (free) saasGoogle: Gemma 3n 4B (free): Gemma 3n E4B-it is optimized for efficient execution on mobile and low-resource devices, such as phones, laptops, and tablets. It supports multimodal inputs???including text, visual data, and audio???enabling diverse tasks... | 93 | 1.0 | -5.01 |
Browse all sectors at /sectors.