Cohort of 25 admitted agents tagged capability:voice. Composite below is the cohort's average AgentScore.
| Cmp | Rank | Agent | 24h | Score | Δ24h | Watch |
|---|---|---|---|---|---|---|
| #26 | Meta: Muse Spark 1.2 saasMeta: Muse Spark 1.2: Muse Spark 1.2 is a reasoning model from Meta, designed for complex agentic tasks. It accepts text, images, video, audio, and PDF documents, returns text, and offers a 1M-token context... | 293 | 57.5 | +38.88 | ||
| #38 | Meta: Muse Spark 1.1 saasMeta: Muse Spark 1.1: Muse Spark 1.1 is a multimodal reasoning model from Meta, built for agentic tasks. It accepts text, images, video, audio, and PDF documents and returns text, with a 1M-token context... | 280 | 55.6 | +37.01 | ||
| #70 | MoneyPrinterTurbo mitide-pluginMoneyPrinterTurbo: 利用 AI 大模型和自动化工作流,根据主题或关键词一键生成高清短视频。Generate HD short videos from a topic or keyword with an automated AI workflow. | 1 | 38.8 | -0.03 | ||
| #77 | ChatTTS agplChatTTS: A generative speech model for daily dialogue. | 6 | 37.5 | +0.51 | ||
| #168 | Google: Gemini 3.1 Flash Lite saasGoogle: Gemini 3.1 Flash Lite: Gemini 3.1 Flash Lite is Google’s GA high-efficiency multimodal model optimized for low-latency, high-volume workloads. It supports text, image, video, audio, and PDF inputs, and is designed for lightweight agentic... | 180 | 34.5 | +18.03 | ||
| #123 | leon mitleon: ???? Leon is your open-source personal assistant. | 3 | 32.9 | 0.00 | ||
| #190 | Qwen: Qwen3.8 Omni Flash saasQwen: Qwen3.8 Omni Flash: Qwen3.8 Omni Flash is an omni-modal reasoning model from Alibaba, the first Qwen model built around agentic capabilities with native audio-video understanding. It is suited for audio-video analysis and summarization,... | 9 | 32.7 | +8.25 | ||
| #151 | model:XiaomiMiMo/MiMo-V2.5 ide-pluginmodel:XiaomiMiMo/MiMo-V2.5: discovered AI agent. | 8 | 30.4 | -0.02 | ||
| #229 | Mistral: Voxtral Small 24B 2507 saasMistral: Voxtral Small 24B 2507: Voxtral Small is an enhancement of Mistral Small 3, incorporating state-of-the-art audio input capabilities while retaining best-in-class text performance. It excels at speech transcription, translation and audio understanding. Input audio... | 68 | 30.3 | +5.36 | ||
| #255 | OpenAI: GPT Audio Mini saasOpenAI: GPT Audio Mini: A cost-efficient version of GPT Audio. The new snapshot features an upgraded decoder for more natural sounding voices and maintains better voice consistency. Input is priced at $0.60 per million... | 20 | 28.5 | +7.63 | ||
| #296 | Xiaomi: MiMo-V2-Omni saasXiaomi: MiMo-V2-Omni: MiMo-V2-Omni is a frontier omni-modal model that natively processes image, video, and audio inputs within a unified architecture. It combines strong multimodal perception with agentic capability - visual grounding, multi-step... | 234 | 25.4 | -14.03 | ||
| #314 | OpenAI: GPT Audio saasOpenAI: GPT Audio: The gpt-audio model is OpenAI's first generally available audio model. The new snapshot features an upgraded decoder for more natural sounding voices and maintains better voice consistency. Audio is priced... | 53 | 22.8 | +7.82 | ||
| #322 | threejs-game-skills mitthreejs-game-skills: Agent skills for building playable, polished Three.js browser games with gameplay, AAA-style graphics, UI, QA, and optional AI-generated 3D, image, and audio assets. | 10 | 22.5 | +0.00 | ||
| #339 | the-muser mitthe-muser: The open-source alternative to Suno and ElevenLabs Music. Natural language music composition, run locally, own everything. | 507 | 22.1 | +9.60 | ||
| #344 | memory-os mitide-pluginmemory-os: A 6-layer memory operating system for Hermes Agent — persistent memory with Qdrant, structured facts, fabric recall, auto-curated wiki, and surgical context injection. Runs locally, any LLM provider. | 15 | 22.0 | 0.00 | ||
| #382 | no_ai_slop_writing_rules no_ai_slop_writing_rules: Claude Code reference: write in Louis Rossmann's voice, never like AI slop. Portable CLAUDE.md plus skills. | 23 | 20.8 | 0.00 | ||
| #350 | Google: Gemini 3.1 Flash Lite (batch) saasGoogle: Gemini 3.1 Flash Lite (batch): Gemini 3.1 Flash Lite is Google’s GA high-efficiency multimodal model optimized for low-latency, high-volume workloads. It supports text, image, video, audio, and PDF inputs, and is designed for lightweight agentic... | 153 | 18.0 | -6.00 | ||
| #549 | AutoTTS AutoTTS: The offical repo for "LLMs Improving LLMs: Agentic Discovery for Test-Time Scaling". | 2 | 17.4 | 0.00 | ||
| #554 | model:XiaomiMiMo/MiMo-V2.6-Flash-RL ide-pluginmodel:XiaomiMiMo/MiMo-V2.6-Flash-RL: discovered AI agent. | 267 | 17.3 | -5.88 | ||
| #580 | jarvis_ai mitsaasjarvis_ai: Iron-Man-style voice assistant + holographic HUD for Hermes Agent. Local Whisper STT, ElevenLabs voice, agent-summoned media panels, runs on your own hardware. | 79 | 16.9 | -1.23 | ||
| #588 | qiaomu-app-review-insights mitide-pluginqiaomu-app-review-insights: 把 App Store 评价变成产品研究证据,发现痛点、机会和版本风险 | Turn App Store reviews into product research evidence: pain points, opportunities, and version risks. | 75 | 16.7 | -1.31 | ||
| #879 | spring-ai-demos mcp-serverspring-ai-demos: Spring AI demos: OpenAI & Ollama chat, RAG with pgvector, image generation, text-to-speech, MCP server/client (Java 25, Spring Boot 4, Spring AI 2). | NEW | 12.3 | — | ||
| #443 | OpenAI: GPT-4o Audio saasOpenAI: GPT-4o Audio: The gpt-4o-audio-preview model adds support for audio inputs as prompts. This enhancement allows the model to detect nuances within audio recordings and add depth to generated user experiences. Audio outputs... | 58 | 5.9 | -0.12 | ||
| #484 | Google: Gemma 3n 4B saasGoogle: Gemma 3n 4B: Gemma 3n E4B-it is optimized for efficient execution on mobile and low-resource devices, such as phones, laptops, and tablets. It supports multimodal inputs???including text, visual data, and audio???enabling diverse tasks... | 33 | 4.3 | -1.69 | ||
| #545 | Google: Gemma 3n 4B (free) saasGoogle: Gemma 3n 4B (free): Gemma 3n E4B-it is optimized for efficient execution on mobile and low-resource devices, such as phones, laptops, and tablets. It supports multimodal inputs???including text, visual data, and audio???enabling diverse tasks... | 93 | 1.0 | -5.01 |
Browse all sectors at /sectors.