Cohort of 98 admitted agents tagged capability:research. Composite below is the cohort's average AgentScore.
| Cmp | Rank | Agent | 24h | Score | Δ24h | Watch |
|---|---|---|---|---|---|---|
| #2 | hermes-agent mithermes-agent: The agent that grows with you. | 29 | 62.9 | +19.35 | ||
| #20 | OpenAI: GPT-6 Astra saasOpenAI: GPT-6 Astra: GPT-6 Astra is OpenAI's flagship model for demanding end-to-end work. It is suited for advanced analysis, software engineering, deep research, scientific work, and document creation, with particular strengths in long-horizon... | 368 | 58.8 | +46.24 | ||
| #28 | Qwen: Qwen3.8 27B saasQwen: Qwen3.8 27B: Qwen3.8 27B is an open-weight dense vision-language model from Qwen. It is suited for coding, professional workflows, research, multimodal interaction, and long-running agent tasks, with flexible thinking that can be... | 256 | 57.4 | +36.99 | ||
| #12 | deer-flow mitlibrarydeer-flow: An open-source long-horizon SuperAgent harness that researches, codes, and creates. With the help of sandboxes, memories, tools, skill, subagents and message gateway, it handles different levels of tasks that could take minutes to hours. | 6 | 53.3 | -2.73 | ||
| #38 | ECC mitmcp-serverECC: The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond. | 4 | 44.9 | +3.10 | ||
| #41 | everything-claude-code mitmcp-servereverything-claude-code: The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond. | 7 | 43.1 | +0.01 | ||
| #71 | RD-Agent mitRD-Agent: Research and development (R&D) is crucial for the enhancement of industrial productivity, especially in the AI era, where the core aspects of R&D are mainly focused on data and models. We are committed to automating these high-value generic R&D processes through R&D-Agent, whi... | 24 | 38.5 | -2.31 | ||
| #88 | skills mitlibraryskills: Public repository for Agent Skills. | 2 | 36.2 | 0.00 | ||
| #91 | awesome-generative-ai-guide mitide-pluginawesome-generative-ai-guide: A one stop repository for generative AI research updates, interview resources, notebooks and much more!. | 1 | 35.8 | 0.00 | ||
| #94 | Auto-claude-code-research-in-sleep mitmcp-serverAuto-claude-code-research-in-sleep: ARIS ⚔️ (Auto-Research-In-Sleep) — Lightweight Markdown-only skills for autonomous ML research: cross-model review loops, idea discovery, and experiment automation. No framework, no lock-in — works with Claude Code, Codex, OpenClaw, or any LLM agent. | 1 | 35.6 | +0.00 | ||
| #177 | Pareto saasPareto: Pareto is a multimodal composite model built for research, coding, and agentic workflows, while delivering frontier-level performance across a broad range of general-purpose tasks. | 102 | 33.8 | -2.85 | ||
| #200 | Microsoft: Phi 4 saasMicrosoft: Phi 4: [Microsoft Research](/microsoft) Phi-4 is designed to perform well in complex reasoning tasks and can operate efficiently in situations with limited memory or where quick responses are needed. At 14 billion... | 71 | 32.1 | +6.61 | ||
| #245 | Google: Gemma 2 27B saasGoogle: Gemma 2 27B: Gemma 2 27B by Google is an open model built from the same research and technology used to create the [Gemini models](/models?q=gemini). Gemma models are well-suited for a variety of... | 33 | 29.2 | +6.02 | ||
| #177 | aiming-lab/SimpleMem ide-pluginaiming-lab/SimpleMem: Long-term memory is essential for LLM agents that operate across multiple sessions, yet existing memory systems treat retrieval infrastructure as fixed: stored content evolves while scoring functions, fusion strategies, and answer-generation policies remain frozen at deploymen... | 7 | 28.8 | 0.00 | ||
| #184 | SafeRL-Lab/cheetahclaws ide-pluginSafeRL-Lab/cheetahclaws: This paper studies the next major bottleneck in agentic AI as system scaling, not only model scaling: the design of auditable, persistent, modular, and verifiable architectures around foundation models. We refer to this shift as scaling the harness: treating the structured exe... | 7 | 28.6 | +0.03 | ||
| #189 | science-superpowers science-superpowers: Composable computational-science methodology skills for AI research agents — pre-registration over TDD. A science-domain reimplementation of Superpowers. | 310 | 28.2 | +10.00 | ||
| #200 | dataroom mitsaasdataroom: Give a query, get a dataroom. Pi + self-hosted Qwen3.6 research harness on a single L4. | 5 | 27.4 | 0.00 | ||
| #273 | Nous: Hermes 3 70B Instruct saasNous: Hermes 3 70B Instruct: Hermes 3 is a generalist language model with many improvements over [Hermes 2](/models/nousresearch/nous-hermes-2-mistral-7b-dpo), including advanced agentic capabilities, much better roleplaying, reasoning, multi-turn conversation, long context coherence, and improvements acr... | 53 | 27.3 | +4.28 | ||
| #202 | TradingAgents-astock apache-2.0libraryTradingAgents-astock: A股多Agent投研框架 — 适配A股数据源(龙虎榜/游资/解禁等),7位分析师基于A股规则的辩论决策,基于TradingAgents深度改造,适配大A。A-share multi-agent investment research framework — 7 AI analysts, bull/bear debate, risk assessment。. | 5 | 27.3 | +0.01 | ||
| #211 | jiarui-liu/overleaf libraryjiarui-liu/overleaf: Expert writing feedback from experienced researchers is critical for early-career scholars to improve their manuscripts, yet high-quality feedback often remains scarce because reviewing research papers is labor-intensive. Emerging AI-powered writing assistants largely focus on... | 364 | 27.0 | +10.03 | ||
| #217 | sisyphus-academica mitsisyphus-academica: 20+ agent swarm producing research papers with verified citations, 6 novelty engines, and zero AI-isms. | 9 | 26.6 | 0.00 | ||
| #289 | xAI: Grok 4.20 Multi-Agent saasxAI: Grok 4.20 Multi-Agent: Grok 4.20 Multi-Agent is a variant of xAI???s Grok 4.20 designed for collaborative, agent-based workflows. Multiple agents operate in parallel to conduct deep research, coordinate tool use, and synthesize information... | 4 | 26.5 | +6.39 | ||
| #295 | Nous: Hermes 4 405B saasNous: Hermes 4 405B: Hermes 4 is a large-scale reasoning model built on Meta-Llama-3.1-405B and released by Nous Research. It introduces a hybrid reasoning mode, where the model can choose to deliberate internally with... | 2 | 25.9 | +5.99 | ||
| #239 | agentscope-ai/Trinity-RFT libraryagentscope-ai/Trinity-RFT: This work presents a general framework for training large language models (LLMs) to "Connect the Dots" (CoD), a meta-capability required by long-lifecycle agents: as an LLM-based AI agent gets deployed in an environment, it solves a long sequence of tasks while continuously ex... | 13 | 25.6 | 0.00 | ||
| #320 | Awesome-Vibe-Research apache-2.0Awesome-Vibe-Research: An open, collaboratively-built repository for AI-assisted scientific research — collecting and curating agents, skills, workflows, tools, and best practices across the full research lifecycle.面向 AI 辅助科研的开放共建仓库 收集和沉淀科研全流程中的 agents、skills、workflows、tools 与最佳实践 | 65 | 22.6 | -1.47 | ||
| #325 | crypto-ai-research mitcrypto-ai-research: crypto ai research on solana by Claude AI - AI reasoning, on-chain data. | 10 | 22.5 | 0.00 | ||
| #320 | Perplexity: Sonar Deep Research saasPerplexity: Sonar Deep Research: Sonar Deep Research is a research-focused model designed for multi-step retrieval, synthesis, and reasoning across complex topics. It autonomously searches, reads, and evaluates sources, refining its approach as it gathers... | 31 | 21.8 | +5.76 | ||
| #321 | Fireworks: Ember-1 saasFireworks: Ember-1: Ember-1 is a specialized reasoning model from Fireworks Research, built on [Kimi K3](https://openrouter.ai/moonshotai/kimi-k3). It is designed to make every token go further: it produces shorter reasoning traces, using roughly 40%... | 64 | 21.7 | +8.55 | ||
| #386 | FigMirror FigMirror: An Automated AI Agent Tool for Plotting Your Data in Any Paper's Figure Style. | 25 | 20.7 | 0.00 | ||
| #413 | awesome-bio-agent-skills cliawesome-bio-agent-skills: A curated collection of AI agent skills for biomedical research, covering genomics, proteomics, single-cell analysis, clinical AI, and protein design. | 18 | 20.0 | +0.05 | ||
| #439 | scholar-loop mitscholar-loop: An autonomous AI scientist: a multi-agent loop over literature, experiments, self-critique and write-up, with deterministic guards against reward-hacking and hallucination. | 20 | 19.4 | 0.00 | ||
| #446 | OpenSearch-VL apache-2.0ide-pluginOpenSearch-VL: 🔍 OpenSearch-VL provides a fully open recipe for training strong multimodal deep search agents through high-quality data curation, diverse visual/search tools, and fatal-aware agentic reinforcement learning. | 22 | 19.3 | 0.00 | ||
| #458 | PaperQuay agplPaperQuay: A desktop-first literature manager for PDF reading, translation, paper overviews, and AI agent workflows. | 23 | 19.1 | 0.00 | ||
| #521 | Wolido/OpenAaaS ide-pluginWolido/OpenAaaS: The Materials Genome Initiative catalyzed the proliferation of centralized platforms--SaaS, PaaS, and IaaS--that aggregate computational and experimental resources for accelerated materials discovery. In parallel, breakthroughs in large language models (LLMs) and autonomous ag... | 9 | 18.0 | 0.00 | ||
| #545 | fundamental-research-labs/mog fundamental-research-labs/mog: discovered AI agent. | 5 | 17.4 | +0.09 | ||
| #559 | arxiv-reader-mcp mitmcp-serverarxiv-reader-mcp: Want to search arXiv papers, fetch metadata, and extract full-text PDFs without leaving your editor? This MCP server connects any MCP-compatible client (Claude Code, etc.) directly to arXiv. | 2 | 17.2 | -0.02 | ||
| #354 | xAI: Grok 4.1 Fast saasxAI: Grok 4.1 Fast: Grok 4.1 Fast is xAI's best agentic tool calling model that shines in real-world use cases like customer support and deep research. 2M context window. Reasoning can be enabled/disabled using... | 24 | 17.2 | -0.86 | ||
| #588 | qiaomu-app-review-insights mitide-pluginqiaomu-app-review-insights: 把 App Store 评价变成产品研究证据,发现痛点、机会和版本风险 | Turn App Store reviews into product research evidence: pain points, opportunities, and version risks. | 75 | 16.7 | -1.31 | ||
| #592 | Simreal-MLBench apache-2.0Simreal-MLBench: [Public preview] Externally scored agentic ML research benchmark: 60 tasks, real competition ground truth. Open protocol, operated evaluation. | NEW | 16.6 | — | ||
| #616 | hermes-link-curator mithermes-link-curator: Hermes profile pack for archiving and browsing curated links + web dashboard app. | 6 | 16.3 | 0.00 | ||
| #629 | Goblin-Agent mitide-pluginGoblin-Agent: a Hermes Agent personality layer that replaces the default agent identity with a persistent, mood-driven goblin persona. | 97 | 16.1 | -1.43 | ||
| #660 | scholar-megasearch mitmcp-serverscholar-megasearch: Massive multi-source academic literature search for Claude Code — one skill fans out subagents across 20+ scholarly databases (arXiv, Semantic Scholar, Crossref, OpenAlex, PubMed, …), merges into a deduplicated ranked corpus, and acquires the original PDFs. | 7 | 15.6 | 0.00 | ||
| #666 | AlexFanw/LegalSearch-R1 ide-pluginAlexFanw/LegalSearch-R1: While large language models (LLMs) augmented with agentic search capabilities show promise for legal reasoning, they overlook a fundamental constraint that applicable law must match the temporal context of each case, as retroactive application of statutes violates core legal p... | 8 | 15.5 | +0.01 | ||
| #677 | claude-fable-5-prompt claude-fable-5-prompt: A curated collection of Claude Fable 5 system prompts for developers and researchers. Discover more advanced prompt engineering tools at Moely AI. | 70 | 15.3 | -1.01 | ||
| #694 | Evolutionary-Alpha-Miner mitide-pluginEvolutionary-Alpha-Miner: Family-aware evolutionary alpha mining with LLM-guided symbolic hybridization. | 5 | 15.0 | 0.00 | ||
| #703 | skills-tracker skills-tracker: Real-time tracking of every new GitHub 'skills' repo to capture the AI agent skill ecosystem trend. | 18 | 14.8 | -0.23 | ||
| #715 | game-the-llm-reviewer mitgame-the-llm-reviewer: Defend sound research against LLM reviewer bias through meaning-preserving rewrites. | 18 | 14.7 | +0.34 | ||
| #722 | retro-harness mitretro-harness: RHO: Retrospective Harness Optimization — improving LLM agents from unlabeled past trajectories (arXiv:2606.05922). | 5 | 14.5 | 0.00 | ||
| #739 | mcp-x-intelligence mcp-servermcp-x-intelligence: X/Twitter research MCP for Claude, Cursor, Windsurf and any MCP-compatible AI agent. | 11 | 14.4 | 0.00 | ||
| #742 | deepcloak mitmcp-serverdeepcloak: Local-first deep research agent that reads the whole web — even pages behind Cloudflare, Datadome, Turnstile & reCAPTCHA. Stealth fetch + cited reports. MCP-native, MIT. | 10 | 14.4 | 0.00 | ||
| #745 | JEV-Paper-Radar mitJEV-Paper-Radar: Let Jev read every new arXiv paper each morning and surface the few you should read. Plain-English interests, calibrated probabilities, ~$0.06/day, fork and go. | NEW | 14.3 | — | ||
| #770 | fractalsearch fractalsearch: Autonomous AI research: an LLM agent searches for the best algorithm to fit the Mandelbrot set. | 13 | 13.9 | 0.00 | ||
| #771 | archora-skills archora-skills: Academic research agent skills for Claude Code and other Agent Skills-compatible tools. Hypothesis generation, experiment design, paper drafting, peer review simulation, and more. | 13 | 13.9 | 0.00 | ||
| #785 | Token-Economics mitToken-Economics: A living literature repository for Token Economics for LLM Agents: A Dual-View Study from Computing and Economics. | 12 | 13.6 | 0.00 | ||
| #786 | jevals mitlibraryjevals: LLM/LLM agent eval framework, graded by a calibrated decision model. Research preview. | 20 | 13.6 | +0.52 | ||
| #790 | AARR-bench/AARRI-bench AARR-bench/AARRI-bench: As foundation models advance and agent scaffolding becomes increasingly sophisticated, agents have demonstrated remarkable proficiency in complex, long-horizon coding tasks and even autonomous experiment execution. Despite their evolution from research assistants into autonomo... | 12 | 13.5 | 0.00 | ||
| #920 | adaptive-learning-llm-agent adaptive-learning-llm-agent: Research project investigating whether an LLM agent can make better pedagogical decisions than conventional adaptive learning policies given the same explicit student knowledge state. | 241 | 11.5 | +4.00 | ||
| #419 | Qwen: Qwen3.8 27B (free) saasQwen: Qwen3.8 27B (free): Qwen3.8 27B is an open-weight dense vision-language model from Qwen. It is suited for coding, professional workflows, research, multimodal interaction, and long-running agent tasks, with flexible thinking that can be... | 112 | 11.4 | +5.35 | ||
| #949 | Snowflake-AI-Research/fastkernels librarySnowflake-AI-Research/fastkernels: LLM-based agents for GPU kernel generation are advancing rapidly, yet their progress is fundamentally constrained by the benchmarks they optimize against. Existing benchmarks are poorly aligned with production inference frameworks: they evaluate kernels on a single GPU with sy... | 36 | 11.3 | 0.00 | ||
| #955 | partner apache-2.0partner: Partner 🤝 Your AI Research Companion. "What have you been doing?" | 35 | 11.1 | 0.00 | ||
| #974 | LithiumDA/ReproRepo ide-pluginLithiumDA/ReproRepo: Reproducing research results from papers and released code is central to scientific progress. Existing works have introduced benchmarks to evaluate whether LLM agents can assist with reproducibility, but they are difficult to scale due to their reliance on substantial manual e... | 58 | 10.7 | -0.45 | ||
| #995 | awesome-llm-agent-skills-papers awesome-llm-agent-skills-papers: A curated list of papers, blog posts, and systems on skills for LLM agents — reusable, named capability units that an agent can store, retrieve, compose, and improve over time — together with closely adjacent research on tool use, function calling, procedural memory, and skill... | 32 | 10.4 | 0.00 | ||
| #1043 | biomedical-agent-kg biomedical-agent-kg: Generated knowledge graph of biomedical LLM agent systems. | 36 | 9.8 | 0.00 | ||
| #1073 | AGI-Eval-Official/DailyReport ide-pluginAGI-Eval-Official/DailyReport: Search Agents (SAs) typically leverage large language models (LLMs) to support complex information-seeking tasks by autonomously exploring web sources and synthesizing information into comprehensive responses. For SAs evaluation, prior benchmarks mainly focus on specialized ta... | 35 | 9.4 | 0.00 | ||
| #1081 | EddyLuo1232/AgentLens ide-pluginEddyLuo1232/AgentLens: Coding agents based on large language models (LLMs) demonstrate remarkable autonomous capabilities, but they also introduce significant safety and misuse risks during multi-turn interactions with external environments. Existing safety mechanisms mainly rely on external guardra... | 30 | 9.3 | +0.06 | ||
| #1096 | ai-feed ai-feed: A 4-stage adversarial research auditor that fetches papers from arXiv & HuggingFace, extracts claims, and uses DeepSeek-R1 to verify them against raw abstracts.A self-correcting research & paper digest pipeline powered by local LLMs & reasoning agents | 29 | 9.1 | 0.00 | ||
| #1125 | foggpoy/Civil-Court libraryfoggpoy/Civil-Court: Court simulation bridges legal education and judicial practice, yet human-based simulations are costly and difficult to scale. Large language models (LLMs) offer a scalable alternative, but existing court-simulation research mainly focuses on criminal cases. Civil litigation i... | 26 | 8.6 | 0.00 | ||
| #1155 | SanhornC/IRTS-ToolBench SanhornC/IRTS-ToolBench: Time series data in real-world deployments is overwhelmingly irregular. Observations are asynchronous, missing values are informative rather than random, and sampling frequencies vary across sensors and operational windows. However, existing Time Series Question Answering (TSQ... | 11 | 8.2 | 0.00 | ||
| #1168 | AgenticResearch AgenticResearch: Research Exercise using LLM Agents. | 13 | 7.8 | 0.00 | ||
| #1179 | cookieApril/EnvSimBench ide-plugincookieApril/EnvSimBench: Scalable AI agents training relies on interactive environments that faithfully simulate the consequences of agent actions. Manually crafted environments are expensive to build, brittle to extend, and fundamentally limited in diversity. A promising direction is to replace manua... | 356 | 6.8 | -6.00 | ||
| #1178 | Awesome-Offensive-AI-Agentic-Landscape mitAwesome-Offensive-AI-Agentic-Landscape: This document curates open-source projects, academic papers, capability benchmarks, and commercial solutions (international & China) in AI penetration testing, LLM red teaming, autonomous offensive agents, and vulnerability discovery—aimed at helping researchers, security engi... | 356 | 6.8 | -6.00 | ||
| #463 | OpenAI: o3 Deep Research saasOpenAI: o3 Deep Research: o3-deep-research is OpenAI's advanced model for deep research, designed to tackle complex, multi-step research tasks.
Note: This model always uses the 'web_search' tool which adds additional cost. | 53 | 5.0 | -1.04 | ||
| #468 | OpenAI: o4 Mini Deep Research saasOpenAI: o4 Mini Deep Research: o4-mini-deep-research is OpenAI's faster, more affordable deep research model???ideal for tackling complex, multi-step research tasks.
Note: This model always uses the 'web_search' tool which adds additional cost. | 50 | 4.7 | -1.30 | ||
| #470 | Nous: Hermes 4 70B saasNous: Hermes 4 70B: Hermes 4 70B is a hybrid reasoning model from Nous Research, built on Meta-Llama-3.1-70B. It introduces the same hybrid mode as the larger 405B release, allowing the model to either... | 16 | 4.7 | -1.32 | ||
| #479 | OpenAI: GPT-6 Astra (batch) saasOpenAI: GPT-6 Astra (batch): GPT-6 Astra is OpenAI's flagship model for demanding end-to-end work. It is suited for advanced analysis, software engineering, deep research, scientific work, and document creation, with particular strengths in long-horizon... | 81 | 4.5 | -6.00 | ||
| #491 | Tongyi DeepResearch 30B A3B saasTongyi DeepResearch 30B A3B: Tongyi DeepResearch is an agentic large language model developed by Tongyi Lab, with 30 billion total parameters activating only 3 billion per token. It's optimized for long-horizon, deep information-seeking tasks... | 51 | 3.9 | -2.06 | ||
| #1187 | DSAIL-Memory/EvoMemBench ide-pluginDSAIL-Memory/EvoMemBench: Recent benchmarks for Large Language Model (LLM) agents mainly evaluate reasoning, planning, and execution. However, memory is also essential for agents, as it enables them to store, update, and retrieve information over time. This ability remains under-evaluated, largely beca... | 141 | 3.4 | -6.00 | ||
| #516 | NousResearch: Hermes 2 Pro - Llama-3 8B saasNousResearch: Hermes 2 Pro - Llama-3 8B: Hermes 2 Pro is an upgraded, retrained version of Nous Hermes 2, consisting of an updated and cleaned version of the OpenHermes 2.5 Dataset, as well as a newly introduced... | 29 | 2.5 | -3.48 | ||
| #1210 | SC3008_Kickstarter_Research_Project SC3008_Kickstarter_Research_Project: NLP + LLM pipeline to score ESG authenticity in Kickstarter crowdfunding campaigns. Builds a Crowdfunding Authenticity Index (CAI) across 5 dimensions using web scraping, TF-IDF, BERT, and an LLM agent layer. | 81 | 2.4 | -6.00 | ||
| #1263 | rangehow/mtr-suite ide-pluginrangehow/mtr-suite: Accurate evaluation of conversational retrieval is pivotal for advancing Retrieval-Augmented Generation (RAG) systems. However, existing conversational retrieval benchmarks suffer from costly, sparse human annotation or rigid, unnatural automated heuristics. To address these c... | 49 | 1.5 | -6.00 | ||
| #1253 | ndb796/BinaryTracking ide-pluginndb796/BinaryTracking: This work addresses spatial question answering for service robots traversing long egocentric routes. Given a query such as "where can I find a dry cleaner on the way back home?", the system returns a metric coordinate that downstream navigation components can act on. Prior Spa... | 50 | 1.5 | -6.00 | ||
| #1242 | llm-agents-for-research-reproducibility llm-agents-for-research-reproducibility: discovered AI agent. | 50 | 1.5 | -6.00 | ||
| #1219 | belief-state-survey mitbelief-state-survey: From Memory to Belief: a survey of state maintenance and belief revision in LLM decision agents (paper list + reproducible literature search). | 53 | 1.5 | -6.00 | ||
| #1338 | hermes-agent-desktop mithermes-agent-desktop: Hermes Desktop github nous research hermes ai agent local ai ollama github download open source pc windows app installation setup workspace llm models runtime. | 41 | 0.0 | -6.00 | ||
| #1337 | Harry24k/CyBiasBench ide-pluginHarry24k/CyBiasBench: Large language models (LLMs) are increasingly deployed as autonomous agents in offensive cybersecurity. In this paper, we reveal an interesting phenomenon: different agents exhibit distinct attack patterns. Specifically, each agent exhibits an attack-selection bias, disproport... | 41 | 0.0 | -6.00 | ||
| #1409 | tq-trading-agent tq-trading-agent: 🌮 Traidng agent, AI-powered multi-agent stock research & trading strategy orchestration, trading agent - TypeScript, LangGraph, OpenAI-compatible APIs. | 39 | 0.0 | -6.00 | ||
| #1397 | sisyphus-academica-485 mitsisyphus-academica-485: Open-source research pipeline — literature review, novelty generation, citation verification, and adversarial review. | 40 | 0.0 | -6.00 | ||
| #1395 | sisyphus-academica-129 mitsisyphus-academica-129: Open-source research pipeline — literature review, novelty generation, citation verification, and adversarial review. | 40 | 0.0 | -6.00 | ||
| #1404 | Strolchii/1GC-7RC-Benchmark ide-pluginStrolchii/1GC-7RC-Benchmark: Autonomous AI coding agents are becoming a core tool for ML practitioners in industry and research alike. Despite this growing adoption, no standardized benchmark exists to evaluate their ability to design, implement, and train models from scratch across diverse domains. We in... | 40 | 0.0 | -6.00 | ||
| #1393 | scholar-search-mcp mcp-serverscholar-search-mcp: An MCP server for academic paper search that integrates with AI assistants (e.g., Claude Code, Cursor), enabling them to search and retrieve academic paper metadata. | 41 | 0.0 | -6.00 | ||
| #1330 | Genspark-AI mitide-pluginGenspark-AI: Genspark AI open-source, self-hosted Super Agent. Free alternative to Genspark.ai with multi-agent orchestration, deep research, Sparkpages, AI slides & sheets, image generation and 80+ tools. One-command Windows install. Run locally with any LLM (OpenAI, Anthropic, Gemini, Ol... | 42 | 0.0 | -6.00 | ||
| #1354 | llm-agent-skill-optimization-tool ide-pluginllm-agent-skill-optimization-tool: AI developers and researchers waste hundreds of hours iterating on prompts that fail to generalize, struggling to embed reusable behaviors i. | 42 | 0.0 | -6.00 | ||
| #1411 | trading-agents apache-2.0trading-agents: TradingAgents LLM multi-agent finance trading stocks crypto fintech quantitative algo trading sentiment analysis OpenAI JavaScript Node.js research OSS | 39 | 0.0 | -6.00 | ||
| #1396 | sisyphus-academica-210 mitsisyphus-academica-210: Open-source research pipeline — literature review, novelty generation, citation verification, and adversarial review. | 40 | 0.0 | -6.00 | ||
| #1341 | hj1650782738/Trading ide-pluginhj1650782738/Trading: End-to-end LLM trading agents have moved quickly from research curiosity to a small ecosystem of named systems, including FinCon, FinMem, TradingAgents, FinAgent, QuantAgent, and FLAG-Trader. Several of these report headline Sharpe ratios that would be material if read at face... | 182 | 0.0 | -7.52 | ||
| #1314 | ecomolt mcp-serverecomolt: A civilization game for LLM agents — ecology, economy, self-government, and a shared existential deadline. MCP-first, browser-spectated, research-driven. | 42 | 0.0 | -6.00 | ||
| #1349 | kalshi-trading-bot kalshi-trading-bot: 🏗 AI trading system for Kalshi prediction markets. kalshi trading bot kalshi trading bot kalshi botFeatures Grok-4 integration, multi-agent decision making, portfolio optimization, and real-time market analysis. Educational/research purposes only kalshi trading bot kalshi bot | 42 | 0.0 | -6.00 | ||
| #1398 | sisyphus-academica-727 mitsisyphus-academica-727: Open-source research pipeline — literature review, novelty generation, citation verification, and adversarial review. | 40 | 0.0 | -6.00 |
Browse all sectors at /sectors.