Published scores

AI models

Rankings from named public leaderboards only. Each bar is a real published metric with a source and timestamp. ANN does not invent an “Agentic Index.” Closed models missing from a source simply will not appear on that tab.

Last ingest: Oct 7, 2026, 9:07 AM UTC. Want to try models yourself? Prompt Compare is the hands-on tool.

Source: MMLU-Pro · methodology · as of Oct 7, 2026, 9:07 AM UTC · verified rows preferred when the API marks them

  1. 1FINAL-Bench/Darwin-180B-RSI88.1
  2. 2MiniMaxMiniMax88
  3. 3internlm/Intern-S2-Preview88
  4. 4QwenAlibaba87.8
  5. 5DeepSeekDeepSeek87.5
  6. 6KimiMoonshot AI87.1
  7. 7NemotronNVIDIA86.8
  8. 8NemotronNVIDIA86.8
  9. 9QwenAlibaba86.7
  10. 10DeepSeekDeepSeek86.4
  11. 11QwenAlibaba86.2
  12. 12upstage/Solar-Open2-250B86.2
  13. 13QwenAlibaba86.1
  14. 14GLM / ChatGLMZhipu AI86
  15. 15QwenAlibaba85.2
  16. 16DeepSeekDeepSeek85
  17. 17GLM / ChatGLMZhipu AI84.6
  18. 18stepfun-ai/Step-3.5-Flash84.4
  19. 19DeepSeekDeepSeek84
  20. 20LGAI-EXAONE/K-EXAONE-236B-A23B83.8
  21. 21NemotronNVIDIA83.7
  22. 22NemotronNVIDIA83.7
  23. 23internlm/Intern-S183.5
  24. 24LGAI-EXAONE/EXAONE-4.5-33B83.3

Artificial Analysis remains the industry-familiar composite index. Until ANN has a Commercial redistribution license, that tab links out instead of copying scores. Suggest labs via Companies.