✦ 今日最强 AI

MODEL RADAR · 少刷新闻,多做事情

今天各家最新模型与能力榜单

30 秒,知道现在该用哪个模型。

最近成功核查:2026-09-15T06:34:08+00:00 · 不代表所有模型都在今天评测

默认先看最新发布,再看有同一指标依据的多名次排行榜。厂商色块是文字品牌标识,不代表合作背书。

LATEST RELEASES

各家最近发布了什么

最新 ≠ 最强;仅列已从官方页面核实的型号。

Anthropic
已发布 · 可用

Claude Fable 5.1

官方日期:2026-08-02

官方称其面向编程、知识工作和长程问题解决;这里只把它列为最新发布,不等于自动获得所有榜单第一。

来源:www.anthropic.com · claude-fable-and-mythos-5-1

Anthropic
已发布 · 可用

Claude Opus 5

官方日期:2026-07-24

官方发布页称其在 Claude 平台、Claude Code 等渠道可用;能力排名另见同口径评测榜。

来源:www.anthropic.com · news/claude-opus-5

DeepSeek
API 可用

DeepSeek-V4.1-Flash

官方日期:2026-09-14

官方 API 文档说明旧版名称的请求由 DeepSeek-V4.1-Flash 提供服务;这是接口文档核实,不是独立能力排名。

来源:DeepSeek API Docs · News / model names

Qwen
已发布 · 安全模型

Qwen3Guard

官方日期:2025-09-23

Qwen 官方博客介绍的安全护栏模型;它不是通用聊天模型,因此不与通用模型直接争夺综合榜名次。

来源:Qwen Blog · Qwen3Guard

厂商核查范围: OpenAI · Anthropic · Google · xAI · Meta · Qwen · DeepSeek · Mistral
未显示的厂商不是没有新模型,而是本轮没有通过来源核验;避免把旧型号冒充最新。

RANKINGS

多名次能力排行榜

同一榜单、同一指标内比较,不跨榜单拼分。

Artificial Analysis Intelligence Index

综合智能指数(max/high effort,数值越高越前) · 数据日期 2026-09-01 · artificialanalysis.ai · articles/claude-fable-5-1

  1. 1Claude Fable 5.166 分
  2. 2Claude Opus 563 分
  3. 3Claude Fable 562 分
  4. 4GPT-5.6 Sol61 分
  5. 5Grok 4.661 分

引用原文:We supported Anthropic with pre-release evaluation of Claude Fable 5.1. At max effort it scores 66 on the Artificial Analysis Intelligence Index, the highest score we have measured, ahead of Claude Opus 5 (max, 63), Claude Fable 5 (max, 62), GPT-5.6 Sol (max, 61) and Grok 4.6 (high, 61). We evaluated the model with Anthropic's 'default' server-side fallback, which routes safety-flagged requests to Claude Opus 4.8 or Claude Opus 5; fallback served ~4% of output tokens across the Intelligence Index. Key takeaways: ➤ Frontier intelligence with improvements across benchmarks: Fable 5.1 gains +4 points on the Intelligence Index over Fable 5. On HLE, Fable 5.1 scores 59.1%, ahead of the previous best of 55.5% from Claude Fable 5. It posts the narrowly highest scores we've seen on Terminal-Bench v2.1 (91.4%) and SciCode (62.0%), and on 𝜏³-Banking it gains 9 points over Fable 5 ➤ 75% cache read price cut, but Fable 5.1 still costs more per task: Anthropic has cut the cache read price from $1 to $0.25 per 1M cached input tokens, with standard

按用途的编辑推荐

综合

编程

推理

性价比

本次核查摘要

本版新增“各家最新模型”与真正多名次榜单:最新卡片只采用官方可用性证据,能力榜仅引用同一份 Artificial Analysis 指标,不把最新发布等同于最强。当前可核实覆盖集中在 Anthropic、DeepSeek、Qwen;OpenAI、Google、xAI、Meta、Mistral 的最新型号待下一轮官方页面核查。

已发布 · 待独立验证

本次没有收录待验证条目。

怎样看这份榜单

“最新发布”与“能力排名”是两件事:上方先看各家最近可核实的模型,下方排行榜只在同一来源、同一指标内排序。本站不是自跑分,也不把不同测试硬拼成总分。

仅有厂商发布、缺少独立证据的模型留在待验证区。传闻与已发布严格分开。来源无法核查时保留旧版并显示原时间,不把旧消息包装成今日更新。

当前为预发布核验版,来源覆盖有限;不强凑前三。本站不收集访客输入,不设登录或追踪器。