Qwen3.8 Max 现已被 Agentic Index 评为综合表现最佳的模型。
Qwen3.8 Max now ranked as the best overall model by agentic index

原始链接: https://artificialanalysis.ai/?intelligence=agentic-index

在 7 月下旬至 8 月初期间,人工智能领域表现活跃,其中最受瞩目的是 **Claude Opus 5** 的发布,它以更具性价比的价格为代理型知识工作设定了新标准。 这一时期伴随着广泛的模型评估,包括 **DeepSeek V4 Flash** 和 **Kimi K3** 的更新,以及 **Ling 3.0**、**Muse Spark 1.2** 和 **Inkling Small** 等新模型的推出——其中 Inkling Small 以极少的参数实现了与前代产品相当的性能,令人印象深刻。 此外,平台更新包括推出了旨在追踪同一模型在不同部署环境下表现的“端点准确度指数”(Endpoint Accuracy Index)。在方法论方面,对“任务成本”(Cost per Task)的计算进行了调整,以提高价格估算的准确性。总体而言,这一时期反映了行业持续优化模型性能、增强推理能力并改善“成本-智能”比率的趋势。

Hacker News 近期的一项讨论审视了关于 Qwen3.8 Max 模型在“Artificial Analysis Agentic Index”中排名第一的说法。 尽管该模型在该特定指标上排名靠前,但评论者指出了几个争议点: * **方法论困惑:** 用户指出该模型未出现在“编码智能体指数(Coding Agent Index)”和“原始智能(raw intelligence)”排名中,引发了关于 Agentic Index 是否能作为衡量通用性能的可靠指标的争论。 * **性能差异:** 用户反馈褒贬不一;一些用户称赞 Qwen 在构建诊断工具和解决复杂 Bug 方面的能力,而另一些用户则认为该模型在处理自动化任务时表现“草率”且不可靠。 * **市场背景:** 讨论还涉及了运行大型开放权重模型的成本,并将其与 GPT 等专有替代方案进行了比较。支持者认为,开放模型通过微调、可移植性以及摆脱单一提供商的定价或审查决策,提供了更长期的价值。 最终,用户对“整体最佳”的标签表示怀疑,认为排名可能会受到速度、成本以及特定基准测试选择等不同权重的影响。
相关文章

原文

New language model evaluation · 6 Aug

Ling 3.0 TinyLing 3.0 Tiny

New article published · 5 Aug

Muse Spark 1.2

New language model evaluation · 5 Aug

Qwen3.8 MaxQwen3.8 Max

New language model evaluation · 5 Aug

Ling-3.0-flashLing-3.0-flash

New language model evaluation · 5 Aug

Muse Spark 1.2 (xhigh)Muse Spark 1.2 (xhigh)

New article published · 4 Aug

Launching the Endpoint Accuracy Index: Same Model, Different Accuracy

New language model evaluation · 3 Aug

G9v3-39A5BG9v3-39A5B

New article published · 31 Jul

DeepSeek V4 Flash 0731 scores 50 on the Artificial Analysis Intelligence Index, 10 points above previous DeepSeek V4 Flash

New language model evaluation · 31 Jul

Celeris-1Celeris-1

New language model evaluation · 31 Jul

DeepSeek V4 Flash 0731 (Reasoning, Max Effort)DeepSeek V4 Flash 0731 (Reasoning, Max Effort)

New article published · 30 Jul

Inkling Small lands within a point of Inkling on the Artificial Analysis Intelligence Index with less than a third of the parameters

Methodology updated · 30 Jul

Artificial AnalysisWe have updated our Cost per Task methodology, resulting in slight absolute increases in cost estimates but with minimal impact on relative positioning.

New language model evaluation · 30 Jul

Kimi K3 (low)Kimi K3 (low)

New language model evaluation · 30 Jul

Inkling SmallInkling Small

New article published · 29 Jul

Agnes AI releases Agnes 2.5 Pro Alpha

New article published · 24 Jul

Claude Opus 5: the new leader in agentic knowledge work

New article published · 24 Jul

Opus 5: Fable 5 level intelligence at a lower cost per task

New language model evaluation · 24 Jul

联系我们 contact @ memedata.com