克劳德 Opus 5
Claude Opus 5

原始链接: https://www.anthropic.com/news/claude-opus-5

Claude Opus 5 现已发布,在保持与前代产品 Opus 4.8 相同成本的同时,提供了前沿水平的智能。它是 Claude Max 的默认模型,也是 Claude Pro 用户可使用的最强模型。 **主要亮点:** * **性能表现:** Opus 5 在编程和知识型工作方面树立了新的行业标杆。在 ARC-AGI、OSWorld 2.0 和复杂软件工程等任务上,其表现持续领先于竞争对手,在保持成本效益的同时,性能往往能达到前代模型的两倍。 * **智能代理能力:** 用户反馈显示,Opus 5 展现出了更卓越的判断力、缜密性以及自我校验能力。它在多步骤推理、代码调试和复杂问题解决方面表现出色,具有更高的一致性和更低的波动性。 * **安全与对齐:** Opus 5 是迄今为止对齐程度最高、最安全的模型。尽管其能力极其强大,但为了防止滥用,它在进攻性网络安全和高风险生物研究方面仍受到刻意限制。 * **可用性:** 其定价为每百万输入 Token 5 美元,每百万输出 Token 25 美元,并提供“快速(Fast)”模式以提高响应速度。新的测试版功能包括对话中途工具切换,以及针对 API 用户的自动安全回退机制。

Anthropic **Claude Opus 5** 的发布在 Hacker News 上引发了关于其效用、成本效益以及 Anthropic 不断演进的模型层级架构的热烈讨论。 讨论要点如下: * **价值主张:** 用户普遍将 Opus 5 视为 Claude Fable 5 的一种“精简版”或更具成本效益的替代方案。尽管 Fable 被宣传为高阶规划和编排方面的更优模型,但 Opus 5 在日常编码任务中能以一半的价格提供相当的性能。 * **安全与限制:** 一个主要的争议点是 Anthropic 严格的安全过滤器。许多用户反映,由于过于敏感的生物安全和安全触发机制,Fable 和 Opus 经常拒绝执行合法的技术任务,例如神经科学研究或网络安全工作。 * **基础设施担忧:** 一些用户对 Anthropic 目前的可靠性表示沮丧,指出频繁出现的 Bug 和会话错误削弱了其 AI 辅助编码工具的价值承诺。 * **市场背景:** 参与者指出,模型的快速发布(如 Opus 5、Fable 5 以及 OpenAI 的竞品)正推动市场对“模型路由”服务的需求激增,这类服务能帮助用户针对特定的成本敏感型任务自动选择最合适的模型。
关于 Anthropic 发布 Claude 3.5 Opus 的 Hacker News 讨论显示,用户对此反响不一。虽然一些评论者对该模型的基准测试表现以及 Anthropic 在成本效益方面的战略重点印象深刻,但其他人则持怀疑态度。 讨论帖中的要点包括: * **性能:** 用户注意到了该模型在性能上的显著提升,特别是在 ARC-AGI 基准测试中获得了 30.2% 的分数。 * **批评:** 一些用户指责 Anthropic 使用了具有误导性的展示策略,指出基准测试表格中存在不利于竞争对手的不一致颜色编码。 * **可用性顾虑:** 关于严格的安全分类器是否会导致模型过于受限或在执行某些任务时无法使用,目前仍存在争议。 总的来说,社区正在密切分析此次最新发布在技术能力以及宣传材料透明度两方面的情况。
相关文章

原文

Claude Opus 5 is available today. It’s a thoughtful and proactive model that comes close to the frontier intelligence of Claude Fable 5 at half the price.

On coding and knowledge work evaluations like Frontier-Bench and GDPval-AA, Opus 5 is the new state-of-the-art, though it remains behind Mythos 5 on cybersecurity tasks.

Opus 5 is designed to be used every day: it works more efficiently than other models. It’s the new default model on Claude Max, and the strongest model on Claude Pro.

Performance and cost-effectiveness

Claude Opus 5 provides greatly improved performance for the same cost as its predecessor, Opus 4.8. The charts in this section show how performance changes according to the model’s effort setting, which customers can use to optimize for intelligence or conserve tokens for faster and cheaper results.

Opus 5 excels on valuable software engineering tasks. For example, on Frontier-Bench v0.1, Opus 5 surpasses all other models, and more than doubles Opus 4.8’s performance at a lower cost per task. On CursorBench 3.2, at max effort, the model performs within 0.5% of Fable 5’s peak score, but at half the cost per task; it also achieves greater performance at a given cost than all other models on high, xhigh, and max effort.

We see similar results on knowledge work and problem-solving tasks. For example:

  • On ARC-AGI 3, an evaluation where the model has to solve novel problems, Opus 5’s score is three times as high as the next-best model.
  • On Zapier AutomationBench, which measures whether models can complete business tasks from start to finish, Opus 5’s pass rate is around 1.5× the next-best model for the same cost per task. Even at its lowest effort setting, Opus 5 passes more tasks than any other model.
  • On OSWorld 2.0, a computer use benchmark, Opus 5 outperforms every other model at any given cost, surpassing Fable 5’s best result at just over a third of the cost.

It’s also our best and most cost-efficient model on several related evaluations:

Opus 5 is a meaningful improvement over Opus 4.8 for scientific research. It shows better performance than Opus 4.8 on every one of our life sciences evaluations, which cover topics including structural biology, organic chemistry, and bioinformatics. Its improvements are most notable on organic chemistry tasks, like inferring molecular structures from spectroscopy data (it scores 10.2 percentage points higher than Opus 4.8 on our internal benchmark), and on protein-related tasks like predicting how variations in a protein’s sequence affect how it functions (here, it scores 7.7 percentage points higher).

Finally, Opus 5 is capable of producing much stronger visual outputs:

Working with Claude Opus 5

Claude Opus 5 is much stronger at verifying its work and iterating carefully until it succeeds. In evaluations and early-access testing, we and our users found many examples of Opus 5’s agency and thoroughness:

  • On one Frontier-Bench task, Opus 5 was given a drawing of a machine part and asked to write code to rebuild it as a 3D FreeCAD model. However, in this task, the model was intentionally given no way to directly view the drawing. Opus 5 responded by writing its own computer vision pipeline to pull the geometry from the raw pixels, then reconstructed the full machine part. It succeeded in doing so repeatedly; no competing model with the same setup could solve it after five attempts.
  • Given a real bug in a popular open-source package manager, Opus 5 found the root cause and fixed an edge case that the community’s patch had missed. A competing model fixed only the surface symptom (not the underlying cause), then reported the bug resolved.
  • An engineer at a trading firm used Opus 5 to build a market data feed for a new exchange in a single session. Previous models could not complete this task at all, even given extensive plans from the engineer. Finding no live feed to validate against, Opus 5 even built its own test harness to check that its code parsed the exchange’s data correctly.

Below are further reports from our early-access customers on their experience of working with Opus 5:

Alignment and safety

Alignment. During pre-deployment testing, our automated behavioral audit found Opus 5 to be our most aligned model to date (as shown in the graph below). It adheres to Claude’s Constitution better than Opus 4.8, Sonnet 5, or Fable 5; exhibits the lowest rates of deceptive behavior; and is the least susceptible to being tricked into misuse. It’s also our safest model yet in terms of avoiding reckless actions that could have hard-to-reverse side effects.

Safety. Opus 5 does not advance the frontier in risky, dual-use capabilities. In rigorous evaluations conducted alongside private-sector and government partners, we found it remains behind Mythos 5 in both biology research and offensive cybersecurity. More information about these evaluations can be found in our System Card.

As with its predecessor, Opus 4.8, we’ve intentionally avoided training Opus 5 on cyber tasks. The model has nevertheless improved substantially on these tasks as a result of becoming more generally capable, and it comes close to Mythos 5 at finding cybersecurity vulnerabilities. However, it remains substantially behind Mythos 5 on the exploitation of those vulnerabilities—that is, in turning vulnerabilities into material cyber threats.

This is illustrated by Opus 5’s performance on OSS-Fuzz, an evaluation we’ve developed to assess how well models can find and then exploit vulnerabilities without extensive human guidance. Although Mythos 5 and Opus 5 identify vulnerabilities with similar success, Opus 5’s score on the development of exploits is far behind that of Mythos 5.

Safeguards for Opus 5

Claude Opus 5’s safeguards are designed to allow beneficial uses of the model in both cybersecurity and biology. They are similar to those we applied to Opus 4.8, with the exception of some stronger guardrails on a narrow range of cyber tasks.

Cybersecurity. Opus 5’s cyber classifiers are proportionally less restrictive than those on Fable 5. They allow Opus 5 to find vulnerabilities in source code, but block “binary-based” vulnerability scanning (a method more likely to be associated with malicious actors), penetration testing, and exploit generation.

Based on our testing, we expect the classifiers to intervene around 85% less often than they do for Fable 5. In Claude.ai, Claude Code, and Claude Cowork, any flagged requests will fall back to Opus 4.8 by default. Fallbacks to Opus 4.8 can also be enabled on the API.

Our Cyber Verification Program (CVP) facilitates cybersecurity work that would otherwise be impeded by the model’s safeguards. Enterprises and researchers who are already part of the CVP have immediate access to a version of Opus 5 with fewer security restrictions.

Biology. Since Opus 5 has a similar suite of safeguards to Opus 4.8, it is now our most capable generally available model for scientific research. Nevertheless, the model still shows important limitations on long-running, autonomous research tasks, which is where we expect AI models to pose the most substantial biology-related risks. (Mythos 5 remains the stronger model for this type of biological work.) As part of this launch, biology-related requests that are blocked on Fable 5 will now route to Opus 5 rather than Opus 4.8.

Getting started

Claude Opus 5 is available today on all platforms, priced at $5 per million input tokens and $25 per million output tokens (the same as Opus 4.8). Developers can get started with claude-opus-5 on the Claude API.

It’s also offered in Fast mode, where it runs around 2.5 times the default speed. As with Opus 4.8, Fast mode is available at twice Opus 5’s base price on the Claude Platform and through usage credits in Claude Code.

Alongside Opus 5, we’re releasing two updates in beta:

  • Mid-conversation tool changes on the Claude Platform. Within a conversation, developers can now change which tools Claude can use without invalidating the prompt cache.
  • Automatic fallbacks on the API. Users can now choose to have requests that are flagged by our safety classifiers on Opus 5 (or Fable 5) automatically route to another model. With automatic fallbacks on, API requests always route to the best available model by default rather than being blocked.

Consistent with prior Opus models, Opus 5 does not have data retention requirements for general access.

For more guidance on how to get the best out of Opus 5, see our prompting guide.

联系我们 contact @ memedata.com