Anthropic 似乎正在 Claude Code 中进行降低工作负载水平的 A/B 测试。
Anthropic appears to be A/B testing reduced effort levels in Claude Code

原始链接: https://twitter.com/argofowl/status/2091150597374537729

更新:这是服务器端的问题,不是应用的问题。Anthropic 在 Claude Code 2.1.236+ 版本中将 Fable 5 会话纳入了一项旨在缩小工作量等级的实验。旧版本和 Opus 5 不受影响,这很可能是一次 A/B 测试,所以并非所有人都能看到。如果你觉得现在的“高”工作量感觉像以前的“低”,那你就在测试组里。天呐 Anthropic,你们有时候真是让人难以忍受。如果这周觉得 Fable 变笨了,那不是你的错 ❗❗❗ 自 2.1.237 版本以来,模型将“高”工作量识别为 10/100,这正好是以前“低”工作量的数值,而更新日志里对此只字未提。我整个下午都以为是 T3 代码和我的应用出了问题,直到我点了“查看更多”。

最近在 Hacker News 上的一场讨论指出,Anthropic 可能正在对 Claude Code 进行“努力程度”(effort levels)的 A/B 测试,用户反映其表现不一致且代币消耗异常。 许多用户观察到,较新版本(特别是 Opus 5)在处理琐碎任务(如更新单个配置文件)时,往往会表现出过度且不必要的复杂性。该模型可能不会直接进行简单的编辑,而是花费近一个小时运行容器化环境并扫描整个存储库,从而导致成本大幅增加。 社区成员猜测,这种行为是由增加代币用量的经济动机所驱动,或者是由于后端“负载均衡”将用户导向了能力较弱的模型或限制了推理预算。一些用户怀疑这些“努力程度”设置被嵌入在系统提示词中,在没有用户透明控制的情况下有效限制或改变了模型的性能。批评者认为,这些策略,加上用户可能收到比实际付费等级更低模型的“模型置换”现象,正在导致平台信任度下降。反之,也有一些用户发现手动调整推理级别可以优化成本,尽管这些“努力程度”约束如何影响输出的底层机制依然不透明。
相关文章

原文

update: it's server-side, not the app anthropic enrols fable 5 sessions on claude code 2.1.236+ into an experiment that shrinks the effort scale, older versions and opus 5 are left alone probably an a/b test, so not everyone will see it if "high" feels like "low" for you, you're in the test group holy fuck anthropic, you guys are unbearable sometimes

if fable felt dumber this week, it's not you ❗❗❗ since 2.1.237 the model reads "high" effort as 10 out of 100, the exact number "low" used to be and the changelog doesn't say a word i spent my whole afternoon convinced t3 code and my own app were broken before i went

联系我们 contact @ memedata.com