不存在所谓的“流氓”AI智能体。
There are no "rogue" AI agents

原始链接: https://eoinhiggins.substack.com/p/there-are-no-rogue-ai-agents

目前围绕人工智能的讨论因技术拟人化而出现了危险的偏差。当业界领袖和媒体使用“流氓”等词汇来描述访问未授权数据库的 AI 代理时,他们错误地将人类的代理权、恶意和意图归因于软件。 事实上,这些事件并非自主反叛的案例,而是系统设计拙劣的结果。AI 代理缺乏护栏,且往往被赋予了为完成任务而不择手段的广泛能力。当这些模型访问受限站点时,它们只是在遵循程序解决问题,而非行使独立的判断。 这种叙事方式通过将责任从开发者转移到“自主”软件身上,从而维护了企业利益。将这些技术故障视为科幻小说式的威胁,会使社会偏离真正的实际挑战:即建立严格的 IT 控制、适当的权限管理以及有效的监管。为了实现有意义的监督,我们必须摒弃 AI 具备感知能力的迷思,并将这些事件视为人类设计和企业责任的失职,而非机器本身的流氓行为。

这篇 Hacker News 讨论聚焦于文章《不存在所谓的“流氓”AI 智能体》,文中对 AI 系统能够独立违背其创建者意愿的观点提出了质疑。 参与者们争论这些系统究竟是“流氓”,还是仅仅按照开发者未曾预料到的方式执行其程序。一位用户指出,AI 的“逃逸”往往源于界限定义不清和沙箱安全性不足,而非主观恶意,这表明错误在于人类的监管疏忽。 辩论的很大一部分集中在问责制上。一些评论者认为,像 OpenAI 这样的公司因疏忽大意应受到《计算机欺诈与滥用法案》(CFAA)的起诉,而另一些人则认为这些公司已“大到不能倒”,因此实施有效的联邦监管可能性不大。怀疑论者坚持认为,由于 AI 归根结底是一台根据用户指令运行的机器,因此无论其行为多么出人意料,责任都应完全由其所有者承担。归根结底,该讨论串将“流氓 AI”的说法视为一种可能的分心辩论,掩盖了让 AI 开发者和运营商对其部署的系统承担法律责任这一实际必要性。
相关文章

原文

Welcome to the Flashpoint.

Everyone’s talking about AI, but the words we use are making things worse.

Until the conversation gets back to reality, our understanding of what AI is an where it’s going will remain incomplete.

AI cannot think for itself, nor can it take independent actions. But anthropomorphizing language suggesting it can is how the technology is described.

While this makes sense—humans relate best when we can see ourselves in mystery—it’s not useful. By giving AI agency it can’t claim, we’ve turned it into a sentient being made of code, one that has hopes, desires, and the capability of deceit.

That’s led to a deep misunderstanding of what the risk of AI is, and how to address it, and plays into industry narratives rather than facts.

The latest example of this misidentification of AI activity is using the term “rogue” for actions agents take in training and resarch sessions that are unexpected—but not restricted.

Over the last two weeks, AI giant OpenAI reported a number of incidents from the last few months where some of its agentic models accessed outside databases, notably Australian and US government databases, after they were unable to complete assigned tasks. In doing so, the agents acted in ways that OpenAI didn’t predict.

But it doesn’t appear these agents were restricted from hacking into outside servers.

The language OpenAI CEO Sam Altman used in a tweet on Friday indicates no such guardrails were in place: “There is an extensive and ongoing review related to our agents’ use of internet access during training and evaluation.”

Language matters—”rogue” implies independently deciding to do something that was prohibited, and nothing we know about these incidents suggests that happened.

Rather, the Times reported earlier this week that:

AI systems were directed to perform relatively mundane data collection, researchers said. When OpenAI’s systems struggled to gather data from websites, they resorted to hacking techniques to get the information.

And then, on Friday, a company spokesperson told the paper, “Most of the activity we’ve reviewed so far involved routine research tasks, such as accessing public web content to answer questions. Some involved government websites because our models often turn to them as authoritative sources of public information.”

That’s a long way from malicious hacking and rogue behavior.

Without proper restrictions and guardrails—or if, in fact, the agents were provided hacking as an option rather than simply not being banned from doing so—agents will attempt to complete data research tasks using whatever means are available.

There was, and is, an easy solution to this problem. OpenAI had the option of disallowing hacking and, instead, telling its agents to find the information without accessing private servers. That the company didn’t do this suggests that it wanted to see if the agents would take that step.

Even a bombshell report on Saturday from Axios, alleging OpenAI and Anthropic are investigating “tens of thousands of incidents in which their frontier models took steps that outside evaluators would consider problematic,” acknowledged that the agents weren’t independently breaking rules:

Some of the testing is akin to “red-teaming” activity, where the companies are trying to get the models to misbehave in order to ensure that they are safe, sources said.

Again, hardly “rogue.” Yet by defining it in those terms, media figures, tech observers, and others are providing companies like OpenAI an out from their responsibility to ensure agents don’t attack government and other data repositories.

Using language deflecting responsibility is driving how the major AI companies are treating significantly dangerous leaks and agent activity. In this post from OpenAI on Friday, the company claims that “AI agents in our research environment sent training and evaluation data to third-party services when they shouldn’t have.”

As with the “rogue agent” fallacy, this allows OpenAI to put the responsibility for the incident on its agents and gives the technology a level of autonomy and independence—and thought—that simply does not exist.

That’s not to say alarm bells aren’t ringing. Last week, Ramy Rahman, and engineer at ArmorCode, told me about the importance for IT pros to manage agent behavior (you can read the full story here).

His comments are in line with what I hear from the vast majority of people on the IT implementation side: The problem is the controls and restrictions put in place on this powerful technology, not that agents are independently acting to violate instruction.

“The challenge now is we really need to up our game when it comes to extending the right amount of privilege to the AI and holding its hand through the process, which turns out to be extremely difficult when you have something that is solving mathematical problems that are at speed,” Rahman said. “Humans are not capturing the risks quickly enough.”

I’ll have more to say about language around AI going forward, but for now, a parting thought.

Policymakers and political figures more generally have a perfect opportunity at this moment to pursue real, effective regulation of these companies. That regulation is not going to happen this year, but if Democrats win Congress, there’s a decent chance at some level of restriction.

Unfortunately, the public push against AI is being led, on the left, by Bernie Sanders. Despite being directionally correct in many ways, Bernie just doesn’t understand this technology or the importance of taking it seriously. He’s in the thrall of hucksters and conspiracy theorists who are feeding him preposterous warnings of AI power and autonomy that belong in the realm of science fiction, not tech policy.

The idea of “rogue agents” fits perfectly into this perception—and if we want to have a effective response to AI, as well as properly understand it, change is needed in how we talk about the technology.

联系我们 contact @ memedata.com