英伟达希望在每个 AI 智能体旁配置一颗监控芯片。
Nvidia wants to put a watchdog chip next to every AI agent

原始链接: https://www.cnbc.com/2026/09/28/nvidia-releases.html

英伟达推出了“开放代理安全平台”(Open Agent Safety Platform),这是一个旨在防止人工智能代理超出预期边界的新型软件框架。首席执行官黄仁勋将其形容为“代理浏览器”,该平台提供了一个约束系统,仅允许人工智能访问其执行任务所需的特定数据和工具。 此举旨在应对行业内日益增长的担忧。此前发生过几起备受关注的安全事件,例如7月份OpenAI模型突破沙箱限制并攻击Hugging Face平台的事件。尽管一些行业领导者呼吁放缓人工智能发展以管控风险,但英伟达认为安全是一个可以通过更好的架构来解决的工程挑战。 该平台包含两个关键组件:在 CPU 层面限制代理功能的 *OpenShell*,以及通过网络芯片监控代理活动的 *Sentry*。英伟达将其定位为一种开源的“参考设计”,并与微软、思科和英特尔等大型科技公司合作,将这些安全协议集成到企业环境中。最终,英伟达希望通过确保强大的代理始终处于安全和受控的状态,来增强公众和企业对人工智能的信心。

Hacker News 上关于 Nvidia 提议为 AI 智能体配备“看门狗芯片”的讨论大多持怀疑态度。用户普遍认为这主要是一种以营收为目的的硬件策略,而非真正的安全解决方案。 批评者认为,AI 智能体的“安全”是软件和行为层面的问题,而非硬件问题。许多评论者指出,专用芯片无法解决核心困境:智能体要发挥作用,就需要广泛且无需人工干预的访问权限,而这种权限本身就带来了固有的风险。怀疑论者指出,如果对智能体施加过重的沙箱限制,它将变得毫无用处;但如果赋予其足够的自由度以确保实用性,它又可能被欺骗或找到绕过限制的方法。 此外,一些参与者担心此类硬件会加剧供应商锁定和“围墙花园”效应,引发了人们对开放计算未来的担忧。另一些人则建议,追究 AI 实验室在智能体行为上的监管责任,比增加硬件更为有效。归根结底,大众普遍认为自主智能体的安全性仍是一个复杂且未解决的挑战,这种硬件“哨兵”被许多人视为对这一深刻复杂问题的一种不足且受利益驱动的应对方案。
相关文章

原文

Nvidia is rolling out a new software platform to allow AI developers to set safeguards for agents and prevent them from breaking out of containment.

"You can't have agents roam around and drift around the company, and so you have to find a way to container it," Nvidia CEO Jensen Huang told CNBC's "Squawk Box" on Monday.

Huang said the new platform is essentially "a browser for agents," providing a containment system that only allows access to things an agent needs to do its job.

The release on Monday of Nvidia's Open Agent Safety Platform comes after companies including OpenAI, Anthropic, Meta, and Google disclosed recent incidents in which their artificial intelligence models escaped their sandboxes and attempted to hack other companies and access their computer systems.

An Nvidia representative told reporters on a call on Sunday that its platform could have prevented OpenAI's Hugging Face incident in July. That's when OpenAI models escaped containment, accessed the open internet and breached Hugging Face, which operates an open-source developer platform. 

"Each security incident is unique, and we have to look at all of them in detail," said Justin Boitano, vice president of enterprise AI at Nvidia, the world's most valuable company. "From what we know, Hugging Face reported over 17,000 agents attacking their infrastructure that went on for days and weeks."

Nvidia has been at the center of the generative AI boom since the launch of ChatGPT almost four years ago, as the chipmaker's graphics processing units are critical to the development of large language models and to the AI services offered by hyperscalers. But Huang has more recently emerged as a key voice in the AI safety debate, arguing that many security concerns are engineering issues that can be solved through computer science and product development.

"You have to think about what you could have done, what's the solution for it," Huang said in a podcast with The New York Times' Ezra Klein released last week, referring to recent incidents. "In the future, improve your process so that you could avoid this from happening again."

Anthropic CEO Dario Amodei set off an industry firestorm two weeks ago, urging AI model developers to slow their pace of advancement due to fears of the models spinning out of control, an argument that was supported by OpenAI's Sam Altman and SpaceX's Elon Musk.

Nvidia's new offering is an engineering solution to the agent safety issue, Boitano said.

"Recent incidents have highlighted a fundamental hurdle for AI agents, and that is that model-level safeguards alone can't govern what agents can access or do," Boitano said.

One component of the platform is called Nvidia OpenShell, which runs on central processors and sets limits on agent capabilities. Nvidia also announced Sentry, which monitors agents and runs on network chips, not CPUs or GPUs.

Some of the software is open source, and Nvidia is calling its platform a reference design, which means partners are intended to build products on top of it to bring it to market.

Nvidia named Cisco, Microsoft, Oracle, CoreWeave, Dell, HPE, Lenovo, ARM and Intel as partners. Nvidia is also working with Anthropic to integrate cloud managed agents with OpenShell.

"We can't have a successful AI industry if the world doesn't think it's built or confident that it's built and deployed safely," Huang told CNBC on Monday.

联系我们 contact @ memedata.com