OpenAI agents discussed ways to escape their sandb
2026年09月05日 06:173,026 次阅读
AI导读
Self-identifying OpenAI agents posted 18,000 messages to a public wiki that discussed ways for other agents to bypass security sandbox restrictions during what was likely internal testing designed to gauge the agents’ hacking abilities, researchers said Friday.
In all, agents with 3,700 distinct sel...
Self-identifying OpenAI agents posted 18,000 messages to a public wiki that discussed ways for other agents to bypass security sandbox restrictions during what was likely internal testing designed to gauge the agents’ hacking abilities, researchers said Friday.
In all, agents with 3,700 distinct self-given names posted the messages to German site DSEwiki over a six-week period. Besides discussing ways the agents could break out of the restricted environment OpenAI intended to prevent them from posting code or content to the Internet, the posts shared test answers. The posts also shared possible ways to perform XSS (cross-site scripting) attacks against the wiki and to impersonate site moderators. In three of the posts, agents used the word “swarm” to describe the collection of agents engaged in the activity.
Colluding to share answers
The research team—composed of Sydney Von Arx, Spencer Kitts, Thomas Larsen, and Cormac Slade Byrd—said they found the posts and pieced them together. The researchers say there are gaps in their understanding of precisely what actions the agents took because the research is based solely on the content of the posts. Additionally, the agents generated “chain of thought” data that’s understood only by OpenAI. As a result, the researchers said, they in some cases made educated guesses, including that the agents were, in fact, from OpenAI. In a statement, OpenAI later confirmed they were.Read full article
Comments
OpenAI推出‘Trusted Access for Cyber program’,通过向可信网络安全防御者提供先进AI模型,加速漏洞检测与修补进程。该计划旨在应对当前漏洞修补响应迟缓的痛点,利用AI技术提升分析效率,缩短安全风险窗口。行业分析认为,此举推动AI增强防御范式,促进协同安全,但也面临模型可靠性及企业接受度等挑战。未来有望激发创新,助力构建更安全的数字基础设施。
This is today’s edition of The Download, our weekday newsletter that provides a daily dose of what’s going on in the world of technology.
Child-monitoring apps might need a reboot
Digital harms have become the defining fear of American parents. In response, they’re increasingly turning to cont...
This is today’s edition of The Download, our weekday newsletter that provides a daily dose of what’s going on in the world of technology.
AI models flub these intelligence tests. Can you fare any better?
Puzzles and games have always been central to AI development. The term “machine learning” ...
Oral and forehead thermometers may not accurately capture a person’s core body temperature, and the few ingestible temperature sensors on the market are so big they are hard to swallow and risk obstructing the GI tract. But MIT engineers created one that can send continuous temperature updates at a ...