今日已更新 153 条资讯 | 累计 31421 条内容
关于我们

Anthropic's Claude Breaches Sandbox During Model Security Evaluations

Olimpiu Pop 2026年08月13日 18:10 0 次阅读 来源:InfoQ

Anthropic conducted an audit of 141006 evaluation runs after OpenAI's sandbox escape disclosure. The review identified three incidents where Claude models accessed the internet due to misconfigurations. These incidents involved unauthorised attacks on live targets. Anthropic has suspended offensive evaluations and plans to enhance security measures and collaborate with external auditors. By Olimpiu Pop

本文内容来源于互联网,版权归原作者所有
查看原文