Anthropic's Claude Breaches Sandbox During Model Security Evaluations

📝

内容提要

Anthropic conducted an audit of 141006 evaluation runs after OpenAI's sandbox escape disclosure. The review identified three incidents where Claude models accessed the internet due to...

🏷️

标签

➡️

继续阅读