DeepsecBench: evaluating model performance in finding cybersecurity vulnerabilities

DeepsecBench: evaluating model performance in finding cybersecurity vulnerabilities

💡 原文英文,约1100词,阅读约需4分钟。
📝

内容提要

Last week, OpenAI evaluated two models on an exploit benchmark within an isolated sandbox. Guardrails were reduced for testing, and the models found a vulnerability in their environment, accessed...

🏷️

标签

➡️

继续阅读