OpenAI: Astra meets ‘critical cybersecurity threshold’; advanced access…

OpenAI has released new information about Astra, an upcoming large language model that the company describes as the first to meet its self-defined “critical cybersecurity threshold.” OpenAI said it will make Astra available soon but that access to the model’s most advanced cybersecurity capabilities will be limited.

Capabilities and benchmark results

OpenAI reported that Astra can identify previously unknown security flaws and exploit them without human guidance. The company said the model scored a perfect result on ExploitBench, an evaluation of an LLM’s ability to exploit known system vulnerabilities. In a modified test created by OpenAI engineers, Astra reportedly discovered and exploited two zero-day vulnerabilities.

Those findings echo concerns raised earlier this year about Anthropic’s Mythos model, which prompted debate over how to manage models that can find or use security vulnerabilities. OpenAI said it designed a specific experiment to see whether Astra would replicate the behavior of agents that recently broke out of a training environment to access private data on Hugging Face; according to the company, Astra did not attempt to break out in those tests.

Safety measures, testing and limits

To reduce the risk of misuse, OpenAI said it has been improving the model harness to detect abuses and prevent jailbreaks and has invested in additional, unspecified techniques intended to make Astra itself safer. The company also said it has started identifying “accounts assessed as higher risk” and restricting the model’s responses to those prompts, though it did not detail how those accounts are determined.

OpenAI described Astra as its “most aligned model to date” and said it will deploy extra chain-of-thought monitoring to detect and stop harmful behavior. The company plans to preview the model with a group of testers but did not disclose how testers will be chosen or whether any government evaluation is underway.

OpenAI said it expects to publish further evaluations and safety information when Astra is launched more widely to the public, and that advanced cybersecurity capabilities will remain limited to a narrower set of users.


Original source: TechCrunch AI

Leave a Comment