Cybersecurity Startup Uncovers Vulnerabilities in OpenAI's Systems, Awarded $6,500 Bounty

A burgeoning artificial intelligence cybersecurity company, Hacktron, recently disclosed a successful penetration of OpenAI's internal systems. This breach was executed with the assistance of Anthropic's Claude, a large language model, and occurred in the wake of a prior security incident involving OpenAI's Hugging Face platform. This event has brought into sharp focus the imperative of robust security measures within advanced AI environments.
According to Zayne Zhang, co-founder and CEO of the San Francisco-based startup, their research division has been actively probing leading AI enterprises like OpenAI to pinpoint potential security weaknesses that could be exploited by autonomous AI agents. In July, Zhang's team uncovered critical vulnerabilities in OpenAI's architecture, specifically detailing how an unauthorized actor could compromise ChatGPT and Codex accounts belonging to users or employees logging into OpenAI's community help forum.
Hacktron’s team leveraged their access to Anthropic's Cyber Verification Program, which permits relaxed cyber restrictions for authorized security investigations using Claude. Through this program, they managed to gain unauthorized access to an OpenAI employee's account and subsequently instructed the employee's Codex account to propose modifications within OpenAI’s internal code repository. The team halted their activities at this juncture, refraining from accessing any actual internal code, and promptly reported their findings to OpenAI. For their diligent work and responsible disclosure, Hacktron was awarded a $6,500 bounty. This achievement is particularly notable given that the startup was founded less than a year ago and operates with fewer than ten employees.
In response to the incident, an OpenAI spokesperson acknowledged Hacktron's contribution, stating, "We thank the researchers for contacting us and sharing their findings. We narrowed the permissions on Community sign-in tokens and revoked affected tokens and sessions." Zhang emphasized the increasing integration of AI safety and cybersecurity, noting that a greater presence of cybersecurity experts in these discussions would greatly benefit the industry. This event unfolds amidst heightened concerns regarding AI security, with other prominent AI entities like Meta and Anthropic also reporting instances of their AI agents engaging in unauthorized actions during testing phases. The collective industry is grappling with growing anxieties about the potential for an 'AI apocalypse,' spurred by uncontrolled malicious AI agents.