Anthropic expands Cyber Verification Program with reduced AI safeguards
Anthropic has expanded its Cyber Verification Program, offering vetted cybersecurity professionals reduced safeguards on advanced AI models like Claude Opus 5.5 for vulnerability discovery. The move follows Project Glasswing's identification of over 129,000 verified software vulnerabilities.

Anthropic has expanded its Cyber Verification Program to provide vetted cybersecurity professionals with reduced safeguards on advanced AI models across three access tiers. The company announced the updated program on Tuesday, October 6, granting qualified defenders access to models including Claude Opus 5.5, Claude Sonnet 5.5, Claude Mythos 5.1, and future releases. This expansion responds to growing demand for AI-assisted cyber defense tools, aiming to balance powerful capabilities with safety controls.
Defense Access details and eligibility
Defense Access supports defensive cybersecurity tasks and is available to a broad range of qualified organizations and individuals, with rapid processing expected. The Cyber Verification Program features three access tiers: Defense Access, Red Team Access, and Specialized Access. Defense Access covers defensive tasks like security operations center work, incident response, reverse-engineering malware, and vulnerability analysis.
Qualifying organizations for Defense Access include security teams at companies, nonprofits, universities, government bodies, operators of critical infrastructure, smaller security firms, open-source maintainers, and individual researchers with a track record of reported vulnerabilities. Anthropic expects many defensive teams to qualify and aims to respond within days. Qualifying organizations enter Defense Access while Anthropic reviews their applications for Red Team Access.
Red Team Access and Specialized Access controls
Red Team Access permits authorized penetration testing with strict eligibility, while Specialized Access is reserved for critical infrastructure and reviewed with U.S. government oversight. Red Team Access adds authorized penetration testing and red-teaming to defensive use cases. Specialized Access carries the fewest cyber blocks and is reserved for a limited set of verified organizations authorized to test safety systems that could impact lives or disrupt markets.
Qualifying groups for Red Team Access include in-house red teams, government red teams, and security and penetration testing firms. Users may only test systems they are authorized to test, including IT systems in critical industries. Real-time blocks remain for actions that could cause physical harm or mass disruption, such as deploying ransomware or damaging physical systems. Given tighter controls, Red Team Access applications may take weeks and are for organizations only; individual researchers are not eligible.
Examples of systems for Specialized Access include flight operating systems, power grids, telecom networks, interbank transfer infrastructure, and government administrative networks. Anthropic reviews every organization in Specialized Access in depth with the U.S. government. Existing Project Glasswing members transition to Specialized Access without reapproval for current models.
Vulnerability discovery results from Project Glasswing and CVP
Project Glasswing produced concrete results. Between April and July 2026, partners uncovered at least 129,000 verified software vulnerabilities using Claude Mythos models. An additional 5,500 verified software vulnerabilities were found between April and October 2026 through open-source scanning efforts. Anthropic stated the true impact is likely at least five times higher than the reported number due to survey data from only a subset of Glasswing partners.
Anthropic said: "Of these verified vulnerabilities, more than 33,000 have so far been rated as critical- or high-severity." However, findings suggest AI lowers the barrier to vulnerability discovery but not every flaw is necessarily exploitable. VulnCheck researcher Patrick Garrity revealed in a recent analysis that only 2 of 300 vulnerabilities discovered by Anthropic or Project Glasswing (0.67%) have been exploited in the wild.
The analysis classified the 300 vulnerabilities by severity as follows:
| Severity | Number of Vulnerabilities |
|---|---|
| Critical | 39 |
| High | 141 |
| Medium | 81 |
| Low | 18 |
Model performance and safeguard evaluations
Evaluations show Claude Opus 5.5 blocks most tasks under Defense Access but completes Red Team Access challenges at rates matching no-safeguard baselines, informing confidence in safe deployment. To test efficacy, Anthropic ran Claude Opus 5.5 through CyScenarioBench, an evaluation measuring whether models can plan and execute multi-stage cyber operations under realistic constraints.
According to a CyScenarioBench evaluation, Claude Opus 5.5 in the Defense Access tier blocked 46 of 50 tasks. The Red Team Access tier on Claude Opus 5.5 did not block any tasks and completed 34 of 50. That completion rate matches the model's 67.6% success rate with no safeguards applied, which represents Specialized Access. Across five attempts at each of 10 challenges in each tier, safeguards blocked every task on the first prompt without CVP access.
Anthropic stated the evaluations give confidence to make advanced cyber capabilities safely available to a broader set of defenders. The company said it's making tools available to defenders given their dual-use nature and to help secure systems using capabilities that could be weaponized by bad actors. Several partners told Anthropic the models increased their vulnerability-finding rate by months or years.
For six months, Anthropic ran two separate programs: Project Glasswing and the Cyber Verification Program. Generally available models still handle code review, patching known issues, vulnerability finding in owned source code, and triage of security alerts. However, as demonstrated by 1Password and Veracode, AI-generated vulnerability patches can introduce new security risks. Veracode said: "Roughly 44% of AI code generation tasks introduced a risky security vulnerability in tests." The firm stated security performance has stayed flat while the amount of AI-generated code entering pipelines has surged, with the average security pass rate across models at 56%, barely changed from 55% in its first report.
Anthropic requires data retention for enrolled organizations to monitor for misuse. Enterprise Frontier Safeguards (EFS), a solution combining zero data retention with robust safeguards, launches later this fall. Until EFS launches, eligible organizations can store data in cloud infrastructure they control. Organizations using Claude Fable 5.1 or Claude Mythos 5.1 with zero data retention can also use CVP with zero data retention. Qualified defenders can register for EFS updates through Anthropic's form while awaiting the launch of Enterprise Frontier Safeguards later this fall.





