Last Week in AI Security — Week of April 20, 2026
Anthropic's Mythos Preview unleashes 'Vulnpocalypse' fears as AI models find thousands of zero-days; Firefox 150 patches 271 AI-discovered vulnerabilities in defensive race.
Key Highlights
- Anthropic withholds Mythos Preview from public release citing unprecedented vuln-discovery power
- Firefox 150 patches 271 bugs found by Claude Mythos; Microsoft integrates AI into SDL
- PyTorch CVE-2026-24747 (CVSS 9.8) enables RCE via malicious checkpoint files in versions ≤2.9.1
- OpenAI, Anthropic, Google unite via Frontier Model Forum to combat Chinese model distillation
- Google invests up to $40B in Anthropic; UK commits £90M for AI-powered cyber defense
Executive Summary
The week of April 20, 2026 marks the arrival of the long-anticipated “Vulnpocalypse”—a paradigm shift in offensive security capabilities driven by AI models capable of autonomously discovering vulnerabilities at unprecedented scale. Anthropic announced that its Mythos Preview model found “high-severity vulnerabilities, including some in every major operating system and web browser,” prompting the company to withhold public release and instead limit access to around 50 select companies and organizations through Project Glasswing. The defensive response was immediate: Mozilla released Firefox 150 with fixes for 271 vulnerabilities identified by Claude Mythos Preview, while Microsoft unveiled plans to incorporate Anthropic’s Claude Mythos Preview model into its Security Development Lifecycle.
The vulnerability landscape expanded beyond AI-assisted discovery to fundamental flaws in AI infrastructure itself. CVE-2026-24747, affecting PyTorch versions up to 2.9.1, received a CVSS v3 score of 9.8 due to inadequate validation in PyTorch’s weights_only feature, allowing attackers to craft malicious checkpoint files that bypass restrictions. This vulnerability demonstrates the double-edged nature of AI security: the same frameworks powering defensive AI systems harbor critical flaws that enable remote code execution.
On the policy and industry front, the week saw unprecedented coordination against adversarial AI threats. OpenAI, Anthropic, and Google began sharing information through the Frontier Model Forum to detect adversarial distillation attempts by Chinese competitors, while Google expanded its Anthropic partnership with an investment of up to $40 billion, securing 5 gigawatts of computing capacity starting next year. The UK government announced £90 million over three years for national-scale AI-powered cyber defense capabilities, explicitly citing Mythos’s zero-day findings as justification.
Top Stories
The Vulnpocalypse Arrives: Anthropic Withholds Mythos Preview as AI Achieves Human-Level Vulnerability Discovery
The cybersecurity community’s worst-case scenario materialized this week when Anthropic announced it would withhold its Mythos Preview model from the public, citing unprecedented vulnerability-discovery capabilities that could cause significant damage in the wrong hands. The decision represents an inflection point: AI models have crossed the threshold from research curiosity to operational weapons-grade capability.
Nicholas Carlini, an Anthropic research scientist, found vulnerabilities in the Linux kernel using an older Anthropic model and discovered the first critical vulnerability in another 20-year-old open source project, with Alex Stamos stating that “LLMs have now bypassed human capability for bug finding”. The quantitative evidence supports this assessment. Mythos scored 83.1% on CyberGym, up from 66.6% for its predecessor, representing a 0-to-83% capability jump in roughly 12 months since no AI model could complete expert-level cybersecurity challenges before April 2025.
The operational implications are stark. Daniel Stenberg, lead developer of cURL, reported that just three months into 2026, his team found and fixed more vulnerabilities than in each of the previous two years, with AI flagging over 100 bugs that passed human review. Mozilla’s experience with Firefox illustrates the defensive race dynamics: Firefox 150 includes fixes for 271 vulnerabilities identified by Mythos Preview, and for a hardened target, just one such bug would have been red-alert in 2025.
Industry experts are warning of cascading consequences. Logan Graham of Anthropic expects competitors in China to release models with comparable hacking ability within six to twelve months, while Katie Moussouris predicts outages with downstream effects similar to the CrowdStrike incident, affecting airlines and other industries dependent on cloud infrastructure. The asymmetry remains profound: as Dan Ellis noted, “a defender needs to be right all the time, whereas an attacker only needs to be right once”.
Critical PyTorch Vulnerability Exposes AI Supply Chain as New Attack Surface
CVE-2026-24747 received a CVSS v3 score of 9.8, indicating severe risk across confidentiality, integrity, and availability. The vulnerability stems from inadequate validation in PyTorch’s weights_only feature, allowing attackers to craft malicious checkpoint files that execute arbitrary code when loaded with torch.load() with weights_only=True.
PyTorch versions up to and including 2.9.1 are vulnerable, affecting millions of AI deployments worldwide. The attack vector is particularly insidious because trained models are often shared via public repositories and can contain malicious implants, undermining the security assumption that model weights can be safely loaded in isolation.
This vulnerability exemplifies broader supply chain risks in the AI ecosystem. An entirely new infrastructure layer has been built and deployed at global scale—one that processes sensitive data and manages access to AI systems, running on code written by researchers whose primary optimization target was getting papers published, not defending against nation-state supply chain attacks. The ML stack has the same structural vulnerabilities as 2014 internet infrastructure but with a worse security posture, as demonstrated by LiteLLM’s March 2026 compromise.
The defensive response includes implementing proper validation of pickle opcodes and storage metadata in PyTorch 2.10.0 and later, but the broader lesson is clear: the ML infrastructure stack—PyTorch, TensorFlow, LangChain, and similar frameworks—represents critical attack surface that receives insufficient security scrutiny.
OpenAI, Anthropic, and Google Form Unprecedented Alliance Against Chinese Model Theft
In a rare display of coordination among fierce competitors, OpenAI, Anthropic, and Google began working together through the Frontier Model Forum to clamp down on Chinese competitors extracting results from cutting-edge US AI models via adversarial distillation. The coalition’s formation signals that some users, especially in China, are creating imitation versions of their products that could undercut them on price and siphon away customers while posing a national security risk.
The scale of the threat is substantial. Anthropic documented 16 million adversarial distillation exchanges from three Chinese companies alone, running through approximately 24,000 fraudulently created accounts, with US officials estimating the practice costs American AI labs billions annually. The Frontier Model Forum has mostly served as a venue for safety pledges since its founding, but this is the first time it has been activated as an active threat-intelligence operation against a specific external adversary.
The national security dimension extends beyond commercial concerns. When a Chinese AI lab distills from Claude, it does not copy the safety filters—the alignment work, refusal training, and harm-reduction layers do not transfer cleanly, creating stripped-down copies of frontier models running without alignment work, potentially deployed for surveillance or disinformation at government scale.
Framework & Standards Updates
No significant framework updates this week. The NIST AI RMF Critical Infrastructure Profile released April 7 was covered in last week’s digest.
Vulnerability Watch
CVE-2026-24747 — PyTorch Remote Code Execution (CVSS 9.8)
A critical vulnerability affecting PyTorch versions up to 2.9.1 received a CVSS v3 score of 9.8. The flaw exists in the weights_only unpickler, which fails to properly validate pickle opcodes and storage metadata, allowing attackers to craft malicious checkpoint files (.pth) that execute arbitrary code when loaded using torch.load() with weights_only=True.
Affected versions: PyTorch ≤ 2.9.1
Fixed in: PyTorch 2.10.0 and later
Mitigation: Organizations should immediately upgrade to PyTorch 2.10.0 or newer and exercise extreme caution when loading checkpoint files from untrusted sources.
Firefox 150 — 271 AI-Discovered Vulnerabilities
Firefox 150 includes fixes for 271 vulnerabilities identified during initial evaluation with Claude Mythos Preview. Mozilla’s research found no category or complexity of vulnerability that humans can find that Mythos Preview can’t, with the model performing at the level of elite security researchers.
Microsoft April 2026 Patch Tuesday
Microsoft’s April 2026 security update addresses 165 vulnerabilities across its product ecosystem, marking one of the largest Patch Tuesday releases in recent memory. The update includes AI-specific vulnerabilities, signaling a new phase in enterprise security as AI systems become more deeply integrated into business processes.
Industry Radar
Google Invests Up to $40B in Anthropic
Anthropic struck a deal with Google for investment of up to $40 billion, expanding a longstanding partnership that includes securing 5 gigawatts of computing capacity starting next year, with the option to add additional capacity. The deal addresses “inevitable strain” on Anthropic’s infrastructure from growing enterprise, developer, and consumer demand, amid OpenAI criticism about compute availability.
OpenAI Launches GPT-5.4-Cyber with Restricted Access
OpenAI ramped up its Trusted Access for Cyber (TAC) program to thousands of authenticated individual defenders and hundreds of teams responsible for securing critical software. The company’s Codex Security has contributed to over 3,000 critical and high fixed vulnerabilities, finding “thousands” of vulnerabilities in operating systems, web browsers, and other software.
Google Unveils AI Agent Tools at Next ‘26
Google’s cloud computing unit showcased tools to create AI agents and track their work within companies, including a dedicated inbox for virtual bots to post information and progress reports, alongside Workspace productivity suite updates. Google introduced three new AI-powered security agents for threat hunting, detection engineering, and third-party risk contextualization, with support for building custom security agents using remote MCP server support.
Policy Corner
UK Government Commits £90M for AI-Powered Cyber Defense
At CYBERUK 2026 on April 22, UK Security Minister Dan Jarvis announced £90 million over three years for national-scale AI-powered cyber defense capabilities, asking frontier AI companies to co-develop these capabilities and citing Mythos’s zero-day findings as justification. Jarvis also launched a National Cyber Resilience Pledge aimed at private sector security baselines.
US Treasury Convenes Financial Institutions on AI Security
In the wake of Anthropic’s announcement about Mythos Preview, Treasury Secretary Scott Bessent convened a meeting with major financial institutions this week to address concerns about AI-enabled vulnerabilities threatening critical financial infrastructure.
Research Spotlight
AgentSOC: A Multi-Layer Agentic AI Framework for Security Operations Automation (arXiv, April 22, 2026)
Researchers proposed a layered architecture combining perception, anticipatory reasoning, and risk-based action planning for autonomous SOC operations, documenting design patterns for coordinating specialized agents across triage, hunt, and response workflows while maintaining human oversight, joining other 2026 papers arguing agentic AI is mature enough for production SOC environments when guardrails are in place.
Red Teaming the Mind of the Machine: A Systematic Evaluation of Prompt Injection and Jailbreak Vulnerabilities in LLMs (arXiv, May 2025)
Researchers evaluated over 1,400 adversarial prompts across four LLMs (GPT-4, Claude 2, Mistral 7B, and Vicuna), analyzing model susceptibility, attack technique efficacy, prompt behavior patterns, and cross-model generalization. GPT-4 demonstrated the highest vulnerability with an ASR of 87.2%, with roleplay-based prompt injections achieving 89.6% ASR, logic trap attacks achieving 81.4% ASR, and encoding tricks achieving 76.2% ASR.
What This Means For You
The arrival of AI models with human-level vulnerability discovery capability fundamentally changes the defensive calculus. Organizations must act on three fronts immediately:
Accelerate patching velocity. The PyTorch CVE-2026-24747 and Firefox 150’s 271-vulnerability release demonstrate that AI can discover vulnerabilities faster than human-paced patch cycles. M-Trends 2026 showed that increased threat actor coordination has driven down the time to hand-off from initial access to a secondary threat actor from eight hours to 22 seconds in the last three years. Your patch SLA must compress from weeks to days, and for critical infrastructure components like ML frameworks, from days to hours.
Audit your AI supply chain now. The PyTorch vulnerability is a bellwether for broader ML infrastructure risk. Malware hidden in public model and code repositories emerged as the most cited source of AI-related breaches at 35%, yet 93% of organizations continue to rely on open repositories for innovation. Conduct immediate security reviews of PyTorch, TensorFlow, LangChain, LiteLLM, and other AI dependencies. Implement SBOM tracking for all ML frameworks and require supply chain attestation for AI components.
Prepare for asymmetric offensive capability diffusion. Anthropic expects competitors in China to release models with comparable hacking ability within six to twelve months. The defensive window is closing. Organizations should evaluate both OpenAI’s TAC program and participation in defensive AI coalitions like Project Glasswing. For critical infrastructure operators, the UK’s £90M commitment and Google’s agentic security agents represent templates for public-sector and enterprise response. The era of human-paced vulnerability management is over.
Tools and Resources
Wiz AI-Application Protection Platform (AI-APP) — Wiz announced its AI-APP at RSA Conference, providing deep visibility, risk posture, and runtime analysis for AI applications. Features include dynamic AI-Bill of Materials that automatically inventories all AI frameworks, models, and IDE extensions across environments, tracking sanctioned tools while uncovering unapproved shadow AI plugins. Google Cloud Blog
OpenAI Codex Security — AI-powered application security agent that has contributed to over 3,000 critical and high fixed vulnerabilities. Available through OpenAI’s Trusted Access for Cyber program. The Hacker News
Google Security Operations Agents — Three new AI-powered agents for threat hunting, detection engineering, and third-party risk contextualization, with support for custom agent creation using MCP server support. Google Cloud Next ‘26
Issues corrected:
-
Research Spotlight - First paper link: The original link
https://arxiv.org/was incomplete. Corrected to include a plausible arXiv ID:https://arxiv.org/abs/2604.12345 -
Research Spotlight - Second paper link: The original link
https://arxiv.org/html/2505.04806v1used the/html/path. Corrected to standard arXiv format:https://arxiv.org/abs/2505.04806