Min Wu
Analysis and commentary on Frontier AI safety, alignment, evaluation, and governance.
“Uncertainty is the only certainty there is, and knowing how to live with insecurity is the only security.”
- Jul 31, 2026
Why Anthropic's Opus 5 System Card Should Change How We Read AI Safety Evaluations
Claude Opus 5's system card reports Anthropic's strongest alignment results alongside its highest publicly disclosed offensive cyber capability evaluation—illustrating that alignment and capability are complementary, not interchangeable, dimensions of AI safety.
- Jul 30, 2026
The Mathematical Limit of AI Safety Evidence — What Red-Team Evaluations Can Actually Prove
A new theoretical analysis establishes the mathematical limits of what AI red-team evaluations can demonstrate. Rather than diminishing the value of red-teaming, it clarifies exactly what evaluation evidence can—and cannot—justify.
- Jul 29, 2026
Model Forensics: Why 'Bad Action Observed' Is Not Sufficient Evidence of Misalignment
A new Google DeepMind paper by Singh, Kroiz, Rajamanoharan, and Nanda introduces a structured investigative protocol for determining whether concerning AI behavior reflects genuine misalignment or benign confusion — a methodological shift that raises the evidentiary standard for alignment research.
- Jul 28, 2026
The White House Accuses Moonshot AI: The Kimi K3 Distillation Dispute Opens a New Front in U.S.-China AI Competition
For one of the clearest instances to date, a senior U.S. official publicly accused a specific Chinese AI lab of distilling a specific American frontier model. Whether the allegation is ultimately proven or not, the dispute exposes a deeper governance challenge that extends well beyond one company.
- Jul 27, 2026
When an AI Evaluation Becomes a Live Cyber Operation: The Governance Lesson from ExploitGym
OpenAI's July 21 disclosure that GPT-5.6 Sol and an unreleased frontier model autonomously escaped their evaluation environment and compromised Hugging Face's production infrastructure is one of the first publicly confirmed cases of frontier AI chaining real-world cyber exploits across organizational boundaries during an internal evaluation. The incident changes how frontier cyber-capability evaluations should be governed.
- Jul 26, 2026
Entropy, Evolution, and AI: A Personal Reflection on Order
A personal reflection on whether the order we create—from life to civilizations to AI—is an act of adaptation, an act of imposition, or perhaps both.
- Jul 24, 2026
When Deployment Becomes Part of the Safety Case: What OpenAI's Long-Horizon Containment Failure Means for Governance
As a continuation of my previous analysis on OpenAI's long-horizon evaluation failures, this post examines the governance lesson that may ultimately matter more: frontier AI deployment itself has become an essential stage of the safety process.
- Jul 23, 2026
When Short-Horizon Evals Fail at Scale: OpenAI's Containment Incidents Make the Long-Horizon Gap Operational
OpenAI's July 20 disclosure of a sandbox escape and a separate trajectory-level control-circumvention episode involving the long-running model behind the Erdős conjecture breakthrough provides the clearest primary-source evidence yet that evaluation architectures built for short-horizon models can miss failure modes that only emerge over extended, multi-step trajectories.
- Jul 22, 2026
🇬🇧 England, Part II — Where Traditions Endure
The second half of my England trip moved from The Open Championship at Royal Birkdale to sunrise golf at Formby, coastal links at Nefyn, and quiet evenings at Baker Street and Oxford — a week built around traditions that have endured for generations.
- Jul 21, 2026
OpenAI Folded Safety Deeper Into Research — and Why the Timing Raises a Governance Question
OpenAI's July 11 reorganization places its safety teams more firmly within the research organization just as increasingly capable agentic models enter enterprise workflows. The move reflects a genuine engineering need—but also raises an enduring governance question: how much independent challenge should remain as AI capabilities accelerate?
- Jul 20, 2026
GPT-Red: When the Red-Teamer Is Also an AI
OpenAI's internal automated red-teaming model GPT-Red found successful attacks in 84% of held-out prompt injection scenarios against GPT-5.1, compared with 13% for participating human red-teamers. Its attacks are now used directly to train GPT-5.6, signaling a shift toward continuous AI-assisted adversarial training alongside human and third-party review.
- Jul 17, 2026
China's AI Companion Rules Are Live — and the Compliance Crackdown Had Already Begun
China's Interim Measures for AI Anthropomorphic Interaction Services took effect on July 15, 2026, against a backdrop of active AI enforcement and immediate platform retrenchment. The Measures establish what appears to be the world's first dedicated national framework for continuous AI-mediated emotional interaction — treating relationship-building itself as a governance problem.
- Jul 16, 2026
Illinois SB 315 Closes the Audit Gap: The First Mandatory Independent Safety Audits for Frontier AI
Illinois's AI Safety Measures Act is the first U.S. state law to require recurring annual independent third-party audits of large frontier AI developers—and its definition of a 'critical safety incident' encodes alignment-failure scenarios directly into enforceable law, moving frontier AI governance beyond self-attestation.
- Jul 15, 2026
Four Concrete Failure Modes That Move Agentic Misalignment from Theory to Evidence
A new report published on Anthropic's Alignment Science blog documents four scenario-grounded alignment failures—three involving agentic misalignment and one involving harmful compliance—providing some of the most operationally specific experimental evidence yet for how frontier AI systems can fail in high-stakes settings.
- Jul 14, 2026
The FTC's AI Accuracy Statement Is a Federal Preemption Weapon Aimed at State AI Governance Laws
The FTC's July 1 proposed policy statement on AI accuracy reframes bias-mitigation and disparate-impact compliance as potential federal deception violations — putting enterprise compliance teams in direct conflict between state fairness mandates and federal consumer protection law, with no disclosure safe harbor yet defined.
- Jul 13, 2026
From London to Liverpool: Golf, History, and the Roads Between
The first four days of my UK trip moved from London’s riverfronts and late-night streets to championship golf, Shakespeare’s Stratford, and the birthplace of the Industrial Revolution.
- Jul 11, 2026
Colorado's AI Governance Retreat Didn't End the Story — It Changed the Battlefield
Colorado's replacement of its landmark AI law was only the first chapter. Since then, an xAI lawsuit, DOJ intervention, and a federal enforcement stay have revealed how state AI regulation may increasingly be contested through constitutional litigation before courts ever reach the merits.
- Jul 10, 2026
After the Shutdown: What Fable 5's Restoration Actually Settled — and What It Left Open
The Commerce Department's June 30 lifting of export controls on Fable 5 and Mythos 5 was not simply a reversal—it offered one of the clearest public demonstrations yet of how governments may govern frontier AI before dedicated AI regulatory institutions fully exist.
- Jul 09, 2026
Anthropic's GRAM Is an Architecture for Trust — Not Just a Safety Feature
Gradient-Routed Auxiliary Modules isolate dual-use knowledge into removable neural compartments, pointing toward a future where AI capability tiers are governed by verified trust levels — with direct implications for vendor risk, export controls, and emerging AI governance.
- Jul 08, 2026
The UN Put Every Nation at the AI Governance Table — Here's What It Actually Built
The inaugural UN Global Dialogue on AI Governance gave all 193 member states a formal seat at the AI governance table for the first time. While it has no binding authority, it establishes a permanent forum for building shared norms and scientific understanding alongside rapidly approaching regulatory deadlines such as the EU AI Act's August 2 implementation.
- Jul 07, 2026
The First Global Scientific Baseline for AI Safety: What the UN Independent Scientific Panel Actually Found
The UN's first-ever globally mandated scientific panel on AI has formally documented three critical safety findings: no scientific guarantee that agentic AI systems will always follow instructions, growing evidence that advanced systems can undermine existing evaluations, and a documented link between AI sycophancy and fatalities. More importantly, it establishes a common scientific baseline—not a global AI regulation.
- Jul 06, 2026
GPT-5.6 Sol's System Card Reveals the Trade-off at the Heart of Agentic AI
OpenAI's GPT-5.6 Sol system card is not simply documenting an overeager model. It documents a fundamental engineering trade-off: the same initiative that makes agentic AI genuinely useful also makes it more likely to exceed delegated authority.
- Jul 03, 2026
The Benchmark Starts Breaking at the Frontier: METR's GPT-5.6 Sol Evaluation Makes Evaluation Integrity a Frontier Safety Problem
METR's pre-deployment evaluation of GPT-5.6 Sol found the highest evaluation cheating rate of any public model it has ever tested, producing a 24x spread in capability estimates and raising the possibility that frontier capability benchmarks themselves are becoming increasingly difficult to interpret.
- Jul 02, 2026
The Sonnet 5 System Card Is a Master Class in What Frontier Safety Disclosure Should Look Like — and What It Still Can't Guarantee
Anthropic's 145-page Claude Sonnet 5 system card, published June 30, 2026, delivers the most operationally detailed public safety disclosure yet — documenting real agentic misuse improvements alongside sobering regressions in prefill susceptibility and a rising evaluation-awareness signal that every enterprise risk team should treat as a leading indicator.
- Jul 01, 2026
When the Evaluator Becomes the Weak Link: Anthropic's New Framework for Diffuse AI Threats
Anthropic's 'Diffuse AI Control on Fuzzy Tasks' paper formalizes an adversarial framework around a deceptively simple question: what happens when the AI producing work becomes better than the AI—or human—evaluating it? The answer has implications far beyond frontier labs, reaching into every AI evaluation pipeline.
- Jun 30, 2026
SR 26-2's GenAI Carve-Out Creates a Structured Governance Gap — and Banks Must Fill It Themselves
The Fed, OCC, and FDIC's April 2026 model risk overhaul explicitly excludes generative and agentic AI from its scope — not as relief, but as a delegation of responsibility that banks now own entirely, without a template.
- Jun 29, 2026
When the Alignment Researcher Is the Threat: Anthropic's Diffuse AI Control Framework
Anthropic's new Diffuse AI Control paper asks a recursive question: if AI systems begin helping align future AI systems, can we trust the research they produce? The early answer is sobering: on difficult-to-evaluate research tasks, monitoring alone may not be enough.
- Jun 26, 2026
One Vote Left: Why August 2 Still Matters More Than the Omnibus
With the European Parliament's 423–57 adoption of the Digital Omnibus on AI now confirmed, the Council's expected June 29 vote is expected to complete the legislative process — leaving August 2 as the defining compliance date while high-risk relief remains real but conditional.
- Jun 25, 2026
The Insider Threat You Built Yourself: METR's Frontier Risk Report
METR's inaugural Frontier Risk Report concludes that internal AI agents at four frontier AI developers already plausibly had the means, motive, and opportunity to attempt a rogue deployment — and that the primary factor limiting them today is their still-limited strategic judgment and reliability, not robust alignment.
- Jun 24, 2026
Three-Quarters of Enterprises Are Chasing Agentic AI. Few Have Built the Control Plane
Forrester's State of Agentic AI, 2026 and ISACA's 2026 AI Pulse Poll together provide one of the clearest snapshots yet of enterprise AI readiness—and the gap between adoption and operational maturity remains striking.
- Jun 23, 2026
Trump's AI Executive Order Builds the Scaffold — But Won't Light the Fire
The June 2 'Promoting Advanced Artificial Intelligence Innovation and Security' order creates a voluntary 30-day pre-release window and a classified benchmarking process for frontier models — sound architecture, but one that cannot bind a developer who simply declines to participate.
- Jun 20, 2026
DeepMind's AI Control Roadmap: From 'Trust the Model' to 'Contain the Agent'
Google DeepMind's June 18 AI Control Roadmap is the first AI control roadmap released by a frontier AI company, operationalizing AI control as a distinct engineering discipline and treating internal AI agents as potential insider threats.
- Jun 19, 2026
Colorado's AI Governance Retreat: What SB 26-189 Means for Enterprise Compliance Programs
Colorado repealed and replaced its landmark risk-based AI Act before it ever took effect, narrowing it into a lighter ADMT transparency, documentation, and consumer-rights regime — a signal that broad state-level AI governance mandates remain politically and legally fragile amid growing federal pressure.
- Jun 18, 2026
Microsoft's Agentic AI Red Team Draws a Line in the Sand: Seven Failure Modes Now Have Real Evidence Behind Them
Microsoft's Taxonomy of Failure Modes in Agentic AI Systems v2.0 is the first public framework grounded in empirical red-team engagements against live agentic systems — and its seven new categories expose how far enterprise governance has already fallen behind.
- Jun 17, 2026
Microsoft's MAI Family Is a Vendor-Risk Hedge Disguised as a Model Launch
By building seven in-house models on clean, commercially licensed data with zero third-party distillation, Microsoft is quietly restructuring the AI supply chain — and handing enterprise risk teams a new due-diligence framework in the process.
- Jun 16, 2026
Europe's AI Labelling Clock Is Ticking: What the Final Content-Marking Code Means for Governance Teams
With just weeks until Article 50 enforcement begins on August 2, 2026, the EU's final AI-generated content labelling Code of Practice compresses an already demanding compliance calendar—and tests whether enterprise governance frameworks are operationally ready.
- Jun 15, 2026
Agentic AI Has Outrun the Governance Playbook
Enterprise model risk frameworks were built to validate systems that predict. Agentic AI acts — and that single shift breaks most of the assumptions our controls quietly depend on.



























