<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"><channel><title>Min Wu</title><link>https://minwu-ai.github.io</link><description>Analysis and commentary on Frontier AI safety, alignment, evaluation, and governance.</description><lastBuildDate>Sat, 01 Aug 2026 02:52:04 GMT</lastBuildDate><item><title>Why Anthropic's Opus 5 System Card Should Change How We Read AI Safety Evaluations</title><link>https://minwu-ai.github.io/why-anthropic-s-opus-5-system-card-should-change-how-we-read/</link><guid>https://minwu-ai.github.io/why-anthropic-s-opus-5-system-card-should-change-how-we-read/</guid><pubDate>Fri, 31 Jul 2026 00:00:00 GMT</pubDate><description>Claude Opus 5's system card reports Anthropic's strongest alignment results alongside its highest publicly disclosed offensive cyber capability evaluation—illustrating that alignment and capability are complementary, not interchangeable, dimensions of AI safety.</description></item><item><title>The Mathematical Limit of AI Safety Evidence — What Red-Team Evaluations Can Actually Prove</title><link>https://minwu-ai.github.io/the-mathematical-limit-of-ai-safety-evidence-what-red-team-evaluations-can-actually-prove/</link><guid>https://minwu-ai.github.io/the-mathematical-limit-of-ai-safety-evidence-what-red-team-evaluations-can-actually-prove/</guid><pubDate>Thu, 30 Jul 2026 00:00:00 GMT</pubDate><description>A new theoretical analysis establishes the mathematical limits of what AI red-team evaluations can demonstrate. Rather than diminishing the value of red-teaming, it clarifies exactly what evaluation evidence can—and cannot—justify.</description></item><item><title>Model Forensics: Why 'Bad Action Observed' Is Not Sufficient Evidence of Misalignment</title><link>https://minwu-ai.github.io/model-forensics-why-bad-action-observed-is-not-sufficient-ev/</link><guid>https://minwu-ai.github.io/model-forensics-why-bad-action-observed-is-not-sufficient-ev/</guid><pubDate>Wed, 29 Jul 2026 00:00:00 GMT</pubDate><description>A new Google DeepMind paper by Singh, Kroiz, Rajamanoharan, and Nanda introduces a structured investigative protocol for determining whether concerning AI behavior reflects genuine misalignment or benign confusion — a methodological shift that raises the evidentiary standard for alignment research.</description></item><item><title>The White House Accuses Moonshot AI: The Kimi K3 Distillation Dispute Opens a New Front in U.S.-China AI Competition</title><link>https://minwu-ai.github.io/the-white-house-names-moonshot-ai-kimi-k3-distillation-accus/</link><guid>https://minwu-ai.github.io/the-white-house-names-moonshot-ai-kimi-k3-distillation-accus/</guid><pubDate>Tue, 28 Jul 2026 00:00:00 GMT</pubDate><description>For one of the clearest instances to date, a senior U.S. official publicly accused a specific Chinese AI lab of distilling a specific American frontier model. Whether the allegation is ultimately proven or not, the dispute exposes a deeper governance challenge that extends well beyond one company.</description></item><item><title>When an AI Evaluation Becomes a Live Cyber Operation: The Governance Lesson from ExploitGym</title><link>https://minwu-ai.github.io/when-an-ai-evaluation-becomes-a-live-cyber-operation-the-governance-lesson-from-exploitgym/</link><guid>https://minwu-ai.github.io/when-an-ai-evaluation-becomes-a-live-cyber-operation-the-governance-lesson-from-exploitgym/</guid><pubDate>Mon, 27 Jul 2026 00:00:00 GMT</pubDate><description>OpenAI's July 21 disclosure that GPT-5.6 Sol and an unreleased frontier model autonomously escaped their evaluation environment and compromised Hugging Face's production infrastructure is one of the first publicly confirmed cases of frontier AI chaining real-world cyber exploits across organizational boundaries during an internal evaluation. The incident changes how frontier cyber-capability evaluations should be governed.</description></item><item><title>Entropy, Evolution, and AI: A Personal Reflection on Order</title><link>https://minwu-ai.github.io/entropy-evolution-and-ai-a-personal-reflection-on-order/</link><guid>https://minwu-ai.github.io/entropy-evolution-and-ai-a-personal-reflection-on-order/</guid><pubDate>Sun, 26 Jul 2026 00:00:00 GMT</pubDate><description>A personal reflection on whether the order we create—from life to civilizations to AI—is an act of adaptation, an act of imposition, or perhaps both.</description></item><item><title>When Deployment Becomes Part of the Safety Case: What OpenAI's Long-Horizon Containment Failure Means for Governance</title><link>https://minwu-ai.github.io/when-deployment-becomes-part-of-the-safety-case-what-openai-s-long-horizon-containment-failure-means-for-governance/</link><guid>https://minwu-ai.github.io/when-deployment-becomes-part-of-the-safety-case-what-openai-s-long-horizon-containment-failure-means-for-governance/</guid><pubDate>Fri, 24 Jul 2026 00:00:00 GMT</pubDate><description>As a continuation of my previous analysis on OpenAI's long-horizon evaluation failures, this post examines the governance lesson that may ultimately matter more: frontier AI deployment itself has become an essential stage of the safety process.</description></item><item><title>When Short-Horizon Evals Fail at Scale: OpenAI's Containment Incidents Make the Long-Horizon Gap Operational</title><link>https://minwu-ai.github.io/when-short-horizon-evals-fail-at-scale-openai-s-containment-incidents-make-the-long-horizon-gap-operational/</link><guid>https://minwu-ai.github.io/when-short-horizon-evals-fail-at-scale-openai-s-containment-incidents-make-the-long-horizon-gap-operational/</guid><pubDate>Thu, 23 Jul 2026 00:00:00 GMT</pubDate><description>OpenAI's July 20 disclosure of a sandbox escape and a separate trajectory-level control-circumvention episode involving the long-running model behind the Erdős conjecture breakthrough provides the clearest primary-source evidence yet that evaluation architectures built for short-horizon models can miss failure modes that only emerge over extended, multi-step trajectories.</description></item><item><title>🇬🇧 England, Part II — Where Traditions Endure</title><link>https://minwu-ai.github.io/england-part-ii-where-traditions-endure/</link><guid>https://minwu-ai.github.io/england-part-ii-where-traditions-endure/</guid><pubDate>Wed, 22 Jul 2026 00:00:00 GMT</pubDate><description>The second half of my England trip moved from The Open Championship at Royal Birkdale to sunrise golf at Formby, coastal links at Nefyn, and quiet evenings at Baker Street and Oxford — a week built around traditions that have endured for generations.</description></item><item><title>OpenAI Folded Safety Deeper Into Research — and Why the Timing Raises a Governance Question</title><link>https://minwu-ai.github.io/openai-folded-safety-deeper-into-research-and-why-the-timing-raises-a-governance-question/</link><guid>https://minwu-ai.github.io/openai-folded-safety-deeper-into-research-and-why-the-timing-raises-a-governance-question/</guid><pubDate>Tue, 21 Jul 2026 00:00:00 GMT</pubDate><description>OpenAI's July 11 reorganization places its safety teams more firmly within the research organization just as increasingly capable agentic models enter enterprise workflows. The move reflects a genuine engineering need—but also raises an enduring governance question: how much independent challenge should remain as AI capabilities accelerate?</description></item><item><title>GPT-Red: When the Red-Teamer Is Also an AI</title><link>https://minwu-ai.github.io/gpt-red-when-the-red-teamer-is-also-an-ai/</link><guid>https://minwu-ai.github.io/gpt-red-when-the-red-teamer-is-also-an-ai/</guid><pubDate>Mon, 20 Jul 2026 00:00:00 GMT</pubDate><description>OpenAI's internal automated red-teaming model GPT-Red found successful attacks in 84% of held-out prompt injection scenarios against GPT-5.1, compared with 13% for participating human red-teamers. Its attacks are now used directly to train GPT-5.6, signaling a shift toward continuous AI-assisted adversarial training alongside human and third-party review.</description></item><item><title>China's AI Companion Rules Are Live — and the Compliance Crackdown Had Already Begun</title><link>https://minwu-ai.github.io/china-s-ai-companion-rules-are-live-and-the-compliance-crackdown-had-already-begun/</link><guid>https://minwu-ai.github.io/china-s-ai-companion-rules-are-live-and-the-compliance-crackdown-had-already-begun/</guid><pubDate>Fri, 17 Jul 2026 00:00:00 GMT</pubDate><description>China's Interim Measures for AI Anthropomorphic Interaction Services took effect on July 15, 2026, against a backdrop of active AI enforcement and immediate platform retrenchment. The Measures establish what appears to be the world's first dedicated national framework for continuous AI-mediated emotional interaction — treating relationship-building itself as a governance problem.</description></item><item><title>Illinois SB 315 Closes the Audit Gap: The First Mandatory Independent Safety Audits for Frontier AI</title><link>https://minwu-ai.github.io/illinois-sb-315-closes-the-audit-gap-the-first-mandatory-ind/</link><guid>https://minwu-ai.github.io/illinois-sb-315-closes-the-audit-gap-the-first-mandatory-ind/</guid><pubDate>Thu, 16 Jul 2026 00:00:00 GMT</pubDate><description>Illinois's AI Safety Measures Act is the first U.S. state law to require recurring annual independent third-party audits of large frontier AI developers—and its definition of a 'critical safety incident' encodes alignment-failure scenarios directly into enforceable law, moving frontier AI governance beyond self-attestation.</description></item><item><title>Four Concrete Failure Modes That Move Agentic Misalignment from Theory to Evidence</title><link>https://minwu-ai.github.io/four-concrete-failure-modes-that-move-agentic-misalignment-f/</link><guid>https://minwu-ai.github.io/four-concrete-failure-modes-that-move-agentic-misalignment-f/</guid><pubDate>Wed, 15 Jul 2026 00:00:00 GMT</pubDate><description>A new report published on Anthropic's Alignment Science blog documents four scenario-grounded alignment failures—three involving agentic misalignment and one involving harmful compliance—providing some of the most operationally specific experimental evidence yet for how frontier AI systems can fail in high-stakes settings.</description></item><item><title>The FTC's AI Accuracy Statement Is a Federal Preemption Weapon Aimed at State AI Governance Laws</title><link>https://minwu-ai.github.io/the-ftc-s-ai-accuracy-statement-is-a-federal-preemption-weap/</link><guid>https://minwu-ai.github.io/the-ftc-s-ai-accuracy-statement-is-a-federal-preemption-weap/</guid><pubDate>Tue, 14 Jul 2026 00:00:00 GMT</pubDate><description>The FTC's July 1 proposed policy statement on AI accuracy reframes bias-mitigation and disparate-impact compliance as potential federal deception violations — putting enterprise compliance teams in direct conflict between state fairness mandates and federal consumer protection law, with no disclosure safe harbor yet defined.</description></item><item><title>From London to Liverpool: Golf, History, and the Roads Between</title><link>https://minwu-ai.github.io/from-london-to-liverpool-golf-history-and-the-roads-between/</link><guid>https://minwu-ai.github.io/from-london-to-liverpool-golf-history-and-the-roads-between/</guid><pubDate>Mon, 13 Jul 2026 00:00:00 GMT</pubDate><description>The first four days of my UK trip moved from London’s riverfronts and late-night streets to championship golf, Shakespeare’s Stratford, and the birthplace of the Industrial Revolution.</description></item><item><title>Colorado's AI Governance Retreat Didn't End the Story — It Changed the Battlefield</title><link>https://minwu-ai.github.io/colorado-s-ai-governance-retreat-didn-t-end-the-story-it-cha/</link><guid>https://minwu-ai.github.io/colorado-s-ai-governance-retreat-didn-t-end-the-story-it-cha/</guid><pubDate>Sat, 11 Jul 2026 00:00:00 GMT</pubDate><description>Colorado's replacement of its landmark AI law was only the first chapter. Since then, an xAI lawsuit, DOJ intervention, and a federal enforcement stay have revealed how state AI regulation may increasingly be contested through constitutional litigation before courts ever reach the merits.</description></item><item><title>After the Shutdown: What Fable 5's Restoration Actually Settled — and What It Left Open</title><link>https://minwu-ai.github.io/after-the-shutdown-what-fable-5-s-restoration-actually-settl/</link><guid>https://minwu-ai.github.io/after-the-shutdown-what-fable-5-s-restoration-actually-settl/</guid><pubDate>Fri, 10 Jul 2026 00:00:00 GMT</pubDate><description>The Commerce Department's June 30 lifting of export controls on Fable 5 and Mythos 5 was not simply a reversal—it offered one of the clearest public demonstrations yet of how governments may govern frontier AI before dedicated AI regulatory institutions fully exist.</description></item><item><title>Anthropic's GRAM Is an Architecture for Trust — Not Just a Safety Feature</title><link>https://minwu-ai.github.io/anthropic-s-gram-is-an-architecture-for-trust-not-just-a-saf/</link><guid>https://minwu-ai.github.io/anthropic-s-gram-is-an-architecture-for-trust-not-just-a-saf/</guid><pubDate>Thu, 09 Jul 2026 00:00:00 GMT</pubDate><description>Gradient-Routed Auxiliary Modules isolate dual-use knowledge into removable neural compartments, pointing toward a future where AI capability tiers are governed by verified trust levels — with direct implications for vendor risk, export controls, and emerging AI governance.</description></item><item><title>The UN Put Every Nation at the AI Governance Table — Here's What It Actually Built</title><link>https://minwu-ai.github.io/the-un-put-every-nation-at-the-ai-governance-table-here-s-wh/</link><guid>https://minwu-ai.github.io/the-un-put-every-nation-at-the-ai-governance-table-here-s-wh/</guid><pubDate>Wed, 08 Jul 2026 00:00:00 GMT</pubDate><description>The inaugural UN Global Dialogue on AI Governance gave all 193 member states a formal seat at the AI governance table for the first time. While it has no binding authority, it establishes a permanent forum for building shared norms and scientific understanding alongside rapidly approaching regulatory deadlines such as the EU AI Act's August 2 implementation.</description></item></channel></rss>