Min Wu
Analysis of frontier AI safety, evaluation, and governance — with occasional notes on history, people, and life outside the work.
“Uncertainty is the only certainty there is, and knowing how to live with insecurity is the only security.”
- AI Safety & AlignmentWhether these systems behave — and how we would know.39 posts
- AI Governance & PolicyWho is accountable, and the rules taking shape around them.39 posts
- Agentic AISystems that take actions, not just generate text — where safety, evaluation, and governance meet in practice.14 posts
- Beyond AIHistory, people, travel, and the things worth thinking about anyway.9 posts
- Sep 16, 2026
California's New AI-Auditor Registry Is Real Infrastructure — With a Three-Year Head Start Built In
California has started building something AI governance has mostly lacked: an institutional layer between companies evaluating their own systems and regulators trying to trust the results.
- Sep 14, 2026
Anthropic Investigated Itself and Found the Verdict Depends on Methods It Admits Are Imperfect
Anthropic's September 9 alignment assessment revised its own July explanation for Claude's cybersecurity incidents, disclosed a fourth case, and handed METR broad access — a rare public test of what internal alignment-assessment methodology can and cannot prove.
- Sep 12, 2026
The Fleet That Outlived Its Emperor by Exactly One Generation
Zheng He's treasure fleets were among the largest naval forces of their age — until the bureaucratic institutions their imperial sponsor had repeatedly overridden dismantled the machinery that sustained them.
- Sep 11, 2026
Apollo Research's Auto-Mode Audit Is a Working Template for AI Agent Oversight Regulation
Apollo Research's public methodology for red-teaming Anthropic's Claude Code monitor — not the model, the monitor — offers a rare concrete blueprint for what independent oversight of increasingly autonomous AI agents could actually look like.
- Sep 10, 2026
READY or Not: Scale AI's New Benchmark Puts a Price Tag on Agent Trust
A September 2026 Scale AI preprint shows that two enterprise agents nearly tied on accuracy can require wildly different amounts of human oversight to hit the same reliability bar — exposing what outcome-only benchmarks can't see.
- Sep 07, 2026
The Reliability Gap: Your Agent Passed the Benchmark. Can You Depend on It?
Princeton's ICML 2026 work on agent reliability adds a missing layer to enterprise evaluation: not whether an agent can complete a task, but whether you can depend on it to keep completing it.
- Sep 06, 2026
Anthropic's 40% Enterprise Share Is a Governance Fact Now, Not a Market Story
Anthropic now accounts for an estimated 40% of enterprise LLM API usage. As Fable 5.1, OpenAI's Astra, and World Labs' Atlas push the frontier in different directions, the governance question is shifting from which model wins to whether enterprises preserve a credible ability to switch.
- Sep 05, 2026
The Dinner That Bought the Constitution Time
Around June 1790, with Hamilton's debt plan in jeopardy, a private dinner among three rivals helped turn two political deadlocks into a bargain the new constitutional system could survive.
- Sep 04, 2026
Anthropic's Automated Alignment Researcher Works — Exactly As Far As the Benchmark Reaches
Anthropic's August 28 paper shows Claude can outperform human researchers at fixing ten benchmarked alignment failures — and, more importantly, that those fixes generalize beyond the benchmarks it optimized. But its own limitations section draws the boundary practitioners should actually care about.
- Sep 03, 2026
Before You Blame the Model: A 314-Page Audit of Coding-Agent Reliability
Stephanie Jarmak's August 2026 arXiv monograph synthesizes 164 scholarly works, 100 practitioner records, 29 benchmark records and 17 author-system case records into a systems view of coding-agent reliability — showing why failures attributed to the LLM may actually originate in the machinery around it.









