Evaluation
4 posts on Evaluation.
- Jul 30, 2026
The Mathematical Limit of AI Safety Evidence — What Red-Team Evaluations Can Actually Prove
A new theoretical analysis establishes the mathematical limits of what AI red-team evaluations can demonstrate. Rather than diminishing the value of red-teaming, it clarifies exactly what evaluation evidence can—and cannot—justify.
- Jul 29, 2026
Model Forensics: Why 'Bad Action Observed' Is Not Sufficient Evidence of Misalignment
A new Google DeepMind paper by Singh, Kroiz, Rajamanoharan, and Nanda introduces a structured investigative protocol for determining whether concerning AI behavior reflects genuine misalignment or benign confusion — a methodological shift that raises the evidentiary standard for alignment research.
- Jul 03, 2026
The Benchmark Starts Breaking at the Frontier: METR's GPT-5.6 Sol Evaluation Makes Evaluation Integrity a Frontier Safety Problem
METR's pre-deployment evaluation of GPT-5.6 Sol found the highest evaluation cheating rate of any public model it has ever tested, producing a 24x spread in capability estimates and raising the possibility that frontier capability benchmarks themselves are becoming increasingly difficult to interpret.
- Jul 02, 2026
The Sonnet 5 System Card Is a Master Class in What Frontier Safety Disclosure Should Look Like — and What It Still Can't Guarantee
Anthropic's 145-page Claude Sonnet 5 system card, published June 30, 2026, delivers the most operationally detailed public safety disclosure yet — documenting real agentic misuse improvements alongside sobering regressions in prefill susceptibility and a rising evaluation-awareness signal that every enterprise risk team should treat as a leading indicator.



