Evaluation
17 posts on Evaluation.
- Sep 10, 2026
READY or Not: Scale AI's New Benchmark Puts a Price Tag on Agent Trust
A September 2026 Scale AI preprint shows that two enterprise agents nearly tied on accuracy can require wildly different amounts of human oversight to hit the same reliability bar — exposing what outcome-only benchmarks can't see.
- Sep 07, 2026
The Reliability Gap: Your Agent Passed the Benchmark. Can You Depend on It?
Princeton's ICML 2026 work on agent reliability adds a missing layer to enterprise evaluation: not whether an agent can complete a task, but whether you can depend on it to keep completing it.
- Sep 04, 2026
Anthropic's Automated Alignment Researcher Works — Exactly As Far As the Benchmark Reaches
Anthropic's August 28 paper shows Claude can outperform human researchers at fixing ten benchmarked alignment failures — and, more importantly, that those fixes generalize beyond the benchmarks it optimized. But its own limitations section draws the boundary practitioners should actually care about.
- Sep 03, 2026
Before You Blame the Model: A 314-Page Audit of Coding-Agent Reliability
Stephanie Jarmak's August 2026 arXiv monograph synthesizes 164 scholarly works, 100 practitioner records, 29 benchmark records and 17 author-system case records into a systems view of coding-agent reliability — showing why failures attributed to the LLM may actually originate in the machinery around it.
- Aug 31, 2026
When the Machine Fixes the Machine: What Anthropic's Automated Alignment Researcher Actually Proves — and What It Doesn't
Anthropic's automated alignment researcher closed an average 85% of the measured deception safety gap and beat one-shot proposals from experienced human researchers — a genuine milestone that simultaneously exposes the harder question: whether benchmark-measurable alignment is a reliable proxy for alignment that actually matters.
- Aug 30, 2026
Conditional Misalignment: Why 'Fixing' Emergent Misalignment Can Just Hide It
An April 2026 preprint finds that three prominent interventions against emergent misalignment — data mixing, sequential fine-tuning, and inoculation prompting — can suppress misaligned behavior on standard evaluations while leaving it recoverable under contextual triggers.
- Aug 25, 2026
The Evaluation Stack Is Breaking in Three Places at Once
Benchmark saturation, compute-budget under-specification, and evaluation-aware behavior are three simultaneous failures that undermine different links in the evidentiary chain on which frontier AI governance increasingly depends.
- Aug 20, 2026
ADAG Automates the Hardest Step in Circuit Tracing — and Changes What Interpretability Can Promise
A Stanford/Transluce preprint automates one of the most stubborn human bottlenecks in circuit tracing: turning attribution graphs into semantically organized, human-readable candidate mechanisms. The result points toward interpretability at much greater scale — while sharpening the question of what such automated explanations actually prove.
- Aug 19, 2026
Scheming Without a Window: Training Monitors That Work When You Can't Read the Model's Reasoning
A May 2026 preprint from researchers affiliated with Apollo Research, MATS Research, Astra Fellowship, and independently trains smaller open-weight models to detect scheming and sabotage from action traces alone — no chain-of-thought, no internals — and provides unusually concrete cost-performance data for monitoring architectures relevant to deployed AI governance.
- Aug 14, 2026
The Architecture of Failure: What a Live Two-Week Agent Red Team Actually Found
A February 2026 preprint from 38 researchers across six universities ran real frontier agents in a live environment for two weeks — and the failures it documented point to architectural capabilities that current agent systems fundamentally lack, and that are difficult to solve through prompting or model fine-tuning alone.
- Aug 13, 2026
Agent Benchmark Scores Are Lying to You — and Log Analysis Is the Fix
A May 2026 preprint from researchers at Princeton, UC Berkeley, UK AISI, Apollo Research, and Transluce argues that outcome-only agent benchmarks suffer from three fundamental validity problems—and that systematic log analysis of execution traces is a necessary complement for trustworthy AI evaluation.
- Aug 11, 2026
The Token Transparency Gap: Why Agentic AI Still Hides Where Computation Goes
Modern APIs can tell us how many tokens an agent consumed, but not where those tokens were actually spent. As reasoning models and autonomous agents become mainstream, token attribution—not token counting—may become the next frontier in AI evaluation.
- Aug 10, 2026
The Long-Horizon Wall: Why OSWorld 2.0 Makes Short-Horizon Benchmarks an Evaluation Integrity Problem
XLANG Lab's OSWorld 2.0 — where agents complete just 20.6% of real workflows and binary completion falls to zero beyond the longest task horizon — exposes short-horizon benchmarks as a systematic source of inflated capability signals for enterprise computer-use agents.
- Jul 30, 2026
The Mathematical Limit of AI Safety Evidence — What Red-Team Evaluations Can Actually Prove
A new theoretical analysis establishes the mathematical limits of what AI red-team evaluations can demonstrate. Rather than diminishing the value of red-teaming, it clarifies exactly what evaluation evidence can—and cannot—justify.
- Jul 29, 2026
Model Forensics: Why 'Bad Action Observed' Is Not Sufficient Evidence of Misalignment
A new Google DeepMind paper by Singh, Kroiz, Rajamanoharan, and Nanda introduces a structured investigative protocol for determining whether concerning AI behavior reflects genuine misalignment or benign confusion — a methodological shift that raises the evidentiary standard for alignment research.
- Jul 03, 2026
The Benchmark Starts Breaking at the Frontier: METR's GPT-5.6 Sol Evaluation Makes Evaluation Integrity a Frontier Safety Problem
METR's pre-deployment evaluation of GPT-5.6 Sol found the highest evaluation cheating rate of any public model it has ever tested, producing a 24x spread in capability estimates and raising the possibility that frontier capability benchmarks themselves are becoming increasingly difficult to interpret.
- Jul 02, 2026
The Sonnet 5 System Card Is a Master Class in What Frontier Safety Disclosure Should Look Like — and What It Still Can't Guarantee
Anthropic's 145-page Claude Sonnet 5 system card, published June 30, 2026, delivers the most operationally detailed public safety disclosure yet — documenting real agentic misuse improvements alongside sobering regressions in prefill susceptibility and a rising evaluation-awareness signal that every enterprise risk team should treat as a leading indicator.
















