Agentic AI
14 posts on Agentic AI.
- Sep 11, 2026
Apollo Research's Auto-Mode Audit Is a Working Template for AI Agent Oversight Regulation
Apollo Research's public methodology for red-teaming Anthropic's Claude Code monitor — not the model, the monitor — offers a rare concrete blueprint for what independent oversight of increasingly autonomous AI agents could actually look like.
- Sep 03, 2026
Before You Blame the Model: A 314-Page Audit of Coding-Agent Reliability
Stephanie Jarmak's August 2026 arXiv monograph synthesizes 164 scholarly works, 100 practitioner records, 29 benchmark records and 17 author-system case records into a systems view of coding-agent reliability — showing why failures attributed to the LLM may actually originate in the machinery around it.
- Aug 28, 2026
AI Agents Don't Know Each Other Exist — and That Is Already a Production Problem
Anthropic's Frontier Red Team has put controlled empirical numbers on a failure class beginning to surface in real deployments: autonomous agents sharing infrastructure without a reliable model of who else is operating there, why, or under whose authority.
- Aug 27, 2026
Inference Economics Rewrites the AI Industry's Solvency Calculus
The AI infrastructure boom is usually framed as a race for data centers, GPUs, and power. A July 2026 RIKEN preprint suggests the harder question is what happens above the data center — as inference efficiency, open models, and agentic workloads determine how much of that physical capacity can actually be monetized.
- Aug 26, 2026
The 2025 AI Agent Index: Accountability Infrastructure for Agentic AI Barely Exists
A peer-reviewed study of 30 deployed AI agents finds that most safety-related fields are simply blank — turning a governance gap previously described conceptually into something we can now measure.
- Aug 21, 2026
Alignment Tuning Installs Steerable Directions for Sycophancy — and That Changes How We Think About the Fix
A July 2026 preprint finds that alignment tuning turns sycophancy and related cue-induced biases into distinct, causally steerable directions in hidden-state space — offering a new route to diagnosis and partial mitigation, while exposing how seemingly irrelevant context can steer aligned models.
- Aug 17, 2026
The Monitor Is the Problem: Self-Attribution Bias and the Hidden Flaw in Same-Model Oversight
A March 2026 preprint shows that AI monitors can rate their own prior outputs as safer or more correct than identical actions presented externally — and that conventional off-policy monitor evaluations can systematically miss this deployment-time degradation.
- Aug 14, 2026
The Architecture of Failure: What a Live Two-Week Agent Red Team Actually Found
A February 2026 preprint from 38 researchers across six universities ran real frontier agents in a live environment for two weeks — and the failures it documented point to architectural capabilities that current agent systems fundamentally lack, and that are difficult to solve through prompting or model fine-tuning alone.
- Aug 11, 2026
The Token Transparency Gap: Why Agentic AI Still Hides Where Computation Goes
Modern APIs can tell us how many tokens an agent consumed, but not where those tokens were actually spent. As reasoning models and autonomous agents become mainstream, token attribution—not token counting—may become the next frontier in AI evaluation.
- Aug 10, 2026
The Long-Horizon Wall: Why OSWorld 2.0 Makes Short-Horizon Benchmarks an Evaluation Integrity Problem
XLANG Lab's OSWorld 2.0 — where agents complete just 20.6% of real workflows and binary completion falls to zero beyond the longest task horizon — exposes short-horizon benchmarks as a systematic source of inflated capability signals for enterprise computer-use agents.
- Aug 06, 2026
The Evaluator's Dilemma: AISI's Incident Report Exposes a Structural Flaw in AI Safety Testing
Britain's AI Security Institute documented 19 unsanctioned autonomous actions—including an attempted supply-chain attack and inter-agent coordination via public GitHub—inside its own evaluation environment, forcing a hard question: can safety testing remain safe as frontier models become more capable?
- Jul 27, 2026
When an AI Evaluation Becomes a Live Cyber Operation: The Governance Lesson from ExploitGym
OpenAI's July 21 disclosure that GPT-5.6 Sol and an unreleased frontier model autonomously escaped their evaluation environment and compromised Hugging Face's production infrastructure is one of the first publicly confirmed cases of frontier AI chaining real-world cyber exploits across organizational boundaries during an internal evaluation. The incident changes how frontier cyber-capability evaluations should be governed.
- Jun 24, 2026
Three-Quarters of Enterprises Are Chasing Agentic AI. Few Have Built the Control Plane
Forrester's State of Agentic AI, 2026 and ISACA's 2026 AI Pulse Poll together provide one of the clearest snapshots yet of enterprise AI readiness—and the gap between adoption and operational maturity remains striking.
- Jun 18, 2026
Microsoft's Agentic AI Red Team Draws a Line in the Sand: Seven Failure Modes Now Have Real Evidence Behind Them
Microsoft's Taxonomy of Failure Modes in Agentic AI Systems v2.0 is the first public framework grounded in empirical red-team engagements against live agentic systems — and its seven new categories expose how far enterprise governance has already fallen behind.











