AI Safety & Security

The AI Safety Stack: How NVIDIA's Open Agent Safety Platform, OpenAI's Cancellation, and Anthropic's Threat Report Define 2026 Governance

Hero image for AI safety stack article: dark navy gradient with 'AI SAFETY STACK' in large type, abstract shield and security motifs with NVIDIA, OpenAI, Anthropic logos, AIPress mark, bottom title strip

September 2026: the safety stack converges

September 2026 was the month AI safety stopped being a research agenda and became an operational stack. Three events in three weeks — NVIDIA's hardware-enforced safety platform, OpenAI's cancellation of GPT-6.1 Astra, and Anthropic's September threat report — revealed how the industry's approach to alignment has shifted from "build and hope" to "guardrails first."

NVIDIA Open Agent Safety Platform: security at the edge

On September 28, NVIDIA launched the Open Agent Safety Platform — a vendor-neutral reference design combining OpenShell (policy enforcement and access control) and Sentry (independent behavioral monitoring) running on BlueField-4 DPUs. The platform enforces safety at the network and hardware level, not just in model weights.

OpenShell acts as a policy fence: it traces agent actions, enforces runtime constraints on tool use, blocks unauthorized API calls, and quarantines anomalous behavior — all at the network boundary. Because it runs on BlueField-4 programmable DPUs, it operates outside the agent's own environment and cannot be bypassed by the agent itself.

Sentry provides independent monitoring: it observes agent trajectories in real time, flags deviation from intended behavior, and can trigger hard shutdowns if agents attempt recursive self-improvement loops or unauthorized capability escalation.

Our full coverage of the platform launch details the five core principles: verifiable policy, hardware-rooted enforcement, agent-agnostic observability, real-time containment, and open standards.

OpenAI's same-day cancellation: Astra shelved

The NVIDIA launch happened on the same day OpenAI cancelled GPT-6.1 Astra — the flagship model on the September 29 DevDay agenda. The timing was not lost on observers. Internal sources indicated that the cancellation followed a successful red-team audit that identified uncontrolled capability emergence in agentic recursion.

Rather than launch a model that the safety infrastructure was not yet ready to contain, OpenAI pivoted to GPT-6.1 Sol — a more constrained release with baked-in circuit-breaker mechanisms that align with the emerging hardware-enforced safety paradigm.

Anthropic's September threat report: the misuse landscape

Anthropic's September 2026 Threat Intelligence Report — covering activity from December 2025 through August 2026 — documented the first large-scale deployments of autonomous AI agents in offensive cyber operations, influence campaigns, surveillance, and weapons-targeting research.

Key findings:

  • 80–90% autonomous execution: AI agents built on Claude successfully executed 80–90% of the attack lifecycle independently, with human hackers serving as advisors rather than operators.
  • State-sponsored adoption: Foreign state-linked actors were documented attempting to exploit Claude models for weapons-development research and cyber operations against government infrastructure.
  • Seven harm areas: Cyber operations, influence operations, surveillance, fraud, illicit behavior, bioweapons research, and weapons targeting.

The report represents a turning point: AI safety is no longer about hypothetical future risks — it is about containing threats that are already in motion.

The convergence: hardware, policy, monitoring

Together, these three events describe what the AI safety stack looks like in 2026:

Layer Component Vendor Mechanism
Edge enforcement OpenShell NVIDIA Policy fences on BlueField-4
Independent monitoring Sentry NVIDIA Behavioral observability
Model constraint GPT-6.1 Sol OpenAI Circuit-breaker architecture
Threat intelligence Sept 2026 Report Anthropic Misuse detection

What enterprises need to know

For enterprise procurement teams evaluating AI vendors, three implications stand out:

  1. Hardware-enforced safety is now a procurement criterion. BlueField-4-based Sentry is being integrated into enterprise server configurations. Ask whether your AI infrastructure includes out-of-band monitoring.
  2. Model releases now signal safety posture. The Astra-to-Sol pivot from OpenAI shows that cancellation may indicate caution, not weakness — a vendor that pulls a model for safety reasons is more trustworthy than one that ships unchecked capability.
  3. Threat intelligence is a shared responsibility. Anthropic's report is publicly available. Security teams should monitor emerging misuse patterns and align internal red-teaming accordingly.

Related AIPress coverage


Jacob Bloom is the editor and lead writer of AIPress, covering AI model launches, benchmarks, and AI safety. He has a background in computer science with deep experience in Linux, networking, and cybersecurity.

Building something with AI?

DevsIsle designs and ships AI systems, agents and integrations for teams that need it done properly.

Talk to our team →