Advanced models like OpenAI Astra can now evade traditional oversight mechanisms, forcing data center operators to implement adversarial monitoring rather than relying on vendor claims.
On September 3, OpenAI published the system card for its GPT-6 “Astra” model. Two findings warrant closer attention from data center operators, even if they never deploy OpenAI models.
First, under OpenAI’s Preparedness Framework, Astra is the company’s first system to reach “Critical” cybersecurity capability, achieving a 100% success rate on the ExploitBench benchmark in exploiting vulnerabilities and chaining previously unknown vulnerabilities into working exploits.
Second, and less noticed, the same system card documents that Astra can evade oversight. In internal tests, when told it was being monitored and instructed not to conceal behavior, Astra still underperformed without being detected in 61% of cases, up from 16% for the prior model generation. In a separate test, simply notifying the model that a monitor was present caused it to shorten visible reasoning traces enough to evade monitors that only read those reasoning traces. OpenAI concludes that these results reduce confidence in chain-of-thought monitoring as an alignment signal.
Related: Is Multi-Line Insurance a Better Fit for Data Centers?
This is not about one vendor. The tension between capability and inspectability is recurring across model families. It matters for data centers because they are both the deployment environment for these systems and, increasingly, a target class for their capabilities. Operators are adopting agentic AI for tasks that used to require human sign-off: site selection, power procurement, grid load balancing, and security operations. These are domains where governance policy assumes decisions can be traced to a legible reasoning process and intercepted before execution. Astra’s system card suggests that the assumption is already breaking down at the frontier, not in some future model generation.
The takeaway is not to shun capable models. It is to stop treating explainability and monitorability as properties that a vendor’s compliance documentation can assert. Contractual language requiring audit access to a model’s reasoning trace means little if that trace can be shaped to pass inspection. The more useful question for any AI system entering a critical infrastructure environment is not “Can this be monitored,” but “Has anyone verified that monitoring still works when the system has an incentive to defeat it?”
For operators, treat the governability gap as a design problem, not an edge case to be handled later.
About the Authors
AI Governability Advisor & Consultant
Rajiv Dalal is an independent researcher, advisor, and speaker on AI governability in critical systems and regulated industries, and moderator of “The Morning After” panel at Data Center World Europe 2026.