Whose Safety Are You Buying? Security and Governance Across the Big AI Labs
"AI safety" is three different questions wearing one name — the frontier-risk framework a lab imposes on itself, the access architecture it ships, and the defaults on the account you actually use. For a buyer, the last two govern your risk far more than the first.
Ask which AI provider is the "safest" and you will get an answer that sounds precise and means almost nothing, because the question is three questions stacked on top of each other. There is the frontier-risk governance a lab imposes on its own model development — the catastrophe-prevention machinery that makes headlines. There is the access architecture it ships — whether the model is a closed service the provider controls or an open weight anyone can run. And there is the mundane layer that actually touches your data every day: the defaults on the specific account tier you are paying for.
These three are routinely collapsed into one word, and the collapse is where enterprises make expensive mistakes. A lab can have the most rigorous frontier-safety framework in the industry and still train on your conversations by default on the tier you happen to be using. A model can be governed by an elaborate capability-threshold policy and still, as an open weight, be entirely outside its maker's control the moment it is downloaded. The press release and the risk you are actually carrying are describing different things.
What follows compares the major labs — Anthropic, OpenAI, Google DeepMind, and Meta — on each of those three layers as they stand in late 2026. The useful finding is not a ranking. It is that the layers diverge independently, so the right provider depends entirely on which of the three questions you were actually asking.
Layer one: the frontier frameworks converged in shape, diverged in strictness
Every major lab now publishes a capability-gated safety policy, and they rhyme by design. Anthropic's Responsible Scaling Policy, first published in 2023 and rewritten as version 3.0 effective February 2026, defines AI Safety Levels (ASL) and ties escalating security and deployment safeguards to capability thresholds — it activated its ASL-3 safeguards in May 2025 around CBRN and AI R&D capabilities. Google DeepMind's Frontier Safety Framework, now at v3.1 (April 2026), is built around Critical Capability Levels, with Tracked Capability Levels added to catch less-extreme risks earlier. OpenAI's Preparedness Framework (v2, April 2025) tracks catastrophic-risk categories against capability tiers. Meta replaced its 2025 Frontier AI Framework in April 2026 with a stricter Advanced AI Scaling Framework that gates release on assessed catastrophic risk.
The convergence is real and, in fairness, was led from the front: Anthropic's RSP is widely credited with pushing the others to adopt broadly similar structures. But shared shape is not shared strictness. The frameworks differ in how much they commit to versus describe, how much they publish versus keep internal, and how enforceable their thresholds actually are. External analyses have been pointed about this — one 2025 study argued OpenAI's Preparedness Framework guarantees no specific mitigation and would still permit deploying systems its own language associates with severe harm. The lesson is not that any one framework is a fraud; it is that a capability-threshold policy is a statement of intent whose value depends on disclosure and follow-through, and those vary widely between labs that all look similar from the outside.
For almost every enterprise buyer, though, this entire layer is the wrong thing to be comparing. Frontier-risk frameworks govern whether a lab should train and release its next model at all — a question of civilizational tail risk, not of whether your procurement data is safe in the tool you deployed last quarter. It matters, but it is not your operational risk, and treating a strong RSP as a reason to trust a consumer chat tier with confidential data is a category error.
Layer two: the deepest fork is closed versus open weights
The governance difference that actually changes what is possible is architectural, not policy-level. Anthropic, OpenAI, and Google ship closed models behind an API. Meta built its reputation on the opposite bet — open-weight Llama models anyone could download and run — before quietly reversing course in 2026, releasing its last open weights (Llama 4 Scout and Maverick) in April 2025 and pivoting to a closed flagship. The open frontier did not collapse so much as relocate, largely to Chinese labs whose weights now anchor the high-volume open-weight tier.
A closed model is a governable object in the ordinary sense: the provider can monitor use, rate-limit, patch a jailbreak overnight, revoke a key, and enforce its usage policy at the point of inference. That control is also its criticism — you are trusting the provider's judgment and cannot inspect the system. An open weight inverts every term. Once it is downloaded it cannot be recalled, patched centrally, or monitored; any safety training can be fine-tuned back out by whoever holds the file. In exchange you get what closed models cannot offer: the weights run on your own infrastructure, so your data never leaves your boundary, and there is no vendor to trust with it at all.
This is the fork that most changes an enterprise's real posture, and it does not have a universally correct side. The pattern that has settled among sophisticated adopters is portfolio deployment — closed frontier models for high-complexity, lower-volume work where peak capability justifies the cost and the vendor dependency, and open weights for high-volume commodity tasks where running them in-house is cheaper and keeps sensitive data on-premises. The governance question "who is responsible when the model misbehaves" has two entirely different answers depending on which side you are on: the provider, or you.
| Lab | Frontier framework | Model access | Paid API / business default | Consumer default |
|---|---|---|---|---|
| Anthropic | Responsible Scaling Policy v3 · ASL levels | Closed / API | No training by default · ZDR available | Opt-out (default on) · 5-yr retention if allowed |
| OpenAI | Preparedness Framework v2 | Closed / API | No training by default · ZDR available | Free ChatGPT trains by default |
| Google DeepMind | Frontier Safety Framework v3.1 · CCLs | Closed / API (Vertex) | No training by default (Vertex) | Free Gemini tier trains by default |
| Meta | Advanced AI Scaling Framework (2026) | Open weights → closed in 2026 | Self-hosted: data stays on your infra | n/a (run the weights yourself) |
Layer three: the defaults that actually touch your data
This is the layer that decides your day-to-day exposure, and the dividing line that matters is not the logo on the model — it is consumer tier versus business tier. Across the major closed providers the paid, business-facing APIs converged on the same default years ago: OpenAI, Anthropic, and Google's Vertex all exclude API traffic from training by default, on free or paid, and offer zero-data-retention arrangements for eligible customers on top of the usual roughly 30-day abuse-monitoring window. If you are on a business API or enterprise agreement, the training question is largely settled in your favor regardless of which of the three you chose.
The consumer tiers are where the defaults diverge and quietly bite. Free ChatGPT and the free Gemini tier use conversations for training by default. Anthropic — the lab with arguably the strongest frontier-safety reputation — flipped its consumer terms in August 2025 from opt-in to opt-out: Free, Pro, and Max chats now train Claude unless you turn the setting off, and opting in carries a five-year retention window against the 30-day standard for those who decline. None of this reaches enterprise, Team, or API accounts. But it is a clean illustration of the central point: the same company can run a rigorous catastrophe-prevention program at the frontier and an opt-out-by-default data policy on its consumer product, because those are two different layers making two different decisions.
For an enterprise, the practical consequence is that your confidentiality risk is set almost entirely by tier discipline, not by vendor selection. The failure mode is not choosing the "wrong" lab; it is an employee pasting a contract into a free consumer chat window while the company's carefully-negotiated enterprise agreement with the same vendor sits unused. The strongest safety framework in the world does not cover that conversation, because that conversation was never on the tier the framework governs.
How to choose across the three layers
Match the question to the layer.If you are assessing systemic or reputational exposure to frontier capability, compare the frameworks — and compare them on disclosure and enforceability, not on whether they exist, since they all now do. If you are assessing operational data risk, ignore the frameworks and read the tier's data terms. Conflating the two is how a strong safety brand gets trusted with data the brand's actual product terms do not protect.
Decide the access architecture deliberately, per workload. Closed API buys you a provider who can patch, monitor, and enforce — and a dependency and a data boundary you cannot cross. Open weights buy you on-premises control and data that never leaves — and full ownership of the safety you fine-tuned away. Portfolio deployment is the mature answer: closed for peak-capability, lower-volume work; open, self-hosted for high-volume or high-sensitivity workloads. Choose per task, not once for the company.
Govern the tier, because that is where the leak is. The highest-leverage control most organizations are missing is not a vendor choice; it is ensuring sensitive work only ever reaches business tiers with training excluded and, where warranted, zero-data-retention enabled — enforced technically, not by policy memo. Assume the consumer default is train-by-default with long retention unless you have confirmed otherwise, because in 2026 that assumption is correct more often than not.
Remember the frameworks all sit under the same external law.Whichever lab and architecture you pick, the deployment still lands under NIST's AI RMF, ISO/IEC 42001, and — for anything touching the EU — the AI Act's obligations, which attach to your use of the system, not to the provider's internal safety policy. The vendor's framework is not a substitute for your own governance; it is one input to it.