CNN, citing people familiar with the conversations, reported that the three companies have held discussions about an AI industry standards body; The Information first reported the talks. The plan under discussion would be substantially different from a conventional voluntary trade group if it gained government backing: Hassabis proposed an organization that could receive frontier models before launch, test them for serious cyber, biological and deceptive capabilities, and eventually make clearance a condition of US deployment.
The supplied report says OpenAI confirmed the effort. CNN’s reporting is less definitive: it says representatives for Google, OpenAI and Anthropic either declined to comment or did not respond. That discrepancy matters. There is evidence of discussions around a proposal, but there is not yet a public joint announcement, charter, budget, board, testing protocol, or agreement that any lab will submit a model for review.
The proposal is for pre-release model access
Hassabis set out his proposal in July for a US-led Frontier AI Standards Body modeled on the Financial Industry Regulatory Authority, the industry-funded organization that oversees US brokerage firms under Securities and Exchange Commission supervision. Axios reported at the time that his proposed body would initially ask frontier labs to provide models voluntarily as much as 30 days before public release.
The evaluations would focus on capabilities with direct relevance to enterprise security: offensive cyber operations, biological misuse and deception. In theory, reviewers would test whether a model can identify vulnerabilities, execute multi-step intrusion workflows, evade safeguards or mislead evaluators about what it can do. Those are not the same as the safety filters users see in a chatbot interface; they are assessments of underlying capability, tool access and the way an AI system behaves in a realistic testing environment.
For developers, this could eventually mean release planning has a new external dependency. A major model launch now involves internal red-teaming, security review, cloud capacity work, documentation and API rollout. If the proposed body ever moves beyond a voluntary arrangement, a pre-deployment evaluation window could become another gating stage for models deemed sufficiently capable. At present, however, neither the definition of frontier nor the threshold that would trigger review has been settled publicly.
The proposal also leaves a difficult technical question unanswered: testing a proprietary AI model without disclosing the evaluation suite in advance. Google DeepMind has recently described double-blind frontier-model evaluations designed to limit benchmark contamination, using protected environments so testing material cannot simply become future training data. That work shows the industry has begun building methods for this kind of review. It does not establish that a common regulator can securely operate them across competing vendors.
Microsoft is already in the overlapping body
The most important omitted context is that OpenAI, Anthropic and Google did not start from zero. In July 2023, those companies and Microsoft launched the Frontier Model Forum, an industry body intended to develop safety practices, support research and share information with governments and other stakeholders. The Forum later added Amazon and Meta.
Microsoft is therefore already part of the existing organization whose remit overlaps considerably with the proposed standards body, despite not being named in the reported discussions. The Forum’s published work includes a voluntary mechanism for sharing safety-relevant information among member companies. Its 2026 incident-response brief says the pilot covers Anthropic, Amazon, Google DeepMind, Meta, Microsoft and OpenAI.
That existing structure has produced information-sharing arrangements and research activity, but it is not a FINRA-like gatekeeper. The Forum does not publicly claim authority to require a model to be submitted before deployment, decide whether a product can launch, publish a binding safety rating, or impose penalties on members. The proposed body’s significance lies precisely in the powers it might acquire beyond the Forum’s voluntary coordination role.
The distinction should temper expectations. A new logo, council or set of best-practice papers would add little to the current landscape. A genuinely independent evaluator with routine prerelease access to frontier models, transparent methodology and a recognized route to government enforcement would be a different institution. The reporting so far does not show that the companies have agreed on any of those features.
Industry funding creates the central credibility test
Hassabis’s framework calls for substantial industry funding, which is understandable on the mechanics: testing advanced models requires scarce security talent, specialized infrastructure and substantial compute. Yet funding by the companies subject to review is also the proposal’s largest governance problem.
The Council on Foreign Relations, in an analysis by former US national-security and AI-policy officials, argued that a FINRA-inspired system could address a real coordination problem but warned that the structure, access rules and handling of sensitive findings would determine whether it earned public trust. The group identified independence and national security as immediate design challenges. A body holding unreleased frontier models, test data and reports on possible cyber weaknesses would itself become a high-value espionage target.
For enterprise defenders, that is more than an abstract policy concern. An evaluator that discovers a model can automate parts of vulnerability research, credential theft, phishing adaptation or exploit development may be handling security-sensitive information before a vendor has deployed mitigations. The model developer, the testing organization and government partners would need explicit rules for containment, reporting and remediation. None have been announced.
The proposal also needs protections against regulatory capture—the risk that the largest model developers write rules that smaller competitors, open-source projects and downstream software companies cannot realistically meet. Cohere chief executive Aiden Gomez has publicly criticized the dominant US labs over precisely that concern, arguing that the dispute is about who writes the guardrails and whose interests they protect. His criticism does not establish that capture will occur, but it identifies the issue that a credible standards body must address in its membership, voting rules, appeal process and public reporting.
An industry body can be useful without becoming a regulator. Shared reporting formats, trusted vulnerability disclosure channels, incident-taxonomy work and reproducible model-evaluation methods would all help security teams compare vendor claims. But a group that sees unreleased proprietary systems while disclosing little about its testing, funding or findings will struggle to be viewed as independent oversight.
What enterprise teams should—and should not—do now
There is no announced change to Azure AI services, Microsoft Copilot, GitHub Copilot, OpenAI APIs, Google Cloud AI services, Anthropic Claude, Windows security baselines or compliance obligations. Administrators should not delay planned deployments or alter procurement controls because of preliminary discussions between AI labs.
They should, however, treat this story as a reminder that external model assurance is still unsettled. Vendor model cards, system cards and safety reports can be useful inputs, but they are vendor-produced evidence and should not replace local testing of identity controls, data boundaries, logging, tool permissions and incident response.
For teams putting AI agents in contact with source repositories, Microsoft 365 data, business applications or Windows endpoints, the controls remain familiar:
- Limit each agent to the narrowest permissions and data scope required for its assigned task.
- Keep human approval in place for consequential actions, especially code changes, external communications, financial workflows and administrative operations.
- Record agent tool calls, authentication events and data access in systems that security teams can investigate.
- Test prompt-injection, data-exfiltration and excessive-permission failures before expanding an agent beyond a controlled pilot.
- Require vendors to explain which model version is in use, what tools it can access, and how quickly they disclose material capability or security changes.
The reported talks could eventually produce common evidence standards that make those procurement questions easier to answer. Today, they have produced only a policy direction: three leading AI labs are considering a mechanism to test the systems they are racing to release. Until the parties publish governance, funding, membership, model-access and enforcement details, the proposed body remains a discussion—not an operational safeguard.