The call, reported by Neowin, Axios and The Herald Business, appears in OpenAI’s proposal, “Building standards for the next phase of AI.” The Herald Business reports that OpenAI wants the United States to develop those standards with other countries, particularly as AI begins performing more AI research itself. These remain recommendations from a model developer, not an international agreement or a compliance deadline.
The useful question is what common standards would make visible. An AI system that completes research tasks under supervision, an agent that acts through a user’s logged-in session, and a hypothetical system that autonomously improves its successors present different oversight problems. OpenAI’s proposal would establish shared measurements and reporting practices around those boundaries—work with practical relevance well before anyone demonstrates fully autonomous self-improvement.
OpenAI’s standards proposal starts with measurement and incident reporting
OpenAI’s central argument is that individual companies and countries should not develop incompatible ways to measure the same emerging risks. Neowin reports that the proposed standards could cover capability evaluations, progress toward recursive self-improvement, circumstances requiring immediate human review, and the classification and reporting of safety incidents. The Next Web separately reports the call for shared incident-reporting rules and international cooperation, although its account explicitly draws on Bloomberg reporting.
Frontier AI here means the highly capable systems at the leading edge of development. The proposal focuses particularly on systems that can contribute to AI research and engineering, where improvements to today’s tools might accelerate development of tomorrow’s models. It does not establish a new category of Windows applications or impose a requirement on every business using an AI assistant.
There are several separate jobs for a technical standard to do:
| Proposed area | What it would make easier to establish |
|---|---|
| Capability measurement | Whether different developers are evaluating comparable abilities under comparable conditions. |
| Research-automation measurement | How much of the research process agents perform, rather than simply how much they are used. |
| Human-review criteria | Which developments should trigger intervention or a decision about continuing work. |
| Incident classification and reporting | Whether organizations mean the same thing when reporting a safety or security failure. |
| Secure threat communication | How sensitive findings can reach governments, developers and infrastructure operators. |
The table describes the intended functions of the proposal, not settled test specifications. Neowin’s reporting supplies the detailed list; Axios independently corroborates the emphasis on shared incident classification and on tracking, reporting and responding to alignment problems. Alignment, in this context, concerns whether a system’s behavior remains consistent with intended human goals and constraints.
An OpenAI official told Axios that the intended incident-reporting scope includes both failures in testing before deployment and incidents involving deployed systems. That detail is attributed to Axios and is not separately confirmed by the other reporting available here. It nevertheless explains an important feature of the proposal: a serious finding inside a research environment could warrant attention even if the affected model never becomes a customer-facing product.
For enterprise readers, comparability is the clearest potential benefit. If suppliers eventually report against common definitions, a buyer could have a better basis for distinguishing an isolated test failure, a repeatable safeguard bypass and an incident with real-world consequences. The proposal does not yet provide that basis, but it identifies the information that would be needed.
OpenAI’s research automation claims make human supervision a measurable boundary
OpenAI supplied a more detailed account of the development behind its policy position in “Research acceleration: The view inside OpenAI,” published September 6, 2026. The company said it had reached its internal goal of an “automated research intern”: a system able to perform well-defined research tasks under human direction, including work that would take a skilled researcher several days. It also described progress toward an automated AI researcher by March 2028.
Those statements are OpenAI’s assessment of its own systems and its development goal. They do not establish that autonomous research has been independently demonstrated across the entire process of creating a frontier model, or that the March 2028 goal will be met.
The distinction is visible in the company’s description of its workflow. OpenAI says people still set research priorities, judge ideas and results, and decide whether to scale, pause or deploy systems. Agents increasingly help with the work between those decisions: producing code, supporting experiments, troubleshooting infrastructure and monitoring research runs.
OpenAI organizes the research process into six broad phases: deciding, designing, building, running, analyzing and communicating. In its internal analysis, agent use increased across all six between January and August 2026, but high-level planning remained a minimal fraction of agent output tokens. That makes it possible for automation to become much more prominent without transferring overall control of the research program to an AI system.
The company’s task-completion figures reinforce that boundary. OpenAI reports that, over a six-month period, more than half of successful tasks estimated to take a human four to eight hours involved at least one human intervention. Its analysis excluded sessions with uncertain outcomes. A successful result therefore cannot automatically be read as evidence that an agent worked independently from beginning to end.
That is a useful lesson for software-development teams evaluating agents today. A completion rate answers only part of the operational question. To judge how much work can safely be delegated, a team also needs to understand how much steering was necessary and which decisions remained with a person. This is an inference from OpenAI’s reported workflow, not a new performance finding about other coding products.
Recursive self-improvement, or RSI, goes further: AI contributes to successive generations of increasingly capable AI, potentially accelerating the process. OpenAI explicitly says that more useful research automation does not mean rapid RSI is necessarily an outcome that should be pursued. In its September 6 update, the company says it does not yet know how to reach aligned, fully automated RSI safely and cannot assume safety work will keep pace with capability improvements.
That admission gives the proposed human-review standards a concrete purpose. They would need to help identify when a system has moved beyond a previously understood level of supervised assistance. Calling every improvement “RSI” would obscure that transition; counting only deployed products could miss it.
OpenAI’s own metrics show why common AI measurements need careful definitions
OpenAI’s internal research update contains striking usage figures, but it also explains why those figures should not be treated as straightforward measures of scientific progress.
By mid-August, the company says its median researcher, ranked by agent usage, was consuming more than $600 per day of inference at API prices. The 90th-percentile user in its research organization consumed more than $7,000 per day on the same basis. These are usage valuations at API prices, not a disclosed statement of OpenAI’s internal cash cost or a purchasing recommendation for outside teams.
The company also reports 3.1 agent-workdays of runtime for every human workday across its research organization, using an eight-hour workday as the comparison. That measures activity. It does not demonstrate that agents contributed 3.1 times as much useful research, replaced an equivalent number of researchers, or accelerated model development by that factor.
OpenAI makes the same broader caution about code and experiments. It says experiments per active experimenter reached a tracking high in August 2026 and that this correlated with increased Codex adoption. However, available computing resources also grew significantly. The report therefore does not isolate coding agents as the sole cause.
These qualifications belong at the center of any standards effort. A common dashboard can still mislead if it standardizes the easiest numbers to collect rather than the outcomes people need to understand. Token consumption, runtime and code production can describe adoption; task completion, intervention requirements and the quality of research outcomes address different questions.
OpenAI itself describes its measurements as preliminary. It notes that easily collected metrics can be difficult to interpret, while measures more directly connected to research progress are harder to develop and validate. It also says that, as some tasks become more automated, the remaining tasks can become more important bottlenecks.
For a business comparing AI-agent claims, the practical implication is to avoid collapsing these measures into one headline productivity figure. A vendor can accurately report greater agent usage without demonstrating an equivalent reduction in elapsed project time. Shared standards would be valuable if they preserved these distinctions instead of giving unlike measurements a common label.
There is also a disclosure trade-off. OpenAI says it wants public tracking of progress toward RSI while protecting security and proprietary information. Common standards could define what must be reported consistently, but the evidence available does not establish how that balance would be enforced. An agreement on terminology would be a useful first step, not a substitute for sufficiently detailed evidence.
ChatGPT Agent testing shows why AI safety overlaps with ordinary cybersecurity
A concrete precedent for government-supported evaluation appears in OpenAI’s September 12, 2025 account of its work with the U.S. Center for AI Standards and Innovation, or CAISI, and the U.K. AI Security Institute. This work predates the September 2026 proposal. It illustrates the kind of technical collaboration OpenAI wants to expand, rather than proving that a new international standards program already exists.
According to OpenAI, CAISI researchers found two vulnerabilities in ChatGPT Agent that, under certain circumstances, could have allowed a sophisticated attacker to bypass protections, control computer systems available to the agent during that session, and impersonate the user on other websites where the user was logged in.
OpenAI says researchers initially believed the vulnerabilities could not be exploited because of the product’s security measures. They subsequently combined conventional software vulnerabilities with an AI-agent hijacking technique to construct a working exploit chain. The company reported an approximately 50 percent success rate for that proof of concept and said it fixed the reported attacks within one business day.
Those numbers describe a particular exercise as reported by OpenAI. They are not a general failure rate for ChatGPT Agent, evidence of a widespread breach, or an outstanding advisory telling customers to apply an unspecified fix. The account also does not establish that every agent with browser access shares those vulnerabilities.
Its broader lesson is nevertheless direct: evaluating a model’s responses alone would not have captured the whole problem. The reported exploit depended on interactions among the agent, software vulnerabilities, security protections and the user’s session. For IT teams, that supports asking whether an assessment covers the deployed system and its access boundaries, rather than only the underlying model.
The U.K. collaboration offers a different example of why testing conditions must accompany results. OpenAI says the institute’s biological-misuse testing produced more than a dozen detailed vulnerability reports, leading to product fixes, policy-enforcement changes and classifier training. Researchers received controlled access that ordinary attackers would not have, including non-public safeguard prototypes and selected configurations with some protections disabled.
Such access can help evaluators locate weaknesses more efficiently. It also means a finding cannot automatically be presented as an attack that any user could reproduce against the public service. Useful reporting must preserve both facts: the weakness informed improvements, and the test conditions materially shaped what researchers could discover.
These examples make incident classification more than administrative housekeeping. A framework needs to distinguish a controlled safeguard test from an exploitable production weakness, while allowing both to inform future defenses. It also needs room for findings that cross the boundary between conventional cybersecurity and AI behavior.
Neowin reports that OpenAI’s proposal includes secure communication channels among governments and critical-infrastructure operators for emerging threats and vulnerabilities. The operational rationale is consistent with the earlier testing: sensitive findings need to reach the organizations able to respond without requiring immediate publication of every technical detail.
Microsoft’s earlier frontier-AI work provides context, not a new compliance mandate
Microsoft has a documented connection to this discussion that does not depend on attaching a Windows hook to an AI policy announcement. In its October 2023 account of frontier-risk governance, OpenAI identified Microsoft, Google DeepMind and Anthropic as its fellow founders of the Frontier Model Forum, an industry body intended to advance safety research and responsible development practices.
That historical account also described a joint Deployment Safety Board with Microsoft. OpenAI said the board approved decisions by either company to deploy models above a certain capability threshold and that GPT-4 was its first eligible deployment. Crucially, the board addressed deployment decisions, rather than earlier decisions about whether to train models of a given size or capability.
The date and scope matter. The 2023 description establishes a precedent for shared governance; it does not confirm the board’s present membership, operating rules or authority in September 2026. OpenAI’s account explicitly allowed for the arrangement and its role to evolve.
The comparison still clarifies what is broader about the new proposal. A deployment review addresses whether a particular system should be released. Standards for research automation, human intervention and pre-deployment incidents would also address what happens while future systems are being built.
According to Neowin, OpenAI wants the United States to work with the existing network of AI safety institutes involving Australia, Canada, Germany, France, Kenya, Japan, Korea, Singapore, India and the United Kingdom. Axios independently reports the proposed use of international safety institutes coordinated through the U.S. center. The country list describes OpenAI’s suggested collaboration network, not evidence that each government has accepted the proposal.
There is a small naming discrepancy in the reporting: Axios calls CAISI the Center for AI “Security” and Innovation, while OpenAI’s primary account uses Center for AI Standards and Innovation. The latter is the name used here. Nothing in the available reporting establishes a separate organization or a rename.
Neowin also reports that OpenAI does not intend the proposed standards to create mandatory pre-release approval or model licensing. Governments would decide how to incorporate technical standards into their own laws. Readers should therefore separate three stages: proposing common practices, adopting a technical standard, and creating binding legal obligations.
For Microsoft customers, no supplied announcement establishes a new Azure, Microsoft 365, GitHub or Windows requirement. The prospective value lies in supplier assurance: more consistent evidence about the advanced models and agents an organization may choose to integrate. Any future product or regulatory obligation would need to be assessed on its own terms.
What OpenAI’s standards proposal means for enterprise AI decisions
Enterprise teams should treat this announcement as a reason to sharpen supplier questions, not as a trigger for an emergency configuration change. The evidence supports reviewing how an AI service is evaluated and supervised; it does not support pausing all agent deployments or claiming that a proposed international framework already certifies a product.
A useful review separates capability, operating conditions and response. Capability concerns what the system can accomplish. Operating conditions include human intervention and the resources available to the agent. Response concerns who investigates an unexpected result and who has authority to stop work. OpenAI’s own research and security accounts show why those questions cannot be answered by a single benchmark.
The announcement also supports patience about formal compliance claims. No adopted test suite, certification procedure or universal human-review threshold is established in the available evidence. An organization can ask suppliers for clearer evidence now without presenting those questions as obligations created by OpenAI’s proposal.
- Treat the September 21 announcement as a policy proposal, not an enacted standard, product certification or new customer deadline.
- When assessing an AI research or coding agent, ask what people had to do for reported tasks to succeed, including interventions and decisions retained by humans.
- Keep usage and output metrics separate from demonstrated productivity gains; OpenAI’s runtime, token and experiment figures measure different aspects of its workflow.
- Ask whether security assessments cover the complete deployed agent and its access boundaries, since OpenAI’s CAISI example combined software vulnerabilities with agent hijacking.
- Distinguish controlled testing incidents from production incidents, and preserve the test conditions when interpreting a supplier’s safety claims.
- Do not assume that Microsoft’s historical frontier-AI governance arrangements establish a new requirement for a current Microsoft service.
OpenAI has put forward a concrete agenda: comparable measurements, clearer human-review boundaries and a shared language for incidents as AI takes on more research work. The useful outcome would be evidence that developers, governments and customers can interpret consistently—not another undifferentiated assurance that a system is safe. Until standards are actually adopted, enterprise decisions still need to rest on the capabilities, access, supervision and incident-handling evidence for the specific system being considered.