Monday, September 28, 2026
AI 인프라 · 뉴스 & 분석
홈 › 정책 › 리포트
정책 · 리포트

Anthropic, OpenAI, and Google are reportedly in discussions to establish an independent AI testing and evaluation body.

Industry coordination on AI testing standards could reduce regulatory uncertainty and create common benchmarks; signals preparation for stricter government oversight.
업계 전문지Slicast · 2026년 9월 27일 12:26 UTC · 미국 · 출처: findarticles.com
중요도 70

Anthropic, OpenAI and Google have reportedly been discussing a shared industry body to develop standards, testing and audits for advanced artificial-intelligence systems before they are released. According to Yahoo News reporting from September 14, citing accounts from The Information and CNN, these talks could give the biggest frontier-model developers a more formal role in deciding what evaluations a model must pass before deployment.

However, there is no announced organization, agreed rulebook or public commitment from the three companies. This distinction matters: a discussion group can explore common language around testing, while an institution would need decisions on membership, financing, technical authority and whether anyone can enforce its findings. The reported proposal arrives as arguments over the speed and safety of AI development have become unusually public, including from Jacob Coxon, a safety researcher formerly employed by both Anthropic and OpenAI.

Yahoo reported that representatives of the companies had met as a working group since at least July and had met again in the preceding week. The talks began before Anthropic chief executive Dario Amodei published his public call to pace frontier AI development on September 6, and before Coxon left Anthropic the following week. The sequence is useful because it separates two related but different developments: according to the report, the companies were already exploring a common standards arrangement before both Amodei's essay and Coxon's resignation, suggesting the safety debate may have added urgency and visibility rather than directly prompting the proposed body.

ABC News's interview with Coxon, published September 13, provides the clearest on-record account of the safety concerns driving scrutiny. Coxon said he resigned over AI-safety concerns, welcomed Amodei's appeal for developers to slow down and argued that neither former employer was acting responsibly. While those are Coxon's assessments rather than findings from an external audit, they illustrate why voluntary testing arrangements are under close examination. Amodei's essay called for a slower pace of development at the frontier but did not establish that Anthropic had adopted a binding slowdown or describe a completed joint program with rivals. Similarly, Coxon noted that conversations about severe AI risks had become more frequent during his time inside the companies, but that observation does not reveal what specific safeguards each company uses or whether they work.

The reported concept is more concrete than a general pledge to develop AI responsibly. Yahoo said Google DeepMind founder Demis Hassabis had proposed a public-private partnership that would evaluate advanced models before release, with industry funding and independent technical experts. Such a body could, in principle, establish common test procedures and publish judgments or recommendations on whether a system met a specified threshold.

A proposed pre-release review system would need to answer fundamental operational questions: what capabilities trigger review, who designs the tests, what evidence is shared, how expert independence is protected, and what happens if a developer proceeds despite an unfavorable result. Testing and auditing perform different functions—testing asks whether a model displays a particular capability or failure under defined conditions, while an audit examines whether the testing process, records, controls and conclusions are credible. None of these details has been agreed publicly. While the reported Hassabis plan's use of independent experts points toward technical credibility, independence depends on more than job titles. Funding, appointment rules, access to models and the treatment of confidential results would all shape who ultimately holds power. A pre-release evaluator also differs from a regulator unless a government grants it legal authority or companies contractually bind themselves to its decisions.

Yahoo reported that OpenAI chief executive Sam Altman favored an industry testing and auditing body and believed leading labs might need to create it without government backing. The same report said the three companies declined to comment to CNN, leaving the account of Altman's internal position and the wider discussions attributed to the reporting rather than confirmed through a joint announcement.

A common testing institution could make it easier for companies to compare results and avoid each developer defining safety for itself. However, it could also concentrate rule-making among companies with the money, compute access and staff needed to operate at the frontier. Yahoo reported that Cohere chief executive Aidan Gomez criticized the prospect of the largest firms setting the rules together, arguing it could put smaller rivals at a disadvantage. The objection is not simply about whether AI needs safeguards; it is about who writes them and how they apply. A demanding evaluation regime may be defensible if its tests are transparent, proportionate and governed independently, but if incumbent developers set costly requirements without meaningful outside representation, the same system can become a barrier to entry.

Political support remains unresolved. Yahoo reported resistance among some companies and said House Speaker Mike Johnson had found no consensus on AI guardrails; it also described a stalled effort around a draft executive order. The companies' reported talks were said to be continuing without administration support. For now, the tangible news is a working-level conversation among three influential developers, not a new regulator or a settled standard. Its credibility will depend on the technical tests it proposes and on whether the labs are willing to accept oversight beyond their own control.

원문 보기
Anthropic, OpenAI, and Google are reportedly… · Slicast