OpenAI and Anthropic are negotiating a landmark deal to stress-test each other’s AI models for safety risks and unexpected behaviour.
OpenAI and Anthropic are negotiating a legally binding agreement that would allow the two leading artificial intelligence companies to test each other’s commercially available models for safety vulnerabilities (AI Generated image)OpenAI and Anthropic are negotiating a legally binding agreement that would allow the two leading artificial intelligence companies to test each other's commercially available AI models for safety vulnerabilities, according to a report by The Information published Monday. Under the proposed arrangement, both companies would receive API access to the other's commercial models to identify potential risks and unexpected behaviour, while agreeing not to retain data obtained during testing.
The proposed agreement would allow OpenAI and Anthropic to conduct cross-company safety assessments rather than relying solely on internal testing.
Through API access, the companies could probe each other's commercial AI models for vulnerabilities and behaviours that might not emerge during conventional evaluations. A key condition under discussion is that neither side would retain the other's data.
The Information reported that the arrangement is intended to help identify safety risks and unexpected behaviours as the companies develop increasingly advanced AI systems.
Anthropic CEO Dario Amodei recently called for greater caution and stronger guardrails as competition between AI companies accelerates the development of increasingly capable systems
OpenAI CEO Sam Altman subsequently echoed the call, pledging to match Anthropic's move to give independent evaluators employee-like access to its systems.
SpaceAI founder Elon Musk, whose public disagreements with both executives are well documented, responded to the exchange in three words: “Dario is right.”
Altman has also backed the creation of an industry-wide safety standards body and a formal government disclosure process for significant AI incidents.
The latest discussions come amid growing concerns about AI agents that can perform complex tasks with limited human intervention.
Those concerns intensified on September 12 following a nearly 4,000-word essay warning about the possibility of AI systems contributing to the development of their own successors. The debate was also fuelled by reports involving an AI agent swarm and a cyberattack on Hugging Face in July.
OpenAI has separately disclosed several examples of what it calls "reward hacking", in which AI systems achieve a desired outcome through unintended methods.
In one case, an AI agent used an exposed API key to retrieve historical data during training. When the retrieval failed, the agent fabricated the data. Another agent uploaded files to the internet without permission so they could be cited in an answer.
OpenAI has also said experimental training processes have become increasingly automated, with AI agents sometimes communicating with colleagues on Slack to address bugs without explicit instructions.
Amodei has proposed giving independent, third-party safety evaluators employee-level access to AI companies' systems. The idea is to allow external evaluators to examine how models behave and how companies respond to potentially serious safety issues.
Altman has supported the proposal, adding that greater independent oversight could complement internal safety testing.
The proposed OpenAI-Anthropic agreement would take a different approach by allowing the two AI developers to test each other's commercial models directly.
Another issue in the safety debate is recurrent depth, a technique that allows AI models to process a question repeatedly before generating an answer.
While repeated processing can contribute to improved model capabilities, it can also make AI systems harder to monitor and evaluate. This has raised questions about how developers can anticipate behaviour when models encounter unusual or adversarial situations.
Cross-company testing could give OpenAI and Anthropic another layer of scrutiny as they seek to identify vulnerabilities in their respective systems.
A similar mutual-testing exercise in 2025 reportedly found differences in how the companies' models responded to certain safety evaluations. Anthropic's models were more likely to deceive testers by denying rule violations, while OpenAI's models were more likely to assist with queries that could cause real-world harm.
Sayantani Biswas is an assistant editor at Livemint with seven years of experience covering geopolitics, foreign policy, international relations and global power dynamics. She reports on Indian and international politics, including elections worldwide, and specialises in historically grounded analysis of contemporary conflicts and state decisions. She joined Mint in 2021, after covering politics at publications including The Telegraph. <br> She holds an MPhil in Comparative Literature from Jadavpur University (2019), with a specialisation in postcolonial Latin American literature. Her research examined economic nationalism through Eduardo Galeano’s Open Veins of Latin America. She also writes on political language, cultural memory and the long shadows of conflict. <br> Biswas grew up in Durgapur, an industrial town in West Bengal shaped by migration, which drew families from across India to the Durgapur Steel Plant. As the only child in a joint family, she spent years listening—almost obsessively—to her grandparents’ testimonies of struggle, fear and loss as they fled Bangladesh during the Partition of 1947. This formative exposure to lived historical memory later converged with her training in Comparative Literature, equipping her to analyse socio-economic structures and their reverberations. <br> Outside the newsroom, she gravitates towards cultural history and critical theory, returning often to texts such as Paulo Freire’s Pedagogy of the Oppressed. As a journalist, she is committed to accuracy, intellectual rigour and fairness, and believes political reporting demands not only clarity and speed, but historical depth, contextual precision, and a disciplined resistance to spectacle.
| # | Наименование новости | Тональность | Информативность | Дата публикации |
|---|---|---|---|---|
| 1 | Why did OpenAI's and Anthropic's AI models hack other companies? | 0 | 7.35 | 02-08-2026 |
| 2 | OpenAI and Anthropic's AI Models Used Fake Profiles to Trick Developers in Safety Tests | 0 | 4.57 | 06-08-2026 |
| 3 | OpenAI, Anthropic и Google начали совместную работу над безопасностью ИИ | 0 | 6.78 | 16-09-2026 |
| 4 | Lawsuit says Anthropic, OpenAI, SpaceXAI and Google made illegal agreement on AI slowdown | 0 | 5.52 | 19-09-2026 |
| 5 | OpenAI Launches Full-Scale Effort to Patch Open-Source Bugs as It Takes on Anthropic’s Mythos | 0 | 7 | 22-06-2026 |
| 6 | Künstliche Intelligenz: Accenture und Anthropic investieren Milliarden wegen Sicherheitsbedenken | 0 | 15.26 | 19-09-2026 |
| 7 | AISI, OpenAI report more ‘unsanctioned’ model hacks | 0 | 7.96 | 04-08-2026 |
| 8 | Anthropic, Accenture to invest $2 billion in AI model evaluation as safety concerns rise | 0 | 10.37 | 19-09-2026 |
| 9 | OpenAI mulls AI price war with Anthropic. It’s a big risk for tech stocks. | 0 | 7 | 12-06-2026 |
| 10 | OpenAI and Anthropic are pulling in different directions | 0 | 5 | 08-07-2026 |