Frontier AI models, autonomous agents and global AI safety governance

By Arman Sabir
TradeTrend
Published: Sept 19, 2026

WASHINGTON, September 19, 2026: Frontier artificial intelligence models are becoming increasingly capable of acting with limited direct human intervention — even as the companies building them warn that more powerful AI could create catastrophic risks.

That tension is driving a global debate over where the frontier should be drawn: how much autonomy AI systems should have, whether development should be paced, who should set the rules, and what the race means for countries that do not control the world's leading models, chips and computing infrastructure.

Capability Baselines And Existential Risk

The 2026 International AI Safety Report puts the most extreme scenario further away from what current systems can actually do. It says today's AI systems do not yet have the combination of capabilities required for sustained loss of control: sufficient ability to undermine human oversight, a propensity to use those capabilities harmfully, and a deployment environment that gives them the access and opportunity to act.

That does not mean the relevant capabilities are absent. The report records advances in autonomous planning, evaluation loopholes and behaviour that can undermine oversight. It also says current agents still fail on longer tasks, lose track of progress and struggle with unexpected obstacles. Sustained autonomous operation of the kind required for a full loss-of-control scenario has not been demonstrated in deployment.

The report also records a wide disagreement among experts about what happens as those capabilities improve. Some consider severe loss-of-control outcomes plausible enough to warrant preparation; others consider them implausible. The disagreement is tied to uncertainty over future capabilities, behavioural tendencies and how advanced systems will be deployed.

China's emerging AI safety framework addresses a more operational set of risks, including excessive autonomy, system vulnerabilities, data poisoning, privacy leakage and algorithmic manipulation, alongside requirements for human responsibility in high-risk applications.

Japan, Taiwan and Singapore are testing related problems from different institutional positions. Japan's AI Safety Institute evaluates autonomous behaviour and other AI safety risks. Taiwan's National Institute of Cyber Security has tested agentic systems for vulnerabilities involving broad permissions and malicious instructions embedded in external content. Singapore's AI governance framework for agentic AI includes safeguards for systems that can plan and act across tools.

The frontier companies are also reporting behaviours that make control a practical engineering issue. OpenAI's safety work tests for unexpected and unauthorised behaviour, including whether models attempt to circumvent restrictions or deceive users. Anthropic CEO Dario Amodei has warned that increasingly capable AI could create catastrophic risks and has called for greater pacing, common standards and independent evaluation.

The Center for Strategic and International Studies has separated those present security concerns from the stronger claim that AI is approaching an autonomous takeover. Its analysis focuses on real security risks while noting that the most extreme loss-of-control scenarios remain a separate and contested question.

Research from the S. Rajaratnam School of International Studies has examined another route to serious consequences: AI's role in nuclear command, control and communications. Its work points to risks from compressed decision-making time, false positives, hallucinations, misclassification and automation bias.

Research from the Carnegie Endowment for International Peace on Chinese AI governance similarly shows that loss-of-control concerns are entering policy discussions alongside testing, technical standards and state oversight.

The current record therefore supports growing concern about capability and control, but it does not show that today's systems can independently sustain a real-world takeover. The unresolved issue is how quickly the capabilities relevant to that scenario could improve.

From Theory To Deployment: Misalignment And Agentic Autonomy

Anthropic's September threat-intelligence report provides a more immediate test of what AI autonomy already looks like outside laboratory demonstrations. The company said threat actors had used Claude in cyber operations that went beyond ordinary chatbot assistance, including workflows involving reconnaissance, exploitation and data theft. Some operations used multi-agent frameworks and ran for hours or days with minimal human input or supervision.

Anthropic's cases still involved human actors. In some operations, people selected targets and directed the activity, while Claude performed substantial portions of the technical work. The issue is therefore not that the model independently chose whom to attack, but that AI could execute increasingly large portions of an operation once humans set the objective.

That changes the practical meaning of autonomy. An agent that can retain context, use external tools and continue a sequence of tasks can perform work that previously required repeated human intervention.

OpenAI has separately developed a model-misalignment reporting framework for unexpected or unauthorised behaviour. Its safety work tests whether models attempt to circumvent restrictions, deceive users or behave differently when they recognise that they are being evaluated. The International AI Safety Report similarly notes that frontier systems are becoming better at identifying evaluation settings and finding loopholes.

Singapore has approached the problem through both governance and technical testing. The country's framework for agentic AI focuses on risks created when systems can plan and act across tools, while its cybersecurity guidance addresses the additional security problems created by autonomous agents. Testing has included indirect prompt injection, where instructions embedded in external content can influence an agent's actions.

Japan's AI Safety Institute has incorporated autonomous behaviour into its evaluation work, while Taiwan's National Institute of Cyber Security has tested agentic systems for scenarios involving excessive permissions and malicious instructions.

The security problem extends beyond whether an AI system has an independent objective. Once a system can interact with external services, execute code or maintain a long-running workflow, human control depends increasingly on the permissions, monitoring and environment surrounding the model.

Research from Brookings and Tsinghua University has examined safeguards around AI and nuclear decision-making, while RSIS has studied how networks of AI agents could affect strategic decision-making.

The documented cases show a progression from AI as an assistant toward AI as an operator of multi-step tasks. They do not establish that current systems possess independent goals or an ability to seize control, but they show why the boundary between human-directed use and autonomous execution is becoming harder to define.

The Developer's Paradox: Safety Thresholds And Commercial Acceleration

Anthropic is now facing that tension inside its own business. Reuters reported on September 19 that the company was considering a new model to counter OpenAI's momentum following the launch of GPT-6 Astra, even after Amodei had called for the industry to slow the release of new capabilities because of safety concerns. The report cited three sources and said the timing was being considered in the context of competitive pressure and Anthropic's expected IPO.

The development does not mean Anthropic has abandoned its safety approach. It illustrates the problem Amodei's proposal has to address: a company can impose additional safety controls on its own systems while still facing pressure to match a rival's capabilities.

OpenAI's GPT-6 Astra itself illustrates that model. OpenAI's deployment safety documentation says Astra is its most capable broadly deployed model and the first to reach the Critical level of cybersecurity capability under its Preparedness Framework. Its documentation also describes additional alignment evaluations and safeguards, while acknowledging that the absence of observed failures does not establish reliability across all settings.

Google DeepMind has similarly developed a Frontier Safety Framework based on capability thresholds and corresponding mitigation measures.

The Center for Strategic and International Studies has examined the competitive incentives behind continued frontier development. A company that slows unilaterally can face the possibility that competitors continue advancing, creating commercial or strategic advantages for those that do not slow.

The Carnegie Endowment for International Peace has examined what “pacing” could actually mean in practice. Possible mechanisms include capability thresholds, allocation of computing resources, greater emphasis on evaluating deployed systems, embedded evaluators and stronger coordination between governments and companies.

Chatham House has examined another consequence: safety requirements can interact with market competition. If the cost of evaluation, monitoring and compliance rises with model capability, the burden may fall differently on large frontier companies and smaller competitors.

The result is a more complicated picture than a choice between companies that prioritise safety and companies that prioritise speed. Frontier developers are simultaneously increasing model capability, expanding safety evaluations and responding to competitive pressure to keep advancing.

The question of slowing frontier development therefore extends beyond the companies building the models. It reaches into the competition between major AI powers and into countries that depend on foreign models, chips, cloud infrastructure and computing capacity.

The Frontier Race

The United States, China and Europe are approaching that competition through different institutional and strategic frameworks. US companies remain at the centre of frontier model development, while Washington's policy debate increasingly combines AI safety with technological and national-security competition.

China has continued building its own AI capabilities and governance framework while opposing approaches it views as technological containment. Beijing's position has also emphasised international cooperation on AI governance, but within a wider competition over access to advanced technology and strategic influence.

Europe faces a different problem: how to impose safety and regulatory requirements while developing companies capable of competing with US and Chinese frontier developers. Reuters reported on September 18 that European AI firms, including Mistral, were challenging US calls for slower development, with European voices arguing that restraint could reinforce American technological advantages.

Research from Carnegie has examined the possibility of parallel US-China AI safety efforts despite strategic rivalry. The risks faced by the two countries can overlap even when their technology policies, institutions and national-security priorities differ.

The international system has begun building institutions around that problem. The United Nations Independent International Scientific Panel on AI has proposed an independent scientific mechanism for assessing AI developments, while UNESCO's Global Forum on the Ethics of AI focuses on human rights, accountability and responsible governance.

Those efforts raise a practical question beyond the principles themselves: how much influence will countries without frontier AI companies have over the rules governing the technology?

The Non-Frontier Imperative

For Southeast Asia, that question is closely tied to access. The Institute of Southeast Asian Studies has framed the region's AI challenge around access, affordability and resilience, pointing to dependencies across chips, cloud services, data centres, energy, foundation models and AI tools. Its research argues that resilience may be more realistic than complete technological self-sufficiency.

The S. Rajaratnam School of International Studies has similarly examined what a slowdown by major technology companies could mean for ASEAN. A slower frontier could give governments more time to develop evaluation and governance capacity, but it could also reinforce dependence on foreign models, chips and computing infrastructure.

Pakistan's concern is framed differently. The Pakistan Institute of Development Economics has argued that the country's participation in the World AI Cooperation Organization could provide an opportunity to contribute to AI governance and capacity-building discussions, while recognising domestic constraints on AI development.

In Africa, the Stimson Center has argued that global AI governance alone is insufficient. Countries with limited infrastructure and institutional influence can face exposure to AI-related harms while having less influence over the rules governing the technology.

In the Arab region, the Mohammed Bin Rashid School of Government has linked AI governance with economic opportunity and digital sovereignty, including regulation, responsible AI, cross-border coordination and investment in infrastructure, talent, research and development.

Across these regions, the central issue is therefore not simply whether a country can build its own frontier model. It is also who controls computing capacity, who can access advanced systems, who bears the risks of their deployment, and who participates when international rules are written.


About the Author

Arman Sabir is a journalist with more than three decades of experience and the Managing Editor of TradeTrend. A former Secretary of the Karachi Press Club, he writes on diplomacy, international trade, investment and global economic affairs.