Risks and controls for multi-agent systems was published on 10 August 2026 by the Department of Industry, Science and Resources, with the AI Safety Institute named as publisher. It runs to 119 pages and was written by Alistair Reid, Simon O’Callaghan, Dustin Venini, Liam Carroll and Tiberio Caetano of Gradient Institute. The department describes it as a world-first systematic technical framework.

Two things about the provenance are worth stating plainly and then leaving alone. It is the same contractor as the 2025 work, which is ordinary rather than suspicious: a department that has funded a team to build an analytical method usually asks the same team to extend it. And the report is copyright Gradient Institute, not the Commonwealth. Neither is a problem. Both are the sort of thing worth knowing when an institute’s first output arrives.

What matters more is that our open question is now closed. The report builds on Gradient Institute’s earlier Risk analysis techniques for governed LLM-based multiagent systems, which the department published on 29 July 2025. The project did continue, the Institute now has work of its own to point at, and the awkwardness we described in July has resolved itself in the most straightforward way available.

What the framework actually does

The subject is narrow and getting less so: what happens when the AI agents one organisation deploys start interacting with agents deployed by its partners, customers, suppliers and strangers. The report’s framing of the problem is the clearest sentence in it: failures can emerge from interactions between agents that no single organisation’s controls can reach.

It sorts deployments into three tiers by the minimum level of governance that can be assumed between interacting agents.

The three governance tiers, as the report defines them
TierWhat it means
Singular governanceOne organisation governs every agent in the system and has unilateral reach over the whole
Federated governanceMultiple organisations deploy into a shared environment under an agreed set of rules. Each loses unilateral reach, and new failures emerge from misaligned incentives
Open environmentsPersistent agents operate through public infrastructure with no central governing authority. Governance, where it exists, emerges polycentrically, and failures appear at population scale

The point of the tiers is not taxonomy. It is that the controls, and who can pull them, change as you move down the list. Each tier brings new failure modes and, in the report’s words, the controls move progressively beyond the deploying organisation’s reach, toward shared frameworks, public infrastructure and collective action. For every risk it examines, it names who is positioned to act.

The part policymakers should read first

Each tier chapter ends with open problems, and the ones in the third tier are of a different kind. The report says so itself: The problems below are therefore collective-action problems: each requires convening, standards work, or funding that no deploying organisation or single operator can supply, and they are the sharpest form of the coordination gaps this report surfaces for policymakers.

The clearest example is identity. If an agent turns up claiming to act for a bank, a supplier or a government service, something has to make that claim checkable. The report finds the technical pieces are already there and the institution is not: the components are mature as standards but no widely trusted issuer exists to make agent credentials meaningful.

Then the analogy that makes it concrete, and it is a good one. The closest precedent is Let’s Encrypt, which solved the equivalent problem for HTTPS by operating a free, automated certificate authority and became near-universal as a result. For agents, the report says, no such issuer exists for agent identity, and no single party has the convening standing or funding model to build one.

That is a government-commissioned report telling its own commissioner that the most important control in its most difficult tier has no owner, no funding model and no obvious candidate. It is not a recommendation, and the report does not make one. It is a gap, stated precisely, in a document the Institute has put its name to.

Two more gaps worth carrying

Vendor fragmentation. At the federated tier the report notes that if major agent vendors adopt different standards, or implement the same ones differently, that becomes an unintentional barrier to interoperability, naming Microsoft, AWS and Google frameworks as the case in point. Its verdict on the missing layer has the same shape as the identity one: the stitching layer that would make cross-vendor identity implementations interoperable has no clear institutional home yet.

Nobody can properly test any of this yet. The report argues a multi-agent system has to be evaluated at system level rather than agent by agent, then sets out why that is hard: the boundaries of the system are not obvious, agents adapt to each other so behaviour is non-stationary, the state grows combinatorially with agent count, outcomes are stochastic so runs must be repeated, and the compute cost can become prohibitive. The consequence it draws is blunt. Existing methodologies and benchmarks focus on simplified simulations, leaving a wide validity gap between what is evaluated and what is deployed.

What this is, and what it is not

It is a framework, not a rulebook. It creates no obligations, and it presents itself as a way for organisations to triage, for policymakers to see where responsibility falls, and for researchers to see which methods are missing. Anyone hoping the Institute’s first act would be to regulate something will not find it here.

What it does supply is the thing Australian AI policy has mostly lacked: a specific, technical account of where the gaps are and who could close each one, published by a body with the government’s name on it. The measure of it will be whether the gaps it names get owners. On the evidence of its own open-problems sections, the report does not expect that to happen by itself.